Nova Patents
CA3140979C

Wake-word detection suppression

Abstract

When enabled, the wake response of a given networked microphone device to a particular wake word causes the given networked microphone device to listen, via a microphone, for a voice command following the particular wake word. Before audio content is played back by a playback device, the playback device detects in the audio content one or more wake words for one or more voice services. The playback device causes one or more networked microphone devices to disable their respective wake response to the detected one or more wake words during the playback of the audio content by the playback device before playing back the audio content via one or more speakers.

CA3140979C, drawing sheet 1
Sheet 1 of 12

Term

11.9 yearsleft in the term

Expires 6 August 2038.

  1. Priority
  2. Filed
  3. Granted
  4. Today
  5. Expires

20 claims: 11 independent, 9 dependent

  1. 1
    What is Claimed is:1. A playback device comprising: a network interface;one or more microphones;one or more processors;data storage having stored therein instructions executable by the one or more processors to cause the playback device to perform functions comprising: receiving audio content for playback by the playback device;providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of the VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;playing back a first portion of the audio content via one or more speakers;detecting, via the local wakeword engine, that a second portion of tire received audio content includes sound data matching one or more particular local wakewords;before the second portion of the received audio content that includes the sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device;and playing back the second portion of the audio content via one or more speakers. Date Reçue/Date Received 2022-07-29
  2. 5
    The playback device of any one of claims 1 to 4, wherein disabling the local wakeword response of Ae local wakeword engine to Ae one or more particular local wakewords during playback of Ae second portion of Ae audio content by Ae playback device comprises:before playing back Ae second portion of Ae audio content, modifying Ae second portion of Ae audio content to incorporate acoustic markers in segments of Ae second portion Aat represent respective local wakewords, wherein Ae acoustic markers causes the local wakeword engine to Date Reçue/Date Received 2022-07-29 disable its local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.
  3. 6
    The playback device of any one of claims 1 to 5, wherein the functions further comprise:detecting, via the VAS wake-word engine, that a third portion of die received audio content includes sound data matching a particular VAS wake word;before the third portion of the received audio content that includes the sound data matching the particular VAS wake word is played back, disabling the VAS wake word response of the VAS wake-word engine to the particular VAS wake word during playback of the third portion of the audio content by the playback device;and playing back the third portion of the audio content via one or more speakers.
  4. 7
    The playback device of any one of claims 1 to 6, wherein the VAS wake-word engine comprises a first VAS wake-word detection algorithm for a first VAS and a second wake-word detection algorithm for a second VAS, and wherein providing the sound data stream representing die received audio content to the VAS wake-word engine comprises:applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the first VAS wake word detection algorithm for the first VAS;and applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, the second VAS wake word detection algorithm for the second VAS.
  5. 8
    A method to be performed by a playback device, the method comprising:receiving audio content for playback by the playback device;providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of tire VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local Date Reçue/Date Received 2022-07-29 wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;playing back a first portion of the audio contait via one or more speakers;detecting, via the local wakeword engine, that a second portion of the received audio content includes sound data matching one or more particular local wakewords;before the second portion of the received audio content that includes die sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device;and playing back the second portion of the audio content via one or more speakers.
  6. 12
    The meAod of any one of claims 8 to 11, wherein disabling Ae local wakeword response of Ae local wakeword engine to Ae one or more particular local wakewords during playback of Ae second portion of Ae audio content by Ae playback device comprises:before playing back Ae second portion of Ae audio content modifying Ae second portion of Ae audio content to incoiporate acoustic markers in segments of Ae second portion Aat represent respective local wakewords, wherein Ae acoustic markers causes Ae local wakeword engine to disable its local wakeword responses to Ae one or more particular local wakewords during playback of Ae second portion of Ae audio content by Ae playback device.
  7. 13
    The meAod of any one of claims 8 to 12, further comprising:detecting, via Ae VAS wake-word engine, Aat a Aird portion of Ae received audio contait includes sound data matching a particular VAS wake word;before Ae Aird portion of Ae received audio content Aat includes Ae sound data matching Ae particular VAS wake word is played back, disabling Ae VAS wake word response of Ae VAS wake-word engine to Ae particular VAS wake word during playback of Ae Aird portion of Ae audio content by Ae playback device;and playing back Ae Aird portion of Ae audio content via one or more speakers.
  8. 14
    The meAod of any one of claims 8 to 13, wherein Ae VAS wake-word engine comprises a first VAS wake-word detection algori Am for a first VAS and a second wake-word detection algori Am for a second VAS, and wherein providing Ae sound data stream representing Ae received audio content to Ae VAS wake-word engine comprises:applying, to Ae sound data stream representing Ae received audio content before Ae audio content is played back by Ae playback device, Ae first VAS wake word detection algori Am for Ae first VAS;and Date Reçue/Date Received 2022-07-29 applying, to the sound data stream representing the received audio content before the audio content is played back by the playback device, die second VAS wake word detection algorithm for the second VAS.
  9. 15
    A tangible, non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors of a playback device, cause the playback device to perform functions comprising:receiving audio content for playback by the playback device;providing a sound data stream representing the received audio content to (i) a voice assistant service (VAS) wake-word engine and (ii) a local wakeword engine, wherein the VAS wake-word engine is operable to (a) generate a VAS wake word response when the VAS wake-word engine detects a VAS wake word in a microphone sound data stream representing sound detected by one or more microphones of the playback device and (b) stream sound data representing the sound detected by the one or more microphones to one or more servers of the VAS when the VAS wake word response is generated, and wherein the local wakeword engine is operable to (a) generate a local wakeword response when the local wakeword engine detects one or more local wakewords in the microphone sound data stream representing sound detected by the one or more microphones and (b) determine an intent of a voice input comprising the one or more local wakewords when the local wakeword response is generated;playing back a first portion of the audio content via one or more speakers;detecting, via the local wakeword engine, that a second portion of the received audio content includes sound data matching one or more particular local wakewords;before the second portion of the received audio content that includes the sound data matching the one or more particular local wakewords is played back, disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of foe second portion of foe audio content by foe playback device;and playing back the second portion of foe audio content via one or more speakers.
  10. 19
    The tangible, non-transitory computer-readable medium of any one of claims 15 to 18, wherein disabling the local wakeword response of the local wakeword engine to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device comprises:before playing back the second portion of the audio content, modifying toe second portion of toe audio content to incorporate acoustic markers in segments of toe second portion that represent respective local wakewords, wherein toe acoustic markers causes toe local wakeword engine to Date Reçue/Date Received 2022-07-29 disable its local wakeword responses to the one or more particular local wakewords during playback of the second portion of the audio content by the playback device.
  11. 20
    The tangible, non-transitoiy computer-readable medium of any one of claims 15 to 19, wherein the functions further comprise:detecting, via the VAS wake-word engine, that a third portion of the received audio content includes sound data matching a particular VAS wake word;before the third portion of the received audio content that includes the sound data matching the particular VAS wake word is played back, disabling the VAS wake word response of the VAS wake-word engine to the particular VAS wake word during playback of the third portion of the audio content by the playback device;and playing back the third portion of the audio content via one or more speakers. Date Reçue/Date Received 2022-07-29