Voice assistants have evolved far beyond simple commands and weather updates. Today, they serve as the primary interface for smart speakers, televisions, streaming devices, connected vehicles, and an expanding ecosystem of AI-powered consumer electronics.
As conversational AI becomes integrated across a growing range of products, the patents covering these technologies are also receiving greater attention. That trend is reflected in a Section 337 investigation instituted by the U.S. International Trade Commission (ITC) on June 5, 2026, where Cerence alleges that Amazon’s Alexa-enabled devices infringe several of its conversational AI patents.
Patent infringement litigation overview
On May 5, 2026, Cerence filed a complaint against Amazon.com, Inc. and Amazon.com Services, LLC alleging infringement of several patents covering five U.S. patents relating to conversational AI and voice-enabled technologies: U.S. Patent No. 8,306,815, U.S. Patent No. 8,320,575, U.S. Patent No. 8,355,484, U.S. Patent No. 9,203,972 and U.S. Patent No. 11,929,073. The accused products span much of Amazon’s Alexa ecosystem, including Echo smart speakers, Fire TV devices, Fire tablets, Echo Auto, smart displays, smart televisions, streaming players, and related Alexa cloud services.
Cerence is represented by Troutman Pepper Locke, while Morgan, Lewis & Bockius serves as counsel for both Amazon.com, Inc. and Amazon.com Services, LLC. Cerence seeks a limited exclusion order and cease-and-desist orders, remedies that could restrict the importation and sale of the accused products if the ITC ultimately finds a violation.
Patents in dispute
The asserted patents cover key technologies that enable modern voice assistants to recognize, process, and respond to spoken requests From improving speech recognition and audio processing to reducing response latency and balancing local and cloud computing, the patents highlight the technical building blocks behind conversational AI.
Improving the reliability of speech-based interactions
Reliable speech recognition remains a fundamental challenge for conversational AI, particularly in noisy environments such as vehicles.
In this context, U.S. Patent No. 8,306,815 describes a system that preprocesses incoming audio through two parallel processing paths. One path is optimized for speech recognition, while the other extracts contextual information such as pitch, speaking volume, ambient noise levels, and the user’s location. By combining recognized speech with these contextual cues, the system can better interpret user intent and deliver more accurate, adaptive voice interactions.

The figure illustrates this dual-path architecture. Incoming audio is first processed by a signal pre-processing unit before being routed separately to the speech recognition unit and the speech dialog control unit. The dialog control unit integrates both speech and contextual information to dynamically adjust system behavior, including speech output, recognition parameters, and location-aware functions. Feedback loops between the dialog control, signal pre-processing, and speech recognition units also enable the system to continuously adapt to changing acoustic conditions, improving the robustness and responsiveness of voice-enabled systems.
The ‘815 patent, titled “Speech dialog control based on signal pre-processing”, was filed on December 6, 2007, and published on November 6, 2012. The patent lists Lars König, Gerhard Uwe Schmidt, and Andreas Löw as inventors. The patent is represented by Sunstein Kann Murphy & Timbers (now part of Merchant & Gould).
Reducing computational requirements for audio processing
U.S. Patent Nos. 8,320,575 and 9,203,972, which belong to the same patent family and both titled “Efficient audio signal processing in the sub-band regime”, are directed to efficient audio signal processing in the sub-band regime. The invention reduces the computational resources required for speech enhancement, making it particularly suitable for real-time voice applications running on resource-constrained devices.
Applications such as hands-free communication, in-vehicle voice assistants, and speech recognition require audio to be processed in real time while operating within the memory and processing limitations of embedded devices. Rather than enhancing every portion of an incoming audio signal, the patented approach processes audio in multiple frequency sub-bands, allowing the system to selectively reduce the amount of data requiring enhancement. This approach helps lower computational demands while maintaining overall speech quality.
In the patented audio communication architecture, microphones capture both the user’s speech and unwanted background noise or acoustic echo from nearby loudspeakers before the signals are routed to a signal processing system. The audio is divided into multiple frequency sub-bands, where enhancement operations such as noise reduction and echo cancellation are performed on selected sub-bands while others are temporarily omitted to reduce processing requirements. The omitted frequency components are subsequently reconstructed, and the processed sub-bands are recombined into a full-band output signal. The invention enables more efficient real-time speech enhancement for voice recognition and hands-free communication systems.
The ‘575 patent was filed on September 30, 2008, and published on November 27, 2012, while the ‘972 patent was filed on September 14, 2012, and published on December 1, 2015. Both patents list Gerhard Uwe Schmidt, Hans-Jörg Köpf, and Günther Wirsching as the inventors. The patents are originally assigned to Nuance Communications, Inc., and represented by Sunstein Kann Murphy &
Timbers.
Reducing perceived latency in voice interactions
While efficient audio processing improves the quality and speed of speech recognition, delivering responses naturally presents another challenge. Voice assistants often require time to recognize speech, interpret user intent, retrieve information, and generate a spoken response, resulting in periods of silence that can make the system appear unresponsive. Rather than relying on generic tones or background audio during these delays, the U.S. Patent No. 8,355,484 describes an approach that introduces natural transitional audio fillers that help maintain the flow of conversation and create a more human-like interaction while the system completes its processing.

The figure illustrates this architecture, where a user’s spoken request passes through automatic speech recognition (ASR), natural language understanding (NLU), a dialog manager, and natural language generation (NLG) before the response is synthesized through a text-to-speech system. As these processes take place, a filler generator is activated immediately after the user finishes speaking, selecting natural vocalizations, such as breaths, coughs, or filled pauses like “um” and “hmmm,” as well as conversational phrases such as “let’s see” from a dedicated database. These transitional cues are played until the final response is ready, after which the system seamlessly delivers the synthesized speech, ultimately enhancing the responsiveness and overall user experience of voice-enabled systems without requiring faster underlying processing.
The patent ‘484, titled “Methods and apparatus for masking latency in text-to-speech systems”, was filed on January 8, 2007, and published on January 15, 2013. The patent lists Ellen Marie Eide, and Wael Mohamed Hamza as inventors. The patent is originally assigned to Nuance Communications, Inc. and represented by Wolf, Greenfield & Sacks.
Balancing local and cloud-based speech processing
Modern voice assistants often rely on cloud-based processing to deliver more accurate speech recognition and natural language understanding, but transmitting audio to remote servers can introduce noticeable delays. In contrast, on-device processing provides faster responses but is typically limited by the device’s computational capabilities. The U.S. Patent No. 11,929,073 addresses this tradeoff through a hybrid arbitration system that determines whether a locally generated recognition result is sufficiently reliable or whether a more comprehensive cloud-based interpretation should be used.

In the illustrated system, a user’s spoken request is processed simultaneously by an on-device automatic speech recognition (ASR) and natural language understanding (NLU) module and a cloud-based ASR/NLU service. While the on-device system quickly generates an initial recognition result, the same speech signal is transmitted to the cloud for more computationally intensive processing. An arbitration module evaluates the confidence of the local result using speech recognition and language understanding metrics. If the confidence exceeds a predefined threshold, the command is executed immediately, providing a faster user experience. Otherwise, the system waits for the cloud-generated result before selecting the final response. By dynamically choosing between embedded and cloud processing, the invention improves the responsiveness of voice assistants while preserving the accuracy of cloud-based speech recognition.
The patent, titled “Hybrid arbitration system”, was filed on October 3, 2022, and published on March 12, 2024. The patent lists Min Tang as the inventor. The patent is represented by Occhiuti & Rohlicek.
Driving the innovation
Originally spun off from Nuance Communications in 2019, Cerence is a leading provider of AI-powered conversational technologies for the automotive industry. Its voice assistants and intelligent in-car experiences have been integrated into millions of vehicles worldwide through partnerships with many of the world’s major automotive manufacturers. The company’s expertise spans speech recognition, natural language understanding, and generative AI solutions designed to enable safer and more intuitive interactions between drivers and their vehicles.
While Cerence is best known for its automotive AI platform, its patent portfolio reflects a broader innovation strategy extending beyond in-vehicle voice assistants. Over the years, the company has expanded its research into technologies supporting intelligent human-machine interaction across connected vehicles, smart homes, mobile devices, and other consumer electronics, demonstrating how its expertise has evolved alongside the broader conversational AI ecosystem.
While the asserted patents provide insight into the technologies at the center of the dispute, they also represent only a small portion of Cerence’s broader innovation strategy. Examining the company’s overall patent portfolio provides additional context on how these technologies fit within its long-term research and development efforts.
Cerence: Patenting Activity
Cerence maintained active patenting throughout the early 2020s, with patent application activity remaining particularly strong for priority years 2020 through 2023. This period coincided with the company’s transition into an independent automotive AI software provider following its 2019 spin-off from Nuance Communications, enabling a dedicated focus on AI-powered mobility solutions for connected, autonomous, electric, and shared vehicles.

During these years, Cerence expanded its technology portfolio through next-generation conversational AI platforms, hybrid embedded-cloud architectures, and cloud-connected mobility services. In 2021, the company introduced multimodal AI capabilities integrating voice, gesture, gaze, and touch, expanded its solutions into automotive, two-wheelers, and building mobility, and later unveiled Cerence Co-Pilot, an AI-powered in-car assistant that combines vehicle sensors with edge and cloud intelligence.
Although recent priority years show fewer published patent filings, this is likely influenced by normal publication lags. Cerence has continued to introduce new AI capabilities, including a 2024 collaboration with Microsoft to integrate generative AI and Azure OpenAI Service into its automotive assistant platform. This suggests the company’s innovation efforts continue to evolve toward large language models, hybrid embedded-cloud architectures, and next-generation conversational AI experiences for connected vehicles.
Cerence: Top Law Firms
Cerence’s patent portfolio has been supported by a mix of U.S. and international patent law firms, reflecting the company’s geographically diverse filing strategy. Occhiuti & Rohlicek and Keenway Patentanwälte Neumann account for the largest share of identified law firm affiliations, followed by Brooks Kushman and Haseltine Lake Kempner. Among individual practitioners, Tilman Taruttis is notably associated with 28 filings through Keenway Patentanwälte Neumann, indicating a significant role in Cerence’s European patent prosecution activities.

Additional firms appearing across the portfolio, including Burr & Forman, Westphal, Mussgnug & Partner, Ohlandt, Greeley, Ruggiero & Perle, and MUHANN Patent & Law Firm, further demonstrate the involvement of patent firms across multiple jurisdictions in supporting Cerence’s intellectual property activities. The diversity of prosecution firms reflects Cerence’s international filing strategy, with U.S. and European counsel supporting protection across its principal innovation markets.
Cerence: Top Technology Areas
Cerence’s patent portfolio is primarily concentrated in G10L and G06F, reflecting the company’s focus on speech technologies and AI-driven computing. These classifications encompass innovations in speech recognition, speech synthesis, natural language processing, machine learning, and software architectures that form the foundation of conversational AI platforms.

Beyond these core technologies, the portfolio demonstrates a strong emphasis on automotive applications through classifications such as B60R, B60W, and B60K, covering vehicle systems, vehicle control, and the integration of intelligent technologies within the cabin environment. Additional activity in H04R, G06V, G08G, H04L, and H04S highlights innovations in audio processing, computer vision, intelligent transportation systems, communications, and acoustic technologies, illustrating the multidisciplinary nature of Cerence’s research and development.
Notably, the technologies represented by these classifications closely align with the patents asserted in the ITC investigation, which relate to speech dialog control, audio signal processing, text-to-speech systems, and hybrid speech recognition architectures. This alignment suggests that the asserted patents are rooted in technology areas that have remained central to Cerence’s innovation strategy.
Innovation in motion
Subsequent patent activity associated with the asserted patents demonstrates that these inventions continue to influence later developments across multiple industries. Organizations citing the asserted patents include consumer electronics manufacturers, automotive companies, and research institutions, reflecting the broad applicability of the underlying innovations.

Samsung Electronics accounts for the largest number of forward-citing patents, followed by Dolby International and ST Case1Tech. Amazon Technologies, Honeywell International, Toyota Motor, and Kyoto University also rank among the leading organizations, while Cerence, through both Cerence Operating and Cerence, continues to build upon technologies related to the asserted patents.
The diversity of organizations represented in the forward citation landscape highlights the broad applicability of the technologies disclosed in the asserted patents. Their adoption across consumer electronics, automotive, and research organizations illustrates the continued development of innovations in speech recognition, audio processing, and conversational AI as voice-enabled systems evolve across a growing range of intelligent devices.
Looking forward
The USITC investigation is still in its early stages, and no determination has been made regarding the alleged patent infringement. As the case progresses, the asserted patents and the technologies they cover are likely to receive closer scrutiny during the proceedings.
Regardless of the outcome, the investigation provides a timely look into Cerence’s conversational AI patent portfolio and the technologies that support modern voice-enabled systems. It also highlights how foundational innovations in speech and language processing continue to play an important role across an expanding range of AI-powered devices.
