What Amazon Nova Sonic was designed to do
Amazon Nova Sonic was Amazon Web Services’ first-generation speech-to-speech foundation model for Amazon Bedrock. Its main purpose was real-time voice conversation: an application could stream a person’s speech to the model and receive generated speech, transcription events, text responses, and tool-related events in return.
Many voice applications are built by connecting separate services for automatic speech recognition, text generation, and text-to-speech. Nova Sonic combined those stages within one model-oriented interaction. That design was intended to reduce orchestration work and support more natural conversations, including multi-turn exchanges and user interruptions.
The model was aimed at voice assistants, customer-service automation, interactive education, language learning, and other speech-enabled enterprise workflows. It was not primarily a general text model, image model, coding model, or media-generation system.
Lifecycle status and position in Amazon’s lineup
Nova Sonic was accessed through Amazon Bedrock with the model identifier amazon.nova-sonic-v1:0. AWS documentation identifies it as a legacy model and records September 14, 2026 as its official end-of-life date. Based on the supplied research dated September 25, 2026, that date has passed, so new deployments should not treat the model as currently available for production use.
Amazon Nova Sonic belongs to the Amazon Nova family, but it is separate from Amazon’s consumer Alexa+ experience and from other AWS services such as Amazon Q, Amazon Bedrock Knowledge Bases, and SageMaker AI. Amazon Nova 2 Sonic is the newer speech-to-speech model in the family. It should be evaluated as a successor with its own model identifier, limits, event formats, regional availability, language coverage, and pricing rather than assumed to be a drop-in alias.
Inputs, outputs, and voice capabilities
Nova Sonic supported both speech and text input. Its direct outputs included speech and text, making it useful when an application needs spoken responses for the user while also retaining text transcripts for interfaces, records, or downstream processing.
- Real-time speech-to-speech conversation
- Speech recognition and streamed text transcription
- Expressive masculine-sounding and feminine-sounding voices
- Text responses alongside generated audio
- Handling of user interruptions while preserving conversational context
- Function calling for external services and enterprise systems
- Retrieval-augmented generation, including Amazon Bedrock Knowledge Bases integration
At launch, AWS documented American and British English. Later documentation listed Spanish, French, Italian, and German as supported languages, with expressive voices across five language groupings: English with US and UK variants, French, Italian, German, and Spanish. Language and voice support should therefore be checked against the relevant AWS documentation rather than assumed to be identical across all deployments.
Streaming behavior and context limits
The primary interaction pattern was bidirectional audio streaming through Amazon Bedrock’s InvokeModelWithBidirectionalStream API. In practical terms, the application could send audio continuously while receiving generated audio, transcription, text, and other events without waiting for a complete recording or a fully completed response.
AWS documented a 300,000-token context window. The context window is the amount of conversation and other supported input the model can consider at once; it is not the same thing as an audio duration limit. Launch documentation also described a default connection limit of eight minutes. Longer conversations could be continued by opening a new connection and supplying relevant prior conversation history as context.
The supplied research does not verify a separate maximum-output-token limit, so no numeric output ceiling should be assumed beyond the documented connection behavior. Developers planning long sessions would need to account for connection renewal, conversation-history transfer, latency, and the cost of processing repeated context.
Function calling and grounded enterprise workflows
Nova Sonic supported function calling, also known as tool use. This allowed a voice application to ask an external system to perform an operation or retrieve information. For example, a customer-service assistant could use a connected function to look up an account, explain a subscription, or begin a supported service workflow.
The model could also use retrieval-augmented generation, including integration with Amazon Bedrock Knowledge Bases. Retrieval-augmented generation supplies relevant information from an external knowledge source at request time. This can help a voice assistant answer questions about enterprise documents or current business information without relying only on what was encoded in the model during training.
Tool use does not mean the model independently has unrestricted access to business systems. The application still needs to define available functions, validate arguments, enforce permissions, handle errors, and decide whether actions require confirmation. Voice interfaces especially benefit from confirmation steps for consequential operations because transcription or interpretation can be wrong.
Pricing and access considerations
Nova Sonic used token-based Amazon Bedrock billing that distinguished speech and text usage. Text input and output charges could apply to activities such as transcription, tool calls, knowledge grounding, and processing conversation history. The supplied research does not verify a static, model-specific legacy price amount, so a precise per-unit price should not be stated.
The model was initially launched in US East (N. Virginia). AWS documentation later listed in-region availability in US East (N. Virginia), Europe (Stockholm), and Asia Pacific (Tokyo). Regional availability, account eligibility, and model lifecycle status mattered because Bedrock models are not necessarily available in every AWS Region.
For a new project, the most important pricing question was not simply the cost of generated speech. Developers also needed to consider streamed input, transcription, text context, tool interactions, retrieved knowledge, repeated history after connection renewal, and the infrastructure required to run a real-time application. Because Nova Sonic has reached its documented end-of-life date, current AWS pricing and availability for a supported successor should be checked instead.
Main strengths and limitations
Nova Sonic’s main technical strength was its focus on live, two-way voice interaction. A unified speech-to-speech approach could simplify an architecture that would otherwise chain speech recognition, a text model, and speech synthesis. Streaming support, expressive voices, interruption handling, and transcription were particularly relevant to conversational applications where waiting for a full audio turn would feel unnatural.
The 300,000-token context window was also substantial for maintaining conversation history or supplying retrieved enterprise information, although large context does not by itself guarantee accurate recall or low cost. Function calling and Knowledge Bases integration extended the model beyond simple scripted conversation.
Its limitations were significant. The model was specialized for speech interaction rather than image or video generation, embedding workloads, batch text processing, or coding-focused work. The eight-minute default connection limit required session-management logic for longer conversations. Availability and language support depended on AWS documentation and region. Most importantly, the model is now legacy and past its recorded end-of-life date, making it unsuitable as the default choice for a new deployment.
Reasoning, coding, speed, and cost trade-offs
AWS positioned Nova Sonic around responsive voice conversations rather than deep reasoning or software development. The supplied editorial evaluation rates its reasoning capability at 5 out of 10 and coding capability at 2 out of 10; these are database evaluations, not provider-published benchmark results. They suggest that the model may be adequate for structured conversational tasks and tool-mediated workflows but should not be selected as a specialist reasoning or coding model.
The same editorial evaluation rates speed at 9 out of 10 and cost at 7 out of 10. These scores are also subjective comparative assessments, not guaranteed latency or price measurements. The high speed assessment reflects the model’s low-latency streaming purpose, while the cost assessment should be interpreted alongside Bedrock’s token billing and the additional processing associated with audio, transcripts, tools, and context.
For a voice assistant, a speech-specialized model may be preferable to a general text model connected to separate speech services when conversational responsiveness and simpler orchestration matter most. Conversely, a general-purpose reasoning model or a dedicated coding model is more appropriate for complex analysis, code generation, or workflows that do not need live audio.
When Nova Sonic would have been appropriate
Before its end of life, Nova Sonic was a reasonable fit for applications with all or most of these requirements:
- Users communicate primarily by speaking rather than typing.
- The application needs responses while the conversation is still in progress.
- Interruptions and natural turn-taking matter.
- The application must provide both spoken responses and text transcripts.
- External functions, account systems, or enterprise knowledge need to be connected.
- The deployment can operate in a supported AWS Region and manage the documented connection duration.
It was a poor fit for new applications that require long-term lifecycle stability, image or video output, high-end coding, batch processing, or a clearly verified current price and availability. For those projects, a currently supported speech model, including the distinct Nova 2 Sonic successor where its documented capabilities meet the requirements, should be evaluated instead.
Bottom line
Amazon Nova Sonic was a focused real-time speech-to-speech model that combined streamed voice interaction, transcription, expressive output, tool use, and enterprise grounding. Its 300,000-token context window and bidirectional streaming design made it relevant to voice assistants and customer-service systems, while its specialization limited its usefulness for coding, image generation, video generation, and general batch workloads.
Its lifecycle now determines the practical verdict: AWS recorded September 14, 2026 as the official end-of-life date, and the supplied research indicates that date has passed. Nova Sonic remains useful to understand as the first-generation Amazon Nova speech model, but new systems should verify a supported replacement rather than build around amazon.nova-sonic-v1:0.

