What is Amazon Nova 2 Sonic?
Amazon Nova 2 Sonic is Amazon's real-time speech-to-speech foundation model for conversational applications. It is available through Amazon Bedrock under the model ID amazon.nova-2-sonic-v1:0. Instead of requiring separate speech-recognition, language-generation, and text-to-speech systems, Nova 2 Sonic combines these stages into one conversational workflow.
In practical terms, an application can stream a user's speech to the model, receive generated speech and text, and maintain a multi-turn conversation with low latency. The model is designed for voice-first experiences rather than image generation, video generation, embeddings, or general offline batch processing.
Amazon positions Nova 2 Sonic for customer-service call automation, voice assistants, telephony workflows, interactive education, language learning, and conversational media applications. It is part of the Amazon Nova model family and is delivered as a managed service through Amazon Bedrock rather than as a self-hosted model.
How Nova 2 Sonic handles voice conversations
Nova 2 Sonic uses Amazon Bedrock's bidirectional streaming operation, InvokeModelWithBidirectionalStream, for its primary real-time workflow. Bidirectional streaming means the client can continue sending conversation events while receiving model events, rather than waiting for a complete request and response before the next turn begins.
The model can return audio responses alongside text responses and transcription events. This allows an application to play speech to a user while also using text for captions, logging, accessibility, or downstream application logic. Text and audio can also be used in the same session, so a conversation does not have to remain exclusively voice-based.
Nova 2 Sonic supports barge-in handling. If the user interrupts the assistant, the application can stop or clear generated audio and continue the conversation. This is important for natural phone and voice-assistant interactions, where waiting for a long response to finish can make the system feel unresponsive.
Core capabilities and supported modalities
- Audio input: The model accepts streamed speech for real-time conversations.
- Text input: Applications can send text and switch between text and audio within the same session.
- Speech output: Nova 2 Sonic generates audio responses for voice interfaces.
- Text output: It can also return text responses and transcription-related events.
- Streaming: Event-based bidirectional streaming supports low-latency, multi-turn conversations.
- Tool use: The model can request external functions, APIs, database queries, and other tools described with JSON schemas.
- Asynchronous tool calling: External operations can run in the background while the conversation continues.
- Cross-modal sessions: Audio and text can be combined without requiring separate conversations.
Tool calling does not mean that Nova 2 Sonic directly operates every external system. The application still has to execute the requested function, return the result, manage permissions, and decide which actions require confirmation. This separation is useful for customer-service systems, appointment workflows, account lookups, and other applications that need model-generated requests connected to controlled business logic.
Languages and voices
Nova 2 Sonic supports seven languages: English, French, Italian, German, Spanish, Portuguese, and Hindi. AWS documents multiple voice options, including feminine-sounding and masculine-sounding voices for supported locales.
The Tiffany and Matthew voices are polyglot voices that can speak across all supported languages. This can help an application maintain a consistent assistant voice when a user changes languages during a session. Language and voice availability should still be checked against the current AWS documentation and configuration supported in the target Region.
Context window and output limits
The model supports a context window of up to 1 million tokens. A context window is the amount of conversational and other input information the model can consider within a request or continuing session. A large window can be useful for long customer-service interactions, extensive conversation history, or applications that need to retain substantial structured context.
Nova 2 Sonic has a maximum output limit of 64,000 tokens. That limit is substantially larger than the amount of speech normally needed for a single voice turn, but it can matter when the same model is used for text-heavy responses or extended generated content. A practical voice application should still control response length because shorter turns generally produce a more usable conversational experience.
A streaming connection is subject to an eight-minute connection limit according to AWS documentation. Longer conversations require an application-level continuation pattern, such as transferring relevant session state into a new connection. The connection limit is therefore an implementation constraint even though the model's context window is large.
Pricing and developer access
Nova 2 Sonic is available through Amazon Bedrock in supported AWS Regions, including US East (N. Virginia), US West (Oregon), Asia Pacific (Tokyo), and Europe (Stockholm). Availability can vary by Region and inference configuration, so deployment teams should verify access before designing around a specific location.
AWS examples identify speech pricing of approximately $0.003 per 1,000 speech-input units and $0.012 per 1,000 speech-output units. These are usage-based charges rather than a fixed monthly subscription. Text input and output may be metered separately, and pricing can vary by modality, Region, and service tier. The current Amazon Bedrock pricing documentation should be checked before production deployment.
The different input and output rates make usage patterns important. An application that generates lengthy spoken answers, keeps users on long calls, or frequently repeats audio may incur more output usage than a short-turn assistant. Cost estimates should model both sides of the conversation, along with any separate AWS services used for application hosting, storage, telephony, databases, or tool execution.
Reasoning and coding profile
Nova 2 Sonic is primarily a real-time conversational speech model. Its editorial reasoning score is 6 out of 10, while its editorial coding score is 4 out of 10. These scores are comparative evaluations for this database, not benchmarks or ratings published by Amazon.
The model can interpret conversational requests, maintain context, decide when a configured tool may be relevant, and produce responses in speech or text. However, the supplied research does not establish that it is optimized for advanced mathematical reasoning, long-form analysis, software engineering, or coding-agent workloads. Teams that need a coding-focused or deeply analytical model may be better served by a different model connected to a voice layer or by a separate specialist workflow.
Speed, cost, and practical trade-offs
Nova 2 Sonic's main advantage is the combination of streaming speech input, generated speech output, and conversational tool use in one managed model workflow. Its editorial speed score is 9 out of 10 and its editorial cost score is 8 out of 10. These are subjective comparative scores, not provider-published performance or pricing ratings.
The model is a strong fit when conversational delay matters more than maximum reasoning depth. A call center assistant, language tutor, or voice-controlled service often benefits more from quick turn-taking and interruption handling than from producing the longest possible written answer. Usage-based speech pricing can also be attractive for targeted voice interactions, but total cost depends on conversation length and the amount of generated audio.
The trade-off is specialization. Nova 2 Sonic is not presented as a general-purpose model for image or video generation, embeddings, offline batch processing, or self-hosted deployment. It also requires application engineering for audio capture, playback, event ordering, session continuation, interruption handling, authentication, and tool execution.
When to choose Amazon Nova 2 Sonic
Choose Nova 2 Sonic when the central product experience is a live, spoken conversation and the application needs more than simple speech transcription. It is particularly suitable for:
- Voice assistants that need multi-turn context and natural interruption handling.
- Customer-service and contact-center automation that can call business tools.
- Telephony applications requiring streamed speech responses.
- Interactive language learning and multilingual practice.
- Voice-enabled scheduling, reservation, support, or account workflows.
- Applications that need both spoken responses and text transcripts.
Another option may be more appropriate when the primary requirement is advanced coding, image or video generation, embeddings, offline batch inference, or complete control over model hosting. A separate model may also be preferable when the application needs a specialized reasoning or coding system and voice is only a secondary interface.
Limitations and lifecycle considerations
Nova 2 Sonic is active and generally available while its documented lifecycle continues, but AWS lists an end-of-life date no sooner than December 2, 2026. Production users should monitor AWS lifecycle notices, release notes, regional availability, and any recommended successor or migration guidance.
Capabilities can also depend on the selected AWS Region, connection configuration, supported voice, application architecture, and the tools made available by the developer. The model does not remove the need for safeguards: external actions should be authenticated and authorized, tool results should be validated, and users may need confirmation before consequential actions are taken.
Overall, Nova 2 Sonic is best understood as a specialized real-time voice foundation model. Its value comes from joining streaming speech, text, multilingual voices, tool calling, and cross-modal conversation in a single Bedrock workflow. It is less suitable as a universal model for every AI task, but it is a focused option for applications where responsive spoken interaction is the primary requirement.

