SeedRealtime

SeedRealtime

by ByteDance Seed · Current; fully rolled out for large-scale audio-visual full-duplex deployment

SeedRealtime is ByteDance Seed’s audio-visual full-duplex model for continuous real-time interaction. It combines speech, video, text, temporal context, proactive scene awareness, natural turn-taking, and tool-assisted responses. It is best suited to live voice-and-vision assistants, but public documentation does not yet specify its exact model ID, context window, output limit, pricing, or complete API.

Text Speech Reasoning Coding
SeedRealtime is designed for conversations that do not fit the usual question-and-answer pattern. Rather than waiting for a user to finish a turn and then processing a static prompt, it continuously interprets speech, video, timing, and context. ByteDance presents it as a full-duplex model for natural audio-visual interaction, with examples including scene-aware reminders, real-time task guidance, visual reference resolution, noise-resistant conversation, and tool-assisted lookups.
Outputs

What SeedRealtime can produce

Text Speech
Inputs

What it can understand

Text Images Audio Video Multimodal input
Capabilities

Supported features

Tool use Web search Streaming Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
9/10 Speed
Specifications

Technical details

Model family SeedRealtime
Model type Multimodal
Release date 2026-08-05
Status Current; fully rolled out for large-scale audio-visual full-duplex deployment
Knowledge cutoff notes

ByteDance’s public SeedRealtime materials do not state an exact knowledge cutoff. The model can use current external information through demonstrated online or tool-assisted interaction, but that does not establish the underlying model cutoff.

Model notes

SeedRealtime is described by ByteDance as a native audio-visual full-duplex LLM that unifies audio, video, text, temporal understanding, conversational timing, and spoken interaction. Official demonstrations include speaker and face association, visual reference resolution, proactive reminders, real-time task guidance, background-noise suppression, and tool-assisted online lookups. The exact public model identifier, API schema, context length, maximum output length, pricing, knowledge cutoff, fine-tuning support, caching, and batch API support are not publicly specified in the reviewed first-party sources. Multimodal output is marked positive because the model natively speaks and produces audio interaction; this does not indicate image or video generation.

Model guide

SeedRealtime: ByteDance’s Full-Duplex Audio-Visual AI for Natural Live Interaction

SeedRealtime is ByteDance Seed’s native audio-visual full-duplex large language model for continuous real-time interaction. It combines audio, video, text, temporal context, conversational timing, and spoken responses to support scene-aware assistants that can listen, watch, respond, pause, and proactively act during an ongoing exchange.

What is SeedRealtime?

SeedRealtime is ByteDance Seed’s native audio-visual full-duplex large language model. ByteDance announced it on August 5, 2026, and describes it as a model for omni-modal natural interaction. In practical terms, it is intended to listen and watch continuously while maintaining an ongoing conversation, rather than treating every exchange as an isolated text prompt.

The model jointly handles audio, video, text, and temporal information. This lets it connect what a person says with what is visible on screen or in the surrounding environment. For example, a spoken reference to “that item” can be interpreted using the current visual scene, while an earlier event in a live stream can remain relevant to a later response.

SeedRealtime is part of ByteDance Seed’s broader foundation-model portfolio, but its role is specific: real-time audio-visual interaction. It is not presented as a general image-generation, video-generation, embedding, or batch text-processing model. ByteDance’s current materials position it alongside other specialized Seed systems, including models for agentic productivity, creative generation, audio creation, and robotics.

How its full-duplex interaction works

In a conventional voice assistant, the system often waits for a clear end point before producing a response. Full-duplex interaction is different: the model can receive incoming audio and visual information while deciding whether to speak, pause, continue listening, or remain silent. This is closer to the timing of a human conversation.

SeedRealtime is designed to evaluate conversational state continuously. That matters in environments where interruptions, background speech, or delays can make an assistant difficult to use. The model’s published demonstrations emphasize more natural turn-taking, fewer false triggers, and the ability to avoid speaking when a response is not yet useful.

The system is also intended to respond to changing visual scenes. A user can set a goal and allow the model to monitor the environment for a relevant event. ByteDance gives examples such as reminding someone when a specified museum artifact appears, recognizing a requested section while a document is being browsed, and identifying an error while a person operates an espresso machine.

Core capabilities and practical examples

Joint audio-visual understanding

SeedRealtime aligns spoken language, visual information, and time within one interaction. This allows it to associate people with voices, interpret objects and menus, follow gestures and faces, and use visual context to resolve ambiguous speech. The distinction is important: the model is not described merely as a text model receiving captions or transcripts. Its purpose is to reason over the relationship between what is said, what is seen, and when events occur.

ByteDance also highlights performance in noisy or crowded settings. Demonstrated functions include identifying speakers and faces, resisting background noise, and retaining information observed earlier in a continuous stream. These are provider-reported capabilities rather than independently verified benchmark results.

Proactive, scene-aware responses

SeedRealtime can monitor a scene and speak when an event matches the user’s objective. This supports assistants that do more than answer direct questions. A learning assistant could explain an object when it enters view; an accessibility tool could announce a relevant visual change; and a live guide could alert a user when a requested location, object, or document section appears.

Proactive behavior must be configured carefully in real applications. An assistant that comments on every detected event would quickly become distracting. SeedRealtime’s emphasis on conversational timing and silence is therefore as important as its ability to recognize content.

Tool-assisted interaction

The model supports tool calls in the demonstrated interaction pattern. A tool call allows the model to request an external operation, such as an online lookup, while using its multimodal understanding to determine what information is needed. This means the model can support real-time multimodal workflows that combine perception, conversation, and external actions.

The available research does not establish a complete public function-calling schema or guarantee autonomous execution of arbitrary actions. Tool support should therefore be understood as a documented capability of the interaction demonstrations, not as evidence of a fully specified, generally available agent API.

Supported modalities and outputs

The documented input modalities are text, audio, and video. The model’s primary output is spoken interaction, alongside text output in the broader language-model interaction. Its audio output is significant because SeedRealtime is designed for direct, low-latency conversation rather than only returning text for a separate speech-synthesis system.

CapabilityAvailable information
Text inputSupported
Audio inputSupported
Video inputSupported
Audio or speech outputSupported; central to the full-duplex design
Text outputSupported
Image generationNot indicated
Video generationNot indicated
Tool useSupported in demonstrated tool-assisted interactions
Streaming or continuous interactionSupported by the model’s real-time design

The absence of image or video generation is a practical boundary. SeedRealtime can interpret visual input, but the supplied specifications do not say that it creates images or videos. Users seeking creative visual generation should consider a dedicated ByteDance Seed model or product instead.

Technical specifications, availability, and pricing

ByteDance states that SeedRealtime has been fully rolled out for large-scale deployment of audio-visual full-duplex technology. The official model page links to Dola and the BytePlus Playground as access points or related interfaces. However, the reviewed first-party material does not provide a conventional public developer model identifier for the exact model.

The following important specifications remain unpublished or unverified for SeedRealtime:

  • Context-window length
  • Maximum output-token limit
  • Public input and output token prices
  • Exact knowledge cutoff
  • Complete API schema
  • Public fine-tuning, caching, and batch-API support

There is therefore no verified public price to report. Access may depend on the specific ByteDance service, account, region, or deployment channel. A user should not assume that availability through Dola or the BytePlus Playground represents identical pricing, limits, or commercial terms across both services.

SeedRealtime can use online or tool-assisted information in demonstrations, but that does not establish a current internal knowledge cutoff. Real-time lookup capability and model training knowledge are separate issues.

Reasoning, coding, speed, and cost positioning

SeedRealtime’s reasoning is primarily multimodal and temporal. It must connect speech, visual evidence, timing, and conversational goals, then decide whether and when to respond. This makes it a strong fit for live interpretation and scene-aware guidance, but it does not make it a substitute for every general-purpose reasoning model.

Its coding capability is not a central design goal. The available evaluation records assign a moderate editorial coding score, but that score is an internal assessment rather than a provider-published benchmark. The same applies to the recorded reasoning and speed scores. The editorial assessment rates SeedRealtime highly for speed and real-time responsiveness, while its cost score is unknown because public pricing was not supplied.

Compared with a conventional text model, SeedRealtime trades the simplicity of turn-based text generation for continuous multimodal interaction. That trade-off is useful when latency, timing, speech, and visual awareness matter. It is less useful when the workload is large-scale batch text generation, software development, embeddings, or any task that depends on clearly documented token limits and predictable API pricing.

Main strengths and limitations

Strengths

  • Native audio-visual interaction rather than a workflow that necessarily separates vision, language, and speech.
  • Continuous full-duplex conversation with attention to pauses, interruptions, and silence.
  • Temporal understanding for connecting current observations with earlier events.
  • Scene-aware proactive behavior, including reminders and live task guidance.
  • Noise-resistant interaction and speaker or face association in demonstrated scenarios.
  • Tool-assisted workflows that can connect multimodal understanding with external lookups.

Limitations

  • The exact public model identifier and complete API specification are not documented in the supplied first-party sources.
  • Context length, maximum output length, pricing, and knowledge cutoff are not publicly specified.
  • There is no supplied evidence of image generation, video generation, embeddings, or broad batch-processing support.
  • Access, account requirements, regional availability, and commercial terms may differ between ByteDance services.
  • Public information does not establish universal fine-tuning, caching, or batch-API support.

When to choose SeedRealtime

Choose SeedRealtime when the application needs an assistant to see and hear a live environment while maintaining a natural spoken exchange. Suitable use cases include real-time voice-and-vision assistants, interactive learning, accessibility support, live explanation, guided procedures, scene-aware reminders, and proactive multimodal collaboration.

It is especially relevant when the assistant must understand references such as “the object on the left,” recognize an event as it happens, or decide whether to interrupt the user. The model’s value comes from combining perception and conversational timing, not simply from generating a high-quality standalone answer.

Another model or service may be more appropriate for batch text generation, code-heavy work, structured language-model pipelines, image or video creation, embeddings, or deployments that require published context limits and token pricing. Within the ByteDance Seed ecosystem, Seed2.1 is positioned more directly toward agentic productivity and coding, while Seedance and Seedream address creative video and image generation. Those models are alternatives by task type, not replacements for SeedRealtime’s live audio-visual interaction role.

Bottom line

SeedRealtime is a specialized real-time interaction model for systems that need to listen, watch, remember the flow of an exchange, and speak at an appropriate moment. Its strongest distinction is the combination of audio-visual understanding, temporal context, proactive scene awareness, and full-duplex conversational timing.

Its main evaluation challenge is not the concept but the incomplete public specification. Developers and buyers cannot yet verify a conventional model ID, context window, output limit, or public price from the supplied materials. For demonstrations and applications centered on live multimodal conversation, it is a compelling ByteDance Seed offering. For predictable, documented, general-purpose API workloads, a more fully specified alternative may be easier to evaluate and operate.


Answers to Frequently Asked Questions

What are SeedRealtime’s main limitations?
Public information does not specify SeedRealtime’s context window, maximum output length, knowledge cutoff, token pricing, or complete API capabilities. There is also no supplied evidence of image generation, video generation, embeddings, or broad batch-processing support. Its demonstrated capabilities and performance claims should therefore be distinguished from independently verified benchmarks or guaranteed product features.
Is SeedRealtime available through an API, and how much does it cost?
ByteDance states that SeedRealtime has been rolled out for large-scale deployment and identifies Dola and the BytePlus Playground as access points or related interfaces. However, the reviewed first-party information does not provide a verified public model identifier, complete API schema, or public pricing. Availability and commercial terms may vary by service, account, region, and deployment channel.
What modalities and outputs does SeedRealtime support?
The documented inputs are text, audio, and video. SeedRealtime supports spoken and text output, with audio output being central to its full-duplex design. It is intended to interpret visual information, but the available specifications do not indicate image generation or video generation capabilities.
What is SeedRealtime?
SeedRealtime is ByteDance Seed’s native audio-visual full-duplex large language model for real-time interaction. It jointly processes audio, video, text, and temporal information to support continuous conversations in which it can listen, observe, remember earlier events, and respond at an appropriate moment.
What can SeedRealtime be used for?
SeedRealtime is designed for live voice-and-vision assistants, interactive learning, accessibility support, live explanations, guided procedures, scene-aware reminders, and proactive multimodal collaboration. It can monitor a visual environment, connect spoken references to objects or events, and use tools such as online lookups in demonstrated workflows.


Sources 3
Provider

About ByteDance Seed