What is Lyria RealTime?
Lyria RealTime is Google DeepMind's experimental model for interactive music creation. Its defining characteristic is continuous generation: instead of accepting a prompt and returning a completed song or short audio clip, it maintains an active music stream that can be influenced during playback.
The model is designed primarily for instrumental music. Developers can describe musical ideas with text, combine several descriptions using weights, and modify musical settings while the model is running. For example, an application could blend ambient pads with electronic percussion, then gradually increase the influence of the percussion or change the musical direction as a live visual scene develops.
Its canonical Gemini API model identifier is models/lyria-realtime-exp. The experimental status is important: access conditions, quotas, behavior, and availability may change, and Google does not publish a model-specific price for Lyria RealTime on the current Gemini API pricing page.
How the streaming model works
Applications connect to Lyria RealTime through a persistent, bidirectional WebSocket session. A WebSocket is a long-lived network connection that allows both sides to send messages without repeatedly starting a new request. In this case, the model sends audio to the application while the application can send prompts and configuration updates back to the model.
The generated audio is delivered as raw 16-bit PCM at 48 kHz in stereo. PCM is an uncompressed representation of audio samples, so the receiving application must be prepared to process, buffer, and play the data. This differs from receiving a conventional compressed music file such as MP3 or a completed downloadable track.
Continuous streaming also creates engineering responsibilities. Network jitter or temporary variation in generation latency can interrupt playback unless the application maintains an audio buffer. Google documents control latency of up to approximately two seconds. Prompt updates therefore should not be treated as guaranteed instantaneous changes.
Prompting and live musical control
Lyria RealTime accepts weighted text prompts. Weighted prompts let developers express several musical concepts and control their relative influence. A session might combine descriptions such as “minimal techno,” “warm analog synthesizers,” and “cinematic tension,” with different weights for each idea.
In addition to descriptive prompts, the model supports music-generation configuration for musical properties including:
- Tempo or BPM
- Musical key or scale
- Density
- Brightness
- Guidance
- Temperature
- Generation mode
These controls are useful when the model is embedded in an interactive application rather than used as a simple prompt box. A MIDI controller, for example, could influence musical parameters, while a visual installation could adjust prompt weights in response to user interaction.
Live steering requires careful transition design. Abruptly changing a prompt can lead to an abrupt musical change. Developers may obtain smoother results by interpolating prompt weights over time or by cross-fading audio. Some changes, including certain BPM or scale adjustments, may require resetting the model context, which can create a hard transition.
Verified technical specifications
| Property | Documented detail |
|---|---|
| Provider | Google DeepMind |
| Model family | Lyria |
| API model ID | models/lyria-realtime-exp |
| Status | Experimental |
| Input | Text weighted prompts and music-generation configuration |
| Output | Raw 16-bit PCM audio |
| Audio format | 48 kHz stereo |
| Connection | Persistent bidirectional WebSocket |
| Documented control latency | Up to approximately two seconds |
| Vocals | Not supported according to the model documentation |
Google does not document a conventional text context window, maximum output-token count, or model-specific audio-duration limit for Lyria RealTime in the supplied research. Those language-model fields are not directly applicable to a continuously controlled music stream. Likewise, no public model-specific input or output price is listed.
Modalities and capabilities
Lyria RealTime accepts text prompts and music-generation settings, then produces audio. The documented input is not a general multimodal input interface: the supplied specifications do not identify image, video, or audio input support. Its output is direct audio rather than text, images, or video.
The model's audio capability is specifically music generation. It is not documented as a speech synthesizer, transcription system, general audio-analysis model, or text-generation model. It also does not provide general-purpose reasoning or coding output. Any reasoning or code used in an application belongs to the surrounding software rather than to Lyria RealTime's generated output.
Tool use and function calling are not documented as model capabilities. The WebSocket connection is an API transport mechanism, not evidence that the model can independently call external tools. Developers can build controls, MIDI integration, visual synchronization, or other application logic around the model, but those features are implemented by the application.
Where Lyria RealTime is strongest
The model's main strength is interactive control over an ongoing instrumental performance. A conventional music generator may be useful when the desired result is a completed clip. Lyria RealTime is more appropriate when the music needs to respond to changing input.
- Continuous playback: It is designed to maintain a stream rather than repeatedly generate disconnected clips.
- Live steering: Developers can send prompt and configuration changes during an active session.
- Musical parameter control: Tempo, key, density, brightness, guidance, temperature, and generation mode provide more direct control than a single descriptive prompt.
- Weighted concepts: Multiple styles, instruments, moods, or other musical ideas can be blended with adjustable influence.
- Performance-oriented output: 48 kHz stereo PCM is suitable for applications that need to process and play audio directly.
- Interactive application design: A persistent bidirectional connection supports tools such as prompt-driven DJ interfaces and MIDI-controlled music systems.
These strengths are practical rather than benchmark-based. The supplied research does not provide comparative listening scores, latency benchmarks against competing systems, or a claim that Lyria RealTime produces better music than other models.
Limitations and trade-offs
Lyria RealTime is not a general-purpose song-generation model. Its documented focus is instrumental music with no vocals. It is therefore a poor fit for a conventional finished song that requires lyrics, lead singing, or a polished vocal arrangement. For those needs, Google's documentation directs users toward newer non-streaming Lyria models rather than this real-time model.
The experimental status introduces uncertainty. Quotas, access, supported features, and operational behavior can change. The absence of a published model-specific price also makes it difficult to calculate a reliable per-minute production cost from the supplied information.
Real-time streaming shifts complexity to the developer. The application needs audio buffering to handle network jitter and variable generation latency. It must also decide how to handle dropped connections, prompt transitions, hard resets, and audio playback. A prototype may be straightforward, but a production-quality experience needs explicit transition and recovery logic.
Musical changes may not be perfectly smooth. Abrupt prompt edits can create abrupt transitions, and changes to BPM or scale may require resetting context. If an application requires precise, repeatable arrangement structure, a fixed audio asset or a non-streaming generation workflow may be easier to manage.
Best use cases
Lyria RealTime is a strong candidate for applications where music should react to people, controls, or events:
- Interactive music installations and audiovisual performances
- Prompt-driven DJ and live remix interfaces
- MIDI-controlled music applications
- AI-assisted improvisation and musical ideation
- Adaptive background music for games, exhibits, or live environments
- Creative tools that let users explore genre, mood, instrumentation, tempo, and intensity in real time
It is especially relevant when the user values immediate musical exploration over a predictable, finalized arrangement. The model can provide a continuously evolving instrumental bed that responds to an application's state.
When to choose Lyria RealTime
Choose Lyria RealTime when the central requirement is live, steerable instrumental audio. It is a suitable choice for developers who need a long-running stream, weighted prompt control, and musical parameters that can change during a session.
Choose another type of model when the requirement is a completed song, vocals, lyrics, a fixed-length export, or a highly repeatable arrangement. A non-streaming music model may be more appropriate for generating finished tracks, while a conventional audio library may be preferable when predictable timing, licensing certainty, and exact repeatability matter more than generative variation.
It is also worth considering the engineering trade-off. Lyria RealTime offers lower-latency interaction than a workflow based on repeatedly requesting complete audio clips, but it requires a live connection, buffering, playback handling, and transition management. Because Google does not publish a model-specific price in the supplied documentation, cost comparisons should be made only after confirming the applicable Gemini API access and billing terms.
Availability and pricing
Lyria RealTime is available through the Gemini API using the experimental identifier models/lyria-realtime-exp. It can also be experienced through Google AI Studio integrations. Availability and quotas may vary because the model is experimental.
No model-specific input price, output price, token price, audio-second price, or request price is listed on the current Gemini API pricing page in the supplied research. Consequently, a precise cost estimate cannot be verified here. Developers should check Google's current Gemini API documentation and account-specific terms before deploying a production application.
Overall, Lyria RealTime occupies a specialized position in Google's Lyria family: it is aimed at real-time instrumental performance and interactive control, not general language work or conventional one-shot song generation. Its value comes from the ability to shape a live musical stream, while its main costs are experimental availability and the implementation work required to make streaming transitions reliable.

