Lyria

Lyria RealTime

by Google DeepMind · Experimental and currently documented through the Gemini API as lyria-realtime-exp

Google DeepMind's experimental Lyria RealTime model continuously streams instrumental music over a WebSocket connection and supports live prompt weighting plus controls for tempo, key, density, brightness, and related musical properties.

Music Reasoning Coding
Lyria RealTime is an experimental Google DeepMind model built for live music generation rather than one-shot track creation. Available through the Gemini API as models/lyria-realtime-exp, it streams instrumental audio continuously and lets an application change musical direction while playback is underway. This makes it suited to interactive performances, prompt-driven DJ tools, MIDI-controlled experiences, and applications that need an evolving soundtrack instead of a finished audio file.
Outputs

What Lyria RealTime can produce

Music
Inputs

What it can understand

Text
Capabilities

Supported features

Streaming Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
Specifications

Technical details

Model family Lyria
Model type Other
Release date 2025-05-20
Status Experimental and currently documented through the Gemini API as lyria-realtime-exp
Knowledge cutoff notes

A knowledge cutoff is not applicable or publicly documented for this continuously controlled music-generation model.

Model notes

The canonical Gemini API identifier is models/lyria-realtime-exp. The model accepts weighted text prompts and music-generation configuration over a persistent bidirectional WebSocket connection and returns raw 16-bit PCM audio at 48 kHz stereo. Google documents a maximum control latency of approximately 2 seconds. The model is experimental and the documentation describes it as intended for instrumental music with no vocals. Applications should implement audio buffering and may need gradual prompt-weight changes or context resets to avoid abrupt transitions. No model-specific input or output price is listed on the current Gemini API pricing page.

Cost

Model pricing

Input Not publicly listed for this model
Output Not publicly listed for this model
Model guide

Lyria RealTime: Live, Steerable Instrumental Music Generation

Lyria RealTime is Google DeepMind's experimental model for continuously generating and interactively controlling instrumental music. It uses a persistent bidirectional WebSocket connection to stream 48 kHz stereo PCM audio while developers adjust weighted text prompts and musical controls such as tempo, key, density, brightness, guidance, temperature, and generation mode.

What is Lyria RealTime?

Lyria RealTime is Google DeepMind's experimental model for interactive music creation. Its defining characteristic is continuous generation: instead of accepting a prompt and returning a completed song or short audio clip, it maintains an active music stream that can be influenced during playback.

The model is designed primarily for instrumental music. Developers can describe musical ideas with text, combine several descriptions using weights, and modify musical settings while the model is running. For example, an application could blend ambient pads with electronic percussion, then gradually increase the influence of the percussion or change the musical direction as a live visual scene develops.

Its canonical Gemini API model identifier is models/lyria-realtime-exp. The experimental status is important: access conditions, quotas, behavior, and availability may change, and Google does not publish a model-specific price for Lyria RealTime on the current Gemini API pricing page.

How the streaming model works

Applications connect to Lyria RealTime through a persistent, bidirectional WebSocket session. A WebSocket is a long-lived network connection that allows both sides to send messages without repeatedly starting a new request. In this case, the model sends audio to the application while the application can send prompts and configuration updates back to the model.

The generated audio is delivered as raw 16-bit PCM at 48 kHz in stereo. PCM is an uncompressed representation of audio samples, so the receiving application must be prepared to process, buffer, and play the data. This differs from receiving a conventional compressed music file such as MP3 or a completed downloadable track.

Continuous streaming also creates engineering responsibilities. Network jitter or temporary variation in generation latency can interrupt playback unless the application maintains an audio buffer. Google documents control latency of up to approximately two seconds. Prompt updates therefore should not be treated as guaranteed instantaneous changes.

Prompting and live musical control

Lyria RealTime accepts weighted text prompts. Weighted prompts let developers express several musical concepts and control their relative influence. A session might combine descriptions such as “minimal techno,” “warm analog synthesizers,” and “cinematic tension,” with different weights for each idea.

In addition to descriptive prompts, the model supports music-generation configuration for musical properties including:

  • Tempo or BPM
  • Musical key or scale
  • Density
  • Brightness
  • Guidance
  • Temperature
  • Generation mode

These controls are useful when the model is embedded in an interactive application rather than used as a simple prompt box. A MIDI controller, for example, could influence musical parameters, while a visual installation could adjust prompt weights in response to user interaction.

Live steering requires careful transition design. Abruptly changing a prompt can lead to an abrupt musical change. Developers may obtain smoother results by interpolating prompt weights over time or by cross-fading audio. Some changes, including certain BPM or scale adjustments, may require resetting the model context, which can create a hard transition.

Verified technical specifications

PropertyDocumented detail
ProviderGoogle DeepMind
Model familyLyria
API model IDmodels/lyria-realtime-exp
StatusExperimental
InputText weighted prompts and music-generation configuration
OutputRaw 16-bit PCM audio
Audio format48 kHz stereo
ConnectionPersistent bidirectional WebSocket
Documented control latencyUp to approximately two seconds
VocalsNot supported according to the model documentation

Google does not document a conventional text context window, maximum output-token count, or model-specific audio-duration limit for Lyria RealTime in the supplied research. Those language-model fields are not directly applicable to a continuously controlled music stream. Likewise, no public model-specific input or output price is listed.

Modalities and capabilities

Lyria RealTime accepts text prompts and music-generation settings, then produces audio. The documented input is not a general multimodal input interface: the supplied specifications do not identify image, video, or audio input support. Its output is direct audio rather than text, images, or video.

The model's audio capability is specifically music generation. It is not documented as a speech synthesizer, transcription system, general audio-analysis model, or text-generation model. It also does not provide general-purpose reasoning or coding output. Any reasoning or code used in an application belongs to the surrounding software rather than to Lyria RealTime's generated output.

Tool use and function calling are not documented as model capabilities. The WebSocket connection is an API transport mechanism, not evidence that the model can independently call external tools. Developers can build controls, MIDI integration, visual synchronization, or other application logic around the model, but those features are implemented by the application.

Where Lyria RealTime is strongest

The model's main strength is interactive control over an ongoing instrumental performance. A conventional music generator may be useful when the desired result is a completed clip. Lyria RealTime is more appropriate when the music needs to respond to changing input.

  • Continuous playback: It is designed to maintain a stream rather than repeatedly generate disconnected clips.
  • Live steering: Developers can send prompt and configuration changes during an active session.
  • Musical parameter control: Tempo, key, density, brightness, guidance, temperature, and generation mode provide more direct control than a single descriptive prompt.
  • Weighted concepts: Multiple styles, instruments, moods, or other musical ideas can be blended with adjustable influence.
  • Performance-oriented output: 48 kHz stereo PCM is suitable for applications that need to process and play audio directly.
  • Interactive application design: A persistent bidirectional connection supports tools such as prompt-driven DJ interfaces and MIDI-controlled music systems.

These strengths are practical rather than benchmark-based. The supplied research does not provide comparative listening scores, latency benchmarks against competing systems, or a claim that Lyria RealTime produces better music than other models.

Limitations and trade-offs

Lyria RealTime is not a general-purpose song-generation model. Its documented focus is instrumental music with no vocals. It is therefore a poor fit for a conventional finished song that requires lyrics, lead singing, or a polished vocal arrangement. For those needs, Google's documentation directs users toward newer non-streaming Lyria models rather than this real-time model.

The experimental status introduces uncertainty. Quotas, access, supported features, and operational behavior can change. The absence of a published model-specific price also makes it difficult to calculate a reliable per-minute production cost from the supplied information.

Real-time streaming shifts complexity to the developer. The application needs audio buffering to handle network jitter and variable generation latency. It must also decide how to handle dropped connections, prompt transitions, hard resets, and audio playback. A prototype may be straightforward, but a production-quality experience needs explicit transition and recovery logic.

Musical changes may not be perfectly smooth. Abrupt prompt edits can create abrupt transitions, and changes to BPM or scale may require resetting context. If an application requires precise, repeatable arrangement structure, a fixed audio asset or a non-streaming generation workflow may be easier to manage.

Best use cases

Lyria RealTime is a strong candidate for applications where music should react to people, controls, or events:

  • Interactive music installations and audiovisual performances
  • Prompt-driven DJ and live remix interfaces
  • MIDI-controlled music applications
  • AI-assisted improvisation and musical ideation
  • Adaptive background music for games, exhibits, or live environments
  • Creative tools that let users explore genre, mood, instrumentation, tempo, and intensity in real time

It is especially relevant when the user values immediate musical exploration over a predictable, finalized arrangement. The model can provide a continuously evolving instrumental bed that responds to an application's state.

When to choose Lyria RealTime

Choose Lyria RealTime when the central requirement is live, steerable instrumental audio. It is a suitable choice for developers who need a long-running stream, weighted prompt control, and musical parameters that can change during a session.

Choose another type of model when the requirement is a completed song, vocals, lyrics, a fixed-length export, or a highly repeatable arrangement. A non-streaming music model may be more appropriate for generating finished tracks, while a conventional audio library may be preferable when predictable timing, licensing certainty, and exact repeatability matter more than generative variation.

It is also worth considering the engineering trade-off. Lyria RealTime offers lower-latency interaction than a workflow based on repeatedly requesting complete audio clips, but it requires a live connection, buffering, playback handling, and transition management. Because Google does not publish a model-specific price in the supplied documentation, cost comparisons should be made only after confirming the applicable Gemini API access and billing terms.

Availability and pricing

Lyria RealTime is available through the Gemini API using the experimental identifier models/lyria-realtime-exp. It can also be experienced through Google AI Studio integrations. Availability and quotas may vary because the model is experimental.

No model-specific input price, output price, token price, audio-second price, or request price is listed on the current Gemini API pricing page in the supplied research. Consequently, a precise cost estimate cannot be verified here. Developers should check Google's current Gemini API documentation and account-specific terms before deploying a production application.

Overall, Lyria RealTime occupies a specialized position in Google's Lyria family: it is aimed at real-time instrumental performance and interactive control, not general language work or conventional one-shot song generation. Its value comes from the ability to shape a live musical stream, while its main costs are experimental availability and the implementation work required to make streaming transitions reliable.


Answers to Frequently Asked Questions

What are the main limitations of Lyria RealTime?
Lyria RealTime is experimental, so access, quotas, features, and behavior may change. Developers must also manage audio buffering, network interruptions, prompt transitions, and control latency of up to approximately two seconds. Google does not currently publish a model-specific price for it on the Gemini API pricing page.
Does Lyria RealTime generate vocals or complete songs?
No. Lyria RealTime is designed for instrumental music and does not support vocals according to the model documentation. It is intended for continuously evolving live audio rather than finalized songs with lyrics, singing, or a fixed arrangement.
What musical controls does Lyria RealTime support?
Lyria RealTime supports weighted text prompts and controls for tempo or BPM, musical key or scale, density, brightness, guidance, temperature, and generation mode. Developers can adjust these settings during an active session, although some changes may require resetting the model context.
What is Lyria RealTime?
Lyria RealTime is Google DeepMind's experimental model for interactive instrumental music generation. It continuously streams music that applications can influence during playback through text prompts and musical configuration settings.
How do developers connect to Lyria RealTime?
Applications connect to Lyria RealTime through a persistent, bidirectional WebSocket session using the Gemini API model identifier models/lyria-realtime-exp. The model sends raw 16-bit PCM audio at 48 kHz in stereo, while the application can send prompts and configuration updates.


Sources 5
Provider

About Google DeepMind