Lyria

Lyria 3.5

by Google DeepMind · Generally available

Lyria 3.5 is Google DeepMind’s specialized music-generation model for creating full-length vocal or instrumental songs from text and image prompts. It outputs 44.1 kHz stereo MP3 audio and lyrics, supports structured song sections and prompt-controlled duration, and costs $0.08 per full song through the Gemini API. Its non-streaming, single-turn workflow makes it more suitable for song drafts and soundtrack creation than for real-time music or iterative editing.

Text Music Reasoning Coding
Lyria 3.5 is a generative audio model built specifically for making songs. A user can describe a musical idea in text or provide an image for inspiration, and the model produces a complete music track with lyrics and, where appropriate, vocals and structured sections such as verses, choruses, and bridges. It is available through the Gemini API at $0.08 per full song, with no free tier documented for the model.
Outputs

What Lyria 3.5 can produce

Text Music
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family Lyria
Model type Other
Context window 131K tokens
Release date 2026-07-29
Status Generally available
Knowledge cutoff notes

Google's official documentation does not publish a conventional training-data knowledge cutoff for Lyria 3.5. The model is a generative music system rather than a general-purpose factual language model.

Model notes

Canonical model ID is lyria-3.5. Lyria 3.5 accepts text and image inputs and produces 44.1 kHz stereo MP3 audio plus text lyrics. It is optimized for full-length songs with verses, choruses, bridges, vocals, and prompt-controlled duration. The model uses a non-streaming generation workflow; real-time music generation is provided by the separate Lyria RealTime model. The Gemini API documentation lists caching, function calling, grounding, structured outputs, code execution, and batch API support as unavailable. Generated audio includes an imperceptible SynthID watermark. Safety filters block requests for specific artist voices and copyrighted lyrics. Music generation is currently single-turn, without iterative editing through follow-up prompts. The model was introduced in Google Flow Music on July 29, 2026, and reached general availability in the Gemini API on September 3, 2026.

Cost

Model pricing

Input $0.08 per full song; no free tier
Output $0.08 per full song
Model guide

Lyria 3.5: Google DeepMind’s Full-Length AI Music Generation Model

Lyria 3.5 is Google DeepMind’s specialized music-generation model for creating full-length songs from text or image prompts. It generates 44.1 kHz stereo MP3 audio and text lyrics, with support for vocals, instrumental arrangements, verses, choruses, bridges, and prompt-controlled duration. It is designed for songwriting and music creation rather than general-purpose text generation, coding, speech, or real-time music interaction.

What is Lyria 3.5?

Lyria 3.5 is Google DeepMind’s specialized model for generating music rather than a general-purpose conversational or language model. Its main task is turning a description of a musical idea into a full-length song. Prompts can specify characteristics such as genre, mood, instrumentation, vocal style, arrangement, or lyrical direction. The model can also accept an image, allowing the visual content to serve as inspiration for the generated music.

The model produces 44.1 kHz stereo MP3 audio along with text lyrics. This combination makes it suitable for song concepts, demos, soundtrack ideas, instrumental pieces, and other creative music workflows where both the rendered track and its lyrical content are useful.

Google DeepMind positions Lyria 3.5 within its Lyria music-model family. The model was introduced in Google Flow Music on July 29, 2026, and reached general availability through the Gemini API on September 3, 2026, according to the supplied research.

What Lyria 3.5 can create

Lyria 3.5 is designed around recognizable song structure rather than short, isolated audio snippets. It can generate music containing sections such as:

  • Verses and choruses
  • Bridges and other structured transitions
  • Lead or supporting vocals
  • Instrumental arrangements
  • Text lyrics
  • Prompt-controlled song duration

For example, a prompt could request an atmospheric electronic song with a gradual verse, a larger chorus, a contrasting bridge, and a particular emotional direction. Another prompt might ask for an instrumental soundtrack inspired by the colors and composition of an uploaded image. These examples describe the kinds of creative instructions supported by the model; the supplied research does not specify a fixed maximum song duration or guarantee that every prompt will produce a particular arrangement.

The output is music audio, not a multitrack project or a collection of separately editable instruments. Users should therefore treat Lyria 3.5 as a song-generation system rather than as a full digital audio workstation.

Inputs and outputs

CapabilityLyria 3.5 support
Text inputSupported
Image inputSupported
Audio inputNot documented as supported
Video inputNot documented as supported
Music audio outputSupported
Text lyric outputSupported
Image or video outputNot supported
Audio format44.1 kHz stereo MP3

The model’s multimodal capability is therefore focused on text-and-image prompting for music generation. It is not a general multimodal assistant that analyzes arbitrary files and responds across many output formats.

Strengths and creative strengths

The main advantage of Lyria 3.5 is its focus on complete musical compositions. Instead of treating music as a short sound effect or isolated loop, it is intended to generate songs with a broader arrangement and a sense of progression. The supplied research specifically identifies improvements in musical coherence, vocals, lyrics, and duration control.

That focus makes the model useful for:

  • Developing song ideas from a written concept
  • Creating vocal or instrumental demos
  • Producing background music and soundtrack drafts
  • Exploring arrangements before recording with human musicians
  • Turning an image, scene, or visual mood board into a musical starting point
  • Generating songwriting experiments with different structures and moods

These strengths should not be confused with a guarantee of professional release-ready results. The supplied sources describe the model’s intended capabilities, but do not provide independent benchmark results or a universal quality rating. Output quality can also depend on the specificity and musical clarity of the prompt.

Technical profile and API behavior

The documented model ID is lyria-3.5. Its listed context length is 131,072 tokens. Context length describes how much input information an API request can accept in the model’s processing window; it should not be interpreted as a maximum song length. The research does not provide a separate maximum output-token value or a fixed maximum duration for generated music.

Lyria 3.5 uses a non-streaming generation workflow. The API generates the requested result rather than delivering the track progressively as a real-time audio stream. This distinction matters for interactive music applications, live accompaniment, and systems that require immediate incremental audio. Google provides a separate Lyria RealTime model for real-time music generation, so Lyria 3.5 is better understood as an on-demand full-song generator.

The model does not support function calling, grounding, code execution, structured outputs, caching, or batch API processing according to the supplied Gemini API documentation. These limitations are expected for a specialized music generator: the model is intended to return a musical result and lyrics, not to act as a general-purpose agent that calls tools or produces machine-validated JSON.

Reasoning, coding, and tool support

Lyria 3.5 is not designed for conventional reasoning or coding tasks. It can interpret a creative music prompt and use that instruction to shape a song, but this is different from the deliberate analytical reasoning offered by general-purpose language models. It should not be selected for research, software development, mathematical problem solving, or long-form factual question answering.

There is no documented tool-use or function-calling capability. The model also does not provide structured JSON output. Applications that need metadata, workflow automation, or strict machine-readable responses may need to place Lyria 3.5 inside a larger system and use a separate model for planning, validation, or orchestration.

The supplied model data includes reasoning, coding, speed, and cost scores. Those are editorial or catalog evaluations, not provider-published benchmark results. They should be treated as comparative guidance rather than formal measurements. In practical terms, the model’s value is concentrated in music generation, not in general reasoning or programming.

Pricing and cost considerations

The documented Gemini API price is $0.08 per full song. The supplied research lists no free tier for Lyria 3.5. Because the unit is a generated song rather than a token count, costs are easier to estimate for a workflow that produces a known number of tracks. For example, ten generated songs would represent $0.80 in model charges at the listed rate, before any other applicable service costs.

The flat per-song model can be attractive for creative experimentation, but repeated generation can add up when users need many variations. A workflow that generates several alternatives for every concept should account for the fact that each full-song result is billable. The research does not specify different prices for different durations, output formats, or API usage tiers, so those details should not be assumed.

Important limitations

Lyria 3.5 has several practical boundaries:

  • It is specialized for music and is not a general-purpose language model.
  • Generation is non-streaming, so it is not intended for real-time musical interaction.
  • Music generation is currently single-turn, without iterative editing through follow-up prompts.
  • It does not provide documented function calling, grounding, code execution, structured outputs, caching, or batch API support.
  • The research does not specify a fixed maximum song duration or maximum output-token limit.
  • It cannot be used as a speech-synthesis or transcription model.
  • Safety filters block requests for specific artist voices and copyrighted lyrics.

Generated audio includes an imperceptible SynthID watermark. This is relevant for applications that need to identify or track AI-generated media. Users should also avoid assuming that a prompt requesting a particular living artist’s voice will be fulfilled; the supplied documentation explicitly notes restrictions around specific artist voices and copyrighted lyrics.

When to choose Lyria 3.5

Choose Lyria 3.5 when the central requirement is generating a complete song from a natural-language or image-based idea. It is a good fit for musicians exploring early concepts, creators producing soundtrack drafts, developers building music-generation features, and users who want a vocal or instrumental result rather than a text explanation about music.

Its full-song orientation is especially useful when structure matters. A request for a verse, chorus, bridge, vocals, and a defined mood aligns more closely with Lyria 3.5’s purpose than a request for a short sound effect or a live musical response.

Another option may be more appropriate in several cases. Use a real-time music model such as Lyria RealTime when an application needs streaming or interactive generation. Use a general-purpose language model when the task involves coding, factual research, tool calls, structured JSON, or extended analytical reasoning. Use a speech model for speech synthesis or transcription. A conventional digital audio workstation or a human production workflow remains more suitable when precise multitrack editing, mixing, instrument isolation, or detailed post-production control is required.

Bottom line

Lyria 3.5 is best evaluated as a focused AI songwriting and music-production model. It accepts text and images, generates 44.1 kHz stereo MP3 music and lyrics, and is designed to create full-length arrangements with vocals, instrumental parts, and recognizable song sections. Its $0.08-per-song API pricing provides a clear unit cost, while its non-streaming, single-turn design limits its suitability for live performance and iterative editing.

The model is most compelling when the desired output is a finished musical draft from a compact creative brief. It is not a replacement for a general AI assistant, a coding model, a speech system, or a real-time audio engine, and those distinctions should guide both model selection and application design.


Answers to Frequently Asked Questions

How does Lyria 3.5 differ from Lyria RealTime?
Lyria 3.5 uses non-streaming, on-demand generation for complete songs, making it suitable for song drafts and soundtrack concepts. Lyria RealTime is intended for real-time or interactive music generation, so it is a better choice for live accompaniment and applications requiring progressive audio output.
How much does Lyria 3.5 cost through the Gemini API?
The documented Gemini API price is $0.08 per generated song. There is no free tier listed in the supplied research, and ten generated songs would cost $0.80 in model charges before any other applicable service costs.
What is Lyria 3.5?
Lyria 3.5 is Google DeepMind’s specialized AI music generation model. It turns text or image-based creative prompts into full-length songs with vocals, instrumental arrangements, structured sections such as verses and choruses, and text lyrics.
What inputs and outputs does Lyria 3.5 support?
Lyria 3.5 accepts text and image inputs. It generates 44.1 kHz stereo MP3 music and text lyrics. Audio and video inputs, image or video outputs, and separate editable multitrack files are not documented as supported.


Sources 7
Provider

About Google DeepMind