What is Lyria 3.5?
Lyria 3.5 is Google DeepMind’s specialized model for generating music rather than a general-purpose conversational or language model. Its main task is turning a description of a musical idea into a full-length song. Prompts can specify characteristics such as genre, mood, instrumentation, vocal style, arrangement, or lyrical direction. The model can also accept an image, allowing the visual content to serve as inspiration for the generated music.
The model produces 44.1 kHz stereo MP3 audio along with text lyrics. This combination makes it suitable for song concepts, demos, soundtrack ideas, instrumental pieces, and other creative music workflows where both the rendered track and its lyrical content are useful.
Google DeepMind positions Lyria 3.5 within its Lyria music-model family. The model was introduced in Google Flow Music on July 29, 2026, and reached general availability through the Gemini API on September 3, 2026, according to the supplied research.
What Lyria 3.5 can create
Lyria 3.5 is designed around recognizable song structure rather than short, isolated audio snippets. It can generate music containing sections such as:
- Verses and choruses
- Bridges and other structured transitions
- Lead or supporting vocals
- Instrumental arrangements
- Text lyrics
- Prompt-controlled song duration
For example, a prompt could request an atmospheric electronic song with a gradual verse, a larger chorus, a contrasting bridge, and a particular emotional direction. Another prompt might ask for an instrumental soundtrack inspired by the colors and composition of an uploaded image. These examples describe the kinds of creative instructions supported by the model; the supplied research does not specify a fixed maximum song duration or guarantee that every prompt will produce a particular arrangement.
The output is music audio, not a multitrack project or a collection of separately editable instruments. Users should therefore treat Lyria 3.5 as a song-generation system rather than as a full digital audio workstation.
Inputs and outputs
| Capability | Lyria 3.5 support |
|---|---|
| Text input | Supported |
| Image input | Supported |
| Audio input | Not documented as supported |
| Video input | Not documented as supported |
| Music audio output | Supported |
| Text lyric output | Supported |
| Image or video output | Not supported |
| Audio format | 44.1 kHz stereo MP3 |
The model’s multimodal capability is therefore focused on text-and-image prompting for music generation. It is not a general multimodal assistant that analyzes arbitrary files and responds across many output formats.
Strengths and creative strengths
The main advantage of Lyria 3.5 is its focus on complete musical compositions. Instead of treating music as a short sound effect or isolated loop, it is intended to generate songs with a broader arrangement and a sense of progression. The supplied research specifically identifies improvements in musical coherence, vocals, lyrics, and duration control.
That focus makes the model useful for:
- Developing song ideas from a written concept
- Creating vocal or instrumental demos
- Producing background music and soundtrack drafts
- Exploring arrangements before recording with human musicians
- Turning an image, scene, or visual mood board into a musical starting point
- Generating songwriting experiments with different structures and moods
These strengths should not be confused with a guarantee of professional release-ready results. The supplied sources describe the model’s intended capabilities, but do not provide independent benchmark results or a universal quality rating. Output quality can also depend on the specificity and musical clarity of the prompt.
Technical profile and API behavior
The documented model ID is lyria-3.5. Its listed context length is 131,072 tokens. Context length describes how much input information an API request can accept in the model’s processing window; it should not be interpreted as a maximum song length. The research does not provide a separate maximum output-token value or a fixed maximum duration for generated music.
Lyria 3.5 uses a non-streaming generation workflow. The API generates the requested result rather than delivering the track progressively as a real-time audio stream. This distinction matters for interactive music applications, live accompaniment, and systems that require immediate incremental audio. Google provides a separate Lyria RealTime model for real-time music generation, so Lyria 3.5 is better understood as an on-demand full-song generator.
The model does not support function calling, grounding, code execution, structured outputs, caching, or batch API processing according to the supplied Gemini API documentation. These limitations are expected for a specialized music generator: the model is intended to return a musical result and lyrics, not to act as a general-purpose agent that calls tools or produces machine-validated JSON.
Reasoning, coding, and tool support
Lyria 3.5 is not designed for conventional reasoning or coding tasks. It can interpret a creative music prompt and use that instruction to shape a song, but this is different from the deliberate analytical reasoning offered by general-purpose language models. It should not be selected for research, software development, mathematical problem solving, or long-form factual question answering.
There is no documented tool-use or function-calling capability. The model also does not provide structured JSON output. Applications that need metadata, workflow automation, or strict machine-readable responses may need to place Lyria 3.5 inside a larger system and use a separate model for planning, validation, or orchestration.
The supplied model data includes reasoning, coding, speed, and cost scores. Those are editorial or catalog evaluations, not provider-published benchmark results. They should be treated as comparative guidance rather than formal measurements. In practical terms, the model’s value is concentrated in music generation, not in general reasoning or programming.
Pricing and cost considerations
The documented Gemini API price is $0.08 per full song. The supplied research lists no free tier for Lyria 3.5. Because the unit is a generated song rather than a token count, costs are easier to estimate for a workflow that produces a known number of tracks. For example, ten generated songs would represent $0.80 in model charges at the listed rate, before any other applicable service costs.
The flat per-song model can be attractive for creative experimentation, but repeated generation can add up when users need many variations. A workflow that generates several alternatives for every concept should account for the fact that each full-song result is billable. The research does not specify different prices for different durations, output formats, or API usage tiers, so those details should not be assumed.
Important limitations
Lyria 3.5 has several practical boundaries:
- It is specialized for music and is not a general-purpose language model.
- Generation is non-streaming, so it is not intended for real-time musical interaction.
- Music generation is currently single-turn, without iterative editing through follow-up prompts.
- It does not provide documented function calling, grounding, code execution, structured outputs, caching, or batch API support.
- The research does not specify a fixed maximum song duration or maximum output-token limit.
- It cannot be used as a speech-synthesis or transcription model.
- Safety filters block requests for specific artist voices and copyrighted lyrics.
Generated audio includes an imperceptible SynthID watermark. This is relevant for applications that need to identify or track AI-generated media. Users should also avoid assuming that a prompt requesting a particular living artist’s voice will be fulfilled; the supplied documentation explicitly notes restrictions around specific artist voices and copyrighted lyrics.
When to choose Lyria 3.5
Choose Lyria 3.5 when the central requirement is generating a complete song from a natural-language or image-based idea. It is a good fit for musicians exploring early concepts, creators producing soundtrack drafts, developers building music-generation features, and users who want a vocal or instrumental result rather than a text explanation about music.
Its full-song orientation is especially useful when structure matters. A request for a verse, chorus, bridge, vocals, and a defined mood aligns more closely with Lyria 3.5’s purpose than a request for a short sound effect or a live musical response.
Another option may be more appropriate in several cases. Use a real-time music model such as Lyria RealTime when an application needs streaming or interactive generation. Use a general-purpose language model when the task involves coding, factual research, tool calls, structured JSON, or extended analytical reasoning. Use a speech model for speech synthesis or transcription. A conventional digital audio workstation or a human production workflow remains more suitable when precise multitrack editing, mixing, instrument isolation, or detailed post-production control is required.
Bottom line
Lyria 3.5 is best evaluated as a focused AI songwriting and music-production model. It accepts text and images, generates 44.1 kHz stereo MP3 music and lyrics, and is designed to create full-length arrangements with vocals, instrumental parts, and recognizable song sections. Its $0.08-per-song API pricing provides a clear unit cost, while its non-streaming, single-turn design limits its suitability for live performance and iterative editing.
The model is most compelling when the desired output is a finished musical draft from a compact creative brief. It is not a replacement for a general AI assistant, a coding model, a speech system, or a real-time audio engine, and those distinctions should guide both model selection and application design.

