Lyria 3

Lyria 3 Pro

by Google DeepMind · Public Preview

Lyria 3 Pro is Google DeepMind's public-preview music model for generating complete songs up to approximately three minutes. It accepts text and image prompts, supports vocals, lyrics, instrumentation, tempo, structure, and multiple vocal languages, and costs $0.08 per full-song generation.

Text Music Reasoning Coding
Lyria 3 Pro is a specialized generative music model from Google DeepMind, available through Google Cloud's generative AI platform. Unlike general-purpose models that focus on text, code, or visual generation, it is designed to produce complete musical compositions from natural-language descriptions and image-based creative references. Users can guide vocals, lyrics, instruments, tempo, arrangement, and song structure, making the model suitable for songwriting, soundtracks, advertising, games, and other creative production workflows.
Outputs

What Lyria 3 Pro can produce

Text Music
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Multimodal output
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
7/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Lyria 3
Model type Other
Release date 2026-03-25
Status Public Preview
Knowledge cutoff notes

Google's publicly available documentation does not specify a knowledge cutoff for Lyria 3 Pro. The model is a generative music system rather than a general-purpose knowledge model.

Model notes

Canonical Vertex AI model ID is lyria-3-pro-preview. The model generates complete compositions up to approximately three minutes and supports text and image conditioning, vocals, user-provided or generated lyrics, structural controls, tempo guidance, and multiple vocal languages. Published pricing is per full-song generation rather than token-based input/output pricing. Outputs include SynthID watermarking and support C2PA content credentials. Editorial scores reflect the model's specialized music-generation role; reasoning and coding scores are intentionally minimal because those are not its intended capabilities.

Cost

Model pricing

Input $0.08 per full-song generation
Model guide

Lyria 3 Pro: Full-Length AI Music Generation from Text and Images

Lyria 3 Pro is Google DeepMind's public-preview music generation model for creating complete, high-fidelity songs of up to approximately three minutes from text and image inputs. It supports vocals, lyrics, instrumentation, tempo, musical structure, multiple vocal languages, SynthID watermarking, and C2PA content credentials.

What is Lyria 3 Pro?

Lyria 3 Pro is Google DeepMind's specialized model for generating complete music. It belongs to the Lyria 3 model family and is offered through Google Cloud as the public-preview model lyria-3-pro-preview. Its primary output is music audio rather than a written response, image, video, or software program.

The model is designed for full-song generation. Google describes compositions of up to approximately three minutes, which places Lyria 3 Pro closer to a songwriting and soundtrack tool than to a system intended only for short musical sketches. A prompt can describe a genre, mood, instrumentation, vocal style, language, tempo, and arrangement, while structural instructions can specify elements such as an intro, verse, chorus, bridge, or transition.

This specialization is important when evaluating the model. Lyria 3 Pro should not be treated as a general-purpose conversational AI. Its value comes from producing music with controllable musical characteristics, not from answering factual questions, writing software, performing web research, or carrying out tool-based workflows.

Where Lyria 3 Pro fits in Google DeepMind's lineup

Lyria 3 Pro sits within Google DeepMind's Lyria 3 family of generative music models. The model is distributed through Google Cloud's generative AI services, so its intended audience includes developers and creative teams building music generation into applications as well as users working with music-production workflows.

The available research identifies the model's public-preview status and canonical model ID, but it does not provide a complete feature-by-feature comparison with every other Lyria model. The reliable positioning is therefore its role as a Pro variant focused on complete compositions, detailed musical direction, and production-oriented use cases. It should be selected for music generation rather than as a replacement for Google's general-purpose language or multimodal models.

What can Lyria 3 Pro create?

Lyria 3 Pro can generate high-fidelity stereo music with vocal or instrumental arrangements. It supports both user-provided lyrics and model-generated lyrics, allowing a prompt to request either an instrumental track or a song with a defined vocal performance.

  • Complete songs: compositions can run for up to approximately three minutes.
  • Text-to-music generation: natural-language prompts describe the desired musical result.
  • Image-to-music generation: an image can provide a creative reference or influence the resulting composition.
  • Vocals and instrumentals: users can request songs with vocals or music without vocals.
  • Lyrics: the model supports user-provided or generated lyrics.
  • Musical structure: prompts can describe intros, verses, choruses, bridges, transitions, and other arrangement elements.
  • Musical controls: prompts can guide genre, mood, instruments, vocal characteristics, tempo, and language.
  • Audio quality: Google describes the output as high-fidelity stereo audio.

The documented vocal languages include English, German, Spanish, French, Hindi, Japanese, Korean, and Portuguese. This makes the model relevant to multilingual songwriting and localized creative content, although the research does not provide comparative quality measurements across those languages.

Inputs and outputs

The main input types are text and images. Text is used to describe the intended song, while an image can act as a visual reference for the musical direction. Google materials also describe image and PDF conditioning in the prompting guidance, but the core model documentation identifies text and image inputs.

The primary output is music audio, including vocal or instrumental content. Lyrics are also identified as an output in the model card. Lyria 3 Pro does not provide conventional long-form text completion, image output, video output, speech synthesis, transcription, embeddings, or structured JSON responses as its main capabilities.

CategorySupported or documented behavior
Model IDlyria-3-pro-preview
Input modalitiesText and images
Primary outputMusic audio
Additional outputLyrics support
Audio format descriptionHigh-fidelity stereo audio
Maximum composition lengthUp to approximately three minutes
Context lengthNot publicly specified in the supplied research
Maximum output tokensNot applicable or not publicly specified for this music-generation model

Prompting and musical control

Lyria 3 Pro is intended to respond to detailed musical direction rather than only broad genre labels. A practical prompt might combine the desired style with the arrangement, instrumentation, vocal treatment, tempo, emotional tone, and song structure. For example, a creator could request an upbeat electronic track with a short instrumental intro, a verse led by female vocals, a larger chorus, a bridge with reduced instrumentation, and a defined tempo.

These controls do not mean that every generated track will follow the requested structure perfectly. The supplied research confirms that the model supports structural and musical guidance, but it does not provide a benchmark for instruction-following accuracy or a guarantee that every element of a prompt will appear exactly as requested. Human review remains important when timing, lyrics, arrangement, or vocal characteristics must meet a precise production brief.

Image conditioning adds a different form of direction. Instead of describing every musical idea in text, a user can provide a visual reference and ask the model to create music inspired by it. This can be useful for concept art, advertising treatments, game scenes, mood boards, or video projects in which the visual material already establishes a tone.

Pricing and availability

Lyria 3 Pro is listed as a public-preview model on Google Cloud. The documented price is $0.08 per full-song generation. This is a per-generation charge, not a token-based input or output price. The supplied information does not specify whether pricing varies by region, account configuration, retries, or additional platform services, so those details should be checked in the applicable Google Cloud billing documentation before deployment.

The model's public-preview status also matters operationally. Preview services can have changing limits, availability, interfaces, or terms. Teams evaluating Lyria 3 Pro for production should confirm current regional availability, quotas, authentication requirements, and usage policies in Google Cloud documentation rather than assuming that preview behavior is permanent.

Safety, authenticity, and rights considerations

Google states that Lyria 3 outputs include SynthID watermarking and support the C2PA standard for signed content metadata. SynthID is intended to help identify content associated with Google's generative systems, while C2PA provides a framework for provenance and content credentials. These measures can help downstream users understand how an audio asset was created, but they do not replace human rights review or legal clearance.

Google also describes filtering for input and output content, including safeguards related to vocal likeness and potential similarity to existing content. Such protections are provider claims about the model's safety design, not a guarantee that every legal or ethical risk has been eliminated. Users remain responsible for checking lyrics, generated melodies, vocal characteristics, samples or references, and intended distribution against applicable copyright, publicity-rights, licensing, and platform requirements.

Main strengths and limitations

The strongest reason to use Lyria 3 Pro is its focus on complete, structured music. It can take a relatively high-level creative brief and produce a song-length result with vocals, lyrics, instrumentation, tempo, and arrangement guidance. Image conditioning broadens the ways a creative team can begin a music concept, while support for multiple vocal languages makes the model relevant beyond English-language projects.

Its limitations follow directly from that specialization. Lyria 3 Pro is not a general-purpose reasoning model and does not provide the capabilities normally associated with chat-oriented systems. The supplied research does not document tool use, function calling, streaming, fine-tuning, caching, batch processing, context length, or a conventional JSON mode. Those values should be treated as unverified rather than assumed to be supported.

  • It is not intended for general text generation or question answering.
  • It is not a coding or software-development model.
  • It should not be selected for web research, transcription, speech-to-speech, or speech synthesis.
  • It does not generate images or videos.
  • Generated tracks may need editing, mixing, mastering, lyric review, and rights validation.
  • The approximately three-minute maximum may be restrictive for longer scores, extended mixes, or complete album-length arrangements.
  • Public-preview availability may involve changing quotas, interfaces, or service conditions.

Reasoning, coding, and tool capabilities

Reasoning and coding are not meaningful strengths of this model. The supplied editorial evaluation assigns minimal scores in both areas because Lyria 3 Pro is a music-generation system, not because it has been benchmarked as a weak general-purpose language model. It is better understood as an audio creation engine that interprets musical instructions.

No tool-use or function-calling capability is documented in the supplied research. The model also has no documented web-search capability or general-purpose action output. Applications may still place the model inside a larger workflow that handles file management, approvals, editing, or publishing, but those surrounding functions should not be confused with capabilities of Lyria 3 Pro itself.

When to choose Lyria 3 Pro

Choose Lyria 3 Pro when the desired result is a complete song or structured music track rather than text or a short sound effect. It is particularly suitable for:

  • Songwriting and early-stage musical ideation.
  • Custom background music for video, games, advertising, and social content.
  • Rapid soundtrack concepts based on a written brief or visual reference.
  • Vocal or instrumental drafts in one of the documented supported languages.
  • Applications that need per-generation music creation through Google Cloud.
  • Creative teams that want to explore several arrangements before commissioning or producing a final recording.

Another type of option may be more appropriate when the project requires precise audio editing, long-form orchestral scoring, deterministic control over individual tracks, professional mastering, speech generation, transcription, or general-purpose reasoning and coding. A human musician, digital audio workstation, conventional sample library, or a model designed specifically for those tasks may provide better control. Similarly, a general multimodal language model is a better choice when music generation is only a small part of a broader workflow involving research, document analysis, coding, or tool calls.

Practical evaluation

Lyria 3 Pro's documented price of $0.08 per full-song generation makes it possible to estimate basic generation costs from the number of complete tracks requested. However, the real cost of a project can also include discarded generations, human selection, lyric changes, editing, mixing, mastering, storage, and downstream distribution. A team should therefore test representative prompts rather than judging the service only by the nominal generation fee.

For evaluation, compare whether the model reliably follows requested structure, preserves the intended mood, produces acceptable vocals and lyrics, and creates output that can be edited into the target production. Check multilingual results separately if the project is not in English. Also verify the treatment of watermarks, C2PA metadata, and rights documentation in the final publishing workflow.

Overall, Lyria 3 Pro is best viewed as a specialized full-song generation model in public preview. Its appeal is the combination of song-length output, text and image conditioning, musical structure controls, vocal and lyric support, and a clear per-generation price. Its boundaries are equally clear: it is not a general AI assistant, and generated music still requires creative judgment, production work, and legal review before commercial use.


Answers to Frequently Asked Questions

What limitations and rights considerations apply to Lyria 3 Pro?
Lyria 3 Pro is a specialized music-generation model rather than a general-purpose AI assistant, coding model, transcription service, or tool-using system. Generated tracks may require editing, mixing, mastering, lyric review, and human quality control. Google states that Lyria 3 outputs include SynthID watermarking and support the C2PA standard, but users remain responsible for copyright, licensing, publicity-rights, and other legal checks before distribution.
How much does Lyria 3 Pro cost?
Lyria 3 Pro is listed on Google Cloud as the public-preview model `lyria-3-pro-preview`, with a documented price of $0.08 per full-song generation. Regional availability, quotas, retries, account configuration, and additional platform charges should be confirmed in the current Google Cloud billing documentation.
How long can a song generated by Lyria 3 Pro be?
Lyria 3 Pro can generate compositions of up to approximately three minutes. This makes it suitable for complete songs, soundtrack concepts, background music, advertising, games, video, and social content, but potentially restrictive for long-form scores or extended mixes.
What is Lyria 3 Pro used for?
Lyria 3 Pro is Google DeepMind's specialized music-generation model for creating complete songs and structured music tracks from text and images. It can generate vocal or instrumental compositions with guidance for genre, mood, instruments, tempo, language, lyrics, and song structure.
What inputs and outputs does Lyria 3 Pro support?
Lyria 3 Pro accepts text prompts and images as creative references. Its primary output is high-fidelity stereo music audio, with support for vocal or instrumental arrangements and user-provided or model-generated lyrics.


Sources 6
Provider

About Google DeepMind