Muse Spark

Muse Voice Transcribe 1.0

by Meta AI · Current and available through Meta Model API

Meta’s Muse Voice Transcribe 1.0 is a specialized speech-to-text model for realtime streams and recorded files. It supports speaker diarization for 20 or more speakers, endpointing, voice activity detection, contextual and keyword biasing, code-switching, 25 evaluated languages, and turn-level timestamps. Pricing is $3 per 1,000 minutes.

Text Reasoning Coding
Muse Voice Transcribe 1.0 is a dedicated audio perception model from Meta for applications that need speech converted into text quickly and reliably. It is designed for live voice agents, meeting and call transcription, captions, dictation, and other speech-to-text workloads rather than general conversation, text generation, or speech synthesis. The model is available through Meta Model API with usage-based pricing of $3 per 1,000 minutes.
Outputs

What Muse Voice Transcribe 1.0 can produce

Text
Inputs

What it can understand

Audio
Capabilities

Supported features

Streaming
Model profile

Performance characteristics

1/10 Reasoning
1/10 Coding
9/10 Speed
9/10 Cost efficiency
Specifications

Technical details

Model family Muse Spark
Model type Other
Context window tokens
Maximum output tokens
Release date 2026-09-01
Status Current and available through Meta Model API
Knowledge cutoff notes

Meta's reviewed first-party documentation does not state a separate knowledge cutoff for this speech-recognition model.

Model notes

The canonical API model ID is muse-voice-transcribe-1.0. It supports a realtime WebSocket endpoint for live audio and a file-transcription endpoint for recordings. Meta documents streaming speaker diarization for 20 or more speakers, native endpointing and voice activity detection, contextual and keyword biasing, 25 evaluated languages with code-switching, and turn-level timestamps. The model returns transcript text and does not synthesize speech or provide a speech-to-speech API. Pricing is usage-based at $3.00 per 1,000 minutes, equivalent to $0.18 per hour. Reasoning and coding scores are low because this is a specialized transcription model rather than a general-purpose language model; speed and cost scores are comparative editorial estimates.

Cost

Model pricing

Input $3.00 per 1,000 minutes ($0.18 per hour)
Model guide

Muse Voice Transcribe 1.0: Meta’s Real-Time Speech-to-Text Model

Muse Voice Transcribe 1.0 is Meta’s specialized speech-recognition model for converting live audio and recordings into text. Available through Meta Model API, it supports real-time WebSocket transcription, file-based transcription, speaker diarization, endpointing, voice activity detection, contextual and keyword biasing, code-switching, 25 evaluated languages, and turn-level timestamps.

What is Muse Voice Transcribe 1.0?

Muse Voice Transcribe 1.0 is Meta’s specialized automatic speech recognition model. Its job is to listen to spoken audio and return a written transcript. The canonical API model ID is muse-voice-transcribe-1.0, and Meta makes it available through Meta Model API for both live and recorded audio workflows.

Unlike a general-purpose language model, Muse Voice Transcribe is not intended to answer questions, write software, generate images, or hold a text conversation. Its output is transcript text. That narrower focus is important: the model is best evaluated as a speech-to-text component that can be placed inside a larger application, such as a voice agent, video-captioning system, meeting recorder, or call-analysis pipeline.

Meta documents two principal access patterns. A realtime WebSocket endpoint is intended for streaming audio as it is captured, while a file-transcription endpoint handles recordings. This allows the same model family to support both interactive use cases and after-the-fact transcription.

Where it fits in Meta’s model catalog

Muse Voice Transcribe 1.0 belongs to Meta’s Muse family and is presented as an audio perception model in Meta Model API. The supplied model information identifies its model family as Muse Spark. It is a focused service rather than a general-purpose text model, and the available documentation does not describe it as a speech-generation or speech-to-speech system.

That positioning separates it from broader AI assistants and language models. Muse Voice Transcribe handles the first stage of many voice applications: turning audio into text. A separate application component could then use the transcript for search, summarization, classification, customer-support workflows, or interaction with another model. Muse itself should not be assumed to perform those downstream tasks.

Core transcription capabilities

The model’s most useful features concern live audio handling and transcript structure:

  • Streaming transcription: audio can be sent through a realtime WebSocket connection so an application can receive transcription during an ongoing conversation or recording.
  • File transcription: recorded audio can be submitted through the file-transcription endpoint for non-live processing.
  • Speaker diarization: the system can distinguish speakers in a conversation. Meta documents streaming diarization for 20 or more speakers, making the feature relevant to meetings, panels, interviews, and group calls.
  • Endpointing: the model can identify likely boundaries between spoken turns. This helps voice applications decide when a person has finished speaking.
  • Voice activity detection: the system can detect portions of an audio stream that contain speech, which is useful for managing live input and reducing unnecessary processing.
  • Contextual and keyword biasing: applications can provide relevant terms or context to improve recognition of names, domain vocabulary, product terminology, or other words that may be difficult to recognize without context.
  • Turn-level timestamps: returned transcript segments can be associated with speaker turns and their timing. The supplied research does not verify word-level timestamps.
  • Code-switching: Meta describes support for conversations that move between languages rather than remaining in only one language throughout.

These features make Muse more than a basic audio-to-text converter. In particular, diarization, endpointing, and voice activity detection address practical problems that arise when transcription is used inside a live application rather than as a simple batch conversion.

Supported input and output

Muse Voice Transcribe accepts audio input and produces text output. The supplied research does not identify text input, image input, video input, or audio output as supported modalities for this model. It therefore should not be treated as a multimodal conversational model simply because it can be integrated into a broader multimodal product.

Meta’s materials describe 25 evaluated languages and code-switching support. The research supplied for this page does not provide the complete language list, language-by-language accuracy figures, or a guarantee that every language has identical performance. Availability and behavior should therefore be checked against the current Meta documentation when language coverage is critical.

The model returns transcript text rather than synthesized speech. It does not provide a speech-to-speech API, and it is not designed for sound-event detection or emotion detection. The supplied research also does not verify word-level timestamping.

How it works in live voice applications

For a live application, audio can be captured from a microphone or another audio source and sent over a realtime WebSocket connection. Muse can use voice activity detection and endpointing to identify when speech is present and when a speaker’s turn is likely to have ended. The application can then use the incoming transcript for captions, an agent response, search, logging, or another downstream action.

Speaker diarization is especially useful when several people share a call or meeting. Instead of receiving only an undifferentiated block of text, an application can associate transcript turns with different speakers. The documented support for 20 or more speakers is a provider-stated capability; actual results can still depend on audio quality, overlapping speech, microphone placement, and the number of voices present.

Contextual and keyword biasing can help when the audio contains uncommon names, technical terms, product identifiers, or industry-specific vocabulary. This does not mean every term will always be recognized correctly. It means the application has a way to provide relevant context instead of relying only on the model’s general recognition behavior.

Pricing and technical limits

Meta’s listed price is $3.00 per 1,000 minutes, equivalent to $0.18 per hour. This is usage-based pricing for transcription time, not a recurring monthly subscription. For example, 10,000 minutes of processed audio would cost $30 at the stated rate before any additional service charges or account-specific conditions.

SpecificationVerified information
ProviderMeta
Model IDmuse-voice-transcribe-1.0
Primary functionSpeech-to-text transcription
InputAudio, including realtime streams and recorded files
OutputText transcript
Pricing$3.00 per 1,000 minutes, or $0.18 per hour
Realtime accessWebSocket endpoint
Recorded audio accessFile-transcription endpoint
Languages25 evaluated languages, with code-switching described by Meta
Speaker supportStreaming diarization for 20 or more speakers is documented
Context lengthNot publicly specified in the supplied research
Maximum output tokensNot publicly specified in the supplied research

Because no context window or maximum output-token limit is supplied, those values should not be inferred from general Meta API limits or from other models. For long recordings, application developers should follow the current transcription documentation regarding connection duration, file size, audio format, rate limits, and segmentation.

Main strengths and trade-offs

Muse Voice Transcribe’s central strength is specialization. It combines streaming access with features that matter in real conversations: turn detection, voice activity detection, speaker attribution, contextual biasing, and timestamps. A general-purpose model that happens to accept audio may not expose the same transcription-oriented controls or may be less suitable for predictable speech-processing pipelines.

Its price also favors high-volume transcription. At the documented rate, the cost is easy to estimate from audio duration, which is useful for captioning, call archives, and meeting records. The editorial speed and cost assessments supplied for this model are both high relative to broader models, reflecting its narrow purpose and usage-based price. These are comparative editorial estimates, not provider-published benchmark scores.

The trade-off is limited scope. Muse does not synthesize speech, conduct a speech-to-speech conversation, generate images or video, execute tools, or serve as a general reasoning and coding model. The supplied editorial reasoning and coding scores are low because transcription is its intended task, not because the model is meant to compete with general language models on those abilities. Developers needing summarization, question answering, workflow execution, or code generation will need another component after transcription.

Best use cases

  • Live captions: convert spoken audio into text while a meeting, broadcast, presentation, or event is taking place.
  • Voice agents: provide the speech-recognition layer for an application that listens to a user before passing the transcript to a conversational model.
  • Meeting transcription: create speaker-aware records for meetings, interviews, panels, or group discussions.
  • Call intelligence: transcribe customer-support or sales calls for search, quality review, compliance workflows, or later analysis.
  • Dictation: turn spoken notes into text when low-latency recognition is more important than broader language-model features.
  • Domain-specific speech: use contextual or keyword biasing for names, product terms, and specialized vocabulary.

In each case, the model should be viewed as the transcription layer. Summaries, sentiment analysis, action-item extraction, database updates, and automated replies would require additional application logic or other AI services.

When to choose Muse Voice Transcribe 1.0

Choose Muse when the main requirement is fast, structured speech-to-text and the application benefits from live streaming, speaker diarization, endpointing, or voice activity detection. It is particularly suitable when predictable per-minute pricing and a focused transcription API matter more than broad generative features.

Another audio-capable model or service may be more appropriate if the application needs speech synthesis, direct speech-to-speech interaction, emotion or sound-event analysis, or a single model that can reason over audio and perform downstream tasks. A general language model may also be preferable when transcription is only a small part of a workflow and the primary requirement is complex reasoning, summarization, coding, or tool use.

Muse is also not the right choice when guaranteed word-level timestamps, a specific unsupported language, or a documented context and file-size limit is a hard requirement that the current documentation does not confirm. In those cases, verify the live Meta API specifications or compare with a transcription provider that publishes the needed limits.

Bottom line

Muse Voice Transcribe 1.0 is a focused Meta model for turning live or recorded speech into text. Its distinguishing capabilities are realtime WebSocket transcription, speaker-aware output, endpointing, voice activity detection, contextual biasing, code-switching, and support for 25 evaluated languages. At $3 per 1,000 minutes, it is aimed at applications that need scalable transcription rather than a general-purpose AI assistant. Its narrow output and lack of documented context or maximum-output limits mean that it should be selected as one component of a voice or audio pipeline, not as a replacement for a broader reasoning model.


Answers to Frequently Asked Questions

What is Muse Voice Transcribe 1.0?
Muse Voice Transcribe 1.0 is Meta’s specialized automatic speech recognition model for converting live or recorded audio into written transcripts. Its canonical API model ID is muse-voice-transcribe-1.0.
How much does Muse Voice Transcribe 1.0 cost?
Meta lists pricing at $3.00 per 1,000 minutes of audio, which is equivalent to $0.18 per hour. For example, processing 10,000 minutes would cost $30 before any additional charges or account-specific conditions.
Is Muse Voice Transcribe 1.0 a general-purpose conversational AI model?
No. Muse Voice Transcribe 1.0 is focused on speech-to-text transcription and does not generate speech, provide speech-to-speech interaction, answer questions, write code, or perform downstream tasks such as summarization without additional application logic or another AI model.
What are the best use cases for Muse Voice Transcribe 1.0?
Common use cases include live captions, voice agents, meeting and interview transcription, call intelligence, dictation, and transcription of domain-specific speech using contextual or keyword biasing.
What features does Muse Voice Transcribe 1.0 support?
The model supports realtime WebSocket transcription, file transcription, speaker diarization, endpointing, voice activity detection, contextual and keyword biasing, turn-level timestamps, and code-switching across 25 evaluated languages.


Sources 8
Provider

About Meta AI