What is Muse Spark 1.3?
Muse Spark 1.3 is a proprietary multimodal reasoning model provided by Meta. Its main role is not image or video creation; it is a text-output model designed to understand several kinds of input and then reason, write, plan, code, or operate tools over multiple steps.
Meta released Muse Spark 1.3 on September 2, 2026. It is positioned for long-horizon work, meaning tasks that require the model to preserve requirements, inspect intermediate results, resolve conflicting information, and continue through a sequence of actions rather than answer a single short question. The model is available through the Meta Model API and Meta’s Muse Code product, and Meta documents compatibility with OpenAI-compatible integrations and agent frameworks through its API endpoint.
Within Meta’s current catalog, Muse Spark 1.3 is a Muse-family model focused on reasoning, coding, and agentic execution. It should not be confused with a general media-generation model: its outputs are text, even though its inputs can include several media types.
Key specifications at a glance
| Specification | Details |
|---|---|
| Provider | Meta |
| Release date | September 2, 2026 |
| Model ID | muse-spark-1.3 |
| Context window | 1,048,576 tokens |
| Maximum output | 131,072 tokens |
| Input | Text, images, video, PDFs, and audio with limited support |
| Output | Text only |
| Tool support | Tool calling and first-party web-search grounding |
| Other controls | Streaming, structured output, caching, and reasoning-effort controls |
| Availability | Meta Model API and Muse Code |
These are the documented specifications supplied for the model. Editorial assessments of reasoning, coding, speed, and cost are comparative estimates rather than ratings published by Meta.
Long-context and multimodal understanding
The model’s 1,048,576-token context window is its most distinctive technical characteristic. A context window is the amount of material the model can consider within one request or continuing interaction. At this size, Muse Spark 1.3 can be used with very large software repositories, lengthy technical documentation, extensive project histories, and multi-stage task instructions without requiring the user to divide everything into small independent prompts.
Muse Spark 1.3 accepts text, images, video, PDFs, and audio. This allows a workflow to combine written requirements with screenshots, design references, documents, demonstrations, or other visual material. However, the audio capability needs an important qualification: Meta states that audio understanding is not fully supported in version 1.3 and may produce degraded results. Audio-focused applications are better served by Muse Spark 1.2 or Muse Voice Transcribe, according to the supplied documentation.
Multimodal input does not mean multimodal generation. Muse Spark 1.3 does not natively produce images, video, audio, speech, or music. Its response channel is text, which can include code, plans, structured data, tool arguments, or explanations.
Coding and agentic workflows
Muse Spark 1.3 is primarily aimed at software engineering assistants and agents that can carry out work through repeated reasoning and tool calls. A coding agent might inspect a repository, identify the relevant files, propose a change, run a tool, review the result, and revise its approach. The model is designed for this type of loop rather than only for generating an isolated code snippet.
Meta says that, compared with Muse Spark 1.2, version 1.3 uses approximately 20% fewer tool calls and 25% fewer tokens in its internal coding comparisons while producing cleaner and less verbose results. These figures are provider claims about internal comparisons, not an independently verified benchmark result, so they should be treated as directional evidence rather than a guarantee for every project.
Useful coding scenarios include navigating large repositories, maintaining requirements across a long implementation, reviewing changes, investigating errors, writing tests, and coordinating browser or computer-use actions. The model’s large context can reduce the need to repeatedly restate project details, while tool calling lets an application connect it to external actions or information sources.
Reasoning, tools, and output controls
The model supports reasoning-effort controls, including a maximum effort level on the Standard tier. In practical terms, these controls allow an application to trade response speed and token usage against more deliberate reasoning. Higher effort may be useful for complex planning, debugging, or tasks with many constraints, while lower effort can be preferable for routine transformations and latency-sensitive interactions.
Muse Spark 1.3 also supports tool calling. Tool calling allows the model to request a defined application function, such as searching a repository, retrieving a document, checking a system, or taking an action in an agent environment. The model does not independently gain unrestricted access to those systems; the surrounding application decides which tools exist, validates arguments, and executes approved calls.
First-party web-search grounding is available through the Meta Model API. This can help applications obtain current external information during a request. Web search does not change the model’s underlying knowledge cutoff, and Meta’s public documentation does not specify a knowledge-cutoff date for Muse Spark 1.3.
Structured output is supported, which is useful when a program needs responses that follow a defined schema rather than free-form prose. Streaming is also supported, allowing partial text to be delivered while a response is being generated. Caching is documented as supported, and cached input is priced separately from ordinary input.
Pricing and deployment options
The standard Meta Model API pricing supplied for Muse Spark 1.3 is:
- Input: $1.25 per million tokens
- Cached input: $0.15 per million tokens
- Output: $4.25 per million tokens
The input and output prices apply to different parts of usage. Long prompts, repository contents, documents, and other material sent to the model contribute to input usage; generated explanations, code, structured responses, and tool arguments contribute to output usage. The maximum documented output limit is 131,072 tokens, although typical applications may request far less.
Meta also offers a separate muse-spark-1.3-contributor variant at substantially lower rates. It is not simply a cheaper billing tier for the same privacy terms: Meta may use prompts and completions from the Contributor variant to train future models. It should therefore be evaluated as a distinct deployment option. The standard model is the more appropriate reference point when comparing ordinary API pricing and data-use expectations.
Web-search grounding adds a separate charge per search query. The supplied research does not specify that additional amount, so it should not be estimated here.
Main strengths and trade-offs
Muse Spark 1.3’s strongest advantage is the combination of a very large context window and support for extended tool-driven work. It can keep more repository, documentation, or task history in view than a model with a smaller context limit. That is especially relevant when a coding task spans many files or requires repeated inspection and correction.
Its other strengths are the combination of multimodal input, structured output, web grounding, streaming, caching, and adjustable reasoning effort. Together, these features make it suitable for applications that need more than a chat response, such as coding agents, document-analysis systems, browser workflows, and software that passes model output into downstream programs.
The trade-off is that high-capability, long-running tasks can consume substantial tokens and may involve multiple tool calls. Although Meta reports improved efficiency over Muse Spark 1.2, a smaller or faster model may still be preferable for simple classification, short rewriting, routine extraction, or high-volume low-complexity requests. Editorially, the supplied research rates Muse Spark 1.3 highly for reasoning and coding, with a comparatively strong cost assessment, but these are comparative editorial scores and not provider-published measurements.
Limitations and when to use another option
The clearest limitation is incomplete audio understanding. Applications centered on reliable transcription or detailed audio interpretation should use Muse Voice Transcribe or, where appropriate, Muse Spark 1.2 as the more suitable alternative identified by Meta’s documentation.
The model is also unsuitable when the application needs native image, video, audio, speech, or music generation. Its multimodal capability is on the input side; its output remains text. A separate generation model or media pipeline would be required for those tasks.
Other gaps remain in the public documentation. Meta does not provide a knowledge-cutoff date for this exact model, and the supplied research does not verify fine-tuning support or a batch API specification. Teams that require one of those features should confirm current documentation before committing to the model.
Choose Muse Spark 1.3 when the task involves large amounts of context, complex coding, multiple reasoning steps, tool orchestration, multimodal document understanding, or current information obtained through web search. Consider another option when the priority is robust audio processing, native media generation, the lowest possible latency for simple work, or a documented fine-tuning or batch interface.
Bottom line
Muse Spark 1.3 is best understood as Meta’s long-context text-generation model for coding agents and other extended workflows. Its one-million-token context window, 131,072-token output ceiling, multimodal inputs, tool calling, structured output, web grounding, and reasoning controls give it a broad technical range. The practical decision is less about whether it supports many features and more about whether the workload benefits from long-running, tool-assisted reasoning. For those workloads it is a strong fit; for audio-centric or media-generation applications, its limitations are decisive.

