Muse Image

Muse Image 1.0

by Meta AI · Current and available through Meta Model API; also available in selected Meta AI consumer experiences

Meta Muse Image 1.0 creates, edits, and composes images from text prompts and reference images. Its agentic workflow can plan tasks, use optional web or image search, run code, and refine results. The model returns image files in formats including WEBP, PNG, and JPEG, supports 1 to 10 outputs per request, and costs $0.01 per successfully generated image through Meta Model API. Context length and exact pixel limits are not fully published.

Image generation Reasoning Coding
Muse Image 1.0 is Meta’s dedicated image-generation model for creating, editing, and combining visual content. Unlike a text-only model, it accepts text prompts and optional reference images, then returns image files. Meta positions it as an agentic image system: rather than simply mapping a prompt directly to pixels, it can plan a visual task, use search or code tools when appropriate, and perform iterative refinement. That makes it particularly relevant for product imagery, creative assets, image variations, and compositions that need to preserve or combine visual references.
Outputs

What Muse Image 1.0 can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Tool use Web search Multimodal output
Model profile

Performance characteristics

8/10 Reasoning
2/10 Coding
7/10 Speed
10/10 Cost efficiency
Specifications

Technical details

Model family Muse Image
Model type Multimodal
Context window tokens
Maximum output tokens
Release date 2026-07-07
Status Current and available through Meta Model API; also available in selected Meta AI consumer experiences
Knowledge cutoff notes

Meta does not publish a conventional fixed knowledge-cutoff date for Muse Image 1.0. The model can optionally use web search and image search during generation, but those tools do not establish or change an underlying model knowledge cutoff.

Model notes

The canonical API model ID is muse-image-1.0, while Meta's model pages generally display the family name Muse Image. It accepts text prompts and optional reference images, and returns images. Supported workflows include generation, editing, composition, and multi-turn refinement. Meta documents built-in image search, web search, shell/code execution, planning, and self-refinement capabilities. The image-generation API supports output formats including WEBP, PNG, and JPEG, and allows between 1 and 10 images per request. The documented image size parameter controls aspect ratio rather than guaranteeing exact returned pixel dimensions; Meta states output quality is supported up to 1600px. Pricing is flat per successfully returned image, regardless of prompt length, reasoning strength, or enabled tools. Failed or safety-filtered outputs are not counted. The editorial scores are comparative estimates rather than official benchmark scores.

Cost

Model pricing

Output $0.01 per successfully generated image
Model guide

Muse Image 1.0: Meta’s Agentic Model for Image Generation, Editing, and Composition

Meta’s Muse Image 1.0 is an image-output model designed for text-to-image generation, image editing, reference-based composition, and iterative visual refinement. Its distinctive feature is an agentic workflow: it can plan a task, use optional web and image search, run code when useful, and refine an image before returning it. Through Meta Model API, it costs $0.01 per successfully generated image.

What is Muse Image 1.0?

Muse Image 1.0 is an image-generation model from Meta, available through the Meta Model API under the model ID muse-image-1.0. It belongs to Meta’s Muse family and is focused on visual creation rather than general-purpose text completion. Meta also associates Muse-family systems with some of its current AI experiences, but this page concerns the dedicated Muse Image 1.0 model.

The model supports text-to-image generation, image editing, multi-image composition, and multi-turn refinement. In practical terms, a user can describe a new image, provide one or more reference images, ask for changes to an existing image, or request that several visual sources be combined into a coherent result.

Its agentic design is the main distinction from a simple prompt-to-image system. According to Meta’s documentation, Muse Image can plan a task, use image search and web search, execute shell or code operations when useful, and apply self-refinement before producing the result. These tools are intended to help with complex visual tasks, although they do not turn the model into a general-purpose text or software-development model.

Core capabilities and practical workflows

Text-to-image generation

Muse Image can create an image from a natural-language description. Prompts can specify the subject, environment, style, composition, lighting, or intended use. The model is therefore suitable for tasks such as concept art, marketing visuals, social content, product scenes, and other creative assets.

Meta documents several output formats, including WEBP, PNG, and JPEG. A request can produce between one and 10 images. The API’s image-size setting controls aspect ratio rather than guaranteeing a precise returned pixel dimension, and Meta states that output quality is supported up to 1600 pixels. Users should therefore treat the size parameter as a composition and format control, not as a promise of a particular exact resolution.

Editing and reference-image workflows

The model accepts optional image inputs as well as text. This enables edits such as changing an object, adjusting a scene, transforming a style, or preserving selected visual characteristics while modifying the rest of the image. Reference images can also anchor a composition when the result needs to contain several subjects, products, or design elements.

Multi-image composition is useful when a prompt alone is not enough to describe the desired result. For example, a workflow might combine a product reference, a background reference, and a person or character reference into one new image. The supplied research does not specify guaranteed identity preservation, pixel-level editing precision, or exact limits on the number of input reference images, so those outcomes should not be assumed.

Planning, search, code, and refinement

Muse Image’s tool support is intended to help it solve visual tasks in stages. Planning can break down a complex request before generation. Image search and web search can provide relevant visual or factual context when enabled. Shell or code execution may help with operations that are easier to perform procedurally. The model can also refine a generation before returning it.

These features are best understood as workflow capabilities, not as a guarantee that every request will invoke every tool. Tool use is supported, but the supplied specifications do not define a fixed execution sequence, tool latency, or a guaranteed number of refinement passes. Web and image search can also introduce grounding, licensing, or source-review considerations for production workflows.

Inputs, outputs, and technical limits

SpecificationVerified detail
Text inputSupported
Image inputSupported, including reference images
Image outputSupported
Audio or video inputNot supported in the supplied model specification
Text outputNot the model’s output type
Output formatsWEBP, PNG, and JPEG are documented
Images per requestBetween 1 and 10
Quality and sizeMeta states support up to 1600 pixels; the size control specifies aspect ratio rather than an exact returned dimension
Context lengthNot published in the supplied research
Maximum output tokensNot applicable or published for this image-output model

Muse Image is multimodal in the specific sense that it can receive text and images and return images. It is not documented here as an audio, video, speech, embedding, or text-completion model. There is also no verified JSON-mode or structured-text-output specification in the supplied material.

Reasoning and coding capabilities

Meta’s documentation describes planning, tool selection, code execution, and self-refinement, so the model can perform intermediate reasoning for visual tasks. This is different from offering a conventional text reasoning model with a visible chain of thought or a token-based answer. The model’s reasoning is mainly useful when deciding how to construct, edit, ground, or improve an image.

The supplied editorial evaluation gives Muse Image a reasoning score of 8 out of 10 and a coding score of 2 out of 10. These are comparative editorial estimates, not Meta-published benchmark results. The relatively low coding assessment reflects the model’s purpose: it may run code as part of an image workflow, but it is not intended for general software development, code completion, or programming conversations.

Pricing and cost trade-offs

Meta lists a flat price of $0.01 per successfully generated image through the Meta Model API. The price is not based on prompt length, reasoning strength, or whether supported tools are enabled. Failed or safety-filtered outputs are not counted according to the supplied documentation.

Because the charge applies per returned image, requesting multiple alternatives increases the total cost. A request for 10 images would therefore create up to 10 separately priced successful outputs. The flat image-based price is easier to estimate than token billing, but it does not by itself describe total workflow cost: search, retries, application infrastructure, storage, and downstream editing may still affect a project’s budget.

The supplied editorial assessment rates Muse Image’s cost at 10 out of 10 and speed at 7 out of 10. These are subjective comparative ratings rather than provider specifications. Agentic planning, search, code execution, and refinement may make complex requests more involved than a minimal direct-generation workflow, so users should test latency on their own prompts rather than infer a fixed response time.

Main strengths and limitations

Strengths

  • One model for several visual workflows: It covers generation, editing, composition, reference-image use, and iterative refinement.
  • Agentic task handling: Planning, optional search, code execution, and self-refinement can help with requests that require more than direct prompt interpretation.
  • Multiple output formats: WEBP, PNG, and JPEG support common web, design, and content-production workflows.
  • Predictable per-image pricing: The advertised $0.01 charge applies to each successfully generated image rather than to prompt or output token counts.
  • Batch alternatives: Up to 10 images can be requested in one operation, which is useful for ideation and variation testing.

Limitations

  • Image-focused scope: It is not the appropriate choice for text-only chat, speech, audio, video, embeddings, or general coding.
  • Unpublished token and context limits: Meta does not provide a conventional context-window or maximum-output-token value in the supplied research.
  • Exact dimensions are not guaranteed: The documented size control concerns aspect ratio, while output quality is stated only up to 1600 pixels.
  • Tool behavior is not fully specified: The documentation supports tools but does not guarantee when they will run, how long they will take, or how many refinement stages will occur.
  • Search-grounded results need review: When web or image search is used, users should review sources, rights, and factual or visual suitability before publication.

When to choose Muse Image 1.0

Choose Muse Image 1.0 when the primary deliverable is an image and the task benefits from references, composition, editing, or iterative improvement. It is a strong fit for product mockups, campaign concepts, visual variations, anchored character or object series, social-media assets, and images that need information or inspiration gathered through search.

Its flat per-successful-image price is also useful when a team wants a straightforward way to estimate generation costs. The model is particularly attractive when one workflow needs both initial creation and subsequent edits instead of separate tools for generation and image manipulation.

A different type of model may be more appropriate for a text response, software development, speech, audio, or video. A simpler image generator may be preferable when the task is a quick uncomplicated illustration and agentic planning or search would add unnecessary complexity. Conversely, a specialized image-editing system may be a better fit when exact pixel-level control, formally defined input limits, or highly predictable dimensions are more important than broad agentic capabilities. The supplied research does not identify a specific competing model with which to make a verified feature-by-feature comparison.

Availability and bottom line

Muse Image 1.0 is currently available through Meta Model API under the model ID muse-image-1.0. Meta also lists it in selected consumer experiences, but the documented developer access is the clearest route for controlled generation and integration.

The model’s defining proposition is not simply that it creates images. It combines image generation with editing, composition, reference inputs, optional search, code-assisted operations, planning, and self-refinement. That combination makes it most compelling for visual tasks with several constraints or stages. Its boundaries are equally important: it is image-output focused, its context and exact dimension limits are not fully published in the supplied material, and its editorial coding and speed assessments should not be mistaken for provider benchmarks.


Answers to Frequently Asked Questions

What is Muse Image 1.0?
Muse Image 1.0 is Meta’s image-generation model available through the Meta Model API under the model ID "muse-image-1.0". It supports text-to-image generation, image editing, multi-image composition, reference images, and multi-turn refinement.
Does Muse Image 1.0 support search, code execution, and image refinement?
Yes. Meta describes Muse Image 1.0 as an agentic model that can plan tasks, use image and web search, execute shell or code operations when useful, and apply self-refinement. However, the documentation does not guarantee which tools will run, how long they will take, or how many refinement stages will occur.
What are the main limitations of Muse Image 1.0?
Muse Image 1.0 is focused on image output rather than text chat, general software development, audio, speech, video, or embeddings. Its exact context length and maximum output-token limit are not published, and the size setting controls aspect ratio rather than guaranteeing an exact pixel dimension. Meta states that output quality is supported up to 1600 pixels.
How much does Muse Image 1.0 cost?
Meta lists a price of $0.01 per successfully generated image through the Meta Model API. Requests for multiple images are charged per successful output, so a request for 10 images can cost up to $0.10. Failed or safety-filtered outputs are not counted according to the supplied documentation.
What can Muse Image 1.0 be used for?
Muse Image 1.0 can create images from text prompts, edit existing images, combine multiple visual references, generate variations, and refine results iteratively. It is suitable for product mockups, marketing visuals, concept art, social-media assets, and other image-focused workflows.


Sources 6
Provider

About Meta AI