What is Muse Image 1.0?
Muse Image 1.0 is an image-generation model from Meta, available through the Meta Model API under the model ID muse-image-1.0. It belongs to Meta’s Muse family and is focused on visual creation rather than general-purpose text completion. Meta also associates Muse-family systems with some of its current AI experiences, but this page concerns the dedicated Muse Image 1.0 model.
The model supports text-to-image generation, image editing, multi-image composition, and multi-turn refinement. In practical terms, a user can describe a new image, provide one or more reference images, ask for changes to an existing image, or request that several visual sources be combined into a coherent result.
Its agentic design is the main distinction from a simple prompt-to-image system. According to Meta’s documentation, Muse Image can plan a task, use image search and web search, execute shell or code operations when useful, and apply self-refinement before producing the result. These tools are intended to help with complex visual tasks, although they do not turn the model into a general-purpose text or software-development model.
Core capabilities and practical workflows
Text-to-image generation
Muse Image can create an image from a natural-language description. Prompts can specify the subject, environment, style, composition, lighting, or intended use. The model is therefore suitable for tasks such as concept art, marketing visuals, social content, product scenes, and other creative assets.
Meta documents several output formats, including WEBP, PNG, and JPEG. A request can produce between one and 10 images. The API’s image-size setting controls aspect ratio rather than guaranteeing a precise returned pixel dimension, and Meta states that output quality is supported up to 1600 pixels. Users should therefore treat the size parameter as a composition and format control, not as a promise of a particular exact resolution.
Editing and reference-image workflows
The model accepts optional image inputs as well as text. This enables edits such as changing an object, adjusting a scene, transforming a style, or preserving selected visual characteristics while modifying the rest of the image. Reference images can also anchor a composition when the result needs to contain several subjects, products, or design elements.
Multi-image composition is useful when a prompt alone is not enough to describe the desired result. For example, a workflow might combine a product reference, a background reference, and a person or character reference into one new image. The supplied research does not specify guaranteed identity preservation, pixel-level editing precision, or exact limits on the number of input reference images, so those outcomes should not be assumed.
Planning, search, code, and refinement
Muse Image’s tool support is intended to help it solve visual tasks in stages. Planning can break down a complex request before generation. Image search and web search can provide relevant visual or factual context when enabled. Shell or code execution may help with operations that are easier to perform procedurally. The model can also refine a generation before returning it.
These features are best understood as workflow capabilities, not as a guarantee that every request will invoke every tool. Tool use is supported, but the supplied specifications do not define a fixed execution sequence, tool latency, or a guaranteed number of refinement passes. Web and image search can also introduce grounding, licensing, or source-review considerations for production workflows.
Inputs, outputs, and technical limits
| Specification | Verified detail |
|---|---|
| Text input | Supported |
| Image input | Supported, including reference images |
| Image output | Supported |
| Audio or video input | Not supported in the supplied model specification |
| Text output | Not the model’s output type |
| Output formats | WEBP, PNG, and JPEG are documented |
| Images per request | Between 1 and 10 |
| Quality and size | Meta states support up to 1600 pixels; the size control specifies aspect ratio rather than an exact returned dimension |
| Context length | Not published in the supplied research |
| Maximum output tokens | Not applicable or published for this image-output model |
Muse Image is multimodal in the specific sense that it can receive text and images and return images. It is not documented here as an audio, video, speech, embedding, or text-completion model. There is also no verified JSON-mode or structured-text-output specification in the supplied material.
Reasoning and coding capabilities
Meta’s documentation describes planning, tool selection, code execution, and self-refinement, so the model can perform intermediate reasoning for visual tasks. This is different from offering a conventional text reasoning model with a visible chain of thought or a token-based answer. The model’s reasoning is mainly useful when deciding how to construct, edit, ground, or improve an image.
The supplied editorial evaluation gives Muse Image a reasoning score of 8 out of 10 and a coding score of 2 out of 10. These are comparative editorial estimates, not Meta-published benchmark results. The relatively low coding assessment reflects the model’s purpose: it may run code as part of an image workflow, but it is not intended for general software development, code completion, or programming conversations.
Pricing and cost trade-offs
Meta lists a flat price of $0.01 per successfully generated image through the Meta Model API. The price is not based on prompt length, reasoning strength, or whether supported tools are enabled. Failed or safety-filtered outputs are not counted according to the supplied documentation.
Because the charge applies per returned image, requesting multiple alternatives increases the total cost. A request for 10 images would therefore create up to 10 separately priced successful outputs. The flat image-based price is easier to estimate than token billing, but it does not by itself describe total workflow cost: search, retries, application infrastructure, storage, and downstream editing may still affect a project’s budget.
The supplied editorial assessment rates Muse Image’s cost at 10 out of 10 and speed at 7 out of 10. These are subjective comparative ratings rather than provider specifications. Agentic planning, search, code execution, and refinement may make complex requests more involved than a minimal direct-generation workflow, so users should test latency on their own prompts rather than infer a fixed response time.
Main strengths and limitations
Strengths
- One model for several visual workflows: It covers generation, editing, composition, reference-image use, and iterative refinement.
- Agentic task handling: Planning, optional search, code execution, and self-refinement can help with requests that require more than direct prompt interpretation.
- Multiple output formats: WEBP, PNG, and JPEG support common web, design, and content-production workflows.
- Predictable per-image pricing: The advertised $0.01 charge applies to each successfully generated image rather than to prompt or output token counts.
- Batch alternatives: Up to 10 images can be requested in one operation, which is useful for ideation and variation testing.
Limitations
- Image-focused scope: It is not the appropriate choice for text-only chat, speech, audio, video, embeddings, or general coding.
- Unpublished token and context limits: Meta does not provide a conventional context-window or maximum-output-token value in the supplied research.
- Exact dimensions are not guaranteed: The documented size control concerns aspect ratio, while output quality is stated only up to 1600 pixels.
- Tool behavior is not fully specified: The documentation supports tools but does not guarantee when they will run, how long they will take, or how many refinement stages will occur.
- Search-grounded results need review: When web or image search is used, users should review sources, rights, and factual or visual suitability before publication.
When to choose Muse Image 1.0
Choose Muse Image 1.0 when the primary deliverable is an image and the task benefits from references, composition, editing, or iterative improvement. It is a strong fit for product mockups, campaign concepts, visual variations, anchored character or object series, social-media assets, and images that need information or inspiration gathered through search.
Its flat per-successful-image price is also useful when a team wants a straightforward way to estimate generation costs. The model is particularly attractive when one workflow needs both initial creation and subsequent edits instead of separate tools for generation and image manipulation.
A different type of model may be more appropriate for a text response, software development, speech, audio, or video. A simpler image generator may be preferable when the task is a quick uncomplicated illustration and agentic planning or search would add unnecessary complexity. Conversely, a specialized image-editing system may be a better fit when exact pixel-level control, formally defined input limits, or highly predictable dimensions are more important than broad agentic capabilities. The supplied research does not identify a specific competing model with which to make a verified feature-by-feature comparison.
Availability and bottom line
Muse Image 1.0 is currently available through Meta Model API under the model ID muse-image-1.0. Meta also lists it in selected consumer experiences, but the documented developer access is the clearest route for controlled generation and integration.
The model’s defining proposition is not simply that it creates images. It combines image generation with editing, composition, reference inputs, optional search, code-assisted operations, planning, and self-refinement. That combination makes it most compelling for visual tasks with several constraints or stages. Its boundaries are equally important: it is image-output focused, its context and exact dimension limits are not fully published in the supplied material, and its editorial coding and speed assessments should not be mistaken for provider benchmarks.

