What is ERNIE-iRAG-Edit?
ERNIE-iRAG-Edit is a specialized image-editing model provided by Baidu. Baidu’s model catalog identifies the canonical model version as ERNIE-iRAG-Edit-1.0, while the documented Qianfan endpoint is ernie-irag-edit. The model focuses on changing an existing image or a selected part of one, rather than generating a completely new image from a text prompt alone.
Its documented editing operations are erase, repaint, and variation. In practical terms, this makes it suitable for removing unwanted objects, replacing selected regions, or producing altered versions of a reference image while retaining its general visual content.
ERNIE-iRAG-Edit is separate from Baidu’s ERNIE-iRAG-1.0 image-generation model. That distinction matters: ERNIE-iRAG-Edit is positioned around modifying source images, while the related generation model is documented for creating images.
Core image-editing capabilities
Erase unwanted content
The erase operation is intended to remove a selected object or area from an image. For example, a workflow could remove an unwanted item from a product photograph, clear a distracting background element, or eliminate part of a scene before further editing.
Because the operation is region-focused, developers can identify the area to process with a mask. A mask is an additional image or data representation that indicates which part of the original image should be edited. The supplied Qianfan examples show image-editing requests containing image, mask, and feature parameters.
Repaint selected regions
Repaint is designed to redraw or replace a selected region. This can be useful when an object, surface, or visual detail needs to be changed without recreating the entire source image. The exact result depends on the source image, the selected area, and the request parameters used by the application.
Create image variations
The variation operation generates an altered version of a reference image. This supports workflows such as exploring alternative visual treatments, producing multiple versions of a creative asset, or making controlled changes while preserving the source image’s broader composition.
These capabilities make the model more comparable to an image-inpainting or image-variation tool than to a general-purpose multimodal assistant. The supplied documentation does not describe ERNIE-iRAG-Edit as a conversational model that explains its edits in text.
Where to access ERNIE-iRAG-Edit
Baidu lists ERNIE-iRAG-Edit in its Qianfan inference ecosystem. The supplied documentation specifically identifies the model as available for batch inference, with the endpoint identifier ernie-irag-edit. Batch inference is useful when an application needs to submit many image-editing jobs rather than waiting for a person to edit each image manually.
Documented image-generation and editing workflows return results asynchronously. Applications should therefore be designed around task submission and result retrieval or polling, rather than assuming that an edited image is returned immediately in the same response. Image and mask formatting, object-storage handling, link expiration, quotas, and account-level concurrency limits also need to be considered during implementation.
Availability may depend on the Qianfan product, account permissions, region, and current service configuration. Baidu also announced that the related AI作画-iRAG版 API would stop accepting new sales and renewals from April 30, 2026. That announcement concerns the named AI作画-iRAG版 product; it does not explicitly establish that every Qianfan endpoint using ERNIE-iRAG-Edit has been discontinued. Current access and commercial terms should therefore be checked in the Qianfan console.
Supported inputs and outputs
| Area | Verified information |
|---|---|
| Primary input | Images, with masks used for region-based editing in documented workflows |
| Primary output | Edited or varied images |
| Audio input | Not documented |
| Video input | Not documented |
| Text output | Not a documented model output; the model is classified as image editing |
| Batch processing | Listed in the Qianfan documentation |
The model’s multimodal behavior is specifically visual: it accepts image-related inputs and produces image results. The supplied research does not verify audio or video editing, speech generation, text-to-video generation, tool calling, function calling, or structured JSON output for this endpoint.
Technical limits and pricing
Baidu’s public documentation supplied for this model does not provide a token context window, maximum output-token limit, knowledge cutoff, or token-based input and output prices. Those omissions are significant for evaluation: ERNIE-iRAG-Edit should not be assessed using language-model limits that Baidu has not published for this image endpoint.
No verified price amount is available in the supplied research. Exact commercial pricing, quotas, concurrency limits, and any image-specific billing rules should be confirmed in the current Qianfan console before deployment. The absence of a listed price here does not mean that the endpoint is free.
Image links or generated results may expire, and asynchronous processing means that storage and retrieval behavior should be included in the application design. Developers should also verify the currently accepted image and mask formats, size restrictions, regional availability, and content-moderation requirements in the live documentation.
Strengths and trade-offs
The model’s main strength is specialization. Instead of asking a general image generator to recreate a complete scene, a developer can use an editing-oriented endpoint for targeted changes to an existing image. Erasure, repainting, and variation cover common production tasks while preserving the source image as the starting point.
Its Qianfan availability and batch-inference support are also relevant for organizations building automated image pipelines. Product catalogs, marketing operations, content moderation workflows, and user-generated-image services may benefit from submitting many consistent editing jobs through an inference service.
The trade-off is limited scope. ERNIE-iRAG-Edit is not documented as a general-purpose language model, reasoning model, coding model, embedding model, speech model, or video model. It also does not have publicly documented token limits or token-based pricing in the supplied sources. A general image-generation option may be more appropriate when there is no source image to edit, while a conversational multimodal assistant may be better for workflows requiring explanations, document analysis, or interactive instruction.
Editorial data associated with the model gives it a low reasoning score and low coding score because those are not its intended functions, while its speed and cost scores are recorded as moderate. These are editorial evaluations, not Baidu-published benchmark results or official performance guarantees.
When to choose ERNIE-iRAG-Edit
Choose ERNIE-iRAG-Edit when the starting point is an image and the desired result is a targeted modification. Suitable examples include:
- Removing unwanted objects from product, marketing, or user-generated images.
- Replacing or repainting a selected region using a mask.
- Generating variations of a reference image for creative exploration.
- Automating repeated image-editing tasks through Qianfan batch inference.
- Building a workflow in which edited image results are retrieved asynchronously and stored for later use.
Another option may be more appropriate when the task begins with text and requires a wholly new image, when the application needs image-and-text conversation, or when audio, video, coding, reasoning, embeddings, or tool execution are central requirements. ERNIE-iRAG-Edit should be selected for its editing workflow, not treated as a general Baidu AI assistant.
Implementation checklist
Before using the model in production, confirm that the intended Qianfan account can access the endpoint and batch-inference service. Prepare a reliable way to provide the source image and, for regional editing, a correctly aligned mask. The application should support asynchronous task handling, including submission, polling or retrieval, failure handling, and result storage.
Teams should also verify current image and mask requirements, quotas, concurrency limits, content policies, link-expiration behavior, and pricing. These operational details can affect the total cost and reliability of a batch pipeline even when the editing operation itself is straightforward.
Finally, do not infer a service guarantee from the model’s catalog listing alone. Baidu’s related AI作画-iRAG版 lifecycle announcement shows that product availability can change, while the supplied sources do not explicitly equate that notice with shutdown of the Qianfan ERNIE-iRAG-Edit endpoint.
Bottom line
ERNIE-iRAG-Edit is a focused Baidu image-editing model for erasing, repainting, and creating variations from reference images. Its strongest fit is a visual pipeline that already has source images and needs repeatable, region-aware modifications. Qianfan batch inference makes it relevant to automated workflows, but developers must validate current access, asynchronous result handling, image and mask requirements, quotas, and pricing. It is not documented as a general conversational, coding, reasoning, audio, or video model.

