GLM-Image

GLM-Image

by Z.ai · Current; available through the Z.ai API and as downloadable open weights

GLM-Image is Z.ai’s hybrid image-generation model for text-to-image creation and image-to-image editing. Its 9B autoregressive generator and 7B diffusion decoder target readable text, posters, diagrams, presentation graphics, and other structured visual layouts. It is available through a $0.015-per-image API and as downloadable open weights, although local inference requires substantial hardware.

Image generation
GLM-Image combines a 9B-parameter autoregressive generator with a 7B diffusion decoder to create images from prompts and reference images. The model is designed to understand global composition while preserving detailed visual structure, including readable text inside generated images. Z.ai offers a hosted API model named glm-image, while the open-weight checkpoint is published as zai-org/GLM-Image.
Outputs

What GLM-Image can produce

Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning Multimodal output
Model profile

Performance characteristics

3/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family GLM-Image
Model type Image Generation
Release date 2026-01-14
Status Current; available through the Z.ai API and as downloadable open weights
Knowledge cutoff notes

A conventional textual knowledge cutoff is not published for this image-generation model. Its autoregressive component is initialized from GLM-4-9B-0414, but that does not establish a separate GLM-Image knowledge cutoff.

Model notes

The hosted API uses the model identifier glm-image, while the downloadable checkpoint is zai-org/GLM-Image. The model combines a 9B autoregressive generator with a 7B diffusion decoder. It supports text-to-image and image-to-image generation, including editing, style transfer, identity preservation, and multi-subject consistency. API image dimensions must be between 512 and 2048 pixels, with width and height divisible by 32. Local inference is memory-intensive; CPU offloading is documented but slows generation. The overall model is released under the MIT License, while incorporated VQ tokenizer and vision components retain Apache-2.0 licensing. Fine-tuning infrastructure is present in the official repository, but detailed model-specific fine-tuning limits are not documented.

Cost

Model pricing

Input $0.015 per image
Output $0.015 per generated image
Model guide

GLM-Image: Z.ai’s Open-Weight Model for Text-Heavy Image Generation

GLM-Image is Z.ai’s hybrid autoregressive and diffusion image-generation model for text-to-image creation and image-to-image editing. Its main differentiator is its focus on accurate text rendering and knowledge-dense compositions such as posters, diagrams, presentation graphics, and multi-panel layouts. It is available through the Z.ai API and as downloadable open weights for local experimentation.

What is GLM-Image?

GLM-Image is an image-generation model from Z.ai, the AI company associated with Zhipu AI. It was released for both hosted use through the Z.ai API and local experimentation through downloadable open weights. Unlike a text-only language model, GLM-Image is primarily designed to produce visual output rather than written answers, code, structured data, audio, or video.

The model supports two main workflows: text-to-image generation and image-to-image transformation. A text prompt can request a new image, while a reference image can guide editing, style transfer, background replacement, subject consistency, identity preservation, or other prompt-directed changes. Its intended use cases include posters, presentation slides, educational diagrams, social-media graphics, and other images where several visual elements and written labels must fit together coherently.

GLM-Image is particularly notable for targeting text-heavy and information-dense visuals. Many image generators can create attractive scenes but struggle with spelling, layout, or multiple lines of readable text. Z.ai positions GLM-Image’s architecture and Glyph Encoder as solutions for more accurate text rendering inside images. That is a provider-supported design goal, not a guarantee that every generated word will be correct.

How the hybrid architecture works

GLM-Image combines two different generation stages. Its 9B-parameter autoregressive component is initialized from GLM-4-9B-0414 and focuses on interpreting the instruction, planning semantics, and establishing the overall composition. Its 7B single-stream DiT diffusion decoder then reconstructs detailed visual content and high-frequency image information.

In practical terms, the first stage is intended to help decide what should appear and where it belongs, while the diffusion stage is responsible for turning that plan into a detailed image. A Glyph Encoder is included to improve the rendering of written characters. This division is important because the model is not simply a conventional diffusion system with a text prompt attached; it uses a language-oriented generation stage alongside a visual decoder.

The architecture also explains an important trade-off. Combining an autoregressive component and a diffusion decoder can support more complex instruction understanding and layout control, but local inference is correspondingly demanding. Users should not treat the downloadable model like a lightweight image filter that will run comfortably on any computer.

Capabilities, inputs, and outputs

GLM-Image accepts text instructions and can use reference images for image-to-image generation. Supported reference-image workflows include editing, style transfer, identity-preserving generation, background replacement, and multi-subject consistency. The exact quality of each transformation depends on the prompt, reference image, and serving implementation.

The model’s native output is an image. It does not provide native text, audio, video, embedding, speech, or tool-action output. It is therefore best understood as a visual-generation model rather than a general-purpose multimodal assistant. Although its autoregressive component is initialized from a language model, that does not make GLM-Image a general text-generation endpoint.

The available research does not specify a conventional context window or maximum output-token limit. Those fields are not applicable in the same way as they are for a text-generation model. For image dimensions, the hosted API supports common aspect ratios and custom width and height values from 512 to 2048 pixels. Both dimensions must be multiples of 32.

API access and local deployment

The hosted Z.ai API exposes the model under the identifier glm-image. The documented price is $0.015 per generated image. API responses provide an image URL, so an application must download the generated file if it needs to store, process, or display the result directly.

This per-image pricing is easier to estimate than token-based billing for a basic image-generation workflow: the main variable is the number of generated images rather than the length of a text completion. However, the supplied documentation does not establish that every possible image-editing or multi-image operation has identical billing behavior, so production users should verify the current API documentation before estimating a larger workload.

For local use, the official checkpoint is published as zai-org/GLM-Image. The implementation is intended to work with the Transformers and Diffusers ecosystem. Z.ai also documents SGLang-compatible endpoints for image generation and image editing. The practical maturity, optimization, and hardware support of those surrounding inference projects may change over time.

Local inference requires substantial GPU resources because both the autoregressive and diffusion components must be served. Z.ai documents CPU offloading as an option for reducing GPU-memory requirements, but offloading moves some work away from the GPU and therefore makes generation slower. The open-weight route is consequently most attractive to users who need local control, experimentation, or integration with their own serving stack and who can accept additional setup and runtime costs.

Where GLM-Image is strongest

  • Text inside images: The Glyph Encoder and model design specifically target more reliable rendering of words and characters than is typical of image generators that treat text as a secondary visual pattern.
  • Dense layouts: Posters, presentation graphics, diagrams, educational illustrations, and multi-panel compositions often require several objects, labels, and relationships to be coordinated in one image.
  • Reference-guided editing: Image-to-image workflows can be used for background replacement, style changes, identity preservation, and maintaining consistency across multiple subjects.
  • Open-weight experimentation: The downloadable checkpoint allows qualified users to inspect, run, and adapt the model locally rather than relying exclusively on a hosted interface.
  • Predictable basic API pricing: The documented hosted price is stated per generated image, which can make simple batch estimates straightforward.

These strengths should be treated as positioning and documented capabilities rather than universal quality guarantees. For example, a model designed to improve text rendering can still produce incorrect words or require several attempts for a complex poster. Results may also vary between the hosted API and local implementations.

Limitations and trade-offs

The largest limitation is the hardware burden of local deployment. The combination of a 9B autoregressive component and a 7B diffusion decoder is substantially more demanding than a small image model or a hosted image-generation tool. CPU offloading expands the range of possible hardware configurations, but it does so at the cost of speed.

The resolution rules are another operational constraint. Hosted output dimensions must remain between 512 and 2048 pixels, and width and height must both be divisible by 32. Requests that do not follow those requirements need to be adjusted before submission. The research does not provide a single maximum file size, generation-time guarantee, or fixed latency figure.

GLM-Image also has a narrow output profile. It cannot replace a language model for drafting copy, a coding model for software development, an audio model for speech, or a video model for animation. It has no documented native tool or function-calling output, and the supplied model information lists streaming, caching, and batch API support as unavailable. Users needing a complete content-production workflow may need to combine it with separate text, editing, or automation tools.

Fine-tuning infrastructure is present in the official repository, but detailed model-specific limits and procedures are not established in the supplied research. Open weights therefore provide an opportunity for research and adaptation, not a promise of a simple or turnkey fine-tuning process. The overall model is released under the MIT License, while incorporated VQ tokenizer and vision components retain Apache-2.0 licensing.

When to choose GLM-Image

Choose GLM-Image when the primary deliverable is a generated or edited image and the image needs meaningful structure, readable labels, or several coordinated visual elements. It is a good candidate for an educational diagram, a promotional poster with headings, a presentation illustration, a social graphic with multiple text blocks, or an image-editing workflow that must preserve a person or subject across changes.

The hosted API is the more practical choice when you want to avoid managing GPU infrastructure and can work with the documented per-image price. The open-weight version is more appropriate when local execution, experimentation, model inspection, or control over deployment is more important than convenience. In either case, test representative prompts before committing to a large production workload, especially if exact spelling or layout is business-critical.

Another type of image generator may be more suitable if the priority is extremely fast generation on modest hardware, a mature consumer application, or a workflow centered on artistic imagery rather than text-heavy composition. A language model is more appropriate for generating the copy that will appear on a poster, while a separate tool may still be needed for final typography and pixel-level design control. GLM-Image’s main reason to be selected is the combination of visual generation, reference-image editing, text rendering, and open-weight availability—not general reasoning, coding, or tool use.

Bottom line

GLM-Image is a specialized Z.ai image model built around a hybrid 9B autoregressive and 7B diffusion architecture. Its clearest practical distinction is its focus on text rendering and structured, information-dense images, supported by both text-to-image and image-to-image workflows. The hosted API costs $0.015 per generated image, while the open-weight release offers local deployment and experimentation.

Its limitations are equally important: local inference is resource-intensive, output dimensions follow strict constraints, and the model does not provide general text, code, audio, video, or tool outputs. For users who need posters, diagrams, presentation graphics, or reference-guided edits—and who can accept those deployment trade-offs—GLM-Image is a focused alternative to more general image-generation options.


Answers to Frequently Asked Questions

What image sizes does GLM-Image support?
The hosted API supports custom widths and heights from 512 to 2048 pixels. Both dimensions must be multiples of 32, so requests that do not meet these requirements need to be adjusted before submission.
Can GLM-Image run locally?
Yes. The open-weight checkpoint is available as zai-org/GLM-Image and is intended to work with the Transformers and Diffusers ecosystems. Local inference requires substantial GPU resources because it serves both the autoregressive and diffusion components, while CPU offloading can reduce GPU-memory needs at the cost of slower generation.
What is GLM-Image?
GLM-Image is an image-generation model from Z.ai designed for text-to-image generation and image-to-image transformation. It is intended for text-heavy and information-dense visuals such as posters, presentation slides, diagrams, and social-media graphics.
How much does the GLM-Image API cost?
The hosted Z.ai API lists GLM-Image at $0.015 per generated image. API responses return an image URL, which applications must download if they need to store, process, or display the file directly.
How does GLM-Image improve text rendering in images?
GLM-Image combines a 9B-parameter autoregressive component, a 7B diffusion decoder, and a Glyph Encoder. This architecture is designed to improve the placement and rendering of written characters, although it does not guarantee perfectly spelled or formatted text in every generation.


Sources 5
Provider

About Z.ai