What NVIDIA Ising Calibration 1 35B-A3B does
NVIDIA Ising Calibration 1 35B-A3B is an open-weight vision-language model designed for a narrow but technically demanding task: interpreting plots produced during quantum-computing calibration experiments. It accepts text prompts together with PNG or JPEG images and returns technical text rather than images, audio, control commands, or other non-text outputs.
A user might provide a calibration plot and ask the model to describe the signal, determine whether the data is suitable for parameter extraction, assess the quality of a fit, identify important experimental values, judge the significance of an observed feature, or classify whether the experiment succeeded. These capabilities make it more comparable to a specialized scientific analysis assistant than to a general-purpose chatbot.
The checkpoint is provided by NVIDIA and is available as NVIDIA/Ising-Calibration-1-35B-A3B on Hugging Face. NVIDIA lists it as version 1.0.0 and distributes it under the NVIDIA Open Model License. Its release date in the supplied model data is April 14, 2026.
Positioning and model design
Ising Calibration 1 35B-A3B is a fine-tuned derivative of Qwen3.5-35B-A3B. That relationship matters because the NVIDIA checkpoint is not presented as a general replacement for every model in the Qwen family. Its additional training is aimed at quantum-calibration plot understanding, with coverage of superconducting-qubit and neutral-atom experiment families.
The model uses a mixture-of-experts, or MoE, architecture. It has approximately 35 billion total parameters, while approximately 3 billion parameters are active for each token according to the supplied model information. In practical terms, the sparse design can reduce the amount of model computation used per token compared with a dense model of the same total size, but it does not make deployment lightweight: NVIDIA lists substantial GPU requirements for inference.
Its context length is 262,144 tokens. This is a large input window for combining prompts, demonstrations, multiple images, and supporting technical material, although the useful image capacity and actual performance still depend on the serving implementation and available GPU memory. The suggested maximum generation length is 16,384 tokens.
Verified specifications and supported modalities
| Specification | Supplied information |
|---|---|
| Provider | NVIDIA |
| Model family | NVIDIA Ising Calibration |
| Base model | Qwen3.5-35B-A3B |
| Architecture | Mixture-of-experts vision-language model |
| Total parameters | Approximately 35 billion |
| Active parameters | Approximately 3 billion per token |
| Context length | 262,144 tokens |
| Precision | BF16 |
| Inputs | Text and PNG or JPEG images |
| Outputs | Technical text |
| Suggested maximum output | 16,384 tokens |
| Model version | 1.0.0 |
| Image output | Not supported |
| Audio and video input or output | Not supported in the supplied specifications |
| Web search and tool use | Not listed as supported capabilities |
The model supports single-image and multi-image analysis according to the supplied model description. This is useful when a calibration workflow requires comparing related plots or supplying several stages of an experiment in one prompt. Its output remains textual, so any extracted values or conclusions should be transferred into laboratory records or downstream software only after review.
Primary quantum-calibration use cases
The model's intended use is analysis of calibration plots rather than open-ended scientific conversation. Typical tasks include describing plotted signals, evaluating whether a measurement is appropriate for extracting parameters, assessing fit quality, extracting experimental values, determining whether an observed feature is significant, and classifying experimental success.
For example, a calibration engineer could provide a plot and ask whether the curve appears suitable for fitting, which parameters are visible, and what experimental conclusion is supported by the plotted evidence. A researcher could submit several related images and request a comparison of signal behavior or a structured explanation of why a fit appears reliable or questionable. The model can also help turn visual experimental evidence into a written technical interpretation.
These applications require domain judgment. The model card's intended tasks do not make its conclusions equivalent to a validated measurement pipeline. A result that appears plausible should be checked against raw data, fitting procedures, instrument settings, and the expectations of a qualified quantum-computing researcher.
Training and reported evaluation
NVIDIA describes two sequential supervised fine-tuning phases. The first used approximately 23.8 thousand in-context-learning-formatted entries to teach processing of multi-image demonstrations. The second used approximately 48.7 thousand zero-shot entries, producing a total of approximately 72.5 thousand training entries. The supplied research describes the data as synthetic and augmented with Qwen3.5-397B-A17B.
On NVIDIA's QCalEval benchmark, the model achieved an overall score of 74.7, compared with 55.5 for the Qwen3.5-35B base model. The benchmark contains 243 entries across 87 scenario types and 22 experiment families. NVIDIA reports particularly strong relative improvement in experimental-conclusion and fit-quality-assessment tasks.
These are provider-reported benchmark results, not a guarantee for every laboratory dataset. The benchmark's score should be treated as evidence of performance on the stated evaluation rather than as proof that the model can reliably interpret all plot formats, experimental conditions, or noisy measurements.
Deployment, performance, and cost considerations
The checkpoint can be downloaded from NVIDIA's Hugging Face repository. The supplied documentation describes loading it with Transformers through an image-text-to-text pipeline or another multimodal model interface. NVIDIA also documents serving it with vLLM and SGLang using OpenAI-compatible chat-completions endpoints.
NVIDIA lists a minimum test configuration of either two NVIDIA L40S GPUs with 48 GB of memory each or one NVIDIA H100 GPU with 80 GB of memory. BF16 inference, FlashAttention, and vLLM are recommended. These requirements place the model in a workstation, laboratory, server, or data-center deployment category rather than the typical consumer-laptop category.
No input or output API prices are supplied for this checkpoint. It is downloadable under the NVIDIA Open Model License, but the total cost of using it still depends on GPU ownership, cloud rental, power, storage, engineering time, and serving infrastructure. The MoE design uses approximately 3 billion active parameters per token, which may improve the computation-to-capability trade-off compared with a dense 35-billion-parameter model, but the model's total size and multimodal processing requirements still create a significant memory and hardware burden.
There is no verified public speed figure in the supplied research. In practice, response speed will depend on the GPU configuration, inference engine, batching, image processing, prompt length, and requested output length. A smaller or more general model may be easier and cheaper to run when specialized quantum-plot analysis is not required.
Reasoning, coding, and tool support
Ising Calibration 1 35B-A3B is designed to perform task-specific reasoning over visual experimental evidence. Its relevant reasoning capabilities include connecting plotted features to requested conclusions, judging fit quality, extracting values, and classifying experimental outcomes. This is specialized analytical reasoning, not a claim that the model is a general-purpose autonomous scientist.
Coding is not the model's primary purpose. It may be useful in a workflow that surrounds the model—for example, a developer can build an application that submits plots and records textual results—but the supplied information does not establish dedicated code-generation capabilities or a coding benchmark. Likewise, tool use, web browsing, function calling, and direct laboratory-control actions are not listed as supported model capabilities. An OpenAI-compatible serving endpoint describes an interface format; it does not by itself prove that the model can call tools or operate instruments.
Strengths and limitations
Where it is strongest
- It is specifically fine-tuned for quantum-computing calibration plots rather than merely being a general vision-language model.
- It supports text plus PNG or JPEG images, including single-image and multi-image analysis.
- Its 262,144-token context window can accommodate substantial prompts, demonstrations, and supporting technical material.
- Its sparse mixture-of-experts design uses approximately 3 billion active parameters per token.
- Downloadable weights allow self-hosted deployment and integration into controlled research environments.
- The reported QCalEval results indicate meaningful gains over the Qwen3.5-35B base model on the evaluated quantum-calibration tasks.
Where it is limited
- It requires substantial GPU memory and deployment expertise, with NVIDIA listing two L40S GPUs or one H100 as minimum test configurations.
- It produces technical text, not direct hardware-control actions or guaranteed machine-readable laboratory commands.
- Its specialization means it is not the obvious choice for general chat, broad coding, web research, audio, video, or image generation.
- The supplied research does not provide pricing, a universal throughput figure, or a separate knowledge-cutoff date.
- Plot interpretation can be affected by image quality, unfamiliar experiment formats, ambiguous labels, and unusual measurement conditions.
- Model-generated conclusions should be checked against raw measurements and domain expertise before changing hardware settings or making operational decisions.
When to choose NVIDIA Ising Calibration 1 35B-A3B
Choose this model when the central problem is interpreting quantum-computing calibration plots and you can provide the GPU infrastructure needed for self-hosted inference. It is especially suitable for research groups, calibration engineers, and developers building tools around superconducting-qubit or neutral-atom experiments. Its specialization may justify the deployment cost when fit assessment, parameter extraction, and experiment-status classification are recurring tasks.
A general vision-language model may be more appropriate when the work spans many scientific domains, ordinary document images, broad visual question answering, or general assistant tasks. A smaller model may be preferable when low latency, low cloud cost, or deployment on limited hardware matters more than domain specialization. A conventional numerical fitting or analysis pipeline remains more appropriate when a result must be deterministic, reproducible, and directly tied to validated measurement algorithms.
The similarly named NVIDIA Ising Calibration 1.5 31B is a separate model and should not be treated as the API or deployment page for this 35B-A3B checkpoint. For this model, the practical distinction is clear: it is a specialized, downloadable vision-language system for quantum-calibration plot analysis, with technical strengths that come alongside high infrastructure requirements and a need for expert validation.

