Granite Code

Granite-3B-Code-Instruct

by IBM watsonx · Retired; withdrawn from IBM watsonx.ai on 2025-07-17

IBM Granite-3B-Code-Instruct is a compact 3B-parameter coding model with a documented 128K context window and Apache 2.0 license. It supports text-based code generation, explanation, repair, and conversion, but its IBM watsonx.ai deployment was withdrawn on July 17, 2025. The model remains relevant for self-hosted coding assistants, research, archival evaluation, and long-context experiments.

Text Reasoning Coding
Granite-3B-Code-Instruct was IBM’s compact open-weight model for programming tasks. It combined instruction tuning for code-related work with a 128,000-token context window, allowing users to provide large files, documentation, or several related modules in one prompt. The model is no longer an actively supported IBM watsonx.ai endpoint, so its current value is mainly in downloadable, self-hosted, historical, and reproducible research use.
Outputs

What Granite-3B-Code-Instruct can produce

Text
Inputs

What it can understand

Text
Capabilities

Supported features

Fine-tuning
Model profile

Performance characteristics

4/10 Reasoning
6/10 Coding
8/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family Granite Code
Model type Coding
Context window 128K tokens
Maximum output 8K tokens
Release date 2024-05-09
Status Retired; withdrawn from IBM watsonx.ai on 2025-07-17
Deprecation date 2025-04-16
Shutdown date 2025-07-17
Knowledge cutoff notes

IBM's publicly available model documentation and model card do not provide a verified exact knowledge-cutoff date for this model.

Model notes

IBM's canonical watsonx.ai model ID is ibm/granite-3b-code-instruct. The related IBM Hugging Face model card uses ibm-granite/granite-3b-code-instruct-128k. IBM lists a 128,000-token input-plus-output context window and an 8,192-token maximum-new-token limit for multitenant serving. IBM published a deprecation date of April 16, 2025 and a withdrawal date of July 17, 2025, recommending granite-3-3-8b-instruct as the replacement. The open model artifacts are Apache 2.0 licensed and remain relevant for self-hosted or research use, but the IBM-hosted endpoint is retired. Editorial scores are comparative estimates, not IBM specifications.

Cost

Model pricing

Input $0.0006 per 1,000 input tokens historically on IBM watsonx.ai
Output $0.0006 per 1,000 output tokens historically on IBM watsonx.ai
Model guide

IBM Granite-3B-Code-Instruct: A Compact, Long-Context Coding Model Now Retired

IBM Granite-3B-Code-Instruct is a 3-billion-parameter, Apache 2.0-licensed instruction-tuned coding model built for code generation, explanation, conversion, debugging, documentation, and repository-scale analysis. Its 128,000-token context window and relatively small size made it suitable for controlled or self-hosted deployments, but IBM withdrew its watsonx.ai deployment on July 17, 2025. The open model artifacts remain relevant for research, evaluation, and self-hosted use.

What is Granite-3B-Code-Instruct?

Granite-3B-Code-Instruct is a decoder-only language model developed by IBM Research and fine-tuned to follow instructions about software and programming. The “3B” designation refers to approximately 3 billion parameters, the learned numerical values that help a model recognize patterns and produce text. Its relatively small size was intended to make coding assistance more practical in lower-resource or controlled environments than much larger models.

The model was designed for code generation, explanation, repair, editing, translation, documentation, and related code-intelligence tasks. A user could ask it to create a function from a description, explain unfamiliar code, convert code between languages, suggest a fix, or analyze a larger collection of files. IBM’s watsonx.ai identifier was ibm/granite-3b-code-instruct. The associated open model card uses ibm-granite/granite-3b-code-instruct-128k.

Granite-3B-Code-Instruct belongs to IBM’s Granite Code family, but it should not be confused with IBM’s current hosted catalog as a whole. Its watsonx.ai deployment has been withdrawn, while the open artifacts remain available for users who can operate the model independently.

Core specifications and the 128K context limit

SpecificationDetails
ProviderIBM
Model familyGranite Code
ParametersApproximately 3 billion
Context window128,000 tokens
Maximum new tokens8,192 in IBM multitenant serving documentation
LicenseApache 2.0
Primary outputText, including source code and explanations
Hosted statusWithdrawn from IBM watsonx.ai on July 17, 2025

The 128,000-token context window is one of the model’s most useful documented characteristics. A context window is the amount of text the model can consider across the input and generated response. In practical terms, the limit can accommodate substantial source files, API documentation, configuration files, or multiple related modules, subject to the serving environment and the actual token count.

The 8,192-token maximum-new-token figure applies to IBM’s documented multitenant serving configuration. It describes the maximum response length in that environment, not a guarantee that every self-hosted implementation will use the same setting. A local deployment may impose different limits based on its serving software, available memory, quantization, and configuration.

Coding capabilities and training focus

IBM positioned Granite Code models for several stages of software development. Granite-3B-Code-Instruct can be used for:

  • Generating functions, scripts, and code snippets from natural-language requirements.
  • Explaining what existing code does and identifying likely problem areas.
  • Repairing or editing code when the user supplies an error, requirement, or failing example.
  • Converting code or concepts between programming languages and frameworks.
  • Writing comments, documentation, and other natural-language descriptions of software.
  • Analyzing large code-related prompts that require several files or supporting documents.

The documented instruction-tuning data included code, mathematics, language, commits, API-calling examples, and synthetic programming datasets. These categories indicate the intended training focus, but they do not constitute a guarantee of correctness for every language, framework, or software task. Generated code still requires review, testing, dependency checks, and security evaluation before it is used in production.

For beginners, the model is best understood as a text generator specialized for programming rather than an autonomous development environment. It can propose code and reasoning in text, but the supplied research does not verify built-in code execution, a current tool-calling interface, or a guaranteed connection to a repository, browser, terminal, or external services.

Modalities, reasoning, and tool support

Granite-3B-Code-Instruct is text-only in both its inputs and outputs. It does not natively accept images, audio, or video, and it does not generate image, audio, or video content. A workflow that needs screenshots, voice, or other media would need a separate multimodal system or an external preprocessing step that converts those inputs into text.

It is a coding-focused instruction model, not a model documented as having a separate extended reasoning mode. The available research does not provide a verified reasoning benchmark or a dedicated reasoning-level specification. Its practical reasoning ability therefore depends on the prompt, the supplied code and documentation, and the complexity of the task. It may perform useful step-by-step analysis, but users should not treat that behavior as proof of reliable correctness.

Tool or function support is not verified in the supplied model information. Although API-calling data was included in the instruction-tuning description, that does not by itself establish a native, structured tool-use feature in the retired hosted endpoint or in self-hosted deployments. Users should distinguish between generating a tool-call-like text format and actually invoking a tool through an integrated runtime.

Availability, lifecycle, and historical pricing

IBM announced a deprecation date of April 16, 2025, followed by withdrawal from watsonx.ai on July 17, 2025. IBM recommended Granite-3.3-8B-Instruct as a migration alternative. Consequently, Granite-3B-Code-Instruct should not be selected for a new project that requires an actively supported IBM-hosted production endpoint without first verifying a current replacement.

The downloadable model artifacts remain useful for self-hosted deployments, archival evaluation, and reproducible research. They are released under the Apache 2.0 license, which generally permits research and commercial use subject to the license and applicable deployment obligations. Open availability does not mean that IBM continues to provide hosted inference, support, uptime, or current model maintenance.

IBM historically listed the model at $0.0006 per 1,000 input tokens and $0.0006 per 1,000 output tokens on watsonx.ai. These were hosted-service prices associated with the earlier deployment and should not be treated as current API prices after withdrawal. Self-hosting has no IBM per-token charge, but it creates infrastructure costs for hardware, storage, serving software, maintenance, electricity, and operational support. Actual cost depends heavily on utilization and deployment design.

Strengths and limitations

Strengths

  • Compact scale: Approximately 3 billion parameters can be easier to deploy and operate than much larger coding models.
  • Long context: The documented 128,000-token window supports large code and documentation prompts.
  • Code specialization: Instruction tuning was directed toward programming, mathematics, commits, APIs, and related software tasks.
  • Permissive licensing: The Apache 2.0 license supports self-hosted research and commercial experimentation subject to its terms.
  • Controlled deployment: Independent hosting can be useful when an organization wants more control over infrastructure and data handling than a hosted endpoint provides.

Limitations

  • Retired hosted service: IBM withdrew the watsonx.ai deployment, so it is not an ordinary choice for a new managed IBM production integration.
  • Older generation: Newer or larger coding models may provide better accuracy, broader language coverage, or stronger performance on complex software tasks.
  • Text-only operation: The model does not directly process images, audio, or video.
  • No verified native tools: The supplied research does not confirm built-in function calling, code execution, web search, or repository access.
  • Deployment responsibility: Self-hosting requires users to manage inference infrastructure, security, scaling, evaluation, and updates.
  • Variable results: Coding quality can differ by programming language, framework, prompt quality, and task complexity.

These limitations matter especially for production software development. A small model can be attractive for cost and latency, but a lower parameter count may also mean less reliable handling of ambiguous requirements, unfamiliar libraries, multi-step debugging, or large architectural changes.

When to choose Granite-3B-Code-Instruct

Choose Granite-3B-Code-Instruct when the main requirement is a compact, open, code-specialized model that can accept long prompts and run under your own control. It is a reasonable candidate for archival experiments, reproducible evaluations, lightweight coding assistants, code explanation, code conversion, and prototypes where Apache 2.0 licensing and self-hosting are important.

Its speed-and-cost position is best understood as a trade-off rather than a guaranteed benchmark result. Compared with larger coding models, a 3-billion-parameter model may require fewer resources and may be faster or less expensive to operate, but it can give up capability on difficult or highly ambiguous tasks. The editorial speed and cost ratings associated with this model are comparative estimates, not IBM-published measurements.

Another option is more appropriate when you need a currently supported IBM-hosted endpoint, verified tool integration, multimodal input, speech or media generation, or consistently strong performance on complex production code. For IBM users specifically, the documented migration recommendation was Granite-3.3-8B-Instruct. For any successor, teams should verify current availability, pricing, context limits, licensing, and supported interfaces rather than assuming that the older model’s specifications carry forward.

Practical verdict

Granite-3B-Code-Instruct remains a technically interesting compact coding model because it combines instruction tuning, a 128K context window, and Apache 2.0 licensing. Its strongest present-day role is independent or research use, not new IBM-hosted production work. The decisive evaluation questions are whether its coding quality is sufficient for the target language and task, whether the available hardware can serve it efficiently, and whether the project can accept a retired model without ongoing provider support.


Answers to Frequently Asked Questions

What license does Granite-3B-Code-Instruct use, and who should choose it?
Granite-3B-Code-Instruct is released under the Apache 2.0 license, which generally supports research and commercial use subject to its terms. It is best suited to users seeking a compact, code-specialized, long-context model for self-hosted assistants, prototypes, reproducible evaluations, or archival research. IBM users requiring a current hosted service were directed toward Granite-3.3-8B-Instruct.
Is Granite-3B-Code-Instruct still available on IBM watsonx.ai?
No. IBM announced deprecation on April 16, 2025, and withdrew Granite-3B-Code-Instruct from watsonx.ai on July 17, 2025. The model artifacts remain available for self-hosting, archival evaluation, and research, but IBM no longer provides the hosted endpoint or ongoing hosted-service support.
What is IBM Granite-3B-Code-Instruct?
IBM Granite-3B-Code-Instruct is a decoder-only language model from IBM Research with approximately 3 billion parameters, fine-tuned for software and programming tasks such as code generation, explanation, repair, editing, translation, and documentation.
What is the context window of Granite-3B-Code-Instruct?
Granite-3B-Code-Instruct has a documented context window of 128,000 tokens, allowing it to process large source files, documentation, configuration files, or multiple related code modules. IBM’s multitenant serving documentation listed a maximum of 8,192 newly generated tokens.


Sources 5
Provider

About IBM watsonx