What is Cohere North Mini Code?
North Mini Code is Cohere's first model in the North family of code-agent models. It is designed for software-engineering systems that do more than answer a coding question in a single response. A coding agent can use the model to inspect a repository, plan a change, issue terminal or shell commands, edit multiple files, run checks, and revise its work across several turns.
The model uses a Mixture-of-Experts, or MoE, architecture. In simple terms, the model contains 30 billion total parameters, but approximately 3 billion are active for each token processed. This can reduce the computation required for each step compared with running every parameter every time. The result is a model positioned for relatively efficient agentic coding while still offering a very large working context.
North Mini Code is not presented as a broad replacement for Cohere's general Command models. Its focus is narrower and more practical: repository-level software engineering, terminal interaction, code generation, debugging, code review, and related workflows in which the model can use external tools.
Key specifications at a glance
| Specification | North Mini Code |
|---|---|
| Provider | Cohere |
| Model ID | north-mini-code-1-0 |
| Model family | North |
| Architecture | Mixture of Experts |
| Total parameters | 30 billion |
| Active parameters | Approximately 3 billion |
| Context window | 256,000 tokens |
| Maximum output | 64,000 tokens |
| License | Apache 2.0 |
| Text output | Yes |
| Image input | Supported according to the supplied model documentation |
| Audio, video, and image output | Not supported |
The 256K-token context window is especially relevant for software projects. It can provide room for repository documentation, source files, logs, test output, and tool results in one extended task. The maximum output of 64K tokens is also suited to substantial code changes or long-form technical responses, although practical applications may still impose their own response limits.
Why it is built for agentic coding
Traditional code completion usually predicts a short continuation from the text immediately surrounding the cursor. North Mini Code is aimed at a broader workflow in which the model helps manage a software task from investigation through validation. It can be used inside an agent harness that gives it access to files, a shell, a repository, or other defined tools.
For example, an agent using North Mini Code might receive a request to add an API endpoint. It could inspect the project's layout, identify the relevant routing and data-access files, propose a plan, edit several files, run tests, read the resulting errors, and make corrections. The model itself does not automatically gain unrestricted access to a computer; the surrounding application must provide tools and enforce permissions. This distinction matters for security and deployment design.
Cohere describes support for repository-level changes, terminal-based agents, code review, system-architecture mapping, sub-agent orchestration, scientific coding, and algorithmic reasoning. The model has been positioned for use with agent systems such as OpenCode and SWE-agent, while its tool-oriented design can also fit other coding-agent implementations.
Coding, reasoning, and tool support
North Mini Code's main capability is software engineering. It is intended to reason over code structure, dependencies, implementation requirements, and test results rather than only generate isolated snippets. This makes it a better fit for tasks where the model must maintain context across multiple files and decide what to inspect or change next.
The supplied research identifies reasoning and tool use as supported capabilities. Cohere's documentation also describes structured outputs, multilingual support, image inputs, safety modes, and citations. Structured outputs can help an agent return predictable fields for plans, actions, or status information, but that should not automatically be treated as proof of a separate JSON-mode feature. The available model record does not verify a distinct JSON-mode capability.
Tool use is implemented through the application or API integration around the model. A developer can expose carefully scoped actions such as reading a file, searching a repository, applying a patch, or running a test command. The model can then select or request those actions as part of a workflow. Production systems should still validate tool arguments, restrict filesystem access, isolate command execution, and require approval for destructive operations.
API, private, and local deployment
North Mini Code is listed as live in Cohere's model catalog and can be accessed through Cohere's Chat API. Cohere also supports production deployment through Model Vault. In addition, the company publishes downloadable weights in BF16, FP8, and W4A16 formats, allowing organizations to evaluate self-hosted or privately managed deployments under the Apache 2.0 license.
The available launch information identifies one H100 GPU as a minimum deployment target for certain supported configurations. This is not a universal hardware requirement: actual needs vary with numerical precision, quantization, sequence length, concurrency, serving software, and the amount of context retained. A quantized deployment may have different memory and performance characteristics from a BF16 deployment.
These deployment choices are important to the model's positioning. Hosted API access is simpler for teams that want to begin quickly and avoid operating inference infrastructure. Model Vault and private deployment are more relevant to organizations that need greater control over data handling, network location, or operational integration. Downloadable weights can provide additional control, but they shift responsibility for hardware, scaling, monitoring, upgrades, and security to the operator.
Pricing and cost trade-offs
Cohere's documentation states that North Mini Code is free for trial and production API keys until applicable Cohere API rate limits are reached. This means the model can be used without a stated per-token charge within the documented limits, but it should not be interpreted as unlimited hosted inference or as a guarantee of unlimited production capacity. Rate limits, account requirements, and usage policies still apply.
Self-hosted inference does not create a per-token Cohere API bill, but it is not cost-free. The operator must pay for GPU capacity, storage, networking, engineering time, observability, maintenance, and power. A single-GPU starting point may be practical for some configurations, while higher concurrency or longer contexts can require more capacity.
The approximately 3-billion-active-parameter design is a useful efficiency trade-off. It may reduce the computation per token compared with a dense model with a similar total size, which can help latency and operating cost. However, active-parameter count alone does not determine real-world speed. Hardware, quantization, prompt length, output length, batching, and agent tool calls all affect the total time and cost of completing a coding task.
Main strengths and limitations
Strengths
- Focused coding design: The model is optimized for software-engineering agents, repository changes, terminal work, and code review.
- Large working context: A 256K-token context can accommodate substantial project material, documentation, logs, and tool results.
- Efficient MoE structure: Approximately 3 billion active parameters may offer a favorable computation trade-off for coding-agent workloads.
- Long responses: A maximum output of 64K tokens supports large patches, explanations, and multi-step task results.
- Deployment flexibility: Teams can use the Cohere API, Model Vault, OpenRouter, or downloadable weights.
- Permissive license: The Apache 2.0 license supports a broad range of development and private-deployment scenarios, subject to the license terms.
Limitations
- Specialized scope: It is primarily a coding model and may be a less suitable choice for broad business conversation, general content generation, or multilingual enterprise work where a general Command model is more appropriate.
- No media generation: The model produces text and can accept documented image input, but it does not generate images, audio, or video.
- Tool access is external: The model does not independently operate a terminal or repository. Those capabilities depend on an agent framework and the tools that framework exposes.
- Infrastructure can be demanding: Private deployment requires suitable GPUs and engineering support, and long contexts or high concurrency can increase resource requirements.
- Hosted access is rate-limited: Free API availability ends at applicable rate limits, so larger production workloads may require account and capacity planning.
- No published knowledge cutoff: Cohere's currently available documentation does not provide a direct knowledge-cutoff date for this model.
Supported inputs and outputs
North Mini Code is primarily a text-generation model. Text is the central input and output modality, and the documented capability set includes image input for supported workflows. This could be useful when a coding task includes a screenshot of an error, interface, diagram, or other visual reference, but the supplied research does not establish broader image-understanding limits or a specific image resolution.
There is no evidence in the supplied specifications that North Mini Code outputs images, audio, video, music, speech, or embeddings. Its structured-output support can help applications consume predictable text responses, while citations and tool use can support more controlled agent workflows when implemented by the surrounding system.
When to choose North Mini Code
Choose North Mini Code when the central problem is software engineering performed over multiple steps. It is a particularly plausible fit for:
- Repository-level feature development and refactoring
- Terminal or shell-based coding agents
- Debugging workflows that require reading logs and running tests
- Automated code review and architecture mapping
- Scientific or algorithmic programming tasks
- Private, local, or on-premises coding assistance
- Organizations that want an Apache 2.0 open-weight coding model
It may be less appropriate when the main requirement is general-purpose enterprise chat, broad multilingual business writing, image generation, speech, or media creation. In those cases, a general model or a modality-specific system may be a better match. Within Cohere's lineup, the research positions Command models as more suitable for broad enterprise generation, while North Mini Code is the more focused option for coding-agent workflows.
Bottom line
North Mini Code combines an agent-focused coding objective with an unusually large context window, long output capacity, and multiple deployment paths. Its 30-billion-parameter MoE architecture activates approximately 3 billion parameters per token, giving it a clear efficiency-oriented design for repository and terminal tasks. The model is most compelling for teams building coding agents or deploying private coding assistance, especially when open weights and Apache 2.0 licensing matter.
Its trade-offs are equally clear: it is specialized rather than a universal assistant, hosted use is subject to rate limits, and self-hosting requires real infrastructure. For software-engineering agents that need to inspect, modify, and validate substantial codebases, those trade-offs may be worthwhile. For general conversation or non-text generation, another model category is likely a better choice.

