What is Amazon Nova Micro?
Amazon Nova Micro is a text-only foundation model provided by Amazon and delivered through Amazon Bedrock. Its purpose is straightforward: handle routine language-processing tasks quickly and inexpensively. The model can read text and generate text, making it suitable for transforming, labeling, summarizing or extracting information from large numbers of inputs.
The canonical Bedrock model identifier is amazon.nova-micro-v1:0. It belongs to the Amazon Nova family, but its positioning is different from larger models intended for more difficult reasoning, complex coding or multimodal workloads. Nova Micro favors throughput, low latency and low cost over maximum capability.
That positioning makes it useful as a production component inside an application. For example, it can classify support tickets before routing them, extract fields from incoming documents, summarize conversations, translate text or provide a first-pass answer to a frequently asked question. These tasks generally benefit more from speed and consistent operating cost than from the broadest possible model intelligence.
Amazon Nova Micro specifications
| Specification | Verified detail |
|---|---|
| Provider | Amazon |
| Model family | Amazon Nova |
| Model ID | amazon.nova-micro-v1:0 |
| Release date | December 5, 2024 |
| Input | Text |
| Output | Text |
| Context window | 128,000 tokens |
| Maximum output | 5,000 tokens |
| Knowledge cutoff | October 2024 |
| Standard input price | $0.035 per 1 million tokens |
| Standard output price | $0.14 per 1 million tokens |
| Provider status | Active on the reviewed model page |
The 128K-token context window is the amount of combined input information the model can consider in a request, subject to the specific request and service constraints. The 5K-token maximum output limits how much text it can produce in one response. These limits make the model suitable for long text inputs and substantial summaries, although a larger context window does not automatically mean that the model will reason equally well over every detail.
Text-only input and output
Nova Micro accepts text and returns text. It does not natively accept images, audio or video, and it does not generate images, audio, speech or video. It is also not an embedding model. This is an important practical distinction: an application that needs direct image understanding, speech processing or visual generation must use a different model or add separate services around Nova Micro.
Its text-only design is also part of its cost and speed advantage. A workflow that only needs to classify text or extract values from text does not pay the complexity or latency associated with multimodal processing. Conversely, converting an image or recording into text first would add another processing step and would not give Nova Micro native access to the original modality.
Capabilities and application features
AWS documents response streaming, guardrails, client-side tool calling and prompt caching for Nova Micro through Amazon Bedrock. Streaming allows an application to receive portions of a response as they are generated instead of waiting for the complete answer. This can improve perceived responsiveness in chat or interactive interfaces, even though it does not increase the model's underlying reasoning ability.
Client-side tool calling
Nova Micro supports client-side tool calling. In this arrangement, the application describes available functions, the model can request one of them, and the application executes the requested operation. The model does not independently browse the web or complete an external action by itself. Your software remains responsible for validating the request, running the tool and deciding whether confirmation is required.
This can be useful for narrowly scoped operations such as looking up an order, querying an internal database or classifying a request before sending it to another system. Tool calling should not be confused with native web search, and web-search support is not documented for this exact model.
Prompt caching
Prompt caching is useful when the same instructions or message content are reused across many requests. AWS documents cache checkpoints in system and message content, with a minimum of 1,000 tokens per checkpoint and up to four checkpoints per request. AWS also states that Amazon Nova models support a maximum of 20,000 tokens for prompt caching.
For example, an application with a long, stable instruction set or repeated reference material may be able to reuse that content rather than processing it as entirely new context every time. The practical benefit depends on the request pattern, region and applicable Bedrock pricing, so caching should be evaluated against the application's actual traffic.
Fine-tuning and batch inference
AWS lists Nova Micro as eligible for fine-tuning in Amazon Bedrock. Fine-tuning can adapt the model to a specialized task or output pattern using a suitable training dataset, but it does not turn the model into a multimodal system or guarantee strict structured output.
The base model is also listed for batch inference. Batch inference is intended for asynchronous processing: an application submits many prompts, and the inputs and outputs are stored through Amazon S3. This is a better fit for workloads such as overnight document processing or large-scale historical classification than for an interactive request that needs an immediate response.
Reasoning and coding position
Amazon Nova Micro is best understood as a lightweight language model rather than a specialist reasoning model. The supplied evaluation data gives it a reasoning score of 4 out of 10 and a coding score of 5 out of 10; these are editorial or catalog assessments, not scores published by Amazon and not standardized benchmark results.
In practical terms, Nova Micro can be appropriate for simple transformations, narrowly defined text-to-SQL tasks, routing logic, extraction and routine code-related text generation. It is less suitable when a task depends on extended multi-step reasoning, difficult debugging, complex software architecture or highly reliable answers to ambiguous questions. For those workloads, a larger reasoning-oriented or more capable model may justify its higher cost and latency.
Pricing and value
The reviewed standard on-demand pricing is $0.035 per 1 million input tokens and $0.14 per 1 million output tokens. Input and output are priced separately, and generated text costs more per token than supplied text under these rates. Actual charges can vary by region, inference tier or workload type, so AWS pricing should be checked before deployment.
At these rates, Nova Micro is aimed at high-volume applications where even small per-request savings matter. A classifier processing many short records, for instance, may benefit from a model optimized for low cost more than from a larger model that produces more elaborate answers. Streaming, caching and batch inference provide additional ways to align the service with the workload: streaming for interactive responsiveness, caching for repeated context and batch inference for asynchronous volume.
The trade-off is capability. A low token price does not make Nova Micro the best choice for every task. If a request requires interpreting images, generating media, solving difficult problems or producing reliably constrained JSON, the cost of switching to a more suitable model may be lower than trying to compensate for Nova Micro's limitations in application code.
Best use cases for Nova Micro
- Text classification and routing: label support requests, sort documents or direct incoming content to the appropriate workflow.
- Summarization: produce concise summaries of text documents, conversations or customer interactions.
- Translation and rewriting: transform text between languages, formats or writing styles.
- Information extraction: identify names, dates, categories or other fields from unstructured text.
- Automated FAQs: answer narrowly scoped, text-based questions where low latency and low cost are important.
- Lightweight retrieval-augmented generation: use retrieved text as context for straightforward answers.
- Text-to-SQL and specialized transformations: handle narrowly defined tasks, particularly when customization is appropriate.
- High-volume asynchronous processing: process large datasets through Bedrock batch inference.
Limitations to plan for
- It has no native image, audio or video input.
- It has no native image, audio, speech or video output.
- It is not an embedding model.
- Structured outputs are documented as unsupported.
- Separate native JSON-mode support was not verified in the reviewed documentation.
- Web-search support is not documented for the exact model.
- It is less appropriate for demanding reasoning, complex coding and difficult analysis than larger models.
- Regional availability and inference-profile support depend on the AWS Region and account configuration.
The structured-output limitation deserves particular attention in production systems. Nova Micro can be asked to format an answer as JSON, but the reviewed research does not establish that it can enforce a JSON Schema constraint. Applications that require machine-validated output should add their own validation and error-handling layer, or select a model with verified structured-output support.
When to choose Amazon Nova Micro
Choose Nova Micro when the task is primarily text-based, the instructions are reasonably well defined and the application needs low latency or economical large-scale processing. It is especially attractive when a workflow performs the same kind of operation repeatedly, such as classification, extraction, summarization or routing.
It may not be the right choice when quality depends on deep reasoning, visual or audio understanding, media generation, strict schema-constrained responses or sophisticated coding. In those cases, a larger Amazon Nova model or another specialized option may be more appropriate. The supplied research specifically identifies Nova Pro and newer Nova 2 models as potential alternatives when difficult reasoning, complex coding, long-form generation or multimodal input matters more than minimum cost and latency.
A sensible architecture can also use Nova Micro selectively. An application might send routine, low-risk requests to Nova Micro while escalating ambiguous or complex cases to a more capable model. That approach preserves the model's cost and speed advantages without treating it as a universal replacement for larger systems.
Availability and Bedrock access
Nova Micro is available through Amazon Bedrock using the in-Region model ID amazon.nova-micro-v1:0. AWS also documents US and EU geographic inference identifiers, including us.amazon.nova-micro-v1:0 and eu.amazon.nova-micro-v1:0. Supported Bedrock interfaces include Invoke, Converse, InvokeModelWithResponseStream and ConverseStream, with AWS recommending the Bedrock Runtime endpoint for new applications.
The current AWS model page labels Nova Micro active. However, the reviewed lifecycle metadata contains older end-of-life wording that predates AWS's September 2026 lifecycle-policy transition. No definitive exact shutdown date was established in the model-specific documentation reviewed, so teams deploying it should monitor AWS's current model card and regional availability information.

