What Amazon Nova Act v1.0 is
Amazon Nova Act v1.0 is a specialized foundation model from Amazon Web Services for browser and user-interface automation. Its job is to operate within a web interface: interpret what is displayed, determine which action is needed, interact with page elements, and continue toward a task objective.
That makes Nova Act different from a conventional text model. A typical chat model primarily returns text in response to a prompt. Nova Act is primarily used to drive actions in a browser session, such as navigating to a site, selecting a control, entering information, collecting results, or handing a task to a person when approval is required.
The model is accessed through the Amazon Nova Act service and SDK. It is not presented in the supplied documentation as a general-purpose, token-priced chat model in Amazon Bedrock. AWS identifies nova-act-v1.0 as the pinable model selection ID for this release.
Release status and position in AWS's lineup
Nova Act v1.0 reached general availability on December 2, 2025. AWS describes generally available model versions as production releases supported for at least one year. The service is listed for the US East (N. Virginia) AWS Region in the reviewed documentation.
Nova Act occupies a focused position in AWS's AI catalog: it is intended for browser-agent execution rather than broad language generation. Amazon's wider AI portfolio includes other services and model families for conversational applications, foundation-model access, and enterprise development, but those are not substitutes for the specific UI automation runtime that Nova Act provides. The nova-act-latest and nova-act-preview identifiers are aliases for model selection and should not be treated as separate Nova Act model records.
How Nova Act works in practice
A Nova Act workflow combines three elements:
- Natural-language task instructions: the developer describes the browser task the model should perform.
- Browser and visual context: the model works with a browser session and the rendered user interface it needs to interpret.
- Deterministic application code: Python code handles variables, branching, validation, secrets, retries, data processing, and integration with other systems.
This hybrid design is important. The model can handle changes in page layout or wording that would be awkward to encode entirely as fixed selectors, while ordinary code can enforce business rules that should not be left to probabilistic model behavior. For example, Nova Act might locate and fill a form, while Python checks the resulting values, decides whether the workflow can proceed, and records the outcome.
The model can also support human-in-the-loop workflows. When a task reaches a sensitive, ambiguous, or approval-dependent step, the workflow can wait for a person rather than continuing automatically. AWS states that time spent waiting for a human response is excluded from agent-hour billing.
Supported inputs and outputs
Nova Act is multimodal in the practical sense that it works with visual browser and UI information as well as text instructions. The reviewed research identifies image input and multimodal input support, with browser-rendered visual context forming the core operating environment. It does not identify audio or video input support.
Its main output is an action or sequence of actions, not a generated image, audio clip, or video. Actions may include clicking, typing, selecting, navigating, extracting information, and interacting with controls on a page. Nova Act can therefore produce machine-executable behavior through its runtime, but it should not be described as a direct image-output or audio-output model.
AWS does not publish a context-window size or maximum output-token limit for Nova Act v1.0 in the supplied sources. Those omissions matter because the model is designed around an active browser workflow rather than a conventional prompt-and-completion interface. Developers should not assume a specific token limit, JSON mode, streaming interface, caching feature, or dedicated batch API unless AWS documents it for the service version being used.
Core capabilities and strengths
Nova Act's strongest capability is interaction with web interfaces. Its documented and described use cases include:
- Browser navigation and UI control
- Repetitive back-office operations across web applications
- Web data collection and information extraction
- Competitive price monitoring
- Travel and reservation research
- Agentic quality assurance and UI regression testing
- Human-supervised enterprise workflows
The model is particularly useful when a task cannot be completed through a stable, simple API or when the relevant process is exposed primarily through a website. It can follow an instruction such as locating information across pages, completing a repetitive administrative process, or checking a website's behavior while Python code manages the surrounding workflow.
Another strength is the separation between flexible model behavior and deterministic control. Developers are not required to express every interaction as rigid code, but they can still validate important outputs and add safeguards around actions that affect data, accounts, or business processes.
Reasoning, coding, and tool use
Nova Act performs task-oriented reasoning: it interprets a goal in the context of a browser and selects the next UI action. This is narrower than the open-ended reasoning expected from a general-purpose language model. The model's reasoning is valuable when page structure, labels, or navigation paths vary, but it remains tied to the current workflow and available browser context.
Tool use is central rather than incidental. Browser interaction itself is the model's primary tool-oriented environment, and the service is designed to generate or perform actions against that environment. The workflow can also call on ordinary application logic and external integrations through Python orchestration.
Nova Act is not a general coding model. Developers write Python around the model, but the model is not positioned as a standalone system for generating large software projects, solving unrestricted programming problems, or providing token-based code completion. The supplied editorial assessment rates its coding suitability at 5 out of 10 because coding is useful for workflow integration but is not the model's central purpose. That score is an editorial evaluation, not an AWS benchmark or provider-published rating.
Pricing and cost trade-offs
AWS prices Nova Act workflows at $4.75 per agent hour. An agent hour measures real-world elapsed time while an agent is working. Parallel agents generate separate charges. Waiting for a human response in a human-in-the-loop workflow is excluded from agent-hour billing according to AWS's pricing description.
This is different from the usual input-token and output-token pricing used for language models. AWS does not publish separate token rates for Nova Act v1.0 in the supplied research. The practical cost therefore depends on how long the browser agent remains active, how many agents run in parallel, and how efficiently the workflow reaches completion.
For short, simple text-generation tasks, a conventional language model may be more economical and easier to operate. Nova Act's pricing becomes more relevant when the alternative is maintaining brittle browser automation, manually completing repetitive processes, or building a large amount of custom UI-specific logic. The right comparison is the total cost and reliability of completing the browser task, not only the apparent price of a model request.
Limitations and unsupported or unspecified features
Nova Act v1.0 has several important boundaries:
- No customer fine-tuning: AWS states that customers cannot directly fine-tune the model. Behavior is customized through instructions, workflow design, Python logic, browser configuration, tools, secrets, and validation.
- No published context or output-token limits: the reviewed sources do not specify a context length or maximum output-token value.
- Not a general-purpose chat model: it is primarily a browser-agent model rather than a broad conversational or writing system.
- No standalone media generation: the research does not support image, audio, video, music, or speech output capabilities.
- Website dependence: reliability can be affected by page redesigns, authentication, browser state, changing content, task ambiguity, and inaccessible controls.
- Unspecified interface features: streaming, caching, JSON mode, structured output, and a dedicated batch API are not specified in the reviewed Nova Act documentation.
These limitations do not necessarily make the model unsuitable for production. They mean that a production workflow should use validation, error handling, access controls, logging, and human escalation instead of assuming that every browser action will be correct.
When to choose Amazon Nova Act v1.0
Choose Nova Act v1.0 when the main problem is reliable interaction with a website or visual user interface. It is a strong candidate for repetitive web operations, browser-based testing, data collection, price monitoring, and workflows that need both flexible UI interpretation and deterministic application control.
It is also appropriate when a human needs to supervise an otherwise automated process. The workflow can perform routine steps, pause for approval, and continue after a person resolves an exception or confirms a consequential action.
Another option may be more appropriate when the task is primarily long-form writing, open-ended conversation, general software development, image generation, speech synthesis, embeddings, or direct model fine-tuning. A conventional API-first integration may also be preferable when a stable service API is available, because direct API calls can be simpler and less sensitive to visual layout changes than browser automation.
In short, Nova Act should be selected for its browser-agent specialization, not because it is a universal replacement for other model types. Its main trade-off is focused action capability and visual UI navigation in exchange for a narrower scope, agent-hour pricing, and dependence on the behavior of the websites it operates.

