SenseNova-MARS

SenseNova-MARS-32B

by SenseTime · Current open-weight model

SenseNova-MARS-32B is SenseTime’s MIT-licensed open-weight multimodal reasoning model for visual deep search and agentic tool use. It accepts text and visual inputs, produces text, and is designed to coordinate image cropping, image search, and text search during multi-step reasoning. Its published configuration specifies a 262,144-token text context length, while official deployment examples support Transformers, vLLM, and SGLang. No hosted pricing or maximum output-token limit was identified for the exact checkpoint.

Text Reasoning Coding
SenseNova-MARS-32B is the larger model in SenseTime’s SenseNova-MARS family. Released on January 29, 2026, it is designed for visual tasks that require iterative reasoning rather than a single image description. The model accepts text and visual inputs, produces text, and can be connected to external search and image-processing tools. It is available for self-hosted deployment through the official SenseNova GitHub and Hugging Face repositories, with examples for Transformers, vLLM, and SGLang.
Outputs

What SenseNova-MARS-32B can produce

Text
Inputs

What it can understand

Text Images Video Multimodal input
Capabilities

Supported features

Tool use
Model profile

Performance characteristics

8/10 Reasoning
6/10 Coding
4/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family SenseNova-MARS
Model type Multimodal
Context window 262K tokens
Release date 2026-01-29
Status Current open-weight model
Knowledge cutoff notes

No direct authoritative knowledge-cutoff date was published for the exact SenseNova-MARS-32B checkpoint. External search tools can provide newer information during use but do not change the model's underlying knowledge cutoff.

Model notes

SenseNova-MARS-32B is an open-weight MIT-licensed checkpoint available from the official SenseNova Hugging Face organization. It is designed for multimodal agentic reasoning and can be connected to text-search, image-search, and image-cropping tools. Tool execution requires surrounding infrastructure and is not equivalent to a built-in provider-hosted web-search API. The published model configuration specifies a 262,144-token text context length. No official exact-model knowledge cutoff, maximum output-token limit, token pricing, deprecation date, or shutdown date was identified. The official repository provides deployment examples for Transformers, vLLM, and SGLang. Editorial scores reflect the model's documented visual reasoning and tool-use positioning; they are not provider-issued ratings.

Model guide

SenseNova-MARS-32B: Open-Weight Model for Visual Deep Search and Agentic Reasoning

SenseNova-MARS-32B is SenseTime’s open-weight, MIT-licensed multimodal reasoning model for visual question answering, fine-grained image understanding, and multi-step workflows involving text search, image search, and image cropping.

What is SenseNova-MARS-32B?

SenseNova-MARS-32B is an open-weight multimodal vision-language model from SenseTime. In practical terms, it is built to interpret images and reason over information gathered during several steps. A conventional vision-language model might answer a question from the pixels it receives. SenseNova-MARS-32B is intended for harder cases in which the system must identify a visual detail, look for supporting information, inspect another image, and then combine the results into a final textual answer.

The model is part of SenseTime’s SenseNova-MARS project and is published through the official OpenSenseNova repositories. The exact checkpoint is available on Hugging Face as sensenova/SenseNova-MARS-32B. Its repository is marked with the MIT license, and the published materials include deployment paths for Hugging Face Transformers, vLLM, and SGLang.

“Open-weight” means that the model weights can be downloaded and operated by users who have suitable infrastructure. It does not mean that SenseTime provides a free hosted API, managed search service, or consumer application for this exact checkpoint. Those parts must be supplied by the deployment environment.

Purpose and position in the SenseNova lineup

SenseNova-MARS-32B occupies a specialized position in SenseTime’s current model catalog. It is not presented as a general consumer chatbot or as an image, video, audio, or speech generation model. Its focus is agentic visual reasoning: using a vision-language model in a loop with tools and external information sources.

The model is especially relevant to developers and researchers building visual research systems. For example, an application could ask it to identify a logo in a photograph, crop the relevant area, search for the logo or organization, perform a text search for additional context, and produce an evidence-based answer. The model’s role is to interpret inputs, decide how the workflow should proceed, and generate the textual result. Search execution and image-cropping operations require surrounding software.

SenseTime describes the project as using reinforcement learning to improve interleaved visual reasoning and tool invocation. The project documentation introduces Batch-Normalized Group Sequence Policy Optimization, or BN-GSPO, as part of its approach to stabilizing training for multi-step tool-using reasoning. These are descriptions of the project’s technical approach, not guarantees that every deployment will achieve the same behavior without appropriate tool definitions and prompting.

Supported inputs and outputs

SenseNova-MARS-32B accepts text and visual information. The published examples demonstrate image-text input, while the configuration includes image and video processing settings. The model’s primary output is text: answers, reasoning content, and tool-oriented instructions.

CapabilityStatusPractical meaning
Text inputSupportedUsers can provide questions, instructions, and search-related tasks.
Image inputSupportedThe model can analyze photographs, documents, objects, scenes, and fine visual details.
Video inputConfiguration indicates video processingVideo-related handling is present in the published configuration, but the supplied materials primarily demonstrate image-text usage.
Text outputSupportedThe model returns textual answers and can produce tool-oriented instructions.
Image, video, audio, or speech outputNot supported nativelyThe checkpoint is not a media-generation model.

Multimodal input should not be confused with multimodal output. SenseNova-MARS-32B can receive visual information, but it does not natively create images, videos, audio, music, or speech.

Reasoning and tool use

The model’s main distinction is its support for multi-step, tool-assisted visual reasoning. The SenseNova-MARS project identifies three important tool categories: text search, image search, and image cropping. These tools can be combined in a sequence rather than used as isolated features.

A useful workflow might begin with a photograph containing a small or partially obscured object. The model can identify the likely object, request a crop of the relevant area, use image search to find visually similar references, and use text search to retrieve background information. It can then combine the observations and retrieved material in a final response. This makes the checkpoint suitable for visual deep-search systems, research assistants, document investigation, and image-grounded question answering.

However, tool use is not the same as built-in web access. The downloaded weights do not independently provide a search index, browser, image-search provider, or current-information service. A developer must connect compatible tools, define their interfaces, execute the calls, and return the results to the model. The model therefore should not be described as having a first-party web-search API or automatically current knowledge.

Context length and deployment requirements

The published configuration specifies a maximum text position length of 262,144 tokens. This is a substantial context window for combining long instructions, retrieved text, visual-analysis steps, and intermediate tool results. It should not be interpreted as a guaranteed maximum for every type of image or video input, because practical limits also depend on the serving framework, image processing configuration, available memory, and the surrounding application.

No provider-published maximum output-token limit was identified for the exact SenseNova-MARS-32B checkpoint. The same is true for an exact knowledge-cutoff date. External search can supply newer information during a workflow, but it does not change the model’s underlying training knowledge.

The published safetensors model repository is approximately 66.7 GB. As a result, local operation requires substantial GPU memory or distributed inference. The project documentation also reports demanding multi-GPU H100 configurations for training and evaluation infrastructure. Inference requirements vary with quantization, batching, sequence length, image processing, and serving configuration, but the model is clearly aimed more at organizations with dedicated hardware than at ordinary consumer laptops or phones.

Reported benchmark performance

SenseTime reports a score of 74.3 on MMSearch and 54.4 on HR-MMSearch for SenseNova-MARS-32B. The model was evaluated across search-oriented and fine-grained visual-understanding benchmarks including MMSearch, HR-MMSearch, FVQA, InfoSeek, SimpleVQA, and LiveVQA. SenseTime’s announcement reports an average score of 69.74 across its highlighted multimodal search and reasoning benchmarks.

These figures are provider-reported results, not independent guarantees. Results can depend on prompts, tool availability, search configuration, image preprocessing, evaluation protocols, and infrastructure. They indicate that the model is targeted at visual search and reasoning tasks; they do not establish that it is the best choice for every language, coding, generation, or low-latency workload.

Strengths and trade-offs

  • Visual deep-search focus: The model is designed for tasks where visual interpretation and external information retrieval must work together.
  • Fine-grained image analysis: Image cropping and image-search workflows are useful when important evidence occupies only a small part of a larger image.
  • Open-weight deployment: Organizations can download the checkpoint and control the serving environment instead of depending on a single hosted endpoint.
  • Long published context: The 262,144-token text position length can support substantial retrieved material and multi-step interaction histories.
  • Flexible serving options: Official examples cover Transformers, vLLM, and SGLang.
  • Infrastructure cost: The model’s size and hardware requirements make it slower and more expensive to operate than smaller models, especially when running long contexts or multiple tool calls.
  • Application complexity: Search and cropping require an orchestration layer. The model weights alone do not deliver a complete research assistant.

Editorially, its reasoning and tool-use positioning are stronger than its suitability for fast, inexpensive general-purpose inference. The available research assigns a reasoning score of 8, coding score of 6, speed score of 4, and cost score of 7; these are editorial evaluations, not ratings published by SenseTime. Coding is not the model’s primary specialty, and no supplied benchmark establishes it as a leading coding model.

Pricing and availability

No provider-hosted input or output pricing was identified for the exact SenseNova-MARS-32B checkpoint. It is available as an open-weight model through the official SenseNova GitHub and Hugging Face repositories rather than as a documented public token-priced API. The direct financial cost therefore depends on hardware, hosting, electricity, storage, engineering, and any third-party inference service used.

The model was released on January 29, 2026. The supplied research identifies it as a current open-weight model and found no official deprecation or shutdown date. Its MIT license permits broad use subject to the applicable license terms and the responsibilities of operating the model and external tools.

Best use cases

  • Visual research systems that combine images with text and external search.
  • Fine-grained analysis of photographs, products, logos, documents, and complex scenes.
  • Multi-step visual question answering requiring image crops and retrieved evidence.
  • Self-hosted enterprise or academic applications where control over model weights and data handling is important.
  • Research into reinforcement learning, agentic reasoning, and tool-using multimodal models.
  • Applications that can provide their own text-search, image-search, and image-processing infrastructure.

When to choose SenseNova-MARS-32B

Choose SenseNova-MARS-32B when the central problem involves visual evidence, detailed image interpretation, and a sequence of external tool calls. It is particularly appropriate when a team can operate substantial GPU infrastructure and wants an open-weight model that can be integrated into a custom visual-search pipeline.

A smaller or hosted model may be more appropriate when low latency, simple deployment, predictable per-request pricing, or ordinary text chat is more important than deep visual investigation. A dedicated image-generation model is a better choice for creating images, while a speech or audio model is needed for native voice output. Similarly, an application requiring a managed web-search service should not assume that SenseNova-MARS-32B supplies one automatically.

Its strongest practical trade-off is capability versus operational cost: the model is designed to reason through complex visual-search tasks, but that specialization brings a large download, demanding hardware, slower likely inference, and the need to build or maintain the tool layer. For teams prepared to manage those requirements, it offers a focused open-weight foundation for visual agent systems. For users seeking a ready-to-use assistant or a simple API, another type of model or service may be a better fit.

Important limitations and unknowns

The supplied official materials do not specify an exact knowledge cutoff, maximum output-token limit, hosted API price, native structured-output guarantee, fine-tuning support, or first-party web-search API for this checkpoint. These values should remain unknown rather than being inferred from the model’s context length, repository examples, or connected tools.

Benchmark results should also be read with care because the model’s reported strengths depend on evaluation setup and tool infrastructure. The model can be connected to current search systems, but retrieved information can be incomplete, inaccurate, or poorly matched to the image. Human review and application-level safeguards remain important for high-impact research and decision-making.


Answers to Frequently Asked Questions

Where can I download SenseNova-MARS-32B and what does it cost?
The exact checkpoint is available on Hugging Face as sensenova/SenseNova-MARS-32B and is published through the official OpenSenseNova repositories under the MIT license. No provider-hosted token pricing was identified for this checkpoint; operating costs depend on hardware, hosting, electricity, storage, engineering, or any third-party inference service.
How much hardware is needed to run SenseNova-MARS-32B?
The published safetensors repository is approximately 66.7 GB, so local deployment requires substantial GPU memory or distributed inference. Actual requirements vary with quantization, batching, context length, image processing, and serving framework. The model is generally better suited to organizations with dedicated GPU infrastructure than to consumer laptops or phones.
What inputs and outputs does SenseNova-MARS-32B support?
The model supports text and visual inputs, with published examples focused on image-text tasks. Its primary output is text, including answers, reasoning content, and tool-oriented instructions. It does not natively generate images, video, audio, music, or speech.
What is SenseNova-MARS-32B designed for?
SenseNova-MARS-32B is an open-weight multimodal vision-language model designed for agentic visual reasoning. It can analyze images, use connected text search, image search, and image-cropping tools, and combine the gathered information into a textual answer.
Does SenseNova-MARS-32B have built-in web search?
No. The model weights do not include a browser, search index, image-search provider, or hosted web-search API. Developers must supply and integrate compatible text-search, image-search, and image-processing tools.


Sources 4
Provider

About SenseTime