SenseNova-Vision

SenseNova-Vision-7B-MoT

by SenseTime · Current open-weight research model

SenseNova-Vision-7B-MoT is SenseTime’s open-weight 7B research model for unified computer vision. It turns text and image instructions into text, image, or mixed outputs for detection, OCR, grounding, segmentation, depth, surface-normal estimation, keypoints, and multi-view geometry. Its main trade-offs are high-memory deployment requirements, no documented hosted pricing or production API, and no published context or maximum-output limits.

Text Image generation Reasoning Coding
SenseNova-Vision-7B-MoT is SenseTime’s open-weight unified vision model for turning many computer-vision tasks into one multimodal generation workflow. It can produce both symbolic results, such as OCR text and bounding boxes, and dense visual outputs, such as segmentation masks, depth maps, and surface normals. That breadth makes it relevant to researchers and developers building vision systems, although its high-memory hardware requirements and lack of documented hosted API limits make it less suitable for casual or low-cost deployment.
Outputs

What SenseNova-Vision-7B-MoT can produce

Text Image generation
Inputs

What it can understand

Text Images Multimodal input
Capabilities

Supported features

Fine-tuning Structured output Multimodal output
Model profile

Performance characteristics

7/10 Reasoning
4/10 Coding
4/10 Speed
7/10 Cost efficiency
Specifications

Technical details

Model family SenseNova-Vision
Model type Multimodal
Release date 2026-07-08
Status Current open-weight research model
Knowledge cutoff notes

No authoritative knowledge-cutoff date is published for the exact SenseNova-Vision-7B-MoT checkpoint.

Model notes

The canonical released checkpoint is SenseNova-Vision-7B-MoT, while SenseNova-Vision is the project/model family name. The model produces text, image, and mixed text-image outputs for structured visual understanding and dense spatial prediction. The official repository provides inference, benchmarking, and training workflows. Its demo documentation reports approximately 38 GB peak process GPU memory on one NVIDIA A800 80GB GPU and recommends an 80GB GPU. The GitHub repository states an Apache-2.0 license, but the Hugging Face checkpoint page currently displays CC-BY-NC-4.0; users should verify the license applicable to the exact artifact. No official hosted token pricing, context limit, maximum output-token limit, or managed API specification was found.

Model guide

SenseNova-Vision-7B-MoT: A Unified Model for Detection, Segmentation and 3D Vision

SenseNova-Vision-7B-MoT is SenseTime’s open-weight 7B mixture-of-transformers model for unified computer vision. Instead of using separate task-specific prediction heads, it treats visual understanding and dense spatial prediction as multimodal generation. The model can accept text and images and return text, images, or mixed responses for tasks including object detection, OCR, visual grounding, keypoint estimation, segmentation, depth and surface-normal prediction, camera-pose estimation, and multi-view reconstruction. It is best suited to research and self-hosted experimentation, not conventional managed API workloads, because official materials do not document hosted pricing, context limits, maximum output tokens, or a production API, and the recommended deployment hardware is an NVIDIA A800 80GB GPU.

What is SenseNova-Vision-7B-MoT?

SenseNova-Vision-7B-MoT is SenseTime’s canonical released checkpoint for the SenseNova-Vision project. It is an open-weight, 7-billion-parameter mixture-of-transformers model designed for a broad set of computer-vision tasks. The model family’s central idea is to use one unified multimodal generation interface instead of building a separate prediction head or decoder for every task.

In practical terms, a user supplies an image together with a natural-language instruction. The model then generates an answer in text, image, or a combination of both. For example, an instruction may ask it to identify an object and return its bounding box, transcribe text in a photograph, produce a segmentation mask, estimate scene depth, or infer camera information from multiple views.

This positioning is different from a conventional image-chat model focused mainly on describing pictures. SenseNova-Vision-7B-MoT is intended as a research-oriented computer-vision system that represents many visual tasks through a shared generation framework.

Where it fits in SenseTime’s lineup

The model belongs to SenseTime’s SenseNova and OpenSenseNova ecosystem. SenseNova is a broader family of foundation models and AI products, while SenseNova-Vision is the project focused specifically on unified computer vision. The 7B-MoT checkpoint is therefore a specialized research model rather than a general-purpose consumer assistant or a standard hosted language-model endpoint.

SenseTime’s OpenSenseNova project provides the model weights, inference code, benchmark tooling, and training workflow. This makes the checkpoint more relevant to teams that need to inspect, adapt, benchmark, or self-host a model than to users looking for a simple web application.

What the model can do

SenseNova-Vision-7B-MoT supports both visual understanding and structured visual prediction. Its documented capabilities include:

  • Object detection: locating objects with bounding boxes, including referring detection based on a textual description.
  • OCR and visual grounding: reading text in images and associating language with particular visual regions.
  • Keypoint estimation: predicting points or keypoints associated with objects or scenes.
  • Segmentation: semantic, panoptic, referring, reasoning, and interactive segmentation.
  • Depth estimation: generating a spatial representation of relative or scene depth.
  • Surface-normal estimation: predicting the orientation of visible surfaces.
  • Multi-view geometry: supporting multi-view reconstruction and camera-pose estimation.

Some results are returned as text. Bounding boxes, categories, OCR strings, keypoints, and camera parameters can be represented symbolically in the generated response. Other results are returned as images. Segmentation masks, depth maps, surface normals, and multi-view point maps are examples of dense spatial outputs that can be represented visually.

How the unified architecture works

Many computer-vision systems use a dedicated architecture for each task: one head for classification, another for detection, and separate decoders for segmentation or depth. SenseNova-Vision takes a different approach. Its training process expresses heterogeneous vision annotations as instruction-and-response examples in native text and image generation spaces.

This means the model learns to interpret a task description and produce the corresponding output format through a common multimodal interface. The project states that it does not require task-specific prediction heads, decoders, or architectural branches. The model is trained on the SenseNova-Vision Corpus, together with auxiliary multimodal data intended to preserve general visual understanding and generation abilities.

For users, the benefit is conceptual and engineering consistency: different vision tasks can be addressed through related instruction-based workflows. However, a unified interface does not automatically mean that every task will match the accuracy, latency, or memory profile of a narrowly specialized model. The supplied research does not provide a complete benchmark comparison against dedicated detectors, OCR systems, or segmentation models.

Supported inputs and outputs

The documented input modalities are text and images. Audio and video input are not listed as supported for this checkpoint. The model’s output modalities are text, images, and mixed text-image responses, which is why it can combine symbolic answers with dense visual predictions.

AreaSupported or documented behavior
Text inputYes; natural-language instructions are used to specify visual tasks.
Image inputYes; images are the model’s primary visual input.
Audio inputNo documented support.
Video inputNo documented support for the checkpoint.
Text outputYes; includes OCR, labels, boxes, points, keypoints, and camera parameters.
Image outputYes; includes masks, depth, surface normals, and other spatial maps.
Mixed outputYes; text and visual results can be combined.

The model is described as supporting structured visual outputs, but the supplied materials do not document a formal JSON mode or a general-purpose function-calling interface. Structured results may therefore depend on the project’s prescribed prompts, output representations, and inference code rather than on a separately documented managed API feature.

Reasoning, coding and tool support

SenseNova-Vision-7B-MoT has meaningful visual reasoning capabilities in the task-specific sense: it can follow instructions for referring detection, reasoning segmentation, geometric inference, and other operations that require connecting language with image content. An editorial reasoning score of 7 out of 10 reflects this broad task coverage; it is not a score published by SenseTime.

The model is not presented as a coding model. An editorial coding score of 4 out of 10 reflects that coding is not its primary purpose. It may be useful to researchers who write surrounding inference or data-processing code, but the checkpoint itself should not be selected as a general software-development assistant.

No documented tool use, web search, browser access, or external function-calling capability is supplied. It should be treated as a model for processing provided visual inputs, not as an autonomous agent that gathers information or operates external services.

Deployment, hardware and speed

The official inference guide recommends an NVIDIA A800 80GB GPU for the default web demo. The research notes approximately 38 GB of peak process GPU memory in one demo configuration on an A800 and recommends an 80GB GPU. Full benchmarking and training require substantially larger resources.

These requirements are important when comparing the model with smaller, specialized, or hosted vision systems. A 7B parameter count may appear moderate compared with very large multimodal models, but the complete inference workflow, visual generation behavior, and implementation requirements can still make local deployment demanding. Users without high-memory GPUs may need a hosted alternative, a smaller task-specific model, or a remote inference environment.

The supplied material does not provide a standardized latency benchmark. An editorial speed score of 4 out of 10 and cost score of 7 out of 10 are evaluations rather than provider-published measurements. The cost score reflects the absence of hosted token pricing, not a guaranteed estimate of total ownership cost. Hardware rental, engineering time, storage, and model adaptation can all affect the real expense of self-hosting.

Pricing, API access and undocumented limits

No official hosted token pricing was identified for SenseNova-Vision-7B-MoT. The model is distributed as an open-weight research checkpoint, so users should distinguish access to the weights from access to a managed inference service. The supplied research also does not identify a separate production API for this exact model.

Several operational limits remain undocumented for the checkpoint. These include the context window, maximum output-token limit, and a formal hosted request quota. The absence of a published value does not mean that the model has unlimited capacity; it means that users must consult the current repository, inference configuration, and checkpoint documentation before designing a workload.

The model supports fine-tuning or training workflows according to the project materials, but the practical requirements are substantial. The training pipeline is intended for research and engineering teams with suitable data, compute, and evaluation processes rather than for casual customization.

Main strengths and limitations

Strengths

  • It unifies a notably wide range of vision tasks under an instruction-based multimodal interface.
  • It can return both symbolic text results and dense image-based predictions.
  • Its coverage extends beyond basic image description to detection, OCR, segmentation, depth, surface normals, and multi-view geometry.
  • Open weights and public inference, benchmarking, and training workflows provide more control than a closed hosted endpoint.
  • Its architecture is designed to avoid separate task-specific prediction heads and decoders.

Limitations

  • Local inference has high memory requirements, with an 80GB GPU recommended for the documented demo configuration.
  • No official hosted price, production API specification, context window, or maximum output-token limit is documented in the supplied research.
  • Audio and video input are not documented for this checkpoint, and the model is not intended for audio or video generation.
  • There is no documented web search, tool use, or general function-calling layer.
  • License information requires care: the project repository states Apache-2.0, while the Hugging Face checkpoint page currently displays CC-BY-NC-4.0. Users should verify the license attached to the exact artifact before commercial deployment.

When to choose SenseNova-Vision-7B-MoT

Choose this model when the project needs a unified research platform for multiple computer-vision tasks and the team can operate high-memory GPU infrastructure. It is particularly relevant for experiments that combine detection, OCR, grounding, segmentation, depth, keypoints, or multi-view geometry without maintaining a completely separate model for each operation.

It may also be a good fit when open weights, access to inference code, and the ability to inspect or modify the training workflow matter more than turnkey deployment. Researchers studying unified multimodal generation can use it as a concrete system for testing whether different visual annotations can be represented through a shared generation framework.

Another option may be more appropriate for low-latency production applications, constrained hardware, or predictable usage billing. A dedicated detection, OCR, or segmentation model may be faster and easier to optimize when only one task is required. A managed multimodal API may be preferable when the team needs published quotas, stable service-level expectations, automatic scaling, or a documented per-request price. A general-purpose language model is also a better choice for coding, web research, and tool-driven workflows.

Overall assessment

SenseNova-Vision-7B-MoT is best understood as an open research checkpoint for unified computer vision, not as a drop-in replacement for every vision API. Its distinctive value is the breadth of tasks expressed through one multimodal generation approach: the same model family can produce text-based detections and OCR results as well as image-based masks, depth maps, surface normals, and geometric outputs.

That breadth comes with practical trade-offs. The model requires substantial hardware, has limited documented service information, and is not positioned as a general assistant with audio, video, web, or tool capabilities. For researchers and experienced vision engineers, those constraints may be acceptable in exchange for open access and a flexible task interface. For users seeking inexpensive, managed, single-purpose inference, a specialized or hosted alternative is likely to be easier to operate.


Answers to Frequently Asked Questions

What are the main limitations of SenseNova-Vision-7B-MoT?
The model has substantial hardware requirements and is not presented as a general coding, web-search, tool-use, or function-calling assistant. Audio and video support are not documented, and benchmark comparisons with dedicated detection, OCR, or segmentation systems are incomplete. License information should also be verified for the exact artifact because the repository states Apache-2.0 while the Hugging Face checkpoint page displays CC-BY-NC-4.0.
Does SenseNova-Vision-7B-MoT have an official API or hosted pricing?
No official hosted token pricing or separate production API for this exact model was identified in the supplied research. It is distributed as an open-weight research checkpoint with inference code and workflows. The context window, maximum output-token limit, and formal hosted request quotas are also not documented.
What hardware is required to run SenseNova-Vision-7B-MoT?
The official inference guide recommends an NVIDIA A800 80GB GPU for the default web demo. One documented configuration used approximately 38 GB of peak process GPU memory on an A800, so local deployment can be demanding. Users without high-memory GPUs may need a hosted service, remote inference environment, or smaller specialized model.
What is SenseNova-Vision-7B-MoT designed to do?
SenseNova-Vision-7B-MoT is an open-weight, 7-billion-parameter mixture-of-transformers model for unified computer vision. It can process images and natural-language instructions to perform tasks such as object detection, OCR, segmentation, depth estimation, surface-normal prediction, keypoint estimation, and multi-view geometry.
What inputs and outputs does SenseNova-Vision-7B-MoT support?
The checkpoint supports text and image inputs. It can generate text outputs such as OCR results, labels, bounding boxes, keypoints, and camera parameters, as well as image outputs such as segmentation masks, depth maps, surface normals, and other spatial representations. Audio and video input are not documented for this checkpoint.


Sources 6
Provider

About SenseTime