SenseNova-SI

SenseNova-SI-1.5-InternVL3-8B

by SenseTime · Current open-weight release

An open-weight 8B InternVL3-based multimodal model from SenseTime’s SenseNova-SI project, focused on image-text understanding, spatial relationships, 3D scene analysis, and solid-geometry reasoning. It supports local deployment but has no verified hosted per-token pricing, native tool layer, or non-text output.

Text Reasoning Coding
SenseNova-SI-1.5-InternVL3-8B is an open-weight multimodal model developed by SenseTime and released through its SenseNova-SI project. The model combines image understanding with text generation, with particular emphasis on spatial relationships, geometric structures, 3D scenes, and visual mathematics. It is intended mainly for local or self-hosted use through frameworks such as Transformers, vLLM, SGLang, and Docker-compatible workflows. There is no verified hosted per-token price for this checkpoint.
Outputs

What SenseNova-SI-1.5-InternVL3-8B can produce

Text
Inputs

What it can understand

Text Images Multimodal input
Model profile

Performance characteristics

7/10 Reasoning
5/10 Coding
6/10 Speed
8/10 Cost efficiency
Specifications

Technical details

Model family SenseNova-SI
Model type Multimodal
Context window 33K tokens
Release date 2026-04-01
Status Current open-weight release
Knowledge cutoff notes

No authoritative model-specific knowledge-cutoff date was published in the reviewed official model card, repository documentation, or configuration.

Model notes

The supplied topic name is a shortened reference. The verified canonical model is SenseNova-SI-1.5-InternVL3-8B. It is an InternVL3-based multimodal checkpoint released by SenseTime's SenseNova-SI project and distributed as open weights under the Apache 2.0 license. The project reports 63.5 on SolidGeo MCQ, 72.7 on SolidMath, 68.9 on Math3D, and an EASI-8 score of 64.4. Official materials provide local deployment examples with Transformers, vLLM, SGLang, and Docker. No official hosted input or output token pricing was identified.

Model guide

SenseNova-SI-1.5-InternVL3-8B: Open-Weight Model for Spatial and Solid-Geometry Reasoning

SenseNova-SI-1.5-InternVL3-8B is an open-weight, 8-billion-parameter multimodal model from SenseTime’s SenseNova-SI project. Built on InternVL3, it accepts text and images and is specialized for spatial intelligence, visual reasoning, 3D scene understanding, and solid-geometry problem solving rather than image, audio, video, or general tool-based generation.

What is SenseNova-SI-1.5-InternVL3-8B?

SenseNova-SI-1.5-InternVL3-8B is an open-weight multimodal language model from SenseTime’s SenseNova-SI research project. Its full name identifies the specific model checkpoint: it belongs to the SenseNova-SI family, uses the InternVL3 multimodal architecture, and contains approximately 8 billion parameters. “SenseNova-SI-1.5” is often used as a shortened name, but the longer identifier is more precise when selecting weights or configuring an inference system.

In practical terms, the model reads text and images and produces text. Its specialization is spatial intelligence: understanding where objects are relative to one another, interpreting visual scenes, analyzing geometric structures, and working through problems involving three-dimensional objects. It is therefore more focused than a general-purpose assistant designed equally for web research, coding agents, office tasks, and creative generation.

Provider and position in the SenseNova lineup

The model is provided by SenseTime through the SenseNova-SI project. SenseNova is SenseTime’s broader AI ecosystem, while SenseNova-SI is the project focused on spatial intelligence and related multimodal reasoning capabilities. This checkpoint is positioned as an open-weight research and deployment model rather than as a consumer chat plan or a managed commercial API model.

The model continues from SenseNova-SI-1.4-InternVL3-8B and is built around InternVL3. Its published configuration identifies an InternVL chat model with a vision encoder and a Qwen2-based language component. The weights are available through Hugging Face under the Apache 2.0 license, which supports local use and research-oriented integration subject to the license terms.

What the model can do

The model accepts both images and text as input. It then returns textual answers, explanations, or reasoning about the supplied content. Typical tasks include asking questions about an image, identifying spatial relationships, interpreting a 3D diagram, and solving a geometry problem presented visually.

  • Image-text question answering
  • Visual question answering
  • Relative-position and spatial-relation analysis
  • 3D scene understanding
  • Solid-geometry reasoning
  • Visual analysis of mathematical or geometric problems

For example, a user could provide a diagram of a solid and ask which face is opposite another face, or provide a scene and ask which object is nearest to a reference object. The model’s value in these cases comes from connecting visual information with a written explanation, rather than merely describing pixels or producing a short label.

Its output is text only. The model does not natively generate images, audio, video, music, embeddings, or speech. The supplied documentation also does not establish first-party support for web search, function calling, structured-output schemas, fine-tuning services, caching, or batch APIs.

Spatial and geometry reasoning performance

SenseNova-SI project documentation reports a score of 63.5 on SolidGeo MCQ, 72.7 on SolidMath, and 68.9 on Math3D. The project also reports an EASI-8 score of 64.4. These are provider or project-reported benchmark results, and they should be understood in the context of the model’s stated specialization.

The results suggest that the model’s main improvement over earlier versions is stronger analysis and solving of solid-geometry problems. They do not establish that it is the best model for general language, coding, long-form writing, agentic workflows, or real-world visual reliability. Benchmark performance can also vary with prompt format, image resolution, inference settings, and evaluation implementation.

Editorially, the model can be viewed as a focused spatial-reasoning checkpoint: its reasoning score is assessed here as relatively strong for its intended use, while its coding ability is more limited than its visual-geometry specialization. Those assessments are editorial evaluations, not provider-published ratings.

Context and technical specifications

SpecificationVerified detail
Model familySenseNova-SI
ArchitectureInternVL3-based multimodal model with a vision encoder and Qwen2-based language component
Parameter countApproximately 8 billion
InputText and images
OutputText
Maximum position length32,768 tokens
Published numerical formatbfloat16 configuration
LicenseApache 2.0
DeploymentLocal or self-hosted inference; examples include Transformers, vLLM, SGLang, and Docker-compatible workflows

The 32,768-token maximum position length describes the published configuration’s context capacity. It is not a guarantee that every deployment will process an image, prompt, and generated answer at the same practical limit. Image resolution, visual-token processing, batching, quantization, and the selected inference framework can affect memory use and throughput.

No maximum output-token value is specified in the supplied research. Users should therefore avoid assuming a particular response-length limit beyond the constraints imposed by the model configuration and the chosen serving stack.

Pricing and access

The checkpoint is distributed as open weights, so no official hosted input-token or output-token price was identified. This is different from a model sold through a metered API: the software and weights may be available for download, but running the model still requires suitable computing resources and an inference environment.

For local use, the effective cost depends on GPU availability, memory requirements, quantization, image resolution, batching, electricity, and infrastructure management. A self-hosted deployment can be attractive for research or controlled workloads because it avoids a per-token provider bill, but it transfers operational responsibility to the user. The supplied materials do not verify a first-party hosted endpoint with published per-token pricing for this specific checkpoint.

Speed, cost, and deployment trade-offs

At approximately 8 billion parameters, the model is smaller than many high-end multimodal systems, which can make local experimentation more practical than deploying a much larger checkpoint. However, the model’s actual speed cannot be inferred from parameter count alone. Vision processing, input image size, precision, quantization, GPU memory, batching, and serving software all affect latency.

The model is best treated as a specialized local inference option. It may be a sensible choice when the workload requires control over model files, repeatable offline evaluation, or adaptation within an existing multimodal pipeline. A hosted general-purpose model may be more convenient when the priority is immediate access, automatic scaling, managed reliability, or built-in tools. Conversely, a larger general model may be preferable when spatial reasoning is only one part of a broader workload and coding, instruction following, or agentic tool use matters more.

Tool use, coding, and general-purpose limitations

The available research does not verify native function calling, web browsing, code execution, or structured JSON output. It can produce text that resembles code or a structured answer, but that should not be confused with a guaranteed schema-enforcement or tool-execution feature. Any external tools would need to be implemented by the surrounding application and connected through an orchestration layer.

Coding is not the model’s primary design goal. It may help explain a geometric algorithm, interpret a visual programming problem, or generate straightforward text-based code, but the supplied research does not provide evidence of high-end software-engineering performance. Developers seeking a managed coding agent, repository operations, web-grounded answers, or reliable function invocation should consider an option designed and documented for those workflows.

Best use cases

  • Research into spatial intelligence and multimodal reasoning
  • Visual question answering involving relative positions or object relationships
  • 3D scene interpretation and geometry education
  • Solid-geometry and visual-mathematics experiments
  • Benchmarking multimodal models on spatial tasks
  • Local or self-hosted inference where open weights are important

It can be especially useful when the input is a diagram, scene, or geometric representation and the desired result is a written explanation. Researchers can also use it as a controlled checkpoint for studying how multimodal models represent spatial relationships.

When to choose this model

Choose SenseNova-SI-1.5-InternVL3-8B when spatial and geometric reasoning are central requirements, open-weight access matters, and you are prepared to operate a local or self-hosted inference stack. Its Apache 2.0 distribution, image-text interface, and documented focus on solid geometry make it a reasonable candidate for research, education, and specialized visual evaluation.

Choose another type of model when you need native image or video creation, audio interaction, web search, reliable tool calling, guaranteed structured outputs, a managed API, or broad general-purpose coding. A larger general multimodal model may offer wider capabilities, while a smaller or more optimized vision model may provide better speed for simple image classification or extraction. Those alternatives may sacrifice some of the spatial-reasoning focus that distinguishes this checkpoint.

Overall assessment

SenseNova-SI-1.5-InternVL3-8B is a specialized multimodal model rather than an all-purpose AI assistant. Its defining use case is turning visual and textual information into explanations involving spatial relationships, 3D scenes, and solid geometry. The reported benchmark results support that positioning, while the open-weight Apache 2.0 release makes the model suitable for local experimentation and research.

Its main limitations are equally important: there is no verified hosted pricing or managed API for this checkpoint, no documented native tool or structured-output layer, no native non-text generation, and no published maximum output-token value in the supplied materials. Users who match the model to spatial reasoning tasks may find those trade-offs acceptable; users seeking a complete production assistant should evaluate broader, managed alternatives.


Answers to Frequently Asked Questions

What are the limitations of SenseNova-SI-1.5-InternVL3-8B?
The model produces text only and does not natively generate images, audio, video, music, embeddings, or speech. The supplied documentation does not verify native web search, function calling, code execution, structured-output schemas, managed API access, or published per-token pricing. It is also not primarily designed for high-end software engineering or general-purpose assistant workflows.
How well does SenseNova-SI-1.5-InternVL3-8B perform on spatial and geometry benchmarks?
SenseNova-SI project documentation reports scores of 63.5 on SolidGeo MCQ, 72.7 on SolidMath, 68.9 on Math3D, and 64.4 on EASI-8. These provider-reported results support the model's focus on spatial and solid-geometry reasoning but do not establish broad superiority in general language, coding, or agentic tasks.
What license does SenseNova-SI-1.5-InternVL3-8B use, and how can it be deployed?
The model weights are available under the Apache 2.0 license. It is intended for local or self-hosted inference, with deployment options that include Transformers, vLLM, SGLang, and Docker-compatible workflows. Users remain responsible for the required hardware, memory, and infrastructure.
What is SenseNova-SI-1.5-InternVL3-8B designed for?
SenseNova-SI-1.5-InternVL3-8B is an open-weight multimodal model specialized in spatial intelligence. It accepts text and images and is designed for visual question answering, spatial-relation analysis, 3D scene understanding, solid-geometry reasoning, and visual mathematics.
What are the main technical specifications of SenseNova-SI-1.5-InternVL3-8B?
The model has approximately 8 billion parameters and uses an InternVL3-based architecture with a vision encoder and Qwen2-based language component. It accepts text and images, produces text, supports a published maximum position length of 32,768 tokens, and is configured for bfloat16.


Sources 4
Provider

About SenseTime