Model catalog

Xiaomi HyperAI Models

Browse the AI models associated with Xiaomi HyperAI. Compare current and historical models by family, capabilities, context window, availability and intended use.

15 models tracked
15 Total models
8 Model families
4 Model types
15 Current / accessible
All models

Xiaomi HyperAI model catalog

Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation

Type Multimodal
Context 8K
Reasoning 5/10
Speed 3/10
Multimodal Audio input Media output
Status

Current open-weight downloadable model

View model →

Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research

Type Multimodal
Context 8K
Reasoning 7/10
Speed 4/10
Multimodal Audio input Media output
Status

Current open-weight model; locally downloadable and usable

Input No official hosted API price published; self-hosted weights
Output No official hosted API price published; self-hosted weights
View model →
Xiaomi HyperAI logo
MiMo-Embodied

MiMo-Embodied-7B

Embodied AI research, autonomous-driving perception and planning, spatial reasoning, affordance prediction, robot navigation, and multimodal visual analysis

Type Multimodal
Context 128K
Reasoning 7/10
Speed 6/10
Multimodal Image input Video input
Status

Active open-weight model

View model →
Xiaomi HyperAI logo
MiMo-V2-Flash

MiMo-V2-Flash-Base

Research, custom post-training, foundation-model experimentation, and self-hosted long-context inference

Type General Purpose
Context 256K
Reasoning 8/10
Speed 9/10
Status

Current open-weight base model

View model →
Xiaomi HyperAI logo
MiMo-V2.5

MiMo-V2.5-ASR

Speech-to-text transcription for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, and multi-speaker conversations

Type Other
Reasoning 3/10
Speed 7/10
Audio input Streaming
Status

Current and accessible through the Xiaomi MiMo API; open-source code and weights available

Input $0.074 per hour of input audio for overseas API usage; ¥0.5 per hour on the domestic pricing page
Output No separate output charge documented; billing is based on input-audio duration
View model →
Xiaomi HyperAI logo
MiMo-V2.5-TTS

MiMo-V2.5-TTS

Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.

Type Other
Context 8K
Reasoning 1/10
Speed 7/10
Media output Streaming
Status

Available

Input Free for a limited time
Output Free for a limited time
View model →

Zero-shot voice cloning, expressive narration, character dialogue, personalized speech, and custom-voice audio production

Type Other
Context 8K
Speed 7/10
Multimodal Audio input Media output
Status

Current; limited-time free access

Input Free for a limited time
Output Free for a limited time
View model →

Custom synthetic voices for narration, characters, podcasts, ASMR, games, assistants, and creative audio production

Type Other
Context 8K
Speed 6/10
Media output Streaming
Status

Current; limited-time free access

Input Free for a limited time
Output Free for a limited time
View model →

Local coding agents, terminal automation, cybersecurity experimentation, visual coding and agentic reinforcement-learning research

Type Multimodal
Context 262K
Reasoning 7/10
Speed 7/10
Multimodal Image input Tool use
Status

Current open-weight SFT checkpoint

Input No official hosted API price; open-weight download
Output No official hosted API price; open-weight download
View model →

High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.

Type Multimodal
Context 1M
Reasoning 8/10
Speed 9/10
Multimodal Image input Audio input
Status

Current; open-weight and available through the Xiaomi MiMo API, MiMo Studio, MiMo Desktop, MiMo Code, and other supported integrations.

Input $0.14 per million tokens for cache-miss input; $0.0028 per million tokens for cache-hit input.
Output $0.28 per million tokens.
View model →
Xiaomi HyperAI logo
MiMo-V2.6

MiMo-V2.6-Pro

Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows

Type Multimodal
Context 1M
Reasoning 9/10
Speed 7/10
Multimodal Image input Audio input
Status

Current; open-source weights and API access available

Input $0.435 per 1M tokens cache miss; $0.0036 per 1M tokens cache hit; batch input $0.2175 per 1M tokens
Output $0.87 per 1M tokens; batch output $0.435 per 1M tokens
View model →

Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows

Type Reasoning
Context 1M
Reasoning 9/10
Speed 10/10
Multimodal Image input Audio input
Status

Current; high-speed inference variant of MiMo-V2.6-Pro

Input $0.036 per million cached-input tokens; $4.35 per million uncached-input tokens
Output $8.70 per million output tokens
View model →

Local image and video understanding, multimodal reasoning, visual question answering, OCR-oriented tasks, visual grounding, and GUI analysis

Type Multimodal
Context 128K
Reasoning 7/10
Speed 7/10
Multimodal Image input Video input
Status

Current open-weight model; publicly available for download and local deployment

Input No official first-party hosted API pricing found; open-weight download
Output No official first-party hosted API pricing found; open-weight download
View model →

Self-hosted image and video understanding, multimodal reasoning research, supervised fine-tuning, and reinforcement-learning experimentation

Type Multimodal
Context 128K
Reasoning 7/10
Speed 6/10
Multimodal Image input Video input
Status

Current open-weight checkpoint; available for download

Input No official hosted API pricing; downloadable weights are available under the MIT license
Output No official hosted API pricing; downloadable weights are available under the MIT license
View model →
Xiaomi HyperAI logo
MiMo V2.5

MiMo-V2.5-Pro

Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents

Type Reasoning
Context 1M
Reasoning 9/10
Speed 7/10
Tool use Web search Fine-tuning
Status

Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time

Input ¥0.025 per million tokens cache hit; ¥3 per million tokens cache miss; equivalent USD prices $0.0036 and $0.435 per million tokens
Output ¥6 per million tokens; equivalent USD price $0.87 per million tokens
View model →