MiMo-Audio-7B-Base
Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation
Current open-weight downloadable model
Browse the AI models associated with Xiaomi HyperAI. Compare current and historical models by family, capabilities, context window, availability and intended use.
Few-shot audio-language research, speech continuation, voice and style conversion, speech translation, speech editing, and audio-text experimentation
Current open-weight downloadable model
Local audio understanding, speech-to-text dialogue, spoken conversational agents, and controllable text-to-speech research
Current open-weight model; locally downloadable and usable
Embodied AI research, autonomous-driving perception and planning, spatial reasoning, affordance prediction, robot navigation, and multimodal visual analysis
Active open-weight model
Research, custom post-training, foundation-model experimentation, and self-hosted long-context inference
Current open-weight base model
Speech-to-text transcription for Chinese, English, regional Chinese dialects, code-switched speech, lyrics, noisy recordings, and multi-speaker conversations
Current and accessible through the Xiaomi MiMo API; open-source code and weights available
Expressive text-to-speech, audiobooks, podcasts, dubbing, character dialogue, voice interfaces, narrated content, and stylized speech or singing.
Available
Zero-shot voice cloning, expressive narration, character dialogue, personalized speech, and custom-voice audio production
Current; limited-time free access
Custom synthetic voices for narration, characters, podcasts, ASMR, games, assistants, and creative audio production
Current; limited-time free access
Local coding agents, terminal automation, cybersecurity experimentation, visual coding and agentic reinforcement-learning research
Current open-weight SFT checkpoint
High-volume multimodal API workloads, coding assistants, agent automation, long-context document and repository analysis, tool-using workflows, and cost-sensitive professional applications.
Current; open-weight and available through the Xiaomi MiMo API, MiMo Studio, MiMo Desktop, MiMo Code, and other supported integrations.
Long-horizon agents, coding, cybersecurity, research, computer use, multimodal analysis, and complex multi-step workflows
Current; open-source weights and API access available
Latency-sensitive multimodal reasoning, coding, tool-use, long-context research, and interactive agent workflows
Current; high-speed inference variant of MiMo-V2.6-Pro
Local image and video understanding, multimodal reasoning, visual question answering, OCR-oriented tasks, visual grounding, and GUI analysis
Current open-weight model; publicly available for download and local deployment
Self-hosted image and video understanding, multimodal reasoning research, supervised fine-tuning, and reinforcement-learning experimentation
Current open-weight checkpoint; available for download
Long-horizon agent workflows, repository-scale coding, complex software engineering, tool-driven automation, and very large documents
Deprecated; API model name scheduled to expire on 2026-10-21 at 10:00 Beijing Time