Kimi-Audio-7B
Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation
Current open-weight base model; not instruction-tuned
Browse the AI models associated with Moonshot AI. Compare current and historical models by family, capabilities, context window, availability and intended use.
Fine-tuning and research on speech recognition, audio understanding, audio classification, audio question answering, and speech-audio generation
Current open-weight base model; not instruction-tuned
Self-hosted speech recognition, audio understanding, audio question answering, audio captioning, and spoken conversational agents
Available as an open-weight self-hosted model
Repository-level software issue resolution, code repair, test generation, coding agents, and self-hosted software engineering workflows
Available open-weight model
Self-hosted image and video understanding, OCR, long-document analysis, visual question answering, screenshot perception, and efficient multimodal applications
Current open-weight/downloadable model; a newer Kimi-VL-A3B-Thinking-2506 variant is recommended for stronger multimodal reasoning
Local multimodal reasoning, mathematical visual question answering, OCR, document understanding, image and video analysis, and research on open-weight vision-language models
Available as downloadable open weights; superseded by Kimi-VL-A3B-Thinking-2506
Open-weight image and video reasoning, OCR, chart interpretation, visual mathematics, long PDFs, high-resolution screenshots and GUI-agent grounding
Current open-weight model; publicly available for self-hosted and third-party inference
Fine-tuning, foundation-model research, custom language systems, coding experiments, and self-hosted inference
Open-weight and downloadable; legacy relative to newer Kimi K2.x models but still accessible from the official Hugging Face repository
Open-weight coding assistants, tool-using agents, general-purpose chat, and self-hosted research deployments
Open-weight checkpoint available; legacy relative to Moonshot AI's current hosted API catalog
Self-hosted reasoning agents, autonomous research, long-horizon tool workflows, coding, and complex multi-step analysis
Retired from Moonshot's direct API; open-weight checkpoint remains available for self-hosted or third-party deployment
Long-horizon software engineering, agentic coding, visual document understanding, tool-using workflows, and multi-agent orchestration
Current; open-source model, available through Kimi, Kimi API, Kimi Code, and downloadable model weights
Long-horizon software engineering, repository-level coding, multi-file refactoring, debugging, coding agents, and tool-driven development workflows
Current and accessible through Kimi API; Kimi Code default service has moved to Kimi K2.8 Preview, while Kimi K2.7 Code remains available through API and Kimi K2.7 Code HighSpeed service.
Long-context software development, repository analysis, code completion, multi-file refactoring, and agentic coding workflows
Preview; fully rolled out in Kimi Code
Multimodal coding, visual debugging, long-context analysis, tool-using agents, and complex research or office workflows
current
Long-context coding, software engineering, multimodal document and video understanding, agentic workflows, technical research, and complex reasoning
Current and available; open-weight model
Long-context text generation, local inference, continued pretraining, research, and fine-tuning
Current open-weight model; publicly downloadable
Long-context text generation, self-hosted assistants, document analysis, research, and efficient local inference
Current open-weight instruction-tuned checkpoint
Self-hosted text generation, language-model research, efficient MoE inference, code and mathematics experimentation, and fine-tuning research
Available open-weight pretrained checkpoint
Self-hosted instruction following, general text generation, research, experimentation, and cost-conscious local inference
Available open-weight checkpoint
Native-resolution image feature extraction, vision-language model backbones, high-resolution document and image understanding, and multimodal research
Current open-weight model; available for local use through Hugging Face Transformers; not deployed by a Hugging Face Inference Provider