Bring your own key: Moonshot's API rates, no markup. Keys and chats stay in your browser. Run Moonshot in parallel with other models, then compare and merge the answers.
Kimi K3
NEWNative multimodal flagship (text, image, video inputs) with thinking on by default. 1M context.
1M
$3
$15
Jul 2026
Kimi K2.7 Code
Code-focused multimodal model (text, image, video inputs) with always-on thinking. ~180 tok/s output (up to 260 in short contexts for highspeed). 256K context.
262K
$0.95
$4
Jun 2026
Kimi K2.7 Code Highspeed
High-speed code variant with ~180 tok/s output (up to 260 in short contexts). Native multimodal with always-on thinking. 256K context.
262K
$1.9
$8
Jun 2026
Kimi K2.6
Native multimodal flagship (text, image, video inputs) with thinking and non-thinking modes. Stronger long-form coding, improved instruction compliance and sel…
262K
$0.95
$4
Apr 2026
Kimi K2.5
Supports vision (images/videos), thinking mode, and Agent tasks. 256K context.
262K
$0.6
$3
Jan 2026
V1 8K Vision (Preview)
Legacy vision model with 8K context. Preview variant - use moonshot-v1-vision for production.
8K
$0.2
$2
Jan 2025
V1 32K Vision (Preview)
Legacy vision model with 32K context. Preview variant - use moonshot-v1-vision for production.
33K
$1
$3
Jan 2025
V1 128K Vision (Preview)
Legacy vision model with 128K context. Preview variant - use moonshot-v1-vision for production.
131K
$2
$5
Jan 2025
V1 128K
Legacy V1 model with 128K context. Deprecated - use Kimi K2 Instruct instead.
131K
$2
$5
Feb 2024
V1 32K
Legacy V1 model with 32K context. Deprecated - use Kimi K2 Instruct instead.
33K
$1
$3
Feb 2024
V1 8K
Legacy V1 model with 8K context. Deprecated - use Kimi K2 Instruct instead.
8K
$0.2
$2
Feb 2024
Kimi K3
NEWNative multimodal flagship (text, image, video inputs) with thinking on by default. 1M context.
Kimi K2.7 Code
Code-focused multimodal model (text, image, video inputs) with always-on thinking. ~180 tok/s output (up to 260 in short contexts for highspeed). 256K context.
Kimi K2.7 Code Highspeed
High-speed code variant with ~180 tok/s output (up to 260 in short contexts). Native multimodal with always-on thinking. 256K context.
Kimi K2.6
Native multimodal flagship (text, image, video inputs) with thinking and non-thinking modes. Stronger long-form coding, improved instruction compliance and sel…
Kimi K2.5
Supports vision (images/videos), thinking mode, and Agent tasks. 256K context.
V1 8K Vision (Preview)
Legacy vision model with 8K context. Preview variant - use moonshot-v1-vision for production.
V1 32K Vision (Preview)
Legacy vision model with 32K context. Preview variant - use moonshot-v1-vision for production.
V1 128K Vision (Preview)
Legacy vision model with 128K context. Preview variant - use moonshot-v1-vision for production.
V1 128K
Legacy V1 model with 128K context. Deprecated - use Kimi K2 Instruct instead.
V1 32K
Legacy V1 model with 32K context. Deprecated - use Kimi K2 Instruct instead.
V1 8K
Legacy V1 model with 8K context. Deprecated - use Kimi K2 Instruct instead.
1
Create an API key at the Moonshot console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Add your Moonshot API key and run the Kimi models at Moonshot's own API rates. Big-AGI adds no markup and no intermediary: billing runs directly between you and Moonshot, and your keys stay in your browser.
Kimi's own app is a solid way to chat, but it only ever shows you Kimi's take. Beam sends the same prompt to Kimi and to GPT, Claude, or Gemini at once, so a long-context read of your codebase or document gets checked against other labs' models before you trust it. You also get parameters Kimi's app does not expose: temperature, system prompt, per-turn model switching, and a real thinking on/off toggle for the Kimi models that support it (the flagship coding model reasons on every turn, so Big-AGI doesn't show a fake switch for it). Your keys stay in your browser instead of sitting on Moonshot's servers.
Turn on Direct Connection and the browser calls Moonshot directly, bypassing the Big-AGI server, whenever your key is client-side and Moonshot allows it. Your keys stay in your browser. Chats are stored locally first and sync only if you turn it on. The AI Inspector opens on any message to show the exact request sent to Moonshot, the token counts, and a cost estimate for that call.
Run Kimi in parallel with Claude, GPT, and Gemini on the same prompt, then reach for Fusions: several strategies that combine, cross-check, and synthesize the parallel answers, which beats just picking the single best one. Parallel runs use more tokens than a single chat.
Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how Moonshot is called.
BIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego