Bring your own key: Fireworks AI's API rates, no markup. Keys stay in your browser. Run Fireworks AI in parallel with other models, then compare and merge the answers.
Muse Glimmer 30B (Vision)
NEWOpen-weights model served on Fireworks AI.
131K
-
-
Aug 2026
Qwen3.8 Max
NEWOpen-weights model served on Fireworks AI.
262K
-
-
Aug 2026
DeepSeek V4 Flash 0731
NEWOfficial release of DeepSeek V4 Flash, superseding the preview, with substantially enhanced agentic capabilities. Ships with a speculative decoding module atta…
1M
$0.14
$0.28
Jul 2026
Kimi K3 (Vision)
NEWMoonshot AI 2.8T-parameter flagship on Kimi Delta Attention, with native visual understanding and a 1M-token context for long-horizon coding and reasoning.
1M
$3
$15
Jul 2026
Kimi K3 Fast (Vision)
NEWFast serving path for Kimi K3: same model and quality, lower latency, higher per-token price.
1M
$4.5
$22.5
Jul 2026
Inkling (Vision)
NEWThinking Machines Lab first open-weights model: a 975B MoE (41B active) trained natively across text, image, and audio, with controllable thinking effort.
1M
-
-
Jul 2026
Intent De 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent De 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent De C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent En 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent En 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent En BC8D96
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent En C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent En VERIFY1
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent Es 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent Es 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
Intent Es C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
-
-
-
Jul 2026
GLM 5.2
Z.ai flagship with 1M-token context and multi-effort coding for long-horizon agentic tasks. New IndexShare architecture and improved MTP layer cut per-token co…
1M
$1.4
$4.4
Jun 2026
GLM 5.2 Fast
Fast serving path for GLM 5.2: same model and quality, lower latency, higher per-token price.
1M
$2.1
$6.6
Jun 2026
Kimi K2.7 Code (Vision)
Coding-focused agentic model built on Kimi K2.6, with better end-to-end completion on long-horizon software engineering and ~30% fewer thinking tokens.
262K
$0.95
$4
Jun 2026
Kimi K2.7 Code Fast (Vision)
Fast serving path for Kimi K2.7 Code: same model and quality, lower latency, higher per-token price.
262K
$1.9
$8
Jun 2026
MiniMax M3
MiniMax 428B MoE (23B active) with Sparse Attention for efficient long context, tuned for long-horizon agentic coding and cowork.
512K
$0.3
$1.2
Jun 2026
Qwen3.7 Plus (Vision)
Alibaba flagship closed model, available outside Alibaba infrastructure exclusively through Fireworks AI.
-
$0.4
$1.6
Jun 2026
NVIDIA Nemotron 3 Ultra NVFP4
NVIDIA frontier-scale hybrid LatentMoE (550B params, 55B active) interleaving Mamba-2 and MoE layers, for multi-step agents and long-context reasoning.
262K
$0.6
$2.4
Jun 2026
DeepSeek V4 Flash
Streamlined DeepSeek open MoE tuned for low-latency, high-throughput inference at 1M-token context, retaining most of Pro reasoning and coding quality.
1M
$0.14
$0.28
Apr 2026
DeepSeek V4 Pro
DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.
1M
$1.74
$3.48
Apr 2026
Kimi K2.6 (Vision)
Moonshot AI native-multimodal agentic model tuned for long-horizon coding, autonomous execution, and swarm task orchestration.
262K
$0.95
$4
Apr 2026
Kimi K2.6 Fast (Vision)
Fast serving path for Kimi K2.6: same model and quality, lower latency, higher per-token price.
262K
$2
$8
Apr 2026
MiniMax M2.7
MiniMax MoE built for complex agent harnesses and elaborate productivity tasks, leveraging Agent Teams, Skills, and dynamic tool search.
197K
$0.3
$1.2
Apr 2026
GLM 5.1
deprecatedZ.ai 754B-parameter MoE built for agentic engineering, with strong coding and sustained performance across long multi-round tasks.
203K
$1.4
$4.4
Mar 2026
glm 5.1 Fast
deprecatedOpen-weights model served on Fireworks AI.
203K
-
-
Mar 2026
Kimi K2.5 (Vision)
deprecatedOpen-weights model served on Fireworks AI.
262K
-
-
Jan 2026
QWEN3 Reranker 8B
deprecatedModel served on Fireworks AI.
41K
-
-
Oct 2025
QWEN3 Embedding 8B
deprecatedModel served on Fireworks AI.
41K
-
-
Aug 2025
GPT-OSS 120B
OpenAI open-weight model for high-reasoning, agentic, general-purpose use that fits on a single H100.
131K
$0.15
$0.6
Aug 2025
GPT-OSS 20B
OpenAI smaller open-weight model for lower-latency, local, and specialized use cases.
131K
$0.07
$0.3
Aug 2025
DeepSeek V4 Pro
DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.
1M
$1.74
$3.48
-
Muse Glimmer 30B (Vision)
NEWOpen-weights model served on Fireworks AI.
Qwen3.8 Max
NEWOpen-weights model served on Fireworks AI.
DeepSeek V4 Flash 0731
NEWOfficial release of DeepSeek V4 Flash, superseding the preview, with substantially enhanced agentic capabilities. Ships with a speculative decoding module atta…
Kimi K3 (Vision)
NEWMoonshot AI 2.8T-parameter flagship on Kimi Delta Attention, with native visual understanding and a 1M-token context for long-horizon coding and reasoning.
Kimi K3 Fast (Vision)
NEWFast serving path for Kimi K3: same model and quality, lower latency, higher per-token price.
Inkling (Vision)
NEWThinking Machines Lab first open-weights model: a 975B MoE (41B active) trained natively across text, image, and audio, with controllable thinking effort.
Intent De 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
Intent De 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
Intent De C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
Intent En 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
Intent En 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
Intent En BC8D96
deprecatedFine-tuned adapter served on Fireworks AI.
Intent En C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
Intent En VERIFY1
deprecatedFine-tuned adapter served on Fireworks AI.
Intent Es 3CD022
deprecatedFine-tuned adapter served on Fireworks AI.
Intent Es 6A34F8
deprecatedFine-tuned adapter served on Fireworks AI.
Intent Es C9EC41
deprecatedFine-tuned adapter served on Fireworks AI.
GLM 5.2
Z.ai flagship with 1M-token context and multi-effort coding for long-horizon agentic tasks. New IndexShare architecture and improved MTP layer cut per-token co…
GLM 5.2 Fast
Fast serving path for GLM 5.2: same model and quality, lower latency, higher per-token price.
Kimi K2.7 Code (Vision)
Coding-focused agentic model built on Kimi K2.6, with better end-to-end completion on long-horizon software engineering and ~30% fewer thinking tokens.
Kimi K2.7 Code Fast (Vision)
Fast serving path for Kimi K2.7 Code: same model and quality, lower latency, higher per-token price.
MiniMax M3
MiniMax 428B MoE (23B active) with Sparse Attention for efficient long context, tuned for long-horizon agentic coding and cowork.
Qwen3.7 Plus (Vision)
Alibaba flagship closed model, available outside Alibaba infrastructure exclusively through Fireworks AI.
NVIDIA Nemotron 3 Ultra NVFP4
NVIDIA frontier-scale hybrid LatentMoE (550B params, 55B active) interleaving Mamba-2 and MoE layers, for multi-step agents and long-context reasoning.
DeepSeek V4 Flash
Streamlined DeepSeek open MoE tuned for low-latency, high-throughput inference at 1M-token context, retaining most of Pro reasoning and coding quality.
DeepSeek V4 Pro
DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.
Kimi K2.6 (Vision)
Moonshot AI native-multimodal agentic model tuned for long-horizon coding, autonomous execution, and swarm task orchestration.
Kimi K2.6 Fast (Vision)
Fast serving path for Kimi K2.6: same model and quality, lower latency, higher per-token price.
MiniMax M2.7
MiniMax MoE built for complex agent harnesses and elaborate productivity tasks, leveraging Agent Teams, Skills, and dynamic tool search.
GLM 5.1
deprecatedZ.ai 754B-parameter MoE built for agentic engineering, with strong coding and sustained performance across long multi-round tasks.
glm 5.1 Fast
deprecatedOpen-weights model served on Fireworks AI.
Kimi K2.5 (Vision)
deprecatedOpen-weights model served on Fireworks AI.
QWEN3 Reranker 8B
deprecatedModel served on Fireworks AI.
QWEN3 Embedding 8B
deprecatedModel served on Fireworks AI.
GPT-OSS 120B
OpenAI open-weight model for high-reasoning, agentic, general-purpose use that fits on a single H100.
GPT-OSS 20B
OpenAI smaller open-weight model for lower-latency, local, and specialized use cases.
DeepSeek V4 Pro
DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.
1
Create an API key at the Fireworks AI console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Add your Fireworks API key over its OpenAI-compatible endpoint and reach the open-model catalog at Fireworks' own rates. Big-AGI adds no markup and no intermediary: the billing relationship runs directly between you and Fireworks.
Fireworks' playground is built for one prompt at a time. Big-AGI turns the same key into a persistent workspace: chats that stick around, personas, and file and image attachments, all layered on top of the raw API. It's also the only place a Fireworks model runs next to Claude, GPT, and Gemini in Beam, with the parameters and the key still yours to control.
Turn on Direct Connection and the browser talks to Fireworks directly, skipping the Big-AGI server, when your key is client-side and Fireworks allows it. Your keys stay in your browser. Chats are stored locally first, and sync only if you turn it on. The AI Inspector shows the exact request, the token counts, and a cost estimate, so you always know what you're billed for.
Put a Fireworks model into a Beam alongside frontier labs, or run a few open models side by side at Fireworks' speed. Fusions then combine, cross-check, and synthesize the parallel answers instead of just picking the best one. Parallel runs use more tokens than a single chat.
Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how Fireworks AI is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego