Use Fireworks AI Models in Big-AGI.

Fireworks exclusives like Alibaba's closed Qwen3.7-Plus, plus per-model fast serving tiers - with the names and prices Fireworks' API never returns, filled in.

All supported Fireworks AI models

ModelContextInputOutputReleased

GLM 5.3 Flash (Vision)

NEW
VisionReasoningTools / functions

First natively-multimodal model of the GLM-5 line: a 320B MoE (18B active) on a new base with hybrid sparse and linear attention, image input, and a 1M-token c…

1M

$0.15

$0.50

Aug 2026

DeepSeek V4 Flash (Vision) Exp

NEW
VisionTools / functions

Open-weights model served on Fireworks AI.

1M

-

-

Aug 2026

GLM 5.3

NEW
ReasoningTools / functions

Z.ai flagship, post-trained on the GLM-5.2 base for frontier coding, cybersecurity and long-horizon agentic work, with a 1M-token context.

1M

$1.40

$4.40

Aug 2026

DeepSeek V4 Pro 0813

NEW
ReasoningTools / functions

Official release of DeepSeek V4 Pro, superseding the preview, with greatly enhanced agentic capabilities, most pronounced in production environments. Ships wit…

1M

$1.32

$3.96

Aug 2026

Qwen3.8 2.4T-A95B (Vision)

NEW
VisionReasoningTools / functions

Open-weights release of the Qwen3.8 flagship: 2.4T sparse MoE with ~95B active parameters, built for multi-day coding runs and self-improving research.

262K

$2.00

$6.00

Aug 2026

NVIDIA Nemotron 3.5 Lightning 30B A3B

NEW
ReasoningTools / functions

NVIDIA hybrid Mamba-Transformer MoE (30B params, 3B active) with a multi-token prediction head, for low-latency agentic serving at long context.

262K

$0.05

$0.20

Aug 2026

Muse Glimmer 30B (Vision)

NEW
VisionReasoningTools / functions

Meta dense 30B distilled from Muse Spark for local agentic work: multi-step reasoning, schema-based tool calling, and image input across 100+ languages.

131K

$0.35

$1.50

Aug 2026

DeepSeek V4 Flash 0731

NEWHOT
ReasoningTools / functions

Official release of DeepSeek V4 Flash, superseding the preview, with substantially enhanced agentic capabilities. Ships with a speculative decoding module atta…

1M

$0.22

$0.66

Jul 2026

Qwen3.8 Max (Vision)

VisionReasoningTools / functions

Alibaba flagship Qwen3.8 tier, available outside Alibaba infrastructure through Fireworks AI.

262K

$2.00

$6.00

Jul 2026

Inkling (Vision)

VisionReasoningTools / functions

Thinking Machines Lab first open-weights model: a 975B MoE (41B active) trained natively across text, image, and audio, with controllable thinking effort.

1M

$1.00

$4.05

Jul 2026

Kimi K3 (Vision)

HOT
VisionReasoningTools / functions

Moonshot AI 2.8T-parameter flagship on Kimi Delta Attention, with native visual understanding and a 1M-token context for long-horizon coding and reasoning.

1M

$3.00

$15.00

Jul 2026

Kimi K3 Fast (Vision)

VisionReasoningTools / functions

Fast serving path for Kimi K3: same model and quality, lower latency, higher per-token price.

1M

$4.50

$22.50

Jul 2026

Intent De 3CD022

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent De 6A34F8

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent De C9EC41

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent En 3CD022

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent En 6A34F8

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent En BC8D96

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent En C9EC41

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent En VERIFY1

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent Es 3CD022

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent Es 6A34F8

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

Intent Es C9EC41

deprecated

Fine-tuned adapter served on Fireworks AI.

-

-

-

Jul 2026

GLM 5.2

ReasoningTools / functions

Z.ai flagship with 1M-token context and multi-effort coding for long-horizon agentic tasks. New IndexShare architecture and improved MTP layer cut per-token co…

1M

$1.40

$4.40

Jun 2026

GLM 5.2 Fast

ReasoningTools / functions

Fast serving path for GLM 5.2: same model and quality, lower latency, higher per-token price.

1M

$2.10

$6.60

Jun 2026

Kimi K2.7 Code (Vision)

VisionReasoningTools / functions

Coding-focused agentic model built on Kimi K2.6, with better end-to-end completion on long-horizon software engineering and ~30% fewer thinking tokens.

262K

$0.95

$4.00

Jun 2026

Kimi K2.7 Code Fast (Vision)

deprecated
VisionTools / functions

Fast serving path for Kimi K2.7 Code: same model and quality, lower latency, higher per-token price.

262K

$1.90

$8.00

Jun 2026

NVIDIA Nemotron 3 Ultra NVFP4

ReasoningTools / functions

NVIDIA frontier-scale hybrid LatentMoE (550B params, 55B active) interleaving Mamba-2 and MoE layers, for multi-step agents and long-context reasoning.

262K

$0.60

$2.40

Jun 2026

Qwen3.7 Plus (Vision)

VisionReasoningTools / functions

Alibaba flagship closed model, available outside Alibaba infrastructure exclusively through Fireworks AI.

-

$0.40

$1.60

Jun 2026

MiniMax M3

ReasoningTools / functions

MiniMax 428B MoE (23B active) with Sparse Attention for efficient long context, tuned for long-horizon agentic coding and cowork.

512K

$0.30

$1.20

May 2026

DeepSeek V4 Flash

deprecated
ReasoningTools / functions

Streamlined DeepSeek open MoE tuned for low-latency, high-throughput inference at 1M-token context, retaining most of Pro reasoning and coding quality.

1M

$0.14

$0.28

Apr 2026

DeepSeek V4 Pro

ReasoningTools / functions

DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.

1M

$1.74

$3.48

Apr 2026

Kimi K2.6 (Vision)

VisionReasoningTools / functions

Moonshot AI native-multimodal agentic model tuned for long-horizon coding, autonomous execution, and swarm task orchestration.

262K

$0.95

$4.00

Apr 2026

Kimi K2.6 Fast (Vision)

deprecated
VisionTools / functions

Fast serving path for Kimi K2.6: same model and quality, lower latency, higher per-token price.

262K

$2.00

$8.00

Apr 2026

GLM 5.1

deprecated
Tools / functions

Z.ai 754B-parameter MoE built for agentic engineering, with strong coding and sustained performance across long multi-round tasks.

203K

$1.40

$4.40

Mar 2026

glm 5.1 Fast

deprecated
Tools / functions

Open-weights model served on Fireworks AI.

203K

-

-

Mar 2026

MiniMax M2.7

ReasoningTools / functions

MiniMax MoE built for complex agent harnesses and elaborate productivity tasks, leveraging Agent Teams, Skills, and dynamic tool search.

197K

$0.30

$1.20

Mar 2026

Kimi K2.5 (Vision)

deprecated
VisionTools / functions

Open-weights model served on Fireworks AI.

262K

-

-

Jan 2026

QWEN3 Reranker 8B

deprecated

Model served on Fireworks AI.

41K

-

-

Oct 2025

QWEN3 Embedding 8B

deprecated

Model served on Fireworks AI.

41K

-

-

Aug 2025

GPT-OSS 120B

ReasoningTools / functions

OpenAI open-weight model for high-reasoning, agentic, general-purpose use that fits on a single H100.

131K

$0.15

$0.60

Aug 2025

GPT-OSS 20B

deprecated

OpenAI smaller open-weight model for lower-latency, local, and specialized use cases.

131K

$0.07

$0.30

Aug 2025

GLM 5.3 Flash (Vision)

NEW
Aug 2026

First natively-multimodal model of the GLM-5 line: a 320B MoE (18B active) on a new base with hybrid sparse and linear attention, image input, and a 1M-token c…

VisionReasoningTools / functions
1M · in $0.15 · out $0.50

DeepSeek V4 Flash (Vision) Exp

NEW
Aug 2026

Open-weights model served on Fireworks AI.

VisionTools / functions
1M · in - · out -

GLM 5.3

NEW
Aug 2026

Z.ai flagship, post-trained on the GLM-5.2 base for frontier coding, cybersecurity and long-horizon agentic work, with a 1M-token context.

ReasoningTools / functions
1M · in $1.40 · out $4.40

DeepSeek V4 Pro 0813

NEW
Aug 2026

Official release of DeepSeek V4 Pro, superseding the preview, with greatly enhanced agentic capabilities, most pronounced in production environments. Ships wit…

ReasoningTools / functions
1M · in $1.32 · out $3.96

Qwen3.8 2.4T-A95B (Vision)

NEW
Aug 2026

Open-weights release of the Qwen3.8 flagship: 2.4T sparse MoE with ~95B active parameters, built for multi-day coding runs and self-improving research.

VisionReasoningTools / functions
262K · in $2.00 · out $6.00

NVIDIA Nemotron 3.5 Lightning 30B A3B

NEW
Aug 2026

NVIDIA hybrid Mamba-Transformer MoE (30B params, 3B active) with a multi-token prediction head, for low-latency agentic serving at long context.

ReasoningTools / functions
262K · in $0.05 · out $0.20

Muse Glimmer 30B (Vision)

NEW
Aug 2026

Meta dense 30B distilled from Muse Spark for local agentic work: multi-step reasoning, schema-based tool calling, and image input across 100+ languages.

VisionReasoningTools / functions
131K · in $0.35 · out $1.50

DeepSeek V4 Flash 0731

NEWHOT
Jul 2026

Official release of DeepSeek V4 Flash, superseding the preview, with substantially enhanced agentic capabilities. Ships with a speculative decoding module atta…

ReasoningTools / functions
1M · in $0.22 · out $0.66

Qwen3.8 Max (Vision)

Jul 2026

Alibaba flagship Qwen3.8 tier, available outside Alibaba infrastructure through Fireworks AI.

VisionReasoningTools / functions
262K · in $2.00 · out $6.00

Inkling (Vision)

Jul 2026

Thinking Machines Lab first open-weights model: a 975B MoE (41B active) trained natively across text, image, and audio, with controllable thinking effort.

VisionReasoningTools / functions
1M · in $1.00 · out $4.05

Kimi K3 (Vision)

HOT
Jul 2026

Moonshot AI 2.8T-parameter flagship on Kimi Delta Attention, with native visual understanding and a 1M-token context for long-horizon coding and reasoning.

VisionReasoningTools / functions
1M · in $3.00 · out $15.00

Kimi K3 Fast (Vision)

Jul 2026

Fast serving path for Kimi K3: same model and quality, lower latency, higher per-token price.

VisionReasoningTools / functions
1M · in $4.50 · out $22.50

Intent De 3CD022

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent De 6A34F8

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent De C9EC41

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent En 3CD022

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent En 6A34F8

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent En BC8D96

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent En C9EC41

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent En VERIFY1

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent Es 3CD022

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent Es 6A34F8

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

Intent Es C9EC41

deprecated
Jul 2026

Fine-tuned adapter served on Fireworks AI.

- · in - · out -

GLM 5.2

Jun 2026

Z.ai flagship with 1M-token context and multi-effort coding for long-horizon agentic tasks. New IndexShare architecture and improved MTP layer cut per-token co…

ReasoningTools / functions
1M · in $1.40 · out $4.40

GLM 5.2 Fast

Jun 2026

Fast serving path for GLM 5.2: same model and quality, lower latency, higher per-token price.

ReasoningTools / functions
1M · in $2.10 · out $6.60

Kimi K2.7 Code (Vision)

Jun 2026

Coding-focused agentic model built on Kimi K2.6, with better end-to-end completion on long-horizon software engineering and ~30% fewer thinking tokens.

VisionReasoningTools / functions
262K · in $0.95 · out $4.00

Kimi K2.7 Code Fast (Vision)

deprecated
Jun 2026

Fast serving path for Kimi K2.7 Code: same model and quality, lower latency, higher per-token price.

VisionTools / functions
262K · in $1.90 · out $8.00

NVIDIA Nemotron 3 Ultra NVFP4

Jun 2026

NVIDIA frontier-scale hybrid LatentMoE (550B params, 55B active) interleaving Mamba-2 and MoE layers, for multi-step agents and long-context reasoning.

ReasoningTools / functions
262K · in $0.60 · out $2.40

Qwen3.7 Plus (Vision)

Jun 2026

Alibaba flagship closed model, available outside Alibaba infrastructure exclusively through Fireworks AI.

VisionReasoningTools / functions
- · in $0.40 · out $1.60

MiniMax M3

May 2026

MiniMax 428B MoE (23B active) with Sparse Attention for efficient long context, tuned for long-horizon agentic coding and cowork.

ReasoningTools / functions
512K · in $0.30 · out $1.20

DeepSeek V4 Flash

deprecated
Apr 2026

Streamlined DeepSeek open MoE tuned for low-latency, high-throughput inference at 1M-token context, retaining most of Pro reasoning and coding quality.

ReasoningTools / functions
1M · in $0.14 · out $0.28

DeepSeek V4 Pro

Apr 2026

DeepSeek flagship open MoE (1.6T params) for frontier reasoning, coding, and long-context work up to 1M tokens. Hybrid attention keeps long contexts efficient.

ReasoningTools / functions
1M · in $1.74 · out $3.48

Kimi K2.6 (Vision)

Apr 2026

Moonshot AI native-multimodal agentic model tuned for long-horizon coding, autonomous execution, and swarm task orchestration.

VisionReasoningTools / functions
262K · in $0.95 · out $4.00

Kimi K2.6 Fast (Vision)

deprecated
Apr 2026

Fast serving path for Kimi K2.6: same model and quality, lower latency, higher per-token price.

VisionTools / functions
262K · in $2.00 · out $8.00

GLM 5.1

deprecated
Mar 2026

Z.ai 754B-parameter MoE built for agentic engineering, with strong coding and sustained performance across long multi-round tasks.

Tools / functions
203K · in $1.40 · out $4.40

glm 5.1 Fast

deprecated
Mar 2026

Open-weights model served on Fireworks AI.

Tools / functions
203K · in - · out -

MiniMax M2.7

Mar 2026

MiniMax MoE built for complex agent harnesses and elaborate productivity tasks, leveraging Agent Teams, Skills, and dynamic tool search.

ReasoningTools / functions
197K · in $0.30 · out $1.20

Kimi K2.5 (Vision)

deprecated
Jan 2026

Open-weights model served on Fireworks AI.

VisionTools / functions
262K · in - · out -

QWEN3 Reranker 8B

deprecated
Oct 2025

Model served on Fireworks AI.

41K · in - · out -

QWEN3 Embedding 8B

deprecated
Aug 2025

Model served on Fireworks AI.

41K · in - · out -

GPT-OSS 120B

Aug 2025

OpenAI open-weight model for high-reasoning, agentic, general-purpose use that fits on a single H100.

ReasoningTools / functions
131K · in $0.15 · out $0.60

GPT-OSS 20B

deprecated
Aug 2025

OpenAI smaller open-weight model for lower-latency, local, and specialized use cases.

131K · in $0.07 · out $0.30
42 models · sorted by release date · prices in USD per 1M tokens · refreshed hourlyCompare every model across vendors ->

Get started in 3 steps

1

Create an API key at the Fireworks AI console.

2

Paste it into Big-AGI's model settings.

3

Start chatting, or Beam it against other models and fuse the answers.

Running Fireworks AI in Big-AGI

Add your Fireworks API key over its OpenAI-compatible endpoint and reach the open-model catalog at Fireworks' own rates. Big-AGI adds no markup and no intermediary: the billing relationship runs directly between you and Fireworks.

  • Your key, your billing. Usage is billed by Fireworks AI to your account.
  • Exclusives, with the names filled in. Fireworks carries models you can't get elsewhere - Alibaba's closed Qwen3.7-Plus among them - and since its API returns no names or prices, Big-AGI supplies both from a curated table.
  • Zero setup beyond the key. Big-AGI recognizes Fireworks endpoints by hostname and turns on the right vision and tool flags per model automatically.
  • Speed as the product. Fireworks' custom inference stack pushes open models to interactive latencies, priced per token.

Why Big-AGI instead of the playground?

Fireworks' playground is built for one prompt at a time. Big-AGI turns the same key into a persistent workspace: chats that stick around, personas, and file and image attachments, all layered on top of the raw API. It's also the only place a Fireworks model runs next to Claude, GPT, and Gemini in Beam, with the parameters and the key still yours to control.

Your keys and your data

Turn on Direct Connection and the browser talks to Fireworks directly, skipping the Big-AGI server, when your key is client-side and Fireworks allows it. Your keys stay in your browser. Chats are stored locally first, and sync only if you turn it on. The AI Inspector shows the exact request, the token counts, and a cost estimate, so you always know what you're billed for.

Fireworks AI in Beam

Put a Fireworks model into a Beam alongside frontier labs, or run a few open models side by side at Fireworks' speed. Fusions then combine, cross-check, and synthesize the parallel answers instead of just picking the best one. Parallel runs use more tokens than a single chat.

Bring your Fireworks AI key. Keep control.

Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Fireworks AI is called.

Launch Big-AGI

© 2026 Token Fabrics·Built with passion in San Diego