Bring your own key: Modular's API rates, no markup. Keys stay in your browser. Run Modular in parallel with other models, then compare and merge the answers.
Modular Cloud is the hosted endpoint of the MAX and Mojo team, now part of Qualcomm: open-weights models on NVIDIA and AMD GPUs, one sk-mod- key, pay-per-token from a prepaid balance. In our spot checks of August 13, 2026 - two ~200-token streaming runs per host from a US connection, not benchmarks - it was the fastest host of every model we compared that day, with MiniMax M3 at ~238 tokens per second where the model's own API measured 64.
The fine print, verified the same day: models are served as NVIDIA NVFP4 4-bit quantized builds (disclosed in the response's model field; our small task checks showed no quality difference against full-precision hosts), credit must be loaded before the first answer, and the same Big-AGI vendor also connects a free self-hosted MAX server.
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.6
$3
Jun 2026
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
1M
$0.3
$1.2
May 2026
Gemma 4 26B A4B
HOTGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.15
$0.6
Apr 2026
Gemma 4 31B
Google Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.25
$0.65
Apr 2026
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
Gemma 4 26B A4B
HOTGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
Gemma 4 31B
Google Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
1
Create an API key at the Modular console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Add your Modular API key and run open models on Modular Cloud at their own per-token rates. Big-AGI adds no markup and no intermediary: billing runs directly between you and Modular, from the prepaid balance on your Modular account. The model list is read live, so when Modular adds a model - they have shipped day-zero releases before - it appears in your picker on the next refresh, with no app update.
The same Big-AGI vendor connects two very different endpoints, and an Endpoint chip picks between them. Modular Cloud is the managed side: shared endpoints on NVIDIA and AMD GPUs, one key, pay per token. MAX self-hosted points the vendor at a MAX server you run yourself - the container is free, runs on NVIDIA, AMD and Apple Silicon, and speaks the same protocol, so whatever it serves shows up in the list automatically. Both can exist side by side, the cloud for the big MoE models and your own machine for everything you want local.
Modular's API does not accept direct browser calls, so requests route through the Big-AGI fast edge servers - there is no Direct Connection toggle on this vendor. Your key stays stored in your browser and travels only to make your requests. Chats are stored locally first and sync only if you turn it on. The AI Inspector shows the exact request, the token counts, and a cost estimate built from Modular's published rates - the estimate never enforces anything; the prepaid balance at the Modular console does.
The models Modular serves - the same open weights available from other hosts - streamed faster here than anywhere else we tried them in our August 2026 checks. That pace suits Beam: put a Modular-served model next to Claude, GPT, and Gemini, let them answer in parallel, then fuse, cross-check, or synthesize the results. One prepaid Modular balance covers several distinct models in the scatter, and parallel runs use more tokens than a single chat.
FAQ
Three causes share that code: the prepaid balance hit zero, the key passed its hard expiration date, or the key is wrong. Top up or recreate the key at console.modular.com - keys cannot be regenerated, only deleted and recreated. Big-AGI fills the model list again as soon as a working key is in place.
Modular runs open models on both NVIDIA and AMD GPUs. The served builds are NVIDIA NVFP4 4-bit quantizations, which the API discloses in the response's model field. In our checks of August 13, 2026 they matched full-precision hosts on small tasks while streaming several times faster than the models' own APIs - though checks that small cannot rule out subtle differences.
Yes. The MAX container is free, runs on NVIDIA, AMD and Apple Silicon, and serves an OpenAI-compatible endpoint. Big-AGI's Modular vendor has a "MAX self-hosted" endpoint mode that takes your server URL, with an API key optional.
Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how Modular is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego