Open models on NVIDIA and AMD GPUs, from the team behind Mojo and MAX. Your key, Modular's rates, no markup.
Modular Cloud is the hosted endpoint of the MAX and Mojo team, now part of Qualcomm: open-weights models on NVIDIA and AMD GPUs, pay per token. Of the three models we compared across hosts, Modular served each one fastest:
| Model | On Modular | Elsewhere, same day |
|---|---|---|
| MiniMax M3 | ~238 tok/s, first token 257 ms | 64 (MiniMax API) · 138 (OpenRouter) |
| Kimi K2.7 Code | ~170 tok/s | 57 (Moonshot API) |
| Gemma 4 31B | ~199 tok/s, first token 231 ms | 132 (OpenRouter) |
Measured August 13, 2026 from a US connection, best of two ~200-token streaming runs; not a benchmark. Re-measured quarterly.
Free to try: new accounts currently include $10 of API credit, enough to try every model on this page.
Models are served as NVIDIA NVFP4 4-bit quantized builds, disclosed in the response's model field. Our small task checks found no quality difference against full-precision hosts.
MAX is the same engine as a free download: it runs on NVIDIA, AMD and Apple Silicon and serves the same OpenAI-compatible API from your own machine - get MAX from Modular. Big-AGI connects to it through the same Modular vendor, as a second endpoint next to the cloud one.
GLM 5.3
NEWZhipu GLM-5.3, post-trained on the GLM-5.2 base for frontier coding and long-horizon agentic work, reasoning on by default (effort control), text-only. 192K co…
192K
$1.40
$4.40
Aug 2026
GLM 5.2
Zhipu GLM-5.2 open-weights coding/agentic MoE (753B, ~40B active), reasoning on by default (effort control), text-only. Served as NVIDIA NVFP4 (4-bit) quantiza…
1M
$1.40
$4.40
Jun 2026
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.60
$3.00
Jun 2026
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
1M
$0.30
$1.20
May 2026
Gemma 4 31B
Google Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.25
$0.65
Apr 2026
Gemma 4 26B A4B
deprecatedGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.15
$0.60
Apr 2026
FLUX.2 Klein 4B
deprecatedModel served via Modular.
no spec
GLM 5.3
NEWZhipu GLM-5.3, post-trained on the GLM-5.2 base for frontier coding and long-horizon agentic work, reasoning on by default (effort control), text-only. 192K co…
GLM 5.2
Zhipu GLM-5.2 open-weights coding/agentic MoE (753B, ~40B active), reasoning on by default (effort control), text-only. Served as NVIDIA NVFP4 (4-bit) quantiza…
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
Gemma 4 31B
Google Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
Gemma 4 26B A4B
deprecatedGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
FLUX.2 Klein 4B
deprecatedModel served via Modular.
1
Create an API key at the Modular console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Modular's API does not accept direct browser calls, so requests go through the Big-AGI fast edge servers - there is no Direct Connection toggle on this vendor. Your key stays in your browser and is sent only with your requests. Chats are stored on your device first and sync only if you turn sync on. The AI Inspector shows every request, the token counts, and the cost estimate.
Beam runs one prompt across several models in parallel, then fuses the answers, and one Modular balance covers several of its models in the same run. Parallel runs use more tokens than a single chat.
FAQ
Yes: Modular Cloud serves the OpenAI chat completions API at api.modular.com/v1, which is how Cursor and OpenCode connect to it. In Big-AGI, Modular is a native vendor: paste the key, no base URL, and the model list loads from Modular's catalog.
Pay per token at Modular's published per-model rates, drawn from prepaid credit you load at console.modular.com; every model's price is in the list on this page. Big-AGI adds no markup: the bill is between you and Modular.
Three causes share that code: the prepaid balance hit zero, the key passed its expiration date, or the key is wrong. Top up or recreate the key at console.modular.com; Big-AGI refills the model list once a working key is in place.
Yes. Switch the Modular vendor's Endpoint to "MAX self-hosted" and enter your server URL, such as http://localhost:8000; an API key is optional. The models your MAX server loads then appear in Big-AGI's list, next to Modular Cloud's.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Modular is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego