Open models on NVIDIA and AMD GPUs, from the team behind Mojo and MAX - the fastest host in our August 2026 spot checks. Your key, Modular's rates, no markup.
Modular Cloud is the hosted endpoint of the MAX and Mojo team, now part of Qualcomm: open-weights models on NVIDIA and AMD GPUs, pay per token. In our spot checks of August 13, 2026 it was the fastest host of every model we compared:
| Model | On Modular | Elsewhere, same day |
|---|---|---|
| MiniMax M3 | ~238 tok/s, first token 257 ms | 64 (MiniMax API) · 138 (OpenRouter) |
| Kimi K2.7 Code | ~170 tok/s | 57 (Moonshot API) |
| Gemma 4 31B | ~199 tok/s, first token 231 ms | 132 (OpenRouter) |
Free to try: new accounts currently include $10 of API credit (promotional, August 2026) - enough to try every model on this page.
Models are served as NVIDIA NVFP4 4-bit quantized builds, disclosed in the response's model field. Our small task checks found no quality difference against full-precision hosts.
MAX is the same engine as a free download: it runs on NVIDIA, AMD and Apple Silicon, and serves the same models from your own machine. Flip the vendor's Endpoint switch to MAX self-hosted, point it at your server, and its models appear in your list - get MAX from Modular.
GLM 5.3
NEWZhipu GLM-5.3, post-trained on the GLM-5.2 base for frontier coding and long-horizon agentic work, reasoning on by default (effort control), text-only. 192K co…
192K
$1.40
$4.40
Aug 2026
GLM 5.2
Zhipu GLM-5.2 open-weights coding/agentic MoE (753B, ~40B active), reasoning on by default (effort control), text-only. Served as NVIDIA NVFP4 (4-bit) quantiza…
164K
$1.40
$4.40
Jun 2026
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.60
$3.00
Jun 2026
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
1M
$0.30
$1.20
May 2026
Gemma 4 26B A4B
deprecatedGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.15
$0.60
Apr 2026
Gemma 4 31B
HOTGoogle Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
262K
$0.25
$0.65
Apr 2026
FLUX.2 Klein 4B
deprecatedModel served via Modular.
-
-
-
-
GLM 5.3
NEWZhipu GLM-5.3, post-trained on the GLM-5.2 base for frontier coding and long-horizon agentic work, reasoning on by default (effort control), text-only. 192K co…
GLM 5.2
Zhipu GLM-5.2 open-weights coding/agentic MoE (753B, ~40B active), reasoning on by default (effort control), text-only. Served as NVIDIA NVFP4 (4-bit) quantiza…
Kimi K2.7 Code
Moonshot Kimi K2.7 Code, agentic coding model with always-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization.
MiniMax M3
1M-context multimodal MoE with default-on reasoning. Served as NVIDIA NVFP4 (4-bit) quantization on Modular Cloud shared endpoints.
Gemma 4 26B A4B
deprecatedGoogle Gemma 4 26B MoE (4B active), text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
Gemma 4 31B
HOTGoogle Gemma 4 31B instruction-tuned, text+image input. Served as NVIDIA NVFP4 (4-bit) quantization.
FLUX.2 Klein 4B
deprecatedModel served via Modular.
1
Create an API key at the Modular console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Modular's API does not accept direct browser calls, so requests go through the Big-AGI fast edge servers - there is no Direct Connection toggle on this vendor. Your key stays in your browser and is sent only with your requests. Chats are stored on your device first and sync only if you turn sync on. The AI Inspector shows every request, the token counts, and the cost estimate.
Beam runs one prompt across several models in parallel, then fuses the answers. Modular's speed means its models usually finish first, and one balance covers several models in the same run - with the cache rate covering the re-sent context. Parallel runs use more tokens than a single chat.
FAQ
Three causes share that code: the prepaid balance hit zero, the key passed its hard expiration date, or the key is wrong. Top up or recreate the key at console.modular.com - keys cannot be regenerated, only deleted and recreated. Big-AGI fills the model list again as soon as a working key is in place.
Modular runs open models on both NVIDIA and AMD GPUs. The served builds are NVIDIA NVFP4 4-bit quantizations, which the API discloses in the response's model field. In our checks of August 13, 2026 they matched full-precision hosts on small tasks while streaming several times faster than the models' own APIs - though checks that small cannot rule out subtle differences.
Yes. The MAX container is free, runs on NVIDIA, AMD and Apple Silicon, and serves an OpenAI-compatible endpoint. Big-AGI's Modular vendor has a "MAX self-hosted" endpoint mode that takes your server URL, with an API key optional.
In our spot checks of August 13, 2026 - two ~200-token streaming runs per host from a US connection - Modular streamed MiniMax M3 at ~238 tokens per second with a 257 ms first token, against 138 tokens per second on OpenRouter's fastest route and 64 on MiniMax's own API. Single-day checks, not benchmarks, but the gap was not close.
Kimi K2.7 Code. In our August 2026 probes it completed a full three-turn tool loop - list a directory, read a file, answer - in about 1.1 seconds of wall clock, with well-formed parallel and streamed tool calls. It always reasons before answering, so budget output tokens above the visible answer.
$0.30 per million input tokens and $1.20 per million output tokens as of August 13, 2026, with cached input billed at $0.06 - repeated context in agents and long chats pays the cache rate. New accounts currently include $10 of free credit (August 2026); after that, billing is prepaid from your Modular account balance. Modular's pricing page is the live source.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Modular is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego