Modular Cloud models in Big-AGI

Open models on NVIDIA and AMD GPUs, from the team behind Mojo and MAX. Your key, Modular's rates, no markup.

Open models, served fast

Modular Cloud is the hosted endpoint of the MAX and Mojo team, now part of Qualcomm: open-weights models on NVIDIA and AMD GPUs, pay per token. Of the three models we compared across hosts, Modular served each one fastest:

ModelOn ModularElsewhere, same day
MiniMax M3~238 tok/s, first token 257 ms64 (MiniMax API) · 138 (OpenRouter)
Kimi K2.7 Code~170 tok/s57 (Moonshot API)
Gemma 4 31B~199 tok/s, first token 231 ms132 (OpenRouter)

Measured August 13, 2026 from a US connection, best of two ~200-token streaming runs; not a benchmark. Re-measured quarterly.

Free to try: new accounts currently include $10 of API credit, enough to try every model on this page.

Models are served as NVIDIA NVFP4 4-bit quantized builds, disclosed in the response's model field. Our small task checks found no quality difference against full-precision hosts.

Web UI for a self-hosted MAX server

MAX is the same engine as a free download: it runs on NVIDIA, AMD and Apple Silicon and serves the same OpenAI-compatible API from your own machine - get MAX from Modular. Big-AGI connects to it through the same Modular vendor, as a second endpoint next to the cloud one.

All supported Modular models

7 models · sorted by release date · prices in USD per 1M tokens · refreshed hourlyCompare every model across vendors ->

Get started in 3 steps

1

Create an API key at the Modular console.

2

Paste it into Big-AGI's model settings.

3

Start chatting, or Beam it against other models and fuse the answers.

Good to know

  • Keys expire on a set date and are recreated at the console, not renewed.
  • Accounts start at 60 requests and 600K tokens per minute, tripling once a payment method is on file.
  • The model list is read live: when Modular ships a new model, it appears in your picker on the next refresh.
  • For coding agents, Kimi K2.7 Code. It always reasons before answering, so budget output tokens above the visible answer.

Your key and your data

Modular's API does not accept direct browser calls, so requests go through the Big-AGI fast edge servers - there is no Direct Connection toggle on this vendor. Your key stays in your browser and is sent only with your requests. Chats are stored on your device first and sync only if you turn sync on. The AI Inspector shows every request, the token counts, and the cost estimate.

Modular in Beam

Beam runs one prompt across several models in parallel, then fuses the answers, and one Modular balance covers several of its models in the same run. Parallel runs use more tokens than a single chat.

Modular questions

FAQ

Is Modular Cloud OpenAI-compatible, and do I need a base URL in Big-AGI?

Yes: Modular Cloud serves the OpenAI chat completions API at api.modular.com/v1, which is how Cursor and OpenCode connect to it. In Big-AGI, Modular is a native vendor: paste the key, no base URL, and the model list loads from Modular's catalog.

How much does Modular Cloud cost?

Pay per token at Modular's published per-model rates, drawn from prepaid credit you load at console.modular.com; every model's price is in the list on this page. Big-AGI adds no markup: the bill is between you and Modular.

Why does my Modular API key return a 401?

Three causes share that code: the prepaid balance hit zero, the key passed its expiration date, or the key is wrong. Top up or recreate the key at console.modular.com; Big-AGI refills the model list once a working key is in place.

Can Big-AGI be the web UI for my own MAX server?

Yes. Switch the Modular vendor's Endpoint to "MAX self-hosted" and enter your server URL, such as http://localhost:8000; an API key is optional. The models your MAX server loads then appear in Big-AGI's list, next to Modular Cloud's.

Checked September 14, 2026

Bring your Modular key. Keep control.

Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Modular is called.

Launch Big-AGI

© 2026 Token Fabrics·Built with passion in San Diego