Bring your own key: free models on NVIDIA's own endpoint, no charges. Keys stay in your browser. Run NVIDIA in parallel with other models, then compare and merge the answers.
Every model below runs at no cost on NVIDIA's hosted catalog: Nemotron 3, GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card. The throttle counts requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop.
Two limits to know: expect roughly 40 requests per minute (NVIDIA publishes no number, and it moves with load), and NVIDIA's trial terms allow it to use your prompts to improve its models and exclude production use. Keep sensitive and production work on a paid provider.
Ising Calibration 1.5 31B
NEWNVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).
131K
Free
Free
Jul 2026
Inkling
NEWThinking Machines reasoning model (preview). Can be unstable under load on the free endpoint.
131K
Free
Free
Jul 2026
Laguna XS 2.1
NEWPoolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.
262K
Free
Free
Jul 2026
GLM 5.2
HOTZhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M).
203K
Free
Free
Jun 2026
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hanging under load.
250K
Free
Free
Jun 2026
Nemotron 3 Ultra 550B
NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.
1M
Free
Free
Jun 2026
Nemotron 3.5 Content Safety
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Jun 2026
MiniMax M3
HOTMiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 512K context (native: 1M).
524K
Free
Free
May 2026
Step 3.7 Flash
StepFun fast multimodal reasoning model, 256K context.
262K
Free
Free
May 2026
Nemotron 3 Nano Omni 30B
Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.
1M
Free
Free
Apr 2026
Mistral Medium 3.5
deprecatedMistral frontier-class multimodal model with adjustable reasoning.
262K
Free
Free
Apr 2026
DeepSeek V4 Pro
deprecatedDeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.
262K
Free
Free
Apr 2026
DeepSeek V4 Flash
deprecatedFast DeepSeek V4 MoE (284B) with 1M context and reasoning.
1M
Free
Free
Apr 2026
Ising Calibration 1 35B
deprecatedNVIDIA quantum-calibration VLM (domain-specific, not for general chat).
262K
Free
Free
Apr 2026
Gemma 4 31B
HOTGoogle open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.
131K
Free
Free
Apr 2026
Nemotron 3 Super 120B
Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.
1M
Free
Free
Mar 2026
Qwen 3.5 397B
deprecatedQwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.
262K
Free
Free
Feb 2026
Nemotron 3 Nano 30B
Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use.
1M
Free
Free
Dec 2025
Riva Translate 4B
Translation-specialized model, 8K context.
8K
Free
Free
Dec 2025
Nemotron Nano 12B v2 VL
Vision-language Nemotron Nano for image and video understanding (verified image input).
131K
Free
Free
Oct 2025
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Oct 2025
Llama 3.3 Nemotron Super 49B v1.5
Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks.
131K
Free
Free
Oct 2025
Nemotron Nano 9B v2
Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use.
128K
Free
Free
Aug 2025
GPT-OSS 20B
OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.
131K
Free
Free
Aug 2025
GPT-OSS 120B
OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.
131K
Free
Free
Aug 2025
Mistral Nemotron
Mistral model post-trained by NVIDIA.
262K
Free
Free
Jun 2025
Nemotron Nano VL 8B
Small vision-language model, 16K context.
16K
Free
Free
May 2025
Llama Guard 4 12B
Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.
66K
Free
Free
Apr 2025
Llama 3.3 Nemotron Super 49B v1
Superseded by v1.5.
131K
Free
Free
Mar 2025
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Jan 2025
Llama 3.3 70B
Meta Llama 3.3 70B instruction-tuned, with tool calling.
131K
Free
Free
Dec 2024
Llama 3.2 1B
Tiny Llama for edge-class tasks.
131K
Free
Free
Sep 2024
Llama 3.2 11B Vision
Llama vision model for image understanding.
131K
Free
Free
Sep 2024
Llama 3.2 3B
Small Llama for lightweight tasks.
131K
Free
Free
Sep 2024
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
33K
Free
Free
Sep 2024
Nemotron Mini 4B
Tiny legacy Nemotron, 4K context.
4K
Free
Free
Sep 2024
Llama 3.1 70B
Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
131K
Free
Free
Jul 2024
Llama 3.1 8B
Small fast Llama for utility tasks, with tool calling.
131K
Free
Free
Jul 2024
Ising Calibration 1.5 31B
NEWNVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).
Inkling
NEWThinking Machines reasoning model (preview). Can be unstable under load on the free endpoint.
Laguna XS 2.1
NEWPoolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.
GLM 5.2
HOTZhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M).
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hanging under load.
Nemotron 3 Ultra 550B
NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.
Nemotron 3.5 Content Safety
NVIDIA content-safety classifier (not a chat model).
MiniMax M3
HOTMiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 512K context (native: 1M).
Step 3.7 Flash
StepFun fast multimodal reasoning model, 256K context.
Nemotron 3 Nano Omni 30B
Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.
Mistral Medium 3.5
deprecatedMistral frontier-class multimodal model with adjustable reasoning.
DeepSeek V4 Pro
deprecatedDeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.
DeepSeek V4 Flash
deprecatedFast DeepSeek V4 MoE (284B) with 1M context and reasoning.
Ising Calibration 1 35B
deprecatedNVIDIA quantum-calibration VLM (domain-specific, not for general chat).
Gemma 4 31B
HOTGoogle open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.
Nemotron 3 Super 120B
Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.
Qwen 3.5 397B
deprecatedQwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.
Nemotron 3 Nano 30B
Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use.
Riva Translate 4B
Translation-specialized model, 8K context.
Nemotron Nano 12B v2 VL
Vision-language Nemotron Nano for image and video understanding (verified image input).
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
Llama 3.3 Nemotron Super 49B v1.5
Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks.
Nemotron Nano 9B v2
Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use.
GPT-OSS 20B
OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.
GPT-OSS 120B
OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.
Mistral Nemotron
Mistral model post-trained by NVIDIA.
Nemotron Nano VL 8B
Small vision-language model, 16K context.
Llama Guard 4 12B
Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.
Llama 3.3 Nemotron Super 49B v1
Superseded by v1.5.
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
Llama 3.3 70B
Meta Llama 3.3 70B instruction-tuned, with tool calling.
Llama 3.2 1B
Tiny Llama for edge-class tasks.
Llama 3.2 11B Vision
Llama vision model for image understanding.
Llama 3.2 3B
Small Llama for lightweight tasks.
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
Nemotron Mini 4B
Tiny legacy Nemotron, 4K context.
Llama 3.1 70B
Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
Llama 3.1 8B
Small fast Llama for utility tasks, with tool calling.
1
Create an API key at the NVIDIA console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
nvapi- key at build.nvidia.com, keeping the Public API Endpoints scope on it.The build.nvidia.com playground is one prompt, one model, no history. The same key here gets persistent chats, personas, attachments, and NVIDIA models in the same conversation as Claude, GPT, or Gemini.
Run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then Fuse: combine, cross-check, and synthesize the parallel answers instead of just picking one.
Point the same service at your own host, a NIM container or any OpenAI-compatible server (http://spark-xxxx.local:8000 style): identical model ids, same protocol. Direct Connection works on local hosts; the hosted endpoint blocks browser-origin calls, so those route through the server. Either way keys stay in your browser and chats are stored locally first.
FAQ
Either the model is retired (NVIDIA's catalog still lists it, but roughly half the ids no longer serve), or your key was created without the "Public API Endpoints" scope, which fails on every model. Fix the key at build.nvidia.com. Big-AGI only lists live-verified models, so a retired id never reaches your picker.
NVIDIA does not publish one. Expect roughly 40 requests per minute per account, moving with load. It counts requests, not tokens, so long prompts and answers are not penalized. Big-AGI paces its calls, so bursts and Beam scatters queue instead of collecting 429s.
Yes, no billing exists on the hosted endpoint. The catch is in the terms, not the invoice: no production use, NVIDIA may use your prompts and outputs to improve its models, and any model can be retired at any time. Keep sensitive and production work on a paid provider.
Often not the native window: the hosted deployment sets its own limit and the catalog pages misstate it. GLM 5.2 serves about 202K here against a 1M native window, and one model silently truncates instead of erroring. Big-AGI measures every limit against the live endpoint: the number in the table is the one you get.
Yes. Set a custom host on the same service: a NIM container or any OpenAI-compatible server (localhost:8000 style) works with identical model ids. Local hosts also unlock Direct Connection, browser straight to your box, which the hosted endpoint blocks.
Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how NVIDIA is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego