Nemotron, GPT-OSS, DeepSeek, GLM - every model here costs $0 on one free nvapi- key. The catch is in the terms, not the invoice.
Every model below runs at no cost on NVIDIA's hosted catalog: Nemotron 3, GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card. The throttle counts requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop.
Two limits to know: expect roughly 40 requests per minute (NVIDIA publishes no number, and it moves with load), and NVIDIA's trial terms allow it to use your prompts to improve its models and exclude production use. Keep sensitive and production work on a paid provider.
DeepSeek V4 Pro 0813
NEWDeepSeek flagship reasoning MoE (official 0813 release) with the full 1M context. Can be slow to cold-start on the free endpoint.
1M
Free
Free
Aug 2026
Nemotron 3.5 Lightning 30B
NEWFastest Nemotron MoE (30B, 3B active), 1M context, reasoning and tool use. Text only.
1M
Free
Free
Aug 2026
Muse Glimmer 30B
NEWMeta open multimodal reasoning model (30B dense) distilled from Muse Spark, with tool calling. Always reasons.
131K
Free
Free
Aug 2026
DeepSeek V4 Flash 0731
NEWFast DeepSeek V4 MoE (284B, 13B active) with the full 1M context and reasoning.
1M
Free
Free
Jul 2026
Riva Translate 4B v2
NEWTranslation-specialized model (37 languages, few-shot prompting), 8K context.
8K
Free
Free
Jul 2026
Ising Calibration 1.5 31B
NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).
131K
Free
Free
Jul 2026
Inkling
deprecatedThinking Machines reasoning model (preview). Can be unstable under load on the free endpoint. Retires on NVIDIA 2026-08-24.
131K
Free
Free
Jul 2026
Kimi K3
HOTMoonshot flagship MoE, natively multimodal (image inputs), with the full 1M context and reasoning. Currently degraded on NVIDIA.
1M
Free
Free
Jul 2026
Laguna XS 2.1
Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.
262K
Free
Free
Jul 2026
GLM 5.2
deprecatedZhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M). Retires on NVIDIA 2026-08-24.
203K
Free
Free
Jun 2026
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hanging under load.
250K
Free
Free
Jun 2026
Nemotron 3 Ultra 550B
HOTNVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.
1M
Free
Free
Jun 2026
Nemotron 3.5 Content Safety
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Jun 2026
MiniMax M3
MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 256K context (native: 1M). Retires on NVIDIA 2026-09-08.
262K
Free
Free
May 2026
Step 3.7 Flash
deprecatedStepFun fast multimodal reasoning model, 256K context.
262K
Free
Free
May 2026
Nemotron 3 Nano Omni 30B
Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.
1M
Free
Free
Apr 2026
Mistral Medium 3.5
deprecatedMistral frontier-class multimodal model with adjustable reasoning.
262K
Free
Free
Apr 2026
DeepSeek V4 Flash
deprecatedFast DeepSeek V4 MoE (284B) with 1M context and reasoning.
1M
Free
Free
Apr 2026
DeepSeek V4 Pro
deprecatedDeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.
262K
Free
Free
Apr 2026
Ising Calibration 1 35B
deprecatedNVIDIA quantum-calibration VLM (domain-specific, not for general chat).
262K
Free
Free
Apr 2026
Gemma 4 31B
Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.
131K
Free
Free
Apr 2026
Nemotron 3 Super 120B
Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.
1M
Free
Free
Mar 2026
Qwen 3.5 397B
deprecatedQwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.
262K
Free
Free
Feb 2026
Nemotron 3 Nano 30B
deprecatedEfficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use. Retires on NVIDIA 2026-08-31.
1M
Free
Free
Dec 2025
Riva Translate 4B v1.1
Translation-specialized model, 8K context.
8K
Free
Free
Dec 2025
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Oct 2025
Nemotron Nano 12B v2 VL
deprecatedVision-language Nemotron Nano for image and video understanding (verified image input). Retires on NVIDIA 2026-08-25.
131K
Free
Free
Oct 2025
Llama 3.3 Nemotron Super 49B v1.5
deprecatedLlama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks. Retires on NVIDIA 2026-08-25.
131K
Free
Free
Oct 2025
Nemotron Nano 9B v2
deprecatedSmall hybrid Mamba-Transformer for fast, cheap reasoning and tool use. Retires on NVIDIA 2026-08-25.
128K
Free
Free
Aug 2025
GPT-OSS 120B
deprecatedOpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.
131K
Free
Free
Aug 2025
GPT-OSS 20B
OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.
131K
Free
Free
Aug 2025
Mistral Nemotron
Mistral model post-trained by NVIDIA.
262K
Free
Free
Jun 2025
Nemotron Nano VL 8B
deprecatedSmall vision-language model, 16K context.
16K
Free
Free
May 2025
Llama Guard 4 12B
Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.
66K
Free
Free
Apr 2025
Llama 3.3 Nemotron Super 49B v1
deprecatedSuperseded by v1.5.
131K
Free
Free
Mar 2025
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
Jan 2025
NemoGuard 8B Topic Control
NVIDIA topic-control guardrail (not a chat model).
131K
Free
Free
Jan 2025
Llama 3.3 70B
deprecatedMeta Llama 3.3 70B instruction-tuned, with tool calling. Retires on NVIDIA 2026-08-25 and is already unresponsive.
131K
Free
Free
Dec 2024
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
33K
Free
Free
Sep 2024
Llama 3.2 11B Vision
Llama vision model for image understanding.
131K
Free
Free
Sep 2024
Llama 3.2 1B
deprecatedTiny Llama for edge-class tasks.
131K
Free
Free
Sep 2024
Llama 3.2 3B
deprecatedSmall Llama for lightweight tasks.
131K
Free
Free
Sep 2024
Nemotron Mini 4B
deprecatedTiny legacy Nemotron, 4K context.
4K
Free
Free
Sep 2024
Llama 3.1 70B
deprecatedMeta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
131K
Free
Free
Jul 2024
Llama 3.1 8B
deprecatedSmall fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.
131K
Free
Free
Jul 2024
DeepSeek V4 Pro 0813
NEWDeepSeek flagship reasoning MoE (official 0813 release) with the full 1M context. Can be slow to cold-start on the free endpoint.
Nemotron 3.5 Lightning 30B
NEWFastest Nemotron MoE (30B, 3B active), 1M context, reasoning and tool use. Text only.
Muse Glimmer 30B
NEWMeta open multimodal reasoning model (30B dense) distilled from Muse Spark, with tool calling. Always reasons.
DeepSeek V4 Flash 0731
NEWFast DeepSeek V4 MoE (284B, 13B active) with the full 1M context and reasoning.
Riva Translate 4B v2
NEWTranslation-specialized model (37 languages, few-shot prompting), 8K context.
Ising Calibration 1.5 31B
NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).
Inkling
deprecatedThinking Machines reasoning model (preview). Can be unstable under load on the free endpoint. Retires on NVIDIA 2026-08-24.
Kimi K3
HOTMoonshot flagship MoE, natively multimodal (image inputs), with the full 1M context and reasoning. Currently degraded on NVIDIA.
Laguna XS 2.1
Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.
GLM 5.2
deprecatedZhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M). Retires on NVIDIA 2026-08-24.
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hanging under load.
Nemotron 3 Ultra 550B
HOTNVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.
Nemotron 3.5 Content Safety
NVIDIA content-safety classifier (not a chat model).
MiniMax M3
MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 256K context (native: 1M). Retires on NVIDIA 2026-09-08.
Step 3.7 Flash
deprecatedStepFun fast multimodal reasoning model, 256K context.
Nemotron 3 Nano Omni 30B
Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.
Mistral Medium 3.5
deprecatedMistral frontier-class multimodal model with adjustable reasoning.
DeepSeek V4 Flash
deprecatedFast DeepSeek V4 MoE (284B) with 1M context and reasoning.
DeepSeek V4 Pro
deprecatedDeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.
Ising Calibration 1 35B
deprecatedNVIDIA quantum-calibration VLM (domain-specific, not for general chat).
Gemma 4 31B
Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.
Nemotron 3 Super 120B
Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.
Qwen 3.5 397B
deprecatedQwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.
Nemotron 3 Nano 30B
deprecatedEfficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use. Retires on NVIDIA 2026-08-31.
Riva Translate 4B v1.1
Translation-specialized model, 8K context.
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
Nemotron Nano 12B v2 VL
deprecatedVision-language Nemotron Nano for image and video understanding (verified image input). Retires on NVIDIA 2026-08-25.
Llama 3.3 Nemotron Super 49B v1.5
deprecatedLlama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks. Retires on NVIDIA 2026-08-25.
Nemotron Nano 9B v2
deprecatedSmall hybrid Mamba-Transformer for fast, cheap reasoning and tool use. Retires on NVIDIA 2026-08-25.
GPT-OSS 120B
deprecatedOpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.
GPT-OSS 20B
OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.
Mistral Nemotron
Mistral model post-trained by NVIDIA.
Nemotron Nano VL 8B
deprecatedSmall vision-language model, 16K context.
Llama Guard 4 12B
Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.
Llama 3.3 Nemotron Super 49B v1
deprecatedSuperseded by v1.5.
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
NemoGuard 8B Topic Control
NVIDIA topic-control guardrail (not a chat model).
Llama 3.3 70B
deprecatedMeta Llama 3.3 70B instruction-tuned, with tool calling. Retires on NVIDIA 2026-08-25 and is already unresponsive.
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
Llama 3.2 11B Vision
Llama vision model for image understanding.
Llama 3.2 1B
deprecatedTiny Llama for edge-class tasks.
Llama 3.2 3B
deprecatedSmall Llama for lightweight tasks.
Nemotron Mini 4B
deprecatedTiny legacy Nemotron, 4K context.
Llama 3.1 70B
deprecatedMeta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
Llama 3.1 8B
deprecatedSmall fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.
1
Create an API key at the NVIDIA console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
nvapi- key at build.nvidia.com, keeping the Public API Endpoints scope on it.The build.nvidia.com playground is one prompt, one model, no history. The same key here gets persistent chats, personas, attachments, and NVIDIA models in the same conversation as Claude, GPT, or Gemini.
Run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then Fuse: combine, cross-check, and synthesize the parallel answers instead of just picking one.
Point the same service at your own host, a NIM container or any OpenAI-compatible server (http://spark-xxxx.local:8000 style): identical model ids, same protocol. Direct Connection works on local hosts; the hosted endpoint blocks browser-origin calls, so those route through the server. Either way keys stay in your browser and chats are stored locally first.
FAQ
Either the model is retired (NVIDIA's catalog still lists it, but roughly half the ids no longer serve), or your key was created without the "Public API Endpoints" scope, which fails on every model. Fix the key at build.nvidia.com. Big-AGI only lists live-verified models, so a retired id never reaches your picker.
NVIDIA does not publish one. Expect roughly 40 requests per minute per account, moving with load. It counts requests, not tokens, so long prompts and answers are not penalized. Big-AGI paces its calls, so bursts and Beam scatters queue instead of collecting 429s.
Yes, no billing exists on the hosted endpoint. The catch is in the terms, not the invoice: no production use, NVIDIA may use your prompts and outputs to improve its models, and any model can be retired at any time. Keep sensitive and production work on a paid provider.
Often not the native window: the hosted deployment sets its own limit and the catalog pages misstate it. GLM 5.2 serves about 202K here against a 1M native window, and one model silently truncates instead of erroring. Big-AGI measures every limit against the live endpoint: the number in the table is the one you get.
Yes. Set a custom host on the same service: a NIM container or any OpenAI-compatible server (localhost:8000 style) works with identical model ids. Local hosts also unlock Direct Connection, browser straight to your box, which the hosted endpoint blocks.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how NVIDIA is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego