Bring your own key: NVIDIA's API rates, no markup. Keys and chats stay in your browser. Run NVIDIA in parallel with other models, then compare and merge the answers.
Every model below runs at no cost on NVIDIA's hosted catalog: NVIDIA's own Nemotron 3, plus GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card, no credits to burn down. The throttle is on requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop you.
The limits are real and worth knowing: NVIDIA does not publish a rate, developers see somewhere around 40 requests per minute, and it moves with load. NVIDIA's trial terms also reserve the right to use your prompts and the answers to improve its products and models, and do not cover production use. Keep sensitive and production work on a paid provider.
Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…
512K
$0.6
$3.6
Jun 2026
Nemotron 3.5 Content Safety (free) ·
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…
128K
-
-
Jun 2026
Nemotron 3 Nano Omni 30B A3b Reasoning Fp8
Nvidia chat model.
131K
-
-
Apr 2026
Nemotron 3 Nano Omni (free) ·
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…
256K
-
-
Apr 2026
Nvidia Nemotron 3 Super 120B A12b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
262K
-
-
Apr 2026
Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…
1M
$0.09
$0.4
Mar 2026
Nvidia Nemotron 3 Super 120B A12b Fp8
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
262K
-
-
Mar 2026
Nvidia Nemotron 3 Nano 30B A3b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
262K
-
-
Dec 2025
Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI system…
262K
$0.05
$0.2
Dec 2025
Nemotron Nano 12B 2 VL (free) ·
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a…
128K
-
-
Oct 2025
Llama 3.3 Nemotron Super 49B V1.5
deprecatedLlama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s…
131K
$0.4
$0.4
Oct 2025
Nvidia Nemotron Nano 9B V2
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-Nano-9B-v2
131K
$0.06
$0.25
Sep 2025
Nemotron Nano 9B V2 (free) ·
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning…
128K
-
-
Sep 2025
Llama 3.1 Nemotron 70B Instruct HF
nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
33K
$0.88
$0.88
Nov 2024
Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…
Nemotron 3.5 Content Safety (free) ·
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…
Nemotron 3 Nano Omni 30B A3b Reasoning Fp8
Nvidia chat model.
Nemotron 3 Nano Omni (free) ·
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…
Nvidia Nemotron 3 Super 120B A12b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16
Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…
Nvidia Nemotron 3 Super 120B A12b Fp8
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8
Nvidia Nemotron 3 Nano 30B A3b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI system…
Nemotron Nano 12B 2 VL (free) ·
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a…
Llama 3.3 Nemotron Super 49B V1.5
deprecatedLlama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s…
Nvidia Nemotron Nano 9B V2
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-Nano-9B-v2
Nemotron Nano 9B V2 (free) ·
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning…
Llama 3.1 Nemotron 70B Instruct HF
nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF
1
Create an API key at the NVIDIA console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Create an nvapi- key at build.nvidia.com, keep the Public API Endpoints scope on it, and paste it into Big-AGI. Nothing is billed for these models, and Big-AGI adds no markup and no intermediary: the key, the account, and the limits are yours, direct from NVIDIA.
build.nvidia.com is a per-model demo box: one prompt, one model, no history. Big-AGI turns the same key into a workspace with persistent chats, personas, and attachments, and puts a free NVIDIA model in the same conversation as Claude, GPT, or Gemini, with the reasoning dials the demo box hides.
The default NVIDIA endpoint refuses browser-origin requests, so Direct Connection is off there and calls route through the Big-AGI server. Point the same service at your own host, a self-hosted NIM or vLLM box such as a DGX Spark serving an OpenAI-compatible port on your LAN, and Direct Connection works. Either way your keys stay in your browser, chats are stored locally first and sync only if you turn it on, and the AI Inspector shows the exact request.
A free catalog earns its keep in Beam: run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then reach for Fusions: several strategies that combine, cross-check, and synthesize the parallel answers, which beats just picking the best one. Parallel runs use more tokens than a single chat, and here they also spend more of that per-minute budget.
Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how NVIDIA is called.
BIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego