Use NVIDIA Models in Big-AGI.

Bring your own key: NVIDIA's API rates, no markup. Keys and chats stay in your browser. Run NVIDIA in parallel with other models, then compare and merge the answers.

Nemotron 3 Ultra
Nemotron 3.5 Content Safety (free) ·
Nemotron 3 Nano Omni 30B A3b Reasoning Fp8

Frontier open models, for $0.

Every model below runs at no cost on NVIDIA's hosted catalog: NVIDIA's own Nemotron 3, plus GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card, no credits to burn down. The throttle is on requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop you.

The limits are real and worth knowing: NVIDIA does not publish a rate, developers see somewhere around 40 requests per minute, and it moves with load. NVIDIA's trial terms also reserve the right to use your prompts and the answers to improve its products and models, and do not cover production use. Keep sensitive and production work on a paid provider.

All supported NVIDIA models

ModelContextInputOutputReleased

Nemotron 3 Ultra

ReasoningTools / functionsWeb search

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…

512K

$0.6

$3.6

Jun 2026

Nemotron 3.5 Content Safety (free) ·

VisionReasoningWeb search

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…

128K

-

-

Jun 2026

Nemotron 3 Nano Omni 30B A3b Reasoning Fp8

Nvidia chat model.

131K

-

-

Apr 2026

Nemotron 3 Nano Omni (free) ·

VisionReasoningTools / functionsWeb search

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…

256K

-

-

Apr 2026

Nvidia Nemotron 3 Super 120B A12b Bf16

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16

262K

-

-

Apr 2026

Nemotron 3 Super

ReasoningTools / functionsWeb search

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…

1M

$0.09

$0.4

Mar 2026

Nvidia Nemotron 3 Super 120B A12b Fp8

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

262K

-

-

Mar 2026

Nvidia Nemotron 3 Nano 30B A3b Bf16

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

262K

-

-

Dec 2025

Nemotron 3 Nano 30B A3B

ReasoningTools / functionsWeb search

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI system…

262K

$0.05

$0.2

Dec 2025

Nemotron Nano 12B 2 VL (free) ·

VisionReasoningTools / functionsWeb search

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a…

128K

-

-

Oct 2025

Llama 3.3 Nemotron Super 49B V1.5

deprecated
ReasoningTools / functionsWeb search

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s…

131K

$0.4

$0.4

Oct 2025

Nvidia Nemotron Nano 9B V2

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-Nano-9B-v2

131K

$0.06

$0.25

Sep 2025

Nemotron Nano 9B V2 (free) ·

ReasoningTools / functionsWeb search

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning…

128K

-

-

Sep 2025

Llama 3.1 Nemotron 70B Instruct HF

nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF

33K

$0.88

$0.88

Nov 2024

Nemotron 3 Ultra

Jun 2026

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…

ReasoningTools / functionsWeb search
512K · in $0.6 · out $3.6

Nemotron 3.5 Content Safety (free) ·

Jun 2026

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…

VisionReasoningWeb search
128K · in - · out -

Nemotron 3 Nano Omni 30B A3b Reasoning Fp8

Apr 2026

Nvidia chat model.

131K · in - · out -

Nemotron 3 Nano Omni (free) ·

Apr 2026

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…

VisionReasoningTools / functionsWeb search
256K · in - · out -

Nvidia Nemotron 3 Super 120B A12b Bf16

Apr 2026

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16

262K · in - · out -

Nemotron 3 Super

Mar 2026

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…

ReasoningTools / functionsWeb search
1M · in $0.09 · out $0.4

Nvidia Nemotron 3 Super 120B A12b Fp8

Mar 2026

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

262K · in - · out -

Nvidia Nemotron 3 Nano 30B A3b Bf16

Dec 2025

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16

262K · in - · out -

Nemotron 3 Nano 30B A3B

Dec 2025

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI system…

ReasoningTools / functionsWeb search
262K · in $0.05 · out $0.2

Nemotron Nano 12B 2 VL (free) ·

Oct 2025

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a…

VisionReasoningTools / functionsWeb search
128K · in - · out -

Llama 3.3 Nemotron Super 49B V1.5

deprecated
Oct 2025

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context. It’s…

ReasoningTools / functionsWeb search
131K · in $0.4 · out $0.4

Nvidia Nemotron Nano 9B V2

Sep 2025

Nvidia chat model. https://huggingface.co/api/models/nvidia/NVIDIA-Nemotron-Nano-9B-v2

131K · in $0.06 · out $0.25

Nemotron Nano 9B V2 (free) ·

Sep 2025

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning…

ReasoningTools / functionsWeb search
128K · in - · out -

Llama 3.1 Nemotron 70B Instruct HF

Nov 2024

nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemotron-70B-Instruct-HF

33K · in $0.88 · out $0.88
14 models · sorted by release date · prices in USD per 1M tokens · refreshed every 30 minutesCompare every model across vendors →

Get started in 3 steps

1

Create an API key at the NVIDIA console.

2

Paste it into Big-AGI's model settings.

3

Start chatting, or Beam it against other models and fuse the answers.

Running NVIDIA in Big-AGI

Create an nvapi- key at build.nvidia.com, keep the Public API Endpoints scope on it, and paste it into Big-AGI. Nothing is billed for these models, and Big-AGI adds no markup and no intermediary: the key, the account, and the limits are yours, direct from NVIDIA.

  • Rate limits, handled. The free endpoint throttles hard and unpredictably, so Big-AGI paces its calls and a Beam scatter queues instead of collecting 429s.
  • A catalog that answers. NVIDIA reserves the right to drop any model at any time, and its model list is a stale superset: roughly half the ids it returns no longer answer. Big-AGI ships a curated list refreshed against the live endpoint, so a dead id never reaches your model picker.
  • Context windows measured, not copied. The catalog pages misstate context length on roughly a quarter of these models, by as much as 8x, so Big-AGI probes the real limit against the endpoint.
  • Reasoning controls wired through. GPT-OSS keeps its native low, medium, and high effort setting; the Nemotron thinking models get a thinking toggle.

Why Big-AGI instead of the catalog playground?

build.nvidia.com is a per-model demo box: one prompt, one model, no history. Big-AGI turns the same key into a workspace with persistent chats, personas, and attachments, and puts a free NVIDIA model in the same conversation as Claude, GPT, or Gemini, with the reasoning dials the demo box hides.

Your keys and your data

The default NVIDIA endpoint refuses browser-origin requests, so Direct Connection is off there and calls route through the Big-AGI server. Point the same service at your own host, a self-hosted NIM or vLLM box such as a DGX Spark serving an OpenAI-compatible port on your LAN, and Direct Connection works. Either way your keys stay in your browser, chats are stored locally first and sync only if you turn it on, and the AI Inspector shows the exact request.

NVIDIA in Beam

A free catalog earns its keep in Beam: run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then reach for Fusions: several strategies that combine, cross-check, and synthesize the parallel answers, which beats just picking the best one. Parallel runs use more tokens than a single chat, and here they also spend more of that per-minute budget.

Bring your NVIDIA key. Keep control.

Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how NVIDIA is called.

© 2026 Token Fabrics·Built with passion in San Diego