Use NVIDIA Models in Big-AGI.

Bring your own key: free models on NVIDIA's own endpoint, no charges. Keys stay in your browser. Run NVIDIA in parallel with other models, then compare and merge the answers.

Ising Calibration 1.5 31B
Inkling
DeepSeek V4 Pro
DeepSeek V4 Flash
GLM 5.2

Frontier open models, for $0.

Every model below runs at no cost on NVIDIA's hosted catalog: Nemotron 3, GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card. The throttle counts requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop.

Two limits to know: expect roughly 40 requests per minute (NVIDIA publishes no number, and it moves with load), and NVIDIA's trial terms allow it to use your prompts to improve its models and exclude production use. Keep sensitive and production work on a paid provider.

All supported NVIDIA models

ModelContextInputOutputReleased

Ising Calibration 1.5 31B

NEW
Vision

NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).

131K

Free

Free

Jul 2026

Inkling

NEW
VisionReasoning

Thinking Machines reasoning model (preview). Can be unstable under load on the free endpoint.

131K

Free

Free

Jul 2026

Laguna XS 2.1

NEW
ReasoningTools / functions

Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.

262K

Free

Free

Jul 2026

GLM 5.2

HOT
ReasoningTools / functions

Zhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M).

203K

Free

Free

Jun 2026

DiffusionGemma 26B

VisionReasoningTools / functions

Experimental diffusion language model. CAUTION: prone to hanging under load.

250K

Free

Free

Jun 2026

Nemotron 3 Ultra 550B

ReasoningTools / functions

NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.

1M

Free

Free

Jun 2026

Nemotron 3.5 Content Safety

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Jun 2026

MiniMax M3

HOT
VisionReasoningTools / functions

MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 512K context (native: 1M).

524K

Free

Free

May 2026

Step 3.7 Flash

VisionReasoningTools / functions

StepFun fast multimodal reasoning model, 256K context.

262K

Free

Free

May 2026

Nemotron 3 Nano Omni 30B

VisionReasoningTools / functions

Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.

1M

Free

Free

Apr 2026

Mistral Medium 3.5

deprecated
VisionReasoningTools / functions

Mistral frontier-class multimodal model with adjustable reasoning.

262K

Free

Free

Apr 2026

DeepSeek V4 Pro

deprecated
ReasoningTools / functions

DeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.

262K

Free

Free

Apr 2026

DeepSeek V4 Flash

deprecated
ReasoningTools / functions

Fast DeepSeek V4 MoE (284B) with 1M context and reasoning.

1M

Free

Free

Apr 2026

Ising Calibration 1 35B

deprecated
VisionReasoning

NVIDIA quantum-calibration VLM (domain-specific, not for general chat).

262K

Free

Free

Apr 2026

Gemma 4 31B

HOT
VisionReasoningTools / functions

Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.

131K

Free

Free

Apr 2026

Nemotron 3 Super 120B

ReasoningTools / functions

Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.

1M

Free

Free

Mar 2026

Qwen 3.5 397B

deprecated
VisionReasoningTools / functions

Qwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.

262K

Free

Free

Feb 2026

Nemotron 3 Nano 30B

ReasoningTools / functions

Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use.

1M

Free

Free

Dec 2025

Riva Translate 4B

Translation-specialized model, 8K context.

8K

Free

Free

Dec 2025

Nemotron Nano 12B v2 VL

VisionReasoningTools / functions

Vision-language Nemotron Nano for image and video understanding (verified image input).

131K

Free

Free

Oct 2025

Nemotron Safety Guard 8B v3

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Oct 2025

Llama 3.3 Nemotron Super 49B v1.5

ReasoningTools / functions

Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks.

131K

Free

Free

Oct 2025

Nemotron Nano 9B v2

ReasoningTools / functions

Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use.

128K

Free

Free

Aug 2025

GPT-OSS 20B

ReasoningTools / functions

OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.

131K

Free

Free

Aug 2025

GPT-OSS 120B

ReasoningTools / functions

OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.

131K

Free

Free

Aug 2025

Mistral Nemotron

ReasoningTools / functions

Mistral model post-trained by NVIDIA.

262K

Free

Free

Jun 2025

Nemotron Nano VL 8B

Vision

Small vision-language model, 16K context.

16K

Free

Free

May 2025

Llama Guard 4 12B

Vision

Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.

66K

Free

Free

Apr 2025

Llama 3.3 Nemotron Super 49B v1

ReasoningTools / functions

Superseded by v1.5.

131K

Free

Free

Mar 2025

NemoGuard 8B Content Safety

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Jan 2025

Llama 3.3 70B

Tools / functions

Meta Llama 3.3 70B instruction-tuned, with tool calling.

131K

Free

Free

Dec 2024

Llama 3.2 1B

Tools / functions

Tiny Llama for edge-class tasks.

131K

Free

Free

Sep 2024

Llama 3.2 11B Vision

Vision

Llama vision model for image understanding.

131K

Free

Free

Sep 2024

Llama 3.2 3B

Tools / functions

Small Llama for lightweight tasks.

131K

Free

Free

Sep 2024

Llama 3.2 90B Vision

Vision

Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).

33K

Free

Free

Sep 2024

Nemotron Mini 4B

Tools / functions

Tiny legacy Nemotron, 4K context.

4K

Free

Free

Sep 2024

Llama 3.1 70B

Tools / functions

Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).

131K

Free

Free

Jul 2024

Llama 3.1 8B

Tools / functions

Small fast Llama for utility tasks, with tool calling.

131K

Free

Free

Jul 2024

Ising Calibration 1.5 31B

NEW
Jul 2026

NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).

Vision
131K · in Free · out Free

Inkling

NEW
Jul 2026

Thinking Machines reasoning model (preview). Can be unstable under load on the free endpoint.

VisionReasoning
131K · in Free · out Free

Laguna XS 2.1

NEW
Jul 2026

Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.

ReasoningTools / functions
262K · in Free · out Free

GLM 5.2

HOT
Jun 2026

Zhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M).

ReasoningTools / functions
203K · in Free · out Free

DiffusionGemma 26B

Jun 2026

Experimental diffusion language model. CAUTION: prone to hanging under load.

VisionReasoningTools / functions
250K · in Free · out Free

Nemotron 3 Ultra 550B

Jun 2026

NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.

ReasoningTools / functions
1M · in Free · out Free

Nemotron 3.5 Content Safety

Jun 2026

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

MiniMax M3

HOT
May 2026

MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 512K context (native: 1M).

VisionReasoningTools / functions
524K · in Free · out Free

Step 3.7 Flash

May 2026

StepFun fast multimodal reasoning model, 256K context.

VisionReasoningTools / functions
262K · in Free · out Free

Nemotron 3 Nano Omni 30B

Apr 2026

Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.

VisionReasoningTools / functions
1M · in Free · out Free

Mistral Medium 3.5

deprecated
Apr 2026

Mistral frontier-class multimodal model with adjustable reasoning.

VisionReasoningTools / functions
262K · in Free · out Free

DeepSeek V4 Pro

deprecated
Apr 2026

DeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.

ReasoningTools / functions
262K · in Free · out Free

DeepSeek V4 Flash

deprecated
Apr 2026

Fast DeepSeek V4 MoE (284B) with 1M context and reasoning.

ReasoningTools / functions
1M · in Free · out Free

Ising Calibration 1 35B

deprecated
Apr 2026

NVIDIA quantum-calibration VLM (domain-specific, not for general chat).

VisionReasoning
262K · in Free · out Free

Gemma 4 31B

HOT
Apr 2026

Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.

VisionReasoningTools / functions
131K · in Free · out Free

Nemotron 3 Super 120B

Mar 2026

Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.

ReasoningTools / functions
1M · in Free · out Free

Qwen 3.5 397B

deprecated
Feb 2026

Qwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.

VisionReasoningTools / functions
262K · in Free · out Free

Nemotron 3 Nano 30B

Dec 2025

Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use.

ReasoningTools / functions
1M · in Free · out Free

Riva Translate 4B

Dec 2025

Translation-specialized model, 8K context.

8K · in Free · out Free

Nemotron Nano 12B v2 VL

Oct 2025

Vision-language Nemotron Nano for image and video understanding (verified image input).

VisionReasoningTools / functions
131K · in Free · out Free

Nemotron Safety Guard 8B v3

Oct 2025

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

Llama 3.3 Nemotron Super 49B v1.5

Oct 2025

Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks.

ReasoningTools / functions
131K · in Free · out Free

Nemotron Nano 9B v2

Aug 2025

Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use.

ReasoningTools / functions
128K · in Free · out Free

GPT-OSS 20B

Aug 2025

OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.

ReasoningTools / functions
131K · in Free · out Free

GPT-OSS 120B

Aug 2025

OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.

ReasoningTools / functions
131K · in Free · out Free

Mistral Nemotron

Jun 2025

Mistral model post-trained by NVIDIA.

ReasoningTools / functions
262K · in Free · out Free

Nemotron Nano VL 8B

May 2025

Small vision-language model, 16K context.

Vision
16K · in Free · out Free

Llama Guard 4 12B

Apr 2025

Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.

Vision
66K · in Free · out Free

Llama 3.3 Nemotron Super 49B v1

Mar 2025

Superseded by v1.5.

ReasoningTools / functions
131K · in Free · out Free

NemoGuard 8B Content Safety

Jan 2025

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

Llama 3.3 70B

Dec 2024

Meta Llama 3.3 70B instruction-tuned, with tool calling.

Tools / functions
131K · in Free · out Free

Llama 3.2 1B

Sep 2024

Tiny Llama for edge-class tasks.

Tools / functions
131K · in Free · out Free

Llama 3.2 11B Vision

Sep 2024

Llama vision model for image understanding.

Vision
131K · in Free · out Free

Llama 3.2 3B

Sep 2024

Small Llama for lightweight tasks.

Tools / functions
131K · in Free · out Free

Llama 3.2 90B Vision

Sep 2024

Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).

Vision
33K · in Free · out Free

Nemotron Mini 4B

Sep 2024

Tiny legacy Nemotron, 4K context.

Tools / functions
4K · in Free · out Free

Llama 3.1 70B

Jul 2024

Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).

Tools / functions
131K · in Free · out Free

Llama 3.1 8B

Jul 2024

Small fast Llama for utility tasks, with tool calling.

Tools / functions
131K · in Free · out Free
38 models · sorted by release date · prices in USD per 1M tokens · refreshed hourlyCompare every model across vendors ->

Get started in 3 steps

1

Create an API key at the NVIDIA console.

2

Paste it into Big-AGI's model settings.

3

Start chatting, or Beam it against other models and fuse the answers.

Setup, in three steps

  1. Create a free nvapi- key at build.nvidia.com, keeping the Public API Endpoints scope on it.
  2. In Big-AGI, add the NVIDIA NIM service and paste the key.
  3. Load the models and chat. No billing, no markup: the key, the account, and the limits are yours, direct from NVIDIA.

The build.nvidia.com playground is one prompt, one model, no history. The same key here gets persistent chats, personas, attachments, and NVIDIA models in the same conversation as Claude, GPT, or Gemini.

What Big-AGI fixes for you

  • Dead models filtered. Roughly half the ids in NVIDIA's catalog no longer answer. The curated list ships only live-verified models, so a dead id never reaches your picker.
  • Real context windows. Measured against the endpoint, not copied from catalog pages, which misstate about a quarter of them.
  • Rate limits paced. Calls queue instead of collecting 429s, including Beam scatters.
  • Reasoning wired through. GPT-OSS effort levels and Nemotron thinking toggles, in the chat controls.

Use it in Beam

Run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then Fuse: combine, cross-check, and synthesize the parallel answers instead of just picking one.

Local NIM and DGX Spark

Point the same service at your own host, a NIM container or any OpenAI-compatible server (http://spark-xxxx.local:8000 style): identical model ids, same protocol. Direct Connection works on local hosts; the hosted endpoint blocks browser-origin calls, so those route through the server. Either way keys stay in your browser and chats are stored locally first.

Which free model for which job?

  • Deep reasoning: Nemotron 3 Ultra 550B, 1M context with tool calling.
  • Fast agent loops: Nemotron 3 Nano 30B (1M) or GPT-OSS 20B: small models keep a loop moving under a per-minute request cap.
  • Coding: GPT-OSS 120B, or DeepSeek V4 Flash when the whole repository has to fit in the prompt.
  • Vision: Nemotron Nano 12B v2 VL for images and video (131K), or MiniMax M3 (524K) for text-heavy pages.
  • Long documents: DeepSeek V4 Flash or Nemotron 3 Super 120B, both listed at 1M.

NVIDIA questions

FAQ

Why does an NVIDIA model return 404 "Not found for account"?

Either the model is retired (NVIDIA's catalog still lists it, but roughly half the ids no longer serve), or your key was created without the "Public API Endpoints" scope, which fails on every model. Fix the key at build.nvidia.com. Big-AGI only lists live-verified models, so a retired id never reaches your picker.

What is the rate limit on integrate.api.nvidia.com?

NVIDIA does not publish one. Expect roughly 40 requests per minute per account, moving with load. It counts requests, not tokens, so long prompts and answers are not penalized. Big-AGI paces its calls, so bursts and Beam scatters queue instead of collecting 429s.

Is the NVIDIA API really free? What is the catch?

Yes, no billing exists on the hosted endpoint. The catch is in the terms, not the invoice: no production use, NVIDIA may use your prompts and outputs to improve its models, and any model can be retired at any time. Keep sensitive and production work on a paid provider.

What context windows do the NVIDIA-hosted models actually have?

Often not the native window: the hosted deployment sets its own limit and the catalog pages misstate it. GLM 5.2 serves about 202K here against a 1M native window, and one model silently truncates instead of erroring. Big-AGI measures every limit against the live endpoint: the number in the table is the one you get.

Can I point it at a local NIM or a DGX Spark?

Yes. Set a custom host on the same service: a NIM container or any OpenAI-compatible server (localhost:8000 style) works with identical model ids. Local hosts also unlock Direct Connection, browser straight to your box, which the hosted endpoint blocks.

Bring your NVIDIA key. Keep control.

Your key, your data, your choice of model. Big-AGI is open source and self-hostable, so you can check exactly how NVIDIA is called.

Launch Big-AGI

© 2026 Token Fabrics·Built with passion in San Diego