Use NVIDIA Models in Big-AGI.

Nemotron, GPT-OSS, DeepSeek, GLM - every model here costs $0 on one free nvapi- key. The catch is in the terms, not the invoice.

Frontier open models, for $0.

Every model below runs at no cost on NVIDIA's hosted catalog: Nemotron 3, GPT-OSS, DeepSeek V4, GLM, MiniMax, Mistral, Gemma, and Llama. One free nvapi- key, no card. The throttle counts requests, not tokens, so long prompts and long answers go through where token-capped free tiers stop.

Two limits to know: expect roughly 40 requests per minute (NVIDIA publishes no number, and it moves with load), and NVIDIA's trial terms allow it to use your prompts to improve its models and exclude production use. Keep sensitive and production work on a paid provider.

All supported NVIDIA models

ModelContextInputOutputReleased

DeepSeek V4 Pro 0813

NEW
ReasoningTools / functions

DeepSeek flagship reasoning MoE (official 0813 release) with the full 1M context. Can be slow to cold-start on the free endpoint.

1M

Free

Free

Aug 2026

Nemotron 3.5 Lightning 30B

NEW
ReasoningTools / functions

Fastest Nemotron MoE (30B, 3B active), 1M context, reasoning and tool use. Text only.

1M

Free

Free

Aug 2026

Muse Glimmer 30B

NEW
VisionReasoningTools / functions

Meta open multimodal reasoning model (30B dense) distilled from Muse Spark, with tool calling. Always reasons.

131K

Free

Free

Aug 2026

DeepSeek V4 Flash 0731

NEW
ReasoningTools / functions

Fast DeepSeek V4 MoE (284B, 13B active) with the full 1M context and reasoning.

1M

Free

Free

Jul 2026

Riva Translate 4B v2

NEW

Translation-specialized model (37 languages, few-shot prompting), 8K context.

8K

Free

Free

Jul 2026

Ising Calibration 1.5 31B

Vision

NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).

131K

Free

Free

Jul 2026

Inkling

deprecated
VisionReasoning

Thinking Machines reasoning model (preview). Can be unstable under load on the free endpoint. Retires on NVIDIA 2026-08-24.

131K

Free

Free

Jul 2026

Kimi K3

HOT
VisionReasoningTools / functions

Moonshot flagship MoE, natively multimodal (image inputs), with the full 1M context and reasoning. Currently degraded on NVIDIA.

1M

Free

Free

Jul 2026

Laguna XS 2.1

ReasoningTools / functions

Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.

262K

Free

Free

Jul 2026

GLM 5.2

deprecated
ReasoningTools / functions

Zhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M). Retires on NVIDIA 2026-08-24.

203K

Free

Free

Jun 2026

DiffusionGemma 26B

VisionReasoningTools / functions

Experimental diffusion language model. CAUTION: prone to hanging under load.

250K

Free

Free

Jun 2026

Nemotron 3 Ultra 550B

HOT
ReasoningTools / functions

NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.

1M

Free

Free

Jun 2026

Nemotron 3.5 Content Safety

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Jun 2026

MiniMax M3

VisionReasoningTools / functions

MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 256K context (native: 1M). Retires on NVIDIA 2026-09-08.

262K

Free

Free

May 2026

Step 3.7 Flash

deprecated
VisionReasoningTools / functions

StepFun fast multimodal reasoning model, 256K context.

262K

Free

Free

May 2026

Nemotron 3 Nano Omni 30B

VisionReasoningTools / functions

Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.

1M

Free

Free

Apr 2026

Mistral Medium 3.5

deprecated
VisionReasoningTools / functions

Mistral frontier-class multimodal model with adjustable reasoning.

262K

Free

Free

Apr 2026

DeepSeek V4 Flash

deprecated
ReasoningTools / functions

Fast DeepSeek V4 MoE (284B) with 1M context and reasoning.

1M

Free

Free

Apr 2026

DeepSeek V4 Pro

deprecated
ReasoningTools / functions

DeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.

262K

Free

Free

Apr 2026

Ising Calibration 1 35B

deprecated
VisionReasoning

NVIDIA quantum-calibration VLM (domain-specific, not for general chat).

262K

Free

Free

Apr 2026

Gemma 4 31B

VisionReasoningTools / functions

Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.

131K

Free

Free

Apr 2026

Nemotron 3 Super 120B

ReasoningTools / functions

Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.

1M

Free

Free

Mar 2026

Qwen 3.5 397B

deprecated
VisionReasoningTools / functions

Qwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.

262K

Free

Free

Feb 2026

Nemotron 3 Nano 30B

deprecated
ReasoningTools / functions

Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use. Retires on NVIDIA 2026-08-31.

1M

Free

Free

Dec 2025

Riva Translate 4B v1.1

Translation-specialized model, 8K context.

8K

Free

Free

Dec 2025

Nemotron Safety Guard 8B v3

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Oct 2025

Nemotron Nano 12B v2 VL

deprecated
VisionReasoningTools / functions

Vision-language Nemotron Nano for image and video understanding (verified image input). Retires on NVIDIA 2026-08-25.

131K

Free

Free

Oct 2025

Llama 3.3 Nemotron Super 49B v1.5

deprecated
ReasoningTools / functions

Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks. Retires on NVIDIA 2026-08-25.

131K

Free

Free

Oct 2025

Nemotron Nano 9B v2

deprecated
ReasoningTools / functions

Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use. Retires on NVIDIA 2026-08-25.

128K

Free

Free

Aug 2025

GPT-OSS 120B

deprecated
ReasoningTools / functions

OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.

131K

Free

Free

Aug 2025

GPT-OSS 20B

ReasoningTools / functions

OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.

131K

Free

Free

Aug 2025

Mistral Nemotron

ReasoningTools / functions

Mistral model post-trained by NVIDIA.

262K

Free

Free

Jun 2025

Nemotron Nano VL 8B

deprecated
Vision

Small vision-language model, 16K context.

16K

Free

Free

May 2025

Llama Guard 4 12B

Vision

Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.

66K

Free

Free

Apr 2025

Llama 3.3 Nemotron Super 49B v1

deprecated
ReasoningTools / functions

Superseded by v1.5.

131K

Free

Free

Mar 2025

NemoGuard 8B Content Safety

NVIDIA content-safety classifier (not a chat model).

131K

Free

Free

Jan 2025

NemoGuard 8B Topic Control

NVIDIA topic-control guardrail (not a chat model).

131K

Free

Free

Jan 2025

Llama 3.3 70B

deprecated
Tools / functions

Meta Llama 3.3 70B instruction-tuned, with tool calling. Retires on NVIDIA 2026-08-25 and is already unresponsive.

131K

Free

Free

Dec 2024

Llama 3.2 90B Vision

Vision

Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).

33K

Free

Free

Sep 2024

Llama 3.2 11B Vision

Vision

Llama vision model for image understanding.

131K

Free

Free

Sep 2024

Llama 3.2 1B

deprecated
Tools / functions

Tiny Llama for edge-class tasks.

131K

Free

Free

Sep 2024

Llama 3.2 3B

deprecated
Tools / functions

Small Llama for lightweight tasks.

131K

Free

Free

Sep 2024

Nemotron Mini 4B

deprecated
Tools / functions

Tiny legacy Nemotron, 4K context.

4K

Free

Free

Sep 2024

Llama 3.1 70B

deprecated
Tools / functions

Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).

131K

Free

Free

Jul 2024

Llama 3.1 8B

deprecated
Tools / functions

Small fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.

131K

Free

Free

Jul 2024

DeepSeek V4 Pro 0813

NEW
Aug 2026

DeepSeek flagship reasoning MoE (official 0813 release) with the full 1M context. Can be slow to cold-start on the free endpoint.

ReasoningTools / functions
1M · in Free · out Free

Nemotron 3.5 Lightning 30B

NEW
Aug 2026

Fastest Nemotron MoE (30B, 3B active), 1M context, reasoning and tool use. Text only.

ReasoningTools / functions
1M · in Free · out Free

Muse Glimmer 30B

NEW
Aug 2026

Meta open multimodal reasoning model (30B dense) distilled from Muse Spark, with tool calling. Always reasons.

VisionReasoningTools / functions
131K · in Free · out Free

DeepSeek V4 Flash 0731

NEW
Jul 2026

Fast DeepSeek V4 MoE (284B, 13B active) with the full 1M context and reasoning.

ReasoningTools / functions
1M · in Free · out Free

Riva Translate 4B v2

NEW
Jul 2026

Translation-specialized model (37 languages, few-shot prompting), 8K context.

8K · in Free · out Free

Ising Calibration 1.5 31B

Jul 2026

NVIDIA quantum-calibration VLM, preview (domain-specific, not for general chat).

Vision
131K · in Free · out Free

Inkling

deprecated
Jul 2026

Thinking Machines reasoning model (preview). Can be unstable under load on the free endpoint. Retires on NVIDIA 2026-08-24.

VisionReasoning
131K · in Free · out Free

Kimi K3

HOT
Jul 2026

Moonshot flagship MoE, natively multimodal (image inputs), with the full 1M context and reasoning. Currently degraded on NVIDIA.

VisionReasoningTools / functions
1M · in Free · out Free

Laguna XS 2.1

Jul 2026

Poolside compact coding-focused reasoning model. Can be slow to cold-start on the free endpoint.

ReasoningTools / functions
262K · in Free · out Free

GLM 5.2

deprecated
Jun 2026

Zhipu GLM 5.2 reasoning model. NVIDIA serves a reduced 202K context (native: 1M). Retires on NVIDIA 2026-08-24.

ReasoningTools / functions
203K · in Free · out Free

DiffusionGemma 26B

Jun 2026

Experimental diffusion language model. CAUTION: prone to hanging under load.

VisionReasoningTools / functions
250K · in Free · out Free

Nemotron 3 Ultra 550B

HOT
Jun 2026

NVIDIA flagship open hybrid Mamba-Transformer MoE (550B, 55B active), 1M context, reasoning and tool use.

ReasoningTools / functions
1M · in Free · out Free

Nemotron 3.5 Content Safety

Jun 2026

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

MiniMax M3

May 2026

MiniMax M3 reasoning model with image inputs. NVIDIA serves a reduced 256K context (native: 1M). Retires on NVIDIA 2026-09-08.

VisionReasoningTools / functions
262K · in Free · out Free

Step 3.7 Flash

deprecated
May 2026

StepFun fast multimodal reasoning model, 256K context.

VisionReasoningTools / functions
262K · in Free · out Free

Nemotron 3 Nano Omni 30B

Apr 2026

Omni-modal Nemotron Nano (image, video and audio inputs), 1M context, reasoning.

VisionReasoningTools / functions
1M · in Free · out Free

Mistral Medium 3.5

deprecated
Apr 2026

Mistral frontier-class multimodal model with adjustable reasoning.

VisionReasoningTools / functions
262K · in Free · out Free

DeepSeek V4 Flash

deprecated
Apr 2026

Fast DeepSeek V4 MoE (284B) with 1M context and reasoning.

ReasoningTools / functions
1M · in Free · out Free

DeepSeek V4 Pro

deprecated
Apr 2026

DeepSeek flagship reasoning MoE. NVIDIA serves a reduced 256K context (native: 1M). Note: often slow or saturated on the free endpoint.

ReasoningTools / functions
262K · in Free · out Free

Ising Calibration 1 35B

deprecated
Apr 2026

NVIDIA quantum-calibration VLM (domain-specific, not for general chat).

VisionReasoning
262K · in Free · out Free

Gemma 4 31B

Apr 2026

Google open multimodal model. CAUTION: NVIDIA serves 131K context and silently truncates longer inputs.

VisionReasoningTools / functions
131K · in Free · out Free

Nemotron 3 Super 120B

Mar 2026

Open hybrid Mamba-Transformer MoE (120B, 12B active), 1M context, reasoning and tool use.

ReasoningTools / functions
1M · in Free · out Free

Qwen 3.5 397B

deprecated
Feb 2026

Qwen 3.5 flagship MoE (multimodal, reasoning). CAUTION: served only to some NVIDIA accounts - most keys get a 404 "Function not found for account" error.

VisionReasoningTools / functions
262K · in Free · out Free

Nemotron 3 Nano 30B

deprecated
Dec 2025

Efficient open MoE (30B, 3B active) for high-volume tasks, 1M context, reasoning and tool use. Retires on NVIDIA 2026-08-31.

ReasoningTools / functions
1M · in Free · out Free

Riva Translate 4B v1.1

Dec 2025

Translation-specialized model, 8K context.

8K · in Free · out Free

Nemotron Safety Guard 8B v3

Oct 2025

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

Nemotron Nano 12B v2 VL

deprecated
Oct 2025

Vision-language Nemotron Nano for image and video understanding (verified image input). Retires on NVIDIA 2026-08-25.

VisionReasoningTools / functions
131K · in Free · out Free

Llama 3.3 Nemotron Super 49B v1.5

deprecated
Oct 2025

Llama 3.3 70B distilled and post-trained by NVIDIA for reasoning and agentic tasks. Retires on NVIDIA 2026-08-25.

ReasoningTools / functions
131K · in Free · out Free

Nemotron Nano 9B v2

deprecated
Aug 2025

Small hybrid Mamba-Transformer for fast, cheap reasoning and tool use. Retires on NVIDIA 2026-08-25.

ReasoningTools / functions
128K · in Free · out Free

GPT-OSS 120B

deprecated
Aug 2025

OpenAI open-weight MoE (117B, 5.1B active) with adjustable reasoning effort.

ReasoningTools / functions
131K · in Free · out Free

GPT-OSS 20B

Aug 2025

OpenAI open-weight MoE (20B, 3.6B active) with adjustable reasoning effort.

ReasoningTools / functions
131K · in Free · out Free

Mistral Nemotron

Jun 2025

Mistral model post-trained by NVIDIA.

ReasoningTools / functions
262K · in Free · out Free

Nemotron Nano VL 8B

deprecated
May 2025

Small vision-language model, 16K context.

Vision
16K · in Free · out Free

Llama Guard 4 12B

Apr 2025

Meta content-safety classifier (not a chat model). NVIDIA serves a reduced 64K context.

Vision
66K · in Free · out Free

Llama 3.3 Nemotron Super 49B v1

deprecated
Mar 2025

Superseded by v1.5.

ReasoningTools / functions
131K · in Free · out Free

NemoGuard 8B Content Safety

Jan 2025

NVIDIA content-safety classifier (not a chat model).

131K · in Free · out Free

NemoGuard 8B Topic Control

Jan 2025

NVIDIA topic-control guardrail (not a chat model).

131K · in Free · out Free

Llama 3.3 70B

deprecated
Dec 2024

Meta Llama 3.3 70B instruction-tuned, with tool calling. Retires on NVIDIA 2026-08-25 and is already unresponsive.

Tools / functions
131K · in Free · out Free

Llama 3.2 90B Vision

Sep 2024

Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).

Vision
33K · in Free · out Free

Llama 3.2 11B Vision

Sep 2024

Llama vision model for image understanding.

Vision
131K · in Free · out Free

Llama 3.2 1B

deprecated
Sep 2024

Tiny Llama for edge-class tasks.

Tools / functions
131K · in Free · out Free

Llama 3.2 3B

deprecated
Sep 2024

Small Llama for lightweight tasks.

Tools / functions
131K · in Free · out Free

Nemotron Mini 4B

deprecated
Sep 2024

Tiny legacy Nemotron, 4K context.

Tools / functions
4K · in Free · out Free

Llama 3.1 70B

deprecated
Jul 2024

Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).

Tools / functions
131K · in Free · out Free

Llama 3.1 8B

deprecated
Jul 2024

Small fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.

Tools / functions
131K · in Free · out Free
45 models · sorted by release date · prices in USD per 1M tokens · refreshed hourlyCompare every model across vendors ->

Get started in 3 steps

1

Create an API key at the NVIDIA console.

2

Paste it into Big-AGI's model settings.

3

Start chatting, or Beam it against other models and fuse the answers.

Setup, in three steps

  1. Create a free nvapi- key at build.nvidia.com, keeping the Public API Endpoints scope on it.
  2. In Big-AGI, add the NVIDIA NIM service and paste the key.
  3. Load the models and chat. No billing, no markup: the key, the account, and the limits are yours, direct from NVIDIA.

The build.nvidia.com playground is one prompt, one model, no history. The same key here gets persistent chats, personas, attachments, and NVIDIA models in the same conversation as Claude, GPT, or Gemini.

What Big-AGI fixes for you

  • Dead models filtered. Roughly half the ids in NVIDIA's catalog no longer answer. The curated list ships only live-verified models, so a dead id never reaches your picker.
  • Real context windows. Measured against the endpoint, not copied from catalog pages, which misstate about a quarter of them.
  • Rate limits paced. Calls queue instead of collecting 429s, including Beam scatters.
  • Reasoning wired through. GPT-OSS effort levels and Nemotron thinking toggles, in the chat controls.

Use it in Beam

Run Nemotron 3 Ultra, GPT-OSS 120B, and DeepSeek V4 on one prompt at no cost, then Fuse: combine, cross-check, and synthesize the parallel answers instead of just picking one.

Local NIM and DGX Spark

Point the same service at your own host, a NIM container or any OpenAI-compatible server (http://spark-xxxx.local:8000 style): identical model ids, same protocol. Direct Connection works on local hosts; the hosted endpoint blocks browser-origin calls, so those route through the server. Either way keys stay in your browser and chats are stored locally first.

Which free model for which job?

  • Deep reasoning: Nemotron 3 Ultra 550B, 1M context with tool calling.
  • Fast agent loops: Nemotron 3 Nano 30B (1M) or GPT-OSS 20B: small models keep a loop moving under a per-minute request cap.
  • Coding: GPT-OSS 120B, or DeepSeek V4 Flash when the whole repository has to fit in the prompt.
  • Vision: Nemotron Nano 12B v2 VL for images and video (131K), or MiniMax M3 (524K) for text-heavy pages.
  • Long documents: DeepSeek V4 Flash or Nemotron 3 Super 120B, both listed at 1M.

NVIDIA questions

FAQ

Why does an NVIDIA model return 404 "Not found for account"?

Either the model is retired (NVIDIA's catalog still lists it, but roughly half the ids no longer serve), or your key was created without the "Public API Endpoints" scope, which fails on every model. Fix the key at build.nvidia.com. Big-AGI only lists live-verified models, so a retired id never reaches your picker.

What is the rate limit on integrate.api.nvidia.com?

NVIDIA does not publish one. Expect roughly 40 requests per minute per account, moving with load. It counts requests, not tokens, so long prompts and answers are not penalized. Big-AGI paces its calls, so bursts and Beam scatters queue instead of collecting 429s.

Is the NVIDIA API really free? What is the catch?

Yes, no billing exists on the hosted endpoint. The catch is in the terms, not the invoice: no production use, NVIDIA may use your prompts and outputs to improve its models, and any model can be retired at any time. Keep sensitive and production work on a paid provider.

What context windows do the NVIDIA-hosted models actually have?

Often not the native window: the hosted deployment sets its own limit and the catalog pages misstate it. GLM 5.2 serves about 202K here against a 1M native window, and one model silently truncates instead of erroring. Big-AGI measures every limit against the live endpoint: the number in the table is the one you get.

Can I point it at a local NIM or a DGX Spark?

Yes. Set a custom host on the same service: a NIM container or any OpenAI-compatible server (localhost:8000 style) works with identical model ids. Local hosts also unlock Direct Connection, browser straight to your box, which the hosted endpoint blocks.

Bring your NVIDIA key. Keep control.

Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how NVIDIA is called.

Launch Big-AGI

© 2026 Token Fabrics·Built with passion in San Diego