The Big-AGI Model Index
If it's on this list, it runs in Big-AGI. Capabilities, context windows, and pricing for every model. Your key, provider rates, no markup.
What this list is: the models Big-AGI indexes, with full specs for each. Big-AGI also connects to any OpenAI-compatible endpoint, every model on OpenRouter, and local runtimes like Ollama, so the models you can run go beyond this list.
Tencent: Hy4 preview
NEWTencent: Hy4 preview is a mixture-of-experts model from Tencent…
1M
$0.83
$2.50
Aug 2026
Ling 3.0 Flash Fin (free)
NEWLing 3.0 Flash Fin is a finance-focused mixture-of-experts mode…
262K
Free
Free
Aug 2026
Gemini Omni 1.1 Flash (video)
NEWGemini Omni 1.1 Flash
197K
$1.50
$17.50
Aug 2026
Qwen: Qwen3.8 Flash
NEWQwen3.8 Flash is a multimodal reasoning model from Alibaba. It…
1M
$0.15
$0.47
Aug 2026
Gemini 3.5 Transcribe
NEWGemini 3.5 Transcribe
131K
-
-
Aug 2026
GLM-5.3 Flash (1M)
NEWHOTMultimodal Flash on a new 320B MoE base (18B activated, hybrid…
1M
$0.15
$0.50
Aug 2026
Meta: Muse Spark 1.2 Contributor
NEWMuse Spark 1.2 contributor tier is a reasoning model from Meta…
1M
$0.10
$0.20
Aug 2026
DeepSeek V4 Flash Vision (Exp)
NEWExperimental vision variant of V4 Flash with 1M context, releas…
1M
$0.44
$1.32
Aug 2026
Tencent: Hy-MT2-30B-A3B
NEWHy-MT2-30B-A3B is Tencent's flagship translation model in the H…
8K
$0.07
$0.30
Aug 2026
Tencent: Hy-MT2-1.8B
NEWHy-MT2-1.8B is a compact 1.8B-parameter translation model from…
8K
$0.04
$0.18
Aug 2026
Ox Alpha
NEWOx Alpha is a reasoning model designed for coding, sustained ag…
1M
Free
Free
Aug 2026
Tencent: Hy-MT2-7B
NEWHy-MT2-7B is a 7B-parameter translation model from Tencent. It…
8K
$0.07
$0.30
Aug 2026
Z.ai: GLM Latest
NEWThis model always redirects to the latest GLM model from Z.ai.
1.3M
$1.19
$4.18
Aug 2026
Qwen3.8 27B
NEWOpen-weights 27B dense vision-language model of the Qwen3.8 lin…
1M
$0.50
$3.00
Aug 2026
GLM-5.3 (1M)
NEWZ.ai 1M-context flagship, post-trained on the GLM-5.2 base for…
1M
$1.40
$4.40
Aug 2026
Gemini 3.7 Flash Video Understanding
NEWGemini 3.7 Flash Video Understanding EAP
1.1M
$0.75
$3.75
Aug 2026
Dots Studio: Dots3-Note Preview (free)
NEWDots3-Note Preview is an open-weight mixture-of-experts model f…
512K
Free
Free
Aug 2026
Google Gemini Flash Latest
NEWThis model always redirects to the latest model in the Google G…
1M
$0.75
$3.75
Aug 2026
Gemini 3.7 Flash
NEWHOTGemini 3.7 Flash
1.1M
$0.75
$3.75
Aug 2026
xAI: Grok Latest
NEWThis model always redirects to the latest Grok model from xAI.
500K
$2.00
$6.00
Aug 2026
Qwen3.8 2.4T-A95B
NEWOpen-weights release of the Qwen3.8 flagship: 2.4T sparse MoE,…
1M
$2.00
$6.00
Aug 2026
Grok 4.6
NEWxAI's frontier model for coding, agentic tasks, and knowledge w…
500K
$2.00
$6.00
Aug 2026
DeepSeek: DeepSeek V4 Pro 0813
NEWHOTDeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model…
1M
$1.32
$3.96
Aug 2026
ByteDance Seed: Seed-2.0-Code
NEWSeed 2.0 Code is a model from ByteDance Seed optimized for agen…
262K
$0.50
$3.00
Aug 2026
ByteDance Seed: Seed 2.1 Turbo
NEWSeed 2.1 Turbo is a multimodal model from ByteDance Seed for co…
262K
$0.50
$2.50
Aug 2026
Sakana: Namazu
NEWSakana Namazu is a Japanese-specialized reasoning model from Sa…
262K
$0.95
$4.00
Aug 2026
NVIDIA: Nemotron 3.5 Lightning
NEWHOTNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts mod…
262K
$0.08
$0.20
Aug 2026
Nemotron 3.5 Lightning 30B
NEWFastest Nemotron MoE (30B, 3B active), 1M context, reasoning an…
1M
Free
Free
Aug 2026
LiquidAI: LFM2.5-2.6B (free)
NEWLFM2.5-2.6B is a compact reasoning model from Liquid AI. It is…
66K
Free
Free
Aug 2026
Upstage: Solar Pro 4
NEWSolar Pro 4 is Upstage's cost-efficient large language model, f…
524K
$0.03
$0.12
Aug 2026
Meta: Muse Glimmer 30B
NEWMuse Glimmer 30B is a dense, open-weight multimodal model from…
131K
$0.30
$1.20
Aug 2026
inclusionAI: Ling 3.0 Tiny (free)
NEWLing 3.0 Tiny is a mixture-of-experts model from InclusionAI, w…
262K
Free
Free
Aug 2026
Meta: Muse Spark 1.2
NEWMuse Spark 1.2 is a reasoning model from Meta, designed for com…
1M
$1.25
$4.25
Aug 2026
Sakana Namazu v1.0
NEWJapanese-specialized LLM built on Moonshot AI's Kimi K2.6 and a…
262K
$0.95
$4.00
Aug 2026
Sakana Namazu
NEWJapanese-specialized LLM with built-in web search and code exec…
262K
$0.95
$4.00
Aug 2026
DeepSeek V4 Flash Latest
NEWThis model always redirects to the latest model in the DeepSeek…
1.3M
$0.03
$0.16
Aug 2026
DeepSeek: DeepSeek V4 Flash 0731
NEWHOTDeepSeek V4 Flash 0731 is a sparse mixture-of-experts model fro…
1.3M
$0.07
$0.18
Jul 2026
Thinking Machines: Inkling Small
NEWInkling Small is an open-weight multimodal mixture-of-experts m…
1M
$0.45
$1.20
Jul 2026
Gemini Robotics-ER 2 Preview
NEWGemini Robotics-ER 2 Preview
197K
$2.00
$10.00
Jul 2026
Riva Translate 4B v2
NEWTranslation-specialized model (37 languages, few-shot prompting…
8K
Free
Free
-
Jul 2026
Claude Opus 5 (Fast)
NEWFast-mode variant of Opus 5 - identical capabilities with highe…
1M
$10.00
$50.00
Jul 2026
Claude Opus 5
NEWHOTStep-change improvement over Opus 4.8 for complex agentic codin…
1M
$5.00
$25.00
Jul 2026
Anthropic: Claude Opus Latest
NEWThis model always redirects to the latest model in the Claude O…
1M
$5.00
$25.00
Jul 2026
Ling-3.0-flash
NEWHOT*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE)…
262K
$0.02
$0.06
Jul 2026
Sakana Fugu Ultra v1.1
NEWMulti-agent conductor system routing 1-3 expert agents for comp…
1M
$5.00
$30.00
Jul 2026
Sakana Fugu Cyber
NEWOrchestrator specialized for cybersecurity reasoning: security…
1M
$6.00
$36.00
Jul 2026
Poolside: Laguna S 2.1
NEWLaguna S 2.1 is the latest coding agent model from Poolside. La…
1M
$0.09
$0.18
Jul 2026
Gemini 3.6 Flash
NEWHOTGemini 3.6 Flash
1.1M
$0.75
$3.75
Jul 2026
Gemini 3.5 Flash-Lite
NEWGemini 3.5 Flash Lite
1.1M
$0.30
$2.50
Jul 2026
Meituan: LongCat 2.0
NEWLongCat 2.0 is a sparse mixture-of-experts language model from…
1M
$0.30
$1.20
Jul 2026
Ising Calibration 1.5 31B
NEWNVIDIA quantum-calibration VLM, preview (domain-specific, not f…
131K
Free
Free
Jul 2026
Qwen3.8 Max
NEWFlagship 2.4T-parameter sparse MoE multimodal model with 1M con…
1M
$2.00
$6.00
Jul 2026
Thinking Machines: Inkling
Inkling is an open-weight multimodal mixture-of-experts model f…
1M
$1.00
$4.05
Jul 2026
Auto Router (Beta)
Auto Router (Beta) is a task-aware router from OpenRouter. It c…
2M
-
-
Jul 2026
MoonshotAI Kimi Latest
This model always redirects to the latest model in the Moonshot…
1M
$2.55
$12.75
Jul 2026
Meta: Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built…
1M
$1.25
$4.25
Jul 2026
Kimi K3
HOTNative multimodal flagship (text, image, video inputs) with thi…
1M
$3.00
$15.00
Jul 2026
Qwen3.7 Flash
Latest fast multimodal model with 1M context, thinking (on by d…
1M
$0.03
$0.13
Jul 2026
Ternary Bonsai 27B
Prism Ml chat model. https://huggingface.co/api/models/prism-ml…
262K
-
-
Jul 2026
Kwaipilot: KAT-Coder-Pro V2.5
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model tha…
262K
$0.74
$2.96
Jul 2026
Kwaipilot: KAT-Coder-Air V2.5 (free)
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model tha…
256K
Free
Free
Jul 2026
OpenAI: GPT-5.6 Terra Pro
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra…
1.1M
$2.00
$12.00
Jul 2026
OpenAI: GPT-5.6 Sol Pro
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, se…
1.1M
$2.00
$10.00
Jul 2026
OpenAI: GPT-5.6 Luna Pro
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna,…
1.1M
$0.20
$1.20
Jul 2026
OpenAI GPT Latest
This model always redirects to the latest model in the OpenAI G…
1.1M
$2.00
$10.00
Jul 2026
GPT-5.6 Terra
Balanced model for efficient, high-volume everyday work. Compet…
1.1M
$2.00
$12.00
Jul 2026
GPT-5.6 Sol
HOTFlagship next-generation model. Strongest yet for agentic codin…
1.1M
$4.00
$20.00
Jul 2026
GPT-5.6 Luna
Fastest, most affordable GPT-5.6 model for high-volume work. St…
1.1M
$0.20
$1.20
Jul 2026
Grok 4.5
xAI's July 2026 flagship with frontier performance on coding, k…
500K
$2.00
$6.00
Jul 2026
AionLabs: Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling sys…
131K
$0.70
$1.40
Jul 2026
AionLabs: Aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system f…
131K
$3.00
$6.00
Jul 2026
Tencent: Hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (…
262K
$0.13
$0.53
Jul 2026
Poolside: Laguna XS 2.1
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B c…
262K
$0.06
$0.12
Jul 2026
Nano Banana 2 Lite
Gemini 3.1 Flash Lite Image.
131K
$0.25
$1.50
Jun 2026
Leanstral 1.5
A mid & post-trained version of mistral small 4 for Lean (26061…
262K
-
-
Jun 2026
labs-leanstral-1-5
A mid & post-trained version of mistral small 4 for Lean (26061…
262K
-
-
Jun 2026
Gemini Omni Flash Preview (video)
Gemini Omni Flash Preview
197K
$1.50
$17.50
Jun 2026
Claude Sonnet 5
HOTBest combination of speed and intelligence, with the largest ga…
1M
$2.00
$10.00
Jun 2026
Anthropic Claude Sonnet Latest
This model always redirects to the latest model in the Anthropi…
1M
$2.00
$10.00
Jun 2026
Qwen3.6 35B A3B Lora
Qwen chat model.
262K
-
-
Jun 2026
Qwen3.5 2B Lora
Qwen chat model.
262K
-
-
Jun 2026
Nex AGI: Nex-N2-Mini
HOTNex-N2-Mini is an open-source agentic mixture-of-experts model…
262K
$0.03
$0.10
Jun 2026
Sakana Fugu
Fast orchestration model routing tasks across a swappable pool…
1M
-
-
Jun 2026
Cohere: North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the…
256K
Free
Free
Jun 2026
Z.ai GLM 5.2
Official Z.ai GLM 5.2 model
1M
$1.40
$4.40
Jun 2026
GLM-5.2 (1M)
Z.ai 1M-context flagship (744B MoE, 40B activated). Agentic cod…
1M
$1.40
$4.40
Jun 2026
Sakana Fugu Ultra v1.0
Multi-agent conductor system routing 1-3 expert agents for comp…
1M
$5.00
$30.00
Jun 2026
Sakana Fugu Ultra
Multi-agent conductor system routing 1-3 expert agents for comp…
1M
$5.00
$30.00
Jun 2026
OpenRouter: Fusion
Fusion turns your prompt into a small multi-model deliberation.…
1M
-
-
Jun 2026
Kimi K2.7 Code Highspeed
High-speed code variant with ~180 tok/s output (up to 260 in sh…
262K
$1.90
$8.00
Jun 2026
Kimi K2.7 Code
Code-focused multimodal model (text, image, video inputs) with…
262K
$0.95
$4.00
Jun 2026
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hangin…
250K
Free
Free
Jun 2026
CO
North Mini Code
Compact agentic coding MoE (30B total, 3B active, Apache 2.0) f…
436K
-
-
Jun 2026
Claude Fable 5
HOTMost capable widely released model for the most demanding reaso…
1M
$10.00
$50.00
Jun 2026
Anthropic: Claude Fable Latest
HOTThis model always redirects to the latest model in the Claude F…
1M
$10.00
$50.00
Jun 2026
Nex AGI: Nex-N2-Pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI,…
262K
$0.25
$1.00
Jun 2026
NVIDIA: Nemotron 3.5 Content Safety (free)
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter mu…
128K
Free
Free
Jun 2026
NVIDIA: Nemotron 3 Ultra
HOTNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orche…
262K
$0.50
$2.20
Jun 2026
Llama 4 Maverick 17B 128E Instruct Nvfp4
Meta chat model. https://huggingface.co/api/models/RedHatAI/Lla…
1M
-
-
Jun 2026
GLM 4.7 FP4
Zai Org chat model.
203K
-
-
-
Jun 2026
Qwen3.7 Plus
Multimodal agent model with 1M context, native thinking, and vi…
1M
$0.40
$1.60
Jun 2026
MiniMax: MiniMax M3
HOTMiniMax-M3 is a multimodal foundation model from MiniMax. It su…
1M
$0.30
$1.20
May 2026
StepFun: Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal M…
262K
$0.20
$1.15
May 2026
Nano Banana Pro
Gemini 3 Pro Image
164K
$2.00
$12.00
May 2026
Nano Banana 2
HOTGemini 3.1 Flash Image.
131K
$0.50
$3.00
May 2026
Claude Opus 4.8
Previous most capable Opus-tier model for complex reasoning and…
1M
$5.00
$25.00
May 2026
Anthropic: Claude Opus 4.8 (Fast)
Fast-mode variant of Opus 4.8 - identical capabilities with hig…
1M
$10.00
$50.00
May 2026
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M conte…
1M
$2.50
$7.50
May 2026
Grok Build 0.1
HOTxAI fast coding model with reasoning, function calling, and str…
256K
$1.00
$2.00
May 2026
CO
Command A Plus
Cohere flagship MoE (218B total, 25B active, Apache 2.0). Agent…
436K
-
-
May 2026
Llama 4 Scout 17B 16E Instruct Fp8 Lora
Meta chat model.
10M
-
-
May 2026
Gemini 3.5 Flash
Gemini 3.5 Flash
1.1M
$1.50
$9.00
May 2026
Antigravity Agent Preview (2026-05)
Preview release of Antigravity Agent (05-2026)
197K
$0.75
$3.75
May 2026
Gemma 4 31B It Lora
Google chat model. https://huggingface.co/api/models/google/gem…
262K
-
-
May 2026
Gemma 3 27B It Lora
Google chat model.
-
-
-
May 2026
Perceptron: Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality visio…
33K
$0.15
$1.50
May 2026
inclusionAI: Ring-2.6-1T
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B act…
262K
$0.08
$0.63
May 2026
Mixtral 8x7B Instruct V0.1 FP8 Lora
Mistral AI chat model.
33K
-
-
-
May 2026
Gemma 3 270M It Lora
Google chat model.
33K
-
-
-
May 2026
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash Lite
1.1M
$0.25
$1.50
May 2026
Llama 3.3 70B Instruct FP8 Lora
Meta chat model.
131K
-
-
-
May 2026
OpenAI: GPT Chat Latest
GPT Chat Latest
400K
$5.00
$30.00
May 2026
ChatGPT Instant
400K
$5.00
$30.00
May 2026
IBM: Granite 4.1 8B
Granite 4.1 8B is a dense, decoder-only 8-billion-parameter lan…
131K
$0.05
$0.10
Apr 2026
Poolside: Laguna XS.2
Laguna XS.2 is the second-generation model in the XS size class…
262K
$0.10
$0.20
Apr 2026
Poolside: Laguna M.1
Laguna M.1 is the flagship coding agent model from Poolside, op…
262K
$0.20
$0.40
Apr 2026
Owl Alpha
Owl Alpha is a high-performance foundation model designed for a…
1M
Free
Free
Apr 2026
NVIDIA: Nemotron 3 Nano Omni (free)
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…
256K
Free
Free
Apr 2026
Nemotron 3 Nano Omni 30B A3b Reasoning Fp8
Nvidia chat model.
131K
-
-
Apr 2026
mistral-vibe-cli-with-tools
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
mistral-vibe-cli-latest
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
mistral-medium-latest
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
mistral-medium-3-5
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
mistral-medium
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
Mistral Medium 3.5
Mistral frontier-class multimodal model with adjustable reasoni…
262K
Free
Free
Apr 2026
Mistral Medium (2604)
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
magistral-medium-latest
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
Apr 2026
Qwen3.6 Max Preview
Qwen3.6 Max preview (never GA; superseded by Qwen3.7 Max). 256K…
262K
$1.30
$7.80
Apr 2026
Qwen3.6 Flash
Fast, cost-effective multimodal model with 1M context, near-fla…
1M
$0.25
$1.50
Apr 2026
Qwen3.6 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
262K
$0.38
$2.25
Apr 2026
Qwen3.6 27B
Open 27B dense multimodal with thinking. 256K context.
262K
$0.60
$3.60
Apr 2026
DeepSeek V4 Pro (0813)
HOTPremium reasoning model with 1M context, released GA by DeepSee…
1M
$1.32
$3.96
Apr 2026
DeepSeek V4 Flash (0731)
HOTFast general-purpose model with 1M context, re-post-trained by…
1M
$0.44
$1.32
Apr 2026
inclusionAI: Ling-2.6-1T
Ling-2.6-1T is an instant (instruct) model from inclusionAI and…
262K
$0.08
$0.63
Apr 2026
GPT-5.5 Pro
Most capable model for complex tasks. Uses more compute for sma…
1.1M
$30.00
$180.00
Apr 2026
GPT-5.5
New baseline for complex production workflows. Stronger task ex…
1.1M
$5.00
$30.00
Apr 2026
Xiaomi: MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong per…
1.1M
$0.44
$0.87
Apr 2026
Xiaomi: MiMo-V2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pr…
1.1M
$0.14
$0.28
Apr 2026
Tencent: Hy3 preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from…
262K
$0.18
$0.60
Apr 2026
Qwen3.6 35B A3b Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3.6…
262K
-
-
Apr 2026
Pareto Code Router
The Pareto Router maintains a tiered shortlist of strong coding…
2M
-
-
Apr 2026
OpenAI: GPT-5.4 Image 2
GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-t…
272K
$8.00
$15.00
Apr 2026
inclusionAI: Ling-2.6-flash
Ling-2.6-flash is an instant (instruct) model from inclusionAI…
262K
$0.01
$0.03
Apr 2026
Gemma 4 E2B-it
Google chat model. https://huggingface.co/api/models/google/gem…
131K
-
-
Apr 2026
Deep Research Preview (2026-04)
Preview release (April 21th, 2026) of Deep Research
197K
$1.25
$10.00
Apr 2026
Deep Research Max Preview (2026-04)
Preview release (April 21st, 2026) of Deep Research Max
197K
$1.25
$10.00
Apr 2026
Kimi K2.6
Native multimodal flagship (text, image, video inputs) with thi…
262K
$0.95
$4.00
Apr 2026
Grok 4.3
xAI's latest flagship model with reasoning and a 1M token conte…
1M
$1.25
$2.50
Apr 2026
Gemma 4 E4B-it
Google chat model. https://huggingface.co/google/gemma-4-E4B-it
131K
-
-
Apr 2026
Claude Opus 4.7
Previous most capable model for complex reasoning and agentic c…
1M
$5.00
$25.00
Apr 2026
Anthropic: Claude Opus 4.7 (Fast)
Fast-mode variant of Opus 4.7 - identical capabilities with hig…
1M
$30.00
$150.00
Apr 2026
Gemini 3.1 Flash TTS Preview
Gemini 3.1 Flash TTS Preview
25K
$1.00
-
Apr 2026
Ising Calibration 1 35B
NVIDIA quantum-calibration VLM (domain-specific, not for genera…
262K
Free
Free
Apr 2026
Gemini Robotics-ER 1.6 Preview
Gemini Robotics-ER 1.6 Preview
197K
$1.00
$5.00
Apr 2026
Nvidia Nemotron 3 Super 120B A12b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
262K
-
-
-
Apr 2026
GLM-5.1
Z.ai flagship (744B MoE, 40B activated). Post-training upgrade…
205K
$1.40
$4.40
Apr 2026
Qwen3.6 Plus
Previous Plus tier, superseded by Qwen3.7 Plus. 1M context, thi…
1M
$0.50
$3.00
Apr 2026
Gemma 4 31B IT
Gemma 4 31B IT
295K
Free
Free
Apr 2026
Gemma 4 26B A4B IT
HOTGemma 4 26B A4B IT
295K
Free
Free
Apr 2026
GLM-5V Turbo
First multimodal GLM-5 model. Vision-based coding agent with im…
205K
$1.20
$4.00
Apr 2026
Arcee AI: Trinity Large Thinking
Trinity Large Thinking is a powerful open source reasoning mode…
262K
$0.25
$0.80
Apr 2026
SpaceXAI: Grok 4.20
Grok 4.20 is a reasoning model from SpaceXAI with industry-lead…
2M
$1.25
$2.50
Mar 2026
Holo3 35B A3b
Hcompany chat model. https://huggingface.co/api/models/Hcompany…
262K
-
-
-
Mar 2026
Google: Lyria 3 Pro Preview
Full-length songs are priced at $0.08 per song. Lyria 3 is Goog…
1M
Free
Free
Mar 2026
Google: Lyria 3 Clip Preview
30 second duration clips are priced at $0.04 per clip. Lyria 3…
1M
Free
Free
Mar 2026
Kwaipilot: KAT-Coder-Pro V2
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKA…
262K
$0.30
$1.20
Mar 2026
Qwen3 30B A3B Instruct 2507 Lora
Qwen chat model.
262K
-
-
-
Mar 2026
DeepSeek V3.1
Deepseek model via OpenAI-Compatible API on AWS Bedrock Mantle
131K
$0.60
$1.70
Mar 2026
Reka Edge
Reka Edge is an extremely efficient 7B multimodal vision-langua…
16K
$0.10
$0.10
Mar 2026
Qwen3 8B Lora
Qwen chat model.
41K
-
-
-
Mar 2026
MiniMax: MiniMax M2.7 (free)
MiniMax-M2.7 is a next-generation large language model designed…
197K
Free
Free
Mar 2026
OpenAI GPT Mini Latest
This model always redirects to the latest model in the OpenAI G…
400K
$0.75
$4.50
Mar 2026
GPT-5.4 Nano
Cheapest GPT-5.4-class model for simple high-volume tasks like…
400K
$0.20
$1.25
Mar 2026
GPT-5.4 Mini
Strongest mini model for coding, computer use, and subagents. G…
400K
$0.75
$4.50
Mar 2026
Qwen3.5 122B A10b Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3.5…
262K
-
-
Mar 2026
mistral-vibe-cli-fast
Mistral Small 4.
262K
$0.15
$0.60
Mar 2026
mistral-small-latest
Mistral Small 4.
262K
$0.15
$0.60
Mar 2026
Mistral Small (2603)
Mistral Small 4.
262K
$0.15
$0.60
Mar 2026
magistral-small-latest
Mistral Small 4.
262K
$0.15
$0.60
Mar 2026
Leanstral (2603)
A mid & post-trained version of mistral small 4 for Lean
197K
-
-
Mar 2026
GLM-5 Turbo
Speed-optimized GLM-5 variant for agent workflows. Enhanced too…
205K
$1.20
$4.00
Mar 2026
Deepseek OCR 2
Deepseek chat model. https://huggingface.co/api/models/deepseek…
8K
-
-
Mar 2026
NVIDIA: Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE mod…
1M
$0.09
$0.40
Mar 2026
Nvidia Nemotron 3 Super 120B A12b Fp8
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
262K
-
-
-
Mar 2026
Qwen: Qwen3.5-9B
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 fa…
262K
$0.10
$0.15
Mar 2026
ByteDance Seed: Seed-2.0-Lite
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhor…
262K
$0.25
$2.00
Mar 2026
SpaceXAI: Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 desi…
2M
$1.25
$2.50
Mar 2026
Grok 4.20 Reasoning
xAI flagship reasoning model with a 1M token context window. De…
1M
$1.25
$2.50
Mar 2026
Grok 4.20 Multi-Agent
Multi-agent model that runs specialized agents in parallel for…
1M
$1.25
$2.50
Mar 2026
Grok 4.20
xAI flagship model with a 1M token context window. Non-reasonin…
1M
$1.25
$2.50
Mar 2026
Qwen3.5 9B Fp8
Qwen chat model. https://huggingface.co/api/models/togethercomp…
262K
-
-
Mar 2026
GPT-5.4 Pro
Most capable model for complex tasks. Uses more compute for sma…
1.1M
$30.00
$180.00
Mar 2026
GPT-5.4
Most capable and efficient frontier model for professional work…
1.1M
$2.50
$15.00
Mar 2026
Inception: Mercury 2
Mercury 2 is an extremely fast reasoning LLM, and the first rea…
128K
$0.25
$0.75
Mar 2026
OpenAI: GPT-5.3 Chat
GPT-5.3 Chat is an update to ChatGPT's most-used model that mak…
128K
$1.75
$14.00
Mar 2026
GPT-5.3 Instant
deprecatedGPT-5.3 Instant model, previously powering ChatGPT.
128K
$1.75
$14.00
Mar 2026
Glm 4.7 Fp8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
203K
-
-
-
Mar 2026
Gemini 3.1 Flash-Lite Preview
Gemini 3.1 Flash Lite Preview
1.1M
$0.25
$1.50
Mar 2026
Nano Banana 2 Preview
Gemini 3.1 Flash Image Preview.
131K
$0.50
$3.00
Feb 2026
ByteDance Seed: Seed-2.0-Mini
Seed-2.0-mini targets latency-sensitive, high-concurrency, and…
262K
$0.10
$0.40
Feb 2026
Qwen3.5 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
262K
$0.25
$2.00
Feb 2026
Qwen3.5 27B
Open 27B dense multimodal with thinking. 256K context.
262K
$0.30
$2.40
Feb 2026
Qwen3.5 122B-A10B
Open 122B-A10B MoE multimodal with thinking. 256K context.
262K
$0.40
$3.20
Feb 2026
Qwen: Qwen3.5-Flash
The Qwen3.5 native vision-language Flash models are built on a…
1M
$0.07
$0.26
Feb 2026
LiquidAI: LFM2-24B-A2B
LFM2-24B-A2B is the largest model in the LFM2 family of hybrid…
128K
$0.03
$0.12
Feb 2026
GPT Audio 1.5
Best voice model for audio in, audio out with Chat Completions.…
128K
$2.50
$10.00
Feb 2026
Qwen3.5 Flash
Former Flash tier, superseded by Qwen3.6/3.7 Flash. 1M context,…
1M
$0.10
$0.40
Feb 2026
AionLabs: Aion-2.0
Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive…
131K
$0.80
$1.60
Feb 2026
Google Gemini Pro Latest
HOTThis model always redirects to the latest model in the Google G…
1M
$2.00
$12.00
Feb 2026
Gemini 3.1 Pro Preview (Custom Tools)
Gemini 3.1 Pro Preview optimized for custom tool usage
1.1M
$2.00
$12.00
Feb 2026
Gemini 3.1 Pro Preview
HOTGemini 3.1 Pro Preview
1.1M
$2.00
$12.00
Feb 2026
CO
Tiny Aya Water
Tiny Aya (3.35B) region-specialized for European and Asia-Pacif…
8K
-
-
-
Feb 2026
CO
Tiny Aya Global
Tiny multilingual research model (3.35B), best balance across 7…
8K
-
-
-
Feb 2026
CO
Tiny Aya Fire
Tiny Aya (3.35B) region-specialized for South Asian languages.…
8K
-
-
-
Feb 2026
CO
Tiny Aya Earth
Tiny Aya (3.35B) region-specialized for West Asian and African…
8K
-
-
-
Feb 2026
Claude Sonnet 4.6
Best combination of speed and intelligence for everyday tasks
1M
$3.00
$15.00
Feb 2026
Qwen3.5 397B-A17B
Open 397B-A17B MoE multimodal with thinking. 256K context.
262K
$0.60
$3.60
Feb 2026
Qwen: Qwen3.5 Plus 2026-02-15
The Qwen3.5 native vision-language series Plus models are built…
1M
$0.26
$1.56
Feb 2026
Qwen3.5 Plus
Former Plus tier, superseded by Qwen3.6/3.7 Plus. 1M context, t…
1M
$0.40
$2.40
Feb 2026
MiniMax M2.5 FP4
MiniMaxAI chat model.
8K
-
-
-
Feb 2026
MiniMax: MiniMax M2.5
MiniMax-M2.5 is a SOTA large language model designed for real-w…
205K
$0.27
$1.08
Feb 2026
GLM 5 Fp4
Zai Org chat model. https://huggingface.co/api/models/togetherc…
203K
-
-
-
Feb 2026
GLM-5
Z.ai flagship foundation model (744B MoE, 40B activated). Desig…
205K
$1.00
$3.20
Feb 2026
Qwen: Qwen3 Max Thinking
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3…
262K
$0.78
$3.90
Feb 2026
GPT-5.3 Codex
Most capable agentic coding model. Combines frontier coding per…
400K
$1.75
$14.00
Feb 2026
Claude Opus 4.6
HOTPrevious most intelligent model for complex agents and coding,…
1M
$5.00
$25.00
Feb 2026
Qwen3 Coder Next
Budget agentic coder on the Qwen3-Next architecture. 256K conte…
262K
$0.30
$1.50
Feb 2026
GLM-OCR (Vision, OCR)
Specialized OCR model for text extraction from images and docum…
131K
$0.03
$0.03
Feb 2026
Free Models Router
The simplest way to get free inference. openrouter/free is a ro…
200K
Free
Free
Feb 2026
StepFun: Step 3.5 Flash
Step 3.5 Flash is StepFun's most capable open-source foundation…
262K
$0.10
$0.30
Jan 2026
Upstage: Solar Pro 3
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) lang…
131K
$0.15
$0.60
Jan 2026
Kimi K2.5
Supports vision (images/videos), thinking mode, and Agent tasks…
262K
$0.60
$3.00
Jan 2026
MiniMax: MiniMax M2-her
MiniMax M2-her is a dialogue-first large language model built f…
66K
$0.30
$1.20
Jan 2026
Writer: Palmyra X5
Palmyra X5 is Writer's most advanced model, purpose-built for b…
1M
$0.60
$6.00
Jan 2026
LiquidAI: LFM2.5-1.2B-Thinking (free)
LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model o…
33K
Free
Free
Jan 2026
LiquidAI: LFM2.5-1.2B-Instruct (free)
LFM2.5-1.2B-Instruct is a compact, high-performance instruction…
33K
Free
Free
-
Jan 2026
GLM-4.7 FlashX
Fast GLM-4.7 variant with priority routing and higher concurren…
205K
$0.07
$0.40
Jan 2026
GLM-4.7 Flash (Free)
Free GLM-4.7 variant. Same model as FlashX but with limited con…
205K
Free
Free
Jan 2026
Z.ai GLM 4.7 (Preview)
Z.ai GLM 4.7 (355B) on Cerebras (~1,000 tok/s). Strong agentic…
131K
$2.25
$2.75
Jan 2026
MiniMax: MiniMax M2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language…
205K
$0.30
$1.20
Dec 2025
ByteDance Seed: Seed 1.6 Flash
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model…
262K
$0.08
$0.30
Dec 2025
ByteDance Seed: Seed 1.6
Seed 1.6 is a general-purpose model released by the ByteDance S…
262K
$0.25
$2.00
Dec 2025
GLM-4.7
Latest-gen GLM model with 200K context. Thinking mode activated…
205K
$0.60
$2.20
Dec 2025
Gemini 3 Flash Preview
Gemini 3 Flash Preview
1.1M
$0.50
$3.00
Dec 2025
Nvidia Nemotron 3 Nano 30B A3b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
262K
-
-
-
Dec 2025
NVIDIA: Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model wi…
262K
$0.05
$0.20
Dec 2025
Riva Translate 4B v1.1
Translation-specialized model, 8K context.
8K
Free
Free
-
Dec 2025
OpenAI: GPT-5.2 Chat
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of t…
128K
$1.75
$14.00
Dec 2025
GPT-5.2 Instant
deprecatedGPT-5.2 Instant model, previously powering ChatGPT.
128K
$1.75
$14.00
Dec 2025
GPT-5.2 Codex
deprecatedGPT-5.2 optimized for long-horizon, agentic coding tasks in Cod…
400K
$1.75
$14.00
Dec 2025
Deep Research Pro Preview
Preview release (December 12th, 2025) of Deep Research Pro
197K
$1.25
$10.00
Dec 2025
AutoGLM Phone
Mobile phone automation agent. Understands phone screens via mu…
131K
Free
Free
Dec 2025
GPT-5.2 Pro
Smartest and most trustworthy option for difficult questions. U…
400K
$21.00
$168.00
Dec 2025
GPT-5.2
Most capable model for professional work and long-running agent…
400K
$1.75
$14.00
Dec 2025
Mistral: Devstral 2 2512
Devstral 2 is a state-of-the-art open-source model by Mistral A…
262K
$0.40
$2.00
Dec 2025
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
262K
$0.40
$2.00
Dec 2025
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
262K
$0.40
$2.00
Dec 2025
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
262K
$0.40
$2.00
Dec 2025
Relace: Relace Search
The relace-search model uses 4-12 `view_file` and `grep` tools…
256K
$1.00
$3.00
Dec 2025
GLM-4.6 V FlashX
Fast vision GLM-4.6 with priority routing and higher concurrenc…
131K
$0.04
$0.40
Dec 2025
GLM-4.6 V Flash (Free)
Free vision GLM-4.6. Same model as FlashX but with limited conc…
131K
Free
Free
Dec 2025
GLM-4.6 V
Vision-enabled GLM-4.6 model. Supports image/video/file inputs,…
131K
$0.30
$0.90
Dec 2025
EssentialAI Rnj-1 Instruct
Essential AI chat model. https://huggingface.co/api/models/toge…
33K
-
-
-
Dec 2025
Body Builder (beta)
Transform your natural language requests into structured OpenRo…
128K
-
-
Dec 2025
mistral-large-latest
Official mistral-large-2512 Mistral AI model
262K
$0.50
$1.50
Dec 2025
ministral-8b-latest
Ministral 3 (a.k.a. Tinystral) 8B Instruct.
262K
$0.15
$0.15
Dec 2025
ministral-3b-latest
Ministral 3 (a.k.a. Tinystral) 3B Instruct.
131K
$0.10
$0.10
Dec 2025
ministral-14b-latest
Ministral 3 (a.k.a. Tinystral) 14B Instruct.
262K
$0.20
$0.20
Dec 2025
Ministral 8b (2512)
Ministral 3 (a.k.a. Tinystral) 8B Instruct.
262K
$0.15
$0.15
Dec 2025
Ministral 3b (2512)
Ministral 3 (a.k.a. Tinystral) 3B Instruct.
131K
$0.10
$0.10
Dec 2025
Ministral 3 14B Instruct 2512
Mistralai chat model. https://huggingface.co/api/models/mistral…
262K
$0.20
$0.20
-
Dec 2025
Ministral 14b (2512)
Ministral 3 (a.k.a. Tinystral) 14B Instruct.
262K
$0.20
$0.20
Dec 2025
Amazon: Nova 2 Lite
Nova 2 Lite is a fast, cost-effective reasoning model for every…
1M
$0.30
$2.50
Dec 2025
Mistral Large (2512)
Official mistral-large-2512 Mistral AI model
262K
$0.50
$1.50
Dec 2025
DeepSeek: DeepSeek V3.2
DeepSeek-V3.2 is a large language model designed to harmonize h…
164K
$0.27
$0.40
Dec 2025
Arcee AI: Trinity Mini
Trinity Mini is a 26B-parameter (3B active) sparse mixture-of-e…
131K
$0.05
$0.15
Dec 2025
Claude Opus 4.5
Previous most intelligent model with advanced reasoning for com…
200K
$5.00
$25.00
Nov 2025
AllenAI: Olmo 3 32B Think
Olmo 3 32B Think is a large-scale, 32-billion-parameter model p…
66K
$0.15
$0.50
Nov 2025
Nano Banana Pro Preview
Gemini 3 Pro Image Preview
164K
$2.00
$12.00
Nov 2025
Nano Banana Pro
Gemini 3 Pro Image Preview
164K
$2.00
$12.00
Nov 2025
GPT-5.1 Codex Max
deprecatedOur most intelligent coding model optimized for long-horizon, a…
400K
$1.25
$10.00
Nov 2025
GPT-5.1 Codex Mini
deprecatedSmaller, faster version of GPT-5.1 Codex for efficient coding t…
400K
$0.25
$2.00
Nov 2025
GPT-5.1 Codex
deprecatedA version of GPT-5.1 optimized for agentic coding tasks in Code…
400K
$1.25
$10.00
Nov 2025
GPT-5.1
The best model for coding and agentic tasks with configurable r…
400K
$1.25
$10.00
Nov 2025
Deep Cogito: Cogito v2.1 671B
Cogito v2.1 671B MoE represents one of the strongest open model…
128K
$1.25
$1.25
Nov 2025
OpenAI: GPT-5.1 Chat
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of th…
128K
$1.25
$10.00
Nov 2025
GPT-5.1 Instant
deprecatedGPT-5.1 Instant with adaptive reasoning. More conversational wi…
128K
$1.25
$10.00
Nov 2025
Qwen3-VL-235B-A22B-Instruct-FP8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-V…
262K
-
-
Nov 2025
MoonshotAI: Kimi K2 Thinking
Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning…
262K
$0.60
$2.50
Nov 2025
Amazon: Nova Premier 1.0
Amazon Nova Premier is the most capable of Amazon’s multimodal…
1M
$2.50
$12.50
Oct 2025
Perplexity: Sonar Pro Search
Exclusively available on the OpenRouter API, Sonar Pro's new Pr…
200K
$3.00
$15.00
Oct 2025
Mistral: Voxtral Small 24B 2507
Voxtral Small is an enhancement of Mistral Small 3, incorporati…
33K
$0.10
$0.30
Oct 2025
OpenAI: gpt-oss-safeguard-20b
gpt-oss-safeguard-20b is a safety reasoning model from OpenAI b…
131K
$0.08
$0.30
Oct 2025
NVIDIA: Nemotron Nano 12B 2 VL (free)
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multim…
128K
Free
Free
Oct 2025
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
-
Oct 2025
Medgemma 27B Text It
Google chat model. https://huggingface.co/api/models/google/med…
131K
-
-
-
Oct 2025
Qwen: Qwen3 VL 32B Instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-langua…
131K
$0.10
$0.42
Oct 2025
MiniMax: MiniMax M2
MiniMax-M2 is a compact, high-efficiency large language model o…
205K
$0.26
$1.02
Oct 2025
IBM: Granite 4.0 Micro
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family…
131K
$0.02
$0.11
Oct 2025
Microsoft: Phi 4 Mini Instruct
Phi-4-mini-instruct is a lightweight open model built upon synt…
131K
$0.08
$0.35
Oct 2025
Qwen3 VL Flash
Budget VL tier below Qwen3 VL Plus. 256K context, thinking, vis…
262K
$0.05
$0.40
Oct 2025
Claude Haiku 4.5
Fastest model with exceptional speed and performance
200K
$1.00
$5.00
Oct 2025
Anthropic Claude Haiku Latest
This model always redirects to the latest model in the Anthropi…
200K
$1.00
$5.00
Oct 2025
Qwen: Qwen3 VL 8B Thinking
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the…
131K
$0.18
$2.10
Oct 2025
Qwen: Qwen3 VL 8B Instruct
Qwen3-VL-8B-Instruct is a multimodal vision-language model from…
262K
$0.12
$0.46
Oct 2025
GPT-5 Search API
Updated web search model in Chat Completions API. 60% cheaper w…
400K
$1.25
$10.00
Oct 2025
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-c…
131K
$0.40
$0.40
Oct 2025
Gemini 2.5 Computer Use Preview 10-2025
Gemini 2.5 Computer Use Preview 10-2025
197K
$1.25
$10.00
Oct 2025
Qwen: Qwen3 VL 30B A3B Thinking
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies st…
262K
$0.20
$2.40
Oct 2025
Qwen: Qwen3 VL 30B A3B Instruct
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies st…
262K
$0.15
$0.60
Oct 2025
GPT-5 Pro
Version of GPT-5 that uses more compute to produce smarter and…
400K
$15.00
$120.00
Oct 2025
GPT Audio Mini
deprecatedCost-efficient audio model. Accepts audio inputs and outputs vi…
128K
$0.60
$2.40
Oct 2025
Nano Banana
Gemini 2.5 Flash Preview Image
66K
$0.30
$2.50
Oct 2025
GLM-4.6
GLM-4.6 model with 200K context, 128K output. Hybrid thinking:…
205K
$0.60
$2.20
Sep 2025
Gemma 3 270M It
Google chat model. https://huggingface.co/api/models/google/gem…
33K
-
-
-
Sep 2025
DeepSeek: DeepSeek V3.2 Exp
DeepSeek-V3.2-Exp is an experimental large language model relea…
164K
$0.27
$0.41
Sep 2025
Claude Sonnet 4.5
HOTPrevious best combination of speed and intelligence for complex…
200K
$3.00
$15.00
Sep 2025
TheDrummer: Cydonia 24B V4.1
Uncensored and creative writing model based on Mistral Small 3.…
131K
$0.30
$0.50
Sep 2025
Relace: Relace Apply 3
Relace Apply 3 is a specialized code-patching LLM that merges A…
256K
$0.85
$1.25
Sep 2025
Google: Gemini 2.5 Flash Lite Preview 09-2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the G…
1M
$0.10
$0.40
Sep 2025
Qwen3 Next 80B A3b Instruct Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-N…
-
-
-
-
Sep 2025
Qwen3 Vl 235b A22b Thinking
Alibaba model (not yet curated).
131K
$0.40
$4.00
Sep 2025
Qwen3 Vl 235b A22b Instruct
Alibaba model (not yet curated).
131K
$0.21
$1.90
Sep 2025
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context…
262K
$1.20
$6.00
Sep 2025
Qwen3 Coder Plus
Agentic coding model with very long context. Tiered pricing by…
1M
$1.00
$5.00
Sep 2025
DeepSeek: DeepSeek V3.1 Terminus
DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maint…
164K
$0.27
$1.00
Sep 2025
Magistral Small (2509)
Our efficient reasoning model released September 2025.
131K
$0.50
$1.50
Sep 2025
Magistral Medium (2509)
Our frontier-class reasoning model release candidate September…
131K
$2.00
$5.00
Sep 2025
GPT-5 Codex
deprecatedA version of GPT-5 optimized for agentic coding in Codex.
400K
$1.25
$10.00
Sep 2025
Qwen3 Next 80b A3b Thinking
Alibaba model (not yet curated).
131K
$0.15
$1.20
Sep 2025
Qwen3 Next 80b A3b Instruct
Alibaba model (not yet curated).
131K
$0.09
$1.10
Sep 2025
NVIDIA: Nemotron Nano 9B V2 (free)
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trai…
128K
Free
Free
Sep 2025
MoonshotAI: Kimi K2 0905
Kimi K2 0905 is the September update of Kimi K2 0711. It is a l…
262K
$0.60
$2.50
Sep 2025
[Groq] Compound Mini (Agentic System)
Lighter Groq agentic AI, same built-in tools but a single tool…
131K
-
-
Sep 2025
[Groq] Compound (Agentic System)
Groq agentic AI with web search, visit website, code execution,…
131K
-
-
Sep 2025
Qwen3 30b A3b Thinking 2507
Alibaba model (not yet curated).
131K
$0.20
$2.40
Aug 2025
GPT Audio
First generally available audio model. Accepts audio inputs and…
128K
$2.50
$10.00
Aug 2025
Nous: Hermes 4 70B
Hermes 4 70B is a hybrid reasoning model from Nous Research, bu…
131K
$0.13
$0.40
Aug 2025
Nous: Hermes 4 405B
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3…
131K
$1.00
$3.00
Aug 2025
DeepSeek: DeepSeek V3.1
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameter…
164K
$0.55
$1.65
Aug 2025
Nemotron Nano 9B v2
Small hybrid Mamba-Transformer for fast, cheap reasoning and to…
128K
Free
Free
Aug 2025
Mistral: Mistral Medium 3.1
Mistral Medium 3.1 is an updated version of Mistral Medium 3, w…
131K
$0.40
$2.00
Aug 2025
Mistral Medium (2508)
Update on Mistral Medium 3 with improved capabilities.
131K
$0.40
$2.00
Aug 2025
GLM-4.5 V
Vision-enabled GLM-4.5 model. 64K context, 16K output, interlea…
66K
$0.60
$1.80
Aug 2025
Qwen3 4B Instruct 2507
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-4…
262K
-
-
-
Aug 2025
AI21: Jamba Large 1.7
Jamba Large 1.7 is the latest model in the Jamba open family, o…
256K
$2.00
$8.00
Aug 2025
OpenAI: GPT-5 Chat
GPT-5 Chat is designed for advanced, natural, multimodal, and c…
128K
$1.25
$10.00
Aug 2025
GPT-5 Nano
Fastest, most cost-efficient version of GPT-5 for summarization…
400K
$0.05
$0.40
Aug 2025
GPT-5 Mini
A faster, more cost-efficient version of GPT-5 for well-defined…
400K
$0.25
$2.00
Aug 2025
GPT-5 ChatGPT
deprecatedGPT-5 model used in ChatGPT.
128K
$1.25
$10.00
Aug 2025
GPT-5
The best model for coding and agentic tasks across domains.
400K
$1.25
$10.00
Aug 2025
OpenAI: gpt-oss-20b (free)
gpt-oss-20b is an open-weight 21B parameter model released by O…
131K
Free
Free
Aug 2025
OpenAI: gpt-oss-120b
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Exper…
131K
$0.04
$0.17
Aug 2025
Claude Opus 4.1
deprecatedPrevious Opus model. Deprecated June 5, 2026, retiring August 5…
200K
$15.00
$75.00
Aug 2025
CO
Command A Translate
Specialized machine translation across 23 languages, with tool…
9K
-
-
Aug 2025
CO
Command A Reasoning
Reasoning-tuned Command A for multi-step agents and hard proble…
289K
$2.50
$10.00
Aug 2025
Qwen: Qwen3 Coder 30B A3B Instruct
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Ex…
262K
$0.07
$0.28
Jul 2025
mistral-code-latest
Our cutting-edge language model for coding released August 2025.
256K
$0.30
$0.90
Jul 2025
mistral-code-fim-latest
Our cutting-edge language model for coding released August 2025.
256K
$0.30
$0.90
Jul 2025
Glm 4.5 Air Fp8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
131K
$0.20
$1.10
-
Jul 2025
codestral-latest
Our cutting-edge language model for coding released August 2025.
256K
$0.30
$0.90
Jul 2025
Codestral (2508)
Our cutting-edge language model for coding released August 2025.
256K
$0.30
$0.90
Jul 2025
Qwen3 30b A3b Instruct 2507
Alibaba model (not yet curated).
131K
$0.05
$0.19
Jul 2025
Qwen3 235B A22b Instruct 2507 Fp8
Together AI chat model. https://huggingface.co/api/models/Qwen/…
262K
-
-
-
Jul 2025
Qwen3 Coder Flash
Former budget coder, superseded by Qwen3 Coder Next (which lack…
1M
$0.30
$1.50
Jul 2025
GLM-4.5 X
Extended GLM-4.5 model. Interleaved thinking.
131K
$2.20
$8.90
Jul 2025
GLM-4.5 Flash (Free)
Free GLM-4.5 variant with limited concurrency. Prior-gen, super…
131K
Free
Free
Jul 2025
GLM-4.5 AirX
Extended lightweight GLM-4.5 variant. Interleaved thinking.
131K
$1.10
$4.50
Jul 2025
GLM-4.5 Air
Lightweight GLM-4.5 variant. Interleaved thinking.
131K
$0.20
$1.10
Jul 2025
GLM-4.5
Prior-gen GLM-4.5 model with 128K context, 96K output. Interlea…
131K
$0.60
$2.20
Jul 2025
Qwen3 235b A22b Thinking 2507
Alibaba model (not yet curated).
131K
$0.23
$2.30
Jul 2025
Qwen3 Coder 480B A35B Instruct Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-C…
262K
$2.00
$2.00
-
Jul 2025
Qwen: Qwen3 Coder 480B A35B
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) co…
262K
$0.30
$1.00
Jul 2025
Qwen3 235B A22B Instruct 2507 FP8 Throughput
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-2…
262K
$0.20
$0.60
-
Jul 2025
Gemini 2.5 Flash-Lite
Stable version of Gemini 2.5 Flash-Lite, released in July of 20…
1.1M
$0.10
$0.40
Jul 2025
ByteDance: UI-TARS 7B
UI-TARS-1.5 is a multimodal vision-language agent optimized for…
128K
$0.10
$0.20
Jul 2025
Qwen: Qwen3 235B A22B Instruct 2507
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tu…
262K
$0.09
$0.35
Jul 2025
voxtral-small-latest
A small audio understanding model released in July 2025
33K
$0.10
$0.40
Jul 2025
voxtral-mini-latest
A mini audio understanding model released in July 2025
33K
$0.04
$0.04
-
Jul 2025
Voxtral Small (2507)
A small audio understanding model released in July 2025
33K
$0.10
$0.40
Jul 2025
Voxtral Mini (2507)
A mini audio understanding model released in July 2025
33K
$0.04
$0.04
-
Jul 2025
Switchpoint Router
Switchpoint AI's router instantly analyzes your request and dir…
131K
$0.85
$3.40
Jul 2025
MoonshotAI: Kimi K2 0711
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) lang…
131K
$0.57
$2.30
Jul 2025
Sarvam M
Sarvamai chat model. https://huggingface.co/api/models/sarvamai…
33K
-
-
-
Jul 2025
Meta Llama 3.1 8B Instruct Awq Int4
Meta chat model. https://huggingface.co/api/models/togethercomp…
131K
-
-
-
Jul 2025
Venice: Uncensored (free)
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-…
33K
Free
Free
Jul 2025
Tencent: Hunyuan A13B Instruct
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE)…
131K
$0.14
$0.57
Jul 2025
Morph: Morph V3 Large
Morph's high-accuracy apply model for complex code edits. ~4,50…
262K
$0.90
$1.90
Jul 2025
Morph: Morph V3 Fast
Morph's fastest apply model for code edits. ~10,500 tokens/sec…
82K
$0.80
$1.20
Jul 2025
CO
Command A Vision
Multimodal Command A for charts, graphs, diagrams, OCR, and doc…
128K
-
-
Jul 2025
Baidu: ERNIE 4.5 VL 424B A47B
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE)…
123K
$0.42
$1.25
Jun 2025
Minimax M1 80K
MiniMaxAI chat model. https://huggingface.co/api/models/togethe…
1M
-
-
-
Jun 2025
o4 Mini Deep Research
Faster, more affordable deep research model for complex, multi-…
200K
$2.00
$8.00
Jun 2025
o3 Deep Research
Our most powerful deep research model for complex, multi-step r…
200K
$10.00
$40.00
Jun 2025
Minimax M1 40K
MiniMaxAI chat model. https://huggingface.co/api/models/togethe…
1M
-
-
-
Jun 2025
Mistral: Mistral Small 3.2 24B
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter…
131K
$0.08
$0.20
Jun 2025
Mistral Small (2506)
Our latest enterprise-grade small model with the latest version…
131K
$0.10
$0.30
Jun 2025
MiniMax: MiniMax M1
MiniMax-M1 is a large-scale, open-weight reasoning model design…
1M
$0.55
$2.20
Jun 2025
Gemini 2.5 Pro
Stable release (June 17th, 2025) of Gemini 2.5 Pro
1.1M
$1.25
$10.00
Jun 2025
Gemini 2.5 Flash
Stable version of Gemini 2.5 Flash, our mid-size multimodal mod…
1.1M
$0.30
$2.50
Jun 2025
Mistral Nemotron
Mistral model post-trained by NVIDIA.
262K
Free
Free
Jun 2025
Magistral Small 2506
Mistralai chat model. https://huggingface.co/api/models/mistral…
41K
-
-
-
Jun 2025
Llama 4 Scout (17Bx16E)
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
262K
-
-
Jun 2025
o3 Pro
Version of o3 with more compute for better responses. Provides…
200K
$20.00
$80.00
Jun 2025
Gemma 2B It
Google chat model. https://huggingface.co/api/models/google/gem…
8K
-
-
-
Jun 2025
Gemma 2 9B It
Google chat model. https://huggingface.co/api/models/google/gem…
8K
-
-
-
Jun 2025
Qwen3 1.7B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-1…
41K
-
-
-
Jun 2025
Qwen3 0.6B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-0…
41K
-
-
-
Jun 2025
Nemotron Nano VL 8B
Small vision-language model, 16K context.
16K
Free
Free
May 2025
Molmo 7B D 0924
Allenai chat model. https://huggingface.co/api/models/allenai/M…
4K
-
-
May 2025
DeepSeek: R1 0528
May 28th update to the original DeepSeek R1 Performance on par…
164K
$0.50
$2.15
May 2025
Mixtral 8X22b Instruct V0.1
Mistralai chat model. https://huggingface.co/api/models/mistral…
66K
-
-
-
May 2025
Anthropic: Claude Sonnet 4
Claude Sonnet 4 significantly enhances the capabilities of its…
1M
$3.00
$15.00
May 2025
Anthropic: Claude Opus 4
Claude Opus 4 is benchmarked as the world’s best coding model,…
200K
$15.00
$75.00
May 2025
Devstral Small 2505
Mistralai chat model. https://huggingface.co/api/models/togethe…
131K
-
-
-
May 2025
Mistral 7B v0.1
Mistralai chat model. https://huggingface.co/api/models/mistral…
33K
-
-
-
May 2025
Google: Gemma 3n 4B
Gemma 3n E4B-it is optimized for efficient execution on mobile…
33K
$0.06
$0.12
May 2025
Google: Gemini 2.5 Pro Preview 06-05
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed f…
1M
$1.25
$10.00
May 2025
Gemini 2.5 Pro Preview TTS
Gemini 2.5 Pro Preview TTS
25K
$1.00
-
May 2025
Gemini 2.5 Flash Preview TTS
Gemini 2.5 Flash Preview TTS
25K
$0.50
-
May 2025
Deepcoder 14B Preview
Togethercomputer chat model. https://huggingface.co/api/models/…
131K
-
-
-
May 2025
mistral-medium-3
Official mistral-medium-latest Mistral AI model
262K
$1.50
$7.50
May 2025
Mistral Medium (2505)
Our frontier-class multimodal model released May 2025.
131K
$0.40
$2.00
May 2025
Google: Gemini 2.5 Pro Preview 05-06
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed f…
1M
$1.25
$10.00
May 2025
Arcee AI: Virtuoso Large
Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B…
131K
$0.75
$1.20
May 2025
Arcee AI: Coder Large
Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct…
33K
$0.50
$0.80
May 2025
Meta: Llama Guard 4 12B
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained…
164K
$0.18
$0.18
Apr 2025
Qwen3 8b
Alibaba model (not yet curated).
131K
$0.12
$0.46
Apr 2025
Qwen3 32b
Alibaba model (not yet curated).
131K
$0.08
$0.28
Apr 2025
Qwen3 30b A3b
Alibaba model (not yet curated).
131K
$0.12
$0.50
Apr 2025
Qwen3 235b A22b
Alibaba model (not yet curated).
131K
$0.46
$1.82
Apr 2025
Qwen3 14b
Alibaba model (not yet curated).
131K
$0.12
$0.24
Apr 2025
Arize AI Qwen 2 1.5B Instruct
Togethercomputer chat model. https://huggingface.co/api/models/…
33K
$0.10
$0.10
-
Apr 2025
o4 Mini
Latest o4-mini model. Optimized for fast, effective reasoning w…
200K
$1.10
$4.40
Apr 2025
o3
A well-rounded and powerful model across domains. Sets a new st…
200K
$2.00
$8.00
Apr 2025
Llama 3.1 405B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
131K
-
-
-
Apr 2025
Qwen2.5 7B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
33K
-
-
-
Apr 2025
Qwen2.5 7B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
131K
-
-
-
Apr 2025
Qwen2.5 72B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
131K
-
-
-
Apr 2025
Qwen2.5 3B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
33K
-
-
-
Apr 2025
Qwen2.5 32B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
33K
-
-
-
Apr 2025
Qwen2.5 32B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
131K
-
-
-
Apr 2025
Qwen2.5 14B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
131K
-
-
-
Apr 2025
Qwen2.5 1.5B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
33K
-
-
-
Apr 2025
Qwen2.5 1.5B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
131K
-
-
-
Apr 2025
Llama 3.2 1B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
131K
-
-
-
Apr 2025
Llama 3.1 70B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
131K
-
-
-
Apr 2025
GPT-4.1 Nano
Fastest, most cost-effective GPT 4.1 model. Delivers exceptiona…
1M
$0.10
$0.40
Apr 2025
GPT-4.1 Mini
Balanced for intelligence, speed, and cost. Matches or exceeds…
1M
$0.40
$1.60
Apr 2025
GPT-4.1
Flagship GPT model for complex tasks. Major improvements on cod…
1M
$2.00
$8.00
Apr 2025
GLM-4 32B (0414) 128K
GLM-4 32B model with 128K context, 16K output.
131K
$0.10
$0.10
Apr 2025
Qwen2 72B Instruct
Togethercomputer chat model. https://huggingface.co/api/models/…
33K
$0.90
$0.90
-
Apr 2025
Cogito V1 Preview Qwen 32B
deepcogito chat model.
131K
-
-
-
Apr 2025
Cogito V1 Preview Qwen 14B
deepcogito chat model.
131K
-
-
-
Apr 2025
Cogito V1 Preview Llama 8B
deepcogito chat model.
131K
-
-
-
Apr 2025
Cogito V1 Preview Llama 70B Turbo
deepcogito chat model.
131K
-
-
-
Apr 2025
Cogito V1 Preview Llama 70B
deepcogito chat model.
131K
-
-
-
Apr 2025
Meta: Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE)…
1.3M
$0.11
$0.34
Apr 2025
Meta: Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimo…
1M
$0.20
$0.70
Apr 2025
Llama 4 Scout Instruct (17Bx16E)
Meta chat model. https://huggingface.co/meta-llama/Llama-4-Scou…
1M
$0.18
$0.59
-
Apr 2025
Gemma 3 1b it
Google chat model.
33K
-
-
-
Apr 2025
DeepSeek R1 Distill Qwen 7B
Deepseek chat model. https://huggingface.co/deepseek-ai/DeepSee…
131K
-
-
-
Apr 2025
meta-llama/Llama-2-7b-chat-hf
Meta chat model. https://huggingface.co/meta-llama/Llama-2-7b-c…
4K
-
-
-
Apr 2025
DeepSeek: DeepSeek V3 0324
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the…
164K
$0.25
$1.00
Mar 2025
o1 Pro
A version of o1 with more compute for better responses. Provide…
200K
$150.00
$600.00
Mar 2025
Llama 3.3 Nemotron Super 49B v1
Superseded by v1.5.
131K
Free
Free
Mar 2025
Mistral: Mistral Small 3.1 24B
Mistral Small 3.1 24B Instruct is an upgraded variant of Mistra…
128K
$0.35
$0.56
Mar 2025
nim/nv-mistralai/mistral-nemo-12b-instruct
NVIDIA chat model.
16K
-
-
-
Mar 2025
nim/mistralai/mixtral-8x7b-instruct-v01
mistralai chat model.
16K
-
-
-
Mar 2025
Google: Gemma 3 4B
Gemma 3 introduces multimodality, supporting vision-language in…
131K
$0.05
$0.10
Mar 2025
Google: Gemma 3 12B
Gemma 3 introduces multimodality, supporting vision-language in…
131K
$0.05
$0.15
Mar 2025
CO
Command A
Cohere's efficient 111B enterprise model for agents, tool use,…
288K
$2.50
$10.00
Mar 2025
Cohere: Command A
Command A is an open-weights 111B parameter model with a 256k c…
256K
$2.50
$10.00
Mar 2025
Reka Flash 3
Reka Flash 3 is a general-purpose, instruction-tuned large lang…
66K
$0.10
$0.20
Mar 2025
nim/nvidia/llama-3.1-nemotron-70b-instruct
NVIDIA chat model.
16K
-
-
-
Mar 2025
Google: Gemma 3 27B
Gemma 3 introduces multimodality, supporting vision-language in…
131K
$0.08
$0.45
Mar 2025
GPT-4o Search Preview
deprecatedGPT-4o model optimized for web search capabilities. Alias still…
128K
$2.50
$10.00
Mar 2025
GPT-4o Mini Search Preview
deprecatedGPT-4o Mini model optimized for web search capabilities. Alias…
128K
$0.15
$0.60
Mar 2025
TheDrummer: Skyfall 36B V2
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501,…
33K
$0.55
$0.80
Mar 2025
nim/mistralai/mixtral-8x22b-instruct-v01
Mistral chat model.
16K
-
-
-
Mar 2025
Meta Llama 3.1 8B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
131K
$0.18
$0.18
-
Mar 2025
Qwen QwQ-32B
Qwen chat model. https://huggingface.co/Qwen/QwQ-32B
131K
$1.20
$1.20
-
Mar 2025
CO
Aya Vision 32B
Open-weights multilingual vision research model (23 languages)…
16K
$0.50
$1.50
Mar 2025
Sonar Reasoning Pro
Premier reasoning model with enhanced multi-step Chain of Thoug…
128K
$2.00
$8.00
Feb 2025
Mistral: Saba
Mistral Saba is a 24B-parameter language model specifically des…
33K
$0.20
$0.60
Feb 2025
Sonar Deep Research
Expert-level research model for exhaustive searches and compreh…
128K
$2.00
$8.00
Feb 2025
Gemini 2.0 Flash 001
Stable version of Gemini 2.0 Flash, our fast and versatile mult…
1.1M
$0.10
$0.40
Feb 2025
AionLabs: Aion-RP 1.0 (8B)
Aion-RP-Llama-3.1-8B ranks the highest in the character evaluat…
33K
$0.80
$1.60
Feb 2025
AionLabs: Aion-1.0-Mini
Aion-1.0-Mini 32B parameter model is a distilled version of the…
131K
$0.70
$1.40
Feb 2025
AionLabs: Aion-1.0
Aion-1.0 is a multi-model system designed for high performance…
131K
$4.00
$8.00
Feb 2025
Qwen: Qwen2.5 VL 72B Instruct
Qwen2.5-VL is proficient in recognizing common objects such as…
128K
$0.25
$0.75
Feb 2025
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M cont…
1M
$0.40
$1.20
Feb 2025
CO
Command R7B Arabic
Command R7B tuned for Modern Standard Arabic and English enterp…
128K
$0.04
$0.15
Feb 2025
o3 Mini
Latest o3-mini model snapshot. High intelligence at the same co…
200K
$1.10
$4.40
Jan 2025
Mistral: Mistral Small 3
Mistral Small 3 is a 24B-parameter language model optimized for…
33K
$0.05
$0.08
Jan 2025
DeepSeek R1 Distill Qwen 14B
DeepSeek chat model. https://huggingface.co/api/models/deepseek…
131K
$1.60
$1.60
-
Jan 2025
DeepSeek R1 Distill Qwen 1.5B
DeepSeek chat model. https://huggingface.co/deepseek-ai/DeepSee…
131K
$0.18
$0.18
-
Jan 2025
DeepSeek: R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B is a distilled large language mod…
8K
$0.80
$0.80
Jan 2025
Sonar Pro
Advanced search model for complex queries and deep content unde…
200K
$3.00
$15.00
Jan 2025
Sonar
Lightweight, cost-effective search model for quick, grounded an…
128K
$1.00
$1.00
Jan 2025
DeepSeek: R1
DeepSeek R1 is here: Performance on par with OpenAI o1, but ope…
64K
$0.70
$2.50
Jan 2025
NemoGuard 8B Topic Control
NVIDIA topic-control guardrail (not a chat model).
131K
Free
Free
-
Jan 2025
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
131K
Free
Free
-
Jan 2025
V1 8K Vision (Preview)
Legacy vision model with 8K context. Sunset on 2026-08-31 - use…
8K
$0.20
$2.00
Jan 2025
V1 32K Vision (Preview)
Legacy vision model with 32K context. Sunset on 2026-08-31 - us…
33K
$1.00
$3.00
Jan 2025
V1 128K Vision (Preview)
Legacy vision model with 128K context. Sunset on 2026-08-31 - u…
131K
$2.00
$5.00
Jan 2025
MiniMax: MiniMax-01
MiniMax-01 is a combines MiniMax-Text-01 for text generation an…
1M
$0.20
$1.10
Jan 2025
Microsoft: Phi 4
Microsoft Research Phi-4 is designed to perform well in complex…
16K
$0.07
$0.14
Jan 2025
Qwen2-VL (72B) Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2-VL-72B-Instr…
33K
$1.20
$1.20
Jan 2025
Sao10K: Llama 3.1 70B Hanami x1
This is Sao10K's experiment over Euryale v2.2.
16K
$3.00
$3.00
-
Jan 2025
DeepSeek: DeepSeek V3
DeepSeek-V3 is the latest model from the DeepSeek team, buildin…
164K
$0.26
$1.03
Dec 2024
Sao10K: Llama 3.3 Euryale 70B
Euryale L3.3 70B is a model focused on creative roleplay from S…
131K
$0.65
$0.75
Dec 2024
o1
Previous full o-series reasoning model.
200K
$15.00
$60.00
Dec 2024
CO
Cohere: Command R7B (12-2024)
Command R7B (12-2024) is a small, fast update of the Command R+…
128K
$0.04
$0.15
Dec 2024
Qwen 2.5 14B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
33K
$0.80
$0.80
-
Dec 2024
Qwen2.5 72B Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-72B-Instru…
33K
$1.20
$1.20
-
Dec 2024
Meta: Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a…
131K
$0.71
$0.71
Dec 2024
Meta Llama 3.3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Llama-3.3-70…
131K
$1.04
$1.04
-
Dec 2024
Meta Llama 3.1 405B Instruct
Meta chat model. https://huggingface.co/meta-llama/Llama-3.1-40…
4K
$3.50
$3.50
-
Dec 2024
[Meta] Llama 3.3 · 70B Versatile
Meta Llama 3.3 (70B params) with GQA. Strong reasoning, coding,…
131K
$0.59
$0.79
Dec 2024
Amazon: Nova Pro 1.0
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon f…
300K
$0.80
$3.20
Dec 2024
Amazon: Nova Micro 1.0
Amazon Nova Micro 1.0 is a text-only model that delivers the lo…
128K
$0.04
$0.14
Dec 2024
Amazon: Nova Lite 1.0
Amazon Nova Lite 1.0 is a very low-cost multimodal model from A…
300K
$0.06
$0.24
Dec 2024
Mistral Large 2407
This is Mistral AI's flagship model, Mistral Large 2 (version m…
131K
$2.00
$6.00
Nov 2024
Qwen 2.5 Coder 32B Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-Coder-32B-…
16K
$0.80
$0.80
-
Nov 2024
Qwen2.5 Coder 32B Instruct
Qwen2.5-Coder is the latest series of Code-Specific Qwen large…
33K
$0.66
$1.00
Nov 2024
Llama 3.1 Nemotron 70B Instruct HF
nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemo…
33K
$0.88
$0.88
-
Nov 2024
TheDrummer: UnslopNemo 12B
UnslopNemo v4.1 is the latest addition from the creator of Roci…
1M
$0.40
$0.40
Nov 2024
Magnum v4 72B
This is a series of models designed to replicate the prose qual…
33K
$3.00
$5.00
Oct 2024
Mistral: Ministral 8B
Ministral 8B is an 8B parameter model featuring a unique interl…
128K
$0.11
$0.11
Oct 2024
Qwen: Qwen2.5 7B Instruct
Qwen2.5 7B is the latest series of Qwen large language models.…
33K
$0.10
$0.20
Oct 2024
Qwen2.5 7B Instruct Turbo
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-7B-Instruct
33K
$0.30
$0.30
-
Oct 2024
Qwen2.5 72B Instruct Turbo
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-72B-Instru…
131K
$1.20
$1.20
-
Oct 2024
Inflection: Inflection 3 Productivity
Inflection 3 Productivity is optimized for following instructio…
8K
$2.50
$10.00
Oct 2024
Inflection: Inflection 3 Pi
Inflection 3 Pi powers Inflection's Pi chatbot, including backs…
8K
$2.50
$10.00
Oct 2024
CO
Aya Expanse 32B
Open-weights multilingual research model covering 23 languages.…
128K
$0.50
$1.50
-
Oct 2024
TheDrummer: Rocinante 12B
Rocinante 12B is designed for engaging storytelling and rich pr…
66K
$0.25
$0.50
Sep 2024
Meta: Llama 3.2 3B Instruct
Llama 3.2 3B is a 3-billion-parameter multilingual large langua…
131K
$0.05
$0.33
Sep 2024
Meta: Llama 3.2 1B Instruct
Llama 3.2 1B is a 1-billion-parameter language model focused on…
60K
$0.03
$0.20
Sep 2024
Meta: Llama 3.2 11B Vision Instruct
Llama 3.2 11B Vision is a multimodal model with 11 billion para…
131K
$0.35
$0.35
Sep 2024
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (…
33K
Free
Free
Sep 2024
Qwen2.5 72B Instruct
Qwen2.5 72B is the latest series of Qwen large language models.…
33K
$0.36
$0.40
Sep 2024
Nemotron Mini 4B
Tiny legacy Nemotron, 4K context.
4K
Free
Free
Sep 2024
CO
Cohere: Command R+ (08-2024)
command-r-plus-08-2024 is an update of the Command R+ with roug…
128K
$2.50
$10.00
Aug 2024
CO
Cohere: Command R (08-2024)
command-r-08-2024 is an update of the Command R with improved p…
128K
$0.15
$0.60
Aug 2024
Sao10K: Llama 3.1 Euryale 70B v2.2
Euryale L3.1 70B v2.2 is a model focused on creative roleplay f…
131K
$0.85
$0.85
Aug 2024
Nous: Hermes 3 70B Instruct
Hermes 3 is a generalist language model with many improvements…
131K
$0.70
$0.70
Aug 2024
Nous: Hermes 3 405B Instruct (free)
Hermes 3 is a generalist language model with many improvements…
131K
Free
Free
Aug 2024
Sao10K: Llama 3 8B Lunaris
Lunaris 8B is a versatile generalist and roleplaying model base…
8K
$0.04
$0.05
Aug 2024
Meta: Llama 3.1 8B Instruct
Meta's latest class of model (Llama 3.1) launched with a variet…
131K
$0.05
$0.08
Jul 2024
Meta: Llama 3.1 70B Instruct
Meta's latest class of model (Llama 3.1) launched with a variet…
131K
$0.40
$0.40
Jul 2024
[Meta] Llama 3.1 · 8B Instant
Meta Llama 3.1 (8B params). Fast, cost-effective for high-volum…
131K
$0.05
$0.08
Jul 2024
Meta Llama 3.1 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
131K
$0.88
$0.88
-
Jul 2024
Mistral: Mistral Nemo
A 12B parameter model with a 128k token context length built by…
131K
$0.02
$0.03
Jul 2024
open-mistral-nemo-2407
Our best multilingual open source model released July 2024.
131K
$0.15
$0.15
Jul 2024
open-mistral-nemo
Our best multilingual open source model released July 2024.
131K
$0.15
$0.15
Jul 2024
GPT-4o mini
Affordable model for fast, lightweight tasks. GPT-4o Mini is ch…
128K
$0.15
$0.60
Jul 2024
Google: Gemma 2 27B
Gemma 2 27B by Google is an open model built from the same rese…
8K
$0.65
$0.65
Jul 2024
Qwen 2 Instruct (1.5B)
Qwen chat model. https://huggingface.co/Qwen/Qwen2-72B-Instruct
33K
$0.02
$0.02
-
Jun 2024
Mistral (7B) Instruct v0.3
mistralai chat model. https://huggingface.co/api/models/mistral…
33K
$0.20
$0.20
-
May 2024
GPT-4o
deprecatedOriginal gpt-4o snapshot from May 13, 2024.
128K
$5.00
$15.00
May 2024
Meta: Llama 3 8B Instruct
Meta's latest class of model (Llama 3) launched with a variety…
8K
$0.14
$0.14
Apr 2024
Meta Llama 3 8B Instruct Reference
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
8K
$0.20
$0.20
-
Apr 2024
Meta Llama 3 8B Instruct
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
8K
$0.20
$0.20
-
Apr 2024
Mistral: Mixtral 8x22B Instruct
Mistral's official instruct fine-tuned version of Mixtral 8x22B…
66K
$2.00
$6.00
Apr 2024
WizardLM-2 8x22B
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model.…
66K
$0.62
$0.62
Apr 2024
GPT-4 Turbo
GPT-4 Turbo with Vision model. Vision requests can now use JSON…
128K
$10.00
$30.00
Apr 2024
Anthropic: Claude 3 Haiku
Claude 3 Haiku is Anthropic's fastest and most compact model fo…
200K
$0.25
$1.25
Mar 2024
Mistral Large
This is Mistral AI's flagship model, Mistral Large 2 (version `…
128K
$2.00
$6.00
Feb 2024
Deepseek Coder 33B Instruct
Deepseek chat model. https://huggingface.co/api/models/deepseek…
16K
$0.80
$0.80
-
Feb 2024
V1 8K
Legacy V1 model with 8K context. Sunset on 2026-08-31 - use Kim…
8K
$0.20
$2.00
Feb 2024
V1 32K
Legacy V1 model with 32K context. Sunset on 2026-08-31 - use Ki…
33K
$1.00
$3.00
Feb 2024
V1 128K
Legacy V1 model with 128K context. Sunset on 2026-08-31 - use K…
131K
$2.00
$5.00
Feb 2024
OpenAI: GPT-4 Turbo Preview
The preview GPT-4 model with improved instruction following, JS…
128K
$10.00
$30.00
Jan 2024
OpenAI: GPT-3.5 Turbo (older v0613)
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and…
4K
$1.00
$2.00
Jan 2024
3.5-Turbo
deprecatedThe latest GPT-3.5 Turbo model with higher accuracy at respondi…
16K
$0.50
$1.50
Jan 2024
Nous Hermes 2 Mixtral 8X7B Dpo
Nousresearch chat model. https://huggingface.co/api/models/Nous…
33K
$0.60
$0.60
-
Jan 2024
Mixtral-8x7B Instruct v0.1
mistralai chat model. https://huggingface.co/mistralai/Mixtral-…
33K
$0.60
$0.60
-
Dec 2023
Auto Router
The Auto Router automatically selects the best model for your p…
2M
-
-
Nov 2023
3.5-Turbo
deprecatedGPT-3.5 Turbo model with improved instruction following, JSON m…
16K
$1.00
$2.00
Nov 2023
OpenAI: GPT-3.5 Turbo Instruct
This model is a variant of GPT-3.5 Turbo tuned for instructiona…
4K
$1.50
$2.00
Sep 2023
Mistral (7B) Instruct v0.1
mistralai chat model. https://huggingface.co/api/models/mistral…
33K
$0.20
$0.20
-
Sep 2023
OpenAI: GPT-3.5 Turbo 16k
This model offers four times the context length of gpt-3.5-turb…
16K
$3.00
$4.00
Aug 2023
Mancer: Weaver (alpha)
An attempt to recreate Claude-style verbosity, but don't expect…
8K
$0.50
$0.75
Aug 2023
ReMM SLERP 13B
A recreation trial of the original MythoMax-L2-B13 but with upd…
6K
$0.45
$0.65
Jul 2023
MythoMax 13B
One of the highest performing and most popular fine-tunes of Ll…
8K
$0.06
$0.06
Jul 2023
GPT-4
Snapshot of gpt-4 from June 13th 2023 with improved function ca…
8K
$30.00
$60.00
Jun 2023
GPT-4
deprecatedSnapshot of gpt-4 from June 13th 2023 with improved function ca…
8K
$30.00
$60.00
Jun 2023
3.5-Turbo
The latest GPT-3.5 Turbo model with higher accuracy at respondi…
16K
$0.50
$1.50
May 2023
? [gpt live transcribe]
Unknown, please let us know the ID. Assuming a context window o…
128K
-
-
-
? [gpt transcribe]
Unknown, please let us know the ID. Assuming a context window o…
128K
-
-
-
? [ra gpt 5.6 sol]
Unknown, please let us know the ID. Assuming a context window o…
128K
-
-
-
AI21 Labs Jamba 1.5 Large
AI21 Labs model via Unsupported API (Bedrock Foundation Model)
-
-
-
-
-
AI21 Labs Jamba 1.5 Mini
AI21 Labs model via Unsupported API (Bedrock Foundation Model)
-
-
-
-
-
Amazon Nova 2 Lite
Amazon model via Converse API (Bedrock Inference Profile)
-
-
-
-
Amazon Nova Lite
Amazon model via Converse API (Bedrock Inference Profile)
-
-
-
-
Amazon Nova Micro
Amazon model via Converse API (Bedrock Inference Profile)
-
-
-
-
Amazon Nova Premier
Amazon model via Converse API (Bedrock Inference Profile)
-
-
-
-
Amazon Nova Pro
Amazon model via Converse API (Bedrock Foundation Model)
10K
-
-
-
Anthropic Claude 3 Sonnet
Anthropic model (Bedrock Inference Profile)
200K
-
-
-
Anthropic Honey
Anthropic model via OpenAI-Compatible API on AWS Bedrock Mantle
131K
-
-
-
-
Cohere Command R
Cohere model via Unsupported API (Bedrock Foundation Model)
-
-
-
-
-
Cohere Command R+
Cohere model via Unsupported API (Bedrock Foundation Model)
-
-
-
-
-
Cohere Embed v4
Cohere model via Converse API (Bedrock Inference Profile)
-
-
-
-
DeepSeek-R1
Deepseek model via Converse API (Bedrock Inference Profile)
-
-
-
-
FLUX.2 Klein 4B
Model served via Modular.
-
-
-
-
Gemma 4 12B It
Google chat model. https://huggingface.co/google/gemma-4-12B-it
262K
-
-
-
GLM 5.2 FP8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
1M
-
-
-
-
GLM 5.2 FP8 Lora
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
1M
-
-
-
-
GLM-5.2 Fast (Alibaba)
Zhipu GLM-5.2 fast-serving tier via Alibaba Model Studio (previ…
1M
$2.80
$8.80
-
Google Gemma 4 26b A4b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
131K
-
-
-
Google Gemma 4 31b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
131K
$0.99
$1.49
-
Google Gemma 4 E2b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
131K
-
-
-
Kimi K2.5 Fp4
Togethercomputer chat model. https://huggingface.co/api/models/…
262K
$0.50
$2.80
-
-
LFM2-24B-A2B
Togethercomputer chat model.
33K
$0.03
$0.12
-
-
LFM2.5-8B-A1B
LiquidAI chat model. https://huggingface.co/api/models/LiquidAI…
128K
$0.03
$0.12
-
-
Meta Llama 3 70B Instruct
Meta model via Converse API (Bedrock Foundation Model)
-
-
-
-
Meta Llama 3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
8K
$0.88
$0.88
-
-
Meta Llama 3 8B Instruct
Meta model via Converse API (Bedrock Foundation Model)
-
-
-
-
Meta Llama 3 8B Instruct Lite
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
8K
$0.14
$0.14
-
-
Meta Llama 3.1 70B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.1 8B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.2 11B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.2 1B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.2 3B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.2 90B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 3.3 70B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 4 Maverick 17B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Meta Llama 4 Scout 17B Instruct
Meta model via Converse API (Bedrock Inference Profile)
-
-
-
-
Mistral AI Devstral 2 123B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
262K
-
-
-
Mistral AI Ministral 14B 3.0
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
262K
-
-
-
Mistral AI Ministral 3 8B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
262K
-
-
-
Mistral AI Ministral 3B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
262K
-
-
-
Mistral AI Mistral 7B Instruct
Mistral AI model via Converse API (Bedrock Foundation Model)
-
-
-
-
Mistral AI Mistral Large (24.02)
Mistral AI model via Converse API (Bedrock Foundation Model)
-
-
-
-
Mistral AI Mistral Large 3
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
262K
-
-
-
Mistral AI Mistral Small (24.02)
Mistral AI model via Converse API (Bedrock Foundation Model)
-
-
-
-
Mistral AI Mixtral 8x7B Instruct
Mistral AI model via Converse API (Bedrock Foundation Model)
-
-
-
-
Mistral AI Voxtral Mini 3B 2507
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
33K
-
-
-
Mistral Pixtral Large 25.02
Mistral model via Converse API (Bedrock Inference Profile)
-
-
-
-
mistral-tiny-latest
Our best multilingual open source model released July 2024.
131K
-
-
-
NVIDIA Nemotron 3 Super 120B A12B
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
262K
-
-
-
NVIDIA Nemotron Nano 12B v2 VL BF16
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
131K
-
-
-
NVIDIA Nemotron Nano 3 30B
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
262K
-
-
-
NVIDIA Nemotron Nano 9B v2
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
131K
-
-
-
Open Mistral Nemo
Our best multilingual open source model released July 2024.
131K
-
-
-
OpenAI GPT OSS Safeguard 120B
OpenAI model via OpenAI-Compatible API (Bedrock Foundation Mode…
131K
-
-
-
OpenAI gpt-oss-120b
OpenAI model via Converse API (Bedrock Foundation Model)
128K
-
-
-
OpenAI gpt-oss-20b
OpenAI model via Converse API (Bedrock Foundation Model)
128K
-
-
-
CO
parse v5.0
New Cohere Model
128K
-
-
-
-
Qvq Max
Alibaba model (not yet curated).
131K
-
-
-
Qwen Coder Plus
Alibaba model (not yet curated).
131K
-
-
-
Qwen Flash
Fast and very low cost with hybrid thinking. 1M context.
1M
$0.05
$0.40
-
Qwen Max
Best quality of the stable commercial line. 32K context.
33K
$1.60
$6.40
-
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M cont…
1M
$0.40
$1.20
-
Qwen Turbo
Fastest and cheapest for simple tasks. 1M context.
1M
$0.05
$0.20
-
Qwen Vl Max
Alibaba model (not yet curated).
131K
-
-
-
Qwen Vl Plus
Alibaba model (not yet curated).
131K
-
-
-
Qwen3 235b A22b Instruct 2507
Alibaba model (not yet curated).
131K
-
-
-
Qwen3 Coder 480b A35b Instruct
Alibaba model (not yet curated).
131K
-
-
-
Qwen3 Coder Next Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-C…
262K
$0.50
$1.20
-
-
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context…
262K
$1.20
$6.00
-
Qwen3 Next 80B A3B
Qwen model via Converse API (Bedrock Foundation Model)
262K
-
-
-
Qwen3 VL 235B A22B
Qwen model via Converse API (Bedrock Foundation Model)
262K
-
-
-
Qwen3 VL Plus
Current vision-language model with strong visual reasoning and…
262K
$0.20
$1.60
-
Qwen3-Coder-30B-A3B-Instruct
Qwen model via Converse API (Bedrock Foundation Model)
262K
-
-
-
Qwen3.5 35B A3B Lora
Qwen chat model.
262K
-
-
-
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M conte…
1M
$2.50
$7.50
-
Qwq Plus
Alibaba model (not yet curated).
131K
-
-
-
TwelveLabs Pegasus v1.2
Twelvelabs model via Converse API (Bedrock Inference Profile)
-
-
-
-
Twelvelabs TwelveLabs Marengo Embed 3.0
Twelvelabs model via Converse API (Bedrock Inference Profile)
-
-
-
-
Twelvelabs TwelveLabs Marengo Embed v2.7
Twelvelabs model via Converse API (Bedrock Inference Profile)
-
-
-
-
Writer Palmyra Vision 7B
Writer model via OpenAI-Compatible API (Bedrock Foundation Mode…
4K
-
-
-
Writer Palmyra X4
Writer model via Converse API (Bedrock Inference Profile)
-
-
-
-
Tencent: Hy4 preview
NEWTencent: Hy4 preview is a mixture-of-experts model from Tencent…
Ling 3.0 Flash Fin (free)
NEWLing 3.0 Flash Fin is a finance-focused mixture-of-experts mode…
Gemini Omni 1.1 Flash (video)
NEWGemini Omni 1.1 Flash
Qwen: Qwen3.8 Flash
NEWQwen3.8 Flash is a multimodal reasoning model from Alibaba. It…
Gemini 3.5 Transcribe
NEWGemini 3.5 Transcribe
GLM-5.3 Flash (1M)
NEWHOTMultimodal Flash on a new 320B MoE base (18B activated, hybrid…
Meta: Muse Spark 1.2 Contributor
NEWMuse Spark 1.2 contributor tier is a reasoning model from Meta…
DeepSeek V4 Flash Vision (Exp)
NEWExperimental vision variant of V4 Flash with 1M context, releas…
Tencent: Hy-MT2-30B-A3B
NEWHy-MT2-30B-A3B is Tencent's flagship translation model in the H…
Tencent: Hy-MT2-1.8B
NEWHy-MT2-1.8B is a compact 1.8B-parameter translation model from…
Ox Alpha
NEWOx Alpha is a reasoning model designed for coding, sustained ag…
Tencent: Hy-MT2-7B
NEWHy-MT2-7B is a 7B-parameter translation model from Tencent. It…
Z.ai: GLM Latest
NEWThis model always redirects to the latest GLM model from Z.ai.
Qwen3.8 27B
NEWOpen-weights 27B dense vision-language model of the Qwen3.8 lin…
GLM-5.3 (1M)
NEWZ.ai 1M-context flagship, post-trained on the GLM-5.2 base for…
Gemini 3.7 Flash Video Understanding
NEWGemini 3.7 Flash Video Understanding EAP
Dots Studio: Dots3-Note Preview (free)
NEWDots3-Note Preview is an open-weight mixture-of-experts model f…
Google Gemini Flash Latest
NEWThis model always redirects to the latest model in the Google G…
Gemini 3.7 Flash
NEWHOTGemini 3.7 Flash
xAI: Grok Latest
NEWThis model always redirects to the latest Grok model from xAI.
Qwen3.8 2.4T-A95B
NEWOpen-weights release of the Qwen3.8 flagship: 2.4T sparse MoE,…
Grok 4.6
NEWxAI's frontier model for coding, agentic tasks, and knowledge w…
DeepSeek: DeepSeek V4 Pro 0813
NEWHOTDeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model…
ByteDance Seed: Seed-2.0-Code
NEWSeed 2.0 Code is a model from ByteDance Seed optimized for agen…
ByteDance Seed: Seed 2.1 Turbo
NEWSeed 2.1 Turbo is a multimodal model from ByteDance Seed for co…
Sakana: Namazu
NEWSakana Namazu is a Japanese-specialized reasoning model from Sa…
NVIDIA: Nemotron 3.5 Lightning
NEWHOTNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts mod…
Nemotron 3.5 Lightning 30B
NEWFastest Nemotron MoE (30B, 3B active), 1M context, reasoning an…
LiquidAI: LFM2.5-2.6B (free)
NEWLFM2.5-2.6B is a compact reasoning model from Liquid AI. It is…
Upstage: Solar Pro 4
NEWSolar Pro 4 is Upstage's cost-efficient large language model, f…
Meta: Muse Glimmer 30B
NEWMuse Glimmer 30B is a dense, open-weight multimodal model from…
inclusionAI: Ling 3.0 Tiny (free)
NEWLing 3.0 Tiny is a mixture-of-experts model from InclusionAI, w…
Meta: Muse Spark 1.2
NEWMuse Spark 1.2 is a reasoning model from Meta, designed for com…
Sakana Namazu v1.0
NEWJapanese-specialized LLM built on Moonshot AI's Kimi K2.6 and a…
Sakana Namazu
NEWJapanese-specialized LLM with built-in web search and code exec…
DeepSeek V4 Flash Latest
NEWThis model always redirects to the latest model in the DeepSeek…
DeepSeek: DeepSeek V4 Flash 0731
NEWHOTDeepSeek V4 Flash 0731 is a sparse mixture-of-experts model fro…
Thinking Machines: Inkling Small
NEWInkling Small is an open-weight multimodal mixture-of-experts m…
Gemini Robotics-ER 2 Preview
NEWGemini Robotics-ER 2 Preview
Riva Translate 4B v2
NEWTranslation-specialized model (37 languages, few-shot prompting…
Claude Opus 5 (Fast)
NEWFast-mode variant of Opus 5 - identical capabilities with highe…
Claude Opus 5
NEWHOTStep-change improvement over Opus 4.8 for complex agentic codin…
Anthropic: Claude Opus Latest
NEWThis model always redirects to the latest model in the Claude O…
Ling-3.0-flash
NEWHOT*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE)…
Sakana Fugu Ultra v1.1
NEWMulti-agent conductor system routing 1-3 expert agents for comp…
Sakana Fugu Cyber
NEWOrchestrator specialized for cybersecurity reasoning: security…
Poolside: Laguna S 2.1
NEWLaguna S 2.1 is the latest coding agent model from Poolside. La…
Gemini 3.6 Flash
NEWHOTGemini 3.6 Flash
Gemini 3.5 Flash-Lite
NEWGemini 3.5 Flash Lite
Meituan: LongCat 2.0
NEWLongCat 2.0 is a sparse mixture-of-experts language model from…
Ising Calibration 1.5 31B
NEWNVIDIA quantum-calibration VLM, preview (domain-specific, not f…
Qwen3.8 Max
NEWFlagship 2.4T-parameter sparse MoE multimodal model with 1M con…
Thinking Machines: Inkling
Inkling is an open-weight multimodal mixture-of-experts model f…
Auto Router (Beta)
Auto Router (Beta) is a task-aware router from OpenRouter. It c…
MoonshotAI Kimi Latest
This model always redirects to the latest model in the Moonshot…
Meta: Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built…
Kimi K3
HOTNative multimodal flagship (text, image, video inputs) with thi…
Qwen3.7 Flash
Latest fast multimodal model with 1M context, thinking (on by d…
Ternary Bonsai 27B
Prism Ml chat model. https://huggingface.co/api/models/prism-ml…
Kwaipilot: KAT-Coder-Pro V2.5
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model tha…
Kwaipilot: KAT-Coder-Air V2.5 (free)
KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model tha…
OpenAI: GPT-5.6 Terra Pro
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra…
OpenAI: GPT-5.6 Sol Pro
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, se…
OpenAI: GPT-5.6 Luna Pro
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna,…
OpenAI GPT Latest
This model always redirects to the latest model in the OpenAI G…
GPT-5.6 Terra
Balanced model for efficient, high-volume everyday work. Compet…
GPT-5.6 Sol
HOTFlagship next-generation model. Strongest yet for agentic codin…
GPT-5.6 Luna
Fastest, most affordable GPT-5.6 model for high-volume work. St…
Grok 4.5
xAI's July 2026 flagship with frontier performance on coding, k…
AionLabs: Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling sys…
AionLabs: Aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system f…
Tencent: Hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (…
Poolside: Laguna XS 2.1
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B c…
Nano Banana 2 Lite
Gemini 3.1 Flash Lite Image.
Leanstral 1.5
A mid & post-trained version of mistral small 4 for Lean (26061…
labs-leanstral-1-5
A mid & post-trained version of mistral small 4 for Lean (26061…
Gemini Omni Flash Preview (video)
Gemini Omni Flash Preview
Claude Sonnet 5
HOTBest combination of speed and intelligence, with the largest ga…
Anthropic Claude Sonnet Latest
This model always redirects to the latest model in the Anthropi…
Qwen3.6 35B A3B Lora
Qwen chat model.
Qwen3.5 2B Lora
Qwen chat model.
Nex AGI: Nex-N2-Mini
HOTNex-N2-Mini is an open-source agentic mixture-of-experts model…
Sakana Fugu
Fast orchestration model routing tasks across a swappable pool…
Cohere: North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the…
Z.ai GLM 5.2
Official Z.ai GLM 5.2 model
GLM-5.2 (1M)
Z.ai 1M-context flagship (744B MoE, 40B activated). Agentic cod…
Sakana Fugu Ultra v1.0
Multi-agent conductor system routing 1-3 expert agents for comp…
Sakana Fugu Ultra
Multi-agent conductor system routing 1-3 expert agents for comp…
OpenRouter: Fusion
Fusion turns your prompt into a small multi-model deliberation.…
Kimi K2.7 Code Highspeed
High-speed code variant with ~180 tok/s output (up to 260 in sh…
Kimi K2.7 Code
Code-focused multimodal model (text, image, video inputs) with…
DiffusionGemma 26B
Experimental diffusion language model. CAUTION: prone to hangin…
CO
North Mini Code
Compact agentic coding MoE (30B total, 3B active, Apache 2.0) f…
Claude Fable 5
HOTMost capable widely released model for the most demanding reaso…
Anthropic: Claude Fable Latest
HOTThis model always redirects to the latest model in the Claude F…
Nex AGI: Nex-N2-Pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI,…
NVIDIA: Nemotron 3.5 Content Safety (free)
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter mu…
NVIDIA: Nemotron 3 Ultra
HOTNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orche…
Llama 4 Maverick 17B 128E Instruct Nvfp4
Meta chat model. https://huggingface.co/api/models/RedHatAI/Lla…
GLM 4.7 FP4
Zai Org chat model.
Qwen3.7 Plus
Multimodal agent model with 1M context, native thinking, and vi…
MiniMax: MiniMax M3
HOTMiniMax-M3 is a multimodal foundation model from MiniMax. It su…
StepFun: Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal M…
Nano Banana Pro
Gemini 3 Pro Image
Nano Banana 2
HOTGemini 3.1 Flash Image.
Claude Opus 4.8
Previous most capable Opus-tier model for complex reasoning and…
Anthropic: Claude Opus 4.8 (Fast)
Fast-mode variant of Opus 4.8 - identical capabilities with hig…
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M conte…
Grok Build 0.1
HOTxAI fast coding model with reasoning, function calling, and str…
CO
Command A Plus
Cohere flagship MoE (218B total, 25B active, Apache 2.0). Agent…
Llama 4 Scout 17B 16E Instruct Fp8 Lora
Meta chat model.
Gemini 3.5 Flash
Gemini 3.5 Flash
Antigravity Agent Preview (2026-05)
Preview release of Antigravity Agent (05-2026)
Gemma 4 31B It Lora
Google chat model. https://huggingface.co/api/models/google/gem…
Gemma 3 27B It Lora
Google chat model.
Perceptron: Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality visio…
inclusionAI: Ring-2.6-1T
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B act…
Mixtral 8x7B Instruct V0.1 FP8 Lora
Mistral AI chat model.
Gemma 3 270M It Lora
Google chat model.
Gemini 3.1 Flash-Lite
Gemini 3.1 Flash Lite
Llama 3.3 70B Instruct FP8 Lora
Meta chat model.
OpenAI: GPT Chat Latest
GPT Chat Latest
ChatGPT Instant
IBM: Granite 4.1 8B
Granite 4.1 8B is a dense, decoder-only 8-billion-parameter lan…
Poolside: Laguna XS.2
Laguna XS.2 is the second-generation model in the XS size class…
Poolside: Laguna M.1
Laguna M.1 is the flagship coding agent model from Poolside, op…
Owl Alpha
Owl Alpha is a high-performance foundation model designed for a…
NVIDIA: Nemotron 3 Nano Omni (free)
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…
Nemotron 3 Nano Omni 30B A3b Reasoning Fp8
Nvidia chat model.
mistral-vibe-cli-with-tools
Official mistral-medium-latest Mistral AI model
mistral-vibe-cli-latest
Official mistral-medium-latest Mistral AI model
mistral-medium-latest
Official mistral-medium-latest Mistral AI model
mistral-medium-3-5
Official mistral-medium-latest Mistral AI model
mistral-medium
Official mistral-medium-latest Mistral AI model
Mistral Medium 3.5
Mistral frontier-class multimodal model with adjustable reasoni…
Mistral Medium (2604)
Official mistral-medium-latest Mistral AI model
magistral-medium-latest
Official mistral-medium-latest Mistral AI model
Qwen3.6 Max Preview
Qwen3.6 Max preview (never GA; superseded by Qwen3.7 Max). 256K…
Qwen3.6 Flash
Fast, cost-effective multimodal model with 1M context, near-fla…
Qwen3.6 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
Qwen3.6 27B
Open 27B dense multimodal with thinking. 256K context.
DeepSeek V4 Pro (0813)
HOTPremium reasoning model with 1M context, released GA by DeepSee…
DeepSeek V4 Flash (0731)
HOTFast general-purpose model with 1M context, re-post-trained by…
inclusionAI: Ling-2.6-1T
Ling-2.6-1T is an instant (instruct) model from inclusionAI and…
GPT-5.5 Pro
Most capable model for complex tasks. Uses more compute for sma…
GPT-5.5
New baseline for complex production workflows. Stronger task ex…
Xiaomi: MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong per…
Xiaomi: MiMo-V2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pr…
Tencent: Hy3 preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from…
Qwen3.6 35B A3b Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3.6…
Pareto Code Router
The Pareto Router maintains a tiered shortlist of strong coding…
OpenAI: GPT-5.4 Image 2
GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-t…
inclusionAI: Ling-2.6-flash
Ling-2.6-flash is an instant (instruct) model from inclusionAI…
Gemma 4 E2B-it
Google chat model. https://huggingface.co/api/models/google/gem…
Deep Research Preview (2026-04)
Preview release (April 21th, 2026) of Deep Research
Deep Research Max Preview (2026-04)
Preview release (April 21st, 2026) of Deep Research Max
Kimi K2.6
Native multimodal flagship (text, image, video inputs) with thi…
Grok 4.3
xAI's latest flagship model with reasoning and a 1M token conte…
Gemma 4 E4B-it
Google chat model. https://huggingface.co/google/gemma-4-E4B-it
Claude Opus 4.7
Previous most capable model for complex reasoning and agentic c…
Anthropic: Claude Opus 4.7 (Fast)
Fast-mode variant of Opus 4.7 - identical capabilities with hig…
Gemini 3.1 Flash TTS Preview
Gemini 3.1 Flash TTS Preview
Ising Calibration 1 35B
NVIDIA quantum-calibration VLM (domain-specific, not for genera…
Gemini Robotics-ER 1.6 Preview
Gemini Robotics-ER 1.6 Preview
Nvidia Nemotron 3 Super 120B A12b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
GLM-5.1
Z.ai flagship (744B MoE, 40B activated). Post-training upgrade…
Qwen3.6 Plus
Previous Plus tier, superseded by Qwen3.7 Plus. 1M context, thi…
Gemma 4 31B IT
Gemma 4 31B IT
Gemma 4 26B A4B IT
HOTGemma 4 26B A4B IT
GLM-5V Turbo
First multimodal GLM-5 model. Vision-based coding agent with im…
Arcee AI: Trinity Large Thinking
Trinity Large Thinking is a powerful open source reasoning mode…
SpaceXAI: Grok 4.20
Grok 4.20 is a reasoning model from SpaceXAI with industry-lead…
Holo3 35B A3b
Hcompany chat model. https://huggingface.co/api/models/Hcompany…
Google: Lyria 3 Pro Preview
Full-length songs are priced at $0.08 per song. Lyria 3 is Goog…
Google: Lyria 3 Clip Preview
30 second duration clips are priced at $0.04 per clip. Lyria 3…
Kwaipilot: KAT-Coder-Pro V2
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKA…
Qwen3 30B A3B Instruct 2507 Lora
Qwen chat model.
DeepSeek V3.1
Deepseek model via OpenAI-Compatible API on AWS Bedrock Mantle
Reka Edge
Reka Edge is an extremely efficient 7B multimodal vision-langua…
Qwen3 8B Lora
Qwen chat model.
MiniMax: MiniMax M2.7 (free)
MiniMax-M2.7 is a next-generation large language model designed…
OpenAI GPT Mini Latest
This model always redirects to the latest model in the OpenAI G…
GPT-5.4 Nano
Cheapest GPT-5.4-class model for simple high-volume tasks like…
GPT-5.4 Mini
Strongest mini model for coding, computer use, and subagents. G…
Qwen3.5 122B A10b Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3.5…
mistral-vibe-cli-fast
Mistral Small 4.
mistral-small-latest
Mistral Small 4.
Mistral Small (2603)
Mistral Small 4.
magistral-small-latest
Mistral Small 4.
Leanstral (2603)
A mid & post-trained version of mistral small 4 for Lean
GLM-5 Turbo
Speed-optimized GLM-5 variant for agent workflows. Enhanced too…
Deepseek OCR 2
Deepseek chat model. https://huggingface.co/api/models/deepseek…
NVIDIA: Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE mod…
Nvidia Nemotron 3 Super 120B A12b Fp8
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
Qwen: Qwen3.5-9B
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 fa…
ByteDance Seed: Seed-2.0-Lite
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhor…
SpaceXAI: Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 desi…
Grok 4.20 Reasoning
xAI flagship reasoning model with a 1M token context window. De…
Grok 4.20 Multi-Agent
Multi-agent model that runs specialized agents in parallel for…
Grok 4.20
xAI flagship model with a 1M token context window. Non-reasonin…
Qwen3.5 9B Fp8
Qwen chat model. https://huggingface.co/api/models/togethercomp…
GPT-5.4 Pro
Most capable model for complex tasks. Uses more compute for sma…
GPT-5.4
Most capable and efficient frontier model for professional work…
Inception: Mercury 2
Mercury 2 is an extremely fast reasoning LLM, and the first rea…
OpenAI: GPT-5.3 Chat
GPT-5.3 Chat is an update to ChatGPT's most-used model that mak…
GPT-5.3 Instant
deprecatedGPT-5.3 Instant model, previously powering ChatGPT.
Glm 4.7 Fp8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
Gemini 3.1 Flash-Lite Preview
Gemini 3.1 Flash Lite Preview
Nano Banana 2 Preview
Gemini 3.1 Flash Image Preview.
ByteDance Seed: Seed-2.0-Mini
Seed-2.0-mini targets latency-sensitive, high-concurrency, and…
Qwen3.5 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
Qwen3.5 27B
Open 27B dense multimodal with thinking. 256K context.
Qwen3.5 122B-A10B
Open 122B-A10B MoE multimodal with thinking. 256K context.
Qwen: Qwen3.5-Flash
The Qwen3.5 native vision-language Flash models are built on a…
LiquidAI: LFM2-24B-A2B
LFM2-24B-A2B is the largest model in the LFM2 family of hybrid…
GPT Audio 1.5
Best voice model for audio in, audio out with Chat Completions.…
Qwen3.5 Flash
Former Flash tier, superseded by Qwen3.6/3.7 Flash. 1M context,…
AionLabs: Aion-2.0
Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive…
Google Gemini Pro Latest
HOTThis model always redirects to the latest model in the Google G…
Gemini 3.1 Pro Preview (Custom Tools)
Gemini 3.1 Pro Preview optimized for custom tool usage
Gemini 3.1 Pro Preview
HOTGemini 3.1 Pro Preview
CO
Tiny Aya Water
Tiny Aya (3.35B) region-specialized for European and Asia-Pacif…
CO
Tiny Aya Global
Tiny multilingual research model (3.35B), best balance across 7…
CO
Tiny Aya Fire
Tiny Aya (3.35B) region-specialized for South Asian languages.…
CO
Tiny Aya Earth
Tiny Aya (3.35B) region-specialized for West Asian and African…
Claude Sonnet 4.6
Best combination of speed and intelligence for everyday tasks
Qwen3.5 397B-A17B
Open 397B-A17B MoE multimodal with thinking. 256K context.
Qwen: Qwen3.5 Plus 2026-02-15
The Qwen3.5 native vision-language series Plus models are built…
Qwen3.5 Plus
Former Plus tier, superseded by Qwen3.6/3.7 Plus. 1M context, t…
MiniMax M2.5 FP4
MiniMaxAI chat model.
MiniMax: MiniMax M2.5
MiniMax-M2.5 is a SOTA large language model designed for real-w…
GLM 5 Fp4
Zai Org chat model. https://huggingface.co/api/models/togetherc…
GLM-5
Z.ai flagship foundation model (744B MoE, 40B activated). Desig…
Qwen: Qwen3 Max Thinking
Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3…
GPT-5.3 Codex
Most capable agentic coding model. Combines frontier coding per…
Claude Opus 4.6
HOTPrevious most intelligent model for complex agents and coding,…
Qwen3 Coder Next
Budget agentic coder on the Qwen3-Next architecture. 256K conte…
GLM-OCR (Vision, OCR)
Specialized OCR model for text extraction from images and docum…
Free Models Router
The simplest way to get free inference. openrouter/free is a ro…
StepFun: Step 3.5 Flash
Step 3.5 Flash is StepFun's most capable open-source foundation…
Upstage: Solar Pro 3
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) lang…
Kimi K2.5
Supports vision (images/videos), thinking mode, and Agent tasks…
MiniMax: MiniMax M2-her
MiniMax M2-her is a dialogue-first large language model built f…
Writer: Palmyra X5
Palmyra X5 is Writer's most advanced model, purpose-built for b…
LiquidAI: LFM2.5-1.2B-Thinking (free)
LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model o…
LiquidAI: LFM2.5-1.2B-Instruct (free)
LFM2.5-1.2B-Instruct is a compact, high-performance instruction…
GLM-4.7 FlashX
Fast GLM-4.7 variant with priority routing and higher concurren…
GLM-4.7 Flash (Free)
Free GLM-4.7 variant. Same model as FlashX but with limited con…
Z.ai GLM 4.7 (Preview)
Z.ai GLM 4.7 (355B) on Cerebras (~1,000 tok/s). Strong agentic…
MiniMax: MiniMax M2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language…
ByteDance Seed: Seed 1.6 Flash
Seed 1.6 Flash is an ultra-fast multimodal deep thinking model…
ByteDance Seed: Seed 1.6
Seed 1.6 is a general-purpose model released by the ByteDance S…
GLM-4.7
Latest-gen GLM model with 200K context. Thinking mode activated…
Gemini 3 Flash Preview
Gemini 3 Flash Preview
Nvidia Nemotron 3 Nano 30B A3b Bf16
Nvidia chat model. https://huggingface.co/api/models/nvidia/NVI…
NVIDIA: Nemotron 3 Nano 30B A3B
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model wi…
Riva Translate 4B v1.1
Translation-specialized model, 8K context.
OpenAI: GPT-5.2 Chat
GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of t…
GPT-5.2 Instant
deprecatedGPT-5.2 Instant model, previously powering ChatGPT.
GPT-5.2 Codex
deprecatedGPT-5.2 optimized for long-horizon, agentic coding tasks in Cod…
Deep Research Pro Preview
Preview release (December 12th, 2025) of Deep Research Pro
AutoGLM Phone
Mobile phone automation agent. Understands phone screens via mu…
GPT-5.2 Pro
Smartest and most trustworthy option for difficult questions. U…
GPT-5.2
Most capable model for professional work and long-running agent…
Mistral: Devstral 2 2512
Devstral 2 is a state-of-the-art open-source model by Mistral A…
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
Devstral 2 (latest)
Official devstral-2512 Mistral AI model
Relace: Relace Search
The relace-search model uses 4-12 `view_file` and `grep` tools…
GLM-4.6 V FlashX
Fast vision GLM-4.6 with priority routing and higher concurrenc…
GLM-4.6 V Flash (Free)
Free vision GLM-4.6. Same model as FlashX but with limited conc…
GLM-4.6 V
Vision-enabled GLM-4.6 model. Supports image/video/file inputs,…
EssentialAI Rnj-1 Instruct
Essential AI chat model. https://huggingface.co/api/models/toge…
Body Builder (beta)
Transform your natural language requests into structured OpenRo…
mistral-large-latest
Official mistral-large-2512 Mistral AI model
ministral-8b-latest
Ministral 3 (a.k.a. Tinystral) 8B Instruct.
ministral-3b-latest
Ministral 3 (a.k.a. Tinystral) 3B Instruct.
ministral-14b-latest
Ministral 3 (a.k.a. Tinystral) 14B Instruct.
Ministral 8b (2512)
Ministral 3 (a.k.a. Tinystral) 8B Instruct.
Ministral 3b (2512)
Ministral 3 (a.k.a. Tinystral) 3B Instruct.
Ministral 3 14B Instruct 2512
Mistralai chat model. https://huggingface.co/api/models/mistral…
Ministral 14b (2512)
Ministral 3 (a.k.a. Tinystral) 14B Instruct.
Amazon: Nova 2 Lite
Nova 2 Lite is a fast, cost-effective reasoning model for every…
Mistral Large (2512)
Official mistral-large-2512 Mistral AI model
DeepSeek: DeepSeek V3.2
DeepSeek-V3.2 is a large language model designed to harmonize h…
Arcee AI: Trinity Mini
Trinity Mini is a 26B-parameter (3B active) sparse mixture-of-e…
Claude Opus 4.5
Previous most intelligent model with advanced reasoning for com…
AllenAI: Olmo 3 32B Think
Olmo 3 32B Think is a large-scale, 32-billion-parameter model p…
Nano Banana Pro Preview
Gemini 3 Pro Image Preview
Nano Banana Pro
Gemini 3 Pro Image Preview
GPT-5.1 Codex Max
deprecatedOur most intelligent coding model optimized for long-horizon, a…
GPT-5.1 Codex Mini
deprecatedSmaller, faster version of GPT-5.1 Codex for efficient coding t…
GPT-5.1 Codex
deprecatedA version of GPT-5.1 optimized for agentic coding tasks in Code…
GPT-5.1
The best model for coding and agentic tasks with configurable r…
Deep Cogito: Cogito v2.1 671B
Cogito v2.1 671B MoE represents one of the strongest open model…
OpenAI: GPT-5.1 Chat
GPT-5.1 Chat (AKA Instant is the fast, lightweight member of th…
GPT-5.1 Instant
deprecatedGPT-5.1 Instant with adaptive reasoning. More conversational wi…
Qwen3-VL-235B-A22B-Instruct-FP8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-V…
MoonshotAI: Kimi K2 Thinking
Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning…
Amazon: Nova Premier 1.0
Amazon Nova Premier is the most capable of Amazon’s multimodal…
Perplexity: Sonar Pro Search
Exclusively available on the OpenRouter API, Sonar Pro's new Pr…
Mistral: Voxtral Small 24B 2507
Voxtral Small is an enhancement of Mistral Small 3, incorporati…
OpenAI: gpt-oss-safeguard-20b
gpt-oss-safeguard-20b is a safety reasoning model from OpenAI b…
NVIDIA: Nemotron Nano 12B 2 VL (free)
NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multim…
Nemotron Safety Guard 8B v3
NVIDIA content-safety classifier (not a chat model).
Medgemma 27B Text It
Google chat model. https://huggingface.co/api/models/google/med…
Qwen: Qwen3 VL 32B Instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-langua…
MiniMax: MiniMax M2
MiniMax-M2 is a compact, high-efficiency large language model o…
IBM: Granite 4.0 Micro
Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family…
Microsoft: Phi 4 Mini Instruct
Phi-4-mini-instruct is a lightweight open model built upon synt…
Qwen3 VL Flash
Budget VL tier below Qwen3 VL Plus. 256K context, thinking, vis…
Claude Haiku 4.5
Fastest model with exceptional speed and performance
Anthropic Claude Haiku Latest
This model always redirects to the latest model in the Anthropi…
Qwen: Qwen3 VL 8B Thinking
Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the…
Qwen: Qwen3 VL 8B Instruct
Qwen3-VL-8B-Instruct is a multimodal vision-language model from…
GPT-5 Search API
Updated web search model in Chat Completions API. 60% cheaper w…
NVIDIA: Llama 3.3 Nemotron Super 49B V1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-c…
Gemini 2.5 Computer Use Preview 10-2025
Gemini 2.5 Computer Use Preview 10-2025
Qwen: Qwen3 VL 30B A3B Thinking
Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies st…
Qwen: Qwen3 VL 30B A3B Instruct
Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies st…
GPT-5 Pro
Version of GPT-5 that uses more compute to produce smarter and…
GPT Audio Mini
deprecatedCost-efficient audio model. Accepts audio inputs and outputs vi…
Nano Banana
Gemini 2.5 Flash Preview Image
GLM-4.6
GLM-4.6 model with 200K context, 128K output. Hybrid thinking:…
Gemma 3 270M It
Google chat model. https://huggingface.co/api/models/google/gem…
DeepSeek: DeepSeek V3.2 Exp
DeepSeek-V3.2-Exp is an experimental large language model relea…
Claude Sonnet 4.5
HOTPrevious best combination of speed and intelligence for complex…
TheDrummer: Cydonia 24B V4.1
Uncensored and creative writing model based on Mistral Small 3.…
Relace: Relace Apply 3
Relace Apply 3 is a specialized code-patching LLM that merges A…
Google: Gemini 2.5 Flash Lite Preview 09-2025
Gemini 2.5 Flash-Lite is a lightweight reasoning model in the G…
Qwen3 Next 80B A3b Instruct Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-N…
Qwen3 Vl 235b A22b Thinking
Alibaba model (not yet curated).
Qwen3 Vl 235b A22b Instruct
Alibaba model (not yet curated).
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context…
Qwen3 Coder Plus
Agentic coding model with very long context. Tiered pricing by…
DeepSeek: DeepSeek V3.1 Terminus
DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maint…
Magistral Small (2509)
Our efficient reasoning model released September 2025.
Magistral Medium (2509)
Our frontier-class reasoning model release candidate September…
GPT-5 Codex
deprecatedA version of GPT-5 optimized for agentic coding in Codex.
Qwen3 Next 80b A3b Thinking
Alibaba model (not yet curated).
Qwen3 Next 80b A3b Instruct
Alibaba model (not yet curated).
NVIDIA: Nemotron Nano 9B V2 (free)
NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trai…
MoonshotAI: Kimi K2 0905
Kimi K2 0905 is the September update of Kimi K2 0711. It is a l…
[Groq] Compound Mini (Agentic System)
Lighter Groq agentic AI, same built-in tools but a single tool…
[Groq] Compound (Agentic System)
Groq agentic AI with web search, visit website, code execution,…
Qwen3 30b A3b Thinking 2507
Alibaba model (not yet curated).
GPT Audio
First generally available audio model. Accepts audio inputs and…
Nous: Hermes 4 70B
Hermes 4 70B is a hybrid reasoning model from Nous Research, bu…
Nous: Hermes 4 405B
Hermes 4 is a large-scale reasoning model built on Meta-Llama-3…
DeepSeek: DeepSeek V3.1
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameter…
Nemotron Nano 9B v2
Small hybrid Mamba-Transformer for fast, cheap reasoning and to…
Mistral: Mistral Medium 3.1
Mistral Medium 3.1 is an updated version of Mistral Medium 3, w…
Mistral Medium (2508)
Update on Mistral Medium 3 with improved capabilities.
GLM-4.5 V
Vision-enabled GLM-4.5 model. 64K context, 16K output, interlea…
Qwen3 4B Instruct 2507
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-4…
AI21: Jamba Large 1.7
Jamba Large 1.7 is the latest model in the Jamba open family, o…
OpenAI: GPT-5 Chat
GPT-5 Chat is designed for advanced, natural, multimodal, and c…
GPT-5 Nano
Fastest, most cost-efficient version of GPT-5 for summarization…
GPT-5 Mini
A faster, more cost-efficient version of GPT-5 for well-defined…
GPT-5 ChatGPT
deprecatedGPT-5 model used in ChatGPT.
GPT-5
The best model for coding and agentic tasks across domains.
OpenAI: gpt-oss-20b (free)
gpt-oss-20b is an open-weight 21B parameter model released by O…
OpenAI: gpt-oss-120b
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Exper…
Claude Opus 4.1
deprecatedPrevious Opus model. Deprecated June 5, 2026, retiring August 5…
CO
Command A Translate
Specialized machine translation across 23 languages, with tool…
CO
Command A Reasoning
Reasoning-tuned Command A for multi-step agents and hard proble…
Qwen: Qwen3 Coder 30B A3B Instruct
Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Ex…
mistral-code-latest
Our cutting-edge language model for coding released August 2025.
mistral-code-fim-latest
Our cutting-edge language model for coding released August 2025.
Glm 4.5 Air Fp8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
codestral-latest
Our cutting-edge language model for coding released August 2025.
Codestral (2508)
Our cutting-edge language model for coding released August 2025.
Qwen3 30b A3b Instruct 2507
Alibaba model (not yet curated).
Qwen3 235B A22b Instruct 2507 Fp8
Together AI chat model. https://huggingface.co/api/models/Qwen/…
Qwen3 Coder Flash
Former budget coder, superseded by Qwen3 Coder Next (which lack…
GLM-4.5 X
Extended GLM-4.5 model. Interleaved thinking.
GLM-4.5 Flash (Free)
Free GLM-4.5 variant with limited concurrency. Prior-gen, super…
GLM-4.5 AirX
Extended lightweight GLM-4.5 variant. Interleaved thinking.
GLM-4.5 Air
Lightweight GLM-4.5 variant. Interleaved thinking.
GLM-4.5
Prior-gen GLM-4.5 model with 128K context, 96K output. Interlea…
Qwen3 235b A22b Thinking 2507
Alibaba model (not yet curated).
Qwen3 Coder 480B A35B Instruct Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-C…
Qwen: Qwen3 Coder 480B A35B
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) co…
Qwen3 235B A22B Instruct 2507 FP8 Throughput
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-2…
Gemini 2.5 Flash-Lite
Stable version of Gemini 2.5 Flash-Lite, released in July of 20…
ByteDance: UI-TARS 7B
UI-TARS-1.5 is a multimodal vision-language agent optimized for…
Qwen: Qwen3 235B A22B Instruct 2507
Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tu…
voxtral-small-latest
A small audio understanding model released in July 2025
voxtral-mini-latest
A mini audio understanding model released in July 2025
Voxtral Small (2507)
A small audio understanding model released in July 2025
Voxtral Mini (2507)
A mini audio understanding model released in July 2025
Switchpoint Router
Switchpoint AI's router instantly analyzes your request and dir…
MoonshotAI: Kimi K2 0711
Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) lang…
Sarvam M
Sarvamai chat model. https://huggingface.co/api/models/sarvamai…
Meta Llama 3.1 8B Instruct Awq Int4
Meta chat model. https://huggingface.co/api/models/togethercomp…
Venice: Uncensored (free)
Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-…
Tencent: Hunyuan A13B Instruct
Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE)…
Morph: Morph V3 Large
Morph's high-accuracy apply model for complex code edits. ~4,50…
Morph: Morph V3 Fast
Morph's fastest apply model for code edits. ~10,500 tokens/sec…
CO
Command A Vision
Multimodal Command A for charts, graphs, diagrams, OCR, and doc…
Baidu: ERNIE 4.5 VL 424B A47B
ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE)…
Minimax M1 80K
MiniMaxAI chat model. https://huggingface.co/api/models/togethe…
o4 Mini Deep Research
Faster, more affordable deep research model for complex, multi-…
o3 Deep Research
Our most powerful deep research model for complex, multi-step r…
Minimax M1 40K
MiniMaxAI chat model. https://huggingface.co/api/models/togethe…
Mistral: Mistral Small 3.2 24B
Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter…
Mistral Small (2506)
Our latest enterprise-grade small model with the latest version…
MiniMax: MiniMax M1
MiniMax-M1 is a large-scale, open-weight reasoning model design…
Gemini 2.5 Pro
Stable release (June 17th, 2025) of Gemini 2.5 Pro
Gemini 2.5 Flash
Stable version of Gemini 2.5 Flash, our mid-size multimodal mod…
Mistral Nemotron
Mistral model post-trained by NVIDIA.
Magistral Small 2506
Mistralai chat model. https://huggingface.co/api/models/mistral…
Llama 4 Scout (17Bx16E)
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
o3 Pro
Version of o3 with more compute for better responses. Provides…
Gemma 2B It
Google chat model. https://huggingface.co/api/models/google/gem…
Gemma 2 9B It
Google chat model. https://huggingface.co/api/models/google/gem…
Qwen3 1.7B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-1…
Qwen3 0.6B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-0…
Nemotron Nano VL 8B
Small vision-language model, 16K context.
Molmo 7B D 0924
Allenai chat model. https://huggingface.co/api/models/allenai/M…
DeepSeek: R1 0528
May 28th update to the original DeepSeek R1 Performance on par…
Mixtral 8X22b Instruct V0.1
Mistralai chat model. https://huggingface.co/api/models/mistral…
Anthropic: Claude Sonnet 4
Claude Sonnet 4 significantly enhances the capabilities of its…
Anthropic: Claude Opus 4
Claude Opus 4 is benchmarked as the world’s best coding model,…
Devstral Small 2505
Mistralai chat model. https://huggingface.co/api/models/togethe…
Mistral 7B v0.1
Mistralai chat model. https://huggingface.co/api/models/mistral…
Google: Gemma 3n 4B
Gemma 3n E4B-it is optimized for efficient execution on mobile…
Google: Gemini 2.5 Pro Preview 06-05
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed f…
Gemini 2.5 Pro Preview TTS
Gemini 2.5 Pro Preview TTS
Gemini 2.5 Flash Preview TTS
Gemini 2.5 Flash Preview TTS
Deepcoder 14B Preview
Togethercomputer chat model. https://huggingface.co/api/models/…
mistral-medium-3
Official mistral-medium-latest Mistral AI model
Mistral Medium (2505)
Our frontier-class multimodal model released May 2025.
Google: Gemini 2.5 Pro Preview 05-06
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed f…
Arcee AI: Virtuoso Large
Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B…
Arcee AI: Coder Large
Coder‑Large is a 32 B‑parameter offspring of Qwen 2.5‑Instruct…
Meta: Llama Guard 4 12B
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained…
Qwen3 8b
Alibaba model (not yet curated).
Qwen3 32b
Alibaba model (not yet curated).
Qwen3 30b A3b
Alibaba model (not yet curated).
Qwen3 235b A22b
Alibaba model (not yet curated).
Qwen3 14b
Alibaba model (not yet curated).
Arize AI Qwen 2 1.5B Instruct
Togethercomputer chat model. https://huggingface.co/api/models/…
o4 Mini
Latest o4-mini model. Optimized for fast, effective reasoning w…
o3
A well-rounded and powerful model across domains. Sets a new st…
Llama 3.1 405B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
Qwen2.5 7B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 7B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 72B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 3B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 32B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 32B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 14B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 1.5B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 1.5B
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Llama 3.2 1B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
Llama 3.1 70B
Meta chat model. https://huggingface.co/api/models/meta-llama/L…
GPT-4.1 Nano
Fastest, most cost-effective GPT 4.1 model. Delivers exceptiona…
GPT-4.1 Mini
Balanced for intelligence, speed, and cost. Matches or exceeds…
GPT-4.1
Flagship GPT model for complex tasks. Major improvements on cod…
GLM-4 32B (0414) 128K
GLM-4 32B model with 128K context, 16K output.
Qwen2 72B Instruct
Togethercomputer chat model. https://huggingface.co/api/models/…
Cogito V1 Preview Qwen 32B
deepcogito chat model.
Cogito V1 Preview Qwen 14B
deepcogito chat model.
Cogito V1 Preview Llama 8B
deepcogito chat model.
Cogito V1 Preview Llama 70B Turbo
deepcogito chat model.
Cogito V1 Preview Llama 70B
deepcogito chat model.
Meta: Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE)…
Meta: Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimo…
Llama 4 Scout Instruct (17Bx16E)
Meta chat model. https://huggingface.co/meta-llama/Llama-4-Scou…
Gemma 3 1b it
Google chat model.
DeepSeek R1 Distill Qwen 7B
Deepseek chat model. https://huggingface.co/deepseek-ai/DeepSee…
meta-llama/Llama-2-7b-chat-hf
Meta chat model. https://huggingface.co/meta-llama/Llama-2-7b-c…
DeepSeek: DeepSeek V3 0324
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the…
o1 Pro
A version of o1 with more compute for better responses. Provide…
Llama 3.3 Nemotron Super 49B v1
Superseded by v1.5.
Mistral: Mistral Small 3.1 24B
Mistral Small 3.1 24B Instruct is an upgraded variant of Mistra…
nim/nv-mistralai/mistral-nemo-12b-instruct
NVIDIA chat model.
nim/mistralai/mixtral-8x7b-instruct-v01
mistralai chat model.
Google: Gemma 3 4B
Gemma 3 introduces multimodality, supporting vision-language in…
Google: Gemma 3 12B
Gemma 3 introduces multimodality, supporting vision-language in…
CO
Command A
Cohere's efficient 111B enterprise model for agents, tool use,…
Cohere: Command A
Command A is an open-weights 111B parameter model with a 256k c…
Reka Flash 3
Reka Flash 3 is a general-purpose, instruction-tuned large lang…
nim/nvidia/llama-3.1-nemotron-70b-instruct
NVIDIA chat model.
Google: Gemma 3 27B
Gemma 3 introduces multimodality, supporting vision-language in…
GPT-4o Search Preview
deprecatedGPT-4o model optimized for web search capabilities. Alias still…
GPT-4o Mini Search Preview
deprecatedGPT-4o Mini model optimized for web search capabilities. Alias…
TheDrummer: Skyfall 36B V2
Skyfall 36B v2 is an enhanced iteration of Mistral Small 2501,…
nim/mistralai/mixtral-8x22b-instruct-v01
Mistral chat model.
Meta Llama 3.1 8B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Qwen QwQ-32B
Qwen chat model. https://huggingface.co/Qwen/QwQ-32B
CO
Aya Vision 32B
Open-weights multilingual vision research model (23 languages)…
Sonar Reasoning Pro
Premier reasoning model with enhanced multi-step Chain of Thoug…
Mistral: Saba
Mistral Saba is a 24B-parameter language model specifically des…
Sonar Deep Research
Expert-level research model for exhaustive searches and compreh…
Gemini 2.0 Flash 001
Stable version of Gemini 2.0 Flash, our fast and versatile mult…
AionLabs: Aion-RP 1.0 (8B)
Aion-RP-Llama-3.1-8B ranks the highest in the character evaluat…
AionLabs: Aion-1.0-Mini
Aion-1.0-Mini 32B parameter model is a distilled version of the…
AionLabs: Aion-1.0
Aion-1.0 is a multi-model system designed for high performance…
Qwen: Qwen2.5 VL 72B Instruct
Qwen2.5-VL is proficient in recognizing common objects such as…
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M cont…
CO
Command R7B Arabic
Command R7B tuned for Modern Standard Arabic and English enterp…
o3 Mini
Latest o3-mini model snapshot. High intelligence at the same co…
Mistral: Mistral Small 3
Mistral Small 3 is a 24B-parameter language model optimized for…
DeepSeek R1 Distill Qwen 14B
DeepSeek chat model. https://huggingface.co/api/models/deepseek…
DeepSeek R1 Distill Qwen 1.5B
DeepSeek chat model. https://huggingface.co/deepseek-ai/DeepSee…
DeepSeek: R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B is a distilled large language mod…
Sonar Pro
Advanced search model for complex queries and deep content unde…
Sonar
Lightweight, cost-effective search model for quick, grounded an…
DeepSeek: R1
DeepSeek R1 is here: Performance on par with OpenAI o1, but ope…
NemoGuard 8B Topic Control
NVIDIA topic-control guardrail (not a chat model).
NemoGuard 8B Content Safety
NVIDIA content-safety classifier (not a chat model).
V1 8K Vision (Preview)
Legacy vision model with 8K context. Sunset on 2026-08-31 - use…
V1 32K Vision (Preview)
Legacy vision model with 32K context. Sunset on 2026-08-31 - us…
V1 128K Vision (Preview)
Legacy vision model with 128K context. Sunset on 2026-08-31 - u…
MiniMax: MiniMax-01
MiniMax-01 is a combines MiniMax-Text-01 for text generation an…
Microsoft: Phi 4
Microsoft Research Phi-4 is designed to perform well in complex…
Qwen2-VL (72B) Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2-VL-72B-Instr…
Sao10K: Llama 3.1 70B Hanami x1
This is Sao10K's experiment over Euryale v2.2.
DeepSeek: DeepSeek V3
DeepSeek-V3 is the latest model from the DeepSeek team, buildin…
Sao10K: Llama 3.3 Euryale 70B
Euryale L3.3 70B is a model focused on creative roleplay from S…
o1
Previous full o-series reasoning model.
CO
Cohere: Command R7B (12-2024)
Command R7B (12-2024) is a small, fast update of the Command R+…
Qwen 2.5 14B Instruct
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen2.5…
Qwen2.5 72B Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-72B-Instru…
Meta: Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a…
Meta Llama 3.3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Llama-3.3-70…
Meta Llama 3.1 405B Instruct
Meta chat model. https://huggingface.co/meta-llama/Llama-3.1-40…
[Meta] Llama 3.3 · 70B Versatile
Meta Llama 3.3 (70B params) with GQA. Strong reasoning, coding,…
Amazon: Nova Pro 1.0
Amazon Nova Pro 1.0 is a capable multimodal model from Amazon f…
Amazon: Nova Micro 1.0
Amazon Nova Micro 1.0 is a text-only model that delivers the lo…
Amazon: Nova Lite 1.0
Amazon Nova Lite 1.0 is a very low-cost multimodal model from A…
Mistral Large 2407
This is Mistral AI's flagship model, Mistral Large 2 (version m…
Qwen 2.5 Coder 32B Instruct
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-Coder-32B-…
Qwen2.5 Coder 32B Instruct
Qwen2.5-Coder is the latest series of Code-Specific Qwen large…
Llama 3.1 Nemotron 70B Instruct HF
nvidia chat model. https://huggingface.co/nvidia/Llama-3.1-Nemo…
TheDrummer: UnslopNemo 12B
UnslopNemo v4.1 is the latest addition from the creator of Roci…
Magnum v4 72B
This is a series of models designed to replicate the prose qual…
Mistral: Ministral 8B
Ministral 8B is an 8B parameter model featuring a unique interl…
Qwen: Qwen2.5 7B Instruct
Qwen2.5 7B is the latest series of Qwen large language models.…
Qwen2.5 7B Instruct Turbo
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-7B-Instruct
Qwen2.5 72B Instruct Turbo
Qwen chat model. https://huggingface.co/Qwen/Qwen2.5-72B-Instru…
Inflection: Inflection 3 Productivity
Inflection 3 Productivity is optimized for following instructio…
Inflection: Inflection 3 Pi
Inflection 3 Pi powers Inflection's Pi chatbot, including backs…
CO
Aya Expanse 32B
Open-weights multilingual research model covering 23 languages.…
TheDrummer: Rocinante 12B
Rocinante 12B is designed for engaging storytelling and rich pr…
Meta: Llama 3.2 3B Instruct
Llama 3.2 3B is a 3-billion-parameter multilingual large langua…
Meta: Llama 3.2 1B Instruct
Llama 3.2 1B is a 1-billion-parameter language model focused on…
Meta: Llama 3.2 11B Vision Instruct
Llama 3.2 11B Vision is a multimodal model with 11 billion para…
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (…
Qwen2.5 72B Instruct
Qwen2.5 72B is the latest series of Qwen large language models.…
Nemotron Mini 4B
Tiny legacy Nemotron, 4K context.
CO
Cohere: Command R+ (08-2024)
command-r-plus-08-2024 is an update of the Command R+ with roug…
CO
Cohere: Command R (08-2024)
command-r-08-2024 is an update of the Command R with improved p…
Sao10K: Llama 3.1 Euryale 70B v2.2
Euryale L3.1 70B v2.2 is a model focused on creative roleplay f…
Nous: Hermes 3 70B Instruct
Hermes 3 is a generalist language model with many improvements…
Nous: Hermes 3 405B Instruct (free)
Hermes 3 is a generalist language model with many improvements…
Sao10K: Llama 3 8B Lunaris
Lunaris 8B is a versatile generalist and roleplaying model base…
Meta: Llama 3.1 8B Instruct
Meta's latest class of model (Llama 3.1) launched with a variet…
Meta: Llama 3.1 70B Instruct
Meta's latest class of model (Llama 3.1) launched with a variet…
[Meta] Llama 3.1 · 8B Instant
Meta Llama 3.1 (8B params). Fast, cost-effective for high-volum…
Meta Llama 3.1 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Mistral: Mistral Nemo
A 12B parameter model with a 128k token context length built by…
open-mistral-nemo-2407
Our best multilingual open source model released July 2024.
open-mistral-nemo
Our best multilingual open source model released July 2024.
GPT-4o mini
Affordable model for fast, lightweight tasks. GPT-4o Mini is ch…
Google: Gemma 2 27B
Gemma 2 27B by Google is an open model built from the same rese…
Qwen 2 Instruct (1.5B)
Qwen chat model. https://huggingface.co/Qwen/Qwen2-72B-Instruct
Mistral (7B) Instruct v0.3
mistralai chat model. https://huggingface.co/api/models/mistral…
GPT-4o
deprecatedOriginal gpt-4o snapshot from May 13, 2024.
Meta: Llama 3 8B Instruct
Meta's latest class of model (Llama 3) launched with a variety…
Meta Llama 3 8B Instruct Reference
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Meta Llama 3 8B Instruct
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Mistral: Mixtral 8x22B Instruct
Mistral's official instruct fine-tuned version of Mixtral 8x22B…
WizardLM-2 8x22B
WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model.…
GPT-4 Turbo
GPT-4 Turbo with Vision model. Vision requests can now use JSON…
Anthropic: Claude 3 Haiku
Claude 3 Haiku is Anthropic's fastest and most compact model fo…
Mistral Large
This is Mistral AI's flagship model, Mistral Large 2 (version `…
Deepseek Coder 33B Instruct
Deepseek chat model. https://huggingface.co/api/models/deepseek…
V1 8K
Legacy V1 model with 8K context. Sunset on 2026-08-31 - use Kim…
V1 32K
Legacy V1 model with 32K context. Sunset on 2026-08-31 - use Ki…
V1 128K
Legacy V1 model with 128K context. Sunset on 2026-08-31 - use K…
OpenAI: GPT-4 Turbo Preview
The preview GPT-4 model with improved instruction following, JS…
OpenAI: GPT-3.5 Turbo (older v0613)
GPT-3.5 Turbo is OpenAI's fastest model. It can understand and…
3.5-Turbo
deprecatedThe latest GPT-3.5 Turbo model with higher accuracy at respondi…
Nous Hermes 2 Mixtral 8X7B Dpo
Nousresearch chat model. https://huggingface.co/api/models/Nous…
Mixtral-8x7B Instruct v0.1
mistralai chat model. https://huggingface.co/mistralai/Mixtral-…
Auto Router
The Auto Router automatically selects the best model for your p…
3.5-Turbo
deprecatedGPT-3.5 Turbo model with improved instruction following, JSON m…
OpenAI: GPT-3.5 Turbo Instruct
This model is a variant of GPT-3.5 Turbo tuned for instructiona…
Mistral (7B) Instruct v0.1
mistralai chat model. https://huggingface.co/api/models/mistral…
OpenAI: GPT-3.5 Turbo 16k
This model offers four times the context length of gpt-3.5-turb…
Mancer: Weaver (alpha)
An attempt to recreate Claude-style verbosity, but don't expect…
ReMM SLERP 13B
A recreation trial of the original MythoMax-L2-B13 but with upd…
MythoMax 13B
One of the highest performing and most popular fine-tunes of Ll…
GPT-4
Snapshot of gpt-4 from June 13th 2023 with improved function ca…
GPT-4
deprecatedSnapshot of gpt-4 from June 13th 2023 with improved function ca…
3.5-Turbo
The latest GPT-3.5 Turbo model with higher accuracy at respondi…
? [gpt live transcribe]
Unknown, please let us know the ID. Assuming a context window o…
? [gpt transcribe]
Unknown, please let us know the ID. Assuming a context window o…
? [ra gpt 5.6 sol]
Unknown, please let us know the ID. Assuming a context window o…
AI21 Labs Jamba 1.5 Large
AI21 Labs model via Unsupported API (Bedrock Foundation Model)
AI21 Labs Jamba 1.5 Mini
AI21 Labs model via Unsupported API (Bedrock Foundation Model)
Amazon Nova 2 Lite
Amazon model via Converse API (Bedrock Inference Profile)
Amazon Nova Lite
Amazon model via Converse API (Bedrock Inference Profile)
Amazon Nova Micro
Amazon model via Converse API (Bedrock Inference Profile)
Amazon Nova Premier
Amazon model via Converse API (Bedrock Inference Profile)
Amazon Nova Pro
Amazon model via Converse API (Bedrock Foundation Model)
Anthropic Claude 3 Sonnet
Anthropic model (Bedrock Inference Profile)
Anthropic Honey
Anthropic model via OpenAI-Compatible API on AWS Bedrock Mantle
Cohere Command R
Cohere model via Unsupported API (Bedrock Foundation Model)
Cohere Command R+
Cohere model via Unsupported API (Bedrock Foundation Model)
Cohere Embed v4
Cohere model via Converse API (Bedrock Inference Profile)
DeepSeek-R1
Deepseek model via Converse API (Bedrock Inference Profile)
FLUX.2 Klein 4B
Model served via Modular.
Gemma 4 12B It
Google chat model. https://huggingface.co/google/gemma-4-12B-it
GLM 5.2 FP8
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
GLM 5.2 FP8 Lora
Zai Org chat model. https://huggingface.co/api/models/zai-org/G…
GLM-5.2 Fast (Alibaba)
Zhipu GLM-5.2 fast-serving tier via Alibaba Model Studio (previ…
Google Gemma 4 26b A4b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
Google Gemma 4 31b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
Google Gemma 4 E2b
Google model via OpenAI-Compatible API on AWS Bedrock Mantle
Kimi K2.5 Fp4
Togethercomputer chat model. https://huggingface.co/api/models/…
LFM2-24B-A2B
Togethercomputer chat model.
LFM2.5-8B-A1B
LiquidAI chat model. https://huggingface.co/api/models/LiquidAI…
Meta Llama 3 70B Instruct
Meta model via Converse API (Bedrock Foundation Model)
Meta Llama 3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Meta Llama 3 8B Instruct
Meta model via Converse API (Bedrock Foundation Model)
Meta Llama 3 8B Instruct Lite
Meta chat model. https://huggingface.co/meta-llama/Meta-Llama-3…
Meta Llama 3.1 70B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.1 8B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.2 11B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.2 1B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.2 3B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.2 90B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 3.3 70B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 4 Maverick 17B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Meta Llama 4 Scout 17B Instruct
Meta model via Converse API (Bedrock Inference Profile)
Mistral AI Devstral 2 123B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral AI Ministral 14B 3.0
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral AI Ministral 3 8B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral AI Ministral 3B
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral AI Mistral 7B Instruct
Mistral AI model via Converse API (Bedrock Foundation Model)
Mistral AI Mistral Large (24.02)
Mistral AI model via Converse API (Bedrock Foundation Model)
Mistral AI Mistral Large 3
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral AI Mistral Small (24.02)
Mistral AI model via Converse API (Bedrock Foundation Model)
Mistral AI Mixtral 8x7B Instruct
Mistral AI model via Converse API (Bedrock Foundation Model)
Mistral AI Voxtral Mini 3B 2507
Mistral AI model via OpenAI-Compatible API (Bedrock Foundation…
Mistral Pixtral Large 25.02
Mistral model via Converse API (Bedrock Inference Profile)
mistral-tiny-latest
Our best multilingual open source model released July 2024.
NVIDIA Nemotron 3 Super 120B A12B
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
NVIDIA Nemotron Nano 12B v2 VL BF16
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
NVIDIA Nemotron Nano 3 30B
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
NVIDIA Nemotron Nano 9B v2
NVIDIA model via OpenAI-Compatible API (Bedrock Foundation Mode…
Open Mistral Nemo
Our best multilingual open source model released July 2024.
OpenAI GPT OSS Safeguard 120B
OpenAI model via OpenAI-Compatible API (Bedrock Foundation Mode…
OpenAI gpt-oss-120b
OpenAI model via Converse API (Bedrock Foundation Model)
OpenAI gpt-oss-20b
OpenAI model via Converse API (Bedrock Foundation Model)
CO
parse v5.0
New Cohere Model
Qvq Max
Alibaba model (not yet curated).
Qwen Coder Plus
Alibaba model (not yet curated).
Qwen Flash
Fast and very low cost with hybrid thinking. 1M context.
Qwen Max
Best quality of the stable commercial line. 32K context.
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M cont…
Qwen Turbo
Fastest and cheapest for simple tasks. 1M context.
Qwen Vl Max
Alibaba model (not yet curated).
Qwen Vl Plus
Alibaba model (not yet curated).
Qwen3 235b A22b Instruct 2507
Alibaba model (not yet curated).
Qwen3 Coder 480b A35b Instruct
Alibaba model (not yet curated).
Qwen3 Coder Next Fp8
Qwen chat model. https://huggingface.co/api/models/Qwen/Qwen3-C…
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context…
Qwen3 Next 80B A3B
Qwen model via Converse API (Bedrock Foundation Model)
Qwen3 VL 235B A22B
Qwen model via Converse API (Bedrock Foundation Model)
Qwen3 VL Plus
Current vision-language model with strong visual reasoning and…
Qwen3-Coder-30B-A3B-Instruct
Qwen model via Converse API (Bedrock Foundation Model)
Qwen3.5 35B A3B Lora
Qwen chat model.
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M conte…
Qwq Plus
Alibaba model (not yet curated).
TwelveLabs Pegasus v1.2
Twelvelabs model via Converse API (Bedrock Inference Profile)
Twelvelabs TwelveLabs Marengo Embed 3.0
Twelvelabs model via Converse API (Bedrock Inference Profile)
Twelvelabs TwelveLabs Marengo Embed v2.7
Twelvelabs model via Converse API (Bedrock Inference Profile)
Writer Palmyra Vision 7B
Writer model via OpenAI-Compatible API (Bedrock Foundation Mode…
Writer Palmyra X4
Writer model via Converse API (Bedrock Inference Profile)
FAQ
Big-AGI connects to 800+ models across 30+ providers, including Anthropic, OpenAI, Google Gemini, SpaceXAI (formerly xAI), DeepSeek, Mistral, Perplexity, Groq, NVIDIA, and OpenRouter. New models are usually added the day they ship. The index above lists the models currently tracked, with capabilities, context windows, and pricing.
You bring your own provider API keys and pay standard provider rates. Big-AGI adds no markup on model usage and never limits your access. Prices above are in USD per 1M tokens, split into input and output. They cover text tokens only and do not include image, audio, or other multimodal input and output, or provider fees for requests, tools, caching, and other services. Prices can change and may be out of date, so always check the provider for current rates. The optional Pro plan ($10.99/mo) covers cross-device sync and does not restrict model usage.
The context window is how much text a model can hold at once, measured in tokens and shown per model. Sort the Context column to rank every model from largest to smallest. Values change as providers ship new versions, so the table always reflects the current numbers.
Yes. Filter and sort the table by vendor, context, price, or capability to compare specs directly. Inside Big-AGI, Beam runs 2 to 24 models in parallel on the same prompt, then compares or merges their answers, so you judge real output, not just specs.
Keys and chats are stored in your browser, local first. With Direct Connection on, requests go straight from your browser to the provider and never pass through the Big-AGI servers (requires a browser-side key and provider CORS support). The Big-AGI core is open source under the MIT license.
Yes. Big-AGI connects to local runtimes like LocalAI, Ollama, and LM Studio, and you can mix local and cloud models in the same chat. Any OpenAI-compatible endpoint is auto-detected by hostname.
Connect your own keys, run models side by side, then compare and merge the answers. Keys stay in your browser.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego