400+ models on one key - with Anthropic, Google, OpenAI and xAI controls restored to their native ranges, not the router's flattened ones.
OpenRouter is the routing layer: one key, hundreds of models, automatic fallback. It is not where you organize the work. Point Big-AGI at that key and the catalog turns into an expert workspace.
| Capability | Through your OpenRouter key | Inside Big-AGI |
|---|---|---|
| Model catalog | Hundreds, one key | The same catalog, same key |
| Reasoning and parameter control | Router-flattened | Vendor-accurate ranges |
| Merging multiple answers | Basic Fusion | Advanced and custom merges |
| Personas and saved prompts | Not the router's job | Built in, with memory |
| Chat history and attachments | Bring your own | Local-first, yours |
| Request transparency | API-level | AI Inspector on every call |
You do not need credits to begin. On October 4, 2026 OpenRouter listed 17 models at $0, five of them NVIDIA Nemotron builds and two Gemma 4 sizes, and Big-AGI runs them the moment your key is linked. Add a Google key for the Gemini free tier and you have a second vendor at no cost. Free usage runs under OpenRouter's own daily and per-minute limits, so it is built for trying things, not production loads. When you outgrow it, the same key unlocks every paid model at OpenRouter's rates.
Step 5 Preview
NEWStep 5 Preview is StepFun's flagship model for agentic work, built on a sparse Mixture-of-Experts architecture (27B active / 600B total parameters). It perform…
1M
$1.00
$2.70
Oct 2026
Claude Haiku 5.5
NEWClaude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Ha…
1M
$0.10
$0.50
Oct 2026
Claude Haiku Latest
NEWThis model always redirects to the latest model in the Claude Haiku family.
1M
$0.10
$0.50
Oct 2026
Nano Banana 2.1
NEWNano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro. It imp…
66K
$1.50
$7.50
Oct 2026
Mistral Large 4
NEWMistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token…
1M
$0.68
$2.09
Oct 2026
Ling 3.1 Flash
NEWLing 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.
262K
Free
Free
Oct 2026
Apodex 1.1 Mini (free)
NEWApodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, cod…
262K
Free
Free
Oct 2026
Pareto 26.10 Preview
NEWPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of g…
1M
$0.80
$3.20
Oct 2026
GPT-6.1 Sol
NEWGPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer us…
1.1M
$2.00
$10.00
Sep 2026
GPT-6.1 Sol Pro
NEWGPT-6.1 Sol Pro is the same underlying model as GPT-6.1 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
1.1M
$2.00
$10.00
Sep 2026
GPT Sol Latest
NEWThis model always redirects to the latest model in the GPT Sol family.
1.1M
$2.00
$10.00
Sep 2026
Claude Sonnet 5.5
NEWHOTClaude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at b…
1M
$2.00
$10.00
Sep 2026
Claude Sonnet Latest
NEWThis model always redirects to the latest model in the Claude Sonnet family.
1M
$2.00
$10.00
Sep 2026
Jev Router
NEWJev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on Jev, TypeSafe's first System One model, a…
1M
-
-
Sep 2026
Perceptron Mk1.5
NEWPerceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optio…
37K
$0.15
$1.50
Sep 2026
Ember-1
NEWEmber-1 is a specialized reasoning model from Fireworks Research, built on Kimi K3. It is designed to make every token go further: it produces shorter reasonin…
1M
$3.00
$15.00
Sep 2026
Solar Mini 4
NEWSolar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is…
524K
$0.05
$0.20
Sep 2026
Aion 3.5
NEWAion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in w…
262K
$3.00
$6.00
Sep 2026
Aion 3.5 Mini
NEWAion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of A…
262K
$0.70
$1.40
Sep 2026
GLM 5.3 Prime
NEWGLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acc…
1M
$2.80
$8.80
Sep 2026
Qwen3.8 Max Prime
NEWQwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, im…
1M
$4.00
$12.00
Sep 2026
Space Bunny Alpha
deprecatedSpace Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustab…
1M
Free
Free
Sep 2026
Claude Opus 5.5
NEWHOTClaude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly stro…
1M
$4.00
$20.00
Sep 2026
GPT-6 Luna Pro
NEWGPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note…
1.1M
$0.10
$0.50
Sep 2026
GPT-6 Luna
NEWGPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads…
1.1M
$0.10
$0.50
Sep 2026
Claude Opus Latest
NEWThis model always redirects to the latest model in the Claude Opus family.
1M
$4.00
$20.00
Sep 2026
Command A+
NEWCommand A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calli…
192K
$0.30
$1.50
Sep 2026
GPT-6 Sol
NEWGPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is su…
1.1M
$2.00
$10.00
Sep 2026
GPT-6 Sol Pro
NEWGPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:*…
1.1M
$2.00
$10.00
Sep 2026
GPT Luna Latest
NEWThis model always redirects to the latest model in the GPT Luna family.
1.1M
$0.10
$0.50
Sep 2026
MiMo-V2.6-Pro
NEWMiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability fo…
1.1M
$0.44
$0.87
Sep 2026
Grok 4.7
NEWGrok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software en…
500K
$2.00
$6.00
Sep 2026
Qwen3.8 Omni Flash
NEWQwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding.…
1M
$0.15
$0.47
Sep 2026
Grok Latest
NEWThis model always redirects to the latest Grok model from xAI.
500K
$2.00
$6.00
Sep 2026
MiMo-V2.6-Flash
NEWMiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated…
1.1M
$0.14
$0.28
Sep 2026
MiMo-V2.6-Pro-UltraSpeed
NEWMiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it ma…
1M
$4.35
$8.70
Sep 2026
Switchyard
NEWSwitchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter market data…
1M
-
-
Sep 2026
GLM 5.3 FlashX
NEWGLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the sam…
1M
$0.37
$1.25
Sep 2026
Ternary Bonsai 2 27B
NEWBonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding w…
262K
$0.08
$0.50
Sep 2026
Pareto
NEWPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of g…
262K
$2.50
$7.50
Sep 2026
Union Alpha
deprecatedUnion Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of genera…
262K
Free
Free
Sep 2026
DeepSeek Pro Latest
NEWThis model always redirects to the latest model in the DeepSeek Pro family.
1M
$0.22
$0.66
Sep 2026
DeepSeek Flash Latest
NEWThis model always redirects to the latest model in the DeepSeek Flash family.
1M
Free
$0.60
Sep 2026
Schematron V2 Small
NEWSchematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. E…
128K
$0.05
$0.23
Sep 2026
Schematron V2 Turbo
NEWSchematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extract…
128K
$0.03
$0.15
Sep 2026
Fugu Max
NEWFugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a…
1M
$2.00
$6.00
Sep 2026
Fugu Ultra v2
NEWFugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration sy…
1M
$5.00
$30.00
Sep 2026
DeepSeek V4.1 Flash
NEWHOTDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It acti…
1M
$0.30
$1.20
Sep 2026
Ling 3.0 Flash VL
NEWLing 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native…
262K
$0.02
$0.06
Sep 2026
Mercury 2.5
NEWMercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces a…
260K
$0.04
$0.15
Sep 2026
Nex-N2.5-Mini
NEWNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can exp…
262K
$0.03
$0.10
Sep 2026
Nex-N2.5-Pro
NEWNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can exp…
262K
$0.08
$0.25
Sep 2026
Ling 3.0 Flash Sante
NEWLing 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124…
262K
$0.04
$0.12
Sep 2026
GPT-6 Astra
NEWGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work,…
1.1M
$10.00
$50.00
Sep 2026
GPT-6 Astra Pro
NEWGPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
1.1M
$10.00
$50.00
Sep 2026
GPT Astra Latest
NEWThis model always redirects to the latest model in the GPT Astra family.
1.1M
$10.00
$50.00
Sep 2026
Gemini 3.8 Flash
NEWHOTGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reas…
1M
$0.75
$3.75
Sep 2026
Gemini Flash Latest
NEWThis model always redirects to the latest model in the Gemini Flash family.
1M
$0.75
$3.75
Sep 2026
Muse Spark 1.3
NEWMuse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of informati…
1M
$1.25
$4.25
Sep 2026
Muse Spark 1.3 Contributor
NEWMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic,…
1M
$0.10
$0.20
Sep 2026
Qwen3.8 Max (0902)
NEWQwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, ima…
1M
$2.00
$6.00
Sep 2026
Claude Fable 5.1
NEWClaude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: lon…
1M
$10.00
$50.00
Sep 2026
Claude Fable Latest
NEWThis model always redirects to the latest model in the Claude Fable family.
1M
$10.00
$50.00
Sep 2026
Granite 4.2 8B
NEWGranite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi…
131K
$0.06
$0.25
Aug 2026
Mercury 2.5 Preview
deprecatedMercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces a…
260K
$0.20
$0.75
Aug 2026
Hy4 preview
NEWTencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-u…
1M
$0.83
$2.50
Aug 2026
Ling 3.0 Flash Fin
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is…
262K
$0.04
$0.12
Aug 2026
Qwen3.8 Flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase…
1M
$0.15
$0.47
Aug 2026
GLM Flash Latest
This model always redirects to the latest model in the GLM Flash family.
1M
$0.04
$0.50
Aug 2026
GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention ar…
1M
$0.15
$0.50
Aug 2026
DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the b…
1M
$0.22
$0.65
Aug 2026
Hy-MT2-1.8B
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, wit…
8K
$0.04
$0.18
Aug 2026
Hy-MT2-30B-A3B
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs…
8K
$0.07
$0.30
Aug 2026
Ox Alpha
deprecatedOx Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, comple…
1M
Free
Free
Aug 2026
Hy-MT2-7B
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows…
8K
$0.07
$0.30
Aug 2026
GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with…
1M
$0.04
$4.80
Aug 2026
Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and lon…
1M
$0.04
$2.30
Aug 2026
Dots3-Note Preview (free)
Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the D…
512K
Free
Free
Aug 2026
GLM Latest
This model always redirects to the latest GLM model from Z.ai.
1M
$0.04
$4.80
Aug 2026
Gemini 3.7 Flash
HOTGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require…
1M
$0.75
$3.75
Aug 2026
DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
1M
$1.32
$3.96
Aug 2026
Grok 4.6
Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by Grok 4.7.
500K
$2.00
$6.00
Aug 2026
Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out…
1M
$2.00
$6.00
Aug 2026
Seed-2.0-Code
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-ag…
262K
$0.50
$3.00
Aug 2026
Seed 2.1 Turbo
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step…
262K
$0.50
$2.50
Aug 2026
Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput age…
262K
$0.07
$0.20
Aug 2026
LFM2.5-2.6B (free)
LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises ag…
66K
Free
Free
Aug 2026
Namazu
deprecatedSakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts…
262K
$0.95
$4.00
Aug 2026
Solar Pro 4
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with s…
524K
$0.09
$0.36
Aug 2026
Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on co…
131K
$0.30
$1.20
Aug 2026
Ling 3.0 Tiny (free)
deprecatedLing 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction…
262K
Free
Free
Aug 2026
Muse Spark 1.2
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, and PDF documents, returns text, and offers…
1M
$1.25
$4.25
Aug 2026
Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully chea…
1M
$0.10
$0.20
Aug 2026
Sakana Namazu
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts…
262K
$0.95
$4.00
Aug 2026
DeepSeek V4 Flash Latest
This model always redirects to the latest model in the DeepSeek V4 Flash family.
1M
$0.01
$1.28
Aug 2026
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suite…
1M
$0.01
$1.28
Jul 2026
Inkling Small
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned…
524K
$0.45
$1.20
Jul 2026
Claude Opus 5
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software ta…
1M
$5.00
$25.00
Jul 2026
Claude Opus 5 (Fast)
deprecatedFast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https:/…
1M
$10.00
$50.00
Jul 2026
Ling 3.0 Flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *to…
262K
$0.02
$0.06
Jul 2026
Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within c…
1M
$0.30
$2.50
Jul 2026
Laguna S 2.1
Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-…
1M
$0.09
$0.18
Jul 2026
Gemini 3.6 Flash
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs…
1M
$0.75
$3.75
Jul 2026
LongCat 2.0
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level…
1M
$0.30
$1.20
Jul 2026
Qwen3.8 Max (0803)
deprecatedQwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to th…
1M
$2.00
$6.00
Jul 2026
Auto Router (Beta)
The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, b…
2M
-
-
Jul 2026
Inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for gene…
524K
$1.00
$4.05
Jul 2026
Kimi K3
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic…
1M
$0.80
$15.00
Jul 2026
Kimi Latest
This model always redirects to the latest model in the Kimi family.
1M
$0.44
$14.89
Jul 2026
Qwen3.7 Flash
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with stre…
1M
$0.03
$0.13
Jul 2026
KAT-Coder-Air V2.5
deprecatedKAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to au…
256K
$0.15
$0.60
Jul 2026
KAT-Coder-Pro V2.5
deprecatedKAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to au…
262K
$0.74
$2.96
Jul 2026
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at…
1.1M
$2.00
$10.00
Jul 2026
GPT-5.6 Luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, an…
1.1M
$0.20
$1.20
Jul 2026
GPT-5.6 Terra
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for every…
1.1M
$2.00
$12.00
Jul 2026
GPT-5.6 Luna Pro
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost…
1.1M
$0.20
$1.20
Jul 2026
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
1.1M
$2.00
$10.00
Jul 2026
GPT-5.6 Terra Pro
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cos…
1.1M
$2.00
$12.00
Jul 2026
OpenAI GPT Latest
deprecatedThis model always redirects to the latest model in the OpenAI GPT family.
1.1M
$2.00
$10.00
Jul 2026
GPT Terra Latest
This model always redirects to the latest model in the GPT Terra family.
1.1M
$2.00
$12.00
Jul 2026
Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, and PDF documents and returns text, with a 1…
1M
$1.25
$4.25
Jul 2026
Grok 4.5
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
500K
$2.00
$6.00
Jul 2026
Aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in w…
131K
$3.00
$6.00
Jul 2026
Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation pr…
131K
$0.70
$1.40
Jul 2026
Hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-wor…
262K
$0.13
$0.53
Jul 2026
Laguna XS 2.1
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026).…
262K
$0.06
$0.12
Jul 2026
Claude Sonnet 5
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking…
1M
$2.00
$10.00
Jun 2026
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and r…
66K
$0.25
$1.50
Jun 2026
Nex-N2-Mini
deprecatedNex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is b…
262K
$0.03
$0.10
Jun 2026
North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B ac…
256K
Free
Free
Jun 2026
GLM 5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent work…
1M
$0.06
$7.00
Jun 2026
Fugu Ultra
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration syste…
1M
$5.00
$30.00
Jun 2026
Fusion
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web f…
1M
-
-
Jun 2026
Kimi K2.7 Code
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long context…
262K
$0.67
$3.35
Jun 2026
Claude Fable 5
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text outpu…
1M
$10.00
$50.00
Jun 2026
Nex-N2-Pro
deprecatedNex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts tex…
262K
$0.25
$1.00
Jun 2026
Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…
262K
$0.50
$2.20
Jun 2026
Nemotron 3.5 Content Safety
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…
131K
$0.20
$0.20
Jun 2026
Qwen3.7 Plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilitie…
1M
$0.32
$1.28
Jun 2026
MiniMax M3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited…
1M
$0.30
$1.20
May 2026
Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reaso…
1M
$5.00
$25.00
May 2026
Claude Opus 4.8 (Fast)
deprecatedFast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: htt…
1M
$10.00
$50.00
May 2026
Nano Banana 2 (Gemini 3.1 Flash Image)
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at…
131K
$0.50
$3.00
May 2026
Nano Banana Pro (Gemini 3 Pro Image)
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly imp…
131K
$2.00
$12.00
May 2026
Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for n…
262K
$0.20
$1.15
May 2026
Qwen3.7 Max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular s…
1M
$1.48
$4.43
May 2026
Grok Build 0.1
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text out…
256K
$1.00
$2.00
May 2026
Gemini 3.5 Flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimize…
1M
$1.50
$9.00
May 2026
Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired wi…
33K
$0.15
$1.50
May 2026
Ring-2.6-1T
deprecatedRing-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and ope…
262K
$0.08
$0.63
May 2026
Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio,…
1M
$0.25
$1.50
May 2026
GPT Chat Latest
GPT Chat Latest
400K
$5.00
$30.00
May 2026
Granite 4.1 8B
deprecatedGranite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window an…
131K
$0.05
$0.10
Apr 2026
Laguna M.1
deprecatedLaguna M.1 is the flagship coding agent model from Poolside, optimized for complex software engineering tasks. Designed for agentic coding workflows, it suppor…
262K
$0.20
$0.40
Apr 2026
Laguna XS.2
deprecatedLaguna XS.2 is the second-generation model in the XS size class from Poolside, their efficient coding agent series. It combines tool calling and reasoning capa…
262K
$0.10
$0.20
Apr 2026
Mistral Medium 3.5
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic…
262K
$1.50
$7.50
Apr 2026
Nemotron 3 Nano Omni (free)
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…
256K
Free
Free
Apr 2026
Owl Alpha
Owl Alpha is a high-performance foundation model designed for agentic workloads. Natively supports tool use, and long-context tasks, with strong performance in…
1M
Free
Free
Apr 2026
Qwen3.6 27B
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities —…
262K
$0.32
$3.20
Apr 2026
Qwen3.6 35B A3B
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hyb…
262K
$0.10
$0.95
Apr 2026
Qwen3.6 Flash
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tier…
1M
$0.19
$1.13
Apr 2026
Qwen3.6 Max Preview
deprecatedQwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total pa…
262K
$1.03
$6.16
Apr 2026
DeepSeek V4 Pro 0423
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context…
1M
$0.21
$0.42
Apr 2026
DeepSeek V4 Flash 0423
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-to…
1M
$0.03
$1.28
Apr 2026
GPT-5.5
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved t…
1.1M
$5.00
$30.00
Apr 2026
GPT-5.5 Pro
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context wind…
1.1M
$30.00
$180.00
Apr 2026
Ling-2.6-1T
deprecatedLing-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast exe…
262K
$0.08
$0.63
Apr 2026
Hy3 preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning le…
262K
$0.18
$0.60
Apr 2026
MiMo-V2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in…
1.1M
$0.14
$0.28
Apr 2026
MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks,…
1.1M
$0.44
$0.87
Apr 2026
GPT-5.4 Image 2
deprecatedGPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, all…
272K
$8.00
$15.00
Apr 2026
Ling-2.6-flash
deprecatedLing-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that requi…
262K
$0.01
$0.03
Apr 2026
Pareto Code Router
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by Artificial Analysis coding percentiles. Set min_coding_score between 0 and 1…
2M
-
-
Apr 2026
Kimi K2.6
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. I…
262K
$0.65
$3.41
Apr 2026
Grok 4.3
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following task…
1M
$1.25
$2.50
Apr 2026
Claude Opus 4.7
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4…
1M
$5.00
$25.00
Apr 2026
Claude Opus 4.7 (Fast)
deprecatedFast-mode variant of Opus 4.7 - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.…
1M
$30.00
$150.00
Apr 2026
GLM 5.1
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around min…
205K
$0.97
$3.04
Apr 2026
Gemma 4 26B A4B
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token du…
262K
$0.07
$0.23
Apr 2026
Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window,…
262K
$0.09
$0.34
Apr 2026
Qwen3.6 Plus
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and…
1M
$0.33
$1.95
Apr 2026
GLM 5V Turbo
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video,…
203K
$1.20
$4.00
Apr 2026
Trinity Large Thinking
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and r…
262K
$0.25
$0.80
Apr 2026
Grok 4.20
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on th…
2M
$1.25
$2.50
Mar 2026
Lyria 3 Clip Preview
deprecated30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, y…
1M
Free
Free
Mar 2026
Lyria 3 Pro Preview
deprecatedFull-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can…
1M
Free
Free
Mar 2026
KAT-Coder-Pro V2
deprecatedKAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integr…
262K
$0.30
$1.20
Mar 2026
Reka Edge
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimize…
16K
$0.10
$0.10
Mar 2026
MiniMax M2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participa…
205K
$0.21
$0.84
Mar 2026
GPT-5.4 Mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inpu…
400K
$0.75
$4.50
Mar 2026
GPT-5.4 Nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and…
400K
$0.20
$1.25
Mar 2026
GPT Mini Latest
This model always redirects to the latest model in the GPT Mini family.
400K
$0.75
$4.50
Mar 2026
Mistral Small 4
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It co…
262K
$0.15
$0.60
Mar 2026
GLM 5 Turbo
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply o…
203K
$1.20
$4.00
Mar 2026
Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…
262K
$0.08
$0.45
Mar 2026
Qwen3.5-9B
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-…
262K
$0.10
$0.15
Mar 2026
Seed-2.0-Lite
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latenc…
262K
$0.25
$2.00
Mar 2026
Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct de…
2M
$1.25
$2.50
Mar 2026
GPT-5.4
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K outp…
1.1M
$2.50
$15.00
Mar 2026
GPT-5.4 Pro
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It…
1.1M
$30.00
$180.00
Mar 2026
Step 5 Preview
NEWStep 5 Preview is StepFun's flagship model for agentic work, built on a sparse Mixture-of-Experts architecture (27B active / 600B total parameters). It perform…
Claude Haiku 5.5
NEWClaude Haiku 5.5 is Anthropic's small, fast model for high-volume, cost-sensitive work such as summarization, subagents, and browser use. It succeeds Claude Ha…
Claude Haiku Latest
NEWThis model always redirects to the latest model in the Claude Haiku family.
Nano Banana 2.1
NEWNano Banana 2.1 (Gemini Nano Banana 2.1) is Google's image generation and editing model on the Flash tier, succeeding Nano Banana 2 and Nano Banana Pro. It imp…
Mistral Large 4
NEWMistral Large 4 is a frontier multimodal (text and image input) model from Mistral AI built for reasoning, coding, and agentic workloads. It offers a 1M-token…
Ling 3.1 Flash
NEWLing 3.1 Flash is a hybrid reasoning mixture-of-experts model from inclusionAI, with 25B active parameters out of 560B total.
Apodex 1.1 Mini (free)
NEWApodex 1.1 Mini is a reasoning-first model from Apodex, built for complex, long-horizon research and forecasting tasks. It works directly with files, data, cod…
Pareto 26.10 Preview
NEWPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of g…
GPT-6.1 Sol
NEWGPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding, computer us…
GPT-6.1 Sol Pro
NEWGPT-6.1 Sol Pro is the same underlying model as GPT-6.1 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
GPT Sol Latest
NEWThis model always redirects to the latest model in the GPT Sol family.
Claude Sonnet 5.5
NEWHOTClaude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially strong at b…
Claude Sonnet Latest
NEWThis model always redirects to the latest model in the Claude Sonnet family.
Jev Router
NEWJev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on Jev, TypeSafe's first System One model, a…
Perceptron Mk1.5
NEWPerceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with text plus optio…
Ember-1
NEWEmber-1 is a specialized reasoning model from Fireworks Research, built on Kimi K3. It is designed to make every token go further: it produces shorter reasonin…
Solar Mini 4
NEWSolar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context window. It is…
Aion 3.5
NEWAion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in w…
Aion 3.5 Mini
NEWAion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of A…
GLM 5.3 Prime
NEWGLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acc…
Qwen3.8 Max Prime
NEWQwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, im…
Space Bunny Alpha
deprecatedSpace Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustab…
Claude Opus 5.5
NEWHOTClaude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly stro…
GPT-6 Luna Pro
NEWGPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note…
GPT-6 Luna
NEWGPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads…
Claude Opus Latest
NEWThis model always redirects to the latest model in the Claude Opus family.
Command A+
NEWCommand A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports native tool calli…
GPT-6 Sol
NEWGPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is su…
GPT-6 Sol Pro
NEWGPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost note:*…
GPT Luna Latest
NEWThis model always redirects to the latest model in the GPT Luna family.
MiMo-V2.6-Pro
NEWMiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability fo…
Grok 4.7
NEWGrok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software en…
Qwen3.8 Omni Flash
NEWQwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding.…
Grok Latest
NEWThis model always redirects to the latest Grok model from xAI.
MiMo-V2.6-Flash
NEWMiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated…
MiMo-V2.6-Pro-UltraSpeed
NEWMiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it ma…
Switchyard
NEWSwitchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter market data…
GLM 5.3 FlashX
NEWGLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the sam…
Ternary Bonsai 2 27B
NEWBonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding w…
Pareto
NEWPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of g…
Union Alpha
deprecatedUnion Alpha is a multimodal model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of genera…
DeepSeek Pro Latest
NEWThis model always redirects to the latest model in the DeepSeek Pro family.
DeepSeek Flash Latest
NEWThis model always redirects to the latest model in the DeepSeek Flash family.
Schematron V2 Small
NEWSchematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and long pages. E…
Schematron V2 Turbo
NEWSchematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction workloads. Extract…
Fugu Max
NEWFugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a…
Fugu Ultra v2
NEWFugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration sy…
DeepSeek V4.1 Flash
NEWHOTDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It acti…
Ling 3.0 Flash VL
NEWLing 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native…
Mercury 2.5
NEWMercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces a…
Nex-N2.5-Mini
NEWNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can exp…
Nex-N2.5-Pro
NEWNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can exp…
Ling 3.0 Flash Sante
NEWLing 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124…
GPT-6 Astra
NEWGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work,…
GPT-6 Astra Pro
NEWGPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
GPT Astra Latest
NEWThis model always redirects to the latest model in the GPT Astra family.
Gemini 3.8 Flash
NEWHOTGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reas…
Gemini Flash Latest
NEWThis model always redirects to the latest model in the Gemini Flash family.
Muse Spark 1.3
NEWMuse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of informati…
Muse Spark 1.3 Contributor
NEWMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic,…
Qwen3.8 Max (0902)
NEWQwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, ima…
Claude Fable 5.1
NEWClaude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: lon…
Claude Fable Latest
NEWThis model always redirects to the latest model in the Claude Fable family.
Granite 4.2 8B
NEWGranite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi…
Mercury 2.5 Preview
deprecatedMercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces a…
Hy4 preview
NEWTencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-u…
Ling 3.0 Flash Fin
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is…
Qwen3.8 Flash
Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase…
GLM Flash Latest
This model always redirects to the latest model in the GLM Flash family.
GLM 5.3 Flash
GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention ar…
DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while matching the b…
Hy-MT2-1.8B
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, wit…
Hy-MT2-30B-A3B
Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and minority-language pairs…
Ox Alpha
deprecatedOx Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, comple…
Hy-MT2-7B
Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows…
GLM 5.3
GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with…
Qwen3.8 27B
Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and lon…
Dots3-Note Preview (free)
Dots3-Note Preview is an open-weight mixture-of-experts model from Dots Studio, with 16B active parameters out of 280B total. It is the lightest model in the D…
GLM Latest
This model always redirects to the latest GLM model from Z.ai.
Gemini 3.7 Flash
HOTGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require…
DeepSeek V4 Pro 0813
DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
Grok 4.6
Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by Grok 4.7.
Qwen3.8 2.4T A95B
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active parameters out…
Seed-2.0-Code
Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks, and coding-ag…
Seed 2.1 Turbo
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step…
Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput age…
LFM2.5-2.6B (free)
LFM2.5-2.6B is a compact reasoning model from Liquid AI. It is suited for agent workflows, data extraction, RAG, and long-context processing. Liquid advises ag…
Namazu
deprecatedSakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts…
Solar Pro 4
Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with s…
Muse Glimmer 30B
Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on co…
Ling 3.0 Tiny (free)
deprecatedLing 3.0 Tiny is a mixture-of-experts model from InclusionAI, with 1.3B active parameters out of 7.9B total. It is designed for responsive agents, instruction…
Muse Spark 1.2
Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, and PDF documents, returns text, and offers…
Muse Spark 1.2 Contributor
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully chea…
Sakana Namazu
Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and business contexts…
DeepSeek V4 Flash Latest
This model always redirects to the latest model in the DeepSeek V4 Flash family.
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suite…
Inkling Small
Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned…
Claude Opus 5
Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software ta…
Claude Opus 5 (Fast)
deprecatedFast-mode variant of Opus 5 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https:/…
Ling 3.0 Flash
*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *to…
Gemini 3.5 Flash Lite
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within c…
Laguna S 2.1
Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-…
Gemini 3.6 Flash
Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs…
LongCat 2.0
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level…
Qwen3.8 Max (0803)
deprecatedQwen3.8 Max (0803) is the August 3, 2026 checkpoint of Qwen3.8 Max, the flagship model in Alibaba's Qwen3.8 series and the general-availability successor to th…
Auto Router (Beta)
The experimental version of our Auto Router where we test new improvements. Use it to get the latest and greatest version of our general purpose auto router, b…
Inkling
Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for gene…
Kimi K3
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic…
Kimi Latest
This model always redirects to the latest model in the Kimi family.
Qwen3.7 Flash
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with stre…
KAT-Coder-Air V2.5
deprecatedKAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to au…
KAT-Coder-Pro V2.5
deprecatedKAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to au…
GPT-5.6 Sol
GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at…
GPT-5.6 Luna
GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, an…
GPT-5.6 Terra
GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for every…
GPT-5.6 Luna Pro
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost…
GPT-5.6 Sol Pro
GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cost no…
GPT-5.6 Terra Pro
GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. **Cos…
OpenAI GPT Latest
deprecatedThis model always redirects to the latest model in the OpenAI GPT family.
GPT Terra Latest
This model always redirects to the latest model in the GPT Terra family.
Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, and PDF documents and returns text, with a 1…
Grok 4.5
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
Aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in w…
Aion-3.0-Mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation pr…
Hy3
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-wor…
Laguna XS 2.1
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026).…
Claude Sonnet 5
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking…
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and r…
Nex-N2-Mini
deprecatedNex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is b…
North Mini Code (free)
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B ac…
GLM 5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent work…
Fugu Ultra
Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration syste…
Fusion
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web f…
Kimi K2.7 Code
MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long context…
Claude Fable 5
Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text outpu…
Nex-N2-Pro
deprecatedNex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts tex…
Nemotron 3 Ultra
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybri…
Nemotron 3.5 Content Safety
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both input…
Qwen3.7 Plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilitie…
MiniMax M3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited…
Claude Opus 4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reaso…
Claude Opus 4.8 (Fast)
deprecatedFast-mode variant of Opus 4.8 - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: htt…
Nano Banana 2 (Gemini 3.1 Flash Image)
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at…
Nano Banana Pro (Gemini 3 Pro Image)
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly imp…
Step 3.7 Flash
Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for n…
Qwen3.7 Max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular s…
Grok Build 0.1
Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text out…
Gemini 3.5 Flash
Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimize…
Perceptron Mk1
Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired wi…
Ring-2.6-1T
deprecatedRing-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and ope…
Gemini 3.1 Flash Lite
Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio,…
GPT Chat Latest
GPT Chat Latest
Granite 4.1 8B
deprecatedGranite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window an…
Laguna M.1
deprecatedLaguna M.1 is the flagship coding agent model from Poolside, optimized for complex software engineering tasks. Designed for agentic coding workflows, it suppor…
Laguna XS.2
deprecatedLaguna XS.2 is the second-generation model in the XS size class from Poolside, their efficient coding agent series. It combines tool calling and reasoning capa…
Mistral Medium 3.5
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic…
Nemotron 3 Nano Omni (free)
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It acce…
Owl Alpha
Owl Alpha is a high-performance foundation model designed for agentic workloads. Natively supports tool use, and long-context tasks, with strong performance in…
Qwen3.6 27B
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities —…
Qwen3.6 35B A3B
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hyb…
Qwen3.6 Flash
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tier…
Qwen3.6 Max Preview
deprecatedQwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total pa…
DeepSeek V4 Pro 0423
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context…
DeepSeek V4 Flash 0423
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-to…
GPT-5.5
GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved t…
GPT-5.5 Pro
GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context wind…
Ling-2.6-1T
deprecatedLing-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast exe…
Hy3 preview
Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning le…
MiMo-V2.5
MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in…
MiMo-V2.5-Pro
MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks,…
GPT-5.4 Image 2
deprecatedGPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, all…
Ling-2.6-flash
deprecatedLing-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that requi…
Pareto Code Router
The Pareto Router maintains a tiered shortlist of strong coding models, ranked by Artificial Analysis coding percentiles. Set min_coding_score between 0 and 1…
Kimi K2.6
Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. I…
Grok 4.3
Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following task…
Claude Opus 4.7
Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4…
Claude Opus 4.7 (Fast)
deprecatedFast-mode variant of Opus 4.7 - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.…
GLM 5.1
GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around min…
Gemma 4 26B A4B
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token du…
Gemma 4 31B
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window,…
Qwen3.6 Plus
Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and…
GLM 5V Turbo
GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video,…
Trinity Large Thinking
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and r…
Grok 4.20
Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on th…
Lyria 3 Clip Preview
deprecated30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, y…
Lyria 3 Pro Preview
deprecatedFull-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can…
KAT-Coder-Pro V2
deprecatedKAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integr…
Reka Edge
Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimize…
MiniMax M2.7
MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participa…
GPT-5.4 Mini
GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inpu…
GPT-5.4 Nano
GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and…
GPT Mini Latest
This model always redirects to the latest model in the GPT Mini family.
Mistral Small 4
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It co…
GLM 5 Turbo
GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply o…
Nemotron 3 Super
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-…
Qwen3.5-9B
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-…
Seed-2.0-Lite
Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latenc…
Grok 4.20 Multi-Agent
Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct de…
GPT-5.4
GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K outp…
GPT-5.4 Pro
GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks. It…
1
Create an API key at the OpenRouter console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
One OpenRouter key, and Big-AGI adds no markup and no intermediary: the billing relationship runs directly between you and OpenRouter. You bring the aggregator, Big-AGI brings the workspace.
Your keys stay in your browser, not on Big-AGI's servers. Chats are stored locally first, and sync only if you turn it on. Nothing is added to your system prompt, and the AI Inspector shows the exact request, the token counts, and a cost estimate for every call.
This is the part a router cannot do for you. Send one prompt to models from different providers at once, Anthropic next to OpenAI next to an open model OpenRouter hosts. Where they agree, you can trust the answer. Where they differ, you have caught something worth a second look. Fusions then combine, cross-check, and synthesize the parallel answers instead of just picking the best one. Big-AGI has shipped and refined this workflow since 2024. Parallel runs use more tokens than a single chat.
Big-AGI is an independent, open-source client. It is not affiliated with or endorsed by OpenRouter.
FAQ
On purpose. OpenRouter caps its no-charge models per minute and per day, so Big-AGI spaces its own calls to them about five seconds apart instead of collecting rate-limit errors. That is pacing, not a fault.
Press Link OpenRouter Key in Big-AGI's setup panel: you authorise at OpenRouter and the key comes back installed, no copy-paste. Or create one at openrouter.ai/keys; it reads sk-or-v1-... and is shown once, so copy it before leaving the page.
Yes. OpenRouter is one of the services Direct Connection supports: turn it on in the OpenRouter setup panel once the key is in the browser, and requests go browser-to-OpenRouter, skipping the Big-AGI edge. It starts off.
Not for the no-charge models: they answer on a low daily cap with no deposit. Real use runs on prepaid credits added at openrouter.ai/settings/credits, and that balance is the spending limit unless auto top-up is on, which refills it.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how OpenRouter is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego