Qwen3.8 Max and the rest of Model Studio at Alibaba's own rates - filtered to just the chat models, with resold models labeled "(Alibaba)".
Qwen3.8 27B
NEWOpen-weights 27B dense vision-language model of the Qwen3.8 line. 1M context, thinking, image/video input.
1M
$0.50
$3.00
Aug 2026
DeepSeek V4 Pro 0813 (Alibaba)
NEWHOTDeepSeek V4 Pro GA (0813) served via Alibaba Model Studio. Much stronger agentic and tool use than the April preview. 1M context, thinking.
1M
$2.40
$4.80
Aug 2026
Qwen3.8 2.4T-A95B
NEWOpen-weights release of the Qwen3.8 flagship: 2.4T sparse MoE, ~95B active. Text-only serving with 1M context and always-on thinking.
1M
$2.00
$6.00
Aug 2026
DeepSeek V4 Flash 0731 (Alibaba)
NEWHOTDeepSeek V4 Flash 0731 revision served via Alibaba Model Studio. 1M context, thinking.
1M
$0.20
$0.40
Jul 2026
Qwen3.8 Max
NEWFlagship 2.4T-parameter sparse MoE multimodal model with 1M context, thinking, and vision/video understanding.
1M
$2.00
$6.00
Jul 2026
Kimi K3 (Alibaba)
HOTMoonshot Kimi K3 flagship served via Alibaba Model Studio. Multimodal, thinking on by default, 1M context.
1M
$3.00
$15.00
Jul 2026
Qwen3.7 Flash
Latest fast multimodal model with 1M context, thinking (on by default), vision, and 128K output.
1M
$0.03
$0.13
Jul 2026
GLM-5.2 (Alibaba)
Zhipu GLM-5.2 served via Alibaba Model Studio. 1M context, thinking.
1M
$1.40
$4.40
Jun 2026
Kimi K2.7 Code (Alibaba)
Moonshot Kimi K2.7 Code served via Alibaba Model Studio. Multimodal, always-on thinking, 256K context.
262K
$0.95
$4.00
Jun 2026
Qwen3.7 Plus
Multimodal agent model with 1M context, native thinking, and vision/video understanding. Lower cost than Max.
1M
$0.40
$1.60
Jun 2026
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M context. Text-only; strong at coding, productivity, and long-horizon autonomous tasks.
1M
$2.50
$7.50
May 2026
Qwen3.6 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
262K
$0.38
$2.25
Apr 2026
Qwen3.6 27B
Open 27B dense multimodal with thinking. 256K context.
262K
$0.60
$3.60
Apr 2026
Qwen3.6 Flash
Fast, cost-effective multimodal model with 1M context, near-flagship quality, vision/video, and built-in tools.
1M
$0.25
$1.50
Apr 2026
Qwen3.6 Max Preview
Qwen3.6 Max preview (never GA; superseded by Qwen3.7 Max). 256K context, thinking, text-only.
262K
$1.30
$7.80
Apr 2026
DeepSeek V4 Flash (Alibaba)
HOTDeepSeek V4 Flash served via Alibaba Model Studio. 1M context, thinking.
1M
$0.20
$0.40
Apr 2026
DeepSeek V4 Pro (Alibaba)
HOTDeepSeek V4 Pro served via Alibaba Model Studio (Alibaba pricing, well above DeepSeek-direct). 1M context, thinking.
1M
$2.40
$4.80
Apr 2026
GLM-5.1 (Alibaba)
Zhipu GLM-5.1 served via Alibaba Model Studio, superseded by GLM-5.2. 200K context, thinking.
203K
$1.40
$4.40
Apr 2026
Qwen3.6 Plus
Previous Plus tier, superseded by Qwen3.7 Plus. 1M context, thinking, vision.
1M
$0.50
$3.00
Apr 2026
Qwen3.5 122B-A10B
Open 122B-A10B MoE multimodal with thinking. 256K context.
262K
$0.40
$3.20
Feb 2026
Qwen3.5 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
262K
$0.25
$2.00
Feb 2026
Qwen3.5 27B
Open 27B dense multimodal with thinking. 256K context.
262K
$0.30
$2.40
Feb 2026
Qwen3.5 Flash
Former Flash tier, superseded by Qwen3.6/3.7 Flash. 1M context, thinking, vision.
1M
$0.10
$0.40
Feb 2026
Qwen3.5 397B-A17B
Open 397B-A17B MoE multimodal with thinking. 256K context.
262K
$0.60
$3.60
Feb 2026
Qwen3.5 Plus
Former Plus tier, superseded by Qwen3.6/3.7 Plus. 1M context, thinking, vision.
1M
$0.40
$2.40
Feb 2026
Qwen3 Coder Next
Budget agentic coder on the Qwen3-Next architecture. 256K context, non-thinking, tiered pricing.
262K
$0.30
$1.50
Feb 2026
DeepSeek V3.2 (Alibaba)
DeepSeek V3.2 served via Alibaba Model Studio (superseded by V4). Thinking.
131K
$0.57
$1.71
Dec 2025
Qwen3 VL Flash
Budget VL tier below Qwen3 VL Plus. 256K context, thinking, vision.
262K
$0.05
$0.40
Oct 2025
Qwen3 Vl 235b A22b Thinking
deprecatedAlibaba model (not yet curated).
131K
-
-
Sep 2025
Qwen3 Vl 235b A22b Instruct
deprecatedAlibaba model (not yet curated).
131K
-
-
Sep 2025
Qwen3 Coder Plus
Agentic coding model with very long context. Tiered pricing by input length (up to 1M).
1M
$1.00
$5.00
Sep 2025
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context; the base id now serves the thinking-capable 2026-01-23 snapshot.
262K
$1.20
$6.00
Sep 2025
Qwen3 Next 80b A3b Thinking
deprecatedAlibaba model (not yet curated).
131K
-
-
Sep 2025
Qwen3 Next 80b A3b Instruct
deprecatedAlibaba model (not yet curated).
131K
-
-
Sep 2025
Qwen3 30b A3b Thinking 2507
deprecatedAlibaba model (not yet curated).
131K
-
-
Aug 2025
Qwen3 30b A3b Instruct 2507
deprecatedAlibaba model (not yet curated).
131K
-
-
Jul 2025
Qwen3 Coder Flash
Former budget coder, superseded by Qwen3 Coder Next (which lacks its 1M context). Non-thinking.
1M
$0.30
$1.50
Jul 2025
Qwen3 235b A22b Thinking 2507
deprecatedAlibaba model (not yet curated).
131K
-
-
Jul 2025
Qwen3 30b A3b
deprecatedAlibaba model (not yet curated).
131K
-
-
Apr 2025
Qwen3 32b
deprecatedAlibaba model (not yet curated).
131K
-
-
Apr 2025
Qwen3 14b
deprecatedAlibaba model (not yet curated).
131K
-
-
Apr 2025
Qwen3 235b A22b
deprecatedAlibaba model (not yet curated).
131K
-
-
Apr 2025
Qwen3 8b
deprecatedAlibaba model (not yet curated).
131K
-
-
Apr 2025
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M context.
1M
$0.40
$1.20
Feb 2025
Qwen3 235b A22b Instruct 2507
deprecatedAlibaba model (not yet curated).
131K
-
-
-
GLM-5.2 Fast (Alibaba)
Zhipu GLM-5.2 fast-serving tier via Alibaba Model Studio (preview). Same model, lower latency, ~2x price.
1M
$2.80
$8.80
-
Qvq Max
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen Coder Plus
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen Flash
Fast and very low cost with hybrid thinking. 1M context.
1M
$0.05
$0.40
-
Qwen Max
Best quality of the stable commercial line. 32K context.
33K
$1.60
$6.40
-
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M context.
1M
$0.40
$1.20
-
Qwen Turbo
Fastest and cheapest for simple tasks. 1M context.
1M
$0.05
$0.20
-
Qwen Vl Max
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen Vl Plus
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M context. Text-only; strong at coding, productivity, and long-horizon autonomous tasks.
1M
$2.50
$7.50
-
Qwen3 Coder 480b A35b Instruct
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context; the base id now serves the thinking-capable 2026-01-23 snapshot.
262K
$1.20
$6.00
-
Qwen3 VL Plus
Current vision-language model with strong visual reasoning and thinking. Tiered pricing by input length (up to 256K).
262K
$0.20
$1.60
-
Qwq Plus
deprecatedAlibaba model (not yet curated).
131K
-
-
-
Qwen3.8 27B
NEWOpen-weights 27B dense vision-language model of the Qwen3.8 line. 1M context, thinking, image/video input.
DeepSeek V4 Pro 0813 (Alibaba)
NEWHOTDeepSeek V4 Pro GA (0813) served via Alibaba Model Studio. Much stronger agentic and tool use than the April preview. 1M context, thinking.
Qwen3.8 2.4T-A95B
NEWOpen-weights release of the Qwen3.8 flagship: 2.4T sparse MoE, ~95B active. Text-only serving with 1M context and always-on thinking.
DeepSeek V4 Flash 0731 (Alibaba)
NEWHOTDeepSeek V4 Flash 0731 revision served via Alibaba Model Studio. 1M context, thinking.
Qwen3.8 Max
NEWFlagship 2.4T-parameter sparse MoE multimodal model with 1M context, thinking, and vision/video understanding.
Kimi K3 (Alibaba)
HOTMoonshot Kimi K3 flagship served via Alibaba Model Studio. Multimodal, thinking on by default, 1M context.
Qwen3.7 Flash
Latest fast multimodal model with 1M context, thinking (on by default), vision, and 128K output.
GLM-5.2 (Alibaba)
Zhipu GLM-5.2 served via Alibaba Model Studio. 1M context, thinking.
Kimi K2.7 Code (Alibaba)
Moonshot Kimi K2.7 Code served via Alibaba Model Studio. Multimodal, always-on thinking, 256K context.
Qwen3.7 Plus
Multimodal agent model with 1M context, native thinking, and vision/video understanding. Lower cost than Max.
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M context. Text-only; strong at coding, productivity, and long-horizon autonomous tasks.
Qwen3.6 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
Qwen3.6 27B
Open 27B dense multimodal with thinking. 256K context.
Qwen3.6 Flash
Fast, cost-effective multimodal model with 1M context, near-flagship quality, vision/video, and built-in tools.
Qwen3.6 Max Preview
Qwen3.6 Max preview (never GA; superseded by Qwen3.7 Max). 256K context, thinking, text-only.
DeepSeek V4 Flash (Alibaba)
HOTDeepSeek V4 Flash served via Alibaba Model Studio. 1M context, thinking.
DeepSeek V4 Pro (Alibaba)
HOTDeepSeek V4 Pro served via Alibaba Model Studio (Alibaba pricing, well above DeepSeek-direct). 1M context, thinking.
GLM-5.1 (Alibaba)
Zhipu GLM-5.1 served via Alibaba Model Studio, superseded by GLM-5.2. 200K context, thinking.
Qwen3.6 Plus
Previous Plus tier, superseded by Qwen3.7 Plus. 1M context, thinking, vision.
Qwen3.5 122B-A10B
Open 122B-A10B MoE multimodal with thinking. 256K context.
Qwen3.5 35B-A3B
Open 35B-A3B MoE multimodal with thinking. 256K context.
Qwen3.5 27B
Open 27B dense multimodal with thinking. 256K context.
Qwen3.5 Flash
Former Flash tier, superseded by Qwen3.6/3.7 Flash. 1M context, thinking, vision.
Qwen3.5 397B-A17B
Open 397B-A17B MoE multimodal with thinking. 256K context.
Qwen3.5 Plus
Former Plus tier, superseded by Qwen3.6/3.7 Plus. 1M context, thinking, vision.
Qwen3 Coder Next
Budget agentic coder on the Qwen3-Next architecture. 256K context, non-thinking, tiered pricing.
DeepSeek V3.2 (Alibaba)
DeepSeek V3.2 served via Alibaba Model Studio (superseded by V4). Thinking.
Qwen3 VL Flash
Budget VL tier below Qwen3 VL Plus. 256K context, thinking, vision.
Qwen3 Vl 235b A22b Thinking
deprecatedAlibaba model (not yet curated).
Qwen3 Vl 235b A22b Instruct
deprecatedAlibaba model (not yet curated).
Qwen3 Coder Plus
Agentic coding model with very long context. Tiered pricing by input length (up to 1M).
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context; the base id now serves the thinking-capable 2026-01-23 snapshot.
Qwen3 Next 80b A3b Thinking
deprecatedAlibaba model (not yet curated).
Qwen3 Next 80b A3b Instruct
deprecatedAlibaba model (not yet curated).
Qwen3 30b A3b Thinking 2507
deprecatedAlibaba model (not yet curated).
Qwen3 30b A3b Instruct 2507
deprecatedAlibaba model (not yet curated).
Qwen3 Coder Flash
Former budget coder, superseded by Qwen3 Coder Next (which lacks its 1M context). Non-thinking.
Qwen3 235b A22b Thinking 2507
deprecatedAlibaba model (not yet curated).
Qwen3 30b A3b
deprecatedAlibaba model (not yet curated).
Qwen3 32b
deprecatedAlibaba model (not yet curated).
Qwen3 14b
deprecatedAlibaba model (not yet curated).
Qwen3 235b A22b
deprecatedAlibaba model (not yet curated).
Qwen3 8b
deprecatedAlibaba model (not yet curated).
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M context.
Qwen3 235b A22b Instruct 2507
deprecatedAlibaba model (not yet curated).
GLM-5.2 Fast (Alibaba)
Zhipu GLM-5.2 fast-serving tier via Alibaba Model Studio (preview). Same model, lower latency, ~2x price.
Qvq Max
deprecatedAlibaba model (not yet curated).
Qwen Coder Plus
deprecatedAlibaba model (not yet curated).
Qwen Flash
Fast and very low cost with hybrid thinking. 1M context.
Qwen Max
Best quality of the stable commercial line. 32K context.
Qwen Plus
Balanced quality, speed, and cost with hybrid thinking. 1M context.
Qwen Turbo
Fastest and cheapest for simple tasks. 1M context.
Qwen Vl Max
deprecatedAlibaba model (not yet curated).
Qwen Vl Plus
deprecatedAlibaba model (not yet curated).
Qwen3.7 Max
Flagship agent model with native extended thinking and 1M context. Text-only; strong at coding, productivity, and long-horizon autonomous tasks.
Qwen3 Coder 480b A35b Instruct
deprecatedAlibaba model (not yet curated).
Qwen3 Max
Retired 2025 flagship, superseded by Qwen3.6+ Max. 256K context; the base id now serves the thinking-capable 2026-01-23 snapshot.
Qwen3 VL Plus
Current vision-language model with strong visual reasoning and thinking. Tiered pricing by input length (up to 256K).
Qwq Plus
deprecatedAlibaba model (not yet curated).
1
Create an API key at the Alibaba console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Add your Alibaba Cloud Model Studio API key and run the Qwen family, plus the other models Alibaba serves, at Alibaba's own API rates. Big-AGI adds no markup and no intermediary: billing runs directly between you and Alibaba Cloud, and your keys stay in your browser.
Alibaba's own Qwen app is free and covers general chat well, but it cannot put Qwen's answer next to GPT's, Claude's, or Gemini's on the same prompt. Beam can: agreement across labs is a stronger signal than any single model's confidence, and disagreement tells you where to dig further. You also get parameters the app hides (temperature, system prompt, per-turn model swaps), and a key that stays in your browser instead of sitting on Alibaba's servers.
Turn on Direct Connection and the browser calls Alibaba directly, bypassing the Big-AGI server, whenever your key is client-side and Alibaba allows it. Your keys stay in your browser. Chats are stored locally first and sync only if you turn it on. The AI Inspector opens on any message to show the exact request sent to Alibaba, the token counts, and a cost estimate for that call.
Run Qwen in parallel with Claude, GPT, and Gemini on the same prompt, then reach for Fusions: several strategies that combine, cross-check, and synthesize the parallel answers, which beats just picking the single best one. Parallel runs use more tokens than a single chat.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Alibaba is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego