Muse Spark on your own key: a million-token context, reasoning effort from minimal to xhigh, web search and Muse Image - with the Contributor models that train on your prompts hidden until you opt in.
Muse Spark 1.3
NEWMuse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of informati…
1M
$1.25
$4.25
Sep 2026
Muse Spark 1.3 Contributor
NEWMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic,…
1M
$0.10
$0.20
Sep 2026
Muse Spark 1.2 Contributor
NEWMuse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully chea…
1M
$0.10
$0.20
Aug 2026
Muse Glimmer 30B
NEWMuse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on co…
131K
$0.30
$1.10
Aug 2026
Muse Spark 1.2
NEWMuse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and…
1M
$1.25
$4.25
Aug 2026
Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, w…
1M
$1.25
$4.25
Jul 2026
Llama 4 Maverick 17B 128E Instruct Nvfp4
Meta chat model. https://huggingface.co/api/models/RedHatAI/Llama-4-Maverick-17B-128E-Instruct-NVFP4
1M
-
-
Jun 2026
Llama 4 Scout 17B 16E Instruct Fp8 Lora
Meta chat model.
10M
-
-
May 2026
Llama 3.3 70B Instruct FP8 Lora
Meta chat model.
131K
-
-
May 2026
Llama 4 Scout (17Bx16E)
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-4-Scout-17B-16E
262K
-
-
Jun 2025
Llama Guard 4 12B
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be use…
164K
$0.18
$0.18
Apr 2025
Llama 3.1 405B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.1-405B
131K
-
-
Apr 2025
Llama 3.2 1B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.2-1B
131K
-
-
Apr 2025
Llama 3.1 70B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.1-70B
131K
-
-
Apr 2025
Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It su…
1.3M
$0.10
$0.30
Apr 2025
Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts…
1M
$0.20
$0.70
Apr 2025
Llama 4 Scout Instruct (17Bx16E)
deprecatedMeta chat model. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct
1M
$0.18
$0.59
Apr 2025
meta-llama/Llama-2-7b-chat-hf
Meta chat model. https://huggingface.co/meta-llama/Llama-2-7b-chat-hf
4K
-
-
Apr 2025
Meta Llama 3.1 8B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct
131K
$0.18
$0.18
Mar 2025
Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 inst…
131K
$0.10
$0.32
Dec 2024
Meta Llama 3.1 405B Instruct
deprecatedMeta chat model. https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct
4K
$3.50
$3.50
Dec 2024
Meta Llama 3.3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
131K
$1.04
$1.04
Dec 2024
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
33K
Free
Free
Sep 2024
Llama 3.2 11B Vision
Llama vision model for image understanding.
131K
Free
Free
Sep 2024
Llama 3.2 1B
Tiny Llama for edge-class tasks.
131K
Free
Free
Sep 2024
Llama 3.2 3B
Small Llama for lightweight tasks.
131K
Free
Free
Sep 2024
Llama 3.1 8B
Small fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.
131K
Free
Free
Jul 2024
Llama 3.1 70B
Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
131K
Free
Free
Jul 2024
Meta Llama 3.1 70B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct
131K
$0.88
$0.88
Jul 2024
Meta Llama 3 8B Instruct Reference
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
8K
$0.20
$0.20
Apr 2024
Llama 3 8B Instruct
deprecatedMeta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue useca…
8K
$0.14
$0.14
Apr 2024
Meta Llama 3 8B Instruct
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
8K
$0.20
$0.20
Apr 2024
Meta Llama 3 70B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct
8K
$0.88
$0.88
-
Meta Llama 3 8B Instruct Lite
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
8K
$0.14
$0.14
-
Muse Spark 1.3
NEWMuse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of informati…
Muse Spark 1.3 Contributor
NEWMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic,…
Muse Spark 1.2 Contributor
NEWMuse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully chea…
Muse Glimmer 30B
NEWMuse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous agents on co…
Muse Spark 1.2
NEWMuse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and…
Muse Spark 1.1
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, w…
Llama 4 Maverick 17B 128E Instruct Nvfp4
Meta chat model. https://huggingface.co/api/models/RedHatAI/Llama-4-Maverick-17B-128E-Instruct-NVFP4
Llama 4 Scout 17B 16E Instruct Fp8 Lora
Meta chat model.
Llama 3.3 70B Instruct FP8 Lora
Meta chat model.
Llama 4 Scout (17Bx16E)
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-4-Scout-17B-16E
Llama Guard 4 12B
Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be use…
Llama 3.1 405B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.1-405B
Llama 3.2 1B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.2-1B
Llama 3.1 70B
Meta chat model. https://huggingface.co/api/models/meta-llama/Llama-3.1-70B
Llama 4 Scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It su…
Llama 4 Maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts…
Llama 4 Scout Instruct (17Bx16E)
deprecatedMeta chat model. https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct
meta-llama/Llama-2-7b-chat-hf
Meta chat model. https://huggingface.co/meta-llama/Llama-2-7b-chat-hf
Meta Llama 3.1 8B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct
Llama 3.3 70B Instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 inst…
Meta Llama 3.1 405B Instruct
deprecatedMeta chat model. https://huggingface.co/meta-llama/Llama-3.1-405B-Instruct
Meta Llama 3.3 70B Instruct Turbo
Meta chat model. https://huggingface.co/meta-llama/Llama-3.3-70B-Instruct
Llama 3.2 90B Vision
Llama large vision model. NVIDIA serves a reduced 32K context (native: 128K).
Llama 3.2 11B Vision
Llama vision model for image understanding.
Llama 3.2 1B
Tiny Llama for edge-class tasks.
Llama 3.2 3B
Small Llama for lightweight tasks.
Llama 3.1 8B
Small fast Llama for utility tasks, with tool calling. Retires on NVIDIA 2026-08-25.
Llama 3.1 70B
Meta Llama 3.1 70B instruction-tuned (superseded by Llama 3.3 70B).
Meta Llama 3.1 70B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3.1-70B-Instruct
Meta Llama 3 8B Instruct Reference
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
Llama 3 8B Instruct
deprecatedMeta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue useca…
Meta Llama 3 8B Instruct
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
Meta Llama 3 70B Instruct Turbo
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct
Meta Llama 3 8B Instruct Lite
deprecatedMeta chat model. https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct
1
Create an API key at the Meta console.
2
Paste it into Big-AGI's model settings.
3
Start chatting, or Beam it against other models and fuse the answers.
Add your Meta Model API key and use the Muse models at Meta's own rates: Big-AGI adds no markup, and Meta bills you directly. Every Muse Spark version is served on two tiers. Standard does not use your prompts for training; Contributor is cheaper in exchange for Meta training on your prompts and completions. Big-AGI hides the Contributor models until you unhide them.
Your key stays in your browser and is sent only with your requests. Chats are stored on your device first and sync only if you turn sync on. The AI Inspector shows every request, the token counts and the cost estimate, including the reasoning tokens Muse Spark spends before answering.
Beam runs Muse Spark next to Claude, GPT, Gemini and anything else you connected, in parallel, then fuses the answers. Parallel runs use more tokens than a single chat, and Muse Spark's reasoning tokens count too.
FAQ
They are the same model. Standard does not use your prompts and completions for training and costs $1.25 per million input tokens and $4.25 per million output tokens as of September 2, 2026. Contributor costs $0.10 and $0.20 in exchange for Meta training future models on your prompts and completions, with lower rate limits. Big-AGI hides the Contributor models until you unhide them in the model list.
Meta returns the same Unauthorized error for a wrong key, a revoked key, and a missing key. Recreate the key at dev.meta.ai/api-keys and paste it into the Meta Model API Key field; keys start with LLM_ and are shown once at creation.
Muse Spark reasons before every reply, and Meta's default reasoning effort is high. In our September 2, 2026 check a two-word reply spent about 340 reasoning tokens and 9 seconds, with the first token after 2.2 seconds. Lower the Reasoning Effort control - minimal to xhigh, there is no off - to trade depth for speed.
Yes. Muse Image 1.0 appears in the model list next to Muse Spark and generates or edits images from text and reference images, at a flat $0.01 per image billed by Meta. It returns one WebP image per request and runs without streaming.
Your key, your data, your choice of model. Big-AGI's Open branch is open source and self-hostable, so you can check exactly how Meta is called.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego