<- All Models

Gemma 4 31B

Google Gemma 4 31B on Cerebras - first multimodal model on wafer-scale inference (~1,850 tok/s). Vision (base64 PNG/JPEG, max 10 images / 10MB), function calling, reasoning (off by default, enable via effort). 131K context (65K free tier), 40K max output.

131K context·$0.99 in·$1.49 out·Apr 2026 release·Vision, Reasoning, Tools / functions

Providers

Where Gemma 4 31B runs in Big-AGI, with each service's own model ID and USD-per-1M-token rates. The summary above prefers the creator's claim; rates below are each service's own.

ServiceModel IDContextInputOutputCache read

Cerebras

delisted

gemma-4-31b

131K

$0.99

$1.49

$0.99

google.gemma-4-31b

131K

-

-

-

Specs as reported by each service · usage over the last 2 days · refreshed hourly

Run Gemma 4 31B in Big-AGI.

Connect your own key on any of the services above, at their rates, no markup - and run it side by side with every other model.

Launch Big-AGI

© 2026 Token Fabrics·Built with passion in San Diego