<- All Models
Google Gemma 4 31B on Cerebras - first multimodal model on wafer-scale inference (~1,850 tok/s). Vision (base64 PNG/JPEG, max 10 images / 10MB), function calling, reasoning (off by default, enable via effort). 131K context (65K free tier), 40K max output.
Available on
Where Gemma 4 31B runs in Big-AGI, with each service's own model ID and USD-per-1M-token rates. The summary above prefers the creator's claim; rates below are each service's own.
Connect your own key on any of the services above, at their rates, no markup - and run it side by side with every other model.
Launch Big-AGIBIG-AGI
Resources
© 2026 Token Fabrics·Built with passion in San Diego