Gemma 3 12B
Text and image understanding, with a memory budget that deserves a closer look.
- Published size
- ~12B
- Active parameters
- ~12B
- Context window
- 128K
- Architecture
- Dense · vision
- License
- Gemma
Deployment considerations
The nominal 16-bit weights alone approach 24 GB. Leave room for runtime allocations, KV cache, and image processing.
This profile uses the nominal 12B model size for approximate weight arithmetic; it is not a measured checkpoint allocation.
Image inputs add processing and memory requirements. Text-only and multimodal workloads should have separate results.
Access to the upstream weights requires accepting the Gemma terms. Verify runtime support for the instruction-tuned multimodal model.
Give your model
some breathing room.
Weights meet or exceed this budget. Consider more memory or a supported lower-precision checkpoint.
Decimal GB; nominal parameter counts where marked ~. Bit widths illustrate weight storage, not validated quantizations. Understand the estimate
No invented leaderboards.
Latency, throughput, and cost per token will appear here after a reproducible run. Until then, this page helps you understand the model—not predict its performance.
Gemma 3 12B deployment FAQ
How much GPU memory does Gemma 3 12B need?
At 16-bit precision, the estimated weight storage is 24 GB. At 8-bit it is 12 GB, and at 4-bit it is 6 GB. These are theoretical weight-only estimates, excluding KV cache, runtime allocations, and quantization metadata. A working deployment needs additional memory and a supported checkpoint.
Has BenchGrid benchmarked Gemma 3 12B?
Not yet. This profile contains publisher specifications and calculated weight-memory estimates. We do not currently publish measured latency, throughput, or cost per token for this model.
Where do these specifications come from?
The specifications are based on the official Google model card linked on this page. Memory estimates use the stated total parameter count, including inactive experts for MoE models. Nominal model sizes are labeled with ~.
Explore your compute options.
Check available hardware, quotas, and current pricing with the provider. These links are not verified deployments or performance recommendations.