Multimodal input and context
According to the model card, Gemma 3 handles text and image input with 128K context for the 12B size and 8192 output tokens.
Open Source Model Profile · unsloth
gemma-3-12b-it is a 12.19B-parameter Gemma 3 multimodal model from unsloth for text and image input with text output.
gemma-3-12b-it is published by unsloth as an image-text-to-text model. Safetensors metadata reports 12,187,325,040 parameters, with Gemma3ForConditionalGeneration and model type gemma3. According to the model card, it follows the Gemma 3 multimodal design for text and image input.
According to the model card, Gemma 3 handles text and image input with 128K context for the 12B size and 8192 output tokens.
The model card documents pipeline-API initialization and single- or multi-GPU use with AutoProcessor and Gemma3ForConditionalGeneration.
The model card says the 12B model was trained on 12 trillion tokens using TPU hardware with JAX and ML Pathways.
Source: unsloth/gemma-3-12b-it
Captured: Unknown. Processed: 2026-09-07T19:35:32.137618+00:00.
Gemma 3 model card Model Page : Gemma Resources and Technical Documentation : Gemma 3 Technical Report Responsible Generative AI Toolkit Gemma on Kaggle Gemma on Vertex Model Garden Terms of Use : Terms Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned vari…
F001F002F003F004F005F006F007F009F010F011F013F014F015F016F018F023F027F031F033F035F036