Multimodal text-image input
The model card describes handling text and image input with text output, including 896 x 896 image normalization at 256 tokens per image.
Open Source Model Profile · RedHatAI
gemma-3-12b-it is a 12.19B-parameter Gemma3 multimodal model from RedHatAI. Its card documents text-image input, 128k context, and over 140 supported languages.
gemma-3-12b-it is published by RedHatAI as an image-text-to-text model with Gemma3ForConditionalGeneration architecture. Safetensors metadata reports 12,187,325,040 parameters, or about 12.19B, and hub tags record a google/gemma-3-12b-pt base. The model card describes the Gemma 3 family as multimodal text-and-image input models with 128k context and broad multilingual coverage.
The model card describes handling text and image input with text output, including 896 x 896 image normalization at 256 tokens per image.
According to the model card, the 12B size uses 128k tokens of total input context with 8192 tokens of output context.
The model card states multilingual support in over 140 languages and training data covering more than 140 languages, with the 12B model trained on 12 trillion tokens.
Source: RedHatAI/gemma-3-12b-it
Captured: Unknown. Processed: 2026-09-07T19:35:51.336566+00:00.
Gemma 3 model card Model Page : Gemma Resources and Technical Documentation : Gemma 3 Technical Report Responsible Generative AI Toolkit Gemma on Kaggle Gemma on Vertex Model Garden Terms of Use : Terms Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned vari…
F001F002F003F004F005F006F007F009F010F013F014F018F031F032F033F035