Multimodal text and image input
According to the model card, the model takes text plus images normalized to 896x896 and encoded to 256 tokens each, and generates text output.
Open Source Model Profile · google
gemma-3-4b-it is a 4.3B-parameter multimodal instruction-tuned model from Google. According to the model card, Gemma 3 handles text and image input with text output, and the 4B size carries a 128K-token input context.
gemma-3-4b-it is published by Google as an instruction-tuned Gemma 3 image-text-to-text model. The captured configuration identifies Gemma3ForConditionalGeneration with model type gemma3, and Safetensors metadata reports 4,300,079,472 parameters. According to the model card, Gemma 3 ships open weights for pre-trained and instruction-tuned variants and suits question answering, summarization, and reasoning over text and images.
According to the model card, the model takes text plus images normalized to 896x896 and encoded to 256 tokens each, and generates text output.
According to the model card, the 4B, 12B, and 27B sizes carry 128K tokens of total input context with 8192 tokens of output context.
Safetensors metadata reports 4,300,079,472 parameters, and the hub lists Transformers support.
According to the model card, the 4B model was trained on 4 trillion tokens spanning web documents in over 140 languages plus mathematics content.
According to the model card, training used TPU hardware (TPUv4p, TPUv5p, TPUv5e) with JAX and ML Pathways.
Source: google/gemma-3-4b-it
Captured: Unknown. Processed: 2026-09-07T19:34:45.734000+00:00.
Gemma 3 model card Model Page : Gemma Resources and Technical Documentation : Gemma 3 Technical Report Responsible Generative AI Toolkit Gemma on Kaggle Gemma on Vertex Model Garden Terms of Use : Terms Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned vari…
F001F002F003F004F005F006F007F008F010F011F013F014F015F016F018F023F027F030F031F033F034F035