Skip to content

EthenEthenEthen

Open Source Model Profile · google

gemma-3-4b-it

gemma-3-4b-it is a 4.3B-parameter multimodal instruction-tuned model from Google. According to the model card, Gemma 3 handles text and image input with text output, and the 4B size carries a 128K-token input context.

Publisher
google
Task
image-text-to-text
Model type
gemma3
License
gemma
Library
transformers
Publication status
Accepted · not indexed

Model overview

gemma-3-4b-it is published by Google as an instruction-tuned Gemma 3 image-text-to-text model. The captured configuration identifies Gemma3ForConditionalGeneration with model type gemma3, and Safetensors metadata reports 4,300,079,472 parameters. According to the model card, Gemma 3 ships open weights for pre-trained and instruction-tuned variants and suits question answering, summarization, and reasoning over text and images.

Recorded capabilities

Multimodal text and image input

According to the model card, the model takes text plus images normalized to 896x896 and encoded to 256 tokens each, and generates text output.

128K input with 8192-token output

According to the model card, the 4B, 12B, and 27B sizes carry 128K tokens of total input context with 8192 tokens of output context.

Captured 4.3B Transformers weights

Safetensors metadata reports 4,300,079,472 parameters, and the hub lists Transformers support.

4-trillion-token training scale

According to the model card, the 4B model was trained on 4 trillion tokens spanning web documents in over 140 languages plus mathematics content.

TPU training with JAX

According to the model card, training used TPU hardware (TPUv4p, TPUv5p, TPUv5e) with JAX and ML Pathways.

Use cases in the source record

  • Question answering, summarization, and reasoning over text and image inputs within the documented 128K-token input and 8192-token output budget.
  • Resource-constrained deployment on laptops, desktops, or private cloud infrastructure, which the publisher describes as feasible given the model's size.
  • Pipeline-API or single and multi-GPU inference runs using the model card's documented Transformers code paths.

Limitations and unknowns

  • According to the model card, responses may reflect training-data gaps and biases and may contain incorrect or outdated factual statements.
  • Evaluation detail in this record is limited to the publisher's assurance-evaluation process description rather than extracted scores.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: google/gemma-3-4b-it

Captured: Unknown. Processed: 2026-09-07T19:34:45.734000+00:00.

Gemma 3 model card Model Page : Gemma Resources and Technical Documentation : Gemma 3 Technical Report Responsible Generative AI Toolkit Gemma on Kaggle Gemma on Vertex Model Garden Terms of Use : Terms Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned vari…

F001F002F003F004F005F006F007F008F010F011F013F014F015F016F018F023F027F030F031F033F034F035