Skip to content

EthenEthenEthen

Open Source Model Profile · unsloth

gemma-3-4b-it

gemma-3-4b-it is a 4.30B-parameter Gemma 3 multimodal model from unsloth. Its model card documents image-text input, 128K context, and google/gemma-3-4b-it lineage.

Publisher
unsloth
Task
image-text-to-text
Model type
gemma3
License
gemma
Library
transformers
Publication status
Accepted · not indexed

Model overview

gemma-3-4b-it is published by unsloth as a Gemma 3 image-text-to-text model. The captured configuration identifies Gemma3ForConditionalGeneration and Safetensors metadata reports 4,300,079,472 parameters. According to the model card, it is a multimodal instruction-tuned variant handling text and image input, with hub tags pointing to google/gemma-3-4b-it as base.

Recorded capabilities

Multimodal 128K context

According to the model card, the 4B model handles text plus 896x896 images with 128K total input tokens and 8,192 output tokens.

Google Gemma 3 lineage tags

Hub tags list google/gemma-3-4b-it as base model and fine-tune, and the card attributes authorship to Google DeepMind.

Pipeline and GPU usage

According to the model card, inference can use the pipeline API or Transformers processor and model classes on single or multi-GPU setups.

Filtered multilingual training

According to the model card, training covered more than 140 languages with CSAM and sensitive-data filtering.

Documented accuracy limits

According to the model card, the model may produce incorrect or outdated facts and reflect socio-cultural biases in training material.

Use cases in the source record

  • Image-grounded question answering and document analysis using the documented text-plus-image input pattern.
  • Text-generation tasks such as summarization and reasoning described in the Gemma 3 card.
  • Pipeline and GPU-based inference experiments following the card's processor and model-loading pattern.

Limitations and unknowns

  • No independent evaluation results were extracted from this record.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Capability, language, and training-scale statements come from the publisher card and have not been independently verified by Ethen.

Source and provenance

Source: unsloth/gemma-3-4b-it

Captured: Unknown. Processed: 2026-09-07T19:35:31.928318+00:00.

Gemma 3 model card Model Page : Gemma Resources and Technical Documentation : Gemma 3 Technical Report Responsible Generative AI Toolkit Gemma on Kaggle Gemma on Vertex Model Garden Terms of Use : Terms Authors : Google DeepMind Model Information Summary description and brief definition of inputs and outputs. Description Gemma is a family of lightweight, state-of-the-art open models from Google, built from the same research and technology used to create the Gemini models. Gemma 3 models are multimodal, handling text and image input and generating text output, with open weights for both pre-trained variants and instruction-tuned vari…

F001F002F003F004F005F006F007F009F010F013F014F015F016F018F021F022F035F036