Skip to content

EthenEthenEthen

Open Source Model Profile · Qwen

Qwen3.5-4B

Qwen3.5-4B is a 4.66B-parameter Qwen vision-language model for image-text-to-text work. Its model card documents 262K native context and vLLM, SGLang, and KTransformers serving paths.

Publisher
Qwen
Task
image-text-to-text
Model type
qwen3_5
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

Qwen3.5-4B is published by Qwen as an image-text-to-text model. The captured configuration identifies Qwen3_5ForConditionalGeneration with model type qwen3_5, and Safetensors metadata reports 4,659,865,088 parameters. According to the model card, this repository holds the post-trained weights, and hub tags record Qwen/Qwen3.5-4B-Base as the base.

Recorded capabilities

Post-trained multimodal weights

According to the model card, this repository holds the post-trained model compatible with Transformers, vLLM, SGLang, and KTransformers.

Vision-language foundation

The model card describes a causal language model with vision encoder, early-fusion multimodal training, 32 layers, and 262,144-token native context extensible to 1,010,000 tokens.

Hybrid attention layout

According to the model card, Gated DeltaNet blocks combine with sparse mixture-of-experts and gated attention, with multi-step multi-token prediction training.

Documented serving and tool use

The model card documents SGLang and vLLM launch commands, thinking and instruct sampling presets, and tool-calling support.

Use cases in the source record

  • Vision-language conversational and tool-calling work through API serving with the documented inference frameworks.
  • Self-hosted SGLang or vLLM deployment using the publisher's launch commands and thinking or instruct sampling presets.

Limitations and unknowns

  • No VRAM or hardware requirement was extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.
  • Benchmark tables and page-listed evaluation entries are publisher- and page-reported values and were not independently verified by Ethen.

Source and provenance

Source: Qwen/Qwen3.5-4B

Captured: Unknown. Processed: 2026-09-07T19:34:36.382084+00:00.

Qwen3.5-4B This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Over recent months, we have intensified our focus on developing foundation models that deliver exceptional utility and performance. Qwen3.5 represents a significant leap forward, integrating breakthroughs in multimodal learning, architectural efficiency, reinforcement learning scale, and global accessibility to empower developers and enterprises with unprecedented capability and efficiency. Qwen3.5…

F001F002F003F004F005F006F007F009F010F011F012F014F015F017F018F019F021F023F024F026F034