Vision-language instruction model
According to the model card, this weight repository covers Qwen3-VL-8B-Instruct with upgrades across text understanding, visual perception, and agent interaction.
Open Source Model Profile · Qwen
Qwen3-VL-8B-Instruct is a vision-language instruction model from Qwen. The captured record reports about 8.77B parameters under apache-2.0 for image-text-to-text use.
Qwen3-VL-8B-Instruct is published by Qwen as an image-text-to-text model. The captured configuration identifies Qwen3VLForConditionalGeneration with model type qwen3_vl, and Safetensors metadata reports about 8.77B parameters under apache-2.0. According to the model card, this is the weight repository for Qwen3-VL-8B-Instruct.
According to the model card, this weight repository covers Qwen3-VL-8B-Instruct with upgrades across text understanding, visual perception, and agent interaction.
According to the model card, the visual-agent capability covers recognizing interface elements, understanding functions, invoking tools, and completing tasks.
According to the model card, Interleaved-MRoPE allocates frequency over time, width, and height to support long-horizon video reasoning.
According to the model card, inference uses chat-template processing with up to 128 new tokens and separate vision-language and text generation hyperparameters.
Source: Qwen/Qwen3-VL-8B-Instruct
Captured: Unknown. Processed: 2026-09-07T19:34:36.123433+00:00.
Qwen3-VL-8B-Instruct Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date. This generation delivers comprehensive upgrades across the board: superior text understanding & generation, deeper visual perception & reasoning, extended context length, enhanced spatial and video dynamics comprehension, and stronger agent interaction capabilities. Available in Dense and MoE architectures that scale from edge to cloud, with Instruct and reasoning‑enhanced Thinking editions for flexible, on‑demand deployment. Key Enhancements: Visual Agent : Operates PC/mobile GUIs—recognizes elements, understands functions, invo…
F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017