Large multimodal MoE checkpoint
Captured configuration and tags identify a glm4v_moe image-text-to-text model with about 107.75B parameters and quantized base-model lineage.
Open Source Model Profile · zai-org
GLM-4.5V-FP8 is a 107.75B-parameter zai-org vision-language model. Its model card describes GLM-4.5V multimodal reasoning built on GLM-4.5-Air.
GLM-4.5V-FP8 is published by zai-org as an image-text-to-text model. The captured configuration identifies Glm4vMoeForConditionalGeneration, and Safetensors metadata reports 107,751,931,136 parameters. According to the model card, it belongs to the GLM-4.5V vision-language line based on GLM-4.5-Air.
Captured configuration and tags identify a glm4v_moe image-text-to-text model with about 107.75B parameters and quantized base-model lineage.
According to the model card, GLM-4.5V covers image, video, document, GUI, chart, grounding, and Thinking Mode behavior.
According to the model card, the publisher documents AutoProcessor and conditional-generation loading with image-plus-text prompting.
Source: zai-org/GLM-4.5V-FP8
Captured: Unknown. Processed: 2026-09-07T19:36:04.293951+00:00.
GLM-4.5V-FP8 👋 Join our Discord communities. 📖 Check out the paper . 📍 Access the GLM-V series models via API on the ZhipuAI Open Platform . Introduction Vision-language models (VLMs) have become a key cornerstone of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs urgently need to enhance reasoning capabilities beyond basic multimodal perception — improving accuracy, comprehensiveness, and intelligence — to enable complex problem solving, long-context understanding, and multimodal agents. Through our open-source work, we aim to explore the technological frontier together with the community while empowe…
F001F002F003F004F005F006F007F009F010F015F016F017F018F019