Skip to content

EthenEthenEthen

Open Source Model Profile · zai-org

GLM-4.5V-FP8

GLM-4.5V-FP8 is a 107.75B-parameter zai-org vision-language model. Its model card describes GLM-4.5V multimodal reasoning built on GLM-4.5-Air.

Publisher
zai-org
Task
image-text-to-text
Model type
glm4v_moe
License
mit
Library
transformers
Publication status
Approved for indexing

Model overview

GLM-4.5V-FP8 is published by zai-org as an image-text-to-text model. The captured configuration identifies Glm4vMoeForConditionalGeneration, and Safetensors metadata reports 107,751,931,136 parameters. According to the model card, it belongs to the GLM-4.5V vision-language line based on GLM-4.5-Air.

Recorded capabilities

Large multimodal MoE checkpoint

Captured configuration and tags identify a glm4v_moe image-text-to-text model with about 107.75B parameters and quantized base-model lineage.

Publisher-described vision reasoning

According to the model card, GLM-4.5V covers image, video, document, GUI, chart, grounding, and Thinking Mode behavior.

Transformers sample workflow

According to the model card, the publisher documents AutoProcessor and conditional-generation loading with image-plus-text prompting.

Use cases in the source record

  • Image, video, document, chart, GUI, and grounding vision-language workflows described in the model card.
  • Transformers-based image-plus-text experiments that follow the card's processor, chat-template, and generation example.

Limitations and unknowns

  • No context-window value was extracted from this record.
  • No evaluation results were extracted; the card's 42-benchmark SOTA statement is a publisher claim.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Capability, scale, and benchmark claims come from the publisher model card and were not independently verified.

Source and provenance

Source: zai-org/GLM-4.5V-FP8

Captured: Unknown. Processed: 2026-09-07T19:36:04.293951+00:00.

GLM-4.5V-FP8 👋 Join our Discord communities. 📖 Check out the paper . 📍 Access the GLM-V series models via API on the ZhipuAI Open Platform . Introduction Vision-language models (VLMs) have become a key cornerstone of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs urgently need to enhance reasoning capabilities beyond basic multimodal perception — improving accuracy, comprehensiveness, and intelligence — to enable complex problem solving, long-context understanding, and multimodal agents. Through our open-source work, we aim to explore the technological frontier together with the community while empowe…

F001F002F003F004F005F006F007F009F010F015F016F017F018F019