128k training context
The model card says GLM-4.6V scales its context window to 128k tokens in training.
Open Source Model Profile · zai-org
GLM-4.6V-FP8 is a 107.75B-parameter vision-language MoE model from zai-org in the GLM-V family. Its model card documents 128k-token training context and native multimodal function calling.
GLM-4.6V-FP8 is published by zai-org as an image-text-to-text model with Glm4vMoeForConditionalGeneration architecture and a glm4v_moe model type. Safetensors metadata reports 107,751,931,136 parameters, or about 107.75B. According to the model card, it belongs to the GLM-V multimodal family, and captured metadata records an MIT license with Transformers compatibility.
The model card says GLM-4.6V scales its context window to 128k tokens in training.
The card describes native vision-driven tool use where images and document pages pass directly as tool inputs.
The card documents AutoProcessor chat-template inference with vLLM or SGLang and published decoding parameters.
Source: zai-org/GLM-4.6V-FP8
Captured: Unknown. Processed: 2026-09-07T19:36:04.375561+00:00.
GLM-4.6V This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning . GLM-4.6V Blog : https://z.ai/blog/glm-4.6v Paper : https://huggingface.co/papers/2507.01006 GitHub Repository : https://github.com/zai-org/GLM-V Online Demo : https://chat.z.ai/ API Access : Z.ai Open Platform Desktop Assistant App : https://huggingface.co/spaces/zai-org/GLM-4.5V-Demo-App Introduction GLM-4.6V series model includes two versions: GLM-4.6V (106B), a foundation model designed for cloud and high-performance cluster scenarios, and…
F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017