About 8.19B parameters
Safetensors metadata reports 8,190,735,360 parameters, or about 8.19B.
Open Source Model Profile · snap-stanford
HumanLM-Opinion is an 8.19B-parameter Qwen3 text-generation model from snap-stanford. Its model card describes HumanLM as a user simulator trained on the Humanual-Opinion benchmark of Reddit opinionated replies.
snap-stanford publishes humanlm-opinion as a transformers text-generation model. Captured config identifies Qwen3ForCausalLM with a qwen3 model type, and Safetensors metadata reports 8,190,735,360 parameters. Card data records apache-2.0. Hub tags include user-simulation, persona, grpo, reinforcement-learning, state-alignment, humanlm, and Qwen/Qwen3-8B as base model and finetune. The model card says this checkpoint is trained on Humanual-Opinion.
Safetensors metadata reports 8,190,735,360 parameters, or about 8.19B.
Captured metadata records an apache-2.0 license.
The model card says HumanLM generates responses that capture underlying user states, and that this checkpoint was trained on Humanual-Opinion Reddit threads.
The card lists base model Qwen3-8B, training method GRPO with state alignment, and training data of 4.6k Reddit users and 46k responses across 1k threads.
Source: snap-stanford/humanlm-opinion
Captured: Unknown. Processed: 2026-09-07T19:36:02.192968+00:00.
HumanLM-Opinion HumanLM is a user simulator that generates responses capturing the underlying states of real users (beliefs, emotions, stance, values, goals, communication style). This checkpoint is trained on the Humanual-Opinion benchmark, which contains Reddit users’ opinionated responses in personal-issue discussion threads. 📄 Paper: HumanLM: Simulating Users with State Alignment Beats Response Imitation 🌐 Project Page: humanlm.stanford.edu Model Details Base Model: Qwen3-8B Training Method: GRPO (Group Relative Policy Optimization) with state alignment Training Data: Humanual-Opinion (4.6k Reddit users, 46k responses across 1…
F001F002F003F004F005F006F007F008F009F010F011F012F014F015F017F018F019