Skip to content

EthenEthenEthen

Open Source Model Profile · snap-stanford

humanlm-opinion

HumanLM-Opinion is an 8.19B-parameter Qwen3 text-generation model from snap-stanford. Its model card describes HumanLM as a user simulator trained on the Humanual-Opinion benchmark of Reddit opinionated replies.

Publisher
snap-stanford
Task
text-generation
Model type
qwen3
License
apache-2.0
Library
transformers
Publication status
Approved for indexing

Model overview

snap-stanford publishes humanlm-opinion as a transformers text-generation model. Captured config identifies Qwen3ForCausalLM with a qwen3 model type, and Safetensors metadata reports 8,190,735,360 parameters. Card data records apache-2.0. Hub tags include user-simulation, persona, grpo, reinforcement-learning, state-alignment, humanlm, and Qwen/Qwen3-8B as base model and finetune. The model card says this checkpoint is trained on Humanual-Opinion.

Recorded capabilities

About 8.19B parameters

Safetensors metadata reports 8,190,735,360 parameters, or about 8.19B.

Apache-2.0 licensing

Captured metadata records an apache-2.0 license.

Publisher user-simulator role

The model card says HumanLM generates responses that capture underlying user states, and that this checkpoint was trained on Humanual-Opinion Reddit threads.

Publisher GRPO note

The card lists base model Qwen3-8B, training method GRPO with state alignment, and training data of 4.6k Reddit users and 46k responses across 1k threads.

Use cases in the source record

  • User-simulation research that, according to the card, models beliefs, emotions, stance, values, goals, and communication style.
  • Publisher-stated intended uses including user research, content testing, AI alignment feedback, and social simulation of opinion dynamics.

Limitations and unknowns

  • No independently verified evaluation results were extracted; user-study win rates, naturalness ratings, and Azure harm scores come from the publisher model card.
  • No context-window value was extracted from configuration metadata.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • Qwen3-8B lineage, GRPO method, Humanual-Opinion dataset counts, and intended-use statements come from the publisher model card and hub tags, not an independently verified Ethen lineage judgement.

Source and provenance

Source: snap-stanford/humanlm-opinion

Captured: Unknown. Processed: 2026-09-07T19:36:02.192968+00:00.

HumanLM-Opinion HumanLM is a user simulator that generates responses capturing the underlying states of real users (beliefs, emotions, stance, values, goals, communication style). This checkpoint is trained on the Humanual-Opinion benchmark, which contains Reddit users’ opinionated responses in personal-issue discussion threads. 📄 Paper: HumanLM: Simulating Users with State Alignment Beats Response Imitation 🌐 Project Page: humanlm.stanford.edu Model Details Base Model: Qwen3-8B Training Method: GRPO (Group Relative Policy Optimization) with state alignment Training Data: Humanual-Opinion (4.6k Reddit users, 46k responses across 1…

F001F002F003F004F005F006F007F008F009F010F011F012F014F015F017F018F019