Skip to content

EthenEthenEthen

Open Source Model Profile · ai-forever

pollux-judge-7b

pollux-judge-7b is a 7.61B-parameter Qwen2-family text-generation fine-tune from ai-forever. According to the model card, it judges Russian-language responses.

Publisher
ai-forever
Task
text-generation
Model type
qwen2
License
mit
Library
transformers
Publication status
Accepted · not indexed

Model overview

pollux-judge-7b is published by ai-forever as a text-generation model. The captured configuration identifies Qwen2ForCausalLM with model type qwen2, and Safetensors metadata reports about 7.61B parameters. According to the model card, it is a POLLUX-project judge for evaluating Russian-language model responses.

Recorded capabilities

Qwen2 text-generation architecture

The captured configuration identifies Qwen2ForCausalLM with model type qwen2 and Transformers support.

Russian LLM-judge purpose

According to the model card, the model evaluates the quality of other models' Russian-language responses against an instruction and explicit criteria.

T-lite-it-1.0 lineage

According to the model card and hub tags, the model was finetuned from t-tech/T-lite-it-1.0.

Documented POLLUX evaluation design

According to the model card, testing uses the POLLUX dataset for in- and out-of-domain checks with Spearman correlation, MAE, and Verdict Confidence.

Use cases in the source record

  • Single-criterion automated scoring of Russian responses against a supplied instruction, rubric, and criterion.
  • In- and out-of-domain judge experiments using the publisher-described POLLUX task and criteria taxonomies.

Limitations and unknowns

  • No Ethen-measured benchmark scores were extracted; card tables are publisher-reported comparisons.
  • According to the model card, multi-criterion use and autonomous criterion selection fall outside the intended single-criterion design and may yield unpredictable results.
  • According to the model card, outputs reflect statistical patterns from pre-training data and may carry over underlying dataset patterns.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: ai-forever/pollux-judge-7b

Captured: Unknown. Processed: 2026-09-07T19:35:53.325656+00:00.

pollux-judge-7b pollux-judge-7b is a 7-billion parameter generative language model specifically designed to evaluate the quality of other language models' responses in Russian. The model assesses answer quality given input instruction, specific criteria and rubrics, providing automated LLM performance evaluation for Russian-language tasks. Model Details Model Description pollux-judge-7b is an integral component of the POLLUX project, a comprehensive initiative dedicated to evaluating the generative capabilities of Large Language Models (LLMs). At the heart of this project lies the POLLUX dataset , which introduces systematic taxonom…

F001F002F003F004F005F006F007F008F010F011F012F013F015F017F018F019F020F021F030F031F032