Skip to content

EthenEthenEthen

Open Source Model Profile · deepvk

USER2-base

USER2-base is a 149M-parameter Russian sentence encoder from deepvk. According to the model card, it supports long-context retrieval and semantic tasks.

Publisher
deepvk
Task
sentence-similarity
Model type
modernbert
License
apache-2.0
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

USER2-base is published by deepvk as a sentence-similarity model. The captured configuration identifies ModernBertModel with model type modernbert, and Safetensors metadata reports 149,014,272 parameters. According to the model card, it is a Russian universal sentence encoder built on RuModernBERT and tuned for retrieval and semantic tasks.

Recorded capabilities

Russian long-context encoder

According to the model card, the RuModernBERT-based encoder supports sentence representation with context up to 8,192 tokens.

Matryoshka truncation

According to the model card, Matryoshka Representation Learning allows smaller embedding sizes with limited quality loss, configured through truncate_dim.

Prefix-guided inputs

According to the model card, the model expects task-specific prefixes, with "classification: " described as the default universal choice.

Use cases in the source record

  • Russian retrieval and semantic search with long-context inputs up to the documented 8,192-token support.
  • Truncated-embedding deployments where Matryoshka representations and task prefixes balance size and task fit.

Limitations and unknowns

  • No context-window value was extracted from structured metadata; the 8,192-token figure is a publisher model-card claim.
  • No evaluation scores were extracted from this record; the card points to MTEB-rus stage comparisons without extracted figures.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: deepvk/USER2-base

Captured: Unknown. Processed: 2026-09-07T19:35:20.859377+00:00.

USER2-base USER2 is a new generation of the U niversal S entence E ncoder for R ussian, designed for sentence representation with long-context support of up to 8,192 tokens. The models are built on top of the RuModernBERT encoders and are fine-tuned for retrieval and semantic tasks. They also support Matryoshka Representation Learning (MRL) — a technique that enables reducing embedding size with minimal loss in representation quality. This is a base model with 149 million parameters. Model Size Context Length Hidden Dim MRL Dims deepvk/USER2-small 34M 8192 384 [32, 64, 128, 256, 384] deepvk/USER2-base 149M 8192 768 [32, 64, 128, 256…

F001F002F003F004F005F006F007F010F011F012F013F015F016F017