Skip to content

EthenEthenEthen

Open Source Model Profile · DeepPavlov

rubert-base-cased-sentence

rubert-base-cased-sentence is a DeepPavlov Russian sentence encoder initialized from RuBERT. Its card describes a 12-layer, 768-hidden cased encoder with mean-pooled sentence representations.

Publisher
DeepPavlov
Task
feature-extraction
Model type
bert
License
Unknown
Library
transformers
Publication status
Accepted · not indexed

Model overview

rubert-base-cased-sentence is published by DeepPavlov for feature-extraction in Russian. The captured configuration identifies BertModel with model type bert and the transformers library. According to the model card, it is a cased 12-layer, 768-hidden, 12-head encoder with 180M parameters for sentence representation.

Recorded capabilities

Russian cased sentence encoder

According to the model card, the encoder is Russian, cased, with 12 layers, 768 hidden size, 12 heads, and 180M parameters.

SNLI and XNLI tuning

The card states it was initialized with RuBERT and fine-tuned on SNLI translated to Russian and the Russian portion of the XNLI development set.

Mean-pooled representations

Sentence representations are mean-pooled token embeddings, described as following the Sentence-BERT approach.

Use cases in the source record

  • Russian sentence-embedding workflows that use mean-pooled token embeddings in the card’s Sentence-BERT manner.
  • Feature-extraction pipelines with the transformers stack for Russian text, matching the record’s feature-extraction and ru markers.

Limitations and unknowns

  • No license value was recorded for this model.
  • No Safetensors-derived parameter count was extracted; 180M is a publisher-reported figure.
  • No evaluation results were extracted from this record.
  • Provider state is historical snapshot data and should be refreshed before being presented as current.

Source and provenance

Source: DeepPavlov/rubert-base-cased-sentence

Captured: Unknown. Processed: 2026-09-07T19:34:30.080626+00:00.

rubert-base-cased-sentence Sentence RuBERT (Russian, cased, 12-layer, 768-hidden, 12-heads, 180M parameters) is a representation‑based sentence encoder for Russian. It is initialized with RuBERT and fine‑tuned on SNLI[1] google-translated to russian and on russian part of XNLI dev set[2]. Sentence representations are mean pooled token embeddings in the same manner as in Sentence‑BERT[3]. [1]: S. R. Bowman, G. Angeli, C. Potts, and C. D. Manning. (2015) A large annotated corpus for learning natural language inference. arXiv preprint arXiv:1508.05326 [2]: Williams A., Bowman S. (2018) XNLI: Evaluating Cross-lingual Sentence Representa…

F001F002F003F004F005F006F007F008F009