Skip to content

EthenEthenEthen

Open Source Model Profile · sergeyzh

BERTA

BERTA is a 128.3M-parameter BERT embedding model from sergeyzh. According to the model card, it distills ai-forever/FRIDA into sergeyzh/LaBSE-ru-turbo for Russian and English sentence embeddings with mean pooling.

Publisher
sergeyzh
Task
sentence-similarity
Model type
bert
License
mit
Library
sentence-transformers
Publication status
Accepted · not indexed

Model overview

BERTA is published by sergeyzh as a sentence-similarity embedding model. The captured configuration identifies BertModel with a bert model type, and Safetensors metadata reports 128345088 parameters. According to the model card, it was produced by distilling 1536-dimensional, 24-layer FRIDA embeddings into 768-dimensional, 12-layer LaBSE-ru-turbo, replacing CLS pooling with mean pooling.

Recorded capabilities

FRIDA-to-LaBSE distillation

According to the model card, FRIDA embeddings were distilled into sergeyzh/LaBSE-ru-turbo without other behavioral changes, covering Russian and English sentences and prefix behavior.

512-token context

According to the model card, the model context matches FRIDA at 512 tokens.

Prefix system with default

The publisher says all prefixes are inherited from FRIDA, with categorize_entailment set as default and per-prefix encodechka scores published for STS, PI, NLI, SA, and TI.

Published similarity examples

According to the model card, example paraphrase, entailment, and query-document pairs report dot-product scores alongside FRIDA comparisons.

Use cases in the source record

  • Russian and English sentence-embedding work for similarity, paraphrase, inference, sentiment, and toxicity tasks with task prefixes.
  • Retrieval experiments that select FRIDA-inherited prefixes such as search_query, search_document, paraphrase, or categorize_entailment.

Limitations and unknowns

  • Distillation and prefix-score details come from a Russian-language publisher card and have not been independently verified by Ethen.
  • No local hardware requirement is provided in the current evidence.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: sergeyzh/BERTA

Captured: Unknown. Processed: 2026-09-07T19:35:30.073201+00:00.

BERTA Модель для расчетов эмбеддингов предложений на русском и английском языках получена методом дистилляции эмбеддингов ai-forever/FRIDA (размер эмбеддингов - 1536, слоёв - 24) в sergeyzh/LaBSE-ru-turbo (размер эмбеддингов - 768, слоёв - 12). Основной режим использования FRIDA - CLS pooling заменен на mean pooling. Каких-либо других изменений поведения модели не производилось. Дистиляция выполнена в максимально возможном объеме - эмбеддинги русских и английских предложений, работа префиксов. Размер контекста модели соответствует FRIDA - 512 токенов. Префиксы Все префиксы унаследованы от FRIDA. Оптимальный (обеспечивающий средние р…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015