Skip to content

EthenEthenEthen

Open Source Model Profile · ipipan

silver-retriever-base-v1.1

silver-retriever-base-v1.1 is a 124.4M-parameter BERT sentence-similarity model from ipipan. Its model card says it encodes Polish sentences or paragraphs into a 768-dimensional dense vector space for document retrieval and semantic search.

Publisher
ipipan
Task
sentence-similarity
Model type
bert
License
cc-by-sa-4.0
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

ipipan publishes silver-retriever-base-v1.1 as a sentence-transformers sentence-similarity model. Captured config identifies BertModel with a bert model type, and Safetensors metadata reports 124,443,394 parameters. Card data records cc-by-sa-4.0. Hub tags include pl, dataset:ipipan/polqa, dataset:ipipan/maupqa, and arxiv:2309.08469. The model card says the checkpoint was initialized from HerBERT-base and fine-tuned on PolQA and MAUPQA.

Recorded capabilities

About 124.4M parameters

Safetensors metadata reports 124,443,394 parameters, or about 124.4M.

CC-BY-SA-4.0 licensing

Captured metadata records a cc-by-sa-4.0 license. The model card also states CC BY-SA 4.0.

768-d Polish embeddings

The model card says the model encodes Polish sentences or paragraphs into a 768-dimensional dense vector space.

Documented input format

The card says questions were prefixed with Pytanie: and passages used title and text joined with </s> . It also says cosine distance usually works better than the dot product used in training.

Use cases in the source record

  • Polish document retrieval and semantic search using the card's 768-dimensional embeddings.
  • Question-passage encoding that prefixes questions with Pytanie: and concatenates passage title and text with </s> , as documented on the card.

Limitations and unknowns

  • No evaluation scores were extracted; captured card fragments include PolQA/NDCG table headers without usable metric values.
  • Provider state is historical snapshot data, not independently refreshed current availability.
  • HerBERT-base initialization, PolQA/MAUPQA training, 768-d embeddings, input-format rules, and max_seq_length 512 come from the publisher model card and have not been independently verified by Ethen.

Source and provenance

Source: ipipan/silver-retriever-base-v1.1

Captured: Unknown. Processed: 2026-09-07T19:34:47.040668+00:00.

Silver Retriever Base (v1.1) Silver Retriever model encodes the Polish sentences or paragraphs into a 768-dimensional dense vector space and can be used for tasks like document retrieval or semantic search. It was initialized from the HerBERT-base model and fine-tuned on the PolQA and MAUPQA datasets for 8,000 steps with a batch size of 8,192. Please refer to the SilverRetriever: Advancing Neural Passage Retrieval for Polish Question Answering for more details. Evaluation Model Average [Acc] Average [NDCG] PolQA [Acc] PolQA [NDCG] Allegro FAQ [Acc] Allegro FAQ [NDCG] Legal Questions [Acc] Legal Questions [NDCG] BM25 74.87 51.81 61.3…

F001F002F003F004F005F006F007F008F009F010F011F012F013F014F015F016F017F021F022