Skip to content

EthenEthenEthen

Open Source Model Profile · newmindai

Mursit-Large-TR-Retrieval

Mursit-Large-TR-Retrieval is a 403.63M-parameter ModernBERT Turkish embedding model from newmindai. Its model card documents retrieval tuning for Turkish legal use.

Publisher
newmindai
Task
sentence-similarity
Model type
modernbert
License
apache-2.0
Library
sentence-transformers
Publication status
Approved for indexing

Model overview

Mursit-Large-TR-Retrieval is published by newmindai as a sentence-similarity model. The captured configuration identifies ModernBertModel, and Safetensors metadata reports 403,629,056 parameters. According to the model card, it is a ModernBERT-large Turkish embedding model post-trained for legal-domain retrieval.

Recorded capabilities

Turkish legal retrieval focus

According to the model card, the model is based on ModernBERT-large with 1,024 embedding dimensions, 2,048-token context, and optimization for Turkish legal applications.

Documented large-scale training

According to the model card, pre-training used about 112.7B Turkish-dominant tokens, followed by contrastive post-training on MS MARCO-TR triplets.

Sentence-transformers and ONNX use

According to the model card, the publisher documents SentenceTransformer loading and ONNX inference with tokenizer and Hub model download.

Use cases in the source record

  • Turkish legal semantic search, retrieval, ranking, contract matching, regulation checking, and case-law discovery described in the model card.
  • Sentence-embedding and ONNX inference workflows that follow the card's SentenceTransformer and runtime examples.

Limitations and unknowns

  • No context-window value beyond the publisher's stated 2,048-token maximum was independently verified.
  • Benchmark figures and rankings are publisher claims and were not independently verified.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: newmindai/Mursit-Large-TR-Retrieval

Captured: Unknown. Processed: 2026-09-07T19:35:59.037561+00:00.

Mursit-Large-TR-Retrieval Model Description Mursit-Large-TR-Retrieval is a large-scale Turkish embedding model pre-trained entirely from scratch on Turkish-dominant corpora and fine-tuned for retrieval tasks. The model is based on ModernBERT-large architecture (403M parameters) and optimized specifically for Turkish legal domain applications. This model achieves strong performance on Turkish retrieval benchmarks with 56.87 MTEB Score and 46.56 Legal Score, ranking among the top Turkish embedding models. Key Features: Pre-trained from scratch on approximately 112.7 billion tokens of Turkish-dominant corpus Post-trained for embedding…

F001F002F003F004F005F006F007F008F009F010F012F013F015F016F017F018F020F021F022F023F024