About 149.6M parameters
Safetensors metadata reports 149,605,633 parameters, or about 149.6M. The model card also describes a ModernBERT-base design with 149M parameters.
Open Source Model Profile · dleemiller
ModernCE-base-sts is a 149.6M-parameter ModernBERT text-classification model from dleemiller. Its model card describes it as a semantic-similarity cross-encoder that compares two texts and outputs a 0-1 score.
dleemiller publishes ModernCE-base-sts as a sentence-transformers text-classification model. Captured config identifies ModernBertForSequenceClassification with a modernbert model type, and Safetensors metadata reports 149,605,633 parameters. Card data records mit. Hub tags include cross-encoder, sts, stsb, stsbenchmark-sts, en, dataset:dleemiller/wiki-sim, dataset:sentence-transformers/stsb, and answerdotai/ModernBERT-base as base model. The card says the checkpoint was pretrained on wiki-sim and fine-tuned on stsb.
Safetensors metadata reports 149,605,633 parameters, or about 149.6M. The model card also describes a ModernBERT-base design with 149M parameters.
Captured metadata records an mit license. The model card also says the model is licensed under the MIT License.
The model card describes a cross-encoder for semantic similarity that outputs a 0-1 score.
The card reports STS-B test Pearson 0.9162 and Spearman 0.9122 for this checkpoint. Those figures are publisher-reported.
Source: dleemiller/ModernCE-base-sts
Captured: Unknown. Processed: 2026-09-07T19:35:21.128762+00:00.
ModernBERT Cross-Encoder: Semantic Similarity (STS) Cross encoders are high performing encoder models that compare two texts and output a 0-1 score. I've found the cross-encoders/roberta-large-stsb model to be very useful in creating evaluators for LLM outputs. They're simple to use, fast and very accurate. Like many people, I was excited about the architecture and training uplift from the ModernBERT architecture ( answerdotai/ModernBERT-base ). So I've applied it to the stsb cross encoder, which is a very handy model. Additionally, I've added pretraining from a much larger semi-synthetic dataset dleemiller/wiki-sim that targets thi…
F001F002F003F004F005F006F007F008F009F010F011F012F016F017F018F019F022F023