distilroberta-base rejection lineage
According to the model card, it is a fine-tuned version of distilroberta-base on combined rejection and normal RLHF response datasets.
Open Source Model Profile · protectai
distilroberta-base-rejection-v1 is an 82.12M-parameter Roberta-family text-classification fine-tune from protectai. According to the model card, it detects LLM rejections as binary 0 normal versus 1 rejection labels.
distilroberta-base-rejection-v1 is published by protectai as a Roberta-based text-classification model. The captured configuration identifies RobertaForSequenceClassification and Safetensors metadata reports 82119938 parameters. According to the model card, it is a fine-tuned version of distilroberta-base for identifying LLM rejections, with card data recording apache-2.0.
According to the model card, it is a fine-tuned version of distilroberta-base on combined rejection and normal RLHF response datasets.
According to the model card, it classifies inputs as 0 for normal outputs and 1 for rejection detected.
According to the model card, it can run through a text-classification pipeline with truncation and max length 512, including use in the LLM Guard NoRefusal Scanner.
According to the model card, training used about 10% rejections and 90% normal outputs, with 3 epochs and a reported final accuracy of about 0.9939.
Card data records apache-2.0 for this repository.
Source: protectai/distilroberta-base-rejection-v1
Captured: Unknown. Processed: 2026-09-07T19:34:56.035961+00:00.
THIS PROJECT HAS BEEN ARCHIVED. This project and its associated code on GitHub are no longer under active development or maintained. Model Card for distilroberta-base-rejection-v1 This model is a fine-tuned version of distilroberta-base on multiple combined datasets of rejections from different LLMs and normal responses from RLHF datasets. It aims to identify rejections in LLMs when the prompt doesn't pass content moderation, classifying inputs into two categories: 0 for normal outputs and 1 for rejection detected. It achieves the following results on the evaluation set: Loss: 0.0544 Accuracy: 0.9887 Recall: 0.9810 Precision: 0.9279…
F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017F018F019