Skip to content

EthenEthenEthen

Open Source Model Profile · protectai

distilroberta-base-rejection-v1

distilroberta-base-rejection-v1 is an 82.12M-parameter Roberta-family text-classification fine-tune from protectai. According to the model card, it detects LLM rejections as binary 0 normal versus 1 rejection labels.

Publisher
protectai
Task
text-classification
Model type
roberta
License
apache-2.0
Library
transformers
Publication status
Accepted · not indexed

Model overview

distilroberta-base-rejection-v1 is published by protectai as a Roberta-based text-classification model. The captured configuration identifies RobertaForSequenceClassification and Safetensors metadata reports 82119938 parameters. According to the model card, it is a fine-tuned version of distilroberta-base for identifying LLM rejections, with card data recording apache-2.0.

Recorded capabilities

distilroberta-base rejection lineage

According to the model card, it is a fine-tuned version of distilroberta-base on combined rejection and normal RLHF response datasets.

Binary rejection classification

According to the model card, it classifies inputs as 0 for normal outputs and 1 for rejection detected.

Documented pipeline and LLM Guard use

According to the model card, it can run through a text-classification pipeline with truncation and max length 512, including use in the LLM Guard NoRefusal Scanner.

Reported training and results

According to the model card, training used about 10% rejections and 90% normal outputs, with 3 epochs and a reported final accuracy of about 0.9939.

Apache-2.0 licensing

Card data records apache-2.0 for this repository.

Use cases in the source record

  • Rejection screening that flags LLM outputs as normal or rejected for moderation-failure review.
  • LLM Guard NoRefusal Scanner workflows where the publisher describes detecting rejected output as a prompt-health signal.

Limitations and unknowns

  • According to the model card, performance depends on training-data nature and quality and may not transfer to unseen styles or topics.
  • According to the model card, the project has been archived and is no longer actively maintained.
  • No independent evaluation results were extracted beyond the publisher-reported training table.
  • Provider state is historical snapshot data, not independently refreshed current availability.

Source and provenance

Source: protectai/distilroberta-base-rejection-v1

Captured: Unknown. Processed: 2026-09-07T19:34:56.035961+00:00.

THIS PROJECT HAS BEEN ARCHIVED. This project and its associated code on GitHub are no longer under active development or maintained. Model Card for distilroberta-base-rejection-v1 This model is a fine-tuned version of distilroberta-base on multiple combined datasets of rejections from different LLMs and normal responses from RLHF datasets. It aims to identify rejections in LLMs when the prompt doesn't pass content moderation, classifying inputs into two categories: 0 for normal outputs and 1 for rejection detected. It achieves the following results on the evaluation set: Loss: 0.0544 Accuracy: 0.9887 Recall: 0.9810 Precision: 0.9279…

F001F002F003F004F005F006F007F010F011F012F013F014F015F016F017F018F019