What makes an LLM verifier trustworthy?
A verifier is trustworthy for a task class when its false-accept and false-reject rates are measured on a gold set and stay stable under adversarial inputs . Better base models do not remove the need for this; they raise the value of measuring it .