TokenPad
Evaluation

LLM Evaluation Rubric Builder

Turn criteria into a weighted rubric a judge can apply consistently.

One criterion per line: name | weight | what a top score means

168 characters3 lines0 tokensor drop a file

Rubric BuilderExact
0Token change
Token change0no change
Input tokens0what you pasted
Output tokens0what you would send
Token cost of this result
Output tokens0
As input$0.00
× 100K requests$0.00

Everything on this page runs in your browser. Nothing you paste is transmitted, because there is no server here to transmit it to.

Result
 

Anchors are what make scores comparable

A rubric that says "rate accuracy from 1 to 5" produces scores that mean different things to different raters and to the same rater on different days. Describing what each level looks like removes most of that drift.

The generated rubric anchors the top, middle and bottom of each scale. Three anchors is usually enough — describing all five levels tends to produce distinctions nobody can actually apply.

Requiring evidence

The output format demands a quote from the answer for every score. This single requirement improves judge reliability more than anything else in the rubric, because a score that cannot be evidenced is usually a score that was guessed.

It also makes disagreements diagnosable: when two raters differ, the quotes show whether they read different things or valued the same thing differently.

Frequently asked questions

How many criteria should a rubric have?
Three to five. Beyond that raters stop discriminating between them and the weights become decorative — everything drifts towards the middle of every scale.
Should weights be equal?
Rarely. If accuracy matters three times more than tone, say so, otherwise a polished wrong answer beats a plain correct one in your aggregate score.

More evaluation tools