Anchors are what make scores comparable
A rubric that says "rate accuracy from 1 to 5" produces scores that mean different things to different raters and to the same rater on different days. Describing what each level looks like removes most of that drift.
The generated rubric anchors the top, middle and bottom of each scale. Three anchors is usually enough — describing all five levels tends to produce distinctions nobody can actually apply.
Requiring evidence
The output format demands a quote from the answer for every score. This single requirement improves judge reliability more than anything else in the rubric, because a score that cannot be evidenced is usually a score that was guessed.
It also makes disagreements diagnosable: when two raters differ, the quotes show whether they read different things or valued the same thing differently.