Why raw agreement misleads
Two raters who both say "pass" ninety percent of the time will agree about eighty-two percent of the time by chance alone. Reporting that as agreement tells you nothing about whether they share a standard.
Kappa subtracts the agreement you would expect from the label distribution, which is why it is the number to quote and raw percentage is not.
The 0.6 rule
Below roughly 0.6, stop and fix the rubric. If two humans cannot agree on what good looks like, an automated judge built on the same criteria cannot either — it will produce consistent numbers that mean nothing.
The disagreement list is the repair tool: read the cases where raters differed and you will usually find one criterion doing all the damage.