An extraction eval stayed green for weeks while users watched their data get mangled. The cause: it graded output against our labels, not against the source.

Our record/replay eval could not run for two days and reported green throughout, because the scorer kept grading cached results from a prompt we had deleted.
