A model-evaluation engineer choosing an answer verifier for a task submits one identical test set to several verifiers and compares their verdicts.
The inferable old practice is running separate scripts and recording verifier results ad hoc, but the candidate offers no evidence of this alternative behaviour.
The material does not state the evaluator's concrete pain, only the practice of evaluating verifiers on one identical test set, so what is lost by not solving it cannot be confirmed.