Skip to content
CDSBench
Corrections

Corrections methodology and log

A benchmark is only worth citing if its errors are visible. When CDSBench finds a mistake in a verified label, a document, a scoring rule or a published report, the fix is recorded here with the previous and current result and the models that were re-run. Nothing is edited silently.

Policy

What happens when an error is found

  1. Fix the record. The label, evidence map, document or rule is corrected in the current benchmark version. If the change is material to comparability, a new benchmark version is frozen instead and the old one stays immutable.
  2. Record the correction. An entry below states what changed, why, the previous and the current result, and the date.
  3. Re-run impacted models. Every model whose score depends on the corrected item is re-scored or re-run; the run ids are listed.
  4. Update scores and reports. Leaderboard numbers change with the re-run. A published report whose numbers changed gets a new report version; its previous result stays visible on the report page under “Report versioning”.
  5. Preserve the audit trail. Old runs, labels and report versions are never deleted.

Report an error: use the case page’s source links to check the record, then write to the address in the private-data notes. Disputed labels are excluded from scoring until resolved.

Log

No corrections recorded

No corrections have been needed yet
This page fills in the first time a published number changes because of an error. An empty log is not a claim of perfection; it means nothing has been reported or found so far.