Skip to content
CDSBench
Model failure · Federal Republic of Aurelia 2028 · CHF 2037 Notes · CDSBench #49 · benchmark v0.1

GPT answered DOC-1#S1; the verified outcome was DOC-1#S4

IncorrectEvidenceFact extractiondeterministically verifiedself-reported confidence 92%
Question

Which section of the record establishes the seniority (ranking) of the instrument? Answer with the section identifier (e.g. DOC-1#S4).

Verified answer
DOC-1#S4 The Status condition states the ranking.
GPT answered
DOC-1#S1 Section DOC-1#S1 establishes the seniority.
Decisive issue
Evidence
Evidence for the label
GPT cited
evidence score 0%
Failure classification

Retrieval failure

  • Retrieval failure. None of the cited sections is among those that establish the answer; the relevant evidence was not retrieved.
  • Wrong fact. Extracted DOC_1S1; record states DOC_1S4.
  • Overconfident. Confidence 0.92 on an incorrect answer.

Classifications are produced by the failure analyzer against the verified label and the evidence map; the model never grades itself. Failure taxonomy →

Everyone else

The same question, every other participant

  • MastermindCorrectanswered DOC-1#S4
  • ClaudeCorrectanswered DOC-1#S4
  • GeminiIncorrectanswered DOC-1#S1its failure page →

Open the full Case Replay with the documents →

Related failures

GPT also missed at least 6 other evidence questions

Benchmark version

v0.1

This page records GPT’s answer from its latest official run on the frozen benchmark v0.1. If the model is re-run and the answer changes, the page shows the new result and the model page keeps the full run history. Corrections policy →

Same kind of question on your own documents?
Mastermind runs the deterministic tests, the skeptic and the evidence verification on your record. Private documents never enter the benchmark.
Analyze your own documents →
GPT on Federal Republic of Aurelia 2028 · CHF 2037 Notes — CDSBench #49 · CDSBench