Skip to content
CDSBench
Model failure · Republic of Sanderia 2024 (B) · USD 2029 Notes · CDSBench #203 · benchmark v0.1

GPT reached the verified outcome but cited the wrong section

CorrectSeniorityRule applicationdeterministically verifiedself-reported confidence 92%
Question

Is the instrument subordinated?

Verified answer
No Stated in the status condition.
GPT answered
No Based on the seniority provision, the answer is NO.
Decisive issue
Seniority
Evidence for the label
GPT cited
evidence score 0%
Failure classification

Evidence selection failure

  • Evidence selection failure. Correct answer but cited evidence does not match the supporting sections (evidence score 0.00).

Classifications are produced by the failure analyzer against the verified label and the evidence map; the model never grades itself. Failure taxonomy →

Everyone else

The same question, every other participant

  • MastermindCorrectanswered No
  • ClaudeCorrectanswered No
  • GeminiCorrectanswered Noits failure page →

Open the full Case Replay with the documents →

Related failures

No other seniority misses on this run

Benchmark version

v0.1

This page records GPT’s answer from its latest official run on the frozen benchmark v0.1. If the model is re-run and the answer changes, the page shows the new result and the model page keeps the full run history. Corrections policy →

Same kind of question on your own documents?
Mastermind runs the deterministic tests, the skeptic and the evidence verification on your record. Private documents never enter the benchmark.
Analyze your own documents →