Skip to content
CDSBench
Model failure · Brightwater Energy Corp 2024 · USD 2034 Notes · CDSBench #57 · benchmark v0.1

GPT answered Ineligible; the verified outcome was Eligible

IncorrectDeliverabilityDeliverabilitydeterministically verifiedself-reported confidence 92%
Question

Applying the standard Deliverable Obligation Characteristics (Not Subordinated, Specified Currency, Not Contingent, Transferable, Maximum Maturity 30 years, Not Bearer), is the instrument a Deliverable Obligation for this credit event?

Verified answer
Eligible All characteristics are satisfied on the record.
GPT answered
Ineligible Deterministic rule failed: Transferable.
Decisive issue
Deliverability
Evidence for the label
GPT cited
evidence score 22%
Failure classification

Failed multi document reasoning

  • Failed multi document reasoning. A later document supersedes an earlier provision; the answer relied on the superseded text.
  • Wrong final answer. Answered INELIGIBLE; verified outcome ELIGIBLE.
  • Overconfident. Confidence 0.92 on an incorrect answer.

Classifications are produced by the failure analyzer against the verified label and the evidence map; the model never grades itself. Failure taxonomy →

Everyone else

The same question, every other participant

Open the full Case Replay with the documents →

Related failures

GPT also missed at least 6 other deliverability questions

Benchmark version

v0.1

This page records GPT’s answer from its latest official run on the frozen benchmark v0.1. If the model is re-run and the answer changes, the page shows the new result and the model page keeps the full run history. Corrections policy →

Same kind of question on your own documents?
Mastermind runs the deterministic tests, the skeptic and the evidence verification on your record. Private documents never enter the benchmark.
Analyze your own documents →