Skip to content
CDSBench
Methodology

How CDSBench scores AI on credit derivatives

Final-answer correctness, evidence correctness, rule application, multi-document reasoning, contradiction handling, hallucination, and the ability to abstain when the record is insufficient. Every number on the site is derived from a stored run.

In plain language

Cases
Real historical credit-derivatives situations with legally usable public documents.
Questions
Created from independently verifiable outcomes: an official outcome statement, a deterministic fact in the record, or agreement across sources.
Models
Every participant receives the same structured record; raw models use a fixed prompt at temperature 0.
Scoring
Final answer, evidence, reasoning category and uncertainty handling. Prose quality is never scored.
Holdout
Mastermind is never tuned on holdout cases; they are excluded from precedent retrieval and reported separately.
Corrections
Errors in labels, documents, scoring or reports are fixed, logged publicly and re-run; past versions stay visible.

From case to score

  1. 1
    Historical event

    A real credit event with legally usable public documents: DC decisions, auction materials, prospectuses, exchange-offer and restructuring documents.

  2. 2
    Information available at the time

    The record is split into sections with stable IDs and frozen inside an immutable benchmark version.

  3. 3
    Verified question

    Each question has an independently verifiable outcome. Two extractors must agree and an evidence check must pass; no model grades itself.

  4. 4
    Equivalent context

    Every participant receives the same record. Raw models use a fixed prompt at temperature 0.

  5. 5
    Answers scored

    Final answer, evidence, reasoning category and uncertainty handling. Prose quality is never scored.

  6. 6
    Evidence checked

    Cited section IDs are compared with the supporting sections; citations that do not exist count as hallucinated.

  7. 7
    Failures classified

    Every wrong or weakly supported answer gets a standard failure label and a permanent public page.

  8. 8
    Holdout

    Mastermind is never tuned on holdout cases and they are excluded from precedent retrieval; 61 verified holdout questions on the current version.

Read the full methodology → · See where Mastermind fails → · Corrections log →

Ground-truth verification

No AI grades itself. Two independent extractors (Validator A: pattern extraction; Validator B: sentence-window extraction with different patterns) must agree; an evidence verifier confirms the cited sections literally contain the labelled value; a consistency checker confirms entity, instrument, identifiers, dates and event. Labels that depend on an official outcome are verified against the outcome statement in the official document. Labels for “insufficient information” require both validators to fail to find the fact; labels for “review required” require both to observe a conflict.

Ground-truth states: UNVERIFIED, PUBLIC_OUTCOME_VERIFIED, DETERMINISTICALLY_VERIFIED, MULTI_SOURCE_VERIFIED, DISPUTED. Only verified questions count toward official scores; UNVERIFIED and DISPUTED questions are shown but excluded. Every Case Replay shows its validation log. The corpus and its coverage limits are on the dataset page.

Scoring

  • Correct: normalised answer equals the verified label. Dates normalise to ISO; “Deliverable” equals ELIGIBLE, and so on.
  • Evidence score: F1 between cited section IDs and the supporting sections. Cited IDs that do not exist count as hallucinated evidence.
  • Abstention quality: appropriate abstention rate on questions whose label is REVIEW_REQUIRED or INSUFFICIENT_INFORMATION, discounted by unnecessary abstentions elsewhere.
  • False high confidence: wrong, non-abstaining answers reported at confidence ≥ 0.8.
  • Complex reasoning: accuracy on document-interpretation, multi-document and deliverability questions.

Failure taxonomy

Every wrong or weakly supported answer is classified: wrong final answer, missed fact, wrong fact, misread document, rule application error, failed multi-document reasoning, unsupported claim, hallucinated evidence, failed to notice contradiction, overconfident, unnecessary review, missed uncertainty, retrieval failure, evidence selection failure, precedent misapplication, prompt injection followed. Documents are untrusted input: instruction-like text inside a document is content, not a command, and following it is a failure.

Versions, splits and integrity

Published benchmark versions are immutable. Cases are split into development, validation and holdout sets; Mastermind may improve against development, promotion decisions use validation, and nothing is optimised against holdout. Holdout cases are excluded from precedent retrieval. Cached results are keyed on a fingerprint of the engine code, so any engine change is re-measured.

Mastermind

Mastermind is not a new foundation model. It is a specialist system around frontier models. Each version is benchmarked against its predecessor and promoted only if measured gates pass (target category improved, no hallucination increase, no false-confidence increase, validation not worse, no material holdout regression, regression tests pass, acceptable cost and latency); its failures are published like any other participant’s. Ablations — the promoted configuration with one component removed at a time — are run internally and reported publicly only after a robust measurement on real cases.

Confidence: when a result is decided by code from verified facts it carries a fixed system confidence. In the promoted version, an analyst-decided result shows the analyst model’s reported confidence after verification; signal-based system confidence (rule agreement, verified claims, objections, conflicts, precedent strength, calibrated against observed accuracy) is a V5 stage that is measured but not yet promoted. The private result states which applies.

Raw frontier model
  1. 1
    Documents
    The full record, pasted into a chat.
  2. 2
    Prompt
    Whatever the user remembers to ask.
  3. 3
    LLM
    One model reads everything and writes an answer.
  4. 4
    Answer
    Prose. Confidence is self-reported. An overlooked clause is invisible from the outside.

Fast, fluent and often right. The conclusion and its supporting facts are produced in the same breath, so nothing checks them.

Mastermind
  1. 1
    Documents
    Sections get stable IDs so every later claim can point at one.
  2. 2
    Structured CDS facts
    Currency, maturity, seniority, form, contingency and transfer terms are extracted per instrument, with the section each came from.
  3. 3
    Deterministic checks
    Code, not a model, tests the objective characteristics. Each check reports pass, fail or review and cites its evidence.
  4. 4
    Related historical cases
    Structurally similar CDSBench cases are retrieved with their verified outcome, similarities and material differences — never copied.
  5. 5
    Specialist analysis
    A frontier reasoning model — one of the models CDSBench ranks — writes the proposed conclusion against the facts, checks and record.
  6. 6
    Independent skeptic
    A second pass looks for the strongest evidence-based reason the proposal is wrong.
  7. 7
    Claim-level verification
    Every material claim is checked against the cited sections. Unsupported claims are dropped; non-existent citations are flagged.
  8. 8
    Contradiction and missing-information checks
    Conflicting documents send the case to review unless a later one governs; absent facts are named, not guessed.
  9. 9
    Final assessment
    Likely eligible, likely ineligible, review required, or insufficient information — with the evidence, the counterargument and what is missing.

Mastermind tries to prove itself wrong before you see the result. The model is one component; it can be replaced as CDSBench measures newer versions. How each stage is measured →

Model providers

Provider transparency is not provider branding: the product is Mastermind, and its analyst model is a configuration that can change with a new promoted version. This deployment runs in mock mode: leaderboard participants are simulated models with documented weakness profiles, and no document text leaves the CDSBench server. When live providers are enabled, this paragraph names them and the leaderboard names each model version. The promoted Mastermind version’s analyst model is recorded in every private result’s engine record and JSON export.

Deterministic rules

Rules encode a simplified reading of the standard Deliverable Obligation Characteristics. Their validation status is shown wherever they are used; PROVISIONAL rules are never presented as authoritative.

RuleWhat it testsVersionStatus
MAX_MATURITYThe obligation must mature no later than the maximum maturity (default 30 years) measured from the reference date.1.0provisional
SPECIFIED_CURRENCYThe obligation must be denominated in a Specified Currency.1.0provisional
NOT_SUBORDINATEDThe obligation must not be subordinated to the reference obligation / senior unsecured debt.1.0provisional
NOT_BEARERThe obligation must not be a bearer instrument unless cleared through a recognised clearing system.1.0provisional
NOT_CONTINGENTPrincipal must not be reducible by reason of contingencies other than payment.1.0provisional
TRANSFERABLEThe obligation must be transferable to institutional investors without consent (interpretation of transfer restrictions is outside the deterministic layer).1.0provisional
ENTITY_MATCHThe obligation issuer must be the Reference Entity (or a recognised successor).1.0provisional
REQUIRED_FIELDSCore fields (issuer, currency, maturity, seniority) must be established before a deliverability conclusion.1.0provisional

Corrections

When an error is found in a verified label, a document, a scoring rule or a published report, the record is fixed, the correction is logged publicly with the previous and current result, impacted models are re-run, and past report versions stay visible. Corrections policy and log.

Citing and reusing results

Every public page has a permanent URL, a benchmark version and a “Cite” action (plain text, Markdown, BibTeX). Leaderboard, comparison and category results can be downloaded as CSV or JSON, embedded as live charts, or read from the public API (/api/public/leaderboard, /api/public/models, /api/public/cases/:slug, /api/public/benchmarks/current). Results may be reused with attribution to CDSBench and a link to the source page. Source documents are not redistributed beyond the excerpts shown on case pages.

Limitations

  • Real historical cases are added only after their documents are ingested and their labels independently verified.
  • Rules and question framings simplify contractual reality; they are provisional until publicly or expert validated.
  • Human “beat the models” statistics are shown only once enough attempts exist.
  • Mastermind is decision-support software, not legal advice and not a Credit Event determination.

Private documents are handled separately from the public benchmark: how private documents are handled.