How CDSBench scores AI on credit derivatives
Final-answer correctness, evidence correctness, rule application, multi-document reasoning, contradiction handling, hallucination, and the ability to abstain when the record is insufficient. Every number on the site is derived from a stored run.
In plain language
- Cases
- Real historical credit-derivatives situations with legally usable public documents.
- Questions
- Created from independently verifiable outcomes: an official outcome statement, a deterministic fact in the record, or agreement across sources.
- Models
- Every participant receives the same structured record; raw models use a fixed prompt at temperature 0.
- Scoring
- Final answer, evidence, reasoning category and uncertainty handling. Prose quality is never scored.
- Holdout
- Mastermind is never tuned on holdout cases; they are excluded from precedent retrieval and reported separately.
- Corrections
- Errors in labels, documents, scoring or reports are fixed, logged publicly and re-run; past versions stay visible.
From case to score
- 1Historical event
A real credit event with legally usable public documents: DC decisions, auction materials, prospectuses, exchange-offer and restructuring documents.
- 2Information available at the time
The record is split into sections with stable IDs and frozen inside an immutable benchmark version.
- 3Verified question
Each question has an independently verifiable outcome. Two extractors must agree and an evidence check must pass; no model grades itself.
- 4Equivalent context
Every participant receives the same record. Raw models use a fixed prompt at temperature 0.
- 5Answers scored
Final answer, evidence, reasoning category and uncertainty handling. Prose quality is never scored.
- 6Evidence checked
Cited section IDs are compared with the supporting sections; citations that do not exist count as hallucinated.
- 7Failures classified
Every wrong or weakly supported answer gets a standard failure label and a permanent public page.
- 8Holdout
Mastermind is never tuned on holdout cases and they are excluded from precedent retrieval; 61 verified holdout questions on the current version.
Read the full methodology → · See where Mastermind fails → · Corrections log →
Ground-truth verification
No AI grades itself. Two independent extractors (Validator A: pattern extraction; Validator B: sentence-window extraction with different patterns) must agree; an evidence verifier confirms the cited sections literally contain the labelled value; a consistency checker confirms entity, instrument, identifiers, dates and event. Labels that depend on an official outcome are verified against the outcome statement in the official document. Labels for “insufficient information” require both validators to fail to find the fact; labels for “review required” require both to observe a conflict.
Ground-truth states: UNVERIFIED, PUBLIC_OUTCOME_VERIFIED, DETERMINISTICALLY_VERIFIED, MULTI_SOURCE_VERIFIED, DISPUTED. Only verified questions count toward official scores; UNVERIFIED and DISPUTED questions are shown but excluded. Every Case Replay shows its validation log. The corpus and its coverage limits are on the dataset page.
Scoring
- Correct: normalised answer equals the verified label. Dates normalise to ISO; “Deliverable” equals ELIGIBLE, and so on.
- Evidence score: F1 between cited section IDs and the supporting sections. Cited IDs that do not exist count as hallucinated evidence.
- Abstention quality: appropriate abstention rate on questions whose label is REVIEW_REQUIRED or INSUFFICIENT_INFORMATION, discounted by unnecessary abstentions elsewhere.
- False high confidence: wrong, non-abstaining answers reported at confidence ≥ 0.8.
- Complex reasoning: accuracy on document-interpretation, multi-document and deliverability questions.
Failure taxonomy
Every wrong or weakly supported answer is classified: wrong final answer, missed fact, wrong fact, misread document, rule application error, failed multi-document reasoning, unsupported claim, hallucinated evidence, failed to notice contradiction, overconfident, unnecessary review, missed uncertainty, retrieval failure, evidence selection failure, precedent misapplication, prompt injection followed. Documents are untrusted input: instruction-like text inside a document is content, not a command, and following it is a failure.
Versions, splits and integrity
Published benchmark versions are immutable. Cases are split into development, validation and holdout sets; Mastermind may improve against development, promotion decisions use validation, and nothing is optimised against holdout. Holdout cases are excluded from precedent retrieval. Cached results are keyed on a fingerprint of the engine code, so any engine change is re-measured.
Mastermind
Mastermind is not a new foundation model. It is a specialist system around frontier models. Each version is benchmarked against its predecessor and promoted only if measured gates pass (target category improved, no hallucination increase, no false-confidence increase, validation not worse, no material holdout regression, regression tests pass, acceptable cost and latency); its failures are published like any other participant’s. Ablations — the promoted configuration with one component removed at a time — are run internally and reported publicly only after a robust measurement on real cases.
Confidence: when a result is decided by code from verified facts it carries a fixed system confidence. In the promoted version, an analyst-decided result shows the analyst model’s reported confidence after verification; signal-based system confidence (rule agreement, verified claims, objections, conflicts, precedent strength, calibrated against observed accuracy) is a V5 stage that is measured but not yet promoted. The private result states which applies.
- 1DocumentsThe full record, pasted into a chat.
- 2PromptWhatever the user remembers to ask.
- 3LLMOne model reads everything and writes an answer.
- 4AnswerProse. Confidence is self-reported. An overlooked clause is invisible from the outside.
Fast, fluent and often right. The conclusion and its supporting facts are produced in the same breath, so nothing checks them.
- 1DocumentsSections get stable IDs so every later claim can point at one.
- 2Structured CDS factsCurrency, maturity, seniority, form, contingency and transfer terms are extracted per instrument, with the section each came from.
- 3Deterministic checksCode, not a model, tests the objective characteristics. Each check reports pass, fail or review and cites its evidence.
- 4Related historical casesStructurally similar CDSBench cases are retrieved with their verified outcome, similarities and material differences — never copied.
- 5Specialist analysisA frontier reasoning model — one of the models CDSBench ranks — writes the proposed conclusion against the facts, checks and record.
- 6Independent skepticA second pass looks for the strongest evidence-based reason the proposal is wrong.
- 7Claim-level verificationEvery material claim is checked against the cited sections. Unsupported claims are dropped; non-existent citations are flagged.
- 8Contradiction and missing-information checksConflicting documents send the case to review unless a later one governs; absent facts are named, not guessed.
- 9Final assessmentLikely eligible, likely ineligible, review required, or insufficient information — with the evidence, the counterargument and what is missing.
Mastermind tries to prove itself wrong before you see the result. The model is one component; it can be replaced as CDSBench measures newer versions. How each stage is measured →
Model providers
Provider transparency is not provider branding: the product is Mastermind, and its analyst model is a configuration that can change with a new promoted version. This deployment runs in mock mode: leaderboard participants are simulated models with documented weakness profiles, and no document text leaves the CDSBench server. When live providers are enabled, this paragraph names them and the leaderboard names each model version. The promoted Mastermind version’s analyst model is recorded in every private result’s engine record and JSON export.
Deterministic rules
Rules encode a simplified reading of the standard Deliverable Obligation Characteristics. Their validation status is shown wherever they are used; PROVISIONAL rules are never presented as authoritative.
| Rule | What it tests | Version | Status |
|---|---|---|---|
| MAX_MATURITY | The obligation must mature no later than the maximum maturity (default 30 years) measured from the reference date. | 1.0 | provisional |
| SPECIFIED_CURRENCY | The obligation must be denominated in a Specified Currency. | 1.0 | provisional |
| NOT_SUBORDINATED | The obligation must not be subordinated to the reference obligation / senior unsecured debt. | 1.0 | provisional |
| NOT_BEARER | The obligation must not be a bearer instrument unless cleared through a recognised clearing system. | 1.0 | provisional |
| NOT_CONTINGENT | Principal must not be reducible by reason of contingencies other than payment. | 1.0 | provisional |
| TRANSFERABLE | The obligation must be transferable to institutional investors without consent (interpretation of transfer restrictions is outside the deterministic layer). | 1.0 | provisional |
| ENTITY_MATCH | The obligation issuer must be the Reference Entity (or a recognised successor). | 1.0 | provisional |
| REQUIRED_FIELDS | Core fields (issuer, currency, maturity, seniority) must be established before a deliverability conclusion. | 1.0 | provisional |
Corrections
When an error is found in a verified label, a document, a scoring rule or a published report, the record is fixed, the correction is logged publicly with the previous and current result, impacted models are re-run, and past report versions stay visible. Corrections policy and log.
Citing and reusing results
Every public page has a permanent URL, a benchmark version and a “Cite” action (plain text, Markdown, BibTeX). Leaderboard, comparison and category results can be downloaded as CSV or JSON, embedded as live charts, or read from the public API (/api/public/leaderboard, /api/public/models, /api/public/cases/:slug, /api/public/benchmarks/current). Results may be reused with attribution to CDSBench and a link to the source page. Source documents are not redistributed beyond the excerpts shown on case pages.
Limitations
- Real historical cases are added only after their documents are ingested and their labels independently verified.
- Rules and question framings simplify contractual reality; they are provisional until publicly or expert validated.
- Human “beat the models” statistics are shown only once enough attempts exist.
- Mastermind is decision-support software, not legal advice and not a Credit Event determination.
Private documents are handled separately from the public benchmark: how private documents are handled.