When credit breaks, Mastermind helps you understand what matters for the CDS.
Something happened to a company or sovereign. What does it mean for the CDS, which instruments matter, what likely qualifies, why, what historical cases are relevant, and what might I be missing?
Upload the documents. Mastermind structures the event and the instrument, runs the objective checks, analyzes deliverability, retrieves related historical cases, verifies the supporting evidence, challenges its own conclusion and shows what remains unresolved.
Mastermind V0 (raw baseline) on CDSBench v0.1: 296 independently verified questions from 296 verified labels; 61 holdout questions never used for tuning. Same record, same scoring as every ranked model. Methodology →
Frontier models are powerful. High-stakes CDS work still needs the structure they do not bring by default.
- 1DocumentsThe full record, pasted into a chat.
- 2PromptWhatever the user remembers to ask.
- 3LLMOne model reads everything and writes an answer.
- 4AnswerProse. Confidence is self-reported. An overlooked clause is invisible from the outside.
Fast, fluent and often right. The conclusion and its supporting facts are produced in the same breath, so nothing checks them.
- 1DocumentsSections get stable IDs so every later claim can point at one.
- 2Structured CDS factsCurrency, maturity, seniority, form, contingency and transfer terms are extracted per instrument, with the section each came from.
- 3Deterministic checksCode, not a model, tests the objective characteristics. Each check reports pass, fail or review and cites its evidence.
- 4Related historical casesStructurally similar CDSBench cases are retrieved with their verified outcome, similarities and material differences — never copied.
- 5Specialist analysisA frontier reasoning model — one of the models CDSBench ranks — writes the proposed conclusion against the facts, checks and record.
- 6Independent skepticA second pass looks for the strongest evidence-based reason the proposal is wrong.
- 7Claim-level verificationEvery material claim is checked against the cited sections. Unsupported claims are dropped; non-existent citations are flagged.
- 8Contradiction and missing-information checksConflicting documents send the case to review unless a later one governs; absent facts are named, not guessed.
- 9Final assessmentLikely eligible, likely ineligible, review required, or insufficient information — with the evidence, the counterargument and what is missing.
Mastermind tries to prove itself wrong before you see the result. The model is one component; it can be replaced as CDSBench measures newer versions. How each stage is measured →
Mastermind is not a new model and not a competitor to the models CDSBench ranks. It calls one of them for analysis and for skepticism, and wraps it in code that extracts the facts, runs the deterministic checks, retrieves related cases and verifies every claim against the record. It is scored by the same rules as everyone else, and its misses are published on the same terms.
Seven layers, one of which is the model.
- Frontier models
- Commodity input. Mastermind runs on them and can change which one it uses.
- CDSBench dataset
- Historical credit events, instruments, documents, verified questions and outcomes, frozen in versions.
- Rules
- Deterministic CDS logic in code: maturity, currency, seniority, form, contingency, transferability, entity match.
- Precedents
- Structured similarity over the historical corpus, with similarities and material differences stated.
- Verification
- Claim-by-claim testing against the cited sections; unsupported claims are dropped or send the result to review.
- Evaluation history
- Every model version's answers, failures and fixes, classified and kept; regressions become tests.
- Workflow
- One structured, repeatable output built for CDS work, not a chat transcript.
The model is one component. The system around it is the product.
Claude missed 28 of 296 verified questions. Here is what Mastermind did on each.
Same record, same question, same scoring. Every row opens the Case Replay where you can read the clause and both answers. Where Mastermind abstained, the record did not support a conclusion, and abstaining is scored as correct only when the verified label agrees.
Mastermind is valuable even where the raw model reaches the same answer: every conclusion arrives with the rule-level checks, the verified evidence and the counterargument, in the same structure on every case.
We publish Mastermind’s failures too.
A system that only showed its wins would not be worth trusting with a live case. Every Mastermind version is benchmarked before promotion; regressions stay listed on the model page.
| Category | Mastermind | Claude |
|---|---|---|
| Transferability | 57% | 57% |
| Uncertainty | 69% | 69% |
| Deliverability | 83% | 83% |
| Currency | 100% | 100% |
20 verified questions wrong or weakly supported
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (ELIGIBLE).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Wrong final answer. Answered YES; verified outcome NO.
- Wrong final answer. Answered YES; verified outcome NO.
- Wrong final answer. Answered YES; verified outcome NO.
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Hallucinated evidence. 1 cited section id(s) do not exist in the case record.
- Wrong final answer. Answered YES; verified outcome NO.
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (ELIGIBLE).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Wrong final answer. Answered YES; verified outcome NO.
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (NO).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (YES).
- Unnecessary review. Abstained (REVIEW_REQUIRED) although the record supports a definite answer (ELIGIBLE).
- Hallucinated evidence. 1 cited section id(s) do not exist in the case record.
- Wrong final answer. Answered YES; verified outcome NO.
- Wrong final answer. Answered YES; verified outcome NO.
Everything you receive, and the evidence behind it.
- 1Likely eligibleReviewOverall assessmentLikely eligible, likely ineligible, review required, or insufficient information — never a false certainty.
- 2Test matrixEvery objective check — currency, maturity, seniority, form, contingency, transferability, entity — with pass, fail or unresolved and the clause it used.
- 3Exact evidenceEach material claim links to the section it rests on. Click through to the clause in your own document.
- 4Related historical casesStructurally similar CDSBench cases with their verified outcome, what is similar, and what is materially different.
- 5Strongest counterargumentWhat could make this conclusion wrong, found by an independent pass that looks for overlooked evidence and unsupported assumptions.
- 6Missing informationWhich document or fact would change the confidence, so you know what to obtain next.
- 7What Mastermind addedThe checks, verifications and challenges the run actually performed, beyond the underlying model.
- Which instruments appear deliverable, which do not, and which need review.
- The one contractual issue that decides it, with the clause.
- Normalized instrument terms with the section each came from.
- Related historical cases and their material differences.
- Every test and claim linked to the governing wording.
- The strongest counterargument and the documents still missing.
Promoted only when it beats its predecessor under measured gates.
| Version | Adds | Overall | Accuracy | Evidence | Halluc. | |
|---|---|---|---|---|---|---|
| Mastermind V0 (raw baseline) | Standardized prompt → frontier model → structured answer. | 91.2 | 91% | 92% | 6.1% | promoted · in production |
| Mastermind V1 (specialist structure) | CDS case schema, specialist prompt, strict JSON, evidence IDs, insufficient-information state. | 91.4 | 91% | 93% | 0.0% | |
| Mastermind V2 (deterministic rules) | Objective logic moved into code: maturity, currency, seniority, identifiers, required fields. | 92.5 | 94% | 97% | 0.0% | |
| Mastermind V3 (skeptic + verification) | Primary analyst → skeptic → claim-level evidence verification → contradiction check → revised result. | 99.2 | 100% | 96% | 0.0% | |
| Mastermind V4 (precedents) | Analogous historical cases retrieved from the CDSBench corpus with similarities/differences; section-level evidence retrieval for large records. | 99.2 | 100% | 96% | 0.0% | |
| Mastermind V5 (full orchestration) | Adds multi-model consensus on contested conclusions, alternative hypotheses, and system confidence computed from pipeline signals with observed calibration. | 99.2 | 100% | 96% | 0.0% |
Benchmark v0.1. Stages listed for versions above the promoted one are built and measured but not yet promoted; private analyses run the promoted version. As frontier models change, CDSBench reruns them and Mastermind’s analyst model can change with a new version — the customer never has to pick a vendor. Full run history, including every regression, on the model page.
A repeatable credit-event analysis process, not access to a model.
- Speed
- Compress the fragmented document-and-research work around one instrument into one structured analysis.
- Coverage
- Every objective characteristic is tested, every time — nothing depends on remembering to ask.
- Historical memory
- Compare the situation with structured historical CDS cases instead of starting from a blank chat.
- Verification
- Every material conclusion is tied back to the clause it rests on. You never have to trust the answer blindly.
- Second opinion
- Mastermind actively looks for reasons its own conclusion may be wrong before you see it.
- Consistency
- The same structure on every case, so a desk or deal team can compare analyses rather than prompts.
- Uncertainty
- What cannot yet be established is stated, with the document that would resolve it.
Leaderboard, every Case Replay, model pages and failures, research reports, follows and alerts, data downloads and the public API.
Your documents and instruments: event, deliverability, tests, evidence, precedents, the skeptic, missing information and exports. The first analyses are free; paid plans open with real historical cases and live provider runs.
The questions a buyer asks before uploading anything.
- I can just use Claude or GPT.
- Yes. CDSBench shows you how Claude performs directly, question by question. Mastermind is for when you want the model wrapped in CDS-specific checks, historical precedent, evidence verification, adversarial review and structured uncertainty — and you want that measured on the same benchmark.
- Is this legal advice or a determination?
- No. Mastermind is decision-support software. It structures the record, runs objective checks, cites the clauses, retrieves precedent and states what is unresolved. It does not make Credit Event determinations, and its deterministic rules are provisional simplifications of the standard characteristics, shown as such.
- Why should I trust your score?
- Outcomes are verified against source documents by independent checks, questions are frozen in versioned benchmarks, every model receives the same record, holdout cases are never used for tuning, every failure has a public page, and errors are corrected in a public log. Methodology →
- What if the models improve?
- Good — CDSBench reruns them. Mastermind is built around replaceable frontier models. The benchmark tells us where generic capability has caught up and where the specialist infrastructure still adds value, version by version.
- What happens to my documents?
- They are parsed on the CDSBench server, stored against the email you give, never listed publicly, never added to the public benchmark and never used for precedent retrieval. CDSBench trains no models. Exactly what the current code does →
Upload the record for one instrument. Every stage appears as it finishes; the result shows the assessment, the tests, the evidence, the related cases, the counterargument, what is missing, and what Mastermind added beyond the model.
correct · abstained · incorrect. Verified outcome tone: Eligible Ineligible