Research entry points, computed from the benchmark
Each collection is a question a professional asks — which cases are hardest, where does every model fail, when is a right answer not trustworthy — answered from stored runs and refreshed automatically after every official run.
- Hardest CDSBench cases25 questions
The verified questions most tested systems got wrong. Model disagreement and low accuracy point to where the record is genuinely hard to read.
- Cases every frontier model missed10 questions
Questions where no ranked model reached the verified outcome. Mastermind's result on each is shown on the same terms, including where it also failed.
- Right answer, unsupported evidence80 findings
Answers that matched the verified outcome but cited sections that did not support it, or cited sections that do not exist. A correct conclusion with the wrong reasoning is not a result a desk can rely on.
- Biggest improvements between model generations0 findings
Where a newer version of a model gained the most accuracy over its predecessor on the same frozen questions, by category.
- Sovereign restructurings15 cases
Every published sovereign restructuring case in the benchmark, with its instruments, documents and model results.
- Transferability disputes30 questions
Every verified transferability question. Transfer restrictions are where models most often reach a confident, wrong conclusion.
- Failure-to-pay cases2 cases
Every published failure-to-pay case, with the deliverability questions that follow from it.