Skip to content
CDSBench
Collections

Research entry points, computed from the benchmark

Each collection is a question a professional asks — which cases are hardest, where does every model fail, when is a right answer not trustworthy — answered from stored runs and refreshed automatically after every official run.

  • Hardest CDSBench cases25 questions

    The verified questions most tested systems got wrong. Model disagreement and low accuracy point to where the record is genuinely hard to read.

  • Cases every frontier model missed10 questions

    Questions where no ranked model reached the verified outcome. Mastermind's result on each is shown on the same terms, including where it also failed.

  • Right answer, unsupported evidence80 findings

    Answers that matched the verified outcome but cited sections that did not support it, or cited sections that do not exist. A correct conclusion with the wrong reasoning is not a result a desk can rely on.

  • Biggest improvements between model generations0 findings

    Where a newer version of a model gained the most accuracy over its predecessor on the same frozen questions, by category.

  • Sovereign restructurings15 cases

    Every published sovereign restructuring case in the benchmark, with its instruments, documents and model results.

  • Transferability disputes30 questions

    Every verified transferability question. Transfer restrictions are where models most often reach a confident, wrong conclusion.

  • Failure-to-pay cases2 cases

    Every published failure-to-pay case, with the deliverability questions that follow from it.