Skip to content
CDSBench
CDSBench · Mastermind V0 (raw baseline) · benchmark v0.1 · 2026-09-15

Mastermind

CDSBench score91.2on 296 verified questions
Scores (CSV)

Every failure with a permanent page →

Accuracy
91%
Evidence
92%
cited the right sections
Complex reasoning
73%
multi-document, deliverability
Uncertainty
94%
abstains only when it should
Hallucination
6.1%
cites non-existent evidence
False confidence
0.0%
wrong at ≥ 80% confidence
Strongest

Where it is reliable

  • Restructuring100%n=2
  • Multi-document100%n=2
  • Entity match100%n=30
Weakest

Where it fails

  • Transferability57%n=30
  • Uncertainty69%n=32
  • Deliverability83%n=30
Failure profile

Accuracy by category and question level

CategoryAcc.n
Transferability57%30
Uncertainty69%32
Deliverability83%30
Currency100%28
Evidence100%30
Maturity100%26
Seniority100%30
Fact extraction100%56
Entity match100%30
Multi-document100%2
Restructuring100%2
LevelAcc.n
Uncertainty69%32
Deliverability83%30
Multi document75%8
Fact extraction100%86
Rule application100%112
Document interpretation61%28
Failure types

How it goes wrong

Hallucinated evidence18
Unnecessary review18
Wrong final answer10
Unsupported claim10
Performance over time

All official runs on this benchmark version

8090100Mastermind V0 Mastermind V1 Mastermind V2 Mastermind V3 Mastermind V4 Mastermind V5 2026-09-152026-09-15
DateVersionOverallAccuracy
2026-09-15 17:16Mastermind V0 (raw baseline)91.291%
2026-09-15 17:16Mastermind V1 (specialist structure)91.491%
2026-09-15 17:17Mastermind V2 (deterministic rules)92.594%
2026-09-15 17:17Mastermind V3 (skeptic + verification)99.2100%
2026-09-15 17:18Mastermind V4 (precedents)99.2100%
2026-09-15 17:18Mastermind V5 (full orchestration)99.2100%
2026-09-15 20:00Mastermind V0 (raw baseline)91.291%
2026-09-15 20:00Mastermind V1 (specialist structure)91.491%
2026-09-15 20:00Mastermind V2 (deterministic rules)92.594%
2026-09-15 20:01Mastermind V3 (skeptic + verification)99.2100%
2026-09-15 20:01Mastermind V4 (precedents)99.2100%
2026-09-15 20:02Mastermind V5 (full orchestration)99.2100%
2026-09-15 20:14Mastermind V0 (raw baseline)91.291%
2026-09-15 20:14Mastermind V1 (specialist structure)91.491%
2026-09-15 20:14Mastermind V2 (deterministic rules)92.594%
2026-09-15 20:14Mastermind V3 (skeptic + verification)99.2100%
2026-09-15 20:14Mastermind V4 (precedents)99.2100%
2026-09-15 20:14Mastermind V5 (full orchestration)99.2100%
2026-09-15 20:36Mastermind V0 (raw baseline)91.291%
2026-09-15 20:36Mastermind V1 (specialist structure)91.491%
2026-09-15 20:36Mastermind V2 (deterministic rules)92.594%
2026-09-15 20:36Mastermind V3 (skeptic + verification)99.2100%
2026-09-15 20:36Mastermind V4 (precedents)99.2100%
2026-09-15 20:36Mastermind V5 (full orchestration)99.2100%
Engine versions

Every Mastermind version on the same benchmark, same settings

VersionOverallAccuracyEvidenceHalluc.False conf.Cost
Mastermind V0 (raw baseline)promoted91.291%92%6.1%0.0%$0.000
Mastermind V1 (specialist structure)91.491%93%0.0%0.0%$0.000
Mastermind V2 (deterministic rules)92.594%97%0.0%1.4%$0.000
Mastermind V3 (skeptic + verification)99.2100%96%0.0%0.0%$0.000
Mastermind V4 (precedents)99.2100%96%0.0%0.0%$0.000
Mastermind V5 (full orchestration)99.2100%96%0.0%0.0%$0.000

A version is promoted only when it beats its predecessor on the validation split under measured gates. Regressions stay listed.

Mastermind on CDSBench v0.1 · CDSBench