Models
Every model version tested
Each page shows where the model is reliable, where it fails, its recent failures with Case Replays, and its score over time.
| Model | Vendor | Versions tested | Best overall | Evidence | Hallucination | Last run |
|---|---|---|---|---|---|---|
| Claude | Anthropic | Claude Sonnet 5 | 91.2 | 92% | 6.1% | 2026-09-15 |
| Gemini | Gemini 2.5 Pro | 58.6 | 21% | 21.3% | 2026-09-15 | |
| GPT | OpenAI | GPT-5 | 61.5 | 21% | 14.9% | 2026-09-15 |
| MastermindCDSBench system | CDSBench | Mastermind V5 (full orchestration), Mastermind V4 (precedents), Mastermind V3 (skeptic + verification), Mastermind V2 (deterministic rules), Mastermind V1 (specialist structure), Mastermind V0 (raw baseline) | 99.2 | 96% | 0.0% | 2026-09-15 |
Head-to-head on the same verified questions: Claude vs GPT · Claude vs Gemini · GPT vs Gemini. Mastermind is CDSBench’s own system and is compared on request from any of those pages, never ranked.