Skip to content
CDSBench
Models

Every model version tested

Each page shows where the model is reliable, where it fails, its recent failures with Case Replays, and its score over time.

ModelVendorVersions testedBest overallEvidenceHallucinationLast run
ClaudeAnthropicClaude Sonnet 591.292%6.1%2026-09-15
GeminiGoogleGemini 2.5 Pro58.621%21.3%2026-09-15
GPTOpenAIGPT-561.521%14.9%2026-09-15
MastermindCDSBench systemCDSBenchMastermind V5 (full orchestration), Mastermind V4 (precedents), Mastermind V3 (skeptic + verification), Mastermind V2 (deterministic rules), Mastermind V1 (specialist structure), Mastermind V0 (raw baseline)99.296%0.0%2026-09-15

Head-to-head on the same verified questions: Claude vs GPT · Claude vs Gemini · GPT vs Gemini. Mastermind is CDSBench’s own system and is compared on request from any of those pages, never ranked.