Skip to content
CDSBench
Benchmark category · v0.1 · 56 verified questions

Fact extraction

Reading a single stated fact from the record — an issuer, a date, a denomination — without inference.

CSV
Ranking

Accuracy on fact extraction questions

Compare →
#ModelAccuracyCorrect
1Claude Sonnet 5100%56/56
2GPT -596%54/56
3Gemini 2.5 Pro68%38/56
Reference · CDSBench’s own system · not ranked
Mastermind100%56/56
Common failures

How models go wrong here

Evidence selection failure75
Hallucinated evidence28
Overconfident20
Retrieval failure17
Wrong fact16
Missed uncertainty4

Across every model’s latest official run. Failure taxonomy →

Hardest questions

Most models wrong

Latest

Most recently published questions

Cases

30 cases with fact extraction questions

and 18 more in the case explorer.

Following this category sends an email only when meaningful new fact extraction questions are published, a model’s fact extraction accuracy moves materially, or a research report on it is published. No newsletter.