Where the data comes from, and what CDSBench does with it
CDSBench publishes a fact, an excerpt or a page only when it has established where the information came from and what it is permitted to do with it. When it cannot, the material stays internal. This page explains the rules at a high level; the enforcement lives in the product.
What the public benchmark is built from
Public cases rest on public-record material: determinations-committee decisions and statements, regulatory and securities filings, court filings, issuer documents and press releases. Each source family carries a rights classification that decides how much of it may appear on CDSBench:
- Synthetic test corpus · CDSBenchsynthetic test material
Publicly accessible does not mean freely reproducible. Where the terms permit only reference, CDSBench shows the derived facts, a citation and a link to the authoritative document — never a mirror of it. Licensed material used internally for verification is not shown, indexed, exported or used to train anything.
How sources are credited
Every public case lists its sources with publisher, document title, date, retrieval date and a link to the original where linking is permitted. Attribution is a statement of origin — publishers named on CDSBench do not endorse it. Model names (GPT, Claude, Gemini) identify the systems measured; no partnership is implied.
Why a benchmark label can be trusted
A question counts in official scoring only when both gates pass: its ground truth is independently verified (two independent extractors plus an evidence check, or an official outcome statement naming the instrument) and its sources permit public benchmark use. A label can be correct and still unpublishable; both must hold. Questions that fail the rights gate are kept out of the leaderboard, Case Replays, search, exports and the API. Methodology →
- Public verified cases
- 30
- Public questions
- 296
- Internal-only questions
- 0
- Synthetic cases
- 30
Public counts everywhere on CDSBench include only rights-cleared, verified, publishable records. Internal-only and licensed material is counted separately and never added to public totals.
Private documents stay private
Documents uploaded for a private analysis are classified customer-private by default: no publication, no public benchmark, no search, no indexing, no training. They are sent to the configured model providers only for that customer’s own analysis. Contributing a private case to the public corpus is a separate, deliberate request that names exactly what may be shared, and it is reviewed before anything becomes a public candidate. What the current code does → · Privacy →
Errors are fixed in public
Factual and benchmark errors are corrected transparently: the record is fixed, affected models are re-scored, reports are regenerated and the change is logged with the previous and current result. Removals for source-rights reasons are logged the same way, without confidential detail. Corrections log →
Copyright, licensing, corrections and takedown
If you are a rights holder and believe material appears on CDSBench without permission, or you want a correction, write to rights@cdsbench.com with the page URL and the document concerned. A takedown request disables the document, removes public excerpts, updates the sitemap and re-scores any benchmark question that depended on it; the internal audit history is preserved. The same address handles licensing questions.