Results

Model performance on Tasks A–D.

No benchmark version has been released yet. This page is generated by 06_analysis once a full cross-vendor run completes.

Table contract

The generator emits one row per model, per benchmark version. Columns:

Model Vendor A: severity acc. B: design acc. C: FP rate (clean) D: citation prec. Parse failures Paraphrased Δ
Rows appear here after the first released run. Every cell is sourced from a manifest.json; no value is entered by hand.
Reading Task C. The false-positive rate is the share of genuinely clean Item 9A disclosures a model incorrectly flags as deficient. Lower is better, and it is the metric that determines whether these models are usable for real control review. Paraphrased Δ is the score change on the paraphrased subset versus verbatim filing text — a large gap indicates memorization rather than reasoning.

Reproducibility

Each published row links to its run manifest: full model string, prompt version and hash, dataset version, item count, and run date. Results from different prompt versions are never shown in the same table.