Model performance on Tasks A–D.
No benchmark version has been released yet. This page is generated by 06_analysis once a full cross-vendor run completes.
The generator emits one row per model, per benchmark version. Columns:
| Model | Vendor | A: severity acc. | B: design acc. | C: FP rate (clean) | D: citation prec. | Parse failures | Paraphrased Δ |
|---|---|---|---|---|---|---|---|
Rows appear here after the first released run. Every cell is sourced from a
manifest.json; no value is entered by hand.
|
|||||||
Each published row links to its run manifest: full model string, prompt version and hash, dataset version, item count, and run date. Results from different prompt versions are never shown in the same table.