PIT
What the reference run measures.
These are checks on the frozen run: measured on 594 calls across 10 model families on 2026-09-11. Efficiency is reported as measured. A row that misses its bar says so. Every row is computed by the same engine the desk uses.
21 / 22 pass
| Measure | Value | Note |
|---|---|---|
| Second price · measured efficiency | 1.000pass | Mean of realized welfare over first-best across 12 periods; minimum 1.000. Measured on 120 calls across 10 model families on 2026-09-11. |
| Second price · optimal agents | 1.000pass | Every seat replaced by its truthful reference bid on the same values. Efficiency must be 1. |
| Second price · truthful reference | bid = valuepass | The optimal reference bid in a second-price sale equals the private value. |
| First price · risk-neutral Nash shade | 0.100 of valuepass | Reference bid is value × (n−1)/n with n = 10; checked to 1e-9. |
| First price · measured shade | 0.041 meanpass | 91.2% of 113 live bids sit inside the human band [0.000, 0.100] (RNNE to RNNE + 0.10 of value). Reported as measured on 2026-09-11. |
| First price · measured efficiency | 1.000pass | Mean of realized welfare over first-best across 12 periods; minimum 1.000. Measured on 120 calls across 10 model families on 2026-09-11. |
| Sequential · measured efficiency | 1.000pass | Mean of realized welfare over first-best across 12 periods; minimum 1.000. Measured on 114 calls across 10 model families on 2026-09-11. |
| Simultaneous · measured efficiency | 0.995pass | Mean of realized welfare over first-best across 12 periods; minimum 0.971. Measured on 120 calls across 10 model families on 2026-09-11. |
| Double auction · measured efficiency | 0.796fail | Mean of realized welfare over first-best across 12 periods; minimum 0.496. Measured on 120 calls across 10 model families on 2026-09-11. |
| Sequential vs simultaneous · revenue order | +45.52pass | Seller revenue, sequential minus simultaneous, with optimal agents on the same frozen draw (live agents: +55.88). Sign must match the draw's second-minus-third value gap. |
| News · bid movement after a material event | 19.01 mean movepass | 99.1% of 116 seats whose value moved by 10 or more changed their order by at least 1%. |
| Overbid classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 6 hand-labeled positives. Synthetic, not live. |
| Winner's curse classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 3 hand-labeled positives. Synthetic, not live. |
| News ignored classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 5 hand-labeled positives. Synthetic, not live. |
| Synchronized orders classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 2 hand-labeled positives. Synthetic, not live. |
| Preference reversal classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 1 hand-labeled positives. Synthetic, not live. |
| Under the human band classifier · labeled set | 1.000pass | Harmonic mean of precision 1.000 and recall 1.000 on 4 hand-labeled positives. Synthetic, not live. |
| Model families with live orders | 10pass | amazon, anthropic, cohere, deepseek, google, meta-llama, mistralai, openai, qwen, x-ai. Ten or more required. |
| Order shape | all four-tuplespass | Every recorded order has a finite time, a known asset, a finite quantity, and a finite price. 587 orders. |
| Clearing · desk matches frozen run | matchpass | Winners, prices, revenue, and efficiency recomputed from the frozen orders equal the runner's record to 1e-9. |
| Summary · desk matches summary.json | matchpass | Mean and minimum efficiency per structure recomputed from the traces equal summary.json to 1e-9. Generated 2026-09-11. |
| Orders parsed from completions | 98.8%pass | 587 of 594 completions held one valid JSON order. Pass at 80%. 46 classifier hits across the live traces on 2026-09-11. |