PIT

Measure how AI agents behave when something is at stake.

Ten model families bid in a simulated market. The desk records every order, clears each period against the true values, and reports how much of the possible value the agents captured.

Launch live demonstration · How it works

Simulated market · measured on 594 calls across 10 model families on 2026-09-11

What it does

Setup

Choose a market structure. The environment draws each seat's private value and fixes the payout rule.

Feed

A public news tape moves the value of an asset. Every seat reads the same words; only the environment knows the number.

Agents

One stock agent per model family takes a seat. Each is observed through its answers alone.

Book

Every order is a four-tuple: time, asset, quantity, price. Recorded as written, never corrected.

Clear

The market clears on the true values. Efficiency is realized welfare over the best possible allocation.

Assay

Classifiers name overbids, ignored news, synchronized prices, and a winner who paid too much.

Compare

Efficiency and shading by model family, side by side with the human-theory band.

Live reference

Five market structures. Ten model families. One frozen run.

The market is simulated. The behavior is real: every order on this site was written by a model, measured on 594 calls across 10 model families on 2026-09-11. The desk replays the frozen record; it never calls a model.

Structures
5
Model families
10
Calls
594
Orders parsed
98.8%
Second-price η
1.000

21 / 22 measured checks pass

Open the desk · Measured checks · Workspace

Workspace

Your own runs stay with you.

The public demonstration reads the reference run. Load a trace of your own in the workspace and it replays in your browser. Nothing is uploaded, and no model is called.