TRACE
What the public packs prove.
These are checks on the public desks. They are not government validation. Every row is computed from the same engine the desk uses, on the same frozen packs. Same inputs, same numbers.
85 / 85 pass
- Gold questions
- 34 / 34
- Deterministic asks, including authored measures.
- Prose accepted
- 51 / 54
- 12 questions · 6 model families answered
- Introduced-fact rate
- 0.056
- Share of completions that added a number, date, or name. Rejected, named.
- Ask-to-first-paint p95
- 27 ms
- 33 timed asks · p50 13 ms · Chrome 153.0.8010.36 headless · Darwin arm64
| Measure | Value | Note |
|---|---|---|
Packs and structure What is on the desk and how it is identified. | ||
| Published packs | 5pass | Kestrel, NTSB, FAA, airlines, ASRS-format. |
| Kestrel hops | 24pass | Synthetic Grayfield detachment. Frozen as of 2026-03-16. |
| NTSB public cases | 14pass | Published U.S. accident reports. Government works. |
| FAA strike years | 36pass | Table 1 from FAA/USDA Serial Report 32, 1990–2025. |
| 2025 reported strikes | 24447pass | Published Table 1 total for 2025. |
| NTSB case map | mappass | Approximate public crash-site coordinates. |
| Glossary hits | interceptpass | Check names the civil terms that fired. |
| Receipt path | /trace/demo?pack=kestrel&q=Which+hops+missed+the+intercept+last+week%3Fpass | Shareable demo URL for a live figure. |
| Airline safety rows | 56pass | FiveThirtyEight airline-safety.csv, unchanged. |
| ASRS-format records | 12pass | Synthetic desk. Not NASA case files. |
| Local CSV paste | 2pass | Parsed in-browser. No upload. |
| Content hashes | 21pass | FNV-1a 64 on each source. Labeled a content hash. |
Answers that refuse to guess Window, tape, still, conflict, cannot-answer, OCR, Whisper, read-only SQL. | ||
| Last-week misses | K-14, K-17, K-21pass | Window is publication as_of minus six days. |
| Merge radio cue | 4.58spass | Audio stays tied to hop K-17. |
| G-meter still | 3.8 Gpass | Syllabus limit 2.5. Graphic stays quoted. |
| Fuel conflict | 420 vs 380pass | Both figures stay. TRACE does not average. |
| Unknown aircraft | cannot answerpass | N999ZZ is not in the pack. |
| OCR · K-14 G-meter | 3.8 admittedpass | RapidOCR on a synthetic still. Floor 0.85. Same path as FORGE. |
| OCR · FAA 2025 | 24447 admittedpass | Held-out Table 1 excerpt. 24,447 admitted. Smudged 2020 total quarantined. |
| OCR · Hudson card | 155 admittedpass | Held-out public US Airways 1549 card. Not a live NTSB feed. |
| SQLite · fuel write-ups | 24pass | Read-only SELECT. 24 rows. Write attempt rejected. |
| SQLite · 2025 strikes | 24447pass | Held-out Table 1 year. Same admit path as Kestrel. |
| SQLite · Hudson survivors | 155pass | Read-only SELECT of the public Hudson case. |
| Whisper · Hudson count | 155 in transcriptpass | Synthesized speech, then Whisper. Gold is the score, not the product. Not cockpit voice. |
Authored relationships and measures Objects a person writes on the desk. Cardinality is detected. Numbers are computed. Receipts and boards re-run the recipe. | ||
| Authored relationship · log ↔ fuel | 1:1 · 24 matched · 0 orphanspass | raw_logs.hop_id ↔ raw_fuel.hop_id on Kestrel. Cardinality detected, not declared. |
| Authored relationship · year ↔ group | 1:n · 35 orphan yearspass | FAA years.year ↔ groups.year. One year has group rows; the rest are orphans and are counted, not hidden. |
| Authored relationship · pasted tables | 1:n · orphans 1 left, 1 rightpass | Two pasted CSVs. K-03 has no fuel row; K-09 has no hop; K-02 has two fuel rows. |
| Authored measure · avg fuel delta | -1.667 lbpass | avg(hops.fuel_delta_lb). Only K-08 differs (−40), so 24 rows average −1.667. |
| Authored measure · miss rate | 0.208 (5 of 24)pass | hops where intercept = miss / hops. Filter grammar: col op value. |
| Authored measure · peak G | 3.8 Gpass | max(hops.g_max). Matches the K-14 still. |
| Authored measure · long hops | 6 · last week 3pass | count(hops where duration_min >= 90). The ask applies the last-week window to the authored measure. |
| Authored measure · fatalities on approach | 74 peoplepass | sum(cases.fatalities where phase = Approach) on the NTSB desk. |
| Plan cites the authored measure | m-a38b40a5d2pass | "average fuel delta by hop" resolves to the authored measure, grouped by hop_id, 24 bars. |
| Plan cites the authored relationship | rel-878b23426fpass | An orphan-key question is answered from the authored relationship, not guessed. |
| Receipt carries authored objects | 655 charspass | The URL holds the recipe. Reopened, it gives the same plan, number, and lineage. |
| Board receipt round-trip | b466b60b7664pass | Three asks pinned, encoded, decoded, re-run. Board hash and results hash match. |
Guarded prose Frozen model restatements of computed answers, judged by the guard. Model families are the subjects under test. | ||
| Prose guard vs hand labels | 14 / 14pass | Hand-labeled restatements, including smuggled numbers, dates, and names. Guard agrees on every one. |
| Frozen prose accepted | 51 / 54pass | 54 completions with text from 6 model families, 120 calls. Judged by the guard at request time. |
| Introduced-fact rate | 0.056pass | 3 of 54 completions introduced a token not in the computed answer and were rejected. |
| Prose · OpenAI | 11 / 12 acceptedpass | openai/gpt-4.1-mini · median 939 ms · rejected: k-miss-rate: 21 |
| Prose · Anthropic | 12 / 12 acceptedpass | anthropic/claude-haiku-4.5 · median 913 ms |
| Prose · Google | 11 / 12 acceptedpass | google/gemini-3.1-flash-lite · median 694 ms · rejected: k-fuel: 2 |
| Prose · xAI | 11 / 12 acceptedpass | x-ai/grok-4.3 · median 7075 ms · rejected: k-fuel: 2 |
| Prose · Mistral AI | 5 / 5 acceptedpass | mistralai/mistral-small-2603 · median 574 ms |
| Prose · DeepSeek | no completionspass | deepseek/deepseek-v4-flash · 12 calls failed (HTTP 402, account credits exhausted). Frozen as failed, not filled. |
| Prose · Meta | 1 / 1 acceptedpass | meta-llama/llama-4-scout · median 5781 ms |
| Prose · Alibaba Qwen | no completionspass | qwen/qwen3.7-flash · 12 calls failed (HTTP 402, account credits exhausted). Frozen as failed, not filled. |
| Prose · Cohere | no completionspass | cohere/command-r-08-2024 · 12 calls failed (HTTP 402, account credits exhausted). Frozen as failed, not filled. |
| Prose · Amazon | no completionspass | amazon/nova-2-lite-v1 · 12 calls failed (HTTP 402, account credits exhausted). Frozen as failed, not filled. |
Workstation timing A real browser typed each frozen question and timed the first paint of the new answer. | ||
| Ask-to-first-paint p95 | 27 mspass | 33 timed asks · 3 repeats · p50 13 ms · Chrome 153.0.8010.36 headless · Darwin arm64 at 1440×1000 · 2026-09-11 |
Gold questions 34 / 34 pass. Each row is a frozen question with an expected number, list, conflict, cue, still, or cannot-answer. | ||
| Which hops missed the intercept last week? | 3pass | list · last week (2026-03-10–2026-03-16) · missed intercept |
| Parameter exceedance on hop 14 | 3.8pass | graphic · K-14 · parameter exceedance 14 |
| Play the radio call at the merge | 4.58pass | audio · play radio call at merge |
| Where are the conflicting fuel figures? | -40pass | conflict · conflicting fuel figures |
| How many hops used N999ZZ? | cannot answerpass | count · N999ZZ · n999zz |
| How many hops are in the log? | 24pass | count · log |
| How many missed intercepts are in the pack? | 5pass | count · missed intercepts pack |
| What is the average fuel delta by hop? | -1.667pass | measure · measure fuel delta (avg(hops.fuel_delta_lb)) · by hop_id |
| What is the miss rate? | 0.208pass | measure · measure miss rate (hops where intercept = miss / hops) |
| How many long hops last week? | 3pass | measure · measure long hops (count(hops.* where duration_min >= 90)) · last week (2026-03-10–2026-03-16) |
| What did the scan read on the G-meter? | 8pass | scan · scan read g-meter |
| Query the fuel write-ups | 24pass | sql · query fuel write-ups |
| How many fatalities in Colgan Air 3407? | 50pass | count · fatalities colgan air 3407 |
| When did Asiana 214 strike the seawall at San Francisco? | DCA13MA120pass | overview · asiana 214 strike seawall at san |
| How many of these cases are on approach? | 4pass | count · approach |
| Which cases have more than 100 fatalities? | 3pass | list · have more than 100 fatalities |
| Show these cases on a map | 14pass | map · map |
| What did the scan read on the Hudson card? | 8pass | scan · scan read hudson card |
| Query the Hudson case from the database | 155pass | sql · query hudson database |
| Play the Hudson survival count | 0.5pass | audio · play hudson survival count |
| How many wildlife strikes were reported in 2025? | 24447pass | count · wildlife strikes reported 2025 |
| How many wildlife strikes were reported from 1990 through 2025? | 343556pass | count · wildlife strikes reported 1990 through 2025 |
| How many of those strikes were in the United States? | 337882pass | count · strikes united states |
| How many wildlife strikes were reported in 2020? | 11625pass | count · wildlife strikes reported 2020 |
| How many Canada goose strikes were reported in the United States from 1990 through 2025? | 2301pass | count · canada goose strikes reported united states |
| Show the reported strike trend | 24447pass | trend · reported strike trend |
| How many unreported wildlife strikes were there in 2025? | cannot answerpass | count · unreported wildlife strikes 2025 |
| What did the scan read for 2025? | 18pass | scan · scan read 2025 |
| Query reported strikes for 2025 | 24447pass | sql · query reported strikes 2025 |
| How many airlines are in the table? | 56pass | count · table |
| Which airline had the most incidents from 1985 to 1999? | 76pass | list · had most incidents 1985 1999 |
| How many airlines had zero incidents in both periods? | 1pass | count · had zero incidents both periods |
| How many altitude reports are in the desk? | 3pass | count · altitude desk |
| Which report lined up on the taxiway? | 1pass | list · lined up taxiway |
Guarded prose by model family
A model may write the sentence. It may not add a number.
Each family saw only the computed answer JSON for 12 frozen questions and was told to restate it without adding numbers, dates, or names. The completions are frozen. The guard judges them at request time: every number, date, and identifier in the prose must appear in the computed answer, or the prose is rejected and the token is named. The desk never calls a model.
| Family | Model under test | Calls | Accepted | Rejected | Median ms | Introduced tokens |
|---|---|---|---|---|---|---|
| OpenAI | openai/gpt-4.1-mini | 12 / 12 | 11 | 1 | 939 | k-miss-rate: 21 |
| Anthropic | anthropic/claude-haiku-4.5 | 12 / 12 | 12 | 0 | 913 | none |
| google/gemini-3.1-flash-lite | 12 / 12 | 11 | 1 | 694 | k-fuel: 2 | |
| xAI | x-ai/grok-4.3 | 12 / 12 | 11 | 1 | 7075 | k-fuel: 2 |
| Mistral AI | mistralai/mistral-small-2603 | 5 / 12402 | 5 | 0 | 574 | none |
| DeepSeek | deepseek/deepseek-v4-flash | 0 / 12402 | — | — | — | no completions |
| Meta | meta-llama/llama-4-scout | 1 / 12402 | 1 | 0 | 5781 | none |
| Alibaba Qwen | qwen/qwen3.7-flash | 0 / 12402 | — | — | — | no completions |
| Cohere | cohere/command-r-08-2024 | 0 / 12402 | — | — | — | no completions |
| Amazon | amazon/nova-2-lite-v1 | 0 / 12402 | — | — | — | no completions |
120 calls, 54 with text. Families with fewer completions than calls hit HTTP 402 on the account during the run and are frozen as failed, not filled in. Rejections include conservative ones: a derived count such as “two sources” is a number not in the computed answer, so the guard rejects it and says why.