The Consensus: Three AIs Argue Over Every Bet We Fire. Can Their Score Predict Our P&L?
Sunday, August 2, 2026
Our betting rules are deliberately dumb: mechanical triggers on ticket-vs-money divergence, no opinions allowed. The Consensus experiment asks whether adding an informed layer on top — actual pitching matchups, injuries, park factors, the stuff the splits can't see — carries any signal at all. We are not letting it touch money. We're making it show its work in public first.
Exactly what happens on every fire
When a Rule A/B/C trigger fires pre-game, a pipeline runs before first pitch (a leak guard refuses to score any game that has started — a score computed after results exist would be worthless):
- Research pass: a search-connected model (Perplexity) pulls the day's real-world context — probable starters and their recent form, lineups, injuries, park and weather.
- The debate: two frontier models argue opposite sides. One (GPT) must make the strongest honest case for the underdog we're about to back; the other (Claude Opus) must make the case for the favorite. Neither knows our position size or record.
- The rank: a third model (Claude Fable) reads both cases blind and produces two numbers: an independent win probability for the dog, and a 1-10 confidence score for the play.
The blind win probability matters because we can compare it to the price we'd actually pay on Polymarket. A panel that says 0.55 on a dog asking 0.46 is claiming nine points of edge; a panel that says 0.48 on a 0.46 ask is calling it a coin flip at fair price.
What would make this a success (and what wouldn't)
The primary test — fixed before the experiment started — is rank versus P&L, not win rate: after ~100 scored plays, do the high-scored fires make more money than the low-scored ones? If yes, the score becomes a sizing input. If no, we'll have bought the answer for $0, because none of this touches a bet. The gate lands around mid-September.
The live ledger
Every fire, its Polymarket ask at score time, the panel's blind probability, the score, and the outcome — updated at every nightly grade:
| Date | Game | Our dog @ ask | Panel win prob | Score | Result |
|---|---|---|---|---|---|
| 2026-08-04 | Twins@Royals | Royals @0.43 | 0.42 | 4/10 | WON |
| 2026-08-04 | Athletics@Reds | Athletics @0.46 | 0.44 | 4/10 | LOST |
| 2026-08-02 | Royals@Rockies | Royals @0.46 | 0.55 | 8/10 | LOST |
| 2026-08-02 | Cardinals@Blue Jays | Cardinals @0.45 | 0.51 | 7/10 | WON |
| 2026-08-01 | Royals@Rockies | Royals @0.46 | 0.48 | 5/10 | LOST |
Scored so far: 5 plays; graded record 2-3 (−0.45u at the recorded asks).
Since August 5 the same panel also scores two other populations — every remaining game on the board as a control arm, and the sides X handicappers take. Those are separate experiments on separate ledgers and they never touch the pre-registered pool above; they're written up in Crossing Two Experiments.
Honesty notes, same as everywhere on this site: the sample is tiny and means nothing yet — the whole point is the preregistered 100-play gate. The panel's scores are logged before results exist and never edited. No bet is placed, sized, or skipped because of this experiment while it's running. Nothing here is betting advice.
← All research