Line Lab Research

Does the Path Predict the Outcome? We Asked the Wrong Question and Found a Better Answer.

The hunch: two markets both at 60¢ — one that rose into it from 20¢, one that fell into it from 80¢ — shouldn't be the same bet. Maybe the path carries information the price hasn't absorbed. It's a clean, testable idea, and the chain lets us test it for free: every resolved market's full price history is public, and its settlement is ground truth. So we took one clean observation per market — the price partway through its life, the move that got it there, and how it actually resolved — across ~4,800 settled binary markets. Here's the whole journey, including the part where we were wrong.

The path hunch: not supported

Within a given price band, we split markets by whether they rose into it (up ≥10 points) or fell into it (down ≥10). If momentum were real, "rose into 60" would resolve Yes more than "fell into 60." It didn't — the cells that moved went the other way, and every one of them was a handful of markets. Noise, not signal. The trajectory, on its own, is already in the price. Idea declined.

But the calibration table wouldn't sit still

Calibration just asks: of all markets priced at X, how many resolve Yes? A perfect market sits on the diagonal — 30¢ markets resolve Yes 30% of the time, no edge. Ours didn't, in one specific place:

PriceResolved YESGap
~5¢~6%+1 (calibrated)
10-20¢~25%+10
20-30¢~39%+14
30-40¢~50%+15
50-90¢small, noisy
~95¢~98%+1 (calibrated)

Cheap longshots — 20-40¢ YES — resolved Yes roughly 8 to 15 points more often than their price. The extremes (5¢, 95¢) are perfectly calibrated, which is exactly what you want to see: the rig isn't just printing "+" everywhere. This is the flip side of a familiar behavior — paying up to short unlikely things, leaving the cheap Yes underbought.

The skeptic's gauntlet

An edge this clean, this easy to find, is usually an artifact. So we tried to kill it. Our first method sampled each market at 66% of its life and tossed out anything already decided — and that filter can manufacture the effect (a market still cheap late in its life is disproportionately one that's still alive to win). So we re-ran it six ways with no such filter:

It held in all of them. The most stubborn number: at one day before close, on markets up to $50M in volume, the 20-40¢ band still resolved Yes about +8 points above its price. That's the cut hardest to explain away as a stale or thinly-traded quote — and it didn't budge. It also survived the date-split, which is the test that has killed every other idea we've chased (our wallet study died on it three times).

The honest wall

Every number above is a midpoint — the middle of the bid and the ask. You don't buy at the midpoint; you buy at the ask, and on a cheap, thin market the gap between them can eat a 10-point edge whole. Historical data physically cannot show us the ask we'd have paid. So the backtest, however clean, cannot answer the only question that matters: does the edge survive the spread you actually cross?

What we're doing about it (in public, before any money)

We don't get to believe it yet. So we built a forward paper-test: every few hours it scans live markets whose ask sits in the 20-40¢ band, records the real price we'd pay — walked through the order book for a real stake, not the touch — along with the spread and the depth, then waits for the market to resolve and scores realized Yes-rate against the price paid. No midpoints, no hindsight. The rule is pre-registered and absolute: not one dollar of real money until 20-40¢-at-the-ask Yes beats the ask over a real forward sample. The first qualifier logged at a 27¢ ask with a one-cent spread — promising, and a sample size of one. We'll let it run for weeks and publish what it says, win or lose. That's the whole method here: a good hunch, an honest gauntlet, and a wall we refuse to bluff past.

Update (July 2): the first 30 verdicts are in — the wall won

We promised to publish the forward result win or lose. It's a lose. The first 30 resolved paper positions — every one filled at the real, order-book-walked ask, spread and all — came back: average price paid 29.3¢, realized Yes-rate 26.7%. That's an edge of −2.6% net of spread (non-sports only, n=26: −2.9%). The backtest said markets like these resolve Yes in the high 30s to 40s; at the executable ask, they resolved Yes below the price we'd have paid. The 8-to-15-point midpoint mirage is fully eaten — by the spread we could never see in historical data, by adverse selection in which asks are actually sittable, or because the effect was a regime artifact all along. From here it doesn't matter which: there is nothing to bet.

Thirty is a small sample and the confidence interval is honest about that — but the pre-registered rule was "no money until the ask-edge confirms," and a negative point estimate is the opposite of confirmation. The paper tracker keeps running for free; we'll re-examine around 60-100 resolutions. If you only remember one thing from this article, make it this: a backtest edge measured at the midpoint is a hypothesis, not an edge — the ask gets a vote, and here it voted no.

Methods: resolved markets and ground-truth settlement from Polymarket's public Gamma API; full price histories from the public CLOB price endpoint; one observation per market to keep samples independent; six eval definitions and four liquidity tiers, no survivorship filter on the robustness pass. The forward test reads the live CLOB order book for the executable ask. Nothing here is betting advice — it's a research log, and the conclusion is "not yet."

← All research