And we grade ourselves by the same rule, in public, every night — including the nights we lose.
Brier score at sixty seconds to expiry, pooled across all 27,528 resolved windows archived since 13 July 2026. Lower is better.
The pipeline behind every number on this page. Both prices are archived while the window is still open — a prediction that arrives after the outcome simply never enters the record.
Why the order matters: grading only counts because step 02 happens before step 03 — every prediction is committed while the answer is still unknown, which is what makes the record impossible to retrofit.
Every rule is replayed over every archived session at one dollar a fill, and reported with a 95% confidence interval. Two are deliberate controls — a coin flip and a buy-the-favourite. If the harness were leaking, the decoys would drift above zero. They do the opposite: the coin flip loses 0.8¢ a fill — the cost of crossing the spread — and the favourite loses 1.8¢ even while winning 91% of its fills, because comfort trades at a premium. That baseline is what every live rule has to clear.
A side panel that sits next to the market page and does one thing: price the same question independently, and show you where its answer and the book's answer stop agreeing.
σ shrinks as √t, so the last ninety seconds move the probability more than the first ten minutes. The panel shows you the countdown because it is an input, not decoration.
Our probability minus the book's, in points. Not a recommendation — the measured size of a disagreement, stated in the unit both sides quote.
Ours and the book's, on one axis, for the life of the window. When they sit on top of each other — which is most of the time — you can see that too.
A multi-exchange feed at roughly a hundred milliseconds, with each source shown separately so you can see when one of them is the outlier rather than the market.
Price to beat, distance in dollars and in σ, σ over the remaining horizon, annualised and realised volatility. Everything the probability was computed from, so you can disagree with it.
The venue layer is built to be swapped — feeds, order book and resolution — so the same measurement can run anywhere binary markets trade. Polymarket today; Kalshi is next on the roadmap and opens the regulated US market.
If you can't check it, you shouldn't trust it — so here it is, including the parts where it breaks.
Crypto never closes, so volatility is annualised over 365×24×3600 seconds. Borrowing the equity market's 252 trading days here is a factor-of-two error, and it survives in a lot of published crypto option maths.
Variance is additive in time, so the standard deviation grows as σ·√t. Over a fifteen-minute window drift is negligible against noise, so we set it to zero rather than pretend to estimate it. Above two hours we say so on screen.
A Gaussian tail underprices large moves by one to two orders of magnitude. We also run Student-t with ν=4, renormalised so its variance still matches σ. LTCM is the reference case: a thin-tailed model in a fat-tailed world, at 25× leverage.
Liquidation cascades are jumps; no diffusion captures them — the fat tail only softens the blow. Sub-minute sampling picks up bid-ask bounce rather than signal. If returns autocorrelate, √t itself breaks. We publish through those nights too.
A binary market resolves against a single published price. That feed is authoritative, and it is also slower than the venues it summarises — between two of its updates it is quoting a price the market has already left. We blend five exchange books into one composite at ~100 ms, recalibrated to the feed's own scale, which usually means we know where the feed is going to land before it lands there.
The pale lines are five exchange books; the amber line is the composite — median of the five, recalibrated to the feed's scale — which is what the model prices off. The blue staircase is the resolution feed — flat until it republishes, then a jump to catch up. Every shaded band is an interval where the two disagree, and the direction of that disagreement is the signal. Measured over the archived session of 13 July 2026: correlation 0.62 between the composite lead and the subsequent feed move, 75% directional hit. The tick sequence below is a seeded illustration of the mechanism, not a recording — the two statistics are the measured ones.
None of these are opinions. They fall straight out of the maths above, and they are the reason a window that looks settled with two minutes left often isn't. Scroll each one to run it.
σ scales as √t, not as t. Burn three quarters of the clock and you have not burned three quarters of the risk — you still carry $90.86 of it, exactly twice the $45.43 a linear reading would give you. It is why the last ninety seconds of a window move the price more than the first ten minutes, and why a market that looks decided usually isn't yet.
Both curves have the same σ — the Student-t is renormalised so its variance matches. They only disagree about how often the extremes arrive, and the disagreement is invisible until it isn't: they track each other to about 2σ, cross 10× at 3.3σ and 100× at 4.1σ. Log scale on the vertical axis, or the tails would be an invisible smear along the floor. This is the LTCM shape: a thin-tailed model in a fat-tailed world.
120 paths of the same model, all starting half a σ below the price to beat, all seeded identically so this figure is the same every time you load it. The reflection principle says P(touch) = 2 × P(finish beyond) for a driftless walk; the sample lands on 74 and 37. That is why the panel quotes both numbers — a window can spend most of its life on the right side of the line and still resolve on the wrong one.
The free tier is meant to be enough to convince you the measurement is real. It is not meant to be enough to act on.
Every window archived and graded the night it resolves. The record is 26 sessions deep and has been public since the first one — read it before you believe a word of this.
Read the nightly record →