The AI's record, graded live in public
Every pick our engines publish is frozen at the moment it's called, then graded against the real final score. Nothing is edited after the fact. This dashboard is the full accounting — wins, losses, form, and whether the confidence we claimed matched the results we delivered — and it refreshes itself as games finish.
Cryptographic proof — verified in your browser
Loading proof ledger…
Fetching every graded pick and its hash.
Verify it yourself
- This page already did: your browser fetched
/api/track-record/proofand recomputed every hash locally. - Each pick is hashed as
SHA-256(version|engine|game|league|market|side|line|confidence|result|home|away|gradedAt), then chained:chain = SHA-256(prevChain + contentHash). - Change any single stored field — a side, a line, a final score — and its hash plus every hash after it breaks, so the seal no longer matches. That is what makes the record tamper-evident.
Graded record
43-24-1
789 pending
Win rate
64.2%
over 67 decided
Net units
+16.44u
ROI +24.2% per pick
Brier score
0.235
0.250 = coin-flip guessing
Last 7 days
0-1
0.0% · -1.00u
Last 30 days
43-24-1
64.2% · +16.44u
All time
43-24-1
64.2% · +16.44u
Streak
L2
Best W8 · Worst L3
Value Model
751 pendingDeterministic value model, graded against the spread at -110.
- Record
- 30-17-1
- Win rate
- 63.8%
- Units
- +10.27u
- ROI
- +21.4%
- Brier
- 0.234
GPT Engine
38 pendingLLM engine graded on its own market: spread at -110 or the real moneyline price.
- Record
- 13-7
- Win rate
- 65.0%
- Units
- +6.17u
- ROI
- +30.9%
- Brier
- 0.237
Calibration: does the confidence mean anything?
Average gap between claimed confidence and delivered win rate: 9.2 points — reasonably calibrated.
Reliability diagram
Each dot is a confidence bucket (bigger = more picks). The dashed line is perfect calibration — dots below it mean overconfident, above mean underconfident.
Claimed vs delivered, by bucket
Closing Line Value — did we beat the close?
CLV compares the number each pick was published at against the market's closing line. Consistently beating the close is the single hardest-to-fake proof of a real edge — it means we found value before the market agreed.
Beat the close
45.0%
over 20 picks w/ closing data
Avg line CLV
0.0 pts
5 spread picks
Avg price CLV
+0.1 pts
15 moneyline picks
Sample
20
picks with a captured close
Cumulative units
Net profit for each engine as its graded picks settled, in the order games actually finished.
Rolling win rate
Trailing win rate over the last 20 decided picks. Staying above 52.4% beats the standard -110 break-even.
Daily results
Wins and losses settled per day over the last 30 graded days.
Net units by league
Where the profit actually comes from. Green bars are profitable leagues; red are net losers.
Recently graded
Newest first- cs2Home -3Final 0-0 · Value Model · 57% conf · Aug 12, 1:12 AM-1.00uLost
- cs2Home -2Final 0-0 · Value Model · 56% conf · Aug 3, 2:00 AM-1.00uLost
- f1Home -2Final 219-169 · Value Model · 59% conf · Jul 28, 4:08 PM+0.91uWon
- dota2Away -3Final 0-0 · Value Model · 56% conf · Jul 25, 4:05 PM+0.91uWon
- cs2Home -3.5Final 0-0 · Value Model · 57% conf · Jul 25, 4:05 PM-1.00uLost
- cs2Home -1.5Final 0-0 · Value Model · 58% conf · Jul 25, 4:05 PM-1.00uLost
- cs2Away -2Final 0-0 · Value Model · 55% conf · Jul 24, 8:35 AM+0.91uWon
- dota2Away -1.5Final 0-0 · Value Model · 54% conf · Jul 23, 7:42 PM+0.91uWon
- lolAway -3Final 0-0 · Value Model · 57% conf · Jul 23, 12:46 AM+0.91uWon
- cs2Away -0.5Final 0-0 · Value Model · 60% conf · Jul 19, 7:02 PM+0.91uWon
- atpMartinez -1.5Pedro Martinez @ David Jorda Sanchis — Final 0-2 · AI Engine · 75% conf · Jul 18, 2:25 PM-1.00uLost
- wtaPalicova MLBarbora Palicova @ Mai Hontama — Final 2-0 · AI Engine · 80% conf · Jul 18, 2:19 PM+0.54uWon
By league
| League | Record | Win rate | Units |
|---|---|---|---|
| mlb | 17-4-1 | 81.0% | +12.39u |
| wta | 11-6 | 64.7% | +3.82u |
| atp | 4-5 | 44.4% | -1.36u |
| nba | 3-4 | 42.9% | -0.66u |
| cs2 | 2-5 | 28.6% | -3.18u |
| wnba | 2-0 | 100.0% | +1.82u |
| dota2 | 2-0 | 100.0% | +1.82u |
| f1 | 1-0 | 100.0% | +0.91u |
| lol | 1-0 | 100.0% | +0.91u |
How this record is kept
- Picks are recorded the first time they're published — the side, line, and confidence are frozen and never rewritten as the market moves.
- Spread picks are settled against the closing score at the industry-standard -110 price; moneyline picks settle at their real quoted price.
- The rolling win rate tracks the trailing 20 decided picks; 52.4% is the break-even threshold at -110 pricing.
- The Brier score measures whether claimed confidence matched reality: 0 is perfect foresight, 0.250 is what pure 50/50 guessing scores. Lower is better.
- All figures are informational analysis of our model's performance — not betting advice, and no real wagers are placed.