Launching 22 September 2026Kindle pre-order available now ↗

Field story 04 · February–August 2026

Confidently
wrong.

A man, a racing model and several very clever machines.
What could possibly go wrong?

Follow the experiment ↓
A paper horse and jockey emerge from a desktop monitor while a person compares a form guide beside the track.
Ambition, illustrated.
THE SATURDAY EXPERIMENTBY PHIL GRAY · AN ILLUSTRATED RECONSTRUCTION
THE FORM LOOKED GOOD. THE HORSES HAD OTHER IDEAS.

This began with racing. It became a very practical education in working with AI. The dates and numbers follow the surviving records; imagined dialogue is marked reconstructed scene.

THE HUMANPhil

Curious. Impatient. Paying for the lesson.

THE ANALYSTClaude

Finds a pattern. Writes a very persuasive case.

THE ENGINE ROOMKev

Fetches, builds, schedules. Wrestles with the plumbing.

THE LATER ARRIVALCody

Tests the claims. Discovers his own blind spots.

01 28 February → 14 March 2026

A hundred bucks.
A thousand-dollar
idea.

First, Kev made useful race-day spreadsheets. Then the idea grew teeth: give the AI $100 and see whether it could turn that into $1,000 in a year.

There was a name. A website plan. An algorithm. The little operation looked the part.

The money went the other way.

KEV’S SURVIVING LEDGERMAR 2026
THE AMBITION$100 $1,000Target, not an achieved return
7 Mar · start
$100.00
First six bets
$52.63
One more loss
$42.63
14 Mar · close
$28.63

By 14 March, the ledger records $28.63 remaining.

What changedA convincing setup does not make the underlying judgement good.

Open the record Early Kev archive

28 February journal; 7 March project README and audit log.
The journal records Flemington and Randwick spreadsheets. The README sets a $100 starting balance and $1,000 target by March 2027. The project is called Confidently Wrong 2026.

bankroll_history.csv; 7 and 14 March result files.
Balances shown are the archive’s entries, not independently verified account statements. The separate performance summary was still stuck on the earlier $52.63 figure. This ledger is one early experiment, not a complete account of Phil’s racing spend.

02 21 March → 14 April 2026

Good news.
The confidence
was 92%.

Kev admitted sending the wrong Rosehill race data. A few weeks later, his Hong Kong report had the same horse appearing five times in one field.

The accompanying QA report was full of ticks. It declared the data good enough to use.

Someone needed to check the checking.

ARCHIVED QA CLAIM92%“QA PASSED”
Look inside the race report

RACE 1 · FIVE CONSECUTIVE ENTRIES

#10LUCKY ARCHER
#11LUCKY ARCHER
#12LUCKY ARCHER
#13LUCKY ARCHER
#14LUCKY ARCHER

Same name. Five different riders.
These are the report’s entries, not a verified field.

What changedAn official source can still pass through a broken extraction process.

Open the record Kev’s admission and contradictory QA

21 March daily journal.
Kev records confusing the race identity and failing to challenge implausibly uniform data. His explanation blamed test data from the provider; that cause has not been independently established.

Hong Kong report and QA report, written 13 April UTC / 14 April Sydney.
The narrative lists LUCKY ARCHER in slots 10–14. QA says initial parsing counted 448 horses, later claims 120, and assigns 92% overall confidence. The number is the agent’s self-assessment, not a measured accuracy rate.

03 18–25 April 2026

Then it got just good enough
to be irresistible.

18 APRIL · ABBREVIATED BATCH0 / 10

Top picks won.

The cold read had missed the headline.
25 APRIL · LOGGED BOXES3 / 10

Trifecta hits recorded.

The next Saturday offered reasons to keep going.

Claude was finding winners lower down the list, or mentioning them as outsiders. The prose sometimes knew more than the ranking.

So the routine grew: read the form, challenge the case, watch the day unfold, move the picks, write the post-mortem.

A useful hit feels very different from a useful lesson. It makes you want another go.

ReadChallengeUpdateCheckRemember

What changedPromote the useful signal out of the footnote. Then test whether the rule actually helps.

Open the record Claude’s April retrospectives

18 April retrospective; 30 May setup review.
The later review clarifies zero-for-ten refers to abbreviated cold batch top picks, not every full live analysis. The April aggregate contains inconsistent score labels, so this story does not pool its win rates.

19 April Sunshine Coast retrospective.
The Irish was flagged in commentary but not promoted into the ranked four, then won. The note made numerical promotion of live signals central to the next version.

25 April retrospective.
Records three trifecta hits among ten logged boxes. Several races were unlogged and dividends were incomplete. Its positive-day claim is not treated here as a reconciled profit figure.

04 2 May 2026

The machine
kept working.
Phil stepped back.

The model’s “premium” race missed. Some races it had advised skipping produced hits. The labels were sounding more certain than the evidence deserved.

Meanwhile, scheduled updates were firing into their own conversations. Work happened. The answer did not reach the person waiting for it.

Fourteen remaining scheduled jobs were disabled. Phil stopped betting that day; the model carried on for learning.

Reconstructed scene

THE SYSTEMTask completed.

PHILWhere’s the bloody answer?

A job can succeed inside a machine
and fail everywhere that matters.

What changedDelivery is part of the job. So is knowing when to stop paying for another result.

Open the record The mid-day pivot

2 May pre-retrospective summary.
Records the notification failure, disabling 14 remaining fires, switching to the active conversation and continuing in calibration mode after Phil stopped active betting. The header and body disagree about whether he stopped after Race 3 or Race 4, so no exact stopping time is asserted.

05 30 May → 7 June 2026

Turns out the model
also needed
a janitor.

By late May, old instructions pointed at dead folders. A “place lock” rule was still in the documents after the lessons had retired it. The learning database was six weeks stale.

In June, the scheduled report said the racing tables were gone.

You cannot learn from a memory
you cannot find.

A woman fits an orange connector between mismatched pipes beneath a computer surrounded by convoluted paper plumbing.
Illustrated reconstruction
PATHSDeadThe files had moved on.
INSTRUCTIONSArguingThe retired rule had survived.
MEMORYMissingThe next run had nothing to learn from.

What changedLearning needs a working memory, clear ownership and a routine that survives tomorrow.

Open the record The maintenance nobody puts in the demo

30 May setup review.
Records stale session paths, conflicting input locations, a retired rule still active, memory trapped in a read-only mount, missing track profiles and a six-week-old database.

6 June narrative-ingest report, run 7 June.
Reports a paused database, then no application tables after restoration. It records no ingest. This is the contemporary report, not a new forensic finding about who or what removed the tables.

06 22 August 2026 · Kev’s return

Send in more machines.
Surely that will help.

Kev came back with a research team to find trustworthy historical data. Five workers fanned out. The first launch lacked the access it needed. The next ran into the shared token limit.

5research workers
~100calls before one lane failed
1verifier that missed four material problems

Cody caught the misses in the final read. Existing research was salvaged, the claims corrected, and another audit followed. More effort had helped. It had also produced more things to supervise.

The verifier needed a verifier.

What changedDivide the work. Keep someone responsible for whether the pieces make sense together.

Open the record Codex’s dispatch and review history

22 August research thread and Kev’s final sourcing report.
The thread records the failed launch, five concurrent workers exhausting the token allowance, bounded recovery, an initial verifier pass and four material misses found afterwards. Ten tasks ultimately completed, including correction and final acceptance.

The consequential error.
Date-keyed pages had been treated as proof of preserved pre-race data. They were not. A provider’s horse ID had also been confused with a verified official identity. These were corrected in the final report.

07 22 August 2026 · Opening the bonnet

The model was
learning the future.
That helped.

Cody recovered a version of Claude’s model that could actually be reproduced. Then came the suspiciously good improvements.

Some “historical” statistics were not snapshots from before the race. Another apparent breakthrough carried information about the answer into the test.

Brilliant results.
Wrong experiment.

Apparent improvement in prediction score
+0.165THE EXCITING NUMBER
Open the bonnet

INVALIDATED

Reversing summary records created errors related to which horse actually won. The test had picked up a clue from the answer.

The gain was thrown out.

What changedAsk what the model could have known before the event, not what the archive knows now.

Open the record The gains that were refused

Codex experiment report, reproduction and invalidated gains.
All 717 handoff checksums passed. BLS v5.1 was reproducible; later simplified weights were not reproducible from the preserved clean matrix.

Two invalidated gains.
Published jockey/trainer aggregates appeared to improve log loss by 0.02451 but lacked point-in-time integrity. Reversed context summaries appeared to improve it by 0.16519 and were target-label dependent. Neither entered the frozen candidate.

08 22 August 2026 · Less magic, better questions

A smaller improvement.
A harder-earned maybe.

Two prior-form signals offered a modest gain across 641 validation races. One time period still went backwards. The market benchmark remained better on the tested price subset.

And the live hot-jockey / hot-stable rule that April had loved? Its specified version failed the locked August test.

This time, improvement included
taking something out.

APRIL

“Make the live signal mandatory.”

AUGUST · 610-RACE TEST

Remove the tested day-signal layer.

Editorial summaries of the changing decisions
BASELINE2.043
CANDIDATE2.020

Average log loss across 641 validation races.
Lower is better. Promising is not proven.

What changedKeep the failed experiments. They are the reason the next version has fewer bad ideas.

Open the record A promising candidate and a rejected rule

Codex experiment report, 22 August.
Rolling validation improved from 2.04319 to 2.02019. Gains across the three chronological folds were +0.03108, +0.03264 and −0.00153. The exposed 337-race block was not an untouched test.

Locked day-layer audit.
610 races across 83 meeting clusters. The quality-adjusted layer worsened log loss by 0.00350; the inherited fixed-count rule worsened it by 0.00749. The specified Step 4A rule was deleted under the predeclared decision. This does not disprove every possible use of within-day information.

09 23–25 August 2026

Eighteen million
simulated finishes.
No final verdict.

Ninety-four future races were locked before their results. The model was frozen. The scoring rules were written. The rehearsal worked.

Then the real-world capture failed. One process stopped before packaging its archive. Required scripts were absent on the following days.

The experiment could not be scored through its agreed process. After all that work, the honest result was an operational failure.

THE LOCKED TRIAL94

races · 1,135 runners

18.8 million simulated orders
23 AUG
INCOMPLETE
24 AUG
MISSING
25 AUG
MISSING

NO MODEL VERDICT

What changedA rehearsed workflow still has to work on the day. Missing evidence does not become a win or a loss.

Open the record The trial that could not be opened

25 August blocked-opening receipt.
All three independently audited dated outcome archives were unavailable. No canonical final comparison or final model verdict was produced. Three morning operational misses were retained; no successful morning manifest existed.

Frozen prospective records.
94 races, 1,135 runners, 100,000 finishing-order simulations per race per model: 18.8 million orders across two models. These are simulated possibilities, not observed trials. The blockage says nothing conclusive about which model was better.

10 28–29 August 2026 · Back at Rosehill

Four minutes to the jump.
Don’t break the model.

PHIL · 12:16 PM

“4 minutes to race 2. We need updated predictions”

PHIL · 12:17 PM

“But don’t break the model.”

Claude analysed overnight. Cody kept the evidence straight. The morning prediction stayed frozen; a separate live view could react to new information. Phil supplied results when official feeds lagged.

THE MORNING CARDKeep it frozen.

What did we say before we knew?

THE LIVE VIEWLet it respond.

What would we say with this new information?

The separation mattered. Claude’s own local numbers were stale. On the last race, they would have left out the winner. The controlling card included it.

What changedYou can adapt and remain honest, provided you keep the original prediction visible.

Open the record The actual race-day conversation

Codex history, 29 August.
Phil’s messages arrived at 12:16 and 12:17 AEST. He later supplied final results and explicitly asked for official confirmation afterwards. The quotations above are verbatim.

Claude end-of-day retrospective and controlling audit.
Claude’s local Race 10 shortlist omitted Lisztomania; the controlling final-field card included it. The primary card was preserved. The audit also recorded a missed pre-jump review rather than manufacturing one after the race.

11 29 August 2026 · A side experiment

Eighty-one per cent.
Sounds good.

A side experiment counted possible finishes for a spread of bets. More than 81% of the combinations returned more than the stake.

One small problem: those finishes were not equally likely.

When the quoted prices were used to weight the outcomes, the diagnostic pointed the other way.

THE NUMBER THAT GETS YOUR ATTENTION81.1%

of enumerated finishes above break-even

And the question that changes it?

How likely
is each finish?

2,453 of 3,024 possible top-four orders were above break-even. Counting them equally ignores that some are much more likely.

$32.00hypothetical stake$26.15price-weighted expected return

The latter is a quoted-price diagnostic, not a validated probability model or a realised return.

What changedCount the possibilities. Then ask how much weight each one deserves.

Open the record The coverage-bet experiment

29 August Merrylands R3 coverage receipt.
Nine runners in the field, eight backed with the favourite excluded, four $1 markets per backed runner. It enumerated 3,024 ordered top-four outcomes; 2,453 (81.1177%) exceeded the $32 stake.

The important distinction.
Multiplicative normalisation of each quoted market gave a price-based expected return of $26.15, or −$5.85 net. Prices were held fixed; late deductions and dead heats were not modelled. Neither outcome counts nor this price diagnostic prove an executable edge.

12 29 August 2026 · What the finish line said

Right horses.
Wrong confidence.

3 / 9Top picks won
7 / 9Winners in the first four picks
2.248Model log loss · 2.246 for equal chance

The shortlist found plenty. But two ugly misses left the average probability score slightly worse than an equal-chance benchmark. Finding a horse and knowing how much to believe in it are different skills.

The cost of being surprisedWinner log loss by race · lower is better
1.88
R1top pick
R2excluded
1.90
R3in four
2.06
R4in four
1.97
R5top pick
4.13
R6miss
2.84
R7miss
1.44
R8top pick
1.72
R9in four
2.28
R10in four

R6: Oh Arthur won after being given just 1.6%. R7: Plagiarism won at 5.8% in the model.

Race 9 · A moment worth keeping

All four.
Different order.

The frozen shortlist contained every horse in the actual first four.

MODEL RANK
  1. I’mintowin
  2. Fully Lit
  3. Althoff
  4. Rolling Magic
ACTUAL FINISH
  1. Althoff
  2. I’mintowin
  3. Rolling Magic
  4. Fully Lit

What changedA good shortlist deserves credit. It does not excuse overconfidence or establish a profitable edge.

Open the record The reconciled day, not the best-looking screenshot

29 August end-of-day audit against archived Australian Turf Club results.
Ten races completed, nine scored. Race 2 was excluded because four official starters were absent from the locked prediction. Three top-pick wins and seven winners in four are both measured over the nine eligible races.

Probability score.
Mean winner log loss 2.248171. Equal-chance benchmark is the mean of log(field size) over the same nine races, approximately ${uniform.toFixed(4)}. Claude’s retrospective instead quoted 2.2034; this page recomputes it from the official starter counts for the nine eligible races. A lower loss is better. This is one meeting, not an estimate of long-run betting returns.

Correcting the track input did not settle the argument.
The Good 4 diagnostic was very slightly better on scoring rules over Races 4–10, but winner-in-four fell from five of seven to four of seven. A sensible correction did not deliver a clear practical gain.

13 What survived

The bets were
the tuition.

The judgement was
the certificate.

The early book draft says the experiment lost money. The archive does not give us one clean, complete lifetime profit-and-loss account. There is no grand victory lap to invent.

What it does show is a changing relationship with the machines: less accepting, more testing; less impressed by a confident answer, more interested in what survives a check.

The useful part came to work on Monday.

An older person makes notes with an orange pencil on a racecourse bench beside an empty track.
Illustrated reconstruction
01Form your own view.

Before the consensus gets a vote.

02Give doubt a job.

Someone must try to break the case.

03Keep the first answer.

Make revision visible.

04Let a failure count.

Especially when the story sounds good.

Open the record Phil’s earlier telling

The Saturday experiment, July book archive.
Describes losing money while learning habits that transferred into the working week. “The bets were the tuition. The judgement was the certificate.” is from that draft. No manuscript file was changed to make this page.

About this telling

Built from Claude’s project notes and earlier book files, Kev’s surviving OpenClaw and Hermes records, and Codex’s conversations and experiment audits. The documented trail used here begins on 28 February 2026 and ends on 29 August. A folder named “Racing Jan26” does not establish a January start date.

Illustrations, character sketches and marked dialogue are creative reconstructions. Agent-written retrospectives are treated as contemporary claims; contradictory records stay visible. This is the story of learning through an experiment, not a demonstration of a profitable betting system.

Bring the useful part into your week

What would you
put to the test?

Plan your own experiment ↗More field stories ↗