Field story 04 · February–August 2026
Confidently
wrong.
A man, a racing model and several very clever machines.
What could possibly go wrong?

This began with racing. It became a very practical education in working with AI. The dates and numbers follow the surviving records; imagined dialogue is marked reconstructed scene.
Curious. Impatient. Paying for the lesson.
Finds a pattern. Writes a very persuasive case.
Fetches, builds, schedules. Wrestles with the plumbing.
Tests the claims. Discovers his own blind spots.
01 28 February → 14 March 2026
A hundred bucks.
A thousand-dollar
idea.
First, Kev made useful race-day spreadsheets. Then the idea grew teeth: give the AI $100 and see whether it could turn that into $1,000 in a year.
There was a name. A website plan. An algorithm. The little operation looked the part.
The money went the other way.
By 14 March, the ledger records $28.63 remaining.
What changedA convincing setup does not make the underlying judgement good.
↳ Open the record Early Kev archive
28 February journal; 7 March project README and audit log.
The journal records Flemington and Randwick spreadsheets. The README sets a $100 starting balance and $1,000 target by March 2027. The project is called Confidently Wrong 2026.
bankroll_history.csv; 7 and 14 March result files.
Balances shown are the archive’s entries, not independently verified account statements. The separate performance summary was still stuck on the earlier $52.63 figure. This ledger is one early experiment, not a complete account of Phil’s racing spend.
02 21 March → 14 April 2026
Good news.
The confidence
was 92%.
Kev admitted sending the wrong Rosehill race data. A few weeks later, his Hong Kong report had the same horse appearing five times in one field.
The accompanying QA report was full of ticks. It declared the data good enough to use.
Someone needed to check the checking.
Look inside the race report
RACE 1 · FIVE CONSECUTIVE ENTRIES
Same name. Five different riders.
These are the report’s entries, not a verified field.
What changedAn official source can still pass through a broken extraction process.
↳ Open the record Kev’s admission and contradictory QA
21 March daily journal.
Kev records confusing the race identity and failing to challenge implausibly uniform data. His explanation blamed test data from the provider; that cause has not been independently established.
Hong Kong report and QA report, written 13 April UTC / 14 April Sydney.
The narrative lists LUCKY ARCHER in slots 10–14. QA says initial parsing counted 448 horses, later claims 120, and assigns 92% overall confidence. The number is the agent’s self-assessment, not a measured accuracy rate.
03 18–25 April 2026
Then it got just good enough
to be irresistible.
Top picks won.
The cold read had missed the headline.Trifecta hits recorded.
The next Saturday offered reasons to keep going.Claude was finding winners lower down the list, or mentioning them as outsiders. The prose sometimes knew more than the ranking.
So the routine grew: read the form, challenge the case, watch the day unfold, move the picks, write the post-mortem.
A useful hit feels very different from a useful lesson. It makes you want another go.
What changedPromote the useful signal out of the footnote. Then test whether the rule actually helps.
↳ Open the record Claude’s April retrospectives
18 April retrospective; 30 May setup review.
The later review clarifies zero-for-ten refers to abbreviated cold batch top picks, not every full live analysis. The April aggregate contains inconsistent score labels, so this story does not pool its win rates.
19 April Sunshine Coast retrospective.
The Irish was flagged in commentary but not promoted into the ranked four, then won. The note made numerical promotion of live signals central to the next version.
25 April retrospective.
Records three trifecta hits among ten logged boxes. Several races were unlogged and dividends were incomplete. Its positive-day claim is not treated here as a reconciled profit figure.
04 2 May 2026
The machine
kept working.
Phil stepped back.
The model’s “premium” race missed. Some races it had advised skipping produced hits. The labels were sounding more certain than the evidence deserved.
Meanwhile, scheduled updates were firing into their own conversations. Work happened. The answer did not reach the person waiting for it.
Fourteen remaining scheduled jobs were disabled. Phil stopped betting that day; the model carried on for learning.
Reconstructed scene
THE SYSTEMTask completed.
PHILWhere’s the bloody answer?
A job can succeed inside a machine
and fail everywhere that matters.
What changedDelivery is part of the job. So is knowing when to stop paying for another result.
↳ Open the record The mid-day pivot
2 May pre-retrospective summary.
Records the notification failure, disabling 14 remaining fires, switching to the active conversation and continuing in calibration mode after Phil stopped active betting. The header and body disagree about whether he stopped after Race 3 or Race 4, so no exact stopping time is asserted.
05 30 May → 7 June 2026
Turns out the model
also needed
a janitor.
By late May, old instructions pointed at dead folders. A “place lock” rule was still in the documents after the lessons had retired it. The learning database was six weeks stale.
In June, the scheduled report said the racing tables were gone.
You cannot learn from a memory
you cannot find.

What changedLearning needs a working memory, clear ownership and a routine that survives tomorrow.
↳ Open the record The maintenance nobody puts in the demo
30 May setup review.
Records stale session paths, conflicting input locations, a retired rule still active, memory trapped in a read-only mount, missing track profiles and a six-week-old database.
6 June narrative-ingest report, run 7 June.
Reports a paused database, then no application tables after restoration. It records no ingest. This is the contemporary report, not a new forensic finding about who or what removed the tables.
06 22 August 2026 · Kev’s return
Send in more machines.
Surely that will help.
Kev came back with a research team to find trustworthy historical data. Five workers fanned out. The first launch lacked the access it needed. The next ran into the shared token limit.
Cody caught the misses in the final read. Existing research was salvaged, the claims corrected, and another audit followed. More effort had helped. It had also produced more things to supervise.
What changedDivide the work. Keep someone responsible for whether the pieces make sense together.
↳ Open the record Codex’s dispatch and review history
22 August research thread and Kev’s final sourcing report.
The thread records the failed launch, five concurrent workers exhausting the token allowance, bounded recovery, an initial verifier pass and four material misses found afterwards. Ten tasks ultimately completed, including correction and final acceptance.
The consequential error.
Date-keyed pages had been treated as proof of preserved pre-race data. They were not. A provider’s horse ID had also been confused with a verified official identity. These were corrected in the final report.
07 22 August 2026 · Opening the bonnet
The model was
learning the future.
That helped.
Cody recovered a version of Claude’s model that could actually be reproduced. Then came the suspiciously good improvements.
Some “historical” statistics were not snapshots from before the race. Another apparent breakthrough carried information about the answer into the test.
Brilliant results.
Wrong experiment.
Open the bonnet
INVALIDATED
Reversing summary records created errors related to which horse actually won. The test had picked up a clue from the answer.
The gain was thrown out.
What changedAsk what the model could have known before the event, not what the archive knows now.
↳ Open the record The gains that were refused
Codex experiment report, reproduction and invalidated gains.
All 717 handoff checksums passed. BLS v5.1 was reproducible; later simplified weights were not reproducible from the preserved clean matrix.
Two invalidated gains.
Published jockey/trainer aggregates appeared to improve log loss by 0.02451 but lacked point-in-time integrity. Reversed context summaries appeared to improve it by 0.16519 and were target-label dependent. Neither entered the frozen candidate.
08 22 August 2026 · Less magic, better questions
A smaller improvement.
A harder-earned maybe.
Two prior-form signals offered a modest gain across 641 validation races. One time period still went backwards. The market benchmark remained better on the tested price subset.
And the live hot-jockey / hot-stable rule that April had loved? Its specified version failed the locked August test.
This time, improvement included
taking something out.
“Make the live signal mandatory.”
Remove the tested day-signal layer.
Average log loss across 641 validation races.
Lower is better. Promising is not proven.
What changedKeep the failed experiments. They are the reason the next version has fewer bad ideas.
↳ Open the record A promising candidate and a rejected rule
Codex experiment report, 22 August.
Rolling validation improved from 2.04319 to 2.02019. Gains across the three chronological folds were +0.03108, +0.03264 and −0.00153. The exposed 337-race block was not an untouched test.
Locked day-layer audit.
610 races across 83 meeting clusters. The quality-adjusted layer worsened log loss by 0.00350; the inherited fixed-count rule worsened it by 0.00749. The specified Step 4A rule was deleted under the predeclared decision. This does not disprove every possible use of within-day information.
09 23–25 August 2026
Eighteen million
simulated finishes.
No final verdict.
Ninety-four future races were locked before their results. The model was frozen. The scoring rules were written. The rehearsal worked.
Then the real-world capture failed. One process stopped before packaging its archive. Required scripts were absent on the following days.
The experiment could not be scored through its agreed process. After all that work, the honest result was an operational failure.
races · 1,135 runners
18.8 million simulated ordersINCOMPLETE24 AUG
MISSING25 AUG
MISSING
NO MODEL VERDICT
What changedA rehearsed workflow still has to work on the day. Missing evidence does not become a win or a loss.
↳ Open the record The trial that could not be opened
25 August blocked-opening receipt.
All three independently audited dated outcome archives were unavailable. No canonical final comparison or final model verdict was produced. Three morning operational misses were retained; no successful morning manifest existed.
Frozen prospective records.
94 races, 1,135 runners, 100,000 finishing-order simulations per race per model: 18.8 million orders across two models. These are simulated possibilities, not observed trials. The blockage says nothing conclusive about which model was better.
10 28–29 August 2026 · Back at Rosehill
Four minutes to the jump.
Don’t break the model.
Claude analysed overnight. Cody kept the evidence straight. The morning prediction stayed frozen; a separate live view could react to new information. Phil supplied results when official feeds lagged.
What did we say before we knew?
What would we say with this new information?
The separation mattered. Claude’s own local numbers were stale. On the last race, they would have left out the winner. The controlling card included it.
What changedYou can adapt and remain honest, provided you keep the original prediction visible.
↳ Open the record The actual race-day conversation
Codex history, 29 August.
Phil’s messages arrived at 12:16 and 12:17 AEST. He later supplied final results and explicitly asked for official confirmation afterwards. The quotations above are verbatim.
Claude end-of-day retrospective and controlling audit.
Claude’s local Race 10 shortlist omitted Lisztomania; the controlling final-field card included it. The primary card was preserved. The audit also recorded a missed pre-jump review rather than manufacturing one after the race.
11 29 August 2026 · A side experiment
Eighty-one per cent.
Sounds good.
A side experiment counted possible finishes for a spread of bets. More than 81% of the combinations returned more than the stake.
One small problem: those finishes were not equally likely.
When the quoted prices were used to weight the outcomes, the diagnostic pointed the other way.
of enumerated finishes above break-even
And the question that changes it?
How likely
is each finish?
2,453 of 3,024 possible top-four orders were above break-even. Counting them equally ignores that some are much more likely.
The latter is a quoted-price diagnostic, not a validated probability model or a realised return.
What changedCount the possibilities. Then ask how much weight each one deserves.
↳ Open the record The coverage-bet experiment
29 August Merrylands R3 coverage receipt.
Nine runners in the field, eight backed with the favourite excluded, four $1 markets per backed runner. It enumerated 3,024 ordered top-four outcomes; 2,453 (81.1177%) exceeded the $32 stake.
The important distinction.
Multiplicative normalisation of each quoted market gave a price-based expected return of $26.15, or −$5.85 net. Prices were held fixed; late deductions and dead heats were not modelled. Neither outcome counts nor this price diagnostic prove an executable edge.
12 29 August 2026 · What the finish line said
Right horses.
Wrong confidence.
The shortlist found plenty. But two ugly misses left the average probability score slightly worse than an equal-chance benchmark. Finding a horse and knowing how much to believe in it are different skills.
R6: Oh Arthur won after being given just 1.6%. R7: Plagiarism won at 5.8% in the model.
Race 9 · A moment worth keeping
All four.
Different order.
The frozen shortlist contained every horse in the actual first four.
- I’mintowin
- Fully Lit
- Althoff
- Rolling Magic
- Althoff
- I’mintowin
- Rolling Magic
- Fully Lit
What changedA good shortlist deserves credit. It does not excuse overconfidence or establish a profitable edge.
↳ Open the record The reconciled day, not the best-looking screenshot
29 August end-of-day audit against archived Australian Turf Club results.
Ten races completed, nine scored. Race 2 was excluded because four official starters were absent from the locked prediction. Three top-pick wins and seven winners in four are both measured over the nine eligible races.
Probability score.
Mean winner log loss 2.248171. Equal-chance benchmark is the mean of log(field size) over the same nine races, approximately ${uniform.toFixed(4)}. Claude’s retrospective instead quoted 2.2034; this page recomputes it from the official starter counts for the nine eligible races. A lower loss is better. This is one meeting, not an estimate of long-run betting returns.
Correcting the track input did not settle the argument.
The Good 4 diagnostic was very slightly better on scoring rules over Races 4–10, but winner-in-four fell from five of seven to four of seven. A sensible correction did not deliver a clear practical gain.
13 What survived
The bets were
the tuition.
The judgement was
the certificate.
The early book draft says the experiment lost money. The archive does not give us one clean, complete lifetime profit-and-loss account. There is no grand victory lap to invent.
What it does show is a changing relationship with the machines: less accepting, more testing; less impressed by a confident answer, more interested in what survives a check.
The useful part came to work on Monday.

Before the consensus gets a vote.
Someone must try to break the case.
Make revision visible.
Especially when the story sounds good.
↳ Open the record Phil’s earlier telling
The Saturday experiment, July book archive.
Describes losing money while learning habits that transferred into the working week. “The bets were the tuition. The judgement was the certificate.” is from that draft. No manuscript file was changed to make this page.
About this telling
Built from Claude’s project notes and earlier book files, Kev’s surviving OpenClaw and Hermes records, and Codex’s conversations and experiment audits. The documented trail used here begins on 28 February 2026 and ends on 29 August. A folder named “Racing Jan26” does not establish a January start date.
Illustrations, character sketches and marked dialogue are creative reconstructions. Agent-written retrospectives are treated as contemporary claims; contradictory records stay visible. This is the story of learning through an experiment, not a demonstration of a profitable betting system.
Bring the useful part into your week