PROMOTEΔ Elo +37±11vs anchor · live 813,000 Elo 1196 · 20238g · sims 1600ops full gate →
glance
DOWN
trainer
813,040
step
8,905 steps/h
19,575,563 samples
window
42.5 ep
0.774
eval policy
train 0.828
913
checkpoints
1491 published
0
games/h
last hour
policy train / eval
series: train policy ·
eval policy
Policy loss train vs eval (by step)
More metrics & experiments
league glance
#0 · 925,000 [07-25a]Elo 1236
405 fit games · stem ckpt-2026-07-25a-00925000
#1 · 913,000 [07-25a]Elo 1222
534 fit games · stem ckpt-2026-07-25a-00913000
#2 · 863,000 [07-25a]Elo 1216
173 fit games · stem ckpt-2026-07-25a-00863000
full board + disable toggles under Detail · 308,268 total games
secondary charts
Draft / value loss (by step)Trainer progress (steps over hours)Replay window (samples over hours)Learning rate (by step, log y)
Detail — league matrix, fleet knobs, decks, mix
fleet full
0
worker jobs
24k evals/s
inference
all servers
0
replay files
≈games
—
eval draft
uniform = 7.14
0.612
eval value
random ≈ 1.0
Fixed-deck league (meta pool, raw policy) — 308,268 games · leader ckpt 925,000 [07-25a] (ckpt-2026-07-23-00813000) (1235.5) · raw-policy play, random seat order · both lineages · updated 2026-07-26T19:04:07Z
Replay window — composition of the 10000-file training window (sampled 150 files at 2026-07-22T18:29:53Z)
—
forced-LB seats
of 0 sampled
0.0
deck energy
p10/50/90 0/0/0
0.0
deck distinct
p10/50/90 0/0/0
48.4%
P0 win share
11736 games in span
94
turns p50
p10/90 87/95
86.6%
draws
forced-LB deck matches: —
what these metrics mean
forced-LB seats — share of drafted decks in the training window
that were forced to copy a real top-leaderboard deck (the force_lbdraft_p
knob; target equals the knob value once the window fully turns over).
These are the "teacher" decks anchoring self-play to the live Kaggle meta.
The team list below shows which leaderboard teams' decks were matched.
deck energy — basic-energy cards per drafted deck (60 cards
total). Pokémon need energy attached to attack; a deck with ~0 energy
can never attack. Real top decks run ~10–14. Near-zero means the
self-drafted decks are non-functional.
deck distinct — distinct card names per deck. Real decks are
built around a plan: ~19–21 distinct cards in focused multiples. ~55+
distinct means essentially random piles (the draft policy expressing no
preferences); ~2–5 means degenerate single-card spam. Healthy is in
between.
P0 win share — fraction of games won by the player who moves
first. Should be ~50%. When games end by running out of cards
("deck-out"), turn order alone decides the winner and this pins near
90% — a signal that outcomes carry no information about play quality,
which poisons value-head training.
turns p50 — median game length. Real fights end in ~15–40
turns by taking prizes; ~95 turns means games are grinding to deck-out
(nobody can attack effectively).
draws — games with no winner. Should be ~0%; rising draws
mean stalling is back.
Reading the panel as one story:
the window is healthy when forced-LB share ≈ the knob, free-drafted decks
drift toward real-deck shape (energy 10–14, distinct ~20), P0 share ≈ 50%,
and turns settle in the 15–40 band. Deviations tell you which part of the
draft→play→outcome chain is broken.
LB deck pool — 19 decks sampled by force_lbdraft_p (teams: 11; kept 19 of 28440 seen)
Yushin Ito— ep cbaf69e8-8000-11f1-b6d5-0242ac130203 · Q 98 · 10 energy · 21 distinct · 9 attackers
10×
Basic {D} Energy
#7
4×
Marnie's Impidimp
#646
4×
Marnie's Grimmsnarl ex
#648
4×
Rare Candy
#1079
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Lillie's Determination
#1227
3×
Dudunsparce
#66
3×
Munkidori
#112
3×
Dunsparce
#305
3×
Dawn
#1231
3×
Spikemuth Gym
#1259
2×
Marnie's Morgrem
#647
2×
Boss’s Orders
#1182
1×
Fezandipiti ex
#140
1×
Budew
#235
1×
Yveltal
#689
1×
Tool Scrapper
#1137
1×
Hero’s Cape
#1159
1×
Xerosic’s Machinations
#1197
1×
Risky Ruins
#1260
bono— ep 07da3a98-8118-11f1-bfd8-0242ac130203 · Q 98 · 10 energy · 19 distinct · 7 attackers
10×
Basic {D} Energy
#7
4×
Munkidori
#112
4×
Marnie's Impidimp
#646
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Lillie's Determination
#1227
4×
Dawn
#1231
4×
Spikemuth Gym
#1259
3×
Dunsparce
#305
3×
Marnie's Morgrem
#647
3×
Marnie's Grimmsnarl ex
#648
3×
Rare Candy
#1079
2×
Dudunsparce
#66
2×
Night Stretcher
#1097
2×
Xerosic’s Machinations
#1197
1×
Marnie's Morpeko
#649
1×
Energy Search
#1119
1×
Energy Recycler
#1139
1×
Hero’s Cape
#1159
THIRD PTCG Club— ep 18ab4dbe-807b-11f1-b095-0242ac130203 · Q 97 · 14 energy · 21 distinct · 6 attackers
7×
Basic {G} Energy
#1
4×
Team Rocket's Energy
#15
4×
Team Rocket's Tarountula
#400
4×
Team Rocket's Spidops
#401
4×
Team Rocket's Transceiver
#1134
4×
Poké Pad
#1152
4×
Team Rocket's Ariana
#1216
4×
Team Rocket's Proton
#1220
3×
Basic {P} Energy
#5
3×
Bug Catching Set
#1094
3×
Team Rocket's Giovanni
#1218
2×
Team Rocket's Articuno
#414
2×
Team Rocket's Mewtwo ex
#431
2×
Team Rocket's Murkrow
#463
2×
Energy Search
#1119
2×
Lillie's Determination
#1227
2×
Team Rocket's Factory
#1257
1×
Team Rocket's Wobbuffet
#432
1×
Night Stretcher
#1097
1×
Hero’s Cape
#1159
1×
Team Rocket's Archer
#1217
THIRD PTCG Club— ep bdbb060e-80e5-11f1-b89a-0242ac130204 · Q 97 · 14 energy · 22 distinct · 6 attackers
7×
Basic {G} Energy
#1
4×
Team Rocket's Energy
#15
4×
Team Rocket's Tarountula
#400
4×
Team Rocket's Spidops
#401
4×
Team Rocket's Transceiver
#1134
4×
Poké Pad
#1152
4×
Team Rocket's Ariana
#1216
4×
Team Rocket's Proton
#1220
3×
Basic {P} Energy
#5
3×
Bug Catching Set
#1094
3×
Team Rocket's Giovanni
#1218
2×
Team Rocket's Articuno
#414
2×
Team Rocket's Mewtwo ex
#431
2×
Team Rocket's Murkrow
#463
2×
Energy Search
#1119
2×
Team Rocket's Watchtower
#1256
1×
Team Rocket's Wobbuffet
#432
1×
Night Stretcher
#1097
1×
Maximum Belt
#1158
1×
Team Rocket's Archer
#1217
1×
Team Rocket's Petrel
#1219
1×
Lillie's Determination
#1227
THIRD PTCG Club— ep 0517fd72-8221-11f1-bce9-0242ac130202 · Q 97 · 14 energy · 23 distinct · 6 attackers
7×
Basic {G} Energy
#1
4×
Team Rocket's Energy
#15
4×
Team Rocket's Tarountula
#400
4×
Team Rocket's Spidops
#401
4×
Team Rocket's Transceiver
#1134
4×
Poké Pad
#1152
4×
Team Rocket's Ariana
#1216
3×
Basic {P} Energy
#5
3×
Bug Catching Set
#1094
3×
Team Rocket's Giovanni
#1218
2×
Team Rocket's Articuno
#414
2×
Team Rocket's Mewtwo ex
#431
2×
Team Rocket's Murkrow
#463
2×
Energy Search
#1119
2×
Team Rocket's Proton
#1220
2×
Lillie's Determination
#1227
2×
Team Rocket's Factory
#1257
1×
Team Rocket's Wobbuffet
#432
1×
Night Stretcher
#1097
1×
Hero’s Cape
#1159
1×
Brave Bangle
#1175
1×
Team Rocket's Archer
#1217
1×
Team Rocket's Petrel
#1219
bono— ep bfb7f99c-805b-11f1-9008-0242ac130204 · Q 95 · 10 energy · 18 distinct · 6 attackers
10×
Basic {D} Energy
#7
4×
Munkidori
#112
4×
Marnie's Impidimp
#646
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Team Rocket's Petrel
#1219
4×
Lillie's Determination
#1227
4×
Spikemuth Gym
#1259
3×
Marnie's Morgrem
#647
3×
Marnie's Grimmsnarl ex
#648
3×
Rare Candy
#1079
3×
Night Stretcher
#1097
2×
Froslass
#104
2×
Snorunt
#860
2×
Handheld Fan
#1161
2×
Boss’s Orders
#1182
1×
Unfair Stamp
#1080
1×
Larry’s Skill
#1206
Luca— ep 9f4629d6-816e-11f1-b042-0242ac130203 · Q 95 · 10 energy · 19 distinct · 6 attackers
10×
Basic {D} Energy
#7
4×
Munkidori
#112
4×
Marnie's Impidimp
#646
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Team Rocket's Petrel
#1219
4×
Lillie's Determination
#1227
4×
Spikemuth Gym
#1259
3×
Marnie's Morgrem
#647
3×
Marnie's Grimmsnarl ex
#648
3×
Rare Candy
#1079
3×
Night Stretcher
#1097
2×
Froslass
#104
2×
Snorunt
#860
2×
Boss’s Orders
#1182
1×
Unfair Stamp
#1080
1×
Pokégear 3.0
#1122
1×
Tool Scrapper
#1137
1×
Dawn
#1231
Oshbocker— ep 12be1536-80a2-11f1-a51c-0242ac130202 · Q 94 · 9 energy · 21 distinct · 5 attackers
9×
Basic {W} Energy
#3
4×
Buddy-Buddy Poffin
#1086
4×
Lillie's Determination
#1227
4×
Surfing Beach
#1262
3×
Snorunt
#860
3×
Mega Froslass ex
#861
3×
Staryu
#1030
3×
Mega Starmie ex
#1031
3×
Pokégear 3.0
#1122
3×
Mega Signal
#1145
3×
Hilda
#1225
3×
Wally's Compassion
#1229
2×
Cinderace
#666
2×
Hand Trimmer
#1087
2×
Boss’s Orders
#1182
2×
Salvatore
#1189
2×
Xerosic’s Machinations
#1197
2×
Cheren
#1224
1×
Night Stretcher
#1097
1×
Poké Pad
#1152
1×
Hero’s Cape
#1159
kashiwashira— ep d4fbd82a-80a4-11f1-b111-0242ac130204 · Q 93 · 12 energy · 20 distinct · 5 attackers
8×
Basic {G} Energy
#1
4×
Team Rocket's Tarountula
#400
4×
Team Rocket's Spidops
#401
4×
Poké Pad
#1152
4×
Team Rocket's Transceiver
#1134
4×
Team Rocket's Ariana
#1216
4×
Team Rocket's Proton
#1220
4×
Team Rocket's Energy
#15
3×
Team Rocket's Mimikyu
#434
3×
Team Rocket's Giovanni
#1218
3×
Bug Catching Set
#1094
3×
Team Rocket's Factory
#1257
3×
Lillie's Determination
#1227
2×
Team Rocket's Mewtwo ex
#431
2×
Team Rocket's Articuno
#414
1×
Team Rocket's Archer
#1217
1×
Hero’s Cape
#1159
1×
Brave Bangle
#1175
1×
Ultra Ball
#1121
1×
Eri
#1186
Oshbocker— ep efeee4d2-80a0-11f1-9ce1-0242ac130204 · Q 93 · 13 energy · 19 distinct · 5 attackers
9×
Basic {G} Energy
#1
4×
Team Rocket's Energy
#15
4×
Team Rocket's Tarountula
#400
4×
Team Rocket's Spidops
#401
4×
Team Rocket's Transceiver
#1134
4×
Poké Pad
#1152
4×
Team Rocket's Ariana
#1216
4×
Team Rocket's Proton
#1220
3×
Team Rocket's Mimikyu
#434
3×
Bug Catching Set
#1094
3×
Team Rocket's Giovanni
#1218
3×
Lillie's Determination
#1227
3×
Team Rocket's Factory
#1257
2×
Team Rocket's Articuno
#414
2×
Team Rocket's Mewtwo ex
#431
1×
Ultra Ball
#1121
1×
Hero’s Cape
#1159
1×
Brave Bangle
#1175
1×
Team Rocket's Archer
#1217
Jack— ep f6c976b0-802c-11f1-8a44-0242ac130202 · Q 92 · 5 energy · 23 distinct · 7 attackers
5×
Basic {G} Energy
#1
4×
Grookey
#89
4×
Thwackey
#90
4×
Applin
#92
4×
Dipplin
#93
4×
Lillie's Determination
#1227
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Bug Catching Set
#1094
4×
Festival Grounds
#1245
2×
Kieran
#1191
2×
Boss’s Orders
#1182
2×
Dawn
#1231
2×
Night Stretcher
#1097
2×
Brave Bangle
#1175
2×
Air Balloon
#1174
1×
Rellor
#73
1×
Rabsca
#74
1×
Shaymin
#343
1×
Black Belt’s Training
#1211
1×
Lana’s Aid
#1184
1×
Switch
#1123
1×
Secret Box
#1092
junlee789— ep 4a25cf44-808b-11f1-ab0f-0242ac130203 · Q 91 · 9 energy · 20 distinct · 6 attackers
5×
Basic {F} Energy
#6
4×
Rock Fighting Energy
#20
4×
Cynthia's Roselia
#341
4×
Cynthia's Gible
#379
4×
Cynthia's Gabite
#380
4×
Buddy-Buddy Poffin
#1086
4×
Fighting Gong
#1142
4×
Poké Pad
#1152
4×
Lillie's Determination
#1227
3×
Cynthia's Roserade
#342
3×
Cynthia's Garchomp ex
#381
3×
Cynthia's Power Weight
#1173
3×
Hilda
#1225
2×
Cynthia's Spiritomb
#387
2×
Night Stretcher
#1097
2×
Boss’s Orders
#1182
2×
Forest of Vitality
#1261
1×
Unfair Stamp
#1080
1×
Xerosic’s Machinations
#1197
1×
Surfer
#1203
Rmy— ep d4fbd82a-80a4-11f1-b111-0242ac130204 · Q 88 · 7 energy · 20 distinct · 6 attackers
4×
Telepath Psychic Energy
#19
4×
Dunsparce
#305
4×
Abra
#741
4×
Kadabra
#742
4×
Alakazam
#743
4×
Rare Candy
#1079
4×
Enhanced Hammer
#1081
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Dawn
#1231
3×
Basic {P} Energy
#5
3×
Dudunsparce
#66
3×
Boss’s Orders
#1182
3×
Xerosic’s Machinations
#1197
3×
Hilda
#1225
1×
Fezandipiti ex
#140
1×
Night Stretcher
#1097
1×
Sacred Ash
#1129
1×
Lana’s Aid
#1184
1×
Neutralization Zone
#1247
Luca— ep 34e8cd36-802a-11f1-ad47-0242ac130204 · Q 88 · 8 energy · 19 distinct · 5 attackers
4×
Telepath Psychic Energy
#19
4×
Dunsparce
#65
4×
Abra
#741
4×
Kadabra
#742
4×
Alakazam
#743
4×
Rare Candy
#1079
4×
Enhanced Hammer
#1081
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Hilda
#1225
4×
Dawn
#1231
3×
Basic {P} Energy
#5
3×
Dudunsparce
#66
3×
Night Stretcher
#1097
2×
Boss’s Orders
#1182
2×
Xerosic’s Machinations
#1197
1×
Enriching Energy
#13
1×
Sacred Ash
#1129
1×
Lana’s Aid
#1184
Yushin Ito— ep 70ef026a-8079-11f1-aa92-0242ac130204 · Q 85 · 7 energy · 22 distinct · 7 attackers
4×
Telepath Psychic Energy
#19
4×
Abra
#741
4×
Kadabra
#742
4×
Alakazam
#743
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Hilda
#1225
4×
Dawn
#1231
3×
Dunsparce
#305
3×
Rare Candy
#1079
3×
Enhanced Hammer
#1081
3×
Boss’s Orders
#1182
3×
Xerosic’s Machinations
#1197
3×
Nighttime Mine
#1266
2×
Basic {P} Energy
#5
2×
Dudunsparce
#66
1×
Enriching Energy
#13
1×
Fezandipiti ex
#140
1×
Shaymin
#343
1×
Night Stretcher
#1097
1×
Sacred Ash
#1129
1×
Lana’s Aid
#1184
Majkel1337— ep 2120df32-8025-11f1-8d74-0242ac130203 · Q 85 · 7 energy · 22 distinct · 7 attackers
4×
Telepath Psychic Energy
#19
4×
Abra
#741
4×
Kadabra
#742
4×
Alakazam
#743
4×
Enhanced Hammer
#1081
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Hilda
#1225
4×
Dawn
#1231
3×
Dunsparce
#305
3×
Rare Candy
#1079
3×
Boss’s Orders
#1182
3×
Xerosic’s Machinations
#1197
2×
Basic {P} Energy
#5
2×
Dudunsparce
#66
2×
Nighttime Mine
#1266
1×
Enriching Energy
#13
1×
Fezandipiti ex
#140
1×
Shaymin
#343
1×
Night Stretcher
#1097
1×
Sacred Ash
#1129
1×
Lana’s Aid
#1184
THIRD PTCG Club— ep cdfe3320-808f-11f1-8725-0242ac130202 · Q 85 · 7 energy · 22 distinct · 6 attackers
4×
Telepath Psychic Energy
#19
4×
Abra
#741
4×
Kadabra
#742
4×
Alakazam
#743
4×
Enhanced Hammer
#1081
4×
Buddy-Buddy Poffin
#1086
4×
Poké Pad
#1152
4×
Hilda
#1225
4×
Dawn
#1231
3×
Dunsparce
#305
3×
Rare Candy
#1079
3×
Boss’s Orders
#1182
3×
Xerosic’s Machinations
#1197
2×
Basic {P} Energy
#5
2×
Dudunsparce
#66
2×
Jamming Tower
#1246
1×
Enriching Energy
#13
1×
Fezandipiti ex
#140
1×
Night Stretcher
#1097
1×
Sacred Ash
#1129
1×
Wondrous Patch
#1146
1×
Lana’s Aid
#1184
Rmy— ep 3d817f96-8094-11f1-b9b4-0242ac130203 · Q 82 · 13 energy · 18 distinct · 3 attackers
4×
Mist Energy
#11
4×
Spiky Energy
#14
4×
Grow Grass Energy
#18
4×
Mega Kangaskhan ex
#756
4×
Buddy-Buddy Poffin
#1086
4×
Crushing Hammer
#1120
4×
Pokégear 3.0
#1122
4×
Switch
#1123
4×
Jumbo Ice Cream
#1147
4×
Xerosic’s Machinations
#1197
4×
Hilda
#1225
4×
Lillie's Determination
#1227
3×
Dwebble
#344
3×
Crustle
#345
2×
Boss’s Orders
#1182
2×
Battle Cage
#1264
1×
Basic {G} Energy
#1
1×
Hero’s Cape
#1159
Oshbocker— ep f3a2589e-8220-11f1-999a-0242ac130203 · Q 82 · 13 energy · 19 distinct · 4 attackers
4×
Mist Energy
#11
4×
Spiky Energy
#14
4×
Grow Grass Energy
#18
4×
Dwebble
#344
4×
Crustle
#345
4×
Mega Kangaskhan ex
#756
4×
Buddy-Buddy Poffin
#1086
4×
Pokégear 3.0
#1122
4×
Switch
#1123
4×
Jumbo Ice Cream
#1147
4×
Xerosic’s Machinations
#1197
4×
Hilda
#1225
4×
Lillie's Determination
#1227
2×
Boss’s Orders
#1182
2×
Battle Cage
#1264
1×
Basic {G} Energy
#1
1×
Shaymin
#343
1×
Hand Trimmer
#1087
1×
Hero’s Cape
#1159
Knobs — fleet config, applied to new workers on save
09:25 replay window 500k -> 10k files; window_files added as a dynamic knob (trainer re-reads knobs.json each refresh; <=0 = unbounded)
09:54 pacing v3 after restart stall (lifetime samples_seen had drifted to 2.1M vs ~28M real): generated-samples counted from replay dir per-file, ckpt high-water floor
06:27 forced-aggression ON at step 325k (ckpt_00325000): force_aggression_p=0.5, force_aggression_c=0.5, turn_count_penalty=0.002; trainer kept running through the flip; fleet converges via container turnover
21:55 forced-energy E range knobs deployed (~step 393k): force_energy_min/max (E ~ uniform inclusive) via knobs.json + driver pass-through + dashboard panel; set 0/30 matching the prior hardcode; worker image rebuilt, driver restarted
23:08 400k pause executed by armed stop_at_400k.sh: ckpt_00400000 published; tg-train, tg-driver, and all tgjob workers killed; inference servers + probe keeper left running
23:37 resume after 400k pause: driver fleet relaunched (250 workers, same params/knobs: energy p=1 E~U(0,30), aggression 0.5/0.5, turn penalty 0.002); trainer container had survived the stop (docker outlives tmux) ~40 steps past 400k, stalled on the pacing gate, unblocked on fresh samples; train.log pipe reattached (tg-trainlog)
07:26 full-distribution forced targets ON: trainer relaunched with --draft-full-softmax --draft-weight 0.2 (draft CE over full 1267-card vocab, zero-target ids now pushed down; draft_loss rescaled, non-comparable with history); resumed mid-generation (400k gen cosine)
07:29 worker fleet cutover to full-candidate recording: masked (forced-aggression) decisions now store the full pre-mask candidate list with visit mass on attacks and zero on non-attacks - direct attack-vs-pass policy targets; trainer needs no change (candidate softmax unchanged)
21:20 pipelines recadenced: probe keeper -> 25k-step grid from 500k (lowest unprobed first); auto-submitter -> 50k step multiple (~2.2 submissions/day of the 5/day cap)
21:54 500k anneal executed by armed anneal_at_500k.sh: knobs -> energy p=0.5 E~U(5,15), aggression p=0.25 c=1.0, window_files 50000; trainer relaunched --gen-steps 50000 (LR sawtooth now 50k-period), generation start at 500000; fleet converged via turnover (flags verified)
08:08 fleet 250 -> 350 workers: post-regime-change per-worker eval demand fell ~30% (16-22 turn prize games vs 100+ turn deck-outs; fewer game-phase evals/game, more container churn), leaving servers ~20% idle at 250; demand re-fills toward the ~106k evals/s serving ceiling (saturation est. ~310 workers)
05:10 650k final push: 8 GPUs (0-3 released; infer4-7 launched on 0-3, ports 5559-5562; infer3 retired from cluster-tools services.toml, GPU 5 dedicated to trainer); draft sims 512 -> 1024; force_energy_p -> 0 (E-range knobs inert); trainer relaunched WITHOUT --draft-full-softmax (draft CE back to legal-set softmax, applies to whole window instantly); selfplay reverted to masked-list recording at forced-aggression decisions (aggression back to indirect; old full-list samples age out over ~1-2 days); aggression knobs unchanged (p=0.25, c=1.0)
06:55 fleet 350 -> 420 workers (lease constant retired; box fully ours): aggregate serving 157k -> 177k evals/s vs ~185k ceiling, batches p50 268, no queueing (read p50 58ms), load 74/384 cores - balanced operating point for 7 servers
05:05 trainer cache fix: parse LRU tracks window size (was 4096 of 49k files = 92% miss), prefetch threads 6, pacing-gate logging added (gate was silent); step CAPACITY 0.54 -> 0.30 sec/step (~11k/h burst while catching up to the pacing cap); steady-state unchanged ~5.6k/h - the silent pacing gate was the binder all along (now logged, engaging continuously); trainer RAM ~208GB (window-resident cache)
06:55 meta_decks.csv first build (build_meta_decks.py): 19 decks from top-11 leaderboard teams, newest 3 days of episodes; ~approximate time, predates the 07:03:33 knob flip
07:03 force_lbdraft_p=0.25 deployed: per-seat forced drafts of curated top-Elo leaderboard decks (19 decks from top-11 teams, newest 3 days of episodes, legality+playability validated via deck_analyzer); DraftConfig.forced_deck legality mask; anchors the selfplay ecology to the live Kaggle meta - degenerate decks must now beat real decks to persist in z [timestamp corrected from knobs_log; initially misrecorded as 09:30]
07:30 meta_decks.csv rebuilt with display stats for the dashboard explorer (same 19 decks; file mtime exact)
07:35 window_files 50000 -> 10000 (knobs_log exact): concentrate training on post-anchoring data; trainer window confirmed at 9,802 files / 2.16M samples by step 795,460
07:55 8th inference server: tg-infer8 colocated on GPU 5 (trainer pacing-bound ~46% duty since the cache fix); fleet 420 -> 470 workers; aggregate serving 177k -> 205k evals/s, games/h ~5.2k -> ~6.6k; steps/h follows via the pacing gate
20:31 checkpoint league launched: GPU 3 dedicated (tg-infer7 retired from fleet; driver on 7 endpoints); pinned inference server per 50k checkpoint (16 servers, --max-batch 64), raw-policy round-robin (sims 0/0, drafts at 128 sims, random seat order), games to league_games.jsonl (never training replay), Elo refit per game to league_elo.json; live dashboard section
04:45 LINEAGE BRANCH: main lineage dead at step 913k (decayed); branched from ckpt_00650000 as ckpt-2026-07-19-branch-* (own metrics file); fresh fenced replay dir replay-branch (all prior replay excluded - dead-branch policy targets); knobs unchanged (energy p=0, lbdraft 0.25, aggression 0.25/1.0, penalty 0.002, window 10k); trainer/servers/exporter on --ckpt-prefix; probe keeper branch-aware (main capped at 650k); league re-keyed by checkpoint name; submitter prefers branch lineage; dashboard lineage tabs (/ = branch, /main = dead)
00:20 LINEAGE 2026-07-20b: re-cut from the 650k weights after the sprint-corrupted 07-19 branch (dead at ~755k); PR #4 merged (snapshot generation mode + eval_search harness) and deployed; trainer launched with --gen-mode snapshot (frozen-window generations, epoch-cap sizing, forfeit pacing); fresh fenced replay dir replay-20b; knobs unchanged; probe/league/submitter/dashboard repointed (dead lineages capped at 650k)
08:40 LINEAGE 2026-07-21c: Gameplay-regime phase cut from ckpt-2026-07-20b-00813000. Trainer: --regime gameplay --offregime-value-frac 0.15, snapshot mode, gen-steps 50000, window 10000, fresh optimizer (regime switch joint->gameplay). Rollouts: drafting off, --fixed-deck-pool meta_decks.csv (19 decks), sims 800, 470 workers. Fresh fence /shared/pokemon-train/replay-21c (results/ subdir was missing initially; workers crash-looped ~8 min with silent exit=0 until created). Exporter repointed to 21c prefix. Trainer image rebuilt with regime support (training-gym 6e3be5c).
09:15 21c consumer catch-up: probe_keeper + league synced to repointed scripts and restarted; winstats repointed to replay-21c; dashboard fixed for regime metrics rows (absent draft_loss) and run_events schema
10:26 Probe keeper fixed: discovery loop only globbed ckpt_*.ts (BRANCH_PFX was a dead variable in the repo copy; box-local uncommitted edits were clobbered by sync) - now lists live-lineage .ts uncapped and caps dead lineages at 650k. League restarted on repointed code; stale 20b league servers >650k killed. First 21c probe (825000) running; first 21c league entrant at 850k (league GRID=50k).
12:02 Probe 21c-00825000 attempt 1 killed at 95 min on a false wedge diagnosis (docker stats undersampled CPU; docker top later showed the orchestrator busy). Probes normally take ~2.5h and current 80-turn games make them longer. Attempt 2 running; keeper retry path verified working (FAILED logged, attempt 2 auto-started). No further kills below 3h runtime.
17:01 League converted to fixed-deck matches (random pair from the 19-deck meta pool per game, raw policy both seats) with fresh state files league_fixed_games/elo; dashboard panel repointed and retitled. Old drafted-league state preserved. Same GPU3 servers. First reading: 21c-850k leads at ~1319 Elo after 36 games (tiny sample).
20:25 League matchmaking: uniform random pairing replaced by Elo-window matchmaking (uncertainty-weighted player A by 1/(1+games), B within +/-250 Elo, 20% uniform mix for rating-graph connectivity, random seat swap). Deployed and verified live.
05:28 LINEAGE 2026-07-22: re-branch from ckpt-2026-07-20b-00813000 (fixed-deck league Elo ~1132 @153g; 21c dead at ~910k, its 90k gameplay-regime steps a statistical tie vs seed). New search regime: sims 1600 + temp_turns 4 (tau=1 first 2 turns/player then greedy; new --temp-turns selfplay flag, knob-ified with sims) + force_aggression_p=0 (mask corrupted recorded pi targets). Data: autogo exact-inclusion manifest fencing (--replay-manifest, /shared/pokemon-train/replay_manifest.txt -> replay-22). Regime gameplay unchanged (draft head frozen, 19-deck pool). Panel-reviewed (4x Opus: throughput/temp/knobs/fencing all GO-WITH-CHANGES, changes applied). Naming rule updated: date-only prefix, letters only for additional same-day branches. Consumers repointed; worker+trainer images rebuilt; marker BRANCH-2026-07-22.md. Design note: workbook design-notes/2026-07-22 Search Strengthening.
05:50 Retroactive lineage rename to date-only convention: ckpt-2026-07-19-branch- -> ckpt-2026-07-19-, ckpt-2026-07-20b- -> ckpt-2026-07-20-, ckpt-2026-07-21c- -> ckpt-2026-07-21-. 930 files renamed (checkpoints + published .ts + metrics/generations ledgers); league game ledgers, elo state, ports, probe rows and attempts rewritten in place (identity/history preserved - e.g. 07-20-00813000 kept its 593-game record); stale league server pins removed and servers re-ensured under new names. Historical records (run_events labels, old BRANCH markers, workbook notes) intentionally left as written. Dashboard: per-lineage static pages (dashboard-.html) with relative tab links added.
07:06 Game-length instrumentation: winstats appends p10/p50/p90 turns + draw/capped share per 10-min sample to turns_history/.jsonl (--fence-name arg); one-off backfill from result-CSV mtimes for dead fences (789k games); dashboard gains per-tab Game length charts (percentile band + draw/capped share) with run-event markers.
07:35 Action-mix instrumentation: pi-expected played-action type shares decoded from replay candidates (option-type one-hot params[0:17] + stop mass params[91]); per-fence 30-min series in action_mix/.jsonl; incremental added to winstats loop; full-history backfill launched in tmux (tg-mixbackfill); dashboard Action mix chart per tab. First live sample: End 63% / Play 17% / Attach 11% - attack share negligible.
08:00 LINEAGE 2026-07-22b: cut from ckpt-2026-07-20-00813000 (fixed-deck league leader, Elo 1123.6 @657g; top-4 within noise, leader + most-validated + identical to 07-22 seed for clean A/B). Single delta vs dead 07-22 (~830k, ~6h): deckout_draw=1 - deck-out endings score z=0 for both seats (new --deckout-draw selfplay flag + knob; result CSV reports draw). Motivation: action-mix instrumentation showed End=63% pi-mass with zero forced passes (49% End even with attack available) - deck-out wins were the paying strategy. Fence replay-22b via manifest; sims 1600/temp_turns 4/aggression 0 carried over; consumers repointed; images rebuilt. Kaggle score 120.8 on the late-07-21 submission corroborates the decayed-tail diagnosis and evidence-selected seeding.
18:39 LINEAGE 2026-07-22c: IMITATION PHASE 1. Cut from ckpt_00650000 (ruled; frozen draft head predates mono-spam). Behavior cloning on the elite leaderboard-episode corpus: 32,907 episodes (per-day p70 + last-3-weeks + hygiene), decoded to neutral logs and imported to replay-22c as 2.77M one-hot .tgr samples (0 skips). Trainer on manifest, regime gameplay, gen 1 = 21,601 steps (8 epochs, static corpus); the min-new-frac block at gen end IS the phase-end signal - do not remediate the waiting trainer. 22b dead at ~905k (credit fix alone: attack 1.8%, turns p50 94, league tie vs seed). Rollout fleet idle by design during IL; winstats/action-mix loops stopped (no rollouts); inference servers up for probe/league. IL z = real outcomes incl. deck-out wins (ruled acceptable; RL fine-tune will use deckout_draw=1). Promotion metric: fixed-deck league Elo of 22c checkpoints vs 650k anchors.
20:09 22c IL value-head overfit RULED ACCEPTED (Dan): train value MSE 0.12 (memorization below the ~0.85 stochasticity floor), eval value rising ~1.16 - EXPECTED, do not remediate. Policy head clean (eval CE 0.887 falling). Rationale: value head is scaffolding; RL fine-tune retrains it on streaming self-play (structurally overfit-immune); phase promotion metric (fixed-deck league, raw policy) and Kaggle bundle do not use value. Pre-registered escalation: if early RL search is value-blind, build the one-position-per-game value corpus over all 187k episodes (AlphaGo-faithful adaptation).
20:42 League replay viewer: selfplay --vis-out writes per-game visualization JSON (engine ApiVisualizeData stream); league passes it, records vis filename per game, prunes to 200; dashboard Recent replays panel + visualizer.html (?game= loader; NOTE playback POSTs the JSON to the competition's hosted viewer - external service, internet required). Host selfplay rebuilt (flag-gated, default identical); worker image and IL trainer untouched.
21:19 22c IL epoch cap raised to 16 (ruled): trainer restarted with --min-new-frac 0 so gen 2 (671601-693202, epochs 9-16) runs over the static corpus; on-box stop armed in tmux tg-stop16 firing on ckpt 693000 (dry-run verified; min-new-frac 0 would otherwise cycle generations forever). League fine grid (12.5k, step%12500 in {0,500}) live for lineages >= 22c. Broad-minus-elite conversion (esp. recent p50-p70 slice, min_score ~1085-1100) flagged as available corpus doubling, not yet ruled.
22:13 IL A/B fork at 675k: 07-22d (branch B) launched on GPU7 - 8 epochs over the min_score>=850 corpus (172k episodes, ~14.2M samples, replay-22d, own manifest); 07-22c (branch A) continues to 16 epochs on elite, stop armed at 693000. League: pool broadened to 419 decks (19 meta + 400 from 850+ bands), per-game deck indices recorded for condition-split Elo. Hypothesis under test (Dan): lower-Elo corpus = worse peak, better generalization.
22:39 Probe keeper DISABLED (ruled): 5h/probe cannot track the 12.5k fine grid across parallel lineages, and its searchful opponent is ill-defined with two trainers writing. Redesign proposal pending Dan's ruling. GPUs 1 and 6 freed (probe endpoints no longer needed).
22:47 League allocation redesigned Kaggle-style (ruled): player-A weight = entry burst (<30 games, top-player weight) then exp((elo-max)/100) ladder with 0.01 floor; 20% uniform mix retained; servers now on-demand for the sampled pair with LRU eviction beyond 30 (last_match.json). Replaces the proposed rule-based retirement: weak/stale checkpoints quiesce emergently; anchors stay sampled by strength. Probe retired; its roles redistributed (deck snapshots + search-gap matches pre-registered, league = universal judge, Kaggle = external anchor).
23:11 IL grand sweep launched (ruled): 5 new arms on freed GPUs - 22e (warm 675k seed, 850+ corpus, 16ep, GPU0), fresh-init 2x2: 22f (elite/8ep GPU1), 22g (elite/16ep GPU2), 22h (850/8ep GPU4), 22i (850/16ep GPU6). With 22c (elite/16 warm) and 22d (850/8 warm) = full warm 2x2 + fresh 2x2. All trainers on the cache-capped image (--cache-files LRU, default 512; 22d measured 263GB -> bounded after restart). Shared corpora via per-arm manifests; null-root bookkeeping for f-i; per-arm stop scripts armed from generation banners (generic stop_lineage_at.sh, dry-run verified). Idle inference servers killed (probe retired). Pre-registered reading rules: fresh-arm losses at cap require the held-out CE still-falling check (under-training vs inferiority); winner selection needs gate CI + confirmation games (max-of-8 bias).
00:24 850-corpus arms (22d/e/h/i) restarted with --cache-files 3000: the 512-file default caused cache thrash on the 2,854-file corpus (uniform batch sampling touches hundreds of files/batch; ~350 steps/h vs ~10k). Full-corpus caching = ~263GB/arm x4 + 2 elite arms ~60GB = ~1.2TB projected vs 2.2TB host. Also recorded: 22c COMPLETE at 693040 (15.93 epochs; eval CE 1.201->0.853; league 663k/675k/688k = 1306/1310/1315 vs ~1100-1125 anchors: imitation +~200 Elo). Twin anomaly: 22c-00650000 seed twin rated 946@318g vs siblings ~1105 - SE recalibration + pool-change reading flagged urgent.
01:02 League exclusion consolidated on disabled.json (league_archived knob removed); league.py redeployed and restarted in fresh tg-league tmux. Incident during deploy: a pkill over ssh self-matched the relaunch command text and killed the tg-league session, then cluster_train.sh sync's --delete rsync removed runs/tg/league.py (sync never copied it — line added, now it does); league was down ~15 min, recovered by direct rsync + relaunch, verified alive and match-playing on HEAD (d83a179). epoch_budgets knob populated {22c:16,22d:8,22e:16,22f:8,22g:16,22h:8,22i:16}; dashboard restarted serving the Experiments overview; complete-status tolerance 0.25 epoch added (grid stops land short of exact epoch arithmetic).
03:50 GEN-STATE BLOWUP remediated (22d/22e): resumes had restored generation state that was wrong for the corpus (22e inherited its SEED's elite-corpus gen_state - seed copies carry gen_state, now a known bug; 22d via restart chain), causing a premature gen boundary then a fresh 50k-horizon cosine at peak 1e-3 on converged clones -> divergence (d: CE 0.83->1.68; e: 0.75->1.62; league convicted d-700000 at Elo 940). ACTIONS: post-divergence ckpts (19 files) moved to checkpoints/attic-genstate-blowup/ (archived, not deleted); d resumed from 693000, e from 694000; ALL four 850-arms now run --gen-steps 12500 (short cosine cycles, elite-arm-like shape; protocol uniform); h/i restarted from 26000 before their 50k trap (a first kill attempt silently failed - old containers survived one batch; verified dead before relaunch); league entrant 22d-00700000 paused via disabled.json (kept in ledger). SIDE EFFECTS for readers: 22d/22e metrics files contain the divergence segment ~693k-707k / ~694k-700k (steps re-run after rewind; dashboards dedupe by step keeping newest); epoch counts for d/e computed from lineage start slightly overstate consumed epochs by the discarded segment; stops re-armed d@783000 e@892000 h@108000 i@217000. Underlying fix queued: strip gen_state on corpus change at resume.
04:44 850 corpus REBUILT VERIFIED (v4): 20,048,862 samples / 18,851,960 records, 0 shard errors, 45/172,117 episodes skipped loudly (June-era engine-incompatible - the true root cause of all prior losses: per-file exceptions aborted xargs batches; import_logs now isolates per-file failures and never overwrites). Four 850 arms relaunched on the verified corpus (d/e from 675000 seeds on data_sig-fixed image - inherited elite gen_state auto-discarded; h/i fresh). CHECKPOINT REGISTRY introduced (scripts/checkpoint_registry.json): per-lineage provenance chains + invalid ranges with documented reasons and incident refs; dashboard renders INVALID badges with reason tooltips on league rows + provenance/status on lineage tabs; preflight requires registry entries. Sprint-corrupted 07-19-700000/750000 additionally frozen (disabled.json now 9 stems). All invalidations now carry documented reasons: 25-epoch sprint (07-19 tail), gen-state blowup (22d/e segments), partial-corpus training (22d/e/h/i first attempts).
06:35 League EVAL REVERTED to the 19-deck meta pool (Dan's ruling - the 2026-07-22 pool broadening was my over-broad interpretation of an instruction aimed at training-side deck exposure, disclosed but not explicitly ratified). Elo now fitted on meta-condition games only (both deck indices <19; broadened-era games remain in the ledger, excluded from the fit) restoring scale continuity with the original condition. Consequence flagged: the generalization hypothesis (broad vs elite corpus off-meta performance) loses its league instrument; needs a separate eval vehicle if pursued (off-meta gauntlet or secondary condition), not built pending ruling. deck_pool_850.csv / league_deck_pool.csv retained on the box for that future use.
07:28 AUDIT REMEDIATION (mechanical, semantics-restoring): (1) league game counts everywhere (elo json, gate SE, dashboard, matchmaking burst) now count FIT-condition games only - displayed counts had overstated rating information up to ~13x for broadened-era arms (CIs ~3.5x overconfident); games_total_per_ckpt added for transparency; era-orphaned players re-enter placement burst on the ruled 19-deck condition. (2) Elo fit iterates to convergence (tol 0.25, converged at ~359 iters vs the old fixed 60) - new low-game arms were ~64-87 Elo below the fixed point; converged IL cluster ~1390-1420 vs old lineages ~1045. AWAITING RULING: orphaned broadened-era games (second-condition fit vs accept); final-stop-checkpoint league qualification. Deferred (audit's survivable list): money-chart per-corpus panels, epoch-count 15% offregime note, toggle lock, headline staleness.
13:48 RULED: (c) broadened-era games quarantined (ledger-only, no secondary fit). Stop targets re-armed grid-aligned (round-up overtrain): d 838000 (8.53ep), e 988000 (16.37ep), g 50000 (18.5ep), h 163000 (8.53ep), i 313000 (16.37ep); 22f RESUMED 21000->25000 (9.26ep - same epoch coordinate as the warm 8ep cell at 675000, clean comparison); 22c left at 693040 (rated 688000 = 15.69ep, 0.24ep gap accepted). Registry gains stop_alignment_policy.
17:18 RULED + launched: (1) LINEAGE 2026-07-23 = the true warm-850 factorial cell (seed ckpt_00650000 directly on 850-v4; the prior 'warm-850' arms d/e were actually elite-clone->850 curriculum arms inherited from the 675k A/B fork - relabeled in registry). LR 2e-4 continuation profile, gen-steps 12500, stop 813000 (8.53ep). (2) 22d RESCUE from 700000 at lr 2e-4 - doubles as the collapse-mechanism test (if clean, peak-LR shock confirmed); armed stop 838000 still exact for the remaining 6.69ep. 22e stays paused pending the rescue verdict. Collapse segment (d 701-771k, e 701-774k, 145 files) archived in attic-cosine-collapse; registry invalid ranges now incarnation-scoped (old partial-corpus ranges marked superseded - same steps validly retrained on v4; stem collision across incarnations noted); valid v4 d-700000 unfrozen. GPUs: 23 on 7, d-rescue on 0.
23:45 CALIBRATION SUBMISSIONS (Dan's go): five forced-deck bundles - 22c-00693000 x decks {18,14} and 22h-00163000 x decks {18,3,14}; deck picks recomputed from league_fixed_games.jsonl fit-condition games (handoff's 1/13/3 @74/74/73 did NOT reproduce on the current ledger - no saved artifact; superseded by pooled per-deck win rates: 22c d18 69.6%/d14 69.3%, 22h d18 64.4%/d3 62.9%/d14 62.7%). Deck 18 shared across both models = paired piloting comparison. Mechanism: FORCED_DECK marker file in bundle -> agent declares deck.csv directly, draft head skipped (agent/main.py + package.py forced_deck flag). 4 submitted 23:36Z (cap: auto-submitter spent slot 1 at 00:11Z); 5th (22h d14) armed for 00:01Z rollover. 22h finished at 163000 ~23:2xZ; final ckpt exported one-shot (docker pokemon-train:latest) + full 22h lineage exported by the interval exporter. Local disk incident: mac data volume 100% full - pruned 161G of stale candidates mirror (.ts re-rsyncable from cluster) + old staging.
04:14 PROBE B (draft-IL representation probe) launched: tg-train-draftprobe on GPU 4, image pokemon-train:latest rebuilt (494e8a) with training-gym a33bf1f --freeze-value (draft regime trains draft_head ONLY; trunk+policy+value frozen; loss = draft CE only). Corpus: elite draft shards /shared/pokemon-train/imitation/draft-elite-v1 built by scripts/imitation/build_draft_shards.py from tier_elite.txt (32907 episodes, 3948840 type-1 samples, one .tgr per episode = episode-grouped split, 633 holdout files = 1.92% via crc32%50, 0 skipped picks). Seed: ckpt-2026-07-22c-00693000 copied to prefix ckpt-2026-07-24dp-. Trainer: snapshot mode, gen-steps 12500, batch 1024, lr 2e-4 (22d continuation profile), prefetch-threads 16. ~3782 steps/epoch; run-until-Dan-stops. First rows: draft_loss 7.60 at warmup, value_loss 0.
05:10 OVERNIGHT PROBE PROGRAM ARMED (Dan: 'take the wheel'). Running: [B] draft-IL Arm A on GPU4 (tg-train-draftprobe, 22c-00693000 seed, draft head only, --freeze-value new flag, elite type-1 corpus 3.95M picks episode-grouped holdout; seed baseline top-1 0.59% -> 64.3% @ +1000 steps - frozen trunk IS decodable for drafting). [D] deck-evolution 40-gen loop (codex gpt-5.6-sol xhigh mutations on POOL COPY, meta_decks untouched; gens 1-3 all rejected - candidates lose to parents, policy co-adapted to training decks; mutation-memory + parent-cooldown added). [C] LLM-pilot: opus-4-8 effort-low lost 0/5 vs frozen 22c on deck 14 ($0.77/game, killswitch $500); effort-high 10-game arm + one-shot LLM-draft arm launched. [A] coverage scan building. Armed conditionals: night operator (22e rescue auto-launch iff 22d completes clean at 838000; league surge +2 workers per freed trainer GPU; broad-tier conversion after first trainer stop - conversion only, no training run); SE twin recalibration (ckpt-2026-07-24t/u-00688000 = byte-identical 22c-00688000 copies, alias-exemption in league.py TWIN_STEMS, rating spread -> new SE calibration); convert cron (daily fenced increments, never auto-adopted). NOT armed: Kaggle slots (5 held for morning slate), packed-tensor build (deferred to daytime), B Arm B (moot at 64% top-1 unless free-running collapses).
04:40 CONVERT CRON deployed (fenced corpus refresh): daily 12:30 UTC crontab on the box (~2h after the 10:30 UTC Mac-launchd puller) runs /shared/pokemon-train/imitation/convert-cron/convert_cron.sh -> convert_daily.py: extract_games refresh, elite filter imported from build_tiers.py (DONE/DONE hygiene + per-day min_score p70 + 21-day window), unconverted ids (ledger convert-cron/converted_ids.txt, seeded from dlog-elite/draft-elite-v1 = 32,907) decoded to dlog-elite-inc//, imported via training/import_logs to replay-elite-inc// (per-file isolation, .tgr header counts verified vs importer totals), type-1 draft shards to draft-elite-inc//; manifest fragment per increment + summary line to imitation/increments.jsonl. NO live manifest is touched - increments sit fenced until adopted by ruling. Proof run: increment 2026-07-24, newest 200 unconverted elite episodes -> 200/200 decoded, 28,397 dlog records, 30,125 replay samples (verified OK), 24,000 draft samples, 0 skips/failures, 65s. 1,191 elite candidates remain unconverted (mostly day 2026-07-22); next scheduled run will fold them.
07:10 STOP-WATCHER INCIDENT + REMEDIATION: (1) tg-stop-d fired at 838000 but docker kill left a ZOMBIE container that kept training to ~846000 (the documented verify-your-kills trap); trainer killed via host pkill -9; 8 post-stop checkpoints (839000-846000) moved to checkpoints/attic-poststop-22d/ (ruled stop artifact ckpt-2026-07-22d-00838000.pt intact). (2) tg-stop-23 was armed on a filename that cannot exist (ckpt-2026-07-2223-00813000.pt - script prefixed '22' to lineage '23'); lineage-23 would have blown through its ruled 813000 stop. All three stop watchers re-armed on stop_at_v2.sh: explicit target path + container + pgrep pattern, docker kill VERIFIED dead with 3x pkill -9 escalation. 22e watcher re-armed for the night-operator-launched run.
21:35 COMPOSITION-SELECTION PROBE: killed tg-dp-evals watcher (draft-IL trainer stopped, nothing to watch; verified dead). Generated 400 candidate decks from ckpt-2026-07-24dp-00820000 (best argmax_gameplay_wr 0.638 in trajectory.jsonl): temps {0.8,1.0,1.3}, first-5-pick sampling, 25% modal-first-pick ban; batch A top-p 0.95, batch B top-p 1.0 (supplementary, logged assumption: batch A yielded only 25 unique). 36 unique after sorted-multiset dedup; 5 exact meta-deck matches flagged (decks 1,4,7,11,18). Screened all 36 at 76 games vs 19-deck pool on frozen 22c server :5698 (2736 games, 0 errors); top-10 by screen LCB confirmed at 304 games (3040 games, 0 errors). Best non-meta candidate 99f6f523a7fd: confirm wr 0.645, combined Wilson LCB 0.571 - between median (51.5%) and deck-18 baseline (73%), does not beat it. Deck 18 itself re-measured 0.64 in this harness vs its 73% ledger baseline (measurement-condition gap, flagged). Outputs: probes/composition-selection/{results.json,decks/,cluster/}; cluster mirror /shared/pokemon-train/probes/composition-selection/. No league, training, or meta_decks.csv changes.
23:33 probe D-v2 (deck-hillclimb-v2) launched: 3 concurrent LLM hillclimbs on top-3 registry view decks; own pool copy at /shared/pokemon-train/probes/deck-hillclimb-v2/; reuses tg-probeD-server GPU 1 port 5698; <=24 nice-15 workers, farm batches serialized; league/meta_decks.csv untouched
12:00 22e COMPLETE at ruled stop 988000 (16.37ep) — with incident: stop watcher was armed with wrong container name (tg-train-e vs tg-train-22e, grammar-class) AND its pkill escalation ran unprivileged vs a root-owned container proc (Operation not permitted) -> 155k steps unauthorized overrun to ~1143000; killed via sudo pkill; 154 post-stop ckpts -> attic-poststop-22e/; ruled final ckpt intact. ALL SWEEP TRAINING NOW COMPLETE. Separately: league Elo fit DIVERGED outright (unconverged at 60k iters, deranged values incl 675000@1620) — damped 1/sqrt(n) iteration inadequate at current scale; replacement with exact Bradley-Terry MLE dispatched; ratings quarantined-as-artifacts until refit lands (twin pair v/w gap 1.8 suggests high-n relative order survived). Winner ruling waits on refit. Two builds designated by Dan and documented: Build 1 ops-robustness/opslib (design-notes 2026-07-25 Build 1), Build 2 autoresearch phased process on winner (design-notes 2026-07-25 Build 2).
16:38 Trainer image pokemon-train:latest rebuilt (b649bd14ee32) with training-gym e099d6f --value-only (gameplay regime trains value_head ONLY; trunk+policy_head+draft_head frozen; loss = value MSE only, mirrors --freeze-value mechanism). Synced train.py only to ~/runs/tg (no full sync). No training launched, no containers touched, no GPU used, no pushes. Verification note: docker run without --gpus still fails under the default nvidia runtime (fabricmanager socket missing); verified with --runtime=runc (--help shows --value-only; VALUE_ONLY present in image code).
22:23 Kaggle submissions: ckpt-2026-07-23-00813000-forceddeck18 and ckpt-2026-07-22e-00700000-forceddeck18 (102.0 MiB bundles, parity suite PASS). Exports produced via pokemon-train:latest with --runtime=runc (CPU). Selection basis: converged BT-MLE league (07-25 17:29Z) top real (non-twin) players, one per arm family.
21:30 SESSION HANDOFF (see workbook '2026-07-25 Session Handoff — Winner Crowned, Triangulation Slate Staged, GPU Outage Recovery In Flight'). GPU OUTAGE: unattended-upgrades replaced NVIDIA driver mid-operation (580.159.03->580.173.02); recovery peeled 4 stale userspace layers: kernel module reload (Richard), nvidia-fabricmanager stale FAILED unit (systemctl restart fixed), ldcache, CDI specs /etc/cdi/nvidia.yaml + /var/run/cdi/nvidia.yaml still referencing .159.03 (regen via nvidia-ctk PENDING at handoff). ALL GPU workloads down, all resumable. STAGED NOT LAUNCHED: 5-arm triangulation slate (25a winner-continuation/missing-cell, 25b winners-weighted, 25c recency, 25d elitemix, 25e de-scarred-curriculum; corpora verified 13.39M/27.64M/27.57M; registry entries + grid stops + pre-disabled intermediates done) + GPU5 chain (25vr value-only repair, 25dr draft retune). Winner ruling data-complete: lineage-23-00813000 pooled ~1262 (twins v/w 0.8 apart); curriculum family confirmed degrading (22d 885, 22e 761 - regression-toward-demonstrator-mean theory, ICML-draft in workbook autoresearch/). Challenger displacement pre-registered at >3 SE. Elo fit = exact Bradley-Terry MLE (301 iters, Spearman .94). apt-mark hold nvidia-* recommended, unruled.
23:39 GPU outage recovery closed and driver stack pinned. Unattended-upgrades replaced the NVIDIA driver 580.159.03->580.173.02 from jammy-security/restricted under a loaded kernel module mid-training; recovery required four independent repairs (kernel module, nvidia-fabricmanager stale FAILED unit, ldconfig, and BOTH CDI snapshots /etc/cdi/nvidia.yaml + /var/run/cdi/nvidia.yaml, the runtime being mode=cdi). Verified recovered 22:32-22:43Z: both CDI specs regenerated (89/89 refs at 173.02), fabricmanager active, container smoke test 'docker run --rm --init --gpus device=0 pokemon-train:latest nvidia-smi -L' returns H200 rc=0. Adds scripts/pin_nvidia_stack.sh (holds the 17 driver-stack packages at the running driver version plus 4 container-toolkit packages; hold/--unhold/--dry-run/--status; verification gate is a simulated dist-upgrade moving nothing in the hold set, exact-match not substring). Adds scripts/fix_nvidia_module_reload.sh with ldconfig + nvidia-ctk cdi regeneration of both specs, a CDI-vs-running-driver gate, /dev/nvidia-uvm existence check before spec generation, hard fabricmanager is-active gates on both paths, and step-7 pinning delegation. Gate 0 no longer exits on a healthy host nvidia-smi (that is this incident's signature) but skips only the kernel reload, making the script re-entrant. Static review by a Fable subagent found both defects. Holds NOT yet applied - requires sudo. apt-daily-upgrade.timer next fires 2026-07-26T06:59Z.
23:57 Two forced-deck submissions filling the unmeasured cells of the {22c-00693000, 23-00813000} x {deck 14, deck 18} factorial: 54986632 = ckpt-2026-07-23-00813000 x deck14, 54986634 = ckpt-2026-07-22c-00693000 x deck18 (both 102.0 MiB, submitted 23:53Z, PENDING). Existing cells: 22c-693000 x deck14 = 892.4 (07-23), 23-813000 x deck18 = 888.1 (07-25). The two 600.0 ledger entries are UNPLAYED placements frozen out by the 2-active rule, not measurements, so both submitted cells were genuinely unmeasured. Caveat for later reading: the pre-existing diagonal was measured in an earlier window against a different opponent field and scores drift materially with accumulated episodes (submission.tar.gz 54986053 moved 754.9 -> 780.4 -> 802.9 within one hour), so window is confounded with the diagonal; the two new cells are same-window with each other. Day's 5 submission slots now used (3 by the judge auto-submitter at 22:22/22:22/23:01, 2 here). Adds daily-top-agents-submitter/submit_20260725b.py. Note: package.build_submission now names the archive .tar.gz itself, so the post-build rename in submit_calibration.py:43 and submit_20260725.py:40 is dead code that raises SameFileError if those scripts are re-run.
00:45 SLATE ARMED: 5 of 6 arms launched through opslib for the first time. Deploy gate first: scripts/cluster_train.sh sync was failing SILENTLY (exit 1, no output on either stream) because drift_check's `rline=$(ssh ... test -f ...)` fails under `set -e` when a local file is absent on the box, aborting before the '[ -n $rline ] || continue' guard written for that case; trigger was the new scripts/eval_probe/draft_top1.py. Fixed with `|| true`; opslib then deployed (14 modules) and ALL 14 import clean under the box's bare Python 3.10 (config.py's tomllib>tomli>fallback chain works). Launched via `python3 -m opslib.launch --trainer --gpu N --replay-manifest ... --execute`: 25a/gpu0/manifest_23 stop 975000, 25b/gpu2/winners stop 763000, 25c/gpu4/recency stop 875000, 25d/gpu6/elitemix stop 875000, 25e/gpu7/manifest_23 stop 838000. Verified ground truth: all 5 containers Up, all 5 GPU leases HELD by lease.sh with the trainer docker command recorded (the foreground-in-tmux lease bug class holds correctly), generated argv uses --gpus device=N (CDI path) and --init. Stop watchers armed via `opslib.stop_at ` in tmux stop-25{a..e}; all 5 confirmed as live processes by pgrep. stop_at refusal paths verified BEFORE arming: bare arm letter 'e' -> exit 2 (matches zero registered lineages; this is the fix for the tg-train-e vs tg-train-22e 155k-step overrun) and off-grid step 975500 -> exit 2. Preflight checks container_matches_derived and ckpt_prefix_matches both pass. NOT LAUNCHED: the GPU5 chain 25vr/25dr, NO-GO on two counts — (1) 25vr seed copy not staged, (2) MEASURED valrepair corpus is 70,386 samples (68,692 training window) vs the spec's 65,807, a ~7% discrepancy not explained by the 1,694-sample holdout and unresolved; registry entries for 25vr/25dr were written with agent-implemented stop semantics (best-holdout-MSE selection with --ckpt-every 500 and a 816000 ceiling for vr; 824000 %1000-aligned for dr) explicitly flagged AWAITING DAN'S CONFIRMATION. Removed scripts/launch_arm_v3.sh (the reverted-CDI one-off) as superseded by opslib.launch.
00:50 Build 1 opslib merged (main fe74f79, docs 29f5d60): scripts/opslib/ stdlib-only package -- names/truth/preflight/verified_kill/launch/owners/pipeline/invariants/hoststate/remote/stop_at/config -- plus fleet_stop.py and chown_shared_to_dan.sh (committed, NOT run; --user still off, 2931 root-owned ckpts would block a resume's metrics append). Call sites migrated: night_operator launch string and stop arming, league server naming/listing/launch and shared-file writes, rollout_driver worker launch/teardown, probe_keeper launch and prefix discovery, cluster_train stop, fix_nvidia kill path, registry_env config load. BOX INTERVENTION: ~/stop_lineage_at.sh, ~/stop_at_v2.sh, ~/stop_at_v3.sh DELETED (verified no tmux server and no watcher process first); ~/stop_at_400k.sh and ~/stop_22c_at_16ep.sh kept as historical one-offs. CORRECTION TO EVENT 2026-07-24T07:10:00Z: that event records 'docker kill left a ZOMBIE container' at 22d. That diagnosis is FALSE and the correction matters because it is what justified the v2 sudo/pkill escalation. Ground truth /shared/pokemon-train/logs/stop_22d.log: 'Error response from daemon: cannot kill container: tg-train-d: No such container: tg-train-d' immediately followed by 'FIRED: tg-train-d stopped'. The container was tg-train-22d; stop_lineage_at.sh built the name from a bare arm letter and piped docker kill through tee without checking rc, so a total failure logged as success. docker inspect of tg-train-22d: Init=None, PidMode='' (private ns), Path=python, Entrypoint=None -- the trainer was PID 1, so a correctly-named docker kill would have worked. No zombie, no PID-namespace escape, no privilege problem. 22d, 22e (155k overrun) and the impossible ckpt-2026-07-2223-00813000.pt target are ONE root cause: the naming grammar. verified_kill therefore has no sudo and no pkill escalation against container PIDs (sudo -n is unavailable on this box; dan is in group docker so docker kill never needs it). Reviewed by 4 adversarial static-analysis passes; fixes in 3a295c8, eaa991d, 3e91938. Two blockers found that the plan did not anticipate: the box has only Python 3.10.12 with no tomllib/tomli (three modules could not have imported there; centralised in opslib/config.py with a fallback parser), and launch.py's docker run -d released the GPU lease seconds after launch because lease.sh holds its flock only while its child lives (leased launches now run docker foreground inside tmux). Not exercised against the cluster by this session.
01:03 25vr LAUNCHED (GPU5, tg-train-25vr, --value-only --ckpt-every 500, ceiling stop armed 816000) after measuring the decision. BASELINE MEASURED (tg-eval-valbase, GPU3, inference-only): winner ckpt-2026-07-23-00813000 value MSE on the clean 1,694-sample valrepair holdout = 0.5087 +/- 0.0214; constant-predictor floor on that same holdout = 1.0000 EXACTLY (847 wins/847 losses/0 draws, every label +/-1). CORRECTION: the '~0.85 coin-flip floor' used in the acceptance gate has no derivation anywhere and is wrong; the measured floor is 1.0000 (0.85 would require ~15% draws; this corpus has 0.05%). The head is INFORMATIVE (halves the floor, unlike 22c's 1.16 which was worse than useless) but BIASED: mean prediction +0.185 vs mean outcome 0.000, ~11 sigma optimism concentrated in the midrange (pred 0.2-0.4 realizes -0.20; every bucket from -0.6 to +0.8 realizes 0.25-0.50 below prediction; extremes well calibrated). Bias affects the TRAINING loop, not the shipped agent: daily-top-agents-submitter/agent/main.py is a CPU-only RAW-POLICY agent (no MCTS, no sims) so the value head does not participate in deployed action selection at all; the repair is a prerequisite for RL/self-play (value drives advantage estimation and search targets at training time), not an improvement to the current submission. RULED ACCEPTANCE GATE (Dan, this session): floor corrected to 1.0000; accept only if holdout MSE clearly beats the 0.5087 BASELINE (not merely the floor), AND |mean pred - mean outcome| materially below 0.185. Corpus 'discrepancy' RESOLVED as unit confusion, not data loss: 65,807 = decision records (one per episode/seat), 70,386 = training samples after import_logs expands per-pick plus terminal STOP (+4,579); the Build 2 design note has the two labels swapped. opslib.launch gained an --extra passthrough (trainer_spec already accepted extra_train_args but no CLI door existed; without it 25vr could only launch via hand-composed docker run). First attempt used nargs=REMAINDER which swallowed --execute into the extras - caught by launch.py's dry-run-by-default, no bad launch occurred - changed to action='append'.
01:07 CORRECTION (ledger integrity note): entry index 84 was edited in place to remove an incorrect rationale clause. The original clause justified the 25vr value-repair arm on MCTS grounds without noting that the DEPLOYED agent runs no search. Ground truth, verified in daily-top-agents-submitter/agent/main.py (module docstring: 'CPU-only raw-policy Kaggle agent'; zero occurrences of sims/mcts/search in the file): the Kaggle submission is a raw forward pass through the policy head, deck drafted once at module import. The '1600 sims' figure belongs to self-play data generation during TRAINING (design-note 2026-07-22 Search Strengthening), not to deployment. Consequence: the league's raw-policy sims 0/0 evaluation regime MATCHES the deployed regime - this is correct design, not a gap. 25vr remains justified solely as an RL prerequisite.
02:13 25vr root-caused and fixed; all six arms now running. FAILURE: 25vr died ~80s after each of two launches. opslib.stop_at correctly aborted ('container liveness absent before target', exit recorded to stop_events.jsonl) rather than waiting on a checkpoint that would never arrive. FORENSIC GAP FOUND AND FIXED: launch.py used --rm with docker-logs-native logging, so a container that DIES takes its diagnostics with it - no train-25vr.log was ever written and nothing on the box explained the death. Added launch._start_log_follower: after the RUNNING assertion, spawn a detached `docker logs -f ` writing to spec.log_path (logs -f replays existing output before following, so attaching post-assertion loses nothing; best-effort, never fails a good launch). The very next launch captured the traceback. ROOT CAUSE: torch optimizer.load_state_dict ValueError 'loaded state dict contains a parameter group that doesn't match the size of optimizer's group'. --value-only builds the optimizer from grad-enabled params only (value head), but the seed ckpt carries optimizer state saved over ALL params. train.py already discarded optimizer state on a regime switch; --value-only changes the trainable set while leaving regime='gameplay' unchanged, so that path did not trigger. FIX (training-gym 726c68e, Dan ruled option 1): extend the discard to any trainable-set change - param_set_switch = regime_switch or (VALUE_ONLY and not sd.get('value_only')); checkpoints now record value_only so a value-only->value-only resume keeps its optimizer state. Image rebuilt (pokemon-train:latest b649bd14ee32 -> f838595d79d6; Dockerfile.train is a COPY onto a cached base, running containers unaffected). Verified on relaunch: 'value-only resume from a full-parameter checkpoint: loaded model weights, fresh optimizer and LR schedule' then 'resumed from ckpt-2026-07-25vr-00813000.pt at step 813000'. SEPARATE FINDING - ZOMBIE CONTAINERS ARE REAL: two throwaway diagnostic containers could not be killed - `docker kill` returned 'tried to kill container, but did not receive an exit event' repeatedly and opslib.verified_kill correctly exited 4 (UNVERIFIED) rather than reporting success. GPU5 showed 4 MiB / 0%, so the processes were gone and only the daemon's bookkeeping was stuck. This retroactively validates the 22d zombie record (run_events ~292) against this session's earlier claim that a correctly-named docker kill always works; verified_kill's UNVERIFIED path is load-bearing and there is no in-band remedy without root. Box load avg ~158-265 under the full slate; ssh round-trips frequently exceed 120s.
02:28 Deck-vs-policy decomposition CORRECTED by range restriction; notes/ established. The earlier claim 'policy 70.5% / deck 17.3% / interaction 12.2%' was fit over ALL 127 rostered checkpoints and is an artifact of including garbage-tier ones - it answers a question nobody asks. Restricted to the decision-relevant range it INVERTS: top-50% 28.0/48.2/23.8, top-25% 22.2/52.2/25.6, top-10 1.4/78.2/20.3. Among top-10, deck regret ~85 Elo vs policy regret ~31 Elo (deck ~3x policy); interaction sd 48->37->32->21 Elo, a comparable third term throughout. Cells thick (median 356 games) so not a power artifact. CROSS-VALIDATION FAILED against the Kaggle 2x2: ckpt-2026-07-22c-00693000 has ZERO league games (never rostered), its lineage-mates 22c-00688000/00675000 show NO deck-18 collapse, and the league interaction has the OPPOSITE SIGN (deck14-deck18 = +28+/-13 at 22c-688k vs +54+/-16 at 23-813k, diff-in-diff -26+/-18, 1000+ games/cell). The Kaggle interaction is also decaying with maturity: 22c/deck18 670.7->731.6 and 23/deck14 919.5->946.5 within ~2h, dropping the implied interaction from +-95 to +-51; 23/deck18=888.1 is frozen at only 3h maturity by the 2-active rule and is therefore systematically undervalued. Treat the Kaggle interaction magnitude as UNVERIFIED. Established notes/ as an approval-gated trust tier per Dan's ruling (AGENTS.md 'Where knowledge lives'): notes/constraints.md (eval box -> no inference search -> raw policy ships -> value head does nothing at deployment; probed platform limits; box facts), notes/approach.md (AGZ self-play -> echo chamber -> IL pivot; signal-chain theory; ruled-out list), notes/end-state.md (co-evolution loop, three-check gate, Dan's 2026-07-26 rulings incl. autoresearch category map and 'convergence is not structurally guaranteed - detect, judge, revert, tune hyperparameters').
02:43 25vr COMPLETE AND ACCEPTED at ckpt-2026-07-25vr-00813500. Both ruled gate criteria met decisively on the clean 1,694-sample valrepair holdout (identical holdout to the baseline: 2 files, crc32%50==0): MSE 0.5087+-0.0214 -> 0.4406+-0.0183 (~3 SE improvement, floor 1.0000); BIAS +0.1850+-0.0167 -> +0.0001+-0.0161, i.e. the ~11-sigma midrange optimism that motivated the arm is eliminated. Holdout curve confirms the selection semantics were load-bearing: eval_value_loss 0.4406 (813500) / 0.4415 / 0.4438 / 0.4510 / 0.4529 / 0.4705 (816000) while train value_loss fell monotonically 0.4948 -> 0.4055 - the BEST checkpoint is the FIRST one at 500 steps (~7.5 epochs of the 68,692-sample window), and the blind 3000-step endpoint is 0.03 worse. The original blind-3000 spec (~45 epochs) would have been far worse. Stop fired end-to-end verified via opslib.stop_at: target_seen then fired_trainer step 816000, verified=true, exit_code=0, stages=['docker kill','poll inspect'], 'observed exited after docker kill' - the first fully verified checkpoint-targeted stop through opslib. 25dr ABANDONED by Dan (registry status set): registered as an RL prerequisite, but the draft head never ships (FORCED_DECK) so it cannot gate RL, and its only remaining rationale - generator quality for deck-candidate sampling - has no consumer yet (co-evolution P2/P3 unbuilt; D-v2 generates via the LLM). Revisit when P2 exists. Git: six local-only commits (launch --extra, ledger correction, AGENTS.md rename, log follower + submodule bump, notes/, 25dr registry) had lost their only remote copy when origin/opslib-build1 was deleted; merged origin/main (content-free - its 3 commits were merges of work already held) and pushed to main as b9e3bf0. training-gym pushed 07ecd92..726c68e (14 commits) through the outward gate with Dan's approval.
03:55 CROWN-DECISION PROGRAM STARTED (Dan ruled all five items). Objective reframed after a clean-context review: policy regret among top-10 checkpoints is ~31 Elo vs deck regret ~85, and the crown is soft-reversible through the promotion gate, so the goal is not to prove a winner but to buy the two measurements never taken - a Kaggle noise floor and a league rating for the actual submitted artifact. (1) Judge auto-submitter kill switch verified PRESENT (data/judge/judge_paused). (2) KAGGLE T+0: submitted ckpt-2026-07-22c-00693000 x forced deck 14 (submit_20260726.py, 102.0 MiB). Rationale: ckpt-2026-07-23-00813000 x deck14 was already ACTIVE at ~3h, so this creates a CONTEMPORANEOUS paired comparison of both crown candidates on one deck against one opponent field; the existing 892.4 for this pair is 2 days old against a different field and is not comparable to today's 949.0. PRE-REGISTERED PREDICTION logged before play: 860-915, point 885. Deck held constant by instrument assignment, NOT by any claim about interaction structure: the league randomizes decks per game and identifies deck/interaction effects with ~356 games/cell, while 3 unreplicated Kaggle cells against 28-61 points of drift have zero power to estimate an interaction (demonstrated: the apparent Kaggle interaction shrank +-95 -> +-51 as cells matured, and the league found the opposite sign). (3) 22c CONTINUATION launched GPU3 (tg-train-22c, manifest replay_manifest.txt, --lr 0.001 --gen-steps 0), verified 'resumed from ckpt-2026-07-22c-00693000.pt at step 693000'; stop armed at 700000. Implements the standing round-up ruling; 50,000 steps total = 18.52 epochs, same overtrain coordinate as 22g's ruled stop. (4) EVALUATION SEMANTICS RULED: added ckpt-2026-07-22c-00693000 to league.py CUTOVER_BASES. It is not a branch point - it is the terminal off-grid checkpoint (693000 % 12500 = 5500) of a lineage predating the round-up policy, and it is the artifact we submitted externally. Verified on the box: roster 140, 693000 now rateable. Rating both 693000 and 700000 also bounds the substitution error constraints.md flags as unmeasured. (5) CACHE EXPERIMENT: restarted 25c on GPU4 with --cache-files 6000 (was 4100); 25d left UNCHANGED as control. Tests the correlational hypothesis that 25c/25d run at ~880 steps/h vs ~9-10k/h because their corpora (5,714 / 5,709 files) overflow the 4100-file LRU parse cache while 25b (2,782) and 25a/25e (4,050) fit. RAM after restart: 748G free of 2267G. OPERATIONAL FINDING: verified_kill reported UNVERIFIED (exit 4) on tg-train-25c with 'docker kill ... did not receive an exit event', but the kill HAD succeeded - steps frozen, 0% GPU util, container absent from docker ps. So this is a FALSE NEGATIVE under load, not a failed kill; opslib fails closed, which is correct, but the five armed watchers may likewise report UNVERIFIED on successful stops. Third occurrence tonight. Six watchers confirmed armed: 22c@700000, 25a@975000, 25b@763000, 25c@875000, 25d@875000, 25e@838000.
05:04 22c CONTINUATION COMPLETE and crown schedule ARMED. 22c reached the ruled round-up target: ckpt-2026-07-22c-00700000.pt exists; stop fired verified via opslib.stop_at (target_seen 04:39:52Z, fired_trainer 04:40:02Z, verified=true, exit_code=0, stages=['docker kill','poll inspect'], 'observed absent after docker kill') - the second fully verified checkpoint-targeted stop through opslib. Both ckpt-2026-07-22c-00693000 (via CUTOVER_BASES) and ckpt-2026-07-22c-00700000 (on the 12500/{0,500} fine grid) are now league-rateable; roster 144. This closes the 'rated != submitted' gap for the crown's rival: the actual submitted artifact 693000 and its round-up 700000 can both be rated, and rating both bounds the substitution error that was previously unmeasured. ARMED UNATTENDED: com.kaggle-pokemon.crown-schedule launchd job loaded (hourly at :05), running crown_schedule.py. Verified under launchd's own environment (not just an interactive shell): exit 0, 'quota probe: 1/5 submissions today', slots not due, judge skipped. Failure branches tested BEFORE arming per the production-code rule: kill switch honoured; --status clean; --dry-run submits nothing; and with no credentials the judge ABSTAINS to disk rather than crashing. KILL SWITCH DECOUPLED (agent decision, surfaced to Dan and accepted): crown_schedule reads data/crown/crown_paused, NOT the judge submitter's data/judge/judge_paused. Rationale: the judge submitter is paused precisely so it cannot spend the 5 daily slots this schedule plans, and clearing a shared switch to start the crown run would have silently re-armed it. Consequence, stated explicitly: loading the job alone would have been inert under the shared switch; the decoupling is what makes it live. It will submit T+5 (23-813000 x deck14 byte-identical replicate, the noise floor) at 08:40Z and T+10 (22c-693000 x deck14) at 13:40Z, then emit a crown PROPOSAL after 18:00Z. Judge is claude-fable-5 at xhigh effort via the Anthropic SDK, streamed; thinking is always-on for that model so the parameter is omitted rather than configured, and no sampling parameters are sent. Judge applies the PRE-REGISTERED rule and abstains rather than forming its own opinion; it proposes only and never mutates the registry (ratified:false). API key deployed to ~/Library/LaunchAgents plist (chmod 600); the repo template keeps a placeholder so no secret is in git. Note for anyone scraping ~/.zshrc for the key: ANTHROPIC_API_KEY is set by indirection there, so a naive grep yields the literal string '$DEEPTUNE_CLAUDE_KEY' rather than the key.
09:55 LEAGUE THROUGHPUT INTERVENTION (Dan-authorized, overnight loop). Measured league throughput was 2.7 games/min with the crown rival ckpt-2026-07-22c-00693000 at 1 game — matchmaking was thrashing across 54 active players on 2 GPUs with ~20s server spin-ups. Dan authorized (a) scoping the roster and (b) giving the league more GPUs. ACTIONS: (1) Roster scoped via the established disabled.json hot toggle (backup at league/disabled.json.bak-20260726-crownscope) to 12 active players: champion ckpt-2026-07-23-00813000, rival ckpt-2026-07-22c-00693000, ckpt-2026-07-22c-00700000, plus anchors ckpt-2026-07-22c-00675000/00688000, ckpt-2026-07-23-00750000/00800000, ckpt-2026-07-22e-00700000, ckpt-2026-07-24w-00813000, ckpt-2026-07-22i-00200000, ckpt-2026-07-22g-00050000, ckpt-2026-07-20-00813000 (high-game anchors spanning Elo ~880-1260 to bridge the scale). Existing ratings untouched; re-enable by restoring the backup. Newly-rostered 25a/25b checkpoints also disabled for the crown window so BURST_BOOST doesn't divert games from the rival. (2) tg-train-25d DIED ON ITS OWN at 09:53:34Z — docker events show die exit 137 (SIGKILL, consistent with kernel OOM: host available RAM had fallen to ~206G; 639G free after death) then auto-remove; NOT stopped by opslib (my verified_kill found no such container, exit 2) and NOT its watcher (stop_25d.log empty; watcher self-exited on absence). Timeline note: league worker w2 on GPU6 was launched at 09:52:55Z (pre-existing, not by this loop) ~40s before the OOM — plausibly the marginal memory pressure. Per Dan's authorization 25d is NOT relaunched; its GPU goes to the league. 25d halts at step ~657840 of ruled 875000, hypothesis unresolved. (3) League now 3 game workers + fit: fit/GPU3, w1/GPU5, w2/GPU6. This loop accidentally launched a duplicate LEAGUE_WORKER=2 (tmux window league:3) and killed it within a minute; the surviving w2 is the pre-existing PID 1601200. PLANNED: when 25b's armed stop fires at 763000 (~12:20Z), verify it and start LEAGUE_WORKER=3 on GPU2. Roster scoping is evaluation-semantics class but explicitly ruled by Dan mid-loop ('Yes feel free to do both, and any other similarly scoped things are authorized').
09:58 CORRECTION + SINGLE-OWNER HANDOFF. Corrects the previous entry: tg-train-25d was NOT OOM-killed — it was killed by the OTHER agent's opslib.verified_kill at ~09:53:34Z, which returned UNVERIFIED (4th observed false-negative of the docker-kill exit-event bug); the die 137 in docker events is docker kill's SIGKILL. The stop remains Dan-authorized and correct; only the attributed cause changes. TWO AGENTS were driving the cluster concurrently for ~15 min (both scoped disabled.json with different keep-sets, both killed 25d, both launched a GPU6 league worker — tmux name collision luckily prevented a double worker; whole-file rewrites of run_events.json were the worst hazard, no entry lost on inspection). Dan ruled the other agent stands down; this loop is sole owner now. RECONCILIATION: the two overlapping disabled.json writes had left ~29 players active (unintended 22e-01xxx tail and legacy ckpt_* entries); one unifying write scoped to the UNION keep-set of both agents, 15 active players: ckpt-2026-07-23-00813000 (champion), ckpt-2026-07-22c-00693000 (rival), 22c-00700000, 22c-00688000, 22c-00675000, 23-00750000, 23-00800000, 22e-00700000, 24v/24w-00813000 (twins, noise calibration), 22i-00200000, 22g-00050000, 20-00813000, ckpt_00650000, ckpt_00050000. Backups: disabled.json.bak (other agent), .bak-20260726-crownscope (this loop, pre-scope), .bak-20260726-unified-pre (pre-unification). MEASURED EFFECT: league throughput 2.7 -> 20 games/min; rival went from 1 game to +11 games in a 2-min window (BURST_BOOST engaging). League: fit/GPU3, w1/GPU5, w2/GPU6. PLANNED: after 25b's armed stop fires at 763000 (~12:20Z), verify from artifacts and start LEAGUE_WORKER=3 on GPU2.
10:02 KAGGLE NOISE-FLOOR OPERATIONALIZATION (Dan delegated the ruling to the overnight loop; crown ratification remains Dan's). Noise floor := |replicate - original| for the byte-identical pair ckpt-2026-07-23-00813000-forceddeck14: original ref 54986632 (frozen at 941.6) vs replicate ref 54996700, evaluated ONLY once the replicate is mature (score moves <5 points across two consecutive hourly checks, or frozen by displacement). Until maturity the floor is UNDEFINED, not zero. Sign-agreement test for the pre-registered rule's 20-50 Elo band: Kaggle agrees in sign iff (champion - rival) same-deck, both scores mature-or-frozen, exceeds the noise floor. State at ruling time: replicate moved 945.3 -> 989.0 in ~35min (immature); same-deck champion-rival gap 989.0 - 885.8 ~ 103 (also moving); replicate pair currently differs by ~47, so no conclusion is drawn yet. Caveat on record: the original froze by displacement after ~4-10h and may have frozen immature, which would OVERSTATE the floor - conservative for the 20-50 band, accepted. Handover from the stood-down agent also on record: verified_kill has 4/4 UNVERIFIED false negatives (always confirm kills from artifacts); run_events.json is whole-file rewrite, single-writer = this loop; 25d checkpoints persist at ~657850 and can resume with --cache-files 6000, watcher deliberately disarmed.
12:43 25B STOP VERIFIED + LEAGUE WORKER 3 STARTED (pre-authorized plan). 25b's armed stop fired at 12:37:18Z on target ckpt-2026-07-25b-00763000.pt; opslib reported UNVERIFIED exit 4 ('did not receive an exit event') — the 5th observed false-negative of this docker bug. Confirmed dead from artifacts: container absent from docker ps, metrics frozen at step 763060 (last write 12:37:11Z), GPU2 0 MiB, target checkpoint file exists. 25b reached its ruled stop exactly; no relaunch. GPU2 handed to the league per Dan's authorization: LEAGUE_WORKER=3 LEAGUE_GPU=2 LEAGUE_FIT=0 LEAGUE_MAX_SERVERS=60 started in tmux league window w3, verified alive (python3 league.py PID 4057677). League now fit/GPU3 + w1/GPU5 + w2/GPU6 + w3/GPU2. Remaining arms on schedule: 25a 917980/975000, 25c 721740/875000, 25e 791560/838000; 3 watchers armed.
17:12 DECK-HILLCLIMB-V2 PHASE 0 (Dan-ruled, this session sole cluster driver for the deck loop). Rulings: evaluator re-pinned rival ckpt-2026-07-22c-00693000 -> champion ckpt-2026-07-23-00813000; promotion gate stays the frozen constitution rule (Wilson LCB, SE_INFLATION 1.5) — Dan explicitly declined the autogo 0.55 copy after applicability review; budgets 50k eval games + 40 LLM proposer calls per night. ACTIONS: orchestrator._patch_cluster_eval now pins CKPT_SRC to the champion (v1 deck-evolution untouched); deck-registry/condition.json evaluator_policy updated (backup condition.json.bak-20260726-prepin); hc-v2 ckpt/ symlink swapped (old symlink in ~/.trash on the box); tg-probeD-server killed for relaunch under the new pin; shared.json pool_subversion 1->2 to separate pre/post-pin measurement conditions (interpretation logged in probes/deck-hillclimb-v2/ESCALATE.md); fresh baseline launched: baseline --jobs 24 --games 40, log baseline-20260726-champion.log. Driver loop armed in-session at :13/:43 each hour; no Kaggle submit path in the loop.
17:31 25E STOP VERIFIED + W4 + 24V SUBMITTED (Dan-ruled search pivot, items 1&2). 25e's stop fired 16:40Z at exactly 838000, UNVERIFIED exit 4 (6th false-negative), confirmed from artifacts: container absent, metrics frozen at 838000, GPU7 0 MiB, ckpt-2026-07-25e-00838000.pt exists. LEAGUE_WORKER=4 started on GPU7 (league now fit/GPU3 + w1/w2/w3/w4 on GPUs 5/6/2/7). SEARCH PIVOT per Dan's ruling: (1) 25-arm intermediates re-enabled in disabled.json (25a-00863000, 25b-00688000, 25b-00725000; active roster 15->18, new arm checkpoints auto-admit as published); (2) ckpt-2026-07-24v-00813000 exported to TorchScript (export_checkpoint.py --once in pokemon-train:latest), synced, and SUBMITTED to Kaggle x forced deck14 (4/5 today, submit_20260726b.py). Purpose: first external score for the league leader (24v 1209.9 vs champion 1203.4 at 16:29Z fit). PRE-REGISTERED PREDICTION: range 920-1000, point 960. Side effect stated: this displaces the champion replicate (ref 54996700) from the active pair, freezing it ~8.6h mature — the champion noise-floor pair (54996700 vs 54986632 frozen 941.6) becomes DEFINED at whatever score it froze on. Deck-search finding also on record: deck-hillclimb-v2 idle since 07-25 22:38Z with tg-probeD-server holding GPU1; restart proposed to Dan, awaiting ruling on evaluator pin (champion vs rival vs post-crown).
17:44 LOCAL DISK FULL — MITIGATED, ESCALATION OPEN. Dan's Mac data volume hit 100% (197 MiB free) during the candidates rsync (exit 10, 'No space left on device'); the 24v .ts had already landed so the submission was unaffected. Risk: the 18:05Z judge and any bundle build (102 MiB each) write locally. MITIGATION: deleted data/staging/* (2.6G of already-submitted bundle staging, deterministically regenerable via committed package.py from candidates+forced-decks; ALLOW_DATA_RM override, logged to gate-triggers.log). Now 2.8 GiB free — judge safe. NOT TOUCHED, NEEDS DAN: data/candidates is 161G (1,509 mirrored .ts exports, re-syncable from /shared/checkpoints/kaggle-pokemon) on a 926G volume that is otherwise ~73% non-project data. Options: prune the mirror to crown-relevant + recent lineages (rsync exclude list), or move it off-volume. Also: candidates.sync_from_cluster() uses rsync --delete of the FULL export dir — every new cluster export grows the Mac mirror unboundedly; suggest a filter. Awaiting ruling.
18:32 25A STOP + W5 + JUDGE INSTRUMENT BUG FIXED. (1) 25a's stop fired 18:07Z at target 975000 (metrics last row 975020), UNVERIFIED exit 4 (7th false-negative), confirmed dead from artifacts: container absent, GPU0 0 MiB, ckpt-2026-07-25a-00975000.pt exists. LEAGUE_WORKER=5 started on GPU0; league now fit + 5 workers (GPUs 0/2/5/6/7 + fit GPU3). Only 25c still training (783k/875000, 1 watcher armed). (2) The 18:05Z judge ABSTAINED (valid outcome, reported to Dan verbatim, NOT overridden) but its abstain cause was an INSTRUMENT BUG, not evidence: crown_schedule.league_ratings() read /shared/pokemon-train/league_fixed_elo.json — path missing the /league/ dir — and parsed key 'ratings' where the real file uses 'elo'. Both fixed (one-line path + parse fallback chain); verified live: returns champion 1196.9, rival 1164.4, 688000 1187.2, 700000 754.6. Classified as remediation restoring the ruled measurement condition (the judge was always meant to read the real fit); the rule text itself untouched. Judge re-runs at 19:05Z on real data. (3) Kaggle state at tick: champion replicate FROZE at 973.4 (displaced by the 24v submission) -> champion-pair noise floor |973.4-941.6| = 31.8; judge computed rival-pair floor 17.0 from 873.6 vs 891.8 (rival replicate still active/moving). 24v first reading 761.4 at ~1h — far below the pre-registered 920-1000, but immature; no conclusion until mature per ruling.
19:05 LEAGUE TORN DOWN (Dan ruled, 19:0xZ). tmux session `league` killed (fit + w1-w5), all 60 tg-league-* containers docker-killed. Verified from artifacts: 0 league.py processes, 0 tg-league containers, GPUs 0/2/3/5/6/7 at 0 MiB. Final pre-teardown fit (16:29Z-19:01Z era): champion 1196.1 (20,234 g), rival 1162.3 (537 g), Delta ~33.8; new finals 25b-763000 1207.8 (466 g), 25a-975000 1160.4 (61 g), 25e-838000 1149.2 (148 g) — 25b's above-champion reading is at 466 games, NOT yet decision-grade. Still running: tg-train-25c on GPU4 (recency-corpus arm, ruled stop 875000, ~790.7k at teardown, watcher armed) and tg-probeD-server on GPU1 (deck loop, other session). League ledger final line count 308,183. Restart procedure unchanged (README run-operations).
19:14 25C STOPPED EARLY (Dan ruled 'take the trainer down too') — RULED EARLY STOP, NOT A FIX: the recency-corpus arm halts at step 793020 of ruled stop 875000, hypothesis partially resolved; checkpoints exist through ckpt-2026-07-25c-00793000.pt; nearest on-grid rateable coordinate BELOW is 788000 (record as substitution per constraints.md terminal-stop fallback). KILL SEQUENCE NOTE: verified_kill returned UNVERIFIED exit 4 (8th occurrence) and for ~1 min the container stayed visible as 'running' (PID present, GPU memory held, metrics frozen, util 0%) — the daemon reaped it late; the follow-up force-remove found no such container. So: even a stuck-visible container after UNVERIFIED can be a slow-reap false negative — wait 60s and re-check before escalating. 25c's watcher self-exited on absence; orphaned stop/lease tmux sessions cleaned. CLUSTER END STATE: all 8 GPUs free except GPU1 (tg-probeD-server, deck loop, 1.2 GiB). No trainers, no league, no watchers. Overnight arms final coordinates: 25a 975000 (ruled stop, done), 25b 763000 (ruled stop, done), 25c 793020 (ruled early stop), 25d 657840 (unresolved), 25e 838000 (ruled stop, done).
20:05 Box cleanup + verified_kill false-negative fix. CLEANUP (re-verified no trainers, no tmux server, all lease flocks free at the moment of mutation): deleted ~/launch_arm_v3.sh (its single $A argument served as both container suffix and checkpoint-prefix branch letter, so no call produced both a correct container name and a correct prefix -- 'a' gave tg-train-a, the tg-train-e shape; superseded by opslib.launch), and the historical one-offs ~/stop_at_400k.sh and ~/stop_22c_at_16ep.sh. Cleared 7 stale leases/*.holder.json (flock-free; lease.sh does not always remove its metadata when the leased container is killed rather than exiting on its own -- cosmetic, lease.sh status now reads 'no leases held'). Repo: deleted merged local branches pr6-merge and pr6-rebase. SLATE OUTCOMES, first production use of opslib.stop_at: 25a stopped at its armed 975000, 25b at 763000, 25e at 838000, 25vr at 816000 -- every arm at its exact armed step with ZERO post-stop checkpoints (compare 22d, which overran to 846000, and 22e, which overran 155,000 steps and left 154 post-stop checkpoints). 25c aborted at 793000 with 'container liveness absent before target' against an 875000 target -- the arm ended early and the watcher said so instead of waiting forever. 25d reached 657000 with no watcher record. 25vr fired in a single stage (docker kill, then poll) with no escalation, because the container name was correct -- the argument against the sudo/pkill ladder, demonstrated. FALSE-NEGATIVE FOUND AND FIXED: 25a, 25b and 25e each recorded verified=false exit_code=4 with stages ['docker kill'] and detail 'docker kill rc=1: ... tried to kill container, but did not receive an exit event'. The daemon signalled the container but did not observe the exit inside its own timeout; all three had in fact stopped correctly. verified_kill was returning at stage 1 on any unrecognised non-zero rc WITHOUT polling docker inspect -- reporting a failure it had not verified, the exact mirror of the 22d watcher reporting a success it had not verified. A night operator gating on exit 0 would have raised three false alarms on three correct stops. Fixed: only 'No such container' short-circuits (naming bug, never escalate); every other non-zero rc falls through to the poll, and an observed EXITED/ABSENT is a verified stop with the CLI's complaint recorded rather than obeyed; UNVERIFIED is reserved for the poll itself being unable to confirm.
21:52 DECK SWEEP LAUNCHED (Dan ruled: sweep before spending the single Kaggle slot). Objective is P(beat 973.4), not learning. Design: opponent pool FROZEN at the 19 meta decks (Dan ruled); new decks from the 07-24/25 episode dumps of current top-LB teams enter as CANDIDATES only; two pilots run SEQUENTIALLY on the single available GPU1 (champion ckpt-2026-07-23-00813000, then ckpt-2026-07-25a-00913000); single-stage full factorial ~400 games/cell then adaptive top-up of the top 4; hard 100-minute wall clock, phases bounded by time not game count. Reuses probes/deck-hillclimb-v2 evaluator + tg-probeD-server (port 5698, GPU1). Measured reference rate: baseline 960 games / 101 s. Built-in calibration: deck 14 vs 18 under the champion has a known external sign (973.4 vs 888.1); if the sweep does not reproduce 14 > 18 the internal ruler is declared untrustworthy for deck selection. Artifacts: /shared/pokemon-train/scratch/decksweep-20260726/results.json. NOT touched: pool.csv, league state, hillclimb KILL file.
22:53 DECK SWEEP COMPLETE + SUBMISSION 55012347 (Dan ruled the submit). Sweep: 25 candidates (19 frozen pool + 6 mined from the 07-24/25 dumps) x 2 pilots, 399 games/cell, ~27,500 games in ~35 min at 534-687 games/min; pool2/25a short at 379 games (two engine crashes 'invalid index 1 >= 1'); phase C top-up cancelled for time. CALIBRATION PASSED: deck14 vs deck18 under the champion = 74.5 Elo internal vs ~85 external, same sign (~0.9 transfer) - first demonstrated internal->external transfer in this project; note it is the frozen-pool deck sweep, NOT league Elo. RESULT: deck 14 is #1 under both pilots (top-4 identical across pilots); no mined top-team deck beat it (best 0.555, 5th of 25). Decision cell deck14: champion 0.597 [0.548,0.643] vs 25a-00913000 0.659 [0.611,0.704], +46 Elo, z=1.81, p=0.070. Corpus-frequency scan: deck14 = 34.5% of 35,964 declarations, 82x deck18; Spearman rho 0.54 vs win rate. SUBMITTED ckpt-2026-07-25a-00913000 x forced deck 14, ref 55012347 at 22:51:02Z, 102.1 MiB, novel pair, pre-registered prediction [950,1050] point 1000. OPEN: tg-probeD-server still pinned to 25a-00913000 (ckpt-dir /shared/pokemon-train/scratch/decksweep-20260726/ckptB), NOT the original champion pin - restore before any further probe use. Artifacts: probes/decksweep-20260726/results.json; workbook entry '2026-07-26 Deck Sweep'.
04:18 SUBMISSION 55012347 RESULT + PROBE PIN REVERTED. ckpt-2026-07-25a-00913000 x forced deck 14 scored 854.3 at 5.3h active, against a pre-registered [950,1050] point 1000 - a miss of 146 below point / 96 below floor. Co-active same-window comparison: 24v-00813000 x deck14 = 900.6, i.e. 913000 scored ~46 BELOW a net the sweep never tested. Methodological cause: the sweep's calibration validated the DECK axis (deck14 vs deck18 under a fixed pilot, 74.5 internal vs ~85 external); the 0.9 transfer coefficient was then extrapolated to a PILOT comparison on a fixed deck, an axis the calibration did not cover. Transfer is established for deck comparisons under a fixed pilot only. PIN REVERT (Dan ruled: keep 913000 unless it underperforms; it did): tg-probeD-server killed and re-run with ckpt-dir /shared/pokemon-train/probes/deck-hillclimb-v2/ckpt (champion ckpt-2026-07-23-00813000), GPU1, port 5698, net host, same image/args as before; verified from artifact - container RUNNING and log line 'loaded ...ckpt-2026-07-23-00813000.pt', 'serving on 0.0.0.0:5698' at 04:18:08Z.
05:29 BENCH-RICHARD DASHBOARD STARTED on cluster port 7072 (Richard's stays on 7071, untouched). nohup python3 ~/runs/bench-richard/dashboard/dashboard.py --port 7072 reading ~/runs/bench-richard/{results/full,replays,results/replay50}; PID 910788; all endpoints verified 200 via curl. Read-only over run artifacts; no /shared/ writes. Code at probes/bench-richard-20260727/dashboard/ (branch bench-richard-pilot). Aggregate finding surfaced: on card-select menus with >=10 options, 68/111 (61.3%) of meow choices were the [0..k] index prefix (56 exactly [0]).
05:53 BENCH-RICHARD DASHBOARD MOVED TO MAC. Cluster instance on port 7072 (PID 910788) stopped via kill; verified port 7072 free on the box; Richard's 7071 untouched. Dashboard now hosted on dans-mac-studio (bind 0.0.0.0:7072, tailnet http://100.87.46.115:7072), two-arm version (okidogi n=1000 + mirror-deck14 n=354 TRUNCATED), data pulled by probes/bench-richard-20260727/sync-data.sh (rsync, idempotent, safe mid-run); start/stop via dashboard/serve.sh. Read-only over run artifacts; no /shared/ writes.