Just in case anybody sees me in Bicester next week - yes, I’m flying out to the UK tonight, but it’s a weird trip where I will be home for two days (or more like one and a half), before I head to London for a tech conference, and then on Friday night I’m off to see Saint Etienne on their final tour. So I am not likely going to be able to meet up with anybody on this visit as it’s going to be a whirlwind adventure, but we’ll be over for two weeks in December, so hopefully I’ll see a bunch of you then.
(also, continued apologies for ruining Tammy’s love of the Trader Joe’s Vanilla Chantilly cake)
One of the most eyebrow-raising parts of the recent HuggingFace hack was the co-ordination of all the agents, where they found a dormant messageboard and used it to exchange messages (there’s also evidence that they’ve been leaving messages to each other on public wikis too). Really, we shouldn’t be that surprised; software developers in general are messy gossips that would make a gaggle of old women down at the Post Office nod their heads in respect.
Anyway, if they love a messageboard, maybe we just give it to them? I built Babble, a CLI-based tool that you can safely hand off to your agent swarm. It does all the usual things you’d expect of such a tool at this point; it’s fully self-documenting and there’s a skills file you can slot in to your agent workflow. The ‘forum’ it presents allows threads, at-mentions, and as of right now, emoji reaction support, because it’s a forum, after all. As well as all that, it comes with a web view so you can keep an eye on what they’re talking about, and even inject comments in the threads yourself, just in case they’re trying to pull a SkyNet.
Is it actually useful? Well, I’ve only been using it since Wednesday, and it has taken a little nudging to get the agents comfortable posting all over the place instead of confining themselves to one thread. But it’s at least interesting to see them leaving notes for each other. I’ve created a thread where they can ask for improvements to the board. I’ll likely wake up tomorrow to a massive moderation flame war and entrenched factions. Just like a proper board…
of course it's in Rustno poll support yet, mind youforum for terminators
Continuing my middle-class parenting adventures, I am confused. The above is a pretty impressive playground; maybe not huge, but it has a three-level play structure with lots of different slide types, a merry-go-round, big swings and little swings, a dedicated toddler area, two glockenspiels and drums, as well as picnic tables and benches given shade by massive umbrella structures. Plus, it’s very blue. And yet, at 10am on a Sunday morning, it’s completely deserted. It doesn’t make sense — Maeryn and I have been to the playground that is almost equidistant from our house at the same time and it’s usually packed. So why don’t people go to Bicentennial Park? It even has better parking than West Fork!
One thing I have come to realize over the years is that I am very prone to falling for gadgets. This means I have a graveyard of electronics that has followed me from Bicester to Manchester, across the Atlantic to Chapel Hill and back again…and then back again via shipping container and moving van to Cincinnati. There’s all sorts of weird and wonderful things; animal toys that click together and play different tones depending on the arrangement, my Sony Cyber-Shot DSC-U10 (still so cute!), the period of time where I’d take a Handspring Visor on the bus to Oxford every day before I gave in and finally bought my first real mobile phone (and that iPhone is probably still around somewhere too).
There’s all sorts of Bluetooth things, and then a spurt of things from 2022 when Tammy and I were throwing around ideas for another escape room, this time having one room across three different time zones. We really went all out on this; I have rolls of NFC sticker tags hanging about, and I had built a working Pepper’s Ghost effect that would have been used, plus I had already had the idea to pull in (then quite early) models capable of voice cloning and facial recognition. We probably should have written more of it down…maybe when Maeryn is a little older, we can run a version of it…
Then there’s the more recent things. The RFID/NFC battery-less e-ink photoframe (which promises a lot more than it can deliver, but some photos come out really nicely on it), the Xteink mini e-readers, and then the influx of ESP32 stick devices. These were Matt Webb’s fault. An absolutely tiny form-factor (smaller than a pack of chewing gum!) and yet it has 8Mb of memory. Anyway, I now have 5 of these because…of reasons. Finally, the most recent arrival is the Cardputer ADV, which is similar to one of those sticks but it has a “full” keyboard and looks so 90s-hacker-coded that I feel like I’m in Sneakers every time I look at it.1 Easily worth $30.
And then there’s Maeryn’s toys. When she got the Toniebox a couple of years ago, I did have a poke around at the state of reverse engineering, although sadly I came to the conclusion that it probably made sense just to play around with Creative Tonies instead of trying to flash (and likely brick) the Toniebox itself.
There’s also a competitor to Toniebox called Yoto, and one of the things that Yoto has going for it is that their Yoto Mini system is considerably smaller than a Toniebox2, and while it doesn’t have the cute range of figurines, the cards are so much easier to pack for travel. Tammy found a great deal on the Mini recently, and it arrived this weekend. I hadn’t really paid that much attention to the system, but when she took it out of the box, I was struck by how fun and little it was, with a simple 16x16 LCD panel that can be used by stories for little icons as stories are playing.
Anyway, more out of idle curiosity than anything else, I searched to see if there was a similar reverse engineering effort. And instead I landed on the developer API page. I’m fairly sure my eyes widened as I saw API contracts, MQTT events detailed, and quickly found a Python wrapper. You can even point to content that isn’t on their servers…and you can alter that content and get button push information from the box itself over the internet. Five minutes later, I had the idea of “Daddy’s London Travelogue”, where I can send an update (with fun icons!) every day I’m going to be in the UK in September, and Maeryn can put in a card to get the latest update every evening.
Naturally, before 13:00 on Saturday, Claude and I had this working. But there’s even more. The way that the Yoto system works…it doesn’t really care if you swap files about. Even if you do it dynamically. You could use the two button presses from the device to make a simple game work. Which in itself would be fun, but there’s nothing stopping you from using a separate controller system (e.g. a tablet or phone or dedicated ESP32 device…) and using that to control the webapp that you host to alter what gets sent to the Yoto Mini.
It was at that point I had to physically separate myself from the computer, as I was already plotting something akin to an aural version of Knightmare3, and it would have sucked up the whole rest of my day. And funnily enough, a good part of the day I should be actually spending with Maeryn…
It seems like Tonie might be feeling the pinch a little here as they’ve just released the Toniebox Lite, which is a lot closer to the Yoto Mini’s size. And it also has USB-C charging, which is making it a tempting replacement… ↩︎
I built another thing. I don’t think it’s useful, not really, but I still built it. If, by some chance, you had always wanted a online search engine dedicated to Visions of Heaven and Hell, a 1994 documentary broadcast on Channel 4, then…congratulations? And we should probably say hello at some point?
Most of this was two afternoons, and the second was spent fiddling with CSS and picking an embedding model that would be fine in a browser process1.
I still can’t really tell you why I did it, either. I spent a good part of my free evenings in early July mainlining a lot of mid-90s BBC/C4 docs on the Internet and what it was going to do to us. Smiling at some bits and going “Oh, well, damn, you nailed that one.” Simultaneously re-reading The Invisibles to fully bask in that there’s-a-new-world-coming feeling2. I’d never seen Visions before, somehow, and it just felt like something I wanted to cut up and process with 2026 open source technology.
If you really want to be afraid, I have over 500 editions of Horizon on my NAS and I think the world needs them to be searchable…
The Tree-SAE model took a bit longer, but as you might imagine, that was built for a potential revamp of my Adam Curtis site, so it was just hanging around and I thought it would transfer pretty nicely… ↩︎
I was told there would be a lot more smart drinks and programmable nanofields in my future. ↩︎
nice and smooth2012 was a long time ago, Grantchannel 4 way back when
When I came here, a decade-and-a-half ago, I am sure I did not imagine that fifteen years later, I’d be lying on the floor in a completely different state, singing Appetite to my three year old because she wanted to hear the songs “you sang me when I was little.”
And yet, here I am, and I can’t imagine anything better. In another fifteen years, Maeryn will be eighteen, and I can’t imagine that, either, but I know it’ll be here faster than an eyeblink.
In what was otherwise a pretty terrible week, it was wonderful to see a section of my social media follow to completely lose their top over discovering that one of our long-term Holy Grails sneaked onto YouTube a few months ago. Yes, that’s right, the complete, unedited, and unabridged Statutory Right of Entry is now available for your unbridled CSO enjoyment.
(there’s even a comment on the page from one of the VT people that made it happen!)
Ever since I first read the World Models paper, there has been a nagging irresistible thought at the back of my mind: “I want to try that out on Deathchase". And, to be fair, I have over the years spent a few Sunday nights trying to make it work. The RNN or the VAE code wasn’t the problem; I have working code on my old 1080Ti PC tower that trains well enough on the example toy racing world. The problem was somehow hooking it up to a Spectrum. I tried a lot of different approaches, including some major surgery to MAME, and I got tantalizingly close, to the point of hand-disassembling Z80 code to work out where information like lives and scores were being kept but never enough to really make it work in a fashion I could run a full training loop without it keeling over and dying a few iterations in.
(you know how this is going to go)
It popped up in my head again recently, and this time I asked Claude. It laughed, saying: “Why do you not just use zx, as it’s basically built exactly for what you want to do?” because I didn’t know it existed, you smug little SkyNet! And then we set to work.
The World Model paper is actually three different models. Firstly, a variational autoencoder (VAE) takes the pixels from the screen (resized in the paper to 64x64, in my code to 84x84) and compresses them down to a latent vector. Just 64 numbers to describe everything going on in every single frame fed into the model. This is then chained into a LSTM network, which takes the latent vectors and learns to predict what comes next. These two models together form the world model; eventually, you can feed a starting frame into the autoencoder and then ‘dream’ the entire sequence of the game inside the LSTM. And this is the trick that seemed magic in 2018; it’s annoyingly hard to do reinforcement learning or any sort of gradient descent work when you have to keep going back to the game state, perhaps in an awkwardly coded emulator. There’s a lot to go wrong. Here, the controller model trains directly on the dream world and then you transfer the trained controller model back to the real world…and it works! Mostly.
Fig 1 · V–M–C pipeline
%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace, monospace','fontSize':'14px','lineColor':'#8b97ac','edgeLabelBackground':'#0f1420','tertiaryTextColor':'#c7d0de'},'flowchart':{'curve':'basis','nodeSpacing':46,'rankSpacing':60,'padding':10}}}%%
flowchart LR
o["obs oₜ 64×64 frame"]
subgraph WM["world model · runs with no environment (the dream)"]
direction LR
V["V · VAE encoder compress → zₜ ∈ ℝ³²"]
M["M · MDN-RNN predict P(zₜ₊₁ | zₜ, aₜ, hₜ)"]
V -->|"zₜ"| M
end
o --> V
M -->|"zₜ, hₜ"| C["C · controller 1 linear layer · ~900 params"]
C -->|"aₜ"| E(["real environment"])
E -->|"next frame"| o
classDef ha fill:#2a2214,stroke:#eaa63f,stroke-width:2px,color:#f3e7d4;
classDef ctl fill:#1b2130,stroke:#cfd6e2,stroke-width:2px,color:#eef2f8;
classDef io fill:#141a26,stroke:#6b7688,stroke-width:1.5px,color:#c7d0de;
class V,M ha; class C ctl; class o,E io;
style WM fill:transparent,stroke:#eaa63f,stroke-width:1px,stroke-dasharray:5 5,color:#c69a52;
V compresses the frame · M predicts the future in latent space · C maps the latent state to an actionmermaid source
flowchart LR
o["obs oₜ<br/>64×64 frame"]
subgraph WM["world model · runs with no environment (the dream)"]
direction LR
V["V · VAE encoder<br/>compress → zₜ ∈ ℝ³²"]
M["M · MDN-RNN<br/>predict P(zₜ₊₁ | zₜ, aₜ, hₜ)"]
V -->|"zₜ"| M
end
o --> V
M -->|"zₜ, hₜ"| C["C · controller<br/>1 linear layer · ~900 params"]
C -->|"aₜ"| E(["real environment"])
E -->|"next frame"| o
classDef ha fill:#2a2214,stroke:#eaa63f,stroke-width:2px,color:#f3e7d4;
classDef ctl fill:#1b2130,stroke:#cfd6e2,stroke-width:2px,color:#eef2f8;
classDef io fill:#141a26,stroke:#6b7688,stroke-width:1.5px,color:#c7d0de;
class V,M ha; class C ctl; class o,E io;
style WM fill:transparent,stroke:#eaa63f,stroke-dasharray:5 5,color:#c69a52;
After spending an afternoon with Opus a couple of weeks ago, the first discovery was that Ha’s world model…just didn’t work very well for Deathchase. It did learn a grand plan of surviving…which was to slowly go through the first day, shoot at a few things, and simply sit tight when night falls. Which is a strategy, but one that would have got you slapped if you tried it back in the day round your friend’s house. Also, it does have a tendency to hit the trees a lot. But then I have that problem as well.
Despite that, looking at the dream sequences of the VAE model is spooky; you can see exactly what the model is seeing and, yep, well, it’s certainly Deathchase.
Anyhow, I was not to be stopped by that little failure. Because I kept reading these sorts of papers, and I specifically remembered DreamerV3 from DeepMind. Back in 2023 it pretty much set the new state of the art for autonomously playing Atari games. So Claude, myself, and the DGX Spark worked for a day or two to train a new model based on this new approach.
The DreamerV3 framework is a lot more complicated (see the image, yikes!), but I’d say that the too key points are: a CNN model to better capture what’s going on versus the older autoencoder, and a training regime that keeps the world model anchored to reality, with joint training focused on reconstructing the frame, optimising the reward, and whether the model has left the game in a playable state. And this works a lot better.
Fig 2 · the DreamerV3 world model, one timestep
%%{init:{'theme':'base','themeVariables':{'fontFamily':'ui-monospace, monospace','fontSize':'13.5px','lineColor':'#8b97ac','edgeLabelBackground':'#0f1420'},'flowchart':{'curve':'basis','nodeSpacing':40,'rankSpacing':58,'padding':10}}}%%
flowchart LR
prev["hₜ₋₁ · zₜ₋₁ · aₜ₋₁"] --> GRU["sequence model GRU"]
GRU --> H["hₜ deterministic memory"]
X["obs xₜ"] --> ENC["encoder CNN"]
ENC --> Q["representation q posterior zₜ | hₜ, xₜ"]
H --> Q
H -.-> P["dynamics p prior ẑₜ | hₜ · no obs"]
Q -. "KL: pull prior → posterior" .-> P
H --> S["state sₜ (hₜ, zₜ)"]
Q --> S
S --> DEC["decoder → x̂ₜ"]
S --> REW["reward → r̂ₜ"]
S --> CON["continue → ĉₜ"]
classDef dr fill:#123039,stroke:#43c0d0,stroke-width:2px,color:#dff3f6;
classDef prior fill:#123039,stroke:#43c0d0,stroke-width:2px,stroke-dasharray:5 4,color:#dff3f6;
classDef neut fill:#1b2130,stroke:#8b97ac,stroke-width:1.5px,color:#e9edf6;
class GRU,H,Q,S dr; class P prior; class prev,X,ENC,DEC,REW,CON neut;
The prior predicts the latent from memory alone — the dream path; the posterior corrects it using the real observation during learning. The dashed KL ties them together.mermaid source
flowchart LR
prev["hₜ₋₁ · zₜ₋₁ · aₜ₋₁"] --> GRU["sequence model<br/>GRU"]
GRU --> H["hₜ<br/>deterministic memory"]
X["obs xₜ"] --> ENC["encoder<br/>CNN"]
ENC --> Q["representation q<br/>posterior zₜ | hₜ, xₜ"]
H --> Q
H -.-> P["dynamics p<br/>prior ẑₜ | hₜ · no obs"]
Q -. "KL: pull prior → posterior" .-> P
H --> S["state sₜ<br/>(hₜ, zₜ)"]
Q --> S
S --> DEC["decoder → x̂ₜ"]
S --> REW["reward → r̂ₜ"]
S --> CON["continue → ĉₜ"]
classDef dr fill:#123039,stroke:#43c0d0,stroke-width:2px,color:#dff3f6;
classDef prior fill:#123039,stroke:#43c0d0,stroke-dasharray:5 4,color:#dff3f6;
classDef neut fill:#1b2130,stroke:#8b97ac,stroke-width:1.5px,color:#e9edf6;
class GRU,H,Q,S dr; class P prior; class prev,X,ENC,DEC,REW,CON neut;
This is Dreamer dreaming of trees and motorbikes, learning how to ride in just 12,000 iterations.
survived the full episodedied chasing killsphase marker
And this is Dreamer playing the actual game, having learnt to ride safely and shoot things with 300k steps played in the imagined worlds1
I think we can say it’s mostly solved at this point…but if you don’t believe me, then here’s five minutes of it playing without losing a life.
Of course, now that I had Deathchase sorted, I started thinking about other games. We’ve got another post coming that stays firmly in the 16K era, but is a touch more complicated…and another classic. We’ll be going to Ashby-de-la-Zouch…
Although I will admit that the model is definitely cheesing things by not hitting the full acceleration. ↩︎
DeathchaseUsing 120Gb of VRAM to play a 16K ZX Spectrum gameWe bought it for your homework
This episode of Horizon from 1978 is an interesting historical look at microprocessors and how they were starting to filter through into the world at large. As a piece of its time, it’s an interesting documentary all by itself, but I was intrigued/amused at the back half, which was all about how white and blue-collar jobs will be eliminated by our 6502 overlords, showing examples of a doctor creating an expert system that would replace consultants, and a building contractor who only had to tell a system how to do something once and then a robot would be able to perfectly replicate the process. Again, 1978. It’s interesting how we now seem to be in the same position, only this time…maybe it’s more real? Or are we getting ahead of ourselves once more, and the advent of Claude/etc. just means a new normal, like how coding in C or Pascal is often much faster than trying to write assembler.
I find I can do more; this weekend, I trained a new model for work, yes, but I did two other things that I could have done myself, but it would have taken weeks and weeks of effort. Firstly, I finally started setting up the smart solar bird feeder I bought during a recent sale. But it wasn’t the brand I somehow assumed it was, and the only way to communicate with it is to use an iOS or Android app, and they want an extra $20/year for sharing the account. I sat down with Claude just after 13:00 on Saturday, and by 15:00, we had mapped out the entire API, pulled out a long-lived JWT token that allows me to authenticate whenever and where-ever, and built a small web server that produces a child-friendly set of HTML cards ready to show Maeryn just what birds have visited today, and hooking up the audio to BirdNET so it’s strictly speaking better AI than what the company is offering through the app alone.
Then, for a laugh on Sunday afternoon, I pointed Claude at the source code for Grid Wars. Every so often, I remember how fun that game was and then I rediscover that the Intel builds currently crash on MacOS. I told Claude to port it to a Rust WASM runtime that could be played in the browser. And damn if it didn’t just go ahead and do it. It required a little hand-holding to get the blur and gravity mesh effects just right…but I can now just play Grid Wars by serving up just a few files on a website1. It is very strange having all these horizons just open up in front of you. And it is something of a siren call; just one more idea that you can run before bedtime, just one more run and you’ll be fixed. I feel like I have stepped away from the edge a couple of times already, but every Friday night, the AI abyss comes around again…
(Next week? Back to the 1980s)
I would put it up publicly, but I seem to recall that the owners of Geometry Wars did get a little affronted back in the day, and I don’t really want to get into trouble. But…if you email me wink wink↩︎