An interpretability project/birds and one small AI

Murmuration

A murmuration is a flock of thousands of starlings turning together like one animal, with none of them in charge. I wanted to see whether a small AI could learn to do it just by watching.

This flock is an illustration with 12,000 simulated birds. Its big shape is scripted, and each bird follows the three rules from part 1 with its seven nearest neighbours. The falcon goes wherever you click in the sky.

heading together0.00
my simulator

The ringed bird's notes

Part 1/The simulator

How little does it take to make a flock?

I started by writing a small simulator. There is no leader in it. Each bird only watches the birds near it and follows three rules, and each rule has a switch.

Switch one off and you can see what it was doing for the group. I've already switched line up off, which is why this flock is a mess. Try switching it back on.

None of those rules says anything about a flock. A flock only shows up when every bird follows all three at once, which is, pretty cool and spooky.

Part 2/The machine

Can a machine learn this just by watching?

Next I recorded my simulator flying flocks of twelve birds and trained a small AI model on the recordings. It is a transformer, the kind of model behind chatbots, only tiny. I'll call it the machine. It never saw my three rules.

All it gets is : where each bird is and which way it is heading. The right half of the sky shows the same moment that way. Its only job is to guess where each bird will be a moment later, and each ring is one guess.

How many birds should it fly?

It only ever saw flocks of twelve. Given fifty, a size it had never seen, it still flies them as one flock, about as tightly as my simulator does.

That made me curious. Did it work out my three rules, or did it find some other way to do the job? To find out, I had to look inside it.

Note: from here on, the flocks in the frame replay what I measured, using the simulator's rules. The trained model itself is not running in your browser. What is real on this page

Part 3/Inside the machine

What is going on inside it?

Here is the short version of how the machine works. For every bird it keeps a short list of numbers, which I'll call that bird's notes. The technical name is its hidden state. The pad in the corner of the sky shows the ringed bird's notes, one page for each round of listening, and you can click another bird to follow that one instead.

Before each move there are two rounds of listening, called attention. Press a step to hold it.

: every bird hears where the other eleven are and which way they are heading, and writes that on its first page.

: it listens again, but now what it hears is the others' first pages. Each of those already holds what that bird heard, so in round two a bird is hearing what the others already heard. It writes that on its second page.

, using only what is on its two pages.

Part 4/The reader

Are my three rules somewhere in those notes?

To check, I built a second, much simpler tool, the kind researchers call a linear probe. On this page it is the reader, the teal bird with the glasses. It never sees the sky. It gets one page of a bird's notes and has to guess what each of my three rules is doing to that bird.

Here it is on , the notes after round one. Every time that page is rewritten, the reader throws three teal rings at the bird, one for each rule, aiming for the tip of that rule's arrow. A ring that lands on the tip is a good guess. The numbers in the corner score how close its guesses are for each rule.

All three rules can be read after a single round of listening. That felt almost too easy, so I kept checking.

On , the notes after round two, line up and stay close are easier to read and keep your distance is slightly harder. All three were already readable after round one, though.

Part 5/A machine that never practised

Does finding a rule mean the machine learned it?

To test that, I gave the reader a machine that has never practised. It still has its random starting numbers and has never watched a single flock.

Here it is. It can't fly at all, and it does nothing the rules ask for, which is why its arrows have gone faint.

Now compare the two columns of scores in the sky. The reader still finds stay close in these notes, and in this run it scores higher here than on the machine that can fly. That makes sense once you remember that every bird is told where the others are, which is most of what stay close needs.

So being able to read a rule in the notes doesn't prove the machine learned it. I needed a different kind of test.

Part 6/Breaking it on purpose

Does the machine use what the reader found?

The other kind of test is to break something and watch what happens, which is called an ablation. I can make one round of listening deaf: every bird hears only itself during that round, and whatever it had already written stays in its notes.

All three rules were readable after round one. If the machine flies from those notes, then losing round two shouldn't matter much.

Your guess first

Which round can the machine lose and still fly?

Now try both.

So I , and the flock fell apart. It looks like my simulator with line up switched off.

The first page of notes hasn't changed, and the reader can still read all three rules there. The machine just wasn't flying from them.

Then I instead. The flock gets sloppier, but it stays a flock, and it keeps half or more of every rule. The round where I could read everything is the one the machine can best do without.

Part 7/Watching it learn

When did it start using what was in its notes?

I saved copies of the machine all through its practice (checkpoints from its training), so I can replay how it learned. Here I follow one rule, stay close, in the notes after round two, and the buttons under this paragraph switch to the other two. Keep scrolling to move through practice. The bottom axis is a log scale, so each mark is ten times the one before.

The reader finds stay close at 85% before any practice. The machine starts out using it at 6% and builds up from there: above half from step 431 and above 80% from step 1,117, both inside the first of its 30 passes through the recordings. It ends at 99%.

Which rule should we follow?
the reader finds it0% the machine uses it0%

Part 8/Four runs

Would I get the same machine if I trained it again?

I trained the same machine four times. The only difference between the runs is the random numbers each one started from, called the seed. All four guess the next move about equally well, so from the outside you can't tell them apart. Each flock has three bars beside it, one for each rule. A full bar means the machine's answer still holds all of that rule's push.

Make one round deaf in all four
Every number, run by run

Then I in all four. The same cut does different damage each time. Run 1 keeps about half or more of every rule. Run 4 keeps less than a third of any of them. Runs 2 and 3 each lose their own mix. All four still mostly fly the same way, just less neatly than before.

With , none of the four flies as a flock any more. Each one ends up heading about as randomly as my simulator does with line up switched off. Watch the blue bars: line up drains to almost nothing in all four at once. That happens even though the reader could find line up after round one in all four runs (87%, 86%, 78% and 89%). The other two bars don't move together. Runs 1 to 3 lose nearly all of keep your distance and stay close as well. Run 4 keeps about 63% of both. Line up is the one rule every run loses.

The end/three questions, three answers

The short version

Question 1

Can a machine that only watches learn to fly like a flock?

Yes.

It guesses the next move almost perfectly and flies 50 birds it was never shown.

Question 2

Can we find the rules inside it?

Yes, easily.

Too easily: some are readable before it has practised at all.

Question 3

Is that where it uses them?

Not always.

Line up is readable after round one, yet deafen round two and it is gone, in every run. The other two rules are split between the rounds differently each time.

How I built it, and what went wrong firstThe code, on GitHub

Under the hood/for the technical reader

What I actually built

12
birds per flock
2
attention layers
414,602
weights
4
training runs
0.99
R² on unseen clips

The setup

Simulator
Boids with the three classic rules: separation (keep your distance), alignment (line up) and cohesion (stay close). There are 12 birds in a 100 × 100 world that wraps at its edges. Rule radii are set to 20, 30, and 40 units, and I clip bird speed to stay between 1 and 5. Each rule's force is capped at 0.3, and a neighbour's weight fades to zero over the last 25% of each radius. At every tick the simulator records each rule's push on each bird.
Data
1,000 recorded clips of 500 ticks each. The model learns from clips 0–699 and is scored on clips 700–999.
Model
I used a small transformer: two attention layers, one token per bird. It has width 128, 4 heads and 414,602 weights. A bird's token holds only its own last move. When one bird attends to another, the layer also gets the gap between them: how far across, how far up, and the distance, taking the shorter way around the wrapping edge. That gap affects both how much a bird listens and what it hears. The model never sees absolute positions. Each training window is turned by a random multiple of 90 degrees and mirrored half the time.
Task
It's predicting every bird's next change in velocity. On the 300 held-out clips, R² is 0.990 to 0.994 across the four runs (1 = perfect, 0 = guessing the average).
Training
40,950 steps of 256 windows each, which is 30 passes over 349,300 windows. There are four runs with seeds 0 to 3, which this page calls runs 1 to 4. Each run saved 107 checkpoints, evenly spaced on a log scale of the step count. Part 7 uses run 1's checkpoints.
Fifty birds
The model only trained on twelve birds, and I never scored its guesses on fifty. I flew it instead, from the same starts as the simulator. After 400 ticks of free flight, heading together (1 = every bird flying the same way) is 0.95 to 0.99 across the four runs, against 0.99 for the simulator. With twelve birds it is 0.96 to 0.97, against 0.97. That is 5 starts with fifty birds and 10 with twelve.

How I measured

The reader
A ridge regression (penalty 1) from one bird's hidden state to one rule's recorded push on that bird. It reads the state either after the first block or after the second. A block is one attention layer plus a small feed-forward layer for each bird. The reader is fitted on clips 0–49 and scored on clips 700–719. "Reads 96%" means R² 0.96 on those clips.
Never practised
Before any training, the model already reads stay close at 85% after the first block in all four runs. That beats the trained model in runs 1 and 3 (77% and 67%) but not in runs 2 and 4 (98% and 99%). The "scores higher" line in part 5 is about run 1.
Deafening
In one attention layer, every bird attends only to itself, and the rest of the model runs as normal. I then use least squares to fit the deafened model's answer as a mix of the three rules' recorded pushes, all three at once. The weight on each rule is how much of that rule is left. The fit leaves out birds held at a speed limit, which keeps around 74% of birds. The page clamps the share to 0–100%. With round two deaf, keep your distance goes below zero in runs 1 to 3 (−11%, −26%, −11%), which means it now pushes the wrong way. Those show as 0%. Run 3's stay close with round one deaf is 121% and shows as 100%.
How far to trust it
The three rules clearly explain 90–97% of what changes when round two is deaf, but only 68–88% when round one is deaf, so the round-one numbers are rougher. The bars show the last checkpoint. After the normal model stops improving, the deaf numbers keep moving. In runs 1 and 4 they swing by 0.09 to 0.45. I re-measured those two runs on three more sets of 20 clips, and all the sets move together, so the swings are real and don't come from which clips I picked. Over the last 17 checkpoints (from step 10,624), two things hold. With round two deaf, line up never keeps more than 18% in any run. Run 4 keeps 20–65% of keep your distance and 39–70% of stay close, while runs 1 to 3 never keep more than 20% of either. Four runs can't tell me whether a fifth would split the work like run 4 or like the other three.

What didn't work first

  • A scaling bug. My first per-bird model scaled each bird slot on its own, so two birds on the same spot looked 3.7 units apart (median over slot pairs). Separation only reaches 20 units, so that mattered. Trained to predict the separation push alone, it scored 0.48. With one scale shared by all birds it scored 0.59 (3 seeds each).
  • Absolute positions. With the scaling fixed, a model that took each bird's absolute position reached 0.72 on next moves (one seed). To know who was near whom, it had to subtract coordinates itself, and near the wrapping edge two neighbours can be nearly 100 units apart in raw coordinates. It never learned to push close birds apart. Where the simulator moves a bird away from a very close neighbour, this model moved it slightly closer. Giving attention the gap between two birds, together with the turns and mirroring, took the score to 0.949 (3 seeds). Its position error after 40 ticks of free flight fell from 1.78 units to 0.56. The turns and mirroring alone only got the old model to 0.83 (one seed). All of this was on the old physics, so you can't compare 0.949 with the final 0.99.
  • Calm flight. Once a flock settles, each of the three pushes is about 3 times bigger than the move it produces, and they nearly cancel out. Trained on calm flight alone, the model scored 0.69 on its training clips and −0.17 on new ones, so it had memorised. Even the simulator's own rules predicted calm flight badly once I put their radii 1% off (R² 0.26), because of the hard cut-off at each radius. With soft edges and the cap, the same test gives 0.92. I stopped trying to fit calm flight on the old physics.
  • One rule doing the work. In the first simulator, alignment did nearly all the work while a flock was forming. Leaving alignment out of a fit of the recorded pushes cost 0.86 of R². Leaving out separation or cohesion cost 0.02 each. Capping each rule's force and fading the radius edges changed those three numbers to 0.67, 0.11 and 0.10. Alignment is still the biggest push, about 66% of the total. Everything on this page uses the new physics.
  • The first deafening. My first way to deafen a round replaced its listening with an average note. For round two, the three rules explained only 4–17% of what changed. That listening also carries each bird's own speed-keeping, because a bird hears itself too. Letting each bird hear only itself kept the speed-keeping in place, and the rules then explained 90–97%. That is the version on this page.
  • Almost none of it was the model’s fault. It was my setup.

The plots behind parts 6 to 8

These are straight from my analysis code. The titles say "seed" where this page says "run": seed N is run N+1, so seed 0 is run 1 and seed 3 is run 4. Each has one panel per rule, with training step on a log scale along the bottom.

Run 1: for each of the three rules, how readable it is in the notes after each round and how much the machine uses it, across training.
Run 1 (seed 0), readable against used. The black line is how much of the rule the machine uses. The solid coloured lines are how much is lost when each round is deaf, and the dashed lines are the reader's score on the notes after each round.
Run 1: how much of each rule is left across training, normally, with round one deaf and with round two deaf.
Run 1 (seed 0), how much of each rule is left when a round is deaf, across training. The bars in part 8 are the right-hand end of these lines.
Run 2: how much of each rule is left across training, normally, with round one deaf and with round two deaf.
Run 2 (seed 1), the same.
Run 3: how much of each rule is left across training, normally, with round one deaf and with round two deaf.
Run 3 (seed 2), the same. Stay close with round one deaf ends at 1.21, which the bars show as 100%.
Run 4: how much of each rule is left across training, normally, with round one deaf and with round two deaf.
Run 4 (seed 3), the same.

What is real on this page

Measured
Every "reads" and "keeps" figure, the never-practised scores, the two curves in part 7, the model's scores, and in part 8 the bars and the heading-together number on each run.
Re-enacted
Every flock in the framed sky. They fly the simulator's three rules, with each rule turned down by the loss I measured, so the trained model is not running here. The guess rings in part 2 are drawn where that stand-in flock really goes next, which is why they always land, and they are only shown for twelve birds because that is where the guesses were scored. The one exception is how tidily the flock flies partway through practice in part 7: that is a smooth ramp I drew, from scattered at the start to tight at the end. Only the two curves there are measured.
Illustrated
The two big flocks at the top and bottom of the page. Their overall shape is scripted, and each bird follows the three rules with its seven nearest neighbours. Also the listening lines and passed pages in part 3 (they show who hears whom, not the real attention weights), the squares in the notes, and where each of the reader's rings lands (how far it misses by is drawn from the real score, so a rule it reads well lands on the tip).