FlyOS — a Fly Connectome, Five Retractions, and the One Claim That Survived
A fruit fly's central nervous system has been mapped completely: 165,122 neurons and 25,563,197 connections, every synapse counted. Nobody designed that diagram — evolution did — and the obvious question is whether the specific way those connections are arranged is doing computational work, or whether any network of the same size and density would behave the same. That second half is the hard part. "The wiring matters" is the kind of claim that sounds self-evidently true and is very easy to get wrong, so the project was built around being able to answer it honestly rather than impressively.
Why Bother Asking
The dataset is real, published, and awkwardly large in a specific way: the weights table has 151,856,684 rows, and only 25,563,197 of them survive filtering to neurons whose reconstruction has actually been proofread — 16.8%. The rest is segmentation debris. FlyOS converts the survivors into a compact binary format (198 MB) with byte-identical reproducibility: two independent implementations, in different languages, agree on every neuron, edge, and synapse count.
On top of that sits a simulator where the same input produces the same output every time, down to a state hash. That is not a nice-to-have. Without it, every comparison below is meaningless, because you cannot tell a structural difference from a random one.
How to Ask It Without Fooling Yourself
The honest comparison is not "real brain versus random noise" — that is a strawman that will always be won. It is random rewrites of the real brain that preserve different amounts of its structure:
| Control | What it preserves |
|---|---|
| Uniform random | only the neuron and connection counts |
| Degree-matched | + every neuron's exact in/out degree and every connection's weight |
| Sign-preserving | + how much excitation vs inhibition flows between cell classes (9 numbers) |
| Class-preserving | + the full cell-type-to-cell-type connection matrix (330 numbers) |
| The real brain | all of the above, plus everything else |
Each rung keeps a bit more of the truth, so instead of one dubious comparison you get a ladder — and you can attribute a behaviour to the kind of structure responsible. Five replicates per family, all byte-reproducible from seeds.
What the Ladder Found
Ignition threshold rises monotonically as structure is destroyed:
| Graph | Threshold |
|---|---|
| Real brain | 0.0064 |
| Degree-matched | 0.0110 |
| Structure-free | 0.0173 |
And the response sharpens at the same time. The real brain behaves like a dimmer — raise the input and activity climbs gradually. Every randomised rewrite behaves like a switch: dead, then abruptly 97% active. The ordering matters more than any single comparison, and it only exists because there are five rungs rather than one.
Testing whether this was all just excitatory/inhibitory balance eliminated that explanation — the sign-preserving control has the same threshold as the degree-matched one. What it revealed instead is that the mechanism splits into independent axes: threshold is set by the coarse cell-type layout, the runaway ceiling by E/I balance, and oscillation strength by fine wiring.
The Result That Stood Out
Feed the network a 20-millisecond pattern of activity, then let it run free. 400 ms later you can still tell which of eight patterns it was — 80 to 95% accurate. Random rewrites of the same brain, at matched activity: 7.5 to 25%.
The confound we had to rule out first was mundane: with only 200 stimulated neurons against a 5,000-neuron readout, chance overlap is 6 neurons — 0.121% of dimensions. That cannot produce 80% against a 12.5% ceiling. The decoded state is propagation, not the stimulus.
Five Times We Were Wrong
This is the part worth reading carefully, because it is the difference between this and a demo that always works.
One. We reported that the network's ~108 ms rhythm "requires biological wiring". Then we noticed every comparison had used the same absolute input level, while the graphs ignite at different levels — so the controls were below their own threshold, silent by construction. Compared fairly, the rhythm is near-threshold behaviour every graph shows. Retracted.
Two. We reported the real brain's threshold as 2.0× lower than a degree-matched rewrite. A finer measurement grid gave 1.71×. Corrected.
Three. We reported that only the biological network could encode a stimulus at all. That was our decoder being too weak — with common-mode removal and a ridge classifier, the controls reach 100% immediate accuracy. What survives is narrower and more interesting: everyone encodes; only the real brain still has it 400 ms later. Retracted.
Four. We hypothesised the memory came from a property of the connection matrix, the standard result in recurrent network theory. It doesn't. The prediction missed by 25×, and the ordering inverted: graphs with virtually identical spectral fingerprints — the four degree-preserving families sit within 0.7% of each other — differ 5× in how well they remember. Hypothesis dead.
The pattern is consistent and worth naming: we kept over-claiming in the direction of "the biology is special", and only better instrumentation caught it. Not once did a result get stronger under scrutiny.
The Test the Survivor Passed
The memory result had an obvious vulnerability. Every neuron had been given identical parameters, and identical units synchronise artificially — so the memory could have been an artefact of that uniformity rather than of the wiring.
So we jittered every neuron's decay and threshold by ±10%, ±20% and ±30% — five runs, each measured at three activity offsets, so fifteen measurements. (Three independent random draws exist only at ±20%.) If it was an artefact, it should collapse.
| Condition | ~1.0% activity | ~1.7% activity | ~2.3% activity |
|---|---|---|---|
| No jitter (original) | — | 80.0% | 95.0% |
| ±10%, seed 42 | 87.5% | 92.5% | 92.5% |
| ±20%, seed 42 | 80.0% | 90.0% | 90.0% |
| ±20%, seed 7 | 82.5% | 90.0% | 100.0% |
| ±20%, seed 13 | 75.0% | 87.5% | 85.0% |
| ±30%, seed 42 | 67.5% | 95.0% | 95.0% |
It survived everything. Then we attacked the other idealisation the project's own assumptions flagged as most distorting: conduction delay. Real neurons take time to signal; in this model, signals arrived instantly, which the assumptions recorded as overestimating synchrony.
| Delay | Low activity | Mid activity | High activity |
|---|---|---|---|
| 1 ms (original) | 95.0% | 80.0% | 95.0% |
| 3 ms | 35.0% | 62.5% | 70.0% |
| 10 ms | 60.0% | 67.5% | 77.5% |
| 30 ms | 100.0% | 100.0% | 100.0% |
Retention survives a 30× range of the assumption most likely to destroy it, and the contrast holds: at 10 ms, the real brain retains 60–77.5% where sign-preserving, class-preserving and degree-matched rewrites reach 7.5–20%, 10–15% and 12.5–15%.
The most interesting thing in that table is a dissociation. Delay makes encoding trivial for every graph — all of them hit 100% at immediate readout — while conferring none of the persistence. What biological wiring does is demonstrably not what delay does.
What the Memory Actually Is
Knowing that the wiring remembers is not knowing how, and a property you can't explain isn't one you can reuse. So we asked the non-statistical question: what does the retained state physically consist of?
Four candidates, each with a distinct signature, fixed before looking: a self-sustaining attractor, a circulating loop, a wave still travelling outward, or a passive decaying echo.
The first answer was a limit cycle with a 160 ms period. The pattern-specific state appeared to recur at exact multiples of 160 ms — 160, 320, 480, 640 — at correlation r = 0.742, while intermediate lags sat near zero. The story was that the input selects which phase the network occupies, and an attracting cycle holds that phase, so the selection persists instead of fading.
That was a satisfying answer. It was also wrong, and the way it was wrong is the most useful thing in this post.
The retraction
We only found it because we varied the measurement instead of the network. Every test up to that point changed the graph and left the analysis fixed at one configuration: 20 ms bins, an 800 ms window. The recurrence analysis reported its peak lag in bins, and 8 bins × 20 ms happens to be exactly 160 ms.
So we changed the bin width and nothing else — same graph, same operating point:
| bin width | peak lag | in milliseconds |
|---|---|---|
| 5 ms | 24 bins | 120 ms |
| 10 ms | 24 bins | 240 ms |
| 20 ms | 8 bins | 160 ms |
| 40 ms | 8 bins | 320 ms |
A real period is fixed in milliseconds no matter how you bin it. This one isn't. At 5 ms bins the correlation at 160 ms is +0.272; at 40 ms bins it is −0.039, which is nothing at all. The "comb at exactly 160/320/480/640 ms" was a comb of multiples of the peak lag in bins, which at 20 ms binning lands on multiples of 160 ms. The limit cycle was an artifact of the analysis grid.
The tell was there and we talked ourselves out of it twice. The "period" had already survived a 10× change in the membrane time constant, a 4× change in conduction delay, an 8× change in the refractory period, and an 8× change in network size — without moving. Nothing physical is invariant to four independent parameters. A null against a well-established law (delay sets oscillation frequency) is a reason to suspect your instrument, not the law.
What the correction leaves
The contrast was never the artifact — only the period was. Measured at three bin widths with activity matched (biological 2.06%, block-301 2.70%, sign-401 2.44%):
| bin width | biological | block-301 | sign-401 |
|---|---|---|---|
| 5 ms | +0.564 | +0.025 | +0.097 |
| 20 ms | +0.746 | +0.011 | +0.136 |
| 40 ms | +0.796 | +0.005 | +0.140 |
Only the location of the peak follows the bin grid. The gap over the matched controls — 5× to 159× — holds at every bin width.
So the honest statement is weaker than a limit cycle and better supported: the biological graph's pattern-specific state stays correlated with itself over hundreds of milliseconds, and the randomisations decorrelate. That is persistence with structure, not a clock.
It is also a better explanation, for a specific reason. "The input selects which phase of a cycle" was a separate claim from the retention measurement, so it could be wrong on its own. "The state retains its pattern-specific structure" is the same fact the decoder is measuring directly. The mechanism no longer sits on top of the result — it is the result.
One more thing worth knowing: the code is distributed, not a labelled line. Five neurons decode at 52.5%, twenty-five at 57.5%, a hundred at 80%, and five hundred saturate at 82.5%. The neurons doing the work sit about one hop from the stimulus — the direct targets — rather than spread across the network.
Which half of that is the wiring, and which is the weights?
"Maintains coherence" is a description, not yet an explanation. Two things could be responsible: the wiring (which neuron talks to which) or the weights (how many synapses on each connection). Every randomised control we'd built so far moved both together, so none of them could tell them apart.
So we built one that moves only the weights — permuting the synapse counts while leaving the wiring byte-for-byte identical — and, separately, rewired edges gradually in a ladder.
The result splits the mechanism cleanly in two.
The wiring is surprisingly hard to break. Rewiring 4% of edges costs essentially nothing. Rewiring 18% — moving 4.5 million connection endpoints — still costs essentially nothing: decoding 85%. Then between 18% and 63% it falls off a cliff to 25%. The substrate that carries the memory doesn't degrade gracefully, it fragments.
The weights decide how much information rides on it. With the wiring held byte-identical and 78% of the synapse counts reassigned, decoding halves, 87.5% down to 40%.
That is the part that corrected our own story: the same wiring, carrying less than half the information. Whatever the retained state is made of, its content depends on how the synapse counts are arranged, not only on which neurons are connected.
the wiring sets how much can be retained · the weights set how much actually is
Caveat, stated plainly: our original version of this section claimed the recurrence signature "survives completely intact, same 160 ms period" under the weight permutation. That half is void now — it was measured with the affected metric, at one bin width. The information half (87.5% → 40%) is a decoding measurement and stands. And the re-run is now done: at three bin widths, with both graphs at steady state, the recurrence signature survives a topology-identical weight permutation (0.412 / 0.836 / 0.917 against the biological 0.565 / 0.750 / 0.804) while retention falls from a mean of 78.3% to 44.2%. So the dissociation holds — as a statement about the recurrence signature, not about a cycle.
Did the result survive being tested properly?
Everything above rested on one control graph per condition. That is not enough to claim anything — with n = 1 the minimum achievable p-value is 0.5, and the field's standard is around a thousand replicates. So we ran it properly.
146 independent randomisations, all at a matched operating point:
| connectome | 77.5% |
| randomisations (n = 146) | 13.7% mean, range 0–30% |
| how many reached the connectome | 0 |
| closest one | 48 points short |
| p | 0.0068 |
Three independent runs agreed to within a point — 12.9%, 13.5%, 13.7%. The effect is real, and it is large. Chance is 12.5%, so the randomisations are sitting at chance while the real wiring is six times above it.
Then we went looking for what does it structurally
This is where the story turns. We had four candidates for a structural property that explains the difference, and all four died:
| candidate | what we measured | result |
|---|---|---|
| largest strongly-connected component | flat at 98.9553–98.9662% — not just across randomisations, but in graphs with no dynamics at all | dead |
| excitatory/inhibitory balance | flat to four decimal places | dead |
| reciprocity | held exactly constant while the behaviour collapsed by 98% | dead |
| assortativity | moved 10% while behaviour fell 98% | dead |
Four structural statistics, all present in the randomisations too. None of them could be what makes the difference — and that was starting to look like the real answer: maybe no structural statistic can see it.
Then one did. The spectral radius of the weighted connectivity separates the connectome from every single randomisation by 2.3 to 3.1×:
| spectral radius | retention | |
|---|---|---|
| connectome | 6031.8 | 77.5% |
| rewired 4% | 5750.7 | 85.0% |
| rewired 18% | 4912.2 | 85.0% |
| rewired 63% | 2591.1 | 25.0% |
| rewired 98% | 1973.5 | 15.0% |
| weights shuffled, wiring untouched | 1934.2 | 44.0% |
That is the first structural quantity in this whole project that detects the difference at all.
And its mechanism turned out to be measurable rather than mysterious: the connectome aims its heavy synapses at well-connected targets. The correlation between an edge's weight and its target's degree is +0.063 in the real brain, decays steadily as you rewire it, and is −0.0003 once you shuffle the weights. That single correlation tracks the spectral radius almost perfectly.
And then it failed — in the most useful possible way
Two of those graphs have essentially the same spectral radius: 1934.2 and 1926.0 — a difference of 0.4%.
They retain at 44% and 12%.
A 3.7× difference in behaviour between two networks with the same structural number. So the spectral radius cannot be what determines the behaviour, no matter how well it correlates.
That failure is the finding:
The connectome has two independent structural properties. The spectral radius follows the weight arrangement. Retention follows the topology. They can be destroyed separately, and when you do, the two measurements move independently.
Rewiring breaks both — which is why the spectral radius looked like the answer for a while. Shuffling the weights with the wiring held byte-identical breaks only the first, and retargets the question: the weights control how much information rides on the structure, and the structure decides whether it survives.
One number that was wrong for a good reason
We also wrote down the wrong mechanism first. Having measured the spectral radius, we explained it as a "degree-to-weight correlation" — which sounds right and is false.
Degree-preserving rewiring swaps whole edges, so every neuron keeps both its number of connections and its total synaptic weight. That correlation is exactly invariant by construction — we measured 0.7899 for seven different graphs, identical to four decimal places, while the spectral radius moved threefold. A quantity that never changes cannot explain one that does.
The correction took ten minutes to test and is now in the record. It is the third time in this project we caught ourselves stating a mechanism we had not actually measured, and we have written the fix into the process rather than promising to be more careful: a mechanism claim must name the quantity it asserts, and that quantity must appear in a table. The first version did not. The corrected one does.
What We Learned
What worked: the method. Pre-registering the decision rule meant the answer was decided before the data arrived. The control ladder localised each effect to a specific kind of structure instead of producing one uninterpretable null. Every permutation null sat at chance (11–15% against a 12.5% floor) throughout, which is what makes the low control numbers trustworthy rather than suspect. And running each claim through a deliberate attempt to break it is what separated four retractions from a confident wrong post.
What didn't: most of our first explanations. Five claims had to be narrowed or withdrawn, and the surviving one is narrower than we originally stated: not "biology encodes and others don't", but "everyone encodes, only the biology retains" — and that retention is not a decaying trace but persistent, structured evolution of the pattern-specific state, which the controls show is the difference between remembering and not, since they are just as active. Our first explanation of that persistence — a 160 ms limit cycle — did not survive being measured two different ways.
The one thing we still can't explain: which structural feature of the wiring sustains the persistence. We now know more than we did — the spectral radius separates the connectome from every randomisation, its mechanism is weight-to-hub targeting, and it is demonstrably not the control parameter. But the quantity that actually carries the memory is still unidentified. Four aggregate statistics and one spectral quantity have been eliminated. That is the honest edge of the result.
The uncomfortable part: our errors were almost all in one direction. Better instrumentation never strengthened a claim; it only ever weakened one. That is a useful prior to carry: if a result about biological specialness feels satisfying, that is a reason to attack it harder, not to write it up.
What we would not claim: that this is a fly brain. It is an anatomical connectome plus an assumed dynamical model — no plasticity, no per-edge delays, no gap junctions, no neuromodulation. The retention result is a statement about what this model does with this wiring. Whether it survives morphology-derived delays is the open question, and uniform delays are not a substitute for them.
FlyOS is at github.com/brsbyrk/flyos — code, data, and the full experimental record, including every retraction in the commit that made it.
