fagi / an atlas of artificial life
ESEN
Play

Preregistered result / 6 Oct 2026

Inheriting beats
learning.

Daughters born with what their mother learned survive 60% of the time; those that learn the same way from scratch, 49%. What one life finds too late to save itself saves its daughters. 1536 paired lives, 224 against 56; with a deliberately damaged program.

01 / Perceive02 / Experience03 / Learn
Plate I / The organismGame illustration · not an experiment

The central question

Do we pass on knowledge,
or our mistakes too?

In Fagi we can inspect each belief, change the world and repeat the conditions of an experiment.

Field notes / Results

What we know so far.

Simulation results. Every conclusion has a scope.

Inherited learning

A repair can reach the next generation.

60% versus 49% survival when inheriting revisions from surviving mothers. Tested with a deliberately damaged program and individual lives.

See what this does and does not show ↗
Cultural transmission

Learning better is not always adapting better.

Reasons reduce harm in a steady world. At an inversion, they leave fewer survivors than verdicts; matching transmitted coverage reduces the gap.

Read the study and replication ↗
Open questions

Evolution does not mean adaptation to each habitat.

Genes change in exploratory runs, but we found neither consistent local adaptation nor new behavior structure outperforming the human design.

Read limitations and negative results ↗
12,800lineage runs across 64 conditions; 200 shared seeds
11 / 11hypotheses supported in the lab; 3 contrasts also confirmed in the study's game version
6 / 6tests supported in the coverage replication; one descriptive prediction did not hold

01Abstract

When a culture passes on reasons ("the sour one made me sick") instead of conclusions ("don't eat the red drop"), does it adapt better to a changing world, or does it get more stuck in false beliefs? We built an artificial organism whose knowledge can be inspected —every belief records where it came from— in a deterministic world, and we manipulated what is transmitted under an equal budget of items.

With the analysis frozen before running, we found that reasons teach better than verdicts while the world stays the same (≈40% fewer harmful bites than verdicts) but leave the fewest survivors among the formats tested when the world inverts: lineages that inherit reasons reach the inversion carrying more myths and leave 22–27 percentage points fewer survivors than verdicts in that generation. Passing on the evidence together with the reasons keeps almost all of the teaching advantage and reduces myths and deaths, without removing them. The effect was confirmed in an abstract lab (11 of 11 hypotheses) and in the full game with body, space and night (3 of 3, same direction and similar size). A preregistered replication that matches the budget by species coverage (6 of 6 supported) sharpens the reading: about four fifths of that trap came from reasons covering more; the difference in myths persists under this matching scheme. Neither bits nor all transmitted information were matched.

Most recent. A second preregistered study (6 October 2026), with four hypotheses, asked whether inheriting what was learned adds to learning in life. It does: with her program broken on purpose, daughters born with their surviving mother's revisions —each backed by evidence— survive 60% of the time, against 49% for daughters that learn the same way from scratch (+11 points, 224 versus 56 discordant pairs). Learning within a life helps little (+2) and selection adds some (+2.4); a 5-point cost with the intact program could not be ruled out (section 07). In the game, three colonies in different habitats over 20 years (~50 generations) evolve towards the world they share —more muscle, less brain— but after six tests we found no colony adapting to its own habitat. We closed that line and report it in full.

The same organism was used for eight other protocols —memory and sleep, concepts, diversity, social information, exploration, adaptive decision-making, caution—, for twelve batteries on its mind and body (scientific night, learned drives, action selector, stomach, evolving body) and for an attempt to have a language model rewrite its code. We report all of it, including what did not support our hypotheses.

02The question we can put to the test

Does passing on reasons instead of conclusions make a culture more adaptable, or more superstitious?

The literature pulls in two directions. In the "wheel" paradigm of Derex and colleagues (2019), passing on a theory did not speed up improvement and narrowed the learner's exploration; in 2025 they showed that transmitted theories guide learners whether they are correct or not. Roberts-Gaal, Bolic and Cushman (2025) found that when the world varies, people prefer to pass on goals and causes rather than procedures. Formal models (Beppu and Griffiths, 2009) predict that passing on beliefs beats passing on behavior. And the "hot stove" of Denrell and March (2001) explains why a learned avoidance is almost never corrected: whoever avoids something stops trying it.

The project proposes a controlled manipulation of what is transmitted under an equal budget, measuring both adaptation after a change and the persistence of false beliefs. Simulation lets us track beliefs and repeat conditions. This does not establish that the question has never been studied or transfer the results directly to animals or humans.

Two rules that guide the whole projectFagi is judged by realism, not survival: dying is fine if the cause is realistic; only artificial deaths are bugs (dying next to known water, getting stuck in a loop, sleeping while critically thirsty). Confirmatory claims require a protocol frozen before looking at the data; exploratory findings are identified as such.

03The world

A 1280 × 860 pixel map that represents about 64 × 43 cm of ground. One second of game time equals about eight minutes in the organism's life. Everything is physics: the world does not know what is "good" or "bad", only what each thing does to the body.

Fruit with hidden chemistry

Each species is a secret mix of compounds that the tongue reads as seven tastes (sweet, umami, salty, sour, astringent, bitter, spicy), and only in the mouth. A bitter species is poisonous with probability 0.7; a non-bitter one, 0.1. Each study map has two poisonous mimics of nutritious species. Old fruit rots and becomes toxic. In the game the map is different: one nectar tree per nest, and poison comes from rotten fruit and from a habitat's toxic tree.

src/chemistry.js · src/taste.js

Water, wind and smell

The pond never runs dry, but deep water traps. The wind turns slowly and each source gives off a scent plume that can only be sensed downwind and within 34 px. Rain, every 4–7 minutes of game time, leaves puddles and wipes out scents and pheromone.

src/wind.js · src/smell.js · src/rain.js

Day, night and temperature

A day lasts 180 s. The air swings around 22 °C (10 °C at dawn, 34 °C in the afternoon). Its body prefers 25 °C, tolerates 15–33 °C and dies below 4 °C. The nest is at 24 °C (give or take what its habitat shifts); shade is 5 °C cooler. At night it keeps 45% of its sight.

src/cycle.js · src/thermal.js

Seasons

A year lasts 3600 s (20 days). Winter takes up 45%: trees yield ×0.02 and the air drops 8 °C. The code allows hot years (+12 °C) and years with fruit far from the nest, pressures that reverse direction; in the game they are off: every winter is the same.

src/seasons.js

Bodily needs

Hunger (≈1250 s to death, one "week"), thirst (≈180 s, one "day"), energy, temperature, sleep pressure, salt and health: 100 points, reduced by stings (12) and poison (10), which slow it down, limit reproduction and kill at zero.

src/needs.js · src/health.js

Life, sex and generations

Four founders per nest (12 in the game). Adult at 360 s, they live 5400 s ±15%, and age at 75% of their lifespan. Eggs hatch in 240 s. Females are slower and have more energy; males the reverse. Daughters inherit genes and the innate program; in the game, also a mark of their parents' organs and the lines of conduct their mother learned with evidence. Inbreeding costs: an egg with inbreeding coefficient F hatches with probability e−1.57·F (Ralls et al., 1988).

src/lifecycle.js · src/reproduction.js · src/program/genome.js

Colonies and habitats

In the game three colonies of up to 30 live, each in a habitat dealt at random per map: cold (air and soil −3 °C), hot (+6 °C) or with a poisonous tree 150–260 px from the nest. Fruit comes every 30 s: food and winter decide how many live. An empty nest is refounded by a pair from the nearest colony at least half full. No two are born alike: speed, reserves, metabolism, thirst, insulation, senses, memory, poison tolerance and lifespan vary ±10%, half of it heritable.

src/habitats.js · src/reproduction.js · src/variation.js

Body, load and rearing

Six heritable organs (brain, gut, muscle, eyes, antennae, size), each with a benefit and a cost. Fruit has weight and some is hard: carrying costs according to muscle × size⅔. A mother lays only after bringing food home herself, picks the strongest male that is not kin, and what she has lived up to each brood marks that brood. She walks slower uphill, warily on ground she barely knows, and sprints home from rain, cold or heat.

src/morph.js · src/load.js · src/gait.js

The organism's modules (cycle, temperature, sex, sleep, habitats, individual variation, inheritance of what was learned) start switched off in the configuration: with them off, the simulation is exactly the one the preregistered studies ran, down to the last random number. The game switches them on at startup.

04The organism

Each file has a single responsibility: movement.js moves but does not decide, decision.js decides but does not draw, render.js draws but does not touch state. The full loop, from what it senses to what it passes on:

Perception sight · smell · taste Program lines in order; first one wins Action eat · drink · carry Body compared before and after Episode reward −1…+1, real Memory value + confidence Readable rules born · self · night · told · saw · inherited Night in the nest sorts the day · predicts · questions for tomorrow Colony trophallaxis · watching · mother to brood
Figure 1. Eating or drinking opens an episode with a snapshot of the body; when it closes, the comparison yields a reward computed from real changes in the body (hunger, thirst, salt, speed and other capacities) plus how much the taste was liked. The only weights set by hand are how much each taste is liked and the penalty for danger or death. Memory turns it into value and confidence; once the belief weighs enough, a readable rule is written (or revised, or withdrawn) with its origin. Dashed lines are processes that can be switched off for experiments.

A hierarchy that comes from a single directive

Everything it does comes from "survive", ordered in five tiers: survive now (drink, eat, draw on the larder, flee extreme heat or cold); endure (rest, sleep, seek shade or shelter, go home at dusk); provide (carry what it does not need to the nest, and chase what it sees); clues (follow a smell, return to a remembered place, zigzag); explore (with no needs or cues, knowing the map is the only thing that prepares for the rest). This hierarchy is not code that only a person can change: it is data the ant carries, one line per behavior, and the first line that responds wins.

Knowledge you can read

Nothing is evaluated with eval. Rules and program are printed as real JavaScript and read back with a closed grammar, so any belief can be inspected, traced to its origin and exported.

// a rule she wrote herself: red left her hungrier, not fed, and slowed her down
rule('avoid-color-red', {"on":["eat"],"when":{"all":["color:red"]},"verdict":"avoid","weight":-0.62,"pro":3,"con":0,"because":[{"sense":"hunger","v":4.2},{"sense":"speed","v":0.6}],"learnedAt":812,"tries":4,"stage":"medium"})

// one of the two caution lines she is born with in the game
conduct('leave-harmed-mostly', {"if":{"harmedMostly":true},"do":"leave"})

// a line of its program: rest before chasing if energy drops below 35
line('rest-before-pursue-energyBelow35', { "tier": "endure", "from": "rest", "over": "pursue", ... })

Social

In the nest, sisters pass rules to each other by trophallaxis when they touch (they trust them at 0.6 of the teller's confidence); watching a sister eat teaches at 0.4; the mother, if alive, teaches her young. Every copy keeps its origin. A told rule that the ant's own experience contradicts more than it confirms dies.

What the game has switched on today

  • Scientific night: every question carries a prediction, the agenda is ordered by expected progress × safety, and answers come back as verdicts.
  • Learned drives: how much food or water is worth at each level of need is learned from the relief she felt (Zhang and Berridge, 2009).
  • Caution when eating: she is born with two lines, taste the new first and leave what harmed her as often as it fed her.
  • Her program, learned and inherited: she tries rival lines, writes her own when the evidence backs them (judged by her reserves, no risky trials at night), and they pass into the egg.
  • Night mind, diurnal sleep, evolving body, habitats and colonies, described above and in section 11.

Off in the game but available in the settings: concepts, a two-stage stomach and the free-flow selector (section 09).

05How it learns, step by step

  1. Try. It is born not knowing which trees give food; it sees appearances, not species.
  2. Feel. Taste exists only in the mouth. A bad bite brings sickness and an aversion to that smell after a single trial, as in Garcia and Koelling (1966).
  3. Remember. Each experience shifts value and confidence with a Rescorla-Wagner-style rule, from short to long term (short-term confidence fades in ≈1 simulated day; long-term, over weeks).
  4. Write. If a belief crosses a threshold it becomes a rule, with hysteresis so it does not flicker.
  5. Sleep. Every night, asleep in the nest, it sorts the day and turns its doubts into questions with a prediction, answered the next day with a test bite, starting with what teaches most without putting it at risk.
  6. Group. With things that have no innate category (only color, shape and texture), it forms concepts that predict new species and withdraws them when they fail; after a surprise, it doubts everything for a while (Behrens et al., 2007). In the studies; off in the game.
  7. Reorder its program. If putting one behavior in front of another saves reserves again and again, it writes that line with its evidence; its daughters inherit it.
  8. Transmit. To sisters and to the next generation, in the format the experiment sets: nothing, verdicts, reasons, or reasons with evidence.

06Main result: reasons versus conclusions

Confirmatory  Protocol frozen at commit 9d222c3 before any run.

Design

  • Unit of analysis: the lineage (12 generations of colonies of 5). Rows from one lineage are never treated as independent.
  • Four formats with the same budget of 4 items: nothing; verdicts (per-fruit conclusions about fruit the teacher knew); reasons (rules about traits); reasons + evidence (rules and the bites that back them).
  • World change at generation 6: none, inversion (poisonous becomes nutritious and vice versa), rotation, or shift to another dimension.
  • Lab: 64 cells × 200 lineages, seeds 1001–1200, 12,800 lineages in 281 s. Formats are paired by seed within each condition. Changing the change type can also change the initial catalog; not every cell shares exactly the same world.
  • Game: 75 lineages per format, seeds 5000–5074, lifespans of 1800 s, inversion at generation 4; ≈40 min on 18 cores.
  • Analysis: paired permutation tests with 10,000 sign flips, Holm correction, 95% bootstrap intervals and effect size dz.
nothing verdicts reasons reasons + evidence
Figure 2. Means per lineage. . Reasons teach better while the world stays the same (left) but leave fewer survivors when it inverts (center), and carry more myths —false rules without personal experience— (right). These contrasts alone do not isolate causal mediation by myths.
nothing verdicts reasons reasons + evidence
Figure 3. The main cell generation by generation: one-trait chemistry, lifespans of 1800 s, inversion at generation 6, means of 200 lineages per format (seeds 1001–1200). Generated with node scripts/site-data.js from the same study code; generation 6 exactly reproduces the preregistered table. Everything runs level until the change: reasons show the largest survival drop, followed by recovery in later generations.

The eleven lab hypotheses

HypothesisContrastMeasureDifference [95% CI]dz
H1 · reasons teach better in a stable worldverdict − reasonharmful bites0.224 [0.204, 0.247]1.46
H2a · reasons leave fewer survivors at the inversionverdict − reasonsurvivors0.274 [0.219, 0.329]0.68
H2b · … fewer than transmitting nothingnothing − reasonsurvivors0.300 [0.245, 0.356]0.73
H3 · reasons carry more myths into the inversionreason − verdictmyths0.863 [0.672, 1.048]0.63
H4a · evidence carries fewer myths than reasonsreason − evidencemyths2.000 [1.809, 2.199]1.41
H4b · evidence leaves more survivors than reasonsevidence − reasonsurvivors0.150 [0.086, 0.211]0.33
H4c · evidence teaches better than verdictsverdict − evidenceharmful bites0.253 [0.234, 0.273]1.79
H5a · the H2a gap is larger under inversion than rotation(v − r) inv − rotsurvivors0.265 [0.209, 0.321]0.65
H5b · … than under dimension shift(v − r) inv − shiftsurvivors0.277 [0.220, 0.333]0.68
H6a · the H2a gap grows with 1800 s lifespans versus 900 s(v − r) 1800 − 900survivors0.206 [0.147, 0.263]0.49
H6b · the H3 difference holds with 900 s lifespansreason − verdictmyths0.771 [0.583, 0.965]0.56
One-sided, Holm over eleven. With 10,000 permutations the smallest reachable p is 0.0001, so every adjusted p is 0.0011. Measures are per ant.

Confirmation in the game

HypothesisnDifference [95% CI]dzp (Holm)lab
H1 · reasons teach better (gen. 1–3)750.332 [0.243, 0.411]0.910.00030.224
H2a · fewer survivors at the inversion (gen. 4)750.223 [0.133, 0.317]0.540.00030.274
H3 · more myths at the end of gen. 4750.827 [0.630, 1.027]0.920.00030.863
Two-sided, Holm over three. The preregistration required lab and game to agree in direction on H1, H2a and H3: they agree, with similar sizes.

What exploration shows (not preregistered)

  • The large drop occurs with the inversion studied. If the poison rotates to another value or moves to another dimension, reasons lineages also carry more myths (1.3–2.5 times), without a comparable survival drop: survival is 98–100% in every format. Only inversion makes the old reason point right at the new food, and rejection turns into starvation.
  • Conjunctive chemistry (poison = color AND smell): same pattern, smaller (80% versus 92% survivors).
  • Recovery. After any change, what is taught to newborns takes longer to correct with reasons (1.7–2.1 generations versus 1.2–1.7).
  • Short lifespans reduce deaths but not myths: the belief is just as false, there is simply no time to starve because of it.
  • After the inversion reasons lineages are again the ones that suffer least harm: the survivors relearn.

Replication with a coverage-matched budget confirmatory

Protocol frozen at a9c3265 (5 October 2026), new seeds 4001–4200, 200 lineages per cell, Holm over six. The budget is no longer counted in items but in coverage: how much of the world what is passed on covers (SOCIAL.cost "coverage"). The item mode reproduces the original study with its original seeds; with new seeds, C0 estimates a 0.232 gap, versus 0.274 in the main study.

TestDifference [95% CI]dz
C0 · H2a replicates counting items0.232 [0.178, 0.285]0.61
C1 · that gap is larger counting items than coverage0.186 [0.129, 0.242]0.45
C2 · at equal coverage, reasons carry more myths0.902 [0.722, 1.076]0.71
C3 · at equal coverage, evidence carries fewer myths than reasons2.071 [1.910, 2.231]1.76
C4 · at equal coverage, evidence teaches better than verdicts0.119 [0.101, 0.136]0.93
C5 · at equal coverage, reasons teach better than verdicts0.029 [0.013, 0.044]0.25
6 of 6 supported; every adjusted p is 0.0006. Results in docs/research/results.md and research/results/coverage-*.
items (replication C0; main H1/H3/H4a) matching coverage (replication)
Figure 4. Each bar is a difference between formats, in paired lineages. Counting items: the replication's H2a (C0) and the main study's H1, H3 and H4a; matching coverage: the preregistered descriptive and C5, C2 and C3. The survival gap and teaching edge shrink; differences in myths remain similar in size.

Reading. With the budget matched by species coverage, reasons' survival trap shrinks by 80% (from 0.232 to 0.046 [0.017, 0.078]; the preregistration predicted an interval including zero; that did not hold: a small gap remains) and their teaching edge nearly vanishes (0.224 → 0.029). What does not change are the myths: reasons still carry more false beliefs, and evidence still cuts them. The comparison indicates that coverage accounts for much of the survival gap in this design. Differences in myths persist, without establishing a universal property of the format.

Sensitivity analysis (7 Oct): in 101 of the game's 600 maps a species tree lay off the map. Dropping those lineages, H1, H2a and H3 still hold (0.304; 0.210; 0.992). This post-hoc analysis does not validate absolute survival levels or the other studies affected by species maps.

A limitation found while reading the resultsEach lineage's catalog is drawn so that every epoch has at least two poisons and two foods, so the world before the change depends on which change comes after. Comparisons between formats are not affected (they are paired within the same world); the H5a/H5b interactions are not a pure common-random-numbers contrast.

Result sourcesdocs/research/results.md · docs/research/preregistration-coverage.md

07Key result: inheriting beats learning (four hypotheses)

Confirmatory  Protocol frozen at commit 2459b1c on 6 October 2026; run that same day without changing a line or deviating.

Does a daughter born with her surviving mother's program revisions live longer than one that learns the same way, but from scratch?

First we needed a world where behavior decides. In the default world every program survives, even one broken on purpose: no learning can show. Sweeping how often trees bear fruit, survival falls off a cliff (96% for the innate program with one fruit every 450 s, 0% at 750 s). At 575 s the innate program lives through 75% of two-hour lives, and breaking its warmth, rest or food lines costs 27–37 points: there, what she does decides whether she lives.

innate program no warmth (its 4 lines moved last) no food (carry and pursue moved last) learns its program (15 s judge)
Figure 5. Survival at 2 hours by how often a tree bears fruit (24 lives per point; 48 at 550, 575 and 590), varied maps, paired by seed. What she does decides only near the edge of the cliff; at 575 s breaking one part of the program costs 27–37 points (docs/research/world-calibration.md).

In that world we moved dusk (going home as the light goes) to the end of her program: in the pilot, survival drops to 44%, and cold becomes the leading cause of death.

Design

  • Four arms: born (does not learn); learns (every life from scratch); inherits (mother drawn from the previous generation's survivors); inherits from any (dead or alive).
  • 32 populations × 16 lives × 6 generations per arm; seeds 20000–51095, never used before. Lives pair across arms by population, generation and slot.
  • How she learns: a moment is worth what it took from her reserves (not 15 s of distress); at night she only tries endure lines; the worse she is, the less she tries. An accepted revision carries its evidence and passes into the egg; it is retired only on evidence against it.
  • Analysis: survival to 7200 s in generations 3–5 (1536 paired lives per contrast), one-sided exact sign test on discordant pairs, α = 0.05. Only H1 is primary.
born learns inherits from any inherits from survivors
Figure 6. Survival by generation, 512 lives per point (docs/research/prereg-lineage-results.md). In generation 0 every learning arm starts equal; from generation 1 the inheriting arms pull away and stay away.
Hypothesispairs (better / worse)one-sided p / H4 bounddifferenceVerdict
H1 (primary) · inherit > learn224 / 56< 0.0001+10.9 ptssupported
H2 · learn > born72 / 390.0011+2.2 ptssupported
H3 · surviving mothers > any mother177 / 1400.022+2.4 ptssupported
H4 · with the program intact, learning costs no more than 5 pts12 / 14bound −0.1090.750 vs 0.771not supported
Generations 3–5. Survival: born 0.47 · learns 0.49 · inherits from any 0.58 · inherits from survivors 0.60. Cold deaths 486 / 517 / 408 / 406.
  • What one life finds too late saves its daughters. Nearly every inheriting daughter carries dusk-before-pursue from birth; in single lives only 69% write it, and late.
  • Learning within a life helps, but little (+2 points). Selection adds a little on top of inheritance (+2.4).
  • H4 was not supported: 96 lives cannot rule out that learning costs 5 points when the program is intact. So the preregistration did not switch it on in the game; it was switched on later after measuring it in the game itself (8 maps × 4 years, exploratory comparison: +0.8 ants per nest, 5 of 8 maps, t = 1.3; this neither demonstrates no cost nor resolves H4).

How we got here (exploratory)

  • Judging by 15 s of distress hurt: learning lowered survival from 0.75 to 0.58. The whole cost was the trials (leaving a fruit in sight for a remembered place), not the lines. Trying less the worse she is stopped the harm, but did not turn it into help.
  • Selection alone changed nothing (24 versus 23 pairs): every life rediscovered the same "memory before pursue" lines; there was no heritable variation to select.
  • Changing the judge did: judged by her reserves and with safe night trials she stops writing those lines and writes the real repair, dusk-before-pursue (0.44 → 0.53 within one life, 9 versus 0 pairs).
What this result does not sayThe gain is the repair of a program broken on purpose. It shows that learning and inheritance recover what was lost; not that they find something the innate program lacks. One world, daughters living alone rather than in a colony, and rewrites within the grammar (one behavior in front of another under one condition).

Result sourcesdocs/research/prereg-lineage-results.md · docs/research/prereg-lineage-inheritance.md

08Try it: the lab in your browser

This is not an animation: it is the study code running in your browser. Each button launches four lineages of 12 generations —one per format— in the same world (same chemistry and catalog, without a simulated spatial map), with the game's brain, body and social code. The seed, parameters and code version allow the experiment to be repeated. Numerical identity across browsers, architectures and Node versions is not guaranteed.

. true in that world false; thickness is how many ants carry it.

The default seed (1033) is a clear case, chosen on purpose from the study's seeds. A single lineage is noisy: with other seeds, such as 1001, reasons lose no one. The study's effect is a mean of 200 lineages per cell (Figure 3). The genealogy on the right is the "phylogeny" of the culture: each row is one specific belief, born in one ant and passed mouth to mouth and mother to daughter; after the inversion, reasons that were true become false without changing a letter.

09Other studies with the same organism

Each protocol is frozen in docs/research/. The detailed reports were removed from the tree on 5 October and remain in history (git show b0815b8^:research/results/<study>/report.md). The LIBERA batteries are in docs/research/libera/ and the language-model plan in research/code-culture/; the largest raw data stayed out of the repository (see DATOS.md in each folder).

Memory, sleep and consolidation frozen

14 conditions × 2 worlds × 120 lives. Learning matters: judgment 0.852 versus 0.500 for a Fagi that does not learn (dz 3.6); it lives 2069 s versus 344 s for a random agent. With fixed learning, 9 of 12 populations go extinct; with full learning, 2 of 12.

Follow-up (4 of 4 supported): interleaved nightly replay worsens judgment (0.928 without it versus 0.904) and consolidation slows adaptation after a change. That is why replay is off by default.

organism-protocol.md · organism2-protocol.md

Concepts 4 of 4

Faced with things that have no innate category, it correctly names new species at first sight (0.967 versus 0.25 by chance, dz 10.2), concepts save it stings (0.042 versus 1.450), it drinks new sap before examining it, and doubting after a surprise helps when the world turns (0.847 versus 0.392).

concepts-protocol.md

Diversity in the face of change 3 of 4

4 founders, 8 species, chemistry inverted halfway through the run. The population recovers (judgment 0.810 at the end versus 0.402 after the change) and keeps genetic diversity (0.189 versus 0.152). What was not supported: the clonal population did not die more from poison (0.132 versus 0.121).

diversity-protocol.md

The misinformed sister 2 of 3

A sister who tells false rules meets the preregistered 0.05 non-inferiority margin for judgment (0.800 versus 0.813); this does not mean zero harm, and only 26% of the false rules adopted are still standing at the end. But a well-informed sister did not reduce the poison eaten either.

social-protocol.md

Explore or return 2 of 6

10 conditions × 120 colonies. Sisters whose first foraging trips paid off explore more later (r = 0.455, 596 sisters) and after a full nest wait longer before foraging again (983 s versus 153 s). Learned choice did not eat more than fixed rules and did not explore more in short-lived worlds.

forage-protocol.md

Caution in the game 2 of 2

Two lines —test new species first, leave what harmed you— raised survival from 0.882 to 0.955 (+0.073 [0.033, 0.111]) with 200 worlds per profile; poison deaths fell from 29 to 4, at the cost of 4 starvation deaths that did not happen before. Today they are the default behavior.

caution-protocol.md

Is there anything to learn without sabotage? exploratory

We tried all 171 possible moves (an innate line put in front of one above it), 24 lives each, in the scarce world and in five more (warm nights, cold, long nights, hot days, rain). None beats the innate order: 102 change nothing and several are lethal. The one hint (memory>thermal on long nights) did not replicate on 48 new seeds.

No improvement over the innate order was found among the moves and worlds examined, partly because its behaviors already learn inside (she goes home at dusk only once darkness means cold to her).

world-calibration.md · scripts/order-screen.js

A world that presses, in the game exploratory

With fruit every 8 s colonies lived at 71–73% of their ceiling and died of old age. At 30 s, food and winter decide how many live and cold kills more than age; in 3 years no map dies out (in 20 years, 2 of 16 do). The autopsy is consistent with exposure and hunger deaths: all 57 in winter, 39 at night, nearly all far from the nest and 38 very hungry: foragers caught out by the cold.

Along the way a bug turned up: water and trees of the further nests fell off the map (583 of 2000 objects) and ants were dying of thirst by the edge. Fixed.

world-calibration.md · scripts/game-world.js

Scientific night 0 of 4

Four frozen batteries (H3, H3b, H3c, H7, 60 lives each) on whether ordering the night's questions by what teaches most speeds up learning: no (AUC +0.001 and +0.015). With 16 species she survives 10% more [3, 18], from putting safety first. What is worth a lot is having an agenda at all: +28% survival against none.

docs/research/libera/notes/bateria-*-resultados.md

Learned drives and stomach 3 of 4

Learning food value meets the 0.02 non-inferiority margin for safe life (difference −0.002) and, as in animals, after the first deprivation she takes longer to go and drink (+61 s [32, 92]); the second prediction did not hold. A two-stage stomach reduces bites and waste compared with no satiety; those outcomes do not differ from the instant stomach. Drives are on in the game; the stomach is not.

docs/research/libera/notes · src/drive.js · src/stomach.js

Free-flow selector null

Four batteries (H2–H2d) against the current hierarchy: ties, 2–3 wins and 2–4 losses on Tyrrell's (1993) 14 requirements. Part of the gap was a bug in the consume bonus; fixed, it ties. As in Bryson (2000), the hierarchy does not lose. The game keeps the hierarchy.

docs/research/libera/notes · src/decision/select.js

A language model rewrites its code line closed

A local LLM rewrote each ant's bite judge and selection chose. With a tournament, texts beat the current one in new worlds (+0.056 [0.019, 0.087]), but through vocabulary: "leave what harmed" shows up in 10 of 10 texts with clear names and in 1 of 20 with blind names. A rule specific to each lineage is learned by none (0 of 10), and a seeded rule is erased by rewriting (70–84%). In these tests, the model appears to contribute prior knowledge without learning the lineage-specific rule.

research/code-culture/ · src/learned/code-judge.js

10What did not work

We publish it with the same detail as what worked. A project whose end goal is for the organism to rewrite its own code has to say clearly where it stands today.

IdeaResultStatus
An adaptive decision model beats the current heuristicIt beats the current version (+0.056 [0.024, 0.089]) but loses to a simple heuristic (−0.030); planning two steps ahead adds nothing (−0.004). Verdict: prefer the heuristic.not supported
Behavior rules written by the ant itself, within one lifeTwo variants: −0.080 and −0.072 versus the current Fagi.worse
… across lineages≈ −0.23; with prohibitions "with an exit", −0.254 and −0.232.worse
Rewriting its program within a lifeWith logs from 24 lives pooled, the innate order is a local optimum. It repairs a sabotaged program only with a lot of evidence: sharing moments between sisters in the nest made it possible (8 of 8 colonies, versus 0 without sharing; development notes, no protocol). The preregistered study in section 07 takes up this question.exploratory
Night mind that proposes rules+0.060 in the lab with color-and-smell chemistry; no gain in the game.exploratory
Nightly consolidation improves judgment0.852 versus 0.838, p 0.229.not supported
Learning her program judged by 15 s of distress0.58 versus 0.75 for the innate program (9 versus 1 pairs, replicated). The cost is the trials; the lines she writes do not help.worse
Selection as the judge (only mothers that live pass their program on)No effect: 24 versus 23 pairs. Every life writes the same lines.not supported
H4 · rule out a 5-point cost with the intact program0.750 versus 0.771; 96 lives cannot rule out a 5-point cost.not supported
Each colony adapts to its habitatSix tests: reciprocal transplant, opposite habitats, 20 years, 24 maps, milder winters, refounding with parties. No local adaptation. Size genes lean bigger in the cold (+0.037, 15 of 23 maps) without holding up on the new maps.line closed
A milder winter keeps colonies from emptyingNo: 6.6% of years empty versus 8.3% and 6.9%. Colonies that empty had already dwindled to 1–5 the year before.not supported
Refounding a nest with 4 or 6 instead of a pairNo difference (4.9 / 4.9 / 3.5% of years empty).not supported
An LLM rewrites her code from her life0 of 10 lineages learn a rule specific to their world, with or without the mother's diary; rewriting erases seeded rules.line closed
Ordering the night's questions by what teaches most0 of 4 batteries: she does not learn faster.not supported
A free-flow selector beats the hierarchyTie across four batteries.null
One-trial crisis plasticityWrites few lines (≈0.1 per life) and none changes an action; survival identical with and without it. An earlier report of 10 of 10 lives does not reproduce; corrected in docs/research/one-shot-crisis-plasticity.md.no effect
Where we stand on self-rewritingToday the organism writes rules about what to eat, with confirmed success, and reorders lines of its program with a fixed grammar. With evidence, it repairs a broken program within a life and its daughters inherit the repair (+11 points, preregistered). Without sabotage, no improvement was found among the 171 reorderings and six worlds examined. This does not prove that no better behavior exists. That it can invent new behavior structure that beats the human design is not yet shown; for that the world must demand a priority nobody programmed, or the grammar must let it build more than an order.

11In the game: colonies that evolve

exploratory  Six heritable organs —brain, gut, muscle, eyes, antennae, size— each a multiplier with a benefit and an energy cost. Costs grow faster than linearly and metabolism scales with mass to the ¾ (Kleiber, 1932). A large brain costs about 18 times as much as muscle per gram (Elia, 1992) and, as in the fish of Kotrschal et al. (2013), shrinks the gut and the offspring. Within a life, each organ can adjust by up to ±25%, but only when experience moves outside the typical.

We compare three forms of inheritance: Darwinian (genes only), Baldwin (the capacity to adjust is inherited) and epigenetic (part of the mother's adjustment is inherited, like the learned avoidance that C. elegans keeps for about 4 generations; Moore et al., 2019).

  • exploratory With cold winters, cold became the cause of half or more of the deaths (before, 100% were old age); epigenetic inheritance had fewer cold deaths. With 3 seeds per cell, no difference is firm.
  • not confirmed The prediction of de Bruin et al. (2026) —inheriting what was lived pays off more when change is predictable— was not confirmed in 90 runs with 10 seeds: 4.4 [−15.8, 24.6], nor with far-fruit years (−11 [−36, 14]). Body size does track cold and hot years in all three modes (temperature-size rule, Atkinson 1994).
  • exploratory With colonies of ~30 drift beat selection. With own provisioning, mate choice by strength, maternal effects and three colonies, muscle already rises in 7 of 8 seeds in just 4 years (t 2.8).

Twenty years in the game

exploratory  The game version measured on 8 October 2026 —three colonies, habitats (cold −3 °C, hot +6 °C, poison close by), fruit every 30 s, evolving body, individual variation, inheritance of what was lived and of what was learned— over 20 years (~50 generations) on 8 maps (seeds 7500–7507, run on 8 October 2026). Genes from year 1 to year 20, all colonies of each map together:

Genechangetmaps in that direction
muscle+0.1605.68 of 8
brain−0.096−4.17 of 8
size+0.0683.58 of 8
Gut, eyes and antennae drift without direction. Carrying fruit that weighs pays for muscle; a costly brain does not pay for itself. The first 20-year run (cold −6 °C, refounding from the fullest nest) gave the same: muscle +0.185, brain −0.128, size +0.061.
cold habitat hot habitat toxic habitat
Figure 7. Mean genes of the living ants of each habitat over 8 maps, at the end of each year; 1 is the founders' value. All three habitats raise muscle and lower brain almost in step: they do not pull apart. Data: node scripts/game-world.js --interval 30 --seeds 8 --seed 7500 --years 20 --json … and scripts/site-game-data.js.

Evolution happens, the same way everywhere. It is adaptation to the world all colonies share. Conduct also passes from mother to daughter: by year 20 each ant carries 0.9–1.0 lines of her own, about 0.75 inherited, and they are the same in every habitat ("memory before pursue", "memory before scent"). In today's game the habitat that suffers is the hot one: 7 alive on average (23% of the ceiling) and empty in 36 of 160 nest-years, against 16.5 in the cold and toxic ones; 91 refoundings on 8 maps.

It is not local. We looked for each colony adapting to its habitat with a reciprocal transplant (every newborn raised in a random nest and followed until death) and with 20-year runs on 24 maps, in which 2 of 16 new maps died out. What showed up was a stranger advantage: in every habitat, young from elsewhere left more offspring than locals (−0.41, t −2.9). We explained it: a local female may not mate with her brothers, often the best males of a small colony, and a newcomer is kin to no one. Once inbreeding depression was added the rule makes sense, and the advantage persists for a reason seen in nature too (genetic rescue). No consistent local adaptation was detected. Repeated with today's game (12 maps, seeds 7400–7411, 3 years then 2 of transplant): local − foreign = −0.19 offspring (t −0.67, locals ahead on 6 of 12 maps). Neither local adaptation nor, any more, a clear stranger advantage.

born in that nest brought from another nest
Figure 8. Reciprocal transplant with today's game: every newborn is raised in her own nest or another at random and followed until death. Means per map (11 maps per habitat). In the toxic habitat those from elsewhere leave more offspring (2.67 vs 1.82); in the cold and hot ones locals are slightly ahead (2.41 vs 2.32; 2.65 vs 2.35). No consistent pattern of "those from here do better here".

Why, as far as we measured: colonies of 12–19 adults dwindle to a handful and are replaced by their neighbours (≈7% of nest-years empty and 115 refoundings on 16 maps in the 24-map run; 91 on 8 maps today, mostly in the hot habitat), erasing what had built up. Bigger colonies (a ceiling of 60) remain the untried remedy. The game's Evolution panel draws these curves live, colony by colony.

Result sourcesdocs/research/world-calibration.md · investigacion/data/game.json

12Method and rigor

Determinism and fingerprints

In research runs, random numbers come from a seed. A lineage depends only on its seed and its options: run alone or inside a batch, it comes out identical byte for byte on the same machine (rounding can differ between ARM and x64 processors or between Node major versions, so fingerprints are kept per machine). A live game is not seeded: it is recorded and replayed from its events. Seeded lives are reduced to a fingerprint of decisions; the tests compare them against fingerprints recorded before each architecture change, so a refactor cannot silently change what the ant does.

Preregistration

Hypotheses, measures, contrasts, direction, correction and success criterion are frozen in a commit before running. Anything that was not there is reported as exploratory. Execution deviations are documented (a container restarted and the runs were relaunched per lineage, with an identity check).

Statistics

In transmission studies the unit is the lineage; in other protocols it is the colony or life according to the design. The inheritance study uses paired lives grouped in populations with shared ancestry: dependence between descendants limits treating them as independent replicates. Cells with common random numbers: the same worlds in every condition, paired contrasts. Paired sign permutations (10,000), Holm for multiple tests, 95% bootstrap, dz and lineages needed for 80% power. Latin hypercube designs over the hand-tuned parameters to see whether an effect keeps its sign.

Lab versus game

The lab removes space, water, nest and larder, but keeps the learning and transmission mechanisms evaluated in that study, with the same code. It runs about 20 lineages of 12 generations per second on 4 cores. Three preregistered contrasts from the main study were checked in the game: H1, H2a and H3. The coverage replication has no equivalent game confirmation here.

13Limitations

  • Original maps and sensitivity. In 101 of the 600 game-confirmation maps a tree was outside the boundary. The three contrasts retain their direction after affected lineages are excluded; absolute levels and other studies are not validated by this check.
  • The main study matched formats by items. With budgets matched by species coverage, the survival trap shrinks by 80%; the myths stay. The 22–27 point figure should be read with that replication (Prystawski et al., 2025, show that channel rate matters).
  • A verdict only covers fruit the teacher knew. That is the nature of a conclusion, and also why it generalizes less.
  • The pilot shaped the hypotheses. Confirmatory tests use new seeds, but some hypotheses (explore or return, diversity) were written after seeing a pilot.
  • Survivorship bias. In the self-rewriting studies deaths are not yet logged, which biases what is learned toward those who survive.
  • It is a toy ant. Parameters are inspired by the literature but tuned by hand; neither the direction nor magnitude of the effects alone establishes validity in animals or human cultures.
  • Small colonies. With 12–19 adults per nest, drift and mixing through refounding outweigh weak local selection; that is why the game shows global evolution and not local.
  • Inheriting what was learned was tested by repairing a sabotage. It does not show that learning finds something the innate program lacks.
  • Studies run on earlier code. The LIBERA batteries (2–3 Oct) and the language-model plan (1–5 Oct) ran on the code of those days; the game has changed since (habitats, colonies, inheritance of what was learned), so their numbers describe that organism.
  • A single author. External review is missing; this site exists partly to get it.

14Roadmap

Done since the previous version: Tyrrell's realism metrics as a baseline; scientific night, learned drives and free-flow selector tested (the first two switched on in the game); budgets matched by species coverage; a body that evolves under real pressures; inheriting what was learned, preregistered.

  1. A world where the innate order is not optimal without sabotage (seasons or maps where the right priority changes): the next step the inheritance preregistration set.
  2. A richer grammar, so entities can build more than an order of lines: the underlying aim is that they rewrite their own behavior.
  3. Bigger colonies (a ceiling of 60) to take up local adaptation again.
  4. More lives for H4: measure whether learning costs with the program intact.
  5. Realism as the main measure in the game, with Tyrrell's 14 requirements alongside survival.

15Reproduce

Many raw rows are not versioned: research/run.js writes next to them a manifest.json with the design, the seeds and the commit, to repeat them with the same revision and runtime environment. The commands below use the checked-out code; reproducing historical figures requires each study’s recorded commit. Node ≥ 22.

# main lab study (12,800 lineages, ~5 min)
node research/run.js --design research/designs/main.json
node research/confirm.js research/results/main

# sensitivity over the hand-tuned parameters
node research/run.js --design research/designs/sensitivity.json
node research/sensitivity.js research/results/sensitivity --a verdict --b rule

# inheriting what was learned (preregistration: docs/research/prereg-lineage-inheritance.md)
for arm in born learn inherit inheritAny; do for p in $(seq 0 31); do
  node scripts/lineage-selection.js --arm $arm --pop $p --seed0 20000 --sabotage dusk \
    --set PROGRAM.judge=1 --set PROGRAM.darkTrials=1 --set PROGRAM.exploreByState=1 \
    > research/prereg-lineage/$arm-$p.json
done; done
node scripts/lineage-analysis.js research/prereg-lineage   # + H4: see the preregistration

# the game's world: colonies, habitats, years
node scripts/game-world.js --interval 30 --seeds 8 --seed 7500 --years 20 --jobs 16   # with the game's current settings

# coverage-budget-matched replication
node research/run.js --design research/designs/coverage-items.json
node research/run.js --design research/designs/coverage-coverage.json
node research/confirm-coverage.js research/results/coverage-items research/results/coverage-coverage

# all tests, including the decision fingerprints
npm test

16References

Grouped by the idea the project takes from each one. The full review is in reports/Razones frente a conclusiones culturales.md, which worked mostly from abstracts.

Transmitting theories and cumulative culture

  1. Derex, M., Bonnefon, J.-F., Boyd, R. and Mesoudi, A. (2019). Causal understanding is not necessary for the improvement of culturally evolving technology. Nature Human Behaviour, 3, 446–452. Passing on a theory did not speed up improvement and narrowed exploration.
  2. Derex, M., Bonnefon, J.-F., Boyd, R., McElreath, R. and Mesoudi, A. (2025). Proceedings of the Royal Society B, 292, 20242499. Transmitted theories guide learners, whether correct or misleading.
  3. Osiurak, F. et al. (2021, 2022). Nature Human Behaviour; Science Advances, 8, eabl7446. Partial replication: understanding grew with performance.
  4. Roberts-Gaal, X., Bolic, A. and Cushman, F. (2025). PNAS, 122(28). When the world varies, people pass on goals and causes, not procedures.
  5. Harris, J. A., Boyd, R. and Wood, B. M. (2021). Current Biology. The causal theories of Hadza archers are incomplete and culturally learned.
  6. Saral, A. S., Derex, M. and Ibáñez de Aldecoa, P. (2026). PNAS Nexus, 5, pgag090. Causal reasoning speeds up early gains and then fades.

Formal models of transmission

  1. Rogers, A. R. (1988). Does biology constrain culture? American Anthropologist. Rogers' paradox: copying does not raise mean fitness in a changing world.
  2. Enquist, M., Eriksson, K. and Ghirlanda, S. (2007). Critical social learning —copy, then check— resolves the paradox.
  3. Rendell, L. et al. (2010). Why copy others? Science, 328, 208–213. Copying wins because demonstrators filter.
  4. Kalish, M., Griffiths, T. and Lewandowsky, S. (2007). Transmission chains converge to the learner's prior biases.
  5. Kirby, S., Cornish, H. and Smith, K. (2008). PNAS, 105, 10681. Rule systems become regular and compressed as they are transmitted.
  6. Beppu, A. and Griffiths, T. (2009). Passing on beliefs beats passing on behavior.
  7. Prystawski, B., Arumugam, D. and Goodman, N. (2025). Channel rate has nonlinear effects: the basis for matching budgets.

Avoidance, the hot stove and epistemic vigilance

  1. Denrell, J. and March, J. G. (2001). Adaptation as information restriction: the hot stove effect. Organization Science, 12, 523–538. Overestimates get corrected; underestimates do not.
  2. Denrell, J. (2007). Psychological Review, 114, 177–187. Adaptive sampling and systematic biases.
  3. Denrell, J. and Le Mens, G. (2017). Management Science, 63, 528–547. Sampling what is popular creates lasting collective illusions.
  4. Toyokawa, W. and Gaissmaier, W. (2022). eLife, 11, e75308. Conformity can rescue a group from the hot stove trap.
  5. Laland, K. N. and Williams, K. (1998). Behavioral Ecology, 9, 493–499. A maladaptive tradition in guppies outlived its founders.
  6. Galef, B. G. (1985). Rats learn socially what to eat, not what to avoid.
  7. Sperber, D. et al. (2010). Epistemic vigilance. Mind & Language, 25, 359–393.
  8. O'Connor, C. and Weatherall, J. O. (2018). European Journal for Philosophy of Science, 8, 855–875. Weighting by trust can produce stable false polarization.
  9. Doyle, J. (1979). A truth maintenance system. Beliefs with recorded justifications: the idea behind origin.

Learning, memory and sleep

  1. Garcia, J. and Koelling, R. A. (1966). Sickness is attributed to taste more than to appearance; aversion after one trial.
  2. Rescorla, R. A. and Wagner, A. R. (1972). The value-learning rule.
  3. Kalat, J. W. and Rozin, P. (1973). Learned safety.
  4. McClelland, J. L., McNaughton, B. L. and O'Reilly, R. C. (1995). Interleaved nightly replay against interference.
  5. Behrens, T. E. J. et al. (2007). Volatility: surprise breeds doubt.
  6. Charnov, E. L. (1976). Optimal foraging, the marginal value theorem. When to leave a patch.
  7. Seeley, T. D. Bees switch tasks when no one unloads them.

Action selection and motivation

  1. Tyrrell, T. (1993). Computational mechanisms for action selection. PhD thesis, University of Edinburgh. Free-flow selection and 14 requirements.
  2. Lorenz, K. (1950); Hinde, R. A. (1956, 1960). The hydraulic model of drive and its critique.
  3. Brooks, R. A. (1986, 1991); Maes, P. (1989). Reactive and subsumption architectures.
  4. Zhang, J. and Berridge, K. C. (2009). Wanting versus liking: W = κ·V.
  5. Keramati, M. and Gutkin, B. (2014). Homeostatic reinforcement learning.
  6. Cabanac, M. (1971). Alliesthesia; Sterling, P. (2012). Allostasis.
  7. Oudeyer, P.-Y. et al. (2007). Curiosity driven by learning progress.

Body, cost and inheritance

  1. Kleiber, M. (1932). Metabolism scales with mass to the ¾.
  2. Elia, M. (1992). Brain costs about 18 times as much as muscle per gram.
  3. Niven, J. E. et al. (2007); Chittka, L. and Niven, J. (2009). The benefits of larger organs saturate.
  4. Kotrschal, A. et al. (2013, 2019). Large brain: smaller gut, fewer offspring, shorter life.
  5. Atkinson, D. (1994). Temperature-size rule.
  6. Hinton, G. E. and Nowlan, S. J. (1987). How learning can guide evolution. Baldwin effect.
  7. Moore, R. S., Kaletsky, R. and Murphy, C. T. (2019). C. elegans inherits a learned avoidance for about 4 generations.
  8. Ralls, K., Ballou, J. D. and Templeton, A. (1988). Conservation Biology, 2, 185. The cost of inbreeding in 40 captive mammal populations: 1.57 lethal equivalents.
  9. de Bruin et al. (2026). Predictable versus unpredictable change for inheriting what was lived.

Related systems

  1. Grand, S. and Cliff, D. (1998). Creatures. The closest precedent of creatures that learn.
  2. Yaeger, L. (1994). PolyWorld.
  3. Ackley, D. and Littman, M. (1991). Interaction between learning and evolution.
  4. Todd, P. M. and Miller, G. F. (1991). Learning toxicity from cues in simulated creatures.
  5. Blount, Z. D., Borland, C. Z. and Lenski, R. E. (2008). PNAS, 105, 7899. Replaying evolution from frozen clones: the logic of paired seeds.
  6. Park, J. S. et al. (2023). Generative agents. Societies of LLM agents.

Some 2025–2026 references were taken from abstracts and are marked as unverified in the repository's research notes.