Flow Lenia · causal emergence programme
Research synthesis · 18 August 2026
Experimental record · working hypotheses · falsification criteria

Flow Lenia causal emergence

Our working criterion is counterfactual contraction: after distinct injuries and microscopic futures, does the system actively return toward one macroscopic endpoint?

01 / definition

What we mean by goal-directed behavior

A snapshot can look organized and still be causally empty. The stronger test is what the system does when reality pushes it off course.

Without contraction

Microscopic history remains loud. Tiny differences amplify, nearby branches fan outward, and interventions leave different scars. Novelty may be abundant, but nothing forces those possibilities to agree.

This is what most of our Flow Lenia branches did. They generated new futures cheaply. They did not yet show that any one future had become privileged.

With contraction

A higher-level organization should begin suppressing irrelevant microscopic variation. Different wounds, noise histories, and local routes are pulled toward the same macroscopic endpoint. The endpoint now explains the present because the system acts to restore it.

The transition of interest is from divergent to contractive counterfactual dynamics.

Observed

The simpler explanations failed

Scalar Φ did not generalize as a precursor or a useful steering target. Most futures dispersed. Bulk mass returned without shape correction. Wound response was strongly state- and direction-dependent.

We suspect

Structural reorganization may precede contraction

Before visible organismality, the effective parts or scale may reorganize. That event may convert an expansive field of possibilities into a basin with one preferred macroscopic future.

What would prove it

Prospective prediction and causal control

The reorganization must prospectively precede multi-wound contraction on unseen families. Inducing it must create the basin; blocking it must erase the basin. Anything less remains a lead.

Counterfactual futures before and after a goal basin forms On the left, several trajectories from nearby and wounded states diverge toward unrelated futures. After a structural reorganization in the middle, trajectories from distinct injuries curve toward one shared attractor on the right. WITHOUT CONTRACTION · FUTURES DIVERGE WITH CONTRACTION · INJURIES CONVERGE STRUCTURAL REORGANIZATION same state injury identity survives shared endpoint injury identity is erased
A · divergence. Nearby and wounded states retain their histories and separate. B · transition. Parts, boundaries, or effective scale reorganize. C · contraction. Different errors lose their identity as trajectories approach one target.
DisturbPush the state awayUse several genuinely different, severity-matched injuries.
BranchChange the microscopic futureReplay multiple fresh futures from each exact wounded state.
MeasureWatch error and spreadDoes target error shrink? Do the injury fans contract?
InterveneMove the organizationInduce or block the suspected precursor and move the basin itself.
02 / evidence

How the experiments changed the hypothesis

Each stage removed a simpler explanation and made the next experiment more demanding.

  1. 01No broad signal

    Confirm the precursor

    In 288 trajectories with 211 complete rows, the equal-family Φ effect was +0.0113, p=.4694. Adding Φ made control log loss worse: 0.686445 → 0.691239.

  2. 02Timing clue failed prospectively

    Test temporal timing

    72 trajectories yielded 2,027 valid episodes, 1,660 with Φ. Broad correlations were near zero and all held-out improvements were effectively zero. Although 3/3 measurable core onsets rose beforehand in exploration, the frozen follow-up on 23 matched onsets gave only +0.00320 [−0.10242, +0.12651].

  3. 03Causal null

    Steer Φ broadly

    Across 1,920 trajectories in 384 blocks, exact-next-checkpoint Φ separation was only +0.0261 [−0.0186, +0.0679]. All five behavior confidence intervals crossed zero.

  4. 04Φ separated; selectivity failed

    Apply an orthogonal pulse

    In 960 trajectories, Φ-R separated by +0.1885 [+0.0726, +0.3133], but the required TDMI nuisance-equivalence gate failed. Restoration reward was −0.2605 [−5.826, +5.319]. Selective Φ manipulation and behavioral benefit were not demonstrated.

  5. 05Futures dispersed

    Demand prospective prediction

    In 1,152 trajectories across 144 groups, late novelty was 0.4425, similarity fell 0.85494 → 0.75380, and outcome Y was −0.04475. Adding Φ-shape worsened held-out MAE: 0.0206824 → 0.0216150.

  6. 06Structural lead

    Inspect the decomposition

    119 fully estimable groups contained 0 positive labels. The only 2 positives were among 25 boundary groups—both f19, both found post hoc. All 144 historical stochastic partition fans still expanded; mean dispersion grew 0.07345 → 0.28385.

  7. 07Assay invalid

    Injure the putative goal

    All 160/160 tomography trajectories closed, but 0/8 states passed the immediate injury gate. Mean D4 was 0.000245, maximum 0.001565. All 128/128 raw repair values and 8/8 contractions were negative, but descriptively only.

  8. 08Direction achieved

    Make connected bites

    In 32 calibration runs, fixed 5% wounds passed 9/16 severity gates and 7.5% wounds passed 4/16. All 48/48 directional pairs separated, minimum 0.03006. Equal mass was not equal injury.

  9. 09Woundability mapped

    Match wound severity

    464 candidate evaluations produced 16 fresh selected wounds, of which only 9/16 entered the band. Yet all 24/24 direction pairs passed, minimum 0.04106. Direction was reliable; this grid and controller did not produce a common severity shell across the four-state panel.

Architecture atlas of 144 branching groups A twelve by twelve grid represents 144 groups. One hundred nineteen gray dots are fully estimable ordinary groups and contain no positive outcomes. Twenty-five blue boundary dots contain twenty-three negatives and two highlighted positives, both in family f19 and selected after outcomes were known. 144 BRANCHING GROUPS BOUNDARY CASES · 2 POSITIVES, 23 NEGATIVES
The exact anomaly

Zero in the smooth regime. Two at the boundary.

The 119 ordinary groups held no positive discovery labels. The 25 boundary groups held the only two, both f19. But 23 boundary groups were negative, both positives were selected after outcomes were visible, and all 144 fans expanded. This is a candidate crack—not a precursor result.

119 ordinary · 0 positive 25 boundary · 23 negative 2 boundary positives · both f19
The evidence now points from scalar magnitude toward changes in response geometry.

Scalar Φ was neither a broad precursor nor a useful held-out predictor, and the intervention studies did not establish it as a selective control target. The remaining lead is structural: rare canalization coincided with decomposition boundaries, and connected wounds revealed strongly state- and direction-dependent susceptibility.

03 / ladder

A hierarchy of goal-directed behavior

“Recovery” hides a hierarchy. We should reserve teleological language for the upper rungs.

Level 01

Bulk homeostasis

Refill material after loss. Bulk-mass repair was positive in all eight tomography states, roughly +0.0331 to +0.0362, with candidates and controls nearly identical.

Observed · generic
Level 02

Shock absorption

Limit how far damage propagates. The two f19 candidates wandered somewhat less in an invalid post-hoc read, but every historical stochastic partition fan still expanded.

Hint only
Level 03

Error correction

Detect a genuine structural error and reduce it. Our first wound was unreadable; the matched controller reached the severity band in only 9/16 cases.

Not cleanly tested
Level 04

Equifinality

Erase injury identity: distinct wounds and noisy futures collapse toward one macroscopic endpoint. This is the operational signature of a goal basin.

Not observed
Level 05

Goal memory

Carry a target that cannot be read from anatomy alone—and reveal it through repair, transplantation, or conflicting histories.

Not tested

The system restores bulk matter. We have not shown that it restores a specific organization.

That distinction is the heart of the programme. Passive stability, threshold-driven regrowth, and material redundancy can all mimic weak “recovery.” A protected goal must do more: it must make several wrong presents become the same right future.

Our target is level four, discovered before visible organismality and then pushed causally into level five.

04 / newest result

What the wound experiments revealed

9/16

Distinct directions, incomplete severity matching

The adaptive controller searched 29 connected-bite doses for each of four directions in four sacrificial states. Nine selected wounds landed inside the frozen D4 band [0.0325, 0.0375]. Only f24 passed all four. Recovery was therefore not authorized.

inside frozen D4 band closest candidate outside band q = fraction of total matter removed
Selected D4 distance and matter fraction for four injury directions in four states. Nine of sixteen selections passed the accepted band.
State Major + Major − Minor + Minor −
f20-s14-c24
3 / 4 in band
0.035535pass
q 7.25%
0.036005pass
q 2.75%
0.029935low
q 5.50%
0.035559pass
q 4.00%
f21-s14-c24
2 / 4 in band
0.035648pass
q 4.25%
0.028726low
q 2.75%
0.035521pass
q 4.25%
0.042057high
q 5.25%
f24-s14-c24
4 / 4 in band
0.033853pass
q 3.50%
0.036254pass
q 4.25%
0.036335pass
q 3.25%
0.037239pass
q 4.50%
f25-s14-c24
0 / 4 in band
0.038168high
q 3.25%
0.037696high
q 4.00%
0.038513high
q 4.25%
0.030684low
q 4.75%
Four-state matched-wound result Four panels plot selected D4 wound distance for major positive, major negative, minor positive, and minor negative directions. The green band spans 0.0325 through 0.0375. State f20 has three points in the band, f21 has two, f24 has four, and f25 has none. f20 · 3/4 accepted D4 band maj+maj−min+min− .0355 · 7.25%.0360 · 2.75%.0299 · 5.50%.0356 · 4.00% f21 · 2/4 accepted D4 band maj+maj−min+min− .0356 · 4.25%.0287 · 2.75%.0355 · 4.25%.0421 · 5.25% f24 · 4/4 accepted D4 band maj+maj−min+min− .0339 · 3.50%.0363 · 4.25%.0363 · 3.25%.0372 · 4.50% f25 · 0/4 accepted D4 band maj+maj−min+min− .0382 · 3.25%.0377 · 4.00%.0385 · 4.25%.0307 · 4.75% selected D4 selected D4 selected D4 selected D4 inside band outside band
Four woundability profiles. The green stripe is the same frozen acceptance band in every panel. f24 placed all four directions inside it; f25 placed none. Under this grid and controller, severity matching did not generalize across the panel.
Two f20 wounds with nearly equal morphological distance require very different removed mass For state f20, the major negative wound removes 2.75 percent of matter and reaches D4 0.036005, while the major positive wound removes 7.25 percent and reaches D4 0.035535. The mass doses differ by a factor of 2.64 despite nearly equal morphological displacement. f20 · TWO NEARLY EQUAL MORPHOLOGICAL WOUNDS 2.5%3.5%4.5%5.5%6.5%7.0% MAJOR − q 2.75% D4 0.036005 MAJOR + q 7.25% D4 0.035535 2.64× DIFFERENT MASS fraction of total matter removed
The clearest example. In the same f20 state, two directions reached almost the same D4 displacement while requiring 2.75% versus 7.25% matter removal. Equal mass is not equal injury; equal injury can demand radically unequal mass.
Angular success

Four distinct wound directions

All 24/24 within-state direction pairs exceeded the frozen 0.01 separation gate. The global minimum was 0.041058. These are not cosmetic variants of one lesion.

Radial failure

No common shell in this panel

Seven selections missed the D4 band under the frozen 29-point search. f25 missed in every direction; f24 passed all four. State and orientation jointly control injury magnitude, although the seven misses may still include coarse-grid interpolation failures.

Bold reading

An anisotropic woundability manifold

The system filters damage through its organization before damage becomes morphology. This is structured susceptibility—not yet repair, but precisely the geometry a repair controller would have to master.

What this does not show The 9/16 result is an outcome-seen engineering confirmation of the wound operator. It does not show causal emergence, recovery, a goal, or that f24 is an organism. Because the frozen all-16 gate failed, the 160-run recovery cohort was correctly not launched.
05 / next tests

Seven decisive experiments

Each experiment tests a different consequence of the hypothesis, with a result that would support it and a result that would end it.

Swing 01Take first

Measure basin formation directly

Hypothesis: agency begins exactly when the response operator becomes contractive—when different errors start losing their identity.

Current basisOur branches mostly dispersed. The newest assay finally generated genuinely distinct directions: 24/24 pairs separated, and f24 supplied four severity-matched wounds.

Decisive testAt dense checkpoints, give several matched wounds, replay common fresh futures, and measure both target-error reduction and cross-injury contraction at D+6, D+7, and D+8.

Decisive resultA state still looks ordinary, yet every valid injury repairs and all injury fans contract toward one future—while matched earlier states keep dispersing.

Swing 02

Track structural reorganization through time

Hypothesis: the basin begins with an abrupt reorganization of parts, boundaries, and effective scale—not with a smooth rise of scalar Φ.

Current basisThe 119 fully estimable groups held zero positives. Both positive groups sat among 25 estimator-boundary cases and belonged to f19. That is a post-hoc anomaly, not proof—but it motivates the test.

Decisive testMeasure every bipartition, minimum-information-partition turnover, entropy boundaries, covariance geometry, and basin contraction at dense checkpoints in unseen families.

Decisive resultThe same structural transition repeatedly appears one checkpoint before contraction begins, even when morphology and scalar Φ remain unremarkable.

Swing 03

Induce and disrupt the candidate mechanism

Hypothesis: if a structural reorganization creates the goal basin, a brief, correctly targeted intervention should install or destroy future-directed control after the intervention ends.

Current basisThe pulse study produced a +0.188 Φ-R separation, but nuisance equivalence failed and restoration benefit was not demonstrated. The next intervention should target the proposed organization directly, not assume that moving its summary statistic is selective.

Decisive testInduce the candidate partition/scale transition in a pre-basin state; block it in a near-transition state; release both; then apply the same blinded wound panel.

Decisive resultInduction creates persistent multi-wound contraction and blockade abolishes it, with matched energy, mass, timing, and sham interventions.

Swing 04

Test for hidden goal memory

Hypothesis: two states can look nearly identical while carrying different target memories. Anatomy may underdetermine the future the system will restore.

Current basisOur branching work showed that small, cryptic differences can become causally important later. The distributed-wound assay barely moved visible D4, yet those hidden differences subsequently amplified.

Decisive testFind morphology-matched states from different histories, blind the selector to history, apply the same calibrated lesion, and ask whether each returns toward its own history-specific endpoint.

Decisive resultSnapshot twins with indistinguishable current anatomy regenerate reproducibly toward different targets—and swapping the relevant history-carrying organization swaps the destination.

Swing 05

Test whether an organizer transfers the target

Hypothesis: a compact region may carry disproportionate information about the whole-system target. Move the region and the destination should move with it.

Current basisConnected bites showed extreme spatial anisotropy: in f20, two nearly equal D4 wounds required 2.75% versus 7.25% mass. Location and orientation can matter more than bulk dose in this panel.

Decisive testLocate regions whose perturbation most changes future contraction; transplant them between donor and recipient states with mass-, geometry-, and damage-matched shams.

Decisive resultThe recipient adopts the donor’s repair destination, while donor removal weakens or redirects its original basin.

Swing 06

Select for multi-wound recovery and replay ancestry

Hypothesis: robust agency may be too rare to await accidentally. Select directly for the behavior, then use the full ancestry to identify when contraction first appears.

Current basisNatural branching gave abundant novelty but vanishingly little reconvergence: just two positive groups in 144, both from one family and both discovered retrospectively.

Decisive testSelect for recovery from four orthogonal matched wounds, common-endpoint convergence across noise futures, environmental robustness, and phenotype maintenance under material turnover.

Decisive resultAn ancestral replay reveals a discrete generation where contraction first appears, preceded by a reproducible structural reorganization that can itself be manipulated.

Swing 07

Test conflicts between learned targets

Hypothesis: if goals are real higher-level causal constraints, incompatible goal memories should interact through dominance, boundary formation, or stable compromise—not merely average locally.

Current basisThe woundability result says organization is directional and distributed. A goal may therefore be separable from current shape yet still compete across the whole field.

Decisive testJoin regions or matched twins carrying different inferred targets, then wound the chimera and map which endpoint, boundary, or hybrid its futures protect.

Decisive resultThe composite resolves toward one target or a stable negotiated form with nonlinear, history-dependent dominance that cannot be explained by local material averaging.

Now · no new runs

Map the 464 response points

Plot every D4(q) curve. Determine whether the seven misses are coarse-grid interpolation errors or real jumps and forbidden severity regions.

Then · engineering

Freeze a state-aware wound controller

Match actual morphological displacement, preserve four directional distinctions, and demonstrate applicability before any recovery outcome is opened.

Then · discovery

Locate the first contraction

Measure dense pre/post checkpoints around candidate reorganizations and compare them with matched ordinary states.

Finally · causation

Write, erase, transplant

Only manipulation can turn an early warning sign into evidence that the emerging organization produces the protected future.

06 / kill criteria

Falsification criteria

These results would reject the corresponding hypothesis and force a change of direction.

  1. No contraction after valid matched wounds

    If diverse, visible, severity-matched injuries continue to fan outward across complete common futures, there is no demonstrated goal basin in those states. Stability and regrowth remain the better explanations.

  2. Structural reorganization does not predict contraction

    If decomposition-boundary events fail to precede contraction in unseen families—or appear just as often without it—abandon this precursor. The f19 anomaly was selection, estimator pathology, or coincidence.

  3. Architecture moves but repair does not

    If inducing or blocking the suspected reorganization leaves recovery competence unchanged under matched controls, our architecture metrics are epiphenomenal. They describe the transition without causing it.

  4. Twins share one injury future

    If snapshot-matched states from different histories always recover toward the same endpoint, there is no evidence for cryptic target memory beyond present anatomy.

  5. The organizer does not carry the destination

    If transplant effects vanish under mass-, geometry-, and damage-matched shams—or alter only local tissue—the target is not portably encoded in that region.

  6. Evolution produces only wound closure

    If selected systems seal local lesions but fail cross-injury equifinality, environmental generalization, or material-turnover tests, we evolved sophisticated homeostasis—not organism-scale agency.

  7. No discontinuity anywhere

    If ancestral replay and dense natural-history sampling show only smooth quantitative improvement with no identifiable birth of counterfactual contraction, the “sudden goal” picture is wrong. Emergence may be gradual—or Flow Lenia may not support it.

07 / belief

The decisive observation

We have not observed causal emergence. We have narrowed the required observation and identified the measurements that do not suffice.

Decisive evidence would be the onset of reproducible contraction toward one endpoint across distinct, matched injuries and future perturbations.

Prediction

A structural reorganization should precede the first checkpoint where wounds and noisy futures begin contracting toward one endpoint.

Causation

Inducing that organization should create the protected future; blocking it should erase the protection; transplanting it may move the destination.

Meaning

At that point, the whole earns causal status because it explains what its parts do next. Matter has begun acting on behalf of a state that does not yet exist.

08 / record

Evidence trail and claim boundary

Claim boundary

We have not found causal emergence, goal-directed repair, autonomous replication, consciousness, or a universal precursor. We found that scalar Φ did not perform as a predictor and was not demonstrated as a selective control target; that rare canalization coincided post hoc with structural estimator boundaries; that a distributed 5% injury was unreadable in the chosen morphology metric; and that connected woundability is strongly state- and direction-dependent. The proposed experiments on this page are falsifiable research commitments built from those results—not results themselves.