Before the replicator
We asked whether the four founder-composition channels become more predictively integrated before stronger compositional restoration. First strict passes were a secondary view.
The broad answer is no. Across the 1,660 estimable valid episodes, our integration score was not a useful robust precursor. A much narrower onset pattern is interesting enough to test prospectively.
Past only. Then future.
At each passage we froze the clock. We measured the previous 65 composition states, then asked how well the lineage restored its composition over the next four passages.
65 states before T
Compare prediction by all four founder channels together with prediction after each of the seven possible splits. Keep the least-integrated split, then subtract the median score from 256 shuffled-future fits.
4 recovery cycles after T
Use the continuous median improvement in parent-to-return similarity—not a binary pass gate.
What could be measured?
We did not lower thresholds or fill missing values. Early passages often lacked a full 65-state history, so availability is much worse exactly where early warning matters most.
Usable Φ by checkpoint time. Before step 128: 291/555 (52%). At steps 384–512: 208/212 (98%). Only 27/56 old strict passing episodes were estimable.
The cloud is flat.
Across all usable episodes, more integrated past dynamics did not reliably mean better future compositional restoration.
Each dot is an episode. Colors are families. The orange line is the overall linear trend.
Essentially zero
Pearson: 0.018, bootstrap range −0.052 to 0.085.
Rank correlation: −0.004, bootstrap range −0.069 to 0.065.
Those ranges comfortably include no relationship. This is the clearest answer to the broad question.
Does Φ add anything beyond 25 ordinary measurements?
We trained without one family—or one passage schedule—and predicted the held-out group. Ordinary measurements included composition, flux, total predictability, mass, growth, motion, fragmentation, component count, and passage timing.
Orange bars show the percentage reduction in prediction error after adding excess Φ. Both gains are real numerically but tiny: 0.18% and 0.07%. Under every aggregate weighting, held-out R² stayed below zero, meaning the full models lost to simply predicting the evaluation-set mean on squared error.
Unseen family
Error fell by 0.18%. Run-balanced R² moved from −0.0277 to −0.0245.
9/12 family folds improved in MAE; the improvement was not large.
Unseen schedule
Error fell by 0.07%. Run-balanced R² moved from −0.1478 to −0.1463.
4/6 schedule folds improved in MAE; only 2/6 improved rank correlation.
Right before strict onset.
The old pass gate is still useful as a landmark. A “core onset” is the first strict pass after an earlier, non-overlapping informative failure. We compared four nearer checkpoints with the four before them.
Only four core onsets existed. Three had complete eight-checkpoint histories and all three rose. The fourth became constant in one founder coordinate, so this estimator was unavailable. Thin grey dots show the sparse at-risk matched comparisons.
Why this is interesting
The three measurable core onsets rose by +0.101, +0.112, and +0.215 excess Φ. The c06 and c10 rises exceeded every available family match, although c06 did so by only 0.010.
Why this is not the result yet
We already knew the onsets. There are only three measurable cases. Matched controls are sparse and sometimes rise too. Across all 22 first-pass trajectories, only five had complete histories: three rose and two fell.
Change the ruler; change the story.
A robust precursor should survive reasonable estimator choices. This one did not.
| Variant | Within-run rank | ΔR² unseen family | ΔR² unseen schedule | Read |
|---|
What we can say.
Yes
- The old
stop_without_phigate ruled out restoration, but it did not explain the temporal structure we could still see. - We built a past-only, GARD-inspired measure over four founder-composition channels.
- That measure is not a meaningful broad predictor of continuous recovery in this cohort.
- Three measurable strict onsets share a concrete pre-onset rise worth testing fresh.
No
- We have not found causal emergence in all of Lenia.
- We have not shown autonomous reproduction.
- We have not shown that Φ causes restoration.
- We have not prospectively predicted an unseen onset yet.
The clean next experiment
- Before simulation, freeze the cohort size, families, schedules, exclusions, missingness handling, matched contrast, and numerical success rule.
- Generate the fresh cohort without inspecting outcomes.
- Measure excess Φ at every passage from the start, using a shorter history only if it was frozen in step 1 so early onsets are observable.
- Freeze the event: first strict pass after a non-overlapping informative failure.
- Freeze one onset score: mean of checkpoints j−3…j minus j−7…j−4.
- Compare onset runs with the predeclared family- and schedule-matched risk set at j.
- Require the rise to meet the frozen rule across held-out families and schedules before calling it an early-warning signal.
If that passes: perturb matched states up or down in Φ while preserving composition, mass, and spatial statistics. Only then start talking about causality.
What stayed fixed.
Analysis
72 frozen trajectories · 360 exact raw files · 65 past states · 256 null shuffles per checkpoint · all seven partitions · 25 controls · leave-one-family-out and leave-one-schedule-out.
Boundary
Post-outcome exploratory. No p-values. No threshold tuning. No imputation. The sealed v2 result remains provenance, not a veto.
Sealed artifact SHA-256: 21e2adf1c50d3cdc835010e3927fc1bffbd18c000f07b1cdb433fe7c37ba966b (exploration report) · a21d65d22968a20d08a2a317628d59418334f72ea376affdbea40bb64b84cdda (point stream). Generated from the completed local artifacts; this HTML is explanatory, not part of the sealed result.