Hodos

Comparing processes as curves of distributions

UNLISTED — Shared, not announced. A working research record: it contains negative results and retracted claims on purpose. Findings here predate their preprint DOIs. Not indexed.
active 31provisional 4negative 16retracted 8total 59

Line of work

Field measured on

 ·  Work applying this engine in a specific domain keeps its own record, its own numbering and its own evidence ladder. Nothing on this page is a clinical claim, and the lane below is not cleared or approved by any regulator: Cordthym — cardiac rhythm. A result in one line is not a claim about another.

foundations (38) — The distance itself, its operators, and what they were measured against. The results the rest of the record stands on.

architecture (5) — Memory-based learning: the retrieve-compose-decode loop, composition, and the coverage law. An architecture claim, not an application result.

cardiac rhythm (5) — Cardiac rhythm application of the architecture, on public patient recordings. Application results, NOT architecture claims, and no clinical claim.

machine telemetry (11) — Prognostics and anomaly attribution on machine sensor data. Method and negatives, not a predictor and not a product.

The four named equations

The hypothesis is the premise; each equation it generates carries its own name, and the name is Hodos plus one word. This matters for reading the record below: the findings were written before the equations were named, so they speak of "the distance", "the ground" or "the surplus". Those are the objects named here.

Hodos Diastema (διάστημα, interval) — How far apart are two processes?
The model-free distance between processes as curves of distributions. It is the equation nearly every finding on this page is measured with — where a finding says "the distance", "D", or "the Fisher–Rao ground", this is what it means. 10.5281/zenodo.21612829

Hodos Symploke (συμπλοκή, interweaving) — What does a relationship create?
The emergence surplus: the joint against the product of the marginals. Cleared all four criteria fixed before it ran. 10.5281/zenodo.21850666

Hodos Systasis (σύστασις, composition) — What are the parts made of?
Relational — and the 'lossy' half is WITHDRAWN. It genuinely is relational: destroying the shared coordinate system collapses it. The earlier claim that it loses to the window's own values was withdrawn on 2026-08-15/16 by our own measurement, not by argument — the paired leave-one-out gap is 4 items of 200, then 4 to 12 of 600, sign-stable but below what this design can resolve, and a gap the design cannot resolve is withdrawn rather than kept as a null. 10.5281/zenodo.21850666

Hodos Chronos (χρόνος, duration) — How much time has a process lived?
Clears four of its five pre-registered bars. The fifth, sampling invariance, did not clear at the registered configuration - 0.759 against a bar of 0.05 - and the paper says so on the same page that announces the other four. 10.5281/zenodo.21861429

Where this has been measured

The same representation is applied in every field below. Whether it helps is a separate question, measured case by case — the fields where it did not are on this page with the same weight as the fields where it did. Generated from manifest.json, so it cannot disagree with the record underneath it. Click a row to filter.

FieldStatus mixWhat the data is
mathematics74 active 3 provisionalIdentities and proofs. No dataset — these either hold or they do not.
speech1512 active 1 provisional 1 negative 1 retractedReal recorded spoken digits, six speakers. The across-speaker split is the hard one: the same word through a voice the system has never heard.
handwriting128 active 3 negative 1 retractedReal pen motion from a graphics tablet. No audio, no shared sensor with speech — which is what makes it a replication rather than a second look.
text83 active 2 negative 3 retractedCharacter-level English. Discrete and already aligned, which is the regime the method is measured to be WEAKEST in.
human activity31 active 1 provisional 1 retractedBody-worn accelerometer windows — sitting, standing, walking.
light21 negative 1 retractedEmission spectra of materials.
sonar32 negative 1 retractedSonar range profiles and Doppler returns.
turbulence11 negativeTurbulent versus calm flow.
cardiac rhythm53 active 2 negativeInter-beat intervals from public patient recordings, per patient.
jet engine93 active 5 negative 1 retractedNASA C-MAPSS turbofan run-to-failure. Simulated, one engine family — it is not flight telemetry and a result here is not a claim about one.
synthetic vehicles21 active 1 negativeMulti-channel systems built deliberately to break the machinery — mixed units, awkward distributions, planted faults with a known answer.
zeta zeros22 negativeThe Riemann zeta zeros and candidate operator spectra. An honest appendix: everything here reproduces known mathematics and proves nothing new.
spacecraft telemetry41 active 1 provisional 2 negativeReal ESA mission telemetry (ESA Anomaly Dataset, CC BY 3.0 IGO), with channel-level ground truth and real, unequal subsystems.
synthetic pairs21 active 1 negativeConstructed pairs of processes with a known coupling, used where the right answer has to be known in advance.
astronomy53 active 1 provisional 1 retractedPantheon+/SH0ES: 1,701 public type-Ia supernovae, twelve measured columns per object. The catalogue ships inside the instrument that reads it. No cosmological claim is made on this data - the clock reads relations; its pre-registered cosmic-time bars did not clear at the registered configuration, and its own face says so.

Every finding, by number

Counts and the table below are generated from manifest.json, so they cannot disagree with the record. Newest finding: 2026-08-12.

FindingFieldStatusRung Date
1A process is literally a curve on a spherefoundationsmathematicsactive52026-07-21
2Time-averaging is quadratically blind to a brief eventfoundationsmathematicsprovisional22026-07-21
3The information-geometry ground beats Euclidean on real speech across voicesfoundationsspeechactive32026-07-23
4The same win replicates on handwriting, with no audio at allfoundationshandwritingactive42026-07-23
5The method fails where identity is static — the predicted boundaryfoundationshuman activityretracted32026-07-23
6The corrected boundary: two lanes, not onefoundationshuman activityactive32026-07-24
7Pitch invariance by construction, not by trainingfoundationsspeechactive52026-07-22
8The distinctive ground does NOT generalise across light, echo and turbulencefoundationslight, sonar, turbulencenegative52026-07-22
9The min-mean distance 'win' was a padding degeneracyfoundationslight, sonarretracted52026-07-22
10Doppler / velocity measurement is not an edgefoundationssonarnegative52026-07-22
11Against a published method: the field's warp is better, our ground still winsfoundationsspeech, handwritingactive62026-07-23
12The win decomposes: trajectory > ground > warpfoundationsspeech, handwritingactive42026-07-24
13The geometry needs 5–10× less data than a standard networkfoundationsspeech, handwritingactive42026-07-25
14It does not forget — and cannot, by constructionfoundationsspeech, handwritingactive52026-07-26
15A compressed latent representation does NOT helpfoundationstextnegative52026-07-25
16Anything between the landed point and the answer must start as a pass-throughfoundationstextactive52026-07-25
17On discrete aligned data the warp HURTS and the ground is neutralfoundationstextretracted52026-07-25
18Hierarchy does not beat a flat representationfoundationsspeech, handwritingnegative52026-07-25
19The geometry is not fp16-safefoundationstextretracted12026-07-24
20Measurement error: a growing label space looked like memory decayingfoundationsspeech, handwritingretracted22026-07-26
21An impossible value exposed a flaw that only a second dataset could revealfoundationshandwritingactive52026-07-26
22The Riemann work reproduces known mathematics and proves nothing newfoundationszeta zerosnegative52026-07-22
23Character spacing is not Riemann level repulsionfoundationszeta zeros, textnegative52026-07-25
24Memory helps a language model, but the geometry is not whyfoundationstextretracted52026-07-25
25The Fisher-Rao ground IS a better retrieval key than L2 -- once characters have a learned metricfoundationstextactive52026-07-26
26The learned character metric encodes linguistic class, and it is not frequencyfoundationstextactive32026-07-26
27The equation predicts the next step better than a trained network, with no trainingfoundationsspeech, handwritingactive52026-07-24
28The closed loop as a learner: retrieval and instant learning hold; the decode leg does not closearchitecturespeechactive32026-08-03
29Order-invariance holds exactly through the composed retrieval patharchitecturespeechactive32026-08-03
30Composition is task-dependent: blending helps the vote and hurts the artifact; the decode gap is structuralarchitecturespeechactive32026-08-04
31The personal condition: the loop closes end-to-end as a function of coverage, with zero trainingarchitecturespeechactive32026-08-04
32The decider is coverage-dependent, and blended evaluators suppress decode scoresarchitecturespeechactive32026-08-04
33The coverage law does NOT transfer to heartbeats: direction survives, magnitude does notcardiac rhythmcardiac rhythmnegative32026-08-05
34Same organ, other side of the boundary: the loop works on cardiac RHYTHM, per patientcardiac rhythmcardiac rhythmactive32026-08-05
35Measured against the field: below the published state of the art, and per-patient memory adds little on averagecardiac rhythmcardiac rhythmnegative32026-08-05
36The gap to the published field was data, not a ceiling: 0.983 at 64 examples per state (pure windows only — see finding 37 for the harder protocol)cardiac rhythmcardiac rhythmactive32026-08-05
37The excluded hard cases, scored: 0.966 on the continuous stream, and onset is caught on the first windowcardiac rhythmcardiac rhythmactive32026-08-05
38Matched pairs on jet engines: something beyond present distance predicts which unit fails sooner — but the differential operators are not what reads itmachine telemetryjet enginenegative42026-08-05
39Per-condition normalisation makes multi-regime fleets measurable at all — and a count caught the defect that would have hidden itmachine telemetryjet engineactive42026-08-05
40The excursion replicated pre-registered on two unseen fleets — and is retired anyway, because part of it was a modelling choicemachine telemetryjet engineretracted42026-08-07
41Attribution on real engines: the null control holds on hardware, and the binary question turns out to be the wrong onemachine telemetryjet engineactive32026-08-07
42Ordering measured against the right null: nine-tenths of the apparent concentration was the representation, not the faultmachine telemetryjet enginenegative32026-08-08
43Mixed units are not the blocker they were assumed to be — plain z-scoring beats both distributional standardisers on channels built to break itmachine telemetrysynthetic vehiclesnegative22026-08-07
44The layered architecture was refuted too broadly: it was the pooled statistic, not the hierarchymachine telemetrysynthetic vehiclesactive22026-08-07
45First contact with real spacecraft telemetry: the obvious instrument alarms on everythingmachine telemetryspacecraft telemetrynegative32026-08-08
46With the channel's own normal variability as the null, it works — and names the right channels at 4.5× the base rate against real ground truthmachine telemetryspacecraft telemetryactive32026-08-08
47The layered architecture does NOT transfer to real, unequal subsystems — flat wins on 40 events and loses on nonemachine telemetryspacecraft telemetrynegative32026-08-08
48Rare-but-normal events do alarm less than anomalies — by a registered margin, and still far too often to be usefulmachine telemetryspacecraft telemetryprovisional32026-08-08
49The premise made falsifiable: destroy the relations and identity goes on handwriting, but not on a jet enginefoundationshandwriting, jet enginenegative32026-08-08
50Hodos Symploke: a quantity for what a relationship creates, with the control built into the equationfoundationssynthetic pairs, handwriting, jet engineactive32026-08-08
51Hodos Systasis: frames derived from relations really are relational — and lose to simply using the windowfoundationshandwriting, jet enginenegative32026-08-08
52Hodos Chronos: intrinsic time is well posed on smooth processes and diverges on real sensor datafoundationssynthetic pairs, jet enginenegative22026-08-08
53Hodos Chronos rebuilt: four of five pre-registered bars clear, and the obstruction that stopped thirteen attempts was an artifact of our own estimatorfoundationsspeech, human activity, mathematicsprovisional22026-08-09
54A clock made of relations, pointed at 1,701 supernovae - and its cosmic-time bars did not clear at the registered configurationfoundationsastronomyactive32026-08-10
55The shipped warp was not a distance: a thing's distance from itself was negative, and grew with its lengthfoundationsastronomy, mathematicsactive32026-08-11
56RETRACTED: 'the whole converts' - the result described a catalogue that no longer existed, and does not reproduce on the current onefoundationsastronomyretracted32026-08-11
57Chronos is oriented, and nothing had ever checked the direction: 63 of 66 relations read differently reversed - and the headline survives either wayfoundationsastronomy, mathematicsactive32026-08-12
58The 'unexplained' detection band was the plant's own fault: |sin| halves the period, and the corrected rule is exact on 58 of 58 cellsfoundationsmathematicsactive22026-08-12
59A second level of time cannot be resolved on this catalogue: at nine level-2 frames the statistic's own calibration is wider than it claimsfoundationsastronomy, mathematicsprovisional22026-08-12

The record, newest first

59A second level of time cannot be resolved on this catalogue: at nine level-2 frames the statistic's own calibration is wider than it claimsfoundationsastronomymathematicsprovisional

2026-08-12 · rung 2 — measured once, synthetic data

Claim

If constitution has no floor, whatever a pairing produces can itself pair - a surplus timeline standing in its own relation would be a SECOND level of time. The registered test runs the nested-surplus gates at the clock's own shape: 1,700 objects give 67 surplus frames give about nine level-2 frames, the only window arithmetic allows. The calibration gate failed there: on synthetic noise the score must sit inside the width it claims (|z| < 2), and it did not - so the ruler's own markings are wider than stated at this depth, no reading taken with it can be placed, and the run discarded itself before scoring any real data. A resolution verdict: the question needs a longer series, and this catalogue cannot ask it.

Evidence

N0 on registered seeds: nested z of -0.66 / +2.98 / +0.82 against a |z| < 2 calibration. Mechanism measured over 12 seeds: the null is over-dispersed, sd 1.27 against a calibrated 1.0. The ledger row records the discard.

Scope / limits

The gate is a CALIBRATION FLOOR, never a claim that anything is unrelated - under the premise nothing is, and a higher level would stand in relation to everything by construction. Failing it says the instrument cannot certify itself at this depth; it neither finds nor refutes a higher time. The premise's own derivation expects nesting to add nothing if constitution is already pairwise - and an unaskable question confirms nothing either way. The nested object is pairs of pairs throughout: one pair's surplus standing in one relation with a third part.

Controls that could have killed it

  • Gate order is the design: the calibration floor runs before the planted oracle and before any real arm, and a run that fails it discards itself with the failure in the ledger.

universe-clock: run_higher_time.py · higher_time.json

58The 'unexplained' detection band was the plant's own fault: |sin| halves the period, and the corrected rule is exact on 58 of 58 cellsfoundationsmathematicsactive

2026-08-12 · rung 2 — measured once, synthetic data

Claim

Which planted periods each window can detect had resisted a formula - several in-band periods scored zero and an out-of-band one scored - and was recorded as unexplained rather than quoted. Measured on a grid of planted arms: the planted cycle uses |sin|, which doubles the frequency, so the arm's true period is HALF its parameter, and the standing formula had been scoring the wrong period (14/36 against the measured band). The corrected rule - some multiple of the effective period fits between the minimum return separation and the window count, AND the departure gate is satisfiable on the cell's own scales - reproduces the band exactly: 36/36 on the measurement grid and 22/22 on fresh cells, with the arithmetic half of every prediction printed before any distance was computed.

Evidence

probe_detectable_band.py: every detected cell's return lags cluster at integer multiples of period/2 (period 22 detected at lags 33/44/55; period 30 at 45). probe_detectable_band_confirm.py: 22/22 fresh cells; a +/-1-frame lag tolerance FAILED on 8 cells whose clusters are +/-2 wide - the tolerance was the miscalibration, and the failure is reported rather than dropped.

Scope / limits

Instrument calibration on planted arms only; no real data enters. A not-detected cell is a statement about the instrument at that window and period, never about data. Consequence applied the same day: the calibration opened one honest extra window (W=7) for the conversion re-measurement above.

Controls that could have killed it

  • Post-hoc fit vs prediction kept separate: the rule was refined on the measured grid, then confirmed on cells the grid never touched, with predictions printed first.

universe-clock: probe_detectable_band.py · probe_detectable_band_confirm.py

57Chronos is oriented, and nothing had ever checked the direction: 63 of 66 relations read differently reversed - and the headline survives either wayfoundationsastronomymathematicsactive

2026-08-12 · rung 3 — measured on real data, one modality

Claim

Chronos's estimator is an oriented function of its two arguments, and the shipped clock had read every relation in one direction chosen by nothing but the order the parts sat in a dictionary. Measured both ways: tau(a,b) differs from tau(b,a) on 63 of 66 relations, median relative difference 0.4485 against a 0.1828 same-direction seed floor - 2.45x, separable. The whole's total reads 113.978 one way and 106.318 the other (6.7% apart), but the share of time living between the things is 73.9% vs 74.2% - the claim that matters is direction-independent. Treatment settled from the framework's own canon: the two readings are two sides of ONE relation, never summed and never maxed - 66 relations carrying 132 numbers, not 132 relations.

Evidence

probe_is_time_directional.py: 63/66 differ, 2.45x the seed floor. run_direction_readings.py (full restatement, both directions, one seed): 61 relations resolved both ways, 2 one way only (each reading a real number one direction and exactly 0 the other), 3 neither. The per-relation convert/cycle classification is direction-blind to the last digit - max absolute difference 0.00e+00 both scores, zero verdicts flip - measured, not assumed from the joint's permutation symmetry.

Scope / limits

One catalogue, one seed for the both-directions restatement (the seed floor arm is what makes the asymmetry reportable). A relation live in one direction and exactly 0 in the other is a resolution verdict about this instrument, never a one-way relation - and never a horizon detection; it is a lead.

Controls that could have killed it

  • The seed floor: an order effect had to be sized against how much the SAME direction moves on a seed change before it could be called real.
  • A rounding defect was caught by two artifacts disagreeing by one count: round(tau, 6) collapsed a measured 2.1e-08 into a stored 0.0 - which is this instrument's under-resolved verdict, i.e. a verdict no measurement produced. All tau values now stored raw, and the face prints tiny readings as tiny, never as 0.00.

universe-clock: probe_is_time_directional.py · run_direction_readings.py · repair_tau_precision.py

56RETRACTED: 'the whole converts' - the result described a catalogue that no longer existed, and does not reproduce on the current onefoundationsastronomyretracted — positive claim

2026-08-11 · rung 3 — measured on real data, one modality

Claim

ORIGINAL WORDING, NOW WITHDRAWN: 'The whole converts, moderately: the state path scores 1.9134 against phase surrogates near 0.99 (min ratio 1.852 over a 1.5 bar) with zero returns - it converts and does not cycle.' Measured 2026-08-11 00:5x, matching a model distinction stated before the run; withdrawn the same day.

Evidence

Registered run: 1.9134, surrogates 0.9912/1.0331/0.9885, min ratio 1.852, returns 0. Withdrawal: five ledger rows marking each whole-level result not current, readable rather than deleted. Re-measurement sweeping W=5/6 (all usable windows, one planted period in band for all): conversion 0.890/0.948 against surrogates ~1.0 - ratios 0.877/0.942 against the 1.5 bar. Extension at W=7 after an instrument calibration opened it: 0.946, min ratio 0.943 - same verdict at the strongest window the path can carry.

Scope / limits

The retraction is about the CATALOGUE CHANGE, not an error in the registered run: the result was correct for the data it was computed on, and that data is no longer what the instrument reads. The small-window null is weak evidence - a 5-7 frame window gives the warp little to align - and transfers to nothing.

Controls that could have killed it

  • Withdrawn figures stay readable so a retraction can name exactly what it retracts.
  • Every usable window swept rather than one picked - reporting a single post-hoc window would be choosing the arm after seeing which arms exist.

Why this was wrong

We claimed this WORKED. Our own later measurement showed it did not, so the claim was withdrawn. The correction goes against the method.

A data-cleaning fix landed between the run and the day's end: a not-measured marker (-9) in one of the twelve parts had been read as a value, and removing those rows shrank the WHOLE's path from 1,701 objects / 67 frames to 712 / 26 - the five whole-level results had all been scored on a path that no longer exists. At 26 frames the pre-registered window cannot be asked at all (15 windows against the 25 the test requires), and the re-measurement at every usable smaller window - W=5, 6, and later 7 - reads no visible conversion (ratios 0.877 / 0.942 / 0.943 against the 1.5 bar). The retraction is about the catalogue change, not an arithmetic error: the original number was correct for the data it was computed on, and that data is not what the instrument reads now.

What SURVIVES the catalogue change: returns are 0.0000 at every live window, on windows that are live precisely because the planted cycle did score returns there - 'the whole does not cycle' replicates while 'the whole converts' did not. One half of the pairing stands on the current data. ⚠ CONTEXT ADDED 2026-09-09, and it weakens the surviving half: the rule that produced those zero returns was later given a known answer of its own - a periodic three-body orbit, which returns exactly by construction - and it could not distinguish a sequenced return from the SAME STATES WITH THEIR ORDER DESTROYED. Figure-eight 0.0603 against block-shuffled controls at 0.0604 / 0.0603 / 0.0602, a ratio of 1.00. It reads REVISITING, not RETURNING. The positive control named above establishes that the rule can FIRE on a planted cycle; it does not establish that the rule can tell a return from a revisit, which is what 'the whole does not cycle' needs. That phrase is therefore withdrawn to: no revisiting at the neighbour scale on that 26-frame path, under a rule with no positive control for a sequenced return. A density statement, never a verdict on cycling. The zero-returns measurement itself is unchanged and no number moves.

universe-clock: run_conversion.py · mark_stale_whole_results.py · run_conversion_small_window.py · run_conversion_w7.py

55The shipped warp was not a distance: a thing's distance from itself was negative, and grew with its lengthfoundationsastronomymathematicsactive

2026-08-11 · rung 3 — measured on real data, one modality

Claim

The clock shipped with raw soft-DTW as its warp. Raw soft-DTW is the undebiased Cuturi-Blondel object: the softmin subtracts a gamma*log term at every cell, so the bias accumulates with path length and self-distance goes negative. Measured on the project's own state path: D(A,A) = -0.469, and -1.235 for a longer A. One of the 66 relations displayed a negative 'moved' value on the shipped face. Fixed to the debiased soft-DTW divergence (exactly 0 on identity); tau was unaffected because tau is Chronos and only the travel figure used Diastema.

Evidence

check_warp_is_a_distance.py: soft_dtw D(A,A) = -0.469 / -1.235; soft_dtw_divergence and dtw both 0.00000 on identity. After the fix, travel spans 3.79-329.53 with zero negatives.

Scope / limits

The engine itself was untouched - its own source names the divergence as the proper non-negative form and dtw as the master distance; the defect was the instrument's configuration choice. Every earlier Diastema-based figure in this lane was computed with the biased warp; the ledger is append-only and records them as they were.

Controls that could have killed it

  • Found because a sizing probe printed a negative number - nothing in the build had checked that the distance obeyed the definition of a distance. When a quantity has a defining property, test the property; do not trust the name in a comment.

universe-clock: check_warp_is_a_distance.py · uclock/three.py

54A clock made of relations, pointed at 1,701 supernovae - and its cosmic-time bars did not clear at the registered configurationfoundationsastronomyactive

2026-08-10 · rung 3 — measured on real data, one modality

Claim

Chronos and Diastema run over every pairwise relation of the twelve measured columns of the Pantheon+/SH0ES catalogue: 12 parts, 66 relations, 85 dials, packaged as a one-file desktop instrument. Its pre-registered cosmic-time bars did not clear at the registered configuration and the face says so: this is a clock OF the relations in one catalogue; whether it reads the universe's age is what those bars could not resolve at this depth. What it reads that survives its controls: most of the whole's accumulated time lives BETWEEN the things rather than inside any of them - roughly 74%, a share that later proved direction-independent.

Evidence

C1 epoch ordering 1.164 against a 1.5 bar - did not clear. C3v2 rate agreement +0.094 against 0.30 - did not clear, and the rise loses to its own surrogates. A4 sampling 0.147 against 0.05 - did not clear, a live confound quoted beside every trend on this data. What passes: the six-part K6 sub-run's resolvable edges land on known astrophysics nobody told it about (host mass vs brightness residual is the published mass step).

Scope / limits

One public catalogue, one epoch ordering, seeds fixed. Frames are windows of 96 objects, not equal spans of time - every per-frame trend is quoted with that units caveat. Nothing here is a cosmological claim.

Controls that could have killed it

  • Phase surrogates at every level: a reading must beat the same path with its ordering destroyed, or it is reported as under-resolved - never as absent.
  • A shuffle measures the instrument's resolution, never whether a relation exists: no arm in this project may be labelled 'unrelated' by construction.
  • Liveness gates with planted arms (a conversion, a cycle, a drift, stasis) run before any real number is scored; a dead instrument refuses to score.
The A4 sampling verdict is the standing confound on every trend measured on this catalogue, restated per finding rather than footnoted once. WORDING 2026-09-07: from 2026-08-10 the title and claim read 'it does not tell cosmic time' and 'its pre-registered cosmic-time bars fail', and the evidence line read 'FAIL' three times - verdicts about the clock where what was measured is that three registered bars did not clear at this configuration. Restated as what this instrument at that configuration could not resolve. No number moves.

universe-clock: uclock/ + run_bars.py + ledger.jsonl

53Hodos Chronos rebuilt: four of five pre-registered bars clear, and the obstruction that stopped thirteen attempts was an artifact of our own estimatorfoundationsspeechhuman activitymathematicsprovisional

2026-08-09 · rung 2 — measured once, synthetic data

Claim

Duration for a system of interacting parts, accumulated from the relations between the parts rather than along a sampling grid. Thirteen constructions traced a single trade-off — a clock could be sure about the direction of time or quiet when nothing happened, never both — and that trade-off was about to be written up as structural. It was not structural. It was a one-sided nonlinearity in our own estimator: `max(eps - null, 0)` clips a quantity resting on a floor, and a clipped quantity rises abruptly and decays back for no reason but the clip, which reads as an arrow of time that is not in the data. The amount must keep the clip; the arrow must not. The fourteenth version reads the amount off the clipped curve and the direction off the unclipped one, and clears four of the five criteria fixed before any of them ran.

Evidence

Bars registered in advance, scored on the shared harness. A0 time-asymmetry (speech): 2.020 / 2.044 / 2.122 / 2.133 across 4 seeds, every p < 1e-12, against a 1.5 bar. A1 stillness (277 human-activity runs, rest vs motion): 2.278 / 2.137 / 2.218 / 2.209 / 2.252 across 5 seeds, every p < 1e-33, against a 2.0 bar. A2 relation-sensitivity: 2.722 at p 7.5e-07 where a magnitude-only measure is arithmetically blind at 0.958. A3 exactness, on parts independent by construction: z = -0.190, i.e. indistinguishable from its own null. A4 sampling invariance did not clear at the registered configuration: 0.759 against a 0.05 bar.

Scope / limits

One construction over the six-part relational graph, scored on speech, human activity and synthetic independent parts. A0 cannot be scored on the activity data at all — it is time-reversible, so the arrow fires at chance there — and A1 needs a dataset with rest in it, so no single dataset carries all the bars. Nothing here has been through a second pair of eyes or an independent reproduction.

Controls that could have killed it

  • THE HARNESS GATES ITSELF AND CAUGHT SIX DEAD INSTRUMENTS OF ITS AUTHOR'S MAKING. ORACLE and FLOOR are required arguments to every bar and the run ABORTS unless the oracle beats the floor by a declared margin. Each of the six would have produced a confident wrong number: an A0 oracle that is symmetric under reversal whenever activity is spread evenly; an A4 oracle whose counter drifts by construction; comparing totals where one arm has eight times fewer frames; and an A3 null whose own oracle sat at z = 9.3 on parts that are independent by construction. The self-test's own planted 'good clock' was wrong first and would have failed while the bench worked perfectly.
  • SEED-CHECKED BEFORE BELIEVING, NOT AFTER. The four-part version read 2.024 on stillness at seed 0 — a pass — and 1.917 / 1.967 / 1.958 / 1.988 on seeds 1-4. That pass was a lucky draw and would have been published. The six-part version clears on all five seeds with a minimum of 2.137, which is why it, and not the four-part version, is the claim.
  • The null took three attempts and each of the first two broke the bar the other passed: a permutation null changes the within-window marginal (right for a global surplus, wrong for a windowed one) and a circular shift re-aligns with the original at every multiple of the period, so on walking the surrogate is nearly as coupled as the truth. Fourier phase randomisation preserves the power spectrum exactly and destroys every alignment including periodic ones. A trade-off between two bars is a signal that the ESTIMATOR is wrong.
  • A different channel pair scores 3.339 on stillness. It is NOT used and is named here because choosing channels after seeing their scores is fitting.
  • Three one-line impossibility results were checked before any data was touched, and they constrain any future direction-sensitive quantity: total rise equals total fall on a curve returning near its start; I(A;B) = I(B;A) by definition; and KL(P||P-transpose) = KL(P-transpose||P) identically, so no statistic of transition counts can see the arrow of time — measured at 1.000 to every printed digit on three datasets.
A4 is not worked around and is not going to be. Its instrument was rebuilt three times — totals to rates, raw striding to anti-aliased decimation, per-array-length to per-real-elapsed-time — and the verdict never moved. There is also an argument that it may not be fully answerable: an arm holding BOTH the physical span AND the samples-per-cell fixed cannot be constructed, because after decimating there are not enough samples left to fill a window of the same sample count. You can hold the span or the cell occupancy, never both. The two span-matched arms agree to 13%, which is still above the 5% bar, so that sharpens the failure rather than removing it. The wider lesson is the one worth carrying: an apparent structural obstruction, traced across thirteen constructions and about to be published as one, was generated by a nonlinearity inside the measuring instrument. Before writing up any obstruction, check whether the estimator is manufacturing the effect being measured.

hodos/chronos.py · validate/_clockbench.py · validate/clockbench/ledger.jsonl · supersedes finding 52

52Hodos Chronos: intrinsic time is well posed on smooth processes and diverges on real sensor datafoundationssynthetic pairsjet enginenegative

2026-08-08 · rung 2 — measured once, synthetic data

Claim

Every other quantity here takes the time axis as given. Defining duration instead as distance travelled through the process geometry gives a process its own clock: a system that sits still accrues no time however long the wall clock runs. The construction is defined by nothing but successive relations between states — but it is only a clock where the process is smooth. On noisy data the reading depends on how often you looked.

Evidence

Relative change in total duration when the sampling rate is halved: smooth curve 0.003, random walk 0.325, real turbofan sensor trajectory 0.500 — rising to 0.022 / 0.690 / 0.884 at stride 8. On real telemetry, halving the sampling rate roughly halves the measured duration. A still process accrues exactly 0.000.

Scope / limits

One ground metric, one estimator, synthetic curves plus one real sensor trajectory. It says nothing about processes observed finely enough to be smooth at the scale of interest.

Controls that could have killed it

  • THE MACHINE CHECK IS THE SAMPLING-RATE TEST and it is the right one: a clock that depends on how often you consulted it is not a clock. It is asserted in the tests for BOTH outcomes — that it holds on a smooth curve and fails on a noisy one — so the boundary cannot be quietly forgotten later.
  • A still process is verified to accrue exactly zero time, which is the arithmetic half of the claim.
The mechanism is the coastline problem: discrete arc length converges for a smooth curve and diverges for a noise-perturbed one, whose true length is unbounded. Smoothing first, or declaring a resolution, would make the check pass — and would reintroduce exactly the kind of unjustified knob that retired the excursion claim on this record. Doing that after seeing the failure would be fitting the control, so it was not done. The useful half of the result is the other column: an intrinsic clock is meaningful precisely where the process is smooth, which is a statement about which processes have one at all.

hodos/chronos.py · tests/test_chronos.py · superseded by finding 53

51Hodos Systasis: frames derived from relations really are relational — and lose to simply using the windowfoundationshandwritingjet enginenegative

2026-08-08 · rung 3 — measured on real data, one modality

Claim

Every domain this engine touches gets its frame from a hand-built adapter, so an engineer picks the axis each time. Building the frame instead from nothing but relations — locating each window by its normalised similarity to a set of reference windows, so no bin means a frequency or a sensor — produces a genuine relational encoding and a worse one. On handwriting it beat the hand-built adapter (0.960 against 0.935) and still lost to a frame built from the window's own absolute values, referencing nothing at all (0.980).

Evidence

Handwriting, 20 classes, 200 items: adapter 0.935 / relational 0.960 / absolute 0.980 / per-item references 0.215. Jet engine, healthy against mid-life, 100 items: adapter 0.800 / relational 0.580 / absolute 0.590, with the relational arm failing its validity bar against a 0.500 floor. One setting across both domains, never tuned per domain.

Scope / limits

One construction of a relational frame, two domains, one downstream engine. It bounds this construction, not the idea that representations can be derived.

Controls that could have killed it

  • THE SHARED-FRAME CONTROL PASSED DECISIVELY AND IS WHAT MAKES THE NEGATIVE INTERPRETABLE: giving every item its OWN reference set, destroying the common coordinate system while changing nothing else, collapses handwriting from 0.960 to 0.215 — a drop of 0.745. So the frames genuinely were being compared relationally. The construction is relational AND lossy, which is a far more precise result than 'it does not work'.
  • THE DECIDING CONTROL is a frame that is forbidden the relation entirely — built from the window's own values with no reference to anything else — rather than an arbitrary variant of the relational one. It wins by 0.020 on handwriting and 0.010 on the engine.
  • A DIAGNOSTIC WORTH RECORDING: replacing the data-derived references with pure NOISE changes almost nothing (-0.010 and -0.040). That is the expected result rather than a defect — a reference set is a coordinate system, not content, and locating something against arbitrary landmarks is still a relational encoding, in the way that a + b = c and one-half plus one-half carry the same relation in different symbols.
The mechanism is the useful part: projecting a window onto a set of landmarks discards more than it organises. The window's own values already contain the local relations among its samples, and the landmark projection is a lossy summary of them. A derived representation is not thereby refuted — this particular derivation is.

hodos/constitution.py · validate/constitution_vs_adapters.py

50Hodos Symploke: a quantity for what a relationship creates, with the control built into the equationfoundationssynthetic pairshandwritingjet engineactive

2026-08-08 · rung 3 — measured on real data, one modality

Claim

What two processes do together, compared against what the same two would do if they did not interact, is a computable quantity in the same geometry the distance uses. At each instant the joint distribution of the pair is compared with the outer product of its own marginals; the surplus is read through the engine's master distance. If the parts genuinely do not interact the two are identical and the surplus vanishes — the control is a term in the equation rather than an arm bolted beside it.

Evidence

Four criteria fixed before running. Independent parts sit inside the quantity's own shuffle null (z = +0.06, p 0.46). Sweeping coupling from 0 to 0.95 gives z = +0.1 / +3.4 / +9.6 / +24.9 / +39.5 — monotone. On y = x squared, where Pearson correlation reads -0.095, it fires at z = +18.8, so it is not a correlation proxy. And on real data it reproduces a split measured by a different method before this equation existed: mean surplus z = +8.08 across handwriting part-pairs against +1.32 across jet-engine sensor pairs, a gap of +6.75 against a pre-registered bar of 5.0.

Scope / limits

Pairwise by construction; a K-part total carries a higher-order remainder this does not compute. The estimator needs many more samples than the joint has cells — a first version at 8 bins over a 16-sample window scored an INDEPENDENT pair at 1.29 against a coupled pair's 1.98, because a joint that sparse cannot resemble a smooth outer product whatever the parts are doing. Reported throughout as a z against its own null, never as a raw surplus: a finite window makes the quantity positive even for independent parts.

Controls that could have killed it

  • THE NON-INTERACTION TERM IS INSIDE THE EQUATION. If the joint factorises at every instant the surplus vanishes by arithmetic. Measured at 2.98e-08 rather than exactly 0, because the ground is 2*arccos(BC) and arccos has an infinite derivative at 1 — that is the ground metric's conditioning, eight orders below a coupled pair's value, and the claim on the record is 'zero to the numerical floor' rather than 'exactly zero'.
  • RELABELLING EITHER PART'S OWN BINS LEAVES IT UNCHANGED, asserted in tests. It reads the relation and not the substrate, and that is provable here rather than measured.
  • THE KNOWN-ANSWER ARM IS THE ONE THAT COULD HAVE KILLED IT — the handwriting-versus-engine split was established by destroying couplings in a separate experiment, before this quantity existed, so agreement is not something the design could have arranged.
  • The shuffle null preserves the shuffled part's marginal EXACTLY, so any drop is attributable to the relation and to nothing else.
PROVENANCE, unsoftened: a divergence between a joint and the product of its marginals IS mutual information in general form, and total correlation, interaction information, transfer entropy and integrated information all occupy this neighbourhood. Tagged ASSEMBLED. The contribution is the composition — the quantity computed with the Fisher-Rao geodesic rather than a Kullback-Leibler divergence, as a trajectory rather than a scalar, inside a process geometry that already supplies the alignment. The literature check behind that was six web sources with no multi-source corroboration, which is weak evidence of absence and is stated as such.

hodos/emergence.py · validate/emergence_known_answer.py

49The premise made falsifiable: destroy the relations and identity goes on handwriting, but not on a jet enginefoundationshandwritingjet enginenegative

2026-08-08 · rung 3 — measured on real data, one modality

Claim

The principle this record sits under says a thing is the pattern of its connections rather than its substrate. Stated as an empirical prediction it can fail, and it does — on one of two modalities. Three surgeries were applied to the same data: destroy temporal ORDER, destroy CROSS-PART COUPLING while leaving every part's own values byte-identical, and destroy the SUBSTRATE with per-part monotone unit changes while leaving every relation intact. On handwriting the prediction lands hard. On jet-engine telemetry it does not: the parts must be POOLED, but they need not INTERACT.

Evidence

HANDWRITING (20 classes, 200 items, majority floor 0.050): intact 0.935. Destroying cross-part coupling — every part keeping its exact values and the representation's time-average exactly unchanged — collapses it to 0.315 (−0.620, p 0.0001). Destroying order alone: 0.400 (−0.535, p 0.0001). Replacing the substrate with per-part monotone transforms that change every number, unit and distributional shape: 0.885, a difference of exactly 10 items in 200. Non-interacting controls, each part classified in ISOLATION and merged assuming no interaction: 0.565 (vote) and 0.755 (independent likelihoods multiplied) — the intact method beats the better of them by +0.180, p 0.0001. JET ENGINE at 55% of life (120 items, floor 0.500): intact 0.775, coupling destroyed 0.758 — a difference of +0.017 at p 0.42, so the cross-part relation carries essentially nothing there — while the non-interacting controls reach only 0.608 and 0.567, so the intact method still beats them by +0.167 (p 0.0059).

Scope / limits

Two modalities, one representation each, leave-one-out nearest-neighbour. This tests the premise's PREDICTION on these data; it does not establish the general ontological claim, and it is not a promotion of the principle to a theory — that requires replication by people with no stake in it. Neither dataset supports the premise on all five pre-registered criteria, so no rung is earned for it.

Controls that could have killed it

  • EVERY SURGERY'S INVARIANTS ARE ASSERTED IN CODE BEFORE ANY ACCURACY EXISTS. The order surgery is checked to leave the multiset of frames bit-identical; the coupling surgery to leave each part's sorted values identical AND the time-average unchanged to 1e-12; the substrate surgery to leave the rank order inside every part identical while provably changing the numbers. A control that is only approximately blind has already cost this project a result.
  • THE CONTROL THAT DECIDES IT, and the first version got it WRONG. Version 1 built ONE feature vector holding every part's statistics and ran nearest-neighbour across it — which can read how the parts sit relative to each other, so it was a relational method with time removed, not a non-relational one. It scored 0.922 and the run concluded, wrongly, that the parts alone already sufficed. Version 2 forbids the parts from being combined: each is classified alone, and the per-part answers are merged under an explicit no-interaction assumption. Those controls score 0.565 and 0.755 on the same task.
  • A HEADROOM GATE, added because version 1 lacked one. Its validity criterion asked only whether the intact arm beat chance; on the jet engine every arm including the control scored 1.000, so nothing could differ and the comparison measured nothing while the gate passed. Version 2 additionally requires no arm at ceiling and a spread between arms.
⚠ ONE CRITERION FAILED BY FLOATING-POINT EPSILON AND IS KEPT AS A FAIL. The substrate criterion demanded that changing every number, unit and shape cost at most 0.05; the measured cost is 187/200 against 177/200 = exactly 0.05, which the comparison reads as 0.050000000000000044 and rejects by 4.4e-17. This record already contains one criterion missed by 9e-18 and the rule has not changed: a registered bar is not renegotiated after seeing the number. Read plainly, substrate-independence TIED its tolerance on handwriting rather than clearing or missing it. ★ The split between the two modalities is the real content, and it is the same two-lane shape this record keeps finding: where identity lives in how the parts move together, destroying the relation destroys the identity and swapping the substrate costs almost nothing; where the signal is carried by the aggregate of many drifting sensors, the relation between them is not what is doing the work.

validate/relational_premise_test_v2.py

48Rare-but-normal events do alarm less than anomalies — by a registered margin, and still far too often to be usefulmachine telemetryspacecraft telemetryprovisional

2026-08-08 · rung 3 — measured on real data, one modality

Claim

82 of this mission's 200 annotated events are rare but NORMAL, and the dataset's authors state that not alarming on them is of high practical importance. Measured: the flat decider alarms on 0.975 of anomalies and 0.817 of rare-nominal events — a difference of +0.158, clearing a pre-registered 0.15 bar at p = 0.0005. The direction is real and registered. The absolute number is not usable.

Evidence

118 anomalies against 82 rare-nominal events, same instrument, same held-out-validated false-alarm rate of 0.025. Flat Bonferroni and FDR both give +0.158 at p = 0.0005 on a label-shuffle permutation. The layered decider goes the WRONG WAY, −0.133 at p = 0.98, and fails.

Scope / limits

One mission, one annotation scheme, and the events are labelled by ESA operators rather than by us. Being 'rare nominal' is a judgement about operational meaning, not a property of the signal, so a method reading the signal alone has no principled route to the distinction beyond magnitude.

Controls that could have killed it

  • Pre-registered margin and alpha, both fixed before the run.
  • Label-shuffle permutation, so a difference between two high rates cannot pass on point estimates alone.
  • The layered decider is reported at its actual value, in the wrong direction, rather than dropped for being inconvenient.
Marked provisional rather than active on purpose. An operator receiving an alarm on 82% of rare-but-normal events is receiving an alarm on almost everything unusual, and the registered margin does not change that — a statistically significant difference between 0.98 and 0.82 is not an operationally useful separation. The honest reading is that the thesis 'learn the individual vehicle so rare-but-normal does not alarm' is SUPPORTED IN DIRECTION and UNSUPPORTED IN MAGNITUDE by this measurement. ESA Anomaly Dataset (Mission1), Zenodo record 15237121, CC BY 3.0 IGO — attribution required, redistribution permitted.

validate/esa_channel_null_test.py

47The layered architecture does NOT transfer to real, unequal subsystems — flat wins on 40 events and loses on nonemachine telemetryspacecraft telemetrynegative

2026-08-08 · rung 3 — measured on real data, one modality

Claim

Every previous test of the subsystem-then-channel hierarchy used four invented subsystems of five channels each. This mission's subsystems are real and unequal — 42, 17, 11 and 6 channels. There the layered test is strictly worse than testing every channel flat: it detects 0.636 of anomalies against 0.975, and the paired comparison is one-sided.

Evidence

Paired McNemar on the same 118 anomaly events: 40 events the flat test catches and the layered one misses, ZERO the other way, p < 0.0001. Attribution follows the same order — layered recall 0.217 at 3.19× base against flat 0.429 at 4.54×. Mechanism measured rather than argued: mean subsystems reaching stage one is 1.271 of 4 on anomalies against 0.020 on held-out nominal windows, so the stage-one gate is specific but far too conservative on the large subsystem.

Scope / limits

One mission, one subsystem partition. This does not refute hierarchy in general — it refutes it for THIS correction on THESE group sizes.

Controls that could have killed it

  • Both deciders scored on the SAME null draws and the same events, so the comparison is paired and a McNemar test is available rather than two intervals that happen not to overlap.
  • Equal false alarm: 0.018 layered against 0.025 flat on held-out windows, so the flat arm's advantage is not bought with alarms.
  • REGISTERED PREDICTION, written before the run: 'two-stage will LOSE ground here, because subsystem_6 holds 42 channels and a maximum over 42 must clear a higher bar than a maximum over 6.' It held.
This is the correction the synthetic finding asked for. A maximum over 42 channels needs a far larger excursion to clear its own null than a maximum over 6, so a real event inside the largest subsystem never passes stage one and is never drilled into — the asymmetry that equal-size invented subsystems hid completely. The general lesson is not about hierarchies: a result measured on groups of equal size carries no information about groups of unequal size, and the synthetic version could not have revealed this. ESA Anomaly Dataset (Mission1), Zenodo record 15237121, CC BY 3.0 IGO — attribution required, redistribution permitted.

validate/esa_channel_null_test.py · supersedes finding 44

46With the channel's own normal variability as the null, it works — and names the right channels at 4.5× the base rate against real ground truthmachine telemetryspacecraft telemetryactive

2026-08-08 · rung 3 — measured on real data, one modality

Claim

Replacing the null with an empirical, per-channel one — how far does THIS channel move between two normal stretches, measured over 1,600 event-free windows of its own history — turns the instrument of finding 45 into a working one. It holds its false-alarm rate on held-out event-free windows, detects 97.5% of annotated anomalies, and identifies which channels were involved at four and a half times the base rate.

Evidence

False alarm 0.025 (flat Bonferroni), 0.028 (FDR), 0.018 (two-stage) on 400 HELD-OUT event-free windows that played no part in building the null. Detection 0.975 of 118 annotated anomalies for both flat deciders. Attribution against the dataset's own channel-level labels: flat Bonferroni recall 0.429 at precision 0.591 — 4.54× the 0.130 base rate — with a channel-label shuffle p of 0.0010; FDR recall 0.562 at precision 0.528 (4.06× base). 76 channels, 200 annotated events, empirical null of 1,600 windows per channel with the resolution checked against the corrected bar before scoring.

Scope / limits

One mission. Windows are resampled to a fixed 256-point grid, so transients shorter than the grid are not resolved. Recall of 0.43 means the majority of labelled channels are still missed; this identifies SOME of the right channels well above chance, not all of them. Real flight telemetry, but one vehicle and one annotation scheme.

Controls that could have killed it

  • THE HELD-OUT SPLIT IS THE GATE. 400 of the 2,000 event-free windows are reserved and never used to build the null, because scoring the false-alarm rate on the same windows that defined the null would pass by construction. This is what separates a real gate from a self-check.
  • The null is built from THAT CHANNEL'S OWN history, so its noise scale, its binning regime and its quirks appear identically in the statistic and in the null and cancel.
  • Attribution is scored against a channel-label shuffle, so 'flags a lot of channels and hits some by luck' cannot pass.
  • Resolution checked before scoring: the smallest reachable p from the empirical null is compared with the Bonferroni bar, because a bar below the instrument's own resolution manufactures a guaranteed zero.
This is the first attribution result in this record measured against REAL per-channel ground truth. Every previous one was either synthetic with a planted answer, or on hardware where the dataset labels fault modes rather than sensors and no per-sensor truth exists at all. The correction that produced it is the project's own standing rule turned on itself: the null must be the method's behaviour on null data, not an idealisation. ESA Anomaly Dataset (Mission1), Zenodo record 15237121, CC BY 3.0 IGO — attribution required, redistribution permitted.

validate/esa_channel_null_test.py

45First contact with real spacecraft telemetry: the obvious instrument alarms on everythingmachine telemetryspacecraft telemetrynegative

2026-08-08 · rung 3 — measured on real data, one modality

Claim

Asking whether a channel's distribution over a window differs from the window before it — a two-sample permutation test, the textbook construction — flags at least one channel on EVERY event-free window of real mission telemetry. False-alarm rate 1.000. It is not broken; it is correct and useless.

Evidence

150 event-free control windows drawn from stretches with no annotated event anywhere in them or in their reference, 76 channels, 2,000 permutations per channel, permutation count checked against the corrected threshold it must resolve (0.000500 against a 0.000658 bar, margin ×1.32). All three deciders — flat Bonferroni, flat FDR, two-stage — flagged 150 of 150. The pre-registered gate declared every other arm of that run VOID rather than failed, and they are reported nowhere.

Scope / limits

One mission, real flight telemetry, 5.5-hour windows on a fixed resampling grid.

Controls that could have killed it

  • The null control was the FIRST criterion in the file and it gated everything behind it. Without it this run would have reported detection of 1.000 on anomalies — a perfect score — and it would have been meaningless, because the same instrument scores 1.000 on windows where nothing happened.
  • Held to the pre-registered rule: a decider failing the null control has its other numbers voided, not merely marked weak.
The diagnosis is the value, and it is the same shape as the saturation already recorded on jet engines. On a real spacecraft every channel genuinely differs between any two adjacent stretches — orbital period, eclipse, thermal cycling, telecommands all move it — so a test with enough power finds a real difference every time. The exchangeability null answers 'did this window differ from the last one', and a monitor needs to ask 'is this unusual for this vehicle'. Those are different questions and only the second one is useful. ESA Anomaly Dataset (Mission1), Zenodo record 15237121, CC BY 3.0 IGO — attribution required, redistribution permitted.

validate/esa_layered_test.py

44The layered architecture was refuted too broadly: it was the pooled statistic, not the hierarchymachine telemetrysynthetic vehiclesactive

2026-08-07 · rung 2 — measured once, synthetic data

Claim

Testing a subsystem by pooling its channels' distributions was measured at 0.03 detection against a flat test's 0.50, and the conclusion drawn was that the layered subsystem-then-channel architecture is refuted. That was too broad. Asking instead whether a subsystem CONTAINS a diverging channel — a maximum over its channels rather than a pooled statistic — recovers it: 0.67 detection at zero false alarms and 4.4 times cheaper.

Evidence

4 subsystems × 5 channels, 30 seeds, criteria fixed before the run. Two-stage MAX 20/30 = 0.67 [0.49, 0.81] with 0/30 false alarms, against flat Bonferroni 0.50 [0.33, 0.67], flat FDR 0.53, and the pooled two-stage version it replaces at 0.03 [0.01, 0.17]. The registered prediction landed: mean subsystems flagged on a fault vehicle goes from 0.033 to 0.867 — stage one now fires at all, which is the whole finding. Cost 219,330 permutation-divergences against 961,200.

Scope / limits

Synthetic vehicles with EQUAL, INVENTED subsystems of five channels each, one faulted channel, all channels standard normal. On real hardware a raw maximum is dominated by the noisiest channel unless it is standardised first — and real subsystems are not equal sizes, which changes the null a maximum is scored against.

Controls that could have killed it

  • The MAX null takes the maximum over the SAME channels after shuffling each in time, so the largest-of-k inflation appears identically in the statistic and in its null and cancels. The null must contain the artefact or it cannot remove it — the same reasoning that fixed the constant-offset artefact.
  • A POST-HOC PAIRED TEST THAT CUTS AGAINST THE HEADLINE, run because the point estimates looked too good: McNemar on the same 30 vehicles gives 6 versus 1 discordant, p = 0.125. Not significant. The Wilson intervals overlap heavily too.
  • The comparison arms were reused from the prior run after verifying its recorded configuration identical, which is what made the test PAIRED and the McNemar possible at all.
The honest claim is therefore 'matches flat at equal false alarm and is 4.4 times cheaper', NOT 'beats it'. The point estimates say otherwise and are wrong to. Against the pooled version it replaces the comparison is not close: 19 versus 0 discordant, p = 0.000004 — that is the comparison carrying this finding.

validate/attribution_max_group.py · validate/attribution_max_paired.py · superseded by finding 47

43Mixed units are not the blocker they were assumed to be — plain z-scoring beats both distributional standardisers on channels built to break itmachine telemetrysynthetic vehiclesnegative

2026-08-07 · rung 2 — measured once, synthetic data

Claim

For months the record carried an untested assumption: that pressures, temperatures, voltages and rates must each become a distribution before this geometry applies. Measured, it is unsupported. On ten channels chosen to be awkward — lognormal flow, heavy-tailed bus voltage, a valve pinned at its rail, an integer counter — plain per-channel z-scoring wins by a wide margin at equal false alarm.

Evidence

At a 20-frame baseline, detection 0.69 for z-score against 0.28 (median/MAD) and 0.27 (probability integral transform), thresholds fitted on held-out no-fault seeds so the comparison is at equal false alarm. Across baseline lengths 20 to 200 z-score runs 0.89 → 0.97 and never once fails its false-alarm gate. A saturation prediction written before the run landed exactly: the PIT is bounded by its baseline sample, so it goes flat at 0.10–0.12 from drift 3 upward while z-score climbs 0.48 → 0.94.

Scope / limits

Synthetic channels designed to be awkward, not real telemetry. The claim is 'not the bottleneck on data built to break it', NOT 'will never be the bottleneck'. z-score handled them at 0.89–0.97, so they may be milder than real mixed-unit hardware.

Controls that could have killed it

  • The fault size was defined in Z-SCORE'S OWN NATIVE SCALE, so a win for the alternatives could not come from a flattering fault definition — the comparison is rigged AGAINST the incumbent, and the incumbent still wins.
  • An all-gaussian control arm gates adoption: a standardiser that wins on messy channels by wrecking clean ones is a different failure, not an improvement.
  • A CONFOUND OF OURS, found and measured rather than argued away: every arm used a 20-frame baseline, and 20 points do not punish the three methods equally — z-score estimates two numbers, the PIT estimates a whole distribution. Re-running across baseline lengths NARROWED the refutation: at 200 frames the PIT reaches z-score to within 0.04. It was starved, not useless. Its false-alarm control is still unstable across lengths (2 of 8 cells out of band), so it is not adoptable either way.
A second hypothesis of ours was proposed and refuted the same hour. Noticing the PIT could only reach six of twelve bins, we built a stretched version to fill the range — and it scored WORSE, 0.72 against 0.93, because stretching moves the baseline noise out in the same proportion as the signal. It is kept in the module marked measured-worse so nobody re-derives it. Three hypotheses were tested that day and two were refuted by their own runs; that ratio is the point of running them.

hodos/hetero.py · validate/hetero_channels.py · validate/hetero_baseline_length.py

42Ordering measured against the right null: nine-tenths of the apparent concentration was the representation, not the faultmachine telemetryjet enginenegative

2026-08-08 · rung 3 — measured on real data, one modality

Claim

Asked which sensor leads on a degrading engine, the method looks strikingly consistent — until it is scored against how consistent it is on HEALTHY engines instead of against a uniform ideal. Against the healthy population's own concentration the honest gap is +0.102, not the 0.95 the weaker comparison implied. What does survive is narrower and real: WHICH sensor leads differs between the two fault modes, even though neither arm concentrates.

Evidence

Validity arm passes on 100 held-out healthy engines with 12 sensors reaching at least 5% of top-1. Primary arm: healthy top-1 entropy 2.509, fault 2.407, gap +0.102 against a pre-registered ≥0.40 bar, p = 0.0105 — FAIL. Discrimination arm: total-variation distance between the two fault modes' leading-sensor distributions 0.280, p = 0.0330 against a ≥0.25 bar — PASS. A third arm on a single marker sensor fails at p = 0.1231.

Scope / limits

Simulated data, two fleets, one engine family. Rankings are read from the cached attribution state rather than recomputed, so this run and the run it corrects are comparable on representation. The p of 0.0105 is NOT to be rounded into significance — the criterion failed as registered and the effect size misses by four times regardless.

Controls that could have killed it

  • THE CONTROL THAT VOIDED THE BEST RESULT OF ITS DAY, and it was written against that result before there was one. Asked whether any sensor ranks first on HEALTHY engines anyway, the answer is yes — one does, on 24 of 200. So concentration in the fault arm is partly a property of the representation, and the two arms that depended on it were declared VOID by the rule fixed before the run.
  • What that voided: a table showing one sensor ranked first on 41 of 100 engines with an entropy of 1.605 against a 2.557 bar. Enormous, clean, and meaningless. Without the control it would have entered this record as 'the method consistently identifies T50'.
  • The pre-registration named the wrong sensors and is kept as written. The sharpest fault-mode difference in the data is a bypass-duct pressure named on 1 of 100 engines of one mode against 17 of 100 of the other — excluded from the declared list precisely because it is constant in the first fleet, which is exactly what makes it the cleanest marker.
This is the corrected form of a result that looked far better before it was controlled properly. The rule it earned: the null must be the method's own behaviour on null data, not a uniform ideal — a method that ALWAYS names the same sensor concentrates perfectly whether or not anything is wrong. The surviving positive is worth stating plainly on its own: ordering discriminates fault modes but does not concentrate.

validate/ordering_healthy_null_run.py

41Attribution on real engines: the null control holds on hardware, and the binary question turns out to be the wrong onemachine telemetryjet engineactive

2026-08-07 · rung 3 — measured on real data, one modality

Claim

The per-channel permutation null — built after an earlier version named channels on pure noise — holds on real degrading hardware: on engines that have NOT degraded the module names a sensor on 4 of 100 and 1 of 100. But on degraded engines it names 89–92% of all sensors, and that is CORRECT rather than broken, which makes the binary 'is this channel significant' question useless on coupled machinery.

Evidence

400 engine-runs (100 fault + 100 healthy on each of two fleets), 15 sensors common to both, permutation count sized from the corrected threshold it has to resolve (401 against a Bonferroni bar of 0.00333, margin ×1.34, checked before running). On unit 1, 14 of the 15 sensors have genuinely moved more than 1 sd from their own baseline — NRf +8.2, T50 +7.5, Ps30 +6.0, Nf +5.8, P30 −5.7. Across all 200 degraded engines the naming rate is 0.89 and 0.92.

Scope / limits

Simulated run-to-failure data, one engine family, no per-sensor ground truth — the dataset labels fault MODES, not sensors. The series is truncated to the healthy reference plus the scored window, so nothing here speaks to WHEN divergence started.

Controls that could have killed it

  • T1, the healthy-engine null, gates everything else and is the load-bearing arm: an earlier uncalibrated version named a channel on 3 of 5 pure-noise seeds, which is what made the null non-optional.
  • The saturation was CHECKED AGAINST THE DATA before anything was built on it. Flagging 14 of 15 sensors looked like a broken control; the sensors had genuinely moved. The saturation is the machine, not the method.
  • Two criteria (within-mode consistency and between-mode discrimination) were expected to fail for that structural reason, were kept anyway, and did fail. Deleting a pre-registered criterion after watching it fail would have been the dishonest move.
The generalisation is the part that matters: every attribution result before this one was measured with ONE faulted channel among healthy ones. That is the easy regime and it is not the regime real hardware is in. A turbofan is thermodynamically coupled — at end of life the true answer to 'which channels diverged' really is 'almost all of them', and a monitor that says so has told an operator nothing. The answerable question on coupled hardware is ORDERING, not significance.

validate/turbofan_attribution.py · hodos/attribution.py

40The excursion replicated pre-registered on two unseen fleets — and is retired anyway, because part of it was a modelling choicemachine telemetryjet engineretracted — positive claim

2026-08-07 · rung 4 — replicated on a second, unrelated modality

Claim

ORIGINAL WORDING, NOW WITHDRAWN: 'How far above its current position an engine had already been — one number, the excursion — predicts which of two matched engines fails sooner, and it replicates pre-registered on two unseen fleets at 0.721 and 0.764, p < 0.0001, against a control blind at exactly 0.500.'

Evidence

The replication is real and its numbers stand. What was missing was a test of whether the quantity survives the arbitrary choice inside it.

Scope / limits

RETIRED as a claim 2026-08-07. The measurement is not disowned; the interpretation is.

Controls that could have killed it

  • The insight that made it testable: the baseline does TWO jobs — it standardises each sensor AND supplies the reference the excursion is measured from — and varying its length moved both at once, which is why 'baseline-sensitive' was never localised.
  • Dropping the REFERENCE role alone collapses the spread across the baseline grid from 0.262 to 0.022, which locates the sensitivity precisely.
  • Three baseline-free formulations built and scored against a pre-registered 0.65 bar.

Why this was wrong

We claimed this WORKED. Our own later measurement showed it did not, so the claim was withdrawn. The correction goes against the method.

The excursion moves with a knob nobody has justified. Across blind configurations of the baseline it spans 0.533 to 0.795 on one fleet — a range that contains both 'strong result' and 'nothing'. The sensitivity was then localised to the REFERENCE distribution, and no baseline-free formulation recovers the signal: 0.535 and 0.604 against a 0.65 bar, with the radius forms sitting near chance at 0.51–0.57. So stability and signal turned out to live in different places — the form that is stable carries nothing, and the form that carries something is unstable. A pre-registered replication and an unresolved dependence on a modelling choice travel together or neither does. ⚠ This does NOT prove the effect is pure artefact: a radius upper-bounds a signed overshoot, so the failure of the baseline-free version proves only that no baseline-free version has been found. Those are different statements and merging them would be a second error.

This is a retraction of a POSITIVE claim — the rarer kind on this page, and the uncomfortable one. It removed the single best number the machine-telemetry track had.

validate/excursion_baseline_free.py · validate/turbofan_fragility.py

39Per-condition normalisation makes multi-regime fleets measurable at all — and a count caught the defect that would have hidden itmachine telemetryjet engineactive

2026-08-05 · rung 4 — replicated on a second, unrelated modality

Claim

On fleets that operate in six different flight conditions, treating the vehicle as one regime scores at chance. Giving each engine its own baseline PER CONDITION, read from operational settings alone, recovers the measurement — and the excursion replicates on two fleets that had never been touched.

Evidence

Control blind at 0.500 EXACTLY on 4 of 4 fleets (482 engines, 6 conditions, 2 fault modes). The naive single-regime treatment collapses to 0.516 and 0.509 — chance — against 0.721 and 0.764 for the per-condition version, permutation p < 0.0001 on both. The margin over the dynamics operators is +0.048 on FD002 against a pre-registered +0.10 bar, so 'decisively beats the operators' is NOT claimed and was recorded as a fail.

Scope / limits

Simulated run-to-failure data, one engine family. Baselines are set from an engine's first 8 flights in each condition; nothing is borrowed from another engine or a population model.

Controls that could have killed it

  • Distance-only control blind by construction, verified at 0.500 exactly on every fleet.
  • A CONDITION-COUNT CHECK THAT GATES. The detector found 7 operating conditions where the documentation says 6, and the first version PRINTED the mismatch and scored anyway. Cause: one condition's second setting sits near a rounding boundary, so noise split it in two. Re-keying took FD002 from 146 to 241 usable engines and flipped the registered criterion from FAIL to PASS. The script now REFUSES to score a fleet whose condition map disagrees with the documentation.
  • Criteria fixed in the file before any FD002/FD004 number existed.
The rule this earned is worth more than the result: a check that reports and does not gate is not a check. It had been printing a real defect into the log for as long as the pipeline existed.

validate/turbofan_conditions.py

38Matched pairs on jet engines: something beyond present distance predicts which unit fails sooner — but the differential operators are not what reads itmachine telemetryjet enginenegative

2026-08-05 · rung 4 — replicated on a second, unrelated modality

Claim

Given two engines equally far from their own baseline at the same age, one of which will fail much sooner, the differential-operator family built for exactly this question (velocity, acceleration, curvature as a three-vector) does NOT tell them apart on a fleet it has never seen. A single path statistic does.

Evidence

45 disjoint matched pairs, same cycle, distance-from-own-baseline matched to 0.091%; sooner-failing engine has mean remaining life 76 cycles against 149. On the held-out fleet FD003 — 100 engines untouched by any earlier choice — the dynamics three-vector scores 0.594 with a sign-flip permutation p = 0.2675, which is chance. Dynamics minus the path baseline is −0.156. A curvature direction FIXED IN ADVANCE scored 0.562 held-out against 0.737 on the fleet it was first spotted in. What did carry the signal is one number off the distance curve: how far above its current position the engine had already been.

Scope / limits

Simulated run-to-failure data from one engine family (NASA C-MAPSS). This is not flight telemetry and a result here is not a claim about spacecraft. One window size; the ordering between arms flips across the window sweep.

Controls that could have killed it

  • THE CONTROL IS BLIND BY CONSTRUCTION, not by hope: matching equalises distance-from-baseline, so a distance-only decider must score 0.500 — and it does, EXACTLY 0.500, on four fleets and 482 engines.
  • Pairs additionally SIGN-BALANCED. Matching equalises magnitude but not sign, and a one-feature rule reads only the sign; the unbalanced version's control sat at 0.578, which is the residual-sign rate to three decimals.
  • The curvature direction was pre-registered before the held-out run — which is the only reason its collapse from 0.737 to 0.562 counts as a refutation instead of an anecdote.
  • Held-out fleet never used for any design choice.
An earlier version of this run PASSED at 0.711 (p = 0.0075) and was withdrawn, twice, both times because of a control rather than a result. The first control had TIME baked into it — it measured distance travelled over a fixed number of cycles, which is speed under another name, so it was the thing it was supposed to be blind to. The second was only approximately blind. The honest programme claim after all of it is narrower than the one this work set out to make: velocity generalises as a DETECTOR; the operators do not characterise change.

validate/turbofan_matched_pairs.py · validate/turbofan_replication.py

37The excluded hard cases, scored: 0.966 on the continuous stream, and onset is caught on the first windowcardiac rhythmcardiac rhythmactive

2026-08-05 · rung 3 — measured on real data, one modality

Claim

Findings 34 and 36 excluded windows straddling a rhythm change — the onsets and offsets a monitor exists to catch — which made their numbers non-comparable to published detectors evaluated on continuous streams. Scoring every window with nothing excluded, a mixed window labelled by its majority class: accuracy 0.966 across 9 patients, inside the published 95-98% band on the harder protocol. Transition windows do cost accuracy (0.902 on transitions vs 0.969 on pure windows) but they do not move the stream result out of band. Separately, onset is not slow: across 14 AF episodes the loop's first AF call comes on the FIRST window of the episode in every case (median lag 0 windows, 90th percentile 0).

Evidence

Per-patient stores of 64 pure windows per state (finding 36's configuration); queries are the remainder of that patient's stream IN ORDER with nothing filtered, up to 150 windows each. Per-patient stream accuracy ranged 0.873-1.000, with the weakest patient the one carrying by far the most transitions (28 of 150 windows). Onset lag measured only on true AF runs of at least 3 windows.

Scope / limits

Per-patient only; no cross-patient claim. One database, seed 0, one window length (64 intervals), majority-label convention for mixed windows. Published detectors differ in episode-level scoring conventions, so this is a fair-protocol comparison rather than an identical-protocol one. 14 episodes is a small sample for the onset claim.

Controls that could have killed it

  • nothing excluded — the exclusion that made the earlier numbers questionable is precisely what this removes
  • transition and pure windows scored separately as well as together, so the cost of the hard cases is a reported number rather than a hidden one
  • queries taken in stream order rather than shuffled, matching how a monitor would actually see them
The pattern across findings 33-37 is the useful object: a null on beat shape, a working system on rhythm, a self-inflicted 'we are behind' from data starvation, and now the excluded hard cases scored and survived. Each step was a caveat someone could have raised, tested by us before anyone raised it. What remains genuinely open and unmeasured: cross-patient transfer, longer horizons than these recordings, and whether the onset result holds on episodes shorter than 3 windows.

validate/afib_transitions.py

36The gap to the published field was data, not a ceiling: 0.983 at 64 examples per state (pure windows only — see finding 37 for the harder protocol)cardiac rhythmcardiac rhythmactive

2026-08-05 · rung 3 — measured on real data, one modality

Claim

Finding 35 recorded the loop BELOW the published 95-98% band at 0.935. Two follow-ups locate why, and it is not the method. First, a tune-on-dev / judge-on-held-out sweep over window length, sub-window and bin count moved held-out accuracy by exactly nothing (0.940 for both the swept-best and the hand-picked default, while the dev-set best rose 0.925 -> 0.945) — the hyperparameters are not the limit, and the dev gain was overfitting. Second, extending the store past the arbitrary 16 examples per state: accuracy 0.935 (16) -> 0.967 (32) -> 0.983 (64), still rising, with the plain-L2 control trailing at every level (0.898/0.942/0.947). The published comparison was between a tuned literature using full training data and a loop given 16 examples; matched for data, the loop reaches the top of that band.

Evidence

Coverage extension on 9 patients (one lacks enough windows at 64), 40 queries each, per-patient stores, same representation and splits as finding 34. Geometry vs plain L2: +3.7 / +2.5 / +3.6 points at 16 / 32 / 64. Tuning experiment: 11 configurations swept on 5 dev patients, verdict measured only on 5 patients never seen by the sweep; dev range 0.900-0.945, held-out difference between chosen and default 0.000.

Scope / limits

IMPORTANT AND UNFLATTERING: windows whose label is not pure are EXCLUDED, so rhythm-transition windows — onset and offset, the genuinely hard part — are not scored here, while published detectors are usually evaluated on continuous streams that include them. The comparison is therefore indicative, not like-for-like, and the number should not be quoted as beating those methods. Per-patient only; no cross-patient claim. 64 examples per state is roughly an hour of that patient's own labelled rhythm. One database, seed 0.

Controls that could have killed it

  • tuning judged ONLY on patients the sweep never saw — the dev/held-out split is what exposed the tuning gain as overfitting
  • plain-L2 control carried at every coverage level, so the improvement cannot be attributed to more data alone
  • the hand-picked default retained as a baseline arm rather than discarded once a better dev config existed
The methodological content is worth more than the number: a result recorded as 'we are behind the field' turned out to be 'we gave it a sixteenth of the data', and the way that was established was by testing the boring explanation (knobs) before the flattering one (data). Both were run before this was written. The behaviour is also exactly what a memory-based learner should show — accuracy rising monotonically with stored experience, no retraining at any point — and it is the same curve shape as the speech coverage law, in a domain where the shape was in question. The exclusion caveat named in this finding's scope is closed by finding 37, and 0.966 there — measured with nothing excluded — is the number that should be quoted, not the 0.983 in this title. The two differ because this one skips the windows where the rhythm changes mid-window, which are the hard ones.

validate/afib_tune_holdout.py

35Measured against the field: below the published state of the art, and per-patient memory adds little on averagecardiac rhythmcardiac rhythmnegative

2026-08-05 · rung 3 — measured on real data, one modality

Claim

Two honest comparisons, both unfavourable. (1) Published RR-interval AF detectors report 95-98% accuracy on this database (sensitivity ~96-97%, specificity ~97-98%); the untuned loop of finding 34 reaches 93.5%, so on raw accuracy this approach is BEHIND the field, not ahead of it. (2) With store size held equal, a store of the patient's OWN windows beats a store built from other patients' windows by only +2.7 points accuracy and +3.7 points specificity on average across 10 patients — both far under the pre-registered 10-point bars. The headline personalisation pitch does not survive its own test.

Evidence

Personal vs population stores, 16 windows per class in BOTH arms so the comparison isolates WHOSE memories rather than how many, identical queries and splits per patient. Mean gains: accuracy +0.027, specificity +0.037, sensitivity -0.003. Four of ten patients showed exactly zero gain because the population store already scored 1.000 on them; three showed small negative gains. State-of-the-art figures are from the published literature on the same database, not re-measured here.

Scope / limits

One database, one window size, one representation, seed 0, per-patient evaluation. The loop is untuned — no window/bin/representation search was run — while published methods are tuned, so the accuracy comparison flatters the field somewhat; that does not change the direction of the result.

Controls that could have killed it

  • equal store size in both arms — the entire point, since a larger personal store would confound WHOSE memories with HOW MANY
  • population store balanced per class and drawn from all other patients
  • specificity reported separately from accuracy, because the clinical complaint about existing monitors is false alarms, and an accuracy gain carried by sensitivity would not answer it
SUPERSEDED IN PART BY FINDING 36: the accuracy half of this finding was a data-starvation artifact, not a ceiling — at 64 examples per state the same loop reaches 0.983. The personalisation-gain half stands. POST-HOC OBSERVATION, explicitly not a claim and not pre-registered: the gains are not spread evenly. Personalisation gain correlates strongly NEGATIVELY with how well the population store already does (r = -0.86 for accuracy, -0.87 for specificity). For the three patients the population store handles poorly (<0.90) the mean gains are +0.100 accuracy and +0.189 specificity, and for the single worst case the population store's specificity of 0.387 — a majority-false-alarm regime — rises to 0.871 with that patient's own memories. For the seven it already handles well, the mean gain is -0.005. That pattern, IF it survives a pre-registered test on held-out patients, would reframe the architecture's role from 'a better detector' to 'a rescue layer for the atypical patients population models fail' — which is a narrower and more defensible claim than the one this finding refutes. Until that test exists, it is a hypothesis, and the recorded result is the negative one.

validate/afib_personal_vs_population.py

34Same organ, other side of the boundary: the loop works on cardiac RHYTHM, per patientcardiac rhythmcardiac rhythmactive

2026-08-05 · rung 3 — measured on real data, one modality

Claim

Finding 33 failed on single-beat shape. Asked instead about RHYTHM — atrial fibrillation, which is defined by irregularity over time and therefore sits on the temporal-evolution side of the two-lane boundary — the same loop works: per-patient recognition of AF vs that patient's own normal rhythm reaches 0.935 at 16 stored windows per state (chance 0.50), and the decode leg lands in-class 0.89-0.92 at every coverage level — the leg that never rose at all on beat morphology. The geometry beats a plain-L2 control on identical windows at all four coverage levels (+3.0/+5.5/+4.3/+3.7 points).

Evidence

10 patients from a public AF database, each containing both their own AF and their own non-AF periods, so the store holds that patient's own examples of each state. 64-interval windows, window-disjoint store/judge/query splits, transition windows excluded (labels must be pure), unblended 1-NN judge, 60 queries per patient. Recognition by coverage (2/4/8/16 windows per state): 0.895/0.888/0.915/0.935. Plain L2 on identical windows: 0.865/0.833/0.872/0.898. Decode: 0.892/0.915/0.907/0.895. BOTH pre-registered criteria FAILED AS WRITTEN and are reported as such: A1 required monotonic rise AND >= 0.85 — the level was cleared comfortably but a 0.7-point dip between the first two coverage points breaks monotonicity; A2 required the geometry to beat L2 by >= 5 points at top coverage — the margin there is 3.7, though the geometry leads at every level and by 5.5 at one of them.

Scope / limits

Per-patient only. No cross-patient claim is made or tested — that is the transfer problem finding 31 leaves open, in a new domain. One database, one window size, seed 0, deltas-of-intervals representation. Two states, chance 0.50.

Controls that could have killed it

  • plain-L2 control on the IDENTICAL windows and splits, so the comparison isolates the metric rather than the representation or the loop
  • transition windows excluded — a window's label must be pure, or a mixed window would score as an error for either arm arbitrarily
  • judge is 1-NN over held-out windows of the same patient, no blended prototypes (per finding 32)
The scientific content is the CONTRAST with finding 33, not the accuracy alone: the same organ, the same patients' data, the same loop, and the same encoder family give a null on beat shape and a working system on rhythm dynamics — which is what the two-lane boundary predicts and is the sharpest confirmation of it so far. The registered failures are kept because they are the honest ones: a criterion that demands strict monotonicity across noisy coverage points, and a 5-point margin demand that a consistent 3-5 point advantage does not meet at the single point where it was checked. Finding 35 places this against the published state of the art and against a population-memory control, and the answer there is not flattering.

validate/coverage_law_afib.py

33The coverage law does NOT transfer to heartbeats: direction survives, magnitude does notcardiac rhythmcardiac rhythmnegative

2026-08-05 · rung 3 — measured on real data, one modality

Claim

Asked of a medical signal — one cardiac patient's heartbeats, 5 rhythm classes — the coverage law fails both pre-registered bars. Recognition does rise with coverage (0.370 -> 0.428 -> 0.517 -> 0.517 over stores of 1/2/4/8 examples per class, against a 0.200 chance floor) but plateaus at roughly half the bar of 0.85, and the decode leg never rises at all (0.407/0.450/0.417/0.467). The direction of the law transfers; its magnitude is modality-dependent.

Evidence

3 seeds per coverage point, beat-disjoint store/query/judge splits, unblended 1-NN judge, the same protocol and criteria as the speech confirmation. Deciders track each other closely (vote 0.370/0.411/0.472/0.511 vs prototype 0.370/0.422/0.511/0.483) — no regime flip is visible here, unlike speech.

Scope / limits

One dataset, one patient, one encoder (the magnitude-weighted slope-bin representation already used for this data). This bounds the coverage law; it does not bound memory-based learning in general, and it says nothing about medical signals whose identity IS carried by temporal evolution.

Controls that could have killed it

  • identical pre-registered bars and protocol to the speech confirmation, frozen before running — the comparison is like-for-like
  • chance floor stated (0.200) so a weak-but-real effect is not read as failure
  • both deciders reported at every coverage point rather than the better one
This is COHERENT with the record's older domain nulls rather than a surprise: an earlier cross-domain test already found ECG a non-result for this metric (Fisher-Rao edging plain L2 by ~0.5 sigma) and named the confound — those domains ran a generic sliding-window frontend rather than a tuned one. The standing two-lane finding predicts it: the geometry earns its keep where identity lives in temporal EVOLUTION, and the identity of a single heartbeat is closer to a shape than to a motion. So the honest statement of the coverage law is now conditional: it holds where the metric holds. The methodological lesson is ours to keep — the prior null was in the record before this experiment was chosen, and reading it first would have predicted the outcome. Finding 34 asks the same organ the OTHER question and gets the opposite answer.

validate/coverage_law_ecg.py

32The decider is coverage-dependent, and blended evaluators suppress decode scoresarchitecturespeechactive

2026-08-04 · rung 3 — measured on real data, one modality

Claim

Which decision rule the loop should use depends on the memory regime. Across speakers (sparse, shifted), classification by consolidated barycentre prototypes beats the distance-weighted k-NN vote by 9.7 points mean over 3 seeds (0.672 vs 0.575; per-seed gains +10.5/+3.5/+15.0) — the loop as first measured in finding 28 was mis-wired. At dense personal coverage the ordering REVERSES: the raw vote wins (0.917/0.858/0.925) over prototypes (0.833/0.850/0.833), because the retrieved neighbourhood is the same speaker's own takes. Neither rule is 'the' decider; the loop should choose by regime.

Evidence

Same protocol and seeds as findings 28/31. Additionally: the decode evaluation itself is sensitive to finding 30's blending result — judging single-speaker responses against mixed-speaker barycentre prototypes suppressed the measured in-class rate (0.750 at t4) versus a 1-NN judge over raw clips (0.800 on the identical responses, registered before running as the final judge variant). Blending drifts artifacts off every real speaker's manifold, and that applies to yardsticks exactly as it applies to outputs.

Scope / limits

One dataset, one modality, 3 seeds. The regime boundary (where the ordering flips between prototype and vote) is not yet mapped — only its two endpoints are measured.

Controls that could have killed it

  • identical splits and seeds across both deciders, so the comparison isolates the decision rule
  • the judge comparison re-scores the SAME 40 decoded responses under both judges rather than resampling
Two practical rules fall out: consolidate for transfer, vote for personalization; and never put a blended artifact on either side of a decode evaluation. The second one bit this project's own pre-registered criterion before it was caught.

validate/closed_loop_confirm.py

31The personal condition: the loop closes end-to-end as a function of coverage, with zero trainingarchitecturespeechactive

2026-08-04 · rung 3 — measured on real data, one modality

Claim

When the store contains the speaker's own examples of each word — the personalization condition the architecture is for — the full retrieve-and-read-out loop crosses both pre-registered bars, on 3 seeds with full query sets: recognition 0.900 mean (0.917/0.858/0.925; bar 0.85) and decoded responses in-class 0.872 mean (0.867/0.850/0.900; bar 0.80, unblended 1-NN judge). Both legs rise monotonically with coverage (recognition 0.575 -> 0.800 -> ~0.9; decode 0.400 -> 0.700 -> ~0.87 as coverage goes none -> 2 -> 4 examples per voice-word pair), and every improvement costs exactly one append — no gradient step exists anywhere in the system.

Evidence

3 seeds x 120 queries each, store = 4 takes per (speaker, digit), queries = held-out takes of the same pairs, yardstick = a third disjoint slice. Mechanism measured, not argued: 0.964 of retrievals return a memory of the SAME speaker (per-seed 0.967/0.967/0.958) — the representation entangles who with what, which is simultaneously why the across-speaker condition fails (finding 30: right word, wrong voice) and why the personal condition works. The coverage curve's earlier points come from the seed-0 probes at sparse and 2-take coverage.

Scope / limits

One dataset (spoken digits), one modality; the across-speaker condition remains open (finding 30) and identity-preserving quotients are the proposed route, not a result. Decode bar was originally hit exactly (0.800) on a 40-query subsample; the full 120-query sets confirm above it.

Controls that could have killed it

  • take-disjoint splits: store, queries, and judge slice share no clips
  • the judge is 1-NN over raw held-out clips — no blended prototypes anywhere in the evaluator (see finding 32's judge note)
  • full query sets rather than the subsample that first hit the bar exactly; per-seed results reported individually
This is the architecture's core promise measured: accuracy grows with every example remembered, both legs cross their bars at 4 examples per word, and learning is memory formation alone. The open frontier is transfer to unseen speakers, which is a representation problem (disentangling voice from content), not a memory problem.

validate/closed_loop_confirm.py

30Composition is task-dependent: blending helps the vote and hurts the artifact; the decode gap is structuralarchitecturespeechactive

2026-08-04 · rung 3 — measured on real data, one modality

Claim

Three pre-registered follow-ups to finding 28's decode gap. (1) Blending hurts generation monotonically: reading out the single nearest memory verbatim lands in-class 0.475; the barycentre of the vote-winning cohort, 0.400; the consolidated all-members class prototype, 0.325 — the more memories are averaged into the artifact, the worse it gets. (2) The gap is structural, not data-starved: the best strategy swept at 5/10/20 examples per class gives 0.475/0.425/0.575 — no monotone rise, nowhere near the registered 0.70 bar. (3) At fixed shortlist width, index recall@5 degrades as the store grows (width 32: 0.980 -> 0.897 -> 0.797 over stores 50/100/200), yet the index arm's task accuracy is at or above the exact path at every width on the larger stores — the prefilter's misses are net-positive for the vote.

Evidence

Decode arms at 5 shots, 40 queries, identical yardstick as finding 28's L5 (held-out-speaker prototypes): nearest 0.475 / raw cohort 0.400 / consolidated 0.325. Shot sweep of the best arm: 0.475 (store 50) / 0.425 (100) / 0.575 (200). Index scaling: exact vote accuracy saturates across speakers at 0.483/0.500/0.483 while the store grows 4x; operating width (first width reproducing exact answers) 32 at store 50, 16 at store 100, none at store 200 — where every tested width scores ABOVE exact (0.550/0.517/0.517/0.500 vs 0.483), so the registered two-sided match criterion fails in the direction of improvement. Registered sublinear-width criterion: FAIL as stated.

Scope / limits

Single seed, one across-speaker split, 40-query decode subsample and 60-query index subsample — at 60 queries one answer = 1.7 points, so the 1-point match criterion effectively demands identical answers. Follow-up probes; any headline use gets the full 3-seed, 200-query treatment first.

Controls that could have killed it

  • all three decode strategies evaluated on identical queries against identical held-out prototypes, so the comparison isolates the composition operator
  • the same fixed exp(-d/median) weighting everywhere; exact retrieval in the decode arms so index quality is not confounded with composition quality
  • the shot sweep reuses the same base sampling per grid point rather than resampling, so store growth is the only variable
Together with finding 28's L1 (the cohort vote beats the single nearest for classification), this is an architecture rule, not a defect report: composition belongs on the recognition side of the loop — blend evidence to decide, read out the nearest memory to produce. The saturating exact accuracy (0.483 at both 50 and 200 entries) says the across-speaker ceiling is set by the speaker shift, not by memory size, consistent with the shot sweep. Provenance: two background runs were killed externally mid-flight; store-50/100 index rows are transcribed from the first run's line-buffered console output (same seed and protocol), the rest computed live by the resume-capable finisher — noted in the result JSON.

validate/closed_loop_scale.py

29Order-invariance holds exactly through the composed retrieval patharchitecturespeechactive

2026-08-03 · rung 3 — measured on real data, one modality

Claim

Building the identical 50-entry store all at once versus one class at a time produces identical per-class vote accuracy on every one of 10 classes — the end state provably does not depend on arrival order, through the full distance-weighted cohort vote, not just 1-NN recall.

Evidence

Per-class accuracies batch vs incremental: 0.350/0.450/0.300/0.150/0.550/0.900/0.400/0.950/0.800/0.550 — equal to the third decimal in all 10 cases. This is the non-degenerate control for finding 28's failed retention criterion: the first class's fall from 1.000 (alone in the store) to 0.350 (nine rivals present) is entirely rival arrival — task difficulty growing — and 0% forgetting.

Scope / limits

Deterministic retrieval over an append-only store makes this near-structural; the measurement confirms the composed vote path introduces no order dependence either. One dataset, one split, seed 0.

Controls that could have killed it

  • identical entry set both ways (same seed, same sampling), so the comparison isolates arrival order
  • evaluated through the same k=5 weighted vote as the main experiment, not through a simpler recall path
The earlier order-invariance result (finding 14's protocol) measured 1-NN recall; this extends it through cohort composition. It also converts finding 28's registered retention FAIL from an open question into a closed one: nothing was forgotten, the baseline was degenerate.

validate/closed_loop_batch_check.py

28The closed loop as a learner: retrieval and instant learning hold; the decode leg does not closearchitecturespeechactive

2026-08-03 · rung 3 — measured on real data, one modality

Claim

The full retrieve-compose-decode loop, run end-to-end with pre-registered criteria on across-speaker spoken digits (5 shots/class): a distance-weighted 5-NN vote beats the single nearest memory by +2.7 points (positive on 3/3 seeds); a random-cohort control collapses to 0.175, so the vote is carried by which memories are retrieved; every class is recognisable the instant its five examples are appended, with zero gradient steps. The generation leg is the measured frontier: response trajectories composed from the vote-winning cohort land in-class only 0.400 of the time — below the vote's own accuracy.

Evidence

Vote 0.575 vs 1-NN 0.548 (mean of seeds 0/1/2; smallest seed gain +0.5pt — the effect is real but modest). Random cohorts weighted by their real distances: 0.175. Instant learning: several classes at 1.000 immediately after append. Composed decode: 16/40 in-class against held-out-speaker prototypes. Two additional criteria FAILED AS REGISTERED and are kept: a retention criterion whose baseline was degenerate (measured the first class when it was the only class — a one-label classifier scores 1.000 for free; the drop it 'detected' is label-space growth, not damage — see finding 29 for the non-degenerate control), and an index-match criterion missed by literal floating-point epsilon (the index changed exactly 2 answers in 200 at shortlist 32, the registered 1-point boundary).

Scope / limits

One dataset, one across-speaker split, 3 seeds for the composition arms and 1 for the rest. Recognition through the loop is the claim; generation through the loop is explicitly NOT closed at this shot count and split.

Controls that could have killed it

  • random-cohort control drawing random candidates but weighting them with their REAL distances (the corrected control form; a flat-weight random control would conflate cohort choice with weighting)
  • one fixed temperature rule for every weighted arm — weights = exp(-d/median(d)) — so no arm gets a peakier weighting than another (the per-arm temperature confound documented in the retrieval findings)
  • exact (exhaustive) retrieval in the composition arms, so composition quality is measured separately from index quality
The pre-registered retention criterion failing because of its own baseline design is reported as what it is: a registration error, kept visible for the same reason retractions are. The decode gap is the working frontier — follow-up arms (consolidated-prototype decode, shot sweep, index scaling) are the next entries.

validate/closed_loop.py

27The equation predicts the next step better than a trained network, with no trainingfoundationsspeechhandwritingactive

2026-07-24 · rung 5 — survives a fair control designed to kill it

Claim

Continuing the curve along its geodesic — a closed-form step with zero learned parameters — predicts the next frame of a real signal 3-16x more accurately than a matched dense network that was trained to do it.

Evidence

Fisher-Rao error, lower is better. AUDIO (FSDD cochlea, F=32): geodesic continuation 0.0054 vs fair dense 0.0860 vs naive repeat-last 0.0866. PEN (handwriting, F=16): 0.1796 vs 0.3863 vs 0.3903. Paired: audio 646 wins / 2 losses, p=4e-190; pen 591/129, p=2e-71. The LEARNED geometry model beats fair dense on audio (p=2e-122) but only TIES on pen (397/353, p=0.12) — reported because it qualifies the headline: learning helps complex dynamics, not smooth real signals.

Scope / limits

Next-step prediction on continuous real sequences, two unrelated modalities. This is one-step continuation, NOT generation over long horizons, and not language.

Controls that could have killed it

  • THE FAIR-ANCHORED BASELINE, added after a red flag rather than before. The first dense baseline scored WORSE than naive repeat-last on audio, which is not a believable result for a trained net. The cause: it predicted an absolute frame while the geometry methods predict a STEP FROM THE LAST FRAME. `std_res` was rebuilt anchored the same way, which roughly halved its error (0.176 -> 0.086) — and the geometry still wins. The control strengthened the opponent and the result survived.
  • naive repeat-last floor — the fair dense net barely beats it on audio, the geometry beats it 16x
  • matched parameter budget across arms
  • paired per-sequence statistics rather than a difference of means
This is the result that separates PREDICTING from RECOGNISING, and it is the strongest form of the project's central claim: the prediction is not learned, it falls out of the geometry. Note what it does NOT say — one step ahead on smooth continuous signals is a long way from language, where the same continuation currently costs 16% perplexity against an ordinary recurrent net (finding 24's successor).

validate/headtohead.py

26The learned character metric encodes linguistic class, and it is not frequencyfoundationstextactive

2026-07-26 · rung 3 — measured on real data, one modality

Claim

Trained only to predict the next character, the model places characters on the simplex so that Fisher-Rao distance tracks linguistic class -- case pairs, vowels, sentence terminators, whitespace -- and this is not explained by how often the characters occur.

Evidence

Off-diagonal Fisher-Rao spread 0.5050 (min 0.1298, max 0.6349) against exactly 0.0000 for one-hot. Correlation with frequency is nil: Spearman +0.009, ~0.0% of rank variance. After residualising distance on the frequency gap, within-class pairs are closer than cross-class pairs by 0.0226; a 10,000-shuffle permutation null over class labels gives z=+4.46, p=0.00010. Tightest classes are the terminators . ? ! (-0.256) and separators , ; : (-0.212).

Scope / limits

One trained checkpoint (HODOS_embed, 30k steps, val ppl 5.29) on TinyShakespeare. Learned embeddings clustering by character class is a KNOWN property of neural language models and is not claimed as novel here. It is recorded because it is the precondition that makes finding 25 interpretable: without it there is no metric for the Fisher-Rao ground to act on.

Controls that could have killed it

  • FREQUENCY CONTROL -- the obvious confound, since vowels and whitespace are all common. Distance is residualised on the log-frequency gap before the class test.
  • PERMUTATION NULL -- 10,000 shuffles of the class labels, so 'looks clustered' cannot be eyeballed into a result.
  • INTERNAL NEGATIVE CONTROL -- the grab-bag 'other' category (digits, quotes, $, &) is not a linguistic class, and it is the ONE category with a POSITIVE residual (+0.117): its members are FURTHER apart than frequency predicts. The non-class is the one that does not cluster.
This is what the 4,225 parameters buy, and it is worth 7x perplexity on its own: the same architecture with characters pinned to one-hot corners scores 37.63 against 5.29 with the positions learned, at matched parameters and identical steps.

validate/char_metric_structure.py

25The Fisher-Rao ground IS a better retrieval key than L2 -- once characters have a learned metricfoundationstextactive

2026-07-26 · rung 5 — survives a fair control designed to kill it

Claim

Given character positions that carry a real metric, retrieving past contexts under the Hodos process distance predicts the next character better than retrieving under plain L2 over the identical frames. The only difference between the two arms is the distance function.

Evidence

Pooled over 5 seeds, 1000 held-out positions, each seed tuning its own temperature and blend weight on its own separate half: Hodos beats L2 by +0.0970 nats, p=2.0e-05. The warp HELPS by 0.0674 nats (p=2.5e-03) and the ground helps by 0.0296 nats (p=0.024) -- both reversing the retracted findings 17 and 24. The same experiment run in the old one-hot space reads -0.0249 nats, p=0.25: indistinguishable from zero, as a constant ground must be.

Scope / limits

k-NN character language modelling on TinyShakespeare against a unigram term. The learned metric is the 4,225-parameter embedding table from the trained HODOS_embed language model, so this shows the geometry works ON a learned metric -- it does not show the metric can be obtained without learning one.

Controls that could have killed it

  • ONE-HOT CONTROL ARM AS A MACHINE CHECK -- pre-registered to read zero, because the ground is provably constant there. It read -0.0249, p=0.25. Had it moved, the run was to be discarded rather than explained.
  • TUNE/TEST SPLIT -- each arm is swept over 9 temperatures x 7 blend weights = 63 settings; selecting the best of 63 on the same positions the verdict is read from would inflate the noisiest arm most. Selection happens on one half, every statistic is reported on the other.
  • REPRESENTATION ATTRIBUTION -- the learned edge exceeds the one-hot edge by +0.1219 nats, p=2.7e-07, position-paired within seed, so the effect is the representation and not the arm.
  • NO PREFILTER -- the exact distance is computed against the entire store. Learned frames are near-uniform (row entropy 4.133 of a 4.174 maximum), so a cheap signature shortlist would have been noise-dominated and would have silently decided the comparison.
  • BATCHED SOFT-DTW ASSERTED EQUAL to hodos.warp.soft_dtw before any result is produced (max abs err 6.6e-15). The first attempt disagreed at 2.4e-06, traced to float32 softmax rows that ground.bhattacharyya renormalises internally and this path did not.
  • random-retrieval floor (every trip of it was non-significant, all p > 0.17)
  • per-arm temperature tuning
This is a CORRECTION of findings 17 and 24, not a new capability claim. It does not make the language model competitive -- the geodesic head still costs 16% perplexity against a matched GRU (5.29 vs 4.56, 30k steps, 4.07M params). What it establishes is narrower and was previously mis-measured: where the geometry has a metric to act on, it acts, and the alignment machinery earns its place rather than costing.

frontier/lm_memory_pooled.py · frontier/lm_memory_pooled_result.json · supersedes finding [17, 24]

24Memory helps a language model, but the geometry is not whyfoundationstextretracted — negative claim

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'Retrieval is what helps; the geometric distance is not better than a flat one for this.'

Evidence

Memory improves by +0.105 nats, p=2.7e-06. Plain L2 retrieval beats Hodos retrieval by 0.121 nats, p=7.35e-04, and this survives per-arm temperature tuning.

Scope / limits

RETRACTED 2026-07-26. The geometry was never given a metric to work with.

Controls that could have killed it

  • plain-L2 arm
  • random-retrieval floor
  • per-arm temperature tuning

Why this was wrong

We claimed this did NOT work, or that it was a limit of the method. Our own later measurement showed otherwise, so the claim was withdrawn. The correction goes in the method's favour — it was recorded against ourselves and was wrong.

Both of these were measured in a space where the Fisher-Rao ground is a CONSTANT FUNCTION. Contexts were built as label-smoothed one-hot rows, and on the actual vocabulary (V=65, eps=0.03) the Fisher-Rao distance between every pair of distinct characters is identical: min 2.9971, mean 2.9971, max 2.9971 -- spread exactly 0.0000. In that space Fisher-Rao and L2 induce the SAME ranking, so 'the ground ties L2' was an algebraic identity, not a measurement, and the warp was the only thing left that could move. The experiment could not have produced any other answer. Re-measured in a non-degenerate metric (finding 25) both conclusions reverse: the ground beats L2 and the warp HELPS. The H1 half of this finding -- that memory helps a language model at all -- survives and is carried into finding 25. Only the control's verdict on the geometry is withdrawn.

Reported as it came out. An H1 pass alone would have been enough to claim a Hodos win; the control is what prevented that.

frontier/lm_memory.py · superseded by finding 25

23Character spacing is not Riemann level repulsionfoundationszeta zerostextnegative

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

The apparent 'empty band' in character separations is an artifact of one-hot encoding, not a repulsion signature.

Evidence

One-hot characters have exactly ONE distinct pairwise distance — all 2080 pairs at 1.498555, spread 0.00e+00. Arbitrary points on the SAME simplex give 1334 distinct distances, so the constraint is the encoding.

Scope / limits

The structural question is genuinely shared — both ask what a distribution of separations looks like near zero. But level repulsion is a statement about a DENSE 1-D spectrum, and 65 points in 65 dimensions is the sparsest possible arrangement.

Controls that could have killed it

  • coordinate-shuffle matched null
  • the repulsion statistic turned out to have NO dynamic range — a dead measurement, not a null
Reported as a failed measurement rather than a null result. A z of 0 against a null with zero spread means the instrument could not move.

validate/char_repulsion_null.py

22The Riemann work reproduces known mathematics and proves nothing newfoundationszeta zerosnegative

2026-07-22 · rung 5 — survives a fair control designed to kill it

Claim

The distance works as a REFEREE — it correctly ranks a quantum-chaotic drum closest to the real Riemann zeros — but every attempt to use it to find a load-bearing arithmetic operator returned a null.

Evidence

Referee ranks GUE 0.0435 closest, GOE 0.0695 second, noise far. Reproduces Montgomery–Odlyzko (known). The relationship-kernel looked load-bearing at 4–6σ against random integers — but a density-matched control scored BETTER (0.062 vs 0.128), so the signal was a density-envelope artifact.

Scope / limits

RH is not proven, not approached. The real limiter is resolution: an arithmetic fingerprint needs ~10³–10⁴ zeros and we have hundreds.

Controls that could have killed it

  • density-matched sharp control — this is what killed the 4σ result
Catching the artifact via the fair control IS the result. Kept as an honest appendix, not a claim.

validate/drum_referee.py

21An impossible value exposed a flaw that only a second dataset could revealfoundationshandwritingactive

2026-07-26 · rung 5 — survives a fair control designed to kill it

Claim

A negative forgetting score — sequential beating joint — is impossible if the store is order-invariant, and chasing that impossibility found a real defect in the comparison.

Evidence

Handwriting reported hodos forgetting −0.083 on one seed. Cause: an incomplete final chunk is dropped from the task list, but the joint comparison arms still iterated ALL classes — so joint faced a 17-way problem while sequential faced 16-way.

Scope / limits

Speech divides evenly into pairs, so nothing was orphaned there and the flaw could ONLY surface on a second modality.

After the fix the net's forgetting ROSE from 0.692 to 0.801 — the bug had been FLATTERING the opponent. Replication on a second modality is not just confirmation of a number; it is a different defect detector.

validate/catastrophic_forgetting.py

20Measurement error: a growing label space looked like memory decayingfoundationsspeechhandwritingretracted — negative claim

2026-07-26 · rung 2 — measured once, synthetic data

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'the geometry forgets too — task-1 accuracy drops 0.500 as later tasks arrive.'

Evidence

The raw drop, which is real but means something else.

Scope / limits

RETRACTED same day. This is a flaw in the METRIC, not in either system.

Controls that could have killed it

  • comparison against the joint-trained ceiling

Why this was wrong

We claimed this did NOT work, or that it was a limit of the method. Our own later measurement showed otherwise, so the claim was withdrawn. The correction goes in the method's favour — it was recorded against ourselves and was wrong.

In class-incremental learning the LABEL SPACE GROWS: task 1 is a 2-way choice when first learned and a 10-way choice at the end. Comparing those two numbers makes the task getting harder look like memory degrading. Measuring against the joint-trained ceiling makes the growing label space cancel, leaving only the cost of having learned in sequence — which for this system is exactly zero. The file's own docstring warned about this trap before the first version was written that way anyway.

validate/catastrophic_forgetting.py

19The geometry is not fp16-safefoundationstextretracted — negative claim

2026-07-24 · rung 1 — reasoned, unmeasured

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'sph_log takes arccos of an inner product very close to 1.0; fp16 resolution there is coarser than the guard clamp, so every Log map collapses to zero and the geometry dies silently.'

Evidence

None — this was reasoned, never measured. Rung 1.

Scope / limits

RETRACTED 2026-07-24 by direct measurement on a T4.

Controls that could have killed it

  • real-text measurement
  • naive-autocast dtype check

Why this was wrong

We claimed this did NOT work, or that it was a limit of the method. Our own later measurement showed otherwise, so the claim was withdrawn. The correction goes in the method's favour — it was recorded against ourselves and was wrong.

Both halves were wrong. The premise is BACKWARDS: characters are label-smoothed one-hots, so two DIFFERENT consecutive characters are nearly ORTHOGONAL (real-text mean inner product 0.096), nowhere near 1.0. The fp16 danger band contains 0.0000 of real pairs and zero pairs lose real motion. And naive autocast never reaches the geometry anyway — the manifold tensors come out float32. The 'safety' code written to guard this imaginary hazard then introduced the only real bug the run found.

The clearest example in this project of why reasoning is not evidence.

frontier/gpu_check/gpu_validate.py

18Hierarchy does not beat a flat representationfoundationsspeechhandwritingnegative

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

Lifting a trajectory into nested levels gives no advantage over the flat version on either modality.

Evidence

Speech: flat 0.920, lifted 0.935 — but the SMOOTHING control also 0.935, identical. Handwriting: lifted 0.870 < flat 0.877.

Scope / limits

FSDD utterances are 24 frames, so a lift yields ~5 — shallow. A null here means 'no benefit on short sequences', NOT 'no benefit'. The case hierarchy exists for is long sequences.

Controls that could have killed it

  • bin-wise smoothing control — MATCHED it exactly, which is what killed it
  • length-matched subsampling control
A clean kill by a control built specifically so it could kill the idea.

kaggle_pack/hodos_science.py

17On discrete aligned data the warp HURTS and the ground is neutralfoundationstextretracted — negative claim

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'For character contexts, the Fisher-Rao ground ties plain L2 exactly, and the entire performance gap is the alignment machinery.'

Evidence

Per-arm tuned: full Hodos 3.1117, ground-without-warp 2.9949, L2 2.9903. Ground vs L2: −0.0046, p=0.586 — a dead tie. The warp costs +0.1168, the whole gap.

Scope / limits

RETRACTED 2026-07-26. The tie was forced by the representation, not observed.

Controls that could have killed it

  • plain-L2 retrieval arm
  • random-retrieval floor — hurts at every setting, proving the gain is real retrieval not smoothing
  • PER-ARM temperature tuning — a single temperature was unfair because the distances scale differently (FR sums linearly in mismatches, L2 as a square root); best temperatures differed by 40×

Why this was wrong

We claimed this did NOT work, or that it was a limit of the method. Our own later measurement showed otherwise, so the claim was withdrawn. The correction goes in the method's favour — it was recorded against ourselves and was wrong.

Both of these were measured in a space where the Fisher-Rao ground is a CONSTANT FUNCTION. Contexts were built as label-smoothed one-hot rows, and on the actual vocabulary (V=65, eps=0.03) the Fisher-Rao distance between every pair of distinct characters is identical: min 2.9971, mean 2.9971, max 2.9971 -- spread exactly 0.0000. In that space Fisher-Rao and L2 induce the SAME ranking, so 'the ground ties L2' was an algebraic identity, not a measurement, and the warp was the only thing left that could move. The experiment could not have produced any other answer. Re-measured in a non-degenerate metric (finding 25) both conclusions reverse: the ground beats L2 and the warp HELPS.

This is a MECHANISM for the demarcation rather than an observation: rigid aligned windows have no timing to align away, so warping can only manufacture false matches. Actionable: on discrete aligned data, use the ground WITHOUT the warp.

frontier/lm_memory_temp.py · frontier/lm_memory_temp_result.json · superseded by finding 25

16Anything between the landed point and the answer must start as a pass-throughfoundationstextactive

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

A randomly-initialised readout layer costs ~21 perplexity — a 2.5× degradation — regardless of its form. Initialised as a pass-through, the cost vanishes entirely.

Evidence

At identical dimension and parameters: no readout 13.528, random-init readout 34.318, identity-init readout 13.187. Both readout FORMS cost the same, so it is not which readout but having one.

Scope / limits

Applies to anything built downstream of a landed point — the memory, reasoning and decoder organs all sit there.

Controls that could have killed it

  • mixture vs logit readout — identical, proving form is not the variable
The same disease the distance-attention layer had in June, and the same cure: start as identity so a layer can only refine, never replace.

frontier/hodos_lm.py

15A compressed latent representation does NOT helpfoundationstextnegative

2026-07-25 · rung 5 — survives a fair control designed to kill it

Claim

Squeezing characters onto a lower-dimensional concept simplex is monotonically worse, at matched parameters, with a working readout.

Evidence

13.545 (65 dims) → 16.089 (32) → 19.418 (16) → 23.113 (8) perplexity. No sweet spot.

Scope / limits

Character language modelling, small local config.

Controls that could have killed it

  • parameter matching to within 1% — without it, a small-K model is simply a SMALLER model
  • sharp-readout arm to separate the readout from the dimension
The first version of this measured through a broken readout (see finding 16) and reached the right conclusion for the wrong reason.

validate/lm_latent_sweep.py

14It does not forget — and cannot, by constructionfoundationsspeechhandwritingactive

2026-07-26 · rung 5 — survives a fair control designed to kill it

Claim

Taught classes in sequence, a standard network loses half to four-fifths of its first-task accuracy. This system loses nothing, because learning is appending to a store rather than overwriting weights.

Evidence

Forgetting measured against the joint-trained ceiling so the growing label space cancels. Speech: net +0.500, hodos +0.000. Handwriting: net +0.801, hodos +0.000. Per-seed hodos [0,0,0] on both. Learning INCREMENTALLY beats the network trained on EVERYTHING AT ONCE: 0.544 vs 0.422 (speech), 0.845 vs 0.702 (handwriting).

Scope / limits

Two modalities, 3 seeds each. Handwriting uses a random per-class holdout (no speaker axis), a weaker split than the across-speaker speech version.

Controls that could have killed it

  • joint-trained ceiling — without it, task difficulty masquerades as forgetting
  • replay (full retrain each step) — also 0.000 forgetting, but costs a retrain every step and still lands lower
  • ORDER INVARIANCE: building the store in REVERSE gives byte-identical answers, 1.000 on every seed
Order invariance at 1.000 is what makes this structural rather than lucky: the final state provably cannot depend on the sequence, so forgetting is impossible rather than merely absent. Two of our own errors were caught here — see findings 20 and 21.

validate/catastrophic_forgetting.py · frontier/forgetting_result.json

13The geometry needs 5–10× less data than a standard networkfoundationsspeechhandwritingactive

2026-07-25 · rung 4 — replicated on a second, unrelated modality

Claim

Class-barycenter prototypes with NO training reach a standard method's 50-example accuracy using 5 examples (speech) or 10 (handwriting).

Evidence

Speech across-speaker: hodos 0.416/0.539/0.654/0.715/0.746 at 1/2/5/10/50 shots vs GRU 0.189/0.291/0.417/0.520/0.624. Handwriting: hodos 0.730→0.894 vs GRU 0.403→0.828. At ONE example the network is near chance (0.189 vs 0.10).

Scope / limits

Two modalities, 5 seeds, 200 test items, full shot sweep to 50.

Controls that could have killed it

  • 1-NN L2
  • mean-pool + logistic regression
  • GRU classifier
  • all arms on identical splits and shots
A pre-registered criterion predicted the gap would SHRINK with data (the signature of a prior). On handwriting it does, monotonically (+0.098→+0.018). On speech it RISES to a peak at 10 shots then falls — the local 5-shot sweep had only seen the rising half. Describing it as a 'prior' without that check would have been wrong.

kaggle_pack/hodos_science.py · kaggle_pack/out/hodos_science_result.json

12The win decomposes: trajectory > ground > warpfoundationsspeechhandwritingactive

2026-07-24 · rung 4 — replicated on a second, unrelated modality

Claim

The advantage over mean-pooling splits cleanly into three ingredients in a fixed order: using the trajectory at all, then the Fisher–Rao ground, then the alignment.

Evidence

FSDD across-speaker 5-shot, 3 seeds: FR+warp 0.657 / FR+nowarp 0.637 / L2raw+warp 0.602 / L2raw+nowarp 0.573 / mean-pool 0.483. Trajectory +9, ground +5–6, warp +2–3. Handwriting: same ordering, +25 / +3.4 / +1.8.

Scope / limits

Same ordering on a sound task AND a non-sound task, so it is not a speech artifact.

Controls that could have killed it

  • L2-on-RAW-frames arm — the earlier L2-on-sqrt arm was a BAD control (≈ Hellinger, same geometry) and produced a spurious 'FR doesn't help'

validate/nn_decompose.py

11Against a published method: the field's warp is better, our ground still winsfoundationsspeechhandwritingactive

2026-07-23 · rung 6 — survives comparison against a published method

Claim

soft-DTW beats our hard-DTW warp — but the Fisher–Rao ground beats Euclidean under BOTH warps on BOTH modalities. The transferable contribution is the ground, and it composes with the field's better alignment.

Evidence

Handwriting z=2.75, errors 0.045→0.025 for soft-DTW over hard-DTW.

Scope / limits

Recommended distance is therefore soft-DTW + Fisher–Rao — adopt the field's warp, keep our geometry.

Controls that could have killed it

  • both warps × both grounds × both modalities
The claim survived contact with the field and got narrower and stronger.

validate/abx_vs_softdtw.py

10Doppler / velocity measurement is not an edgefoundationssonarnegative

2026-07-22 · rung 5 — survives a fair control designed to kill it

Claim

The conceptual link (Doppler is a time-warp, so the method should measure velocity) is real and elegant, but empirically purpose-built DSP measures it better.

Evidence

Matched filter is perfect (0.000 clean); Hodos ~0.13. On the velocity PROFILE — supposedly the structural edge — the Hodos warp-slope is too noisy (0.14–0.20) and loses even to a FLAT single-α baseline (0.07–0.12).

Scope / limits

Two fair benchmarks. Stopped after two rather than trying a third variant.

Controls that could have killed it

  • matched-filter baseline
  • flat-α baseline

validate/doppler_benchmark.py

9The min-mean distance 'win' was a padding degeneracyfoundationslightsonarretracted — positive claim

2026-07-22 · rung 5 — survives a fair control designed to kill it

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'min-MEAN + Fisher–Rao beats standard DTW on light by +0.21 (0.61→0.82), and it survives 6→12/class.'

Evidence

The apparent win.

Scope / limits

RETRACTED — the metric tracked one condition better than the other with no mechanism to explain it.

Controls that could have killed it

  • mechanism investigation, scratchpad/diag_minmean.py

Why this was wrong

We claimed this WORKED. Our own later measurement showed it did not, so the claim was withdrawn. The correction goes against the method.

min-mean degenerates into long, padded, cheap-repeat paths (path padding 0.00 for standard DTW vs 0.61 for min-mean) which makes it behave MORE like plain averaging (correlation with mean-pool 0.779 → 0.845). Light is precisely the task where averaging wins — so it 'won' for the wrong reason, and the same collapse HURT on echo where timing is the signal. The object is Marzal–Vidal's Normalized Edit Distance (1993), proven non-metric. RETIRED.

An unexplained win-here/lose-there is a red flag, not a result. Explain the mechanism BEFORE banking a win.

hodos/warp.py

8The distinctive ground does NOT generalise across light, echo and turbulencefoundationslightsonarturbulencenegative

2026-07-22 · rung 5 — survives a fair control designed to kill it

Claim

The Fisher–Rao ground shows no edge over plain L2 on light spectra, sonar range-profiles or turbulence. It ties; it does not beat.

Evidence

1-NN LOO, 3 seeds, ground ablation, noise swept to break the ceiling. At 18/class: light 0.45 vs 0.44, echo 0.73 vs 0.75, turbulence 0.62 vs 0.68. The ground earns its keep in 0 of 3 modalities.

Scope / limits

Recognition only. Does not test barycenter, morph, pooling or the physics reproduction.

Controls that could have killed it

  • 18/class robustness re-check after a 6/class small-sample artifact reversed the apparent result
This killed the universal-distinctiveness claim. It is the reason the paper's claim is narrow.

validate/crossmodal_benchmark.py

7Pitch invariance by construction, not by trainingfoundationsspeechactive

2026-07-22 · rung 5 — survives a fair control designed to kill it

Claim

Quotienting the process space by the global bin-shift group gives a distance between ORBITS, making recognition invariant to transposition exactly — not approximately, and with no examples required.

Evidence

Same vowel shifted 6 bins: distance collapses 61× (0.4611→0.0075) while a DIFFERENT vowel is preserved (0.2911) — a 39× separation. The invariant recogniser REPAIRS a real base-engine misclassification.

Scope / limits

Exact-zero only for an exact array roll; independently synthesised transpositions carry edge-tail truncation (D_inv ≈ 7.5e-3, not 1e-16). Pad-mode invariance is approximate.

Controls that could have killed it

  • theorem verified |ΔD| = 4.4e-16
  • W1 claim self-corrected mid-build
Shipped into Auris, closing that project's own documented pitch-sensitivity limit.

hodos/quotient.py

6The corrected boundary: two lanes, not onefoundationshuman activityactive

2026-07-24 · rung 3 — measured on real data, one modality

Claim

The GROUND earns its keep wherever identity lives in the relative-space pattern — INCLUDING static posture. The WARP earns its keep only where identity lives in temporal motion. These are separate lanes; the earlier single-axis 'motion vs static' framing conflated them.

Evidence

With a representation that preserves relative space (distribution over TOTAL acceleration orientation), subject-disjoint sit-vs-stand = 0.758 against chance 0.500. Fisher–Rao ground contributes +5.8 points over Euclidean on a task with NO motion; the warp contributes 0, correctly.

Scope / limits

One dataset (UCI HAR). The two-lane structure is supported by findings 3, 4 and 17 jointly. PRIOR ART, added 2026-07-26: the same demarcation is already stated in the pooling literature — orderless pooling suffices for stationary or weakly nonstationary signals with little discriminative temporal structure, while order-aware pooling is needed where the cues are transient or phase-dependent. This finding independently REPRODUCES a known boundary; it does not establish a new one. It stays active because it is measured and it is true.

Controls that could have killed it

  • subject-disjoint split

validate/nn_har_orientation.py · supersedes finding 5

5The method fails where identity is static — the predicted boundaryfoundationshuman activityretracted — negative claim

2026-07-23 · rung 3 — measured on real data, one modality

Claim

ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'Fisher–Rao does not beat Euclidean on human-activity windows because activity identity in a short window is static — exactly what the principle predicts.'

Evidence

HAR: averaging won, 0.24 vs 0.26–0.27.

Scope / limits

RETRACTED 2026-07-24. The null was OUR OWN BROKEN INPUT, not a property of the method.

Why this was wrong

We claimed this did NOT work, or that it was a limit of the method. Our own later measurement showed otherwise, so the claim was withdrawn. The correction goes in the method's favour — it was recorded against ourselves and was wrong.

The adapter fed gravity-removed acceleration reduced to a MAGNITUDE — a representation with direction deleted. Sitting and standing are both 'at rest, no direction', so under that input they were identical BY CONSTRUCTION. The method was not failing on static identity; it was succeeding on an input from which the answer had already been removed. A null produced by a crippled representation is not a boundary.

validate/nn_har_orientation.py · superseded by finding 6

4The same win replicates on handwriting, with no audio at allfoundationshandwritingactive

2026-07-23 · rung 4 — replicated on a second, unrelated modality

Claim

The identical claim holds on pen motion — a completely different modality — and more strongly than on speech.

Evidence

UCI Character Trajectories, 2858 real handwritten characters, 4000 triples: paired McNemar z≈4.95 (Fisher–Rao fixes 49 errors, breaks 10).

Scope / limits

Single-writer temporal recognition, not an across-body handwriting split. Not a published leaderboard.

Controls that could have killed it

  • paired McNemar
  • fresh seed
This is what took finding 3 from 'a sound trick' to a modality-independent claim. The method was never about sound.

validate/abx_pen.py

3The information-geometry ground beats Euclidean on real speech across voicesfoundationsspeechactive

2026-07-23 · rung 3 — measured on real data, one modality

Claim

Recognising the same spoken word through a DIFFERENT human voice, the Fisher–Rao ground reduces error ~17% vs Euclidean.

Evidence

FSDD, 2000 triples, fresh seed: 0.1500±0.0080 vs 0.1805±0.0086 (~2.6σ).

Scope / limits

Word-level ABX on FSDD, not the published phone-level ZeroSpeech leaderboard. Within-speaker is a TIE — the N=400 within-speaker 'win' did NOT replicate and was a small-sample mirage.

Controls that could have killed it

  • confirm-at-scale with a fresh seed
  • binomial standard errors

validate/abx_fsdd.py

2Time-averaging is quadratically blind to a brief eventfoundationsmathematicsprovisional

2026-07-21 · rung 2 — measured once, synthetic data

Claim

Averaging a signal over time loses a short event at O(ρ²) while the curve distance loses it at Θ(ρ) — so pooling destroys exactly what a process comparison keeps.

Evidence

Log-log slopes 1.90 vs 1.00.

Scope / limits

Holds under an interior/support condition (corrected on review, 2026-07-22). A novelty check found the underlying effect is KNOWN prior art in detection theory and the DTW-averaging literature — the exponent statement may be a small novel instance at most. Do NOT claim as a new theorem.

Controls that could have killed it

  • novelty check vs literature, 2026-07-22

hodos/

1A process is literally a curve on a spherefoundationsmathematicsactive

2026-07-21 · rung 5 — survives a fair control designed to kill it

Claim

The map p ↦ √p sends the simplex onto the unit sphere, under which Fisher–Rao becomes ordinary great-circle geometry. A signal over time is therefore a curve on a sphere, and comparing signals is comparing curves.

Evidence

2·geodesic == fisher_rao to 1e-9; chi² ≈ ¼·d_FR² with correlation 0.987 on real audio frames.

Scope / limits

Mathematical identity, not an empirical claim. Prior art (information geometry); reused, not invented.

hodos/sphere.py

Every equation, with an honest label on each

Every equation the project rests on, with an honest label on each: ASSEMBLED from known parts, KNOWN and merely reused, or OURS. Read ASSEMBLED carefully — it is not a lesser category. Nearly every method in this field is a composition, and §1, the Hodos distance itself, is ASSEMBLED: the contribution is which ground it is composed with, and that composition is not something we found elsewhere. What the labels buy is that nothing here can be dismissed by pointing at a part of it.

A signal becomes a process P = (pⁱ, …, pᵀ), each frame a probability distribution over F bins. Δ^{F−1} is the probability simplex. BC(p,q) = Σᵇ √(pᵇ qᵇ) is the Bhattacharyya coefficient.

§1 The master distance

ASSEMBLED — the composition IS the contribution

D(P,Q) = ( min_{π ∈ Π_β}  Σ_{(i,j)∈π}  g(p⁽ⁱ⁾, q⁽ʲ⁾) )  /  |π★|

Plain: Line the two signals up in time so the total per-step difference is smallest, then divide by the length of that alignment.

Provenance: Dynamic Time Warping (Sakoe–Chiba 1978) composed with an information-geometry ground g. The warp is standard and reused; the ground is classical; THE COMPOSITION IS THE CONTRIBUTION — an elastic distance whose local cost is the exact Fisher–Rao geodesic between distributions, with closed-form geodesic operators on top. A literature check on 2026-07-26 found the neighbouring families (OTW, TiOT, TAOT, Riemannian Time Warping) but not this composition as a working engine with these operators. D is an alignment-tolerant divergence, NOT a metric; ordinary DTW fails the triangle inequality.

Correction: An earlier version of this sheet presented the minimum-MEAN alignment as 'the distinctive object'. That is RETIRED: minimising the mean directly is the Normalized Edit Distance (Marzal–Vidal 1993), whose optimum pads the path with cheap repeats and collapses toward plain averaging. Its one apparent benchmark win was that padding artifact (finding 9).

§2 The ground metrics, frame against frame

KNOWN

chi2(p,q)       = ½ Σᵇ (pᵇ − qᵇ)² / (pᵇ + qᵇ)
fisher_rao(p,q) = 2·arccos( BC(p,q) )
hellinger(p,q)  = √( 1 − BC(p,q) )
wasserstein1    = Σᵇ | CDF_p(b) − CDF_q(b) |

Plain: Five interchangeable ways to measure the difference between two single frames. All closed-form, none learned.

Provenance: All established — Le Cam / Topsøe, Rao 1945, Amari, Bhattacharyya, optimal transport. We reuse them; we do not claim them. Our only move here is making the ground pluggable in one engine.

§3 The sphere lemma — a process is literally a curve on a sphere

KNOWN, load-bearing framing

φ(p) = √p          maps the simplex onto the positive orthant of S^{F−1}
‖√p‖² = Σᵇ pᵇ = 1
geodesic(p,q) = arccos( BC(p,q) )
d_FR(p,q) = 2·arccos( BC(p,q) )

Plain: Taking the square root of a distribution puts it on a sphere, and the distance between distributions becomes the plain great-circle angle. So a signal over time is a curve on a sphere, and comparing signals is comparing curves.

Provenance: Classical information geometry (Rao 1945; Amari). Known — used as the exact geometric home of the engine. It is what makes the closed-form geodesics and means below possible.

§4 Geodesic operations on the sphere

KNOWN math, our operators

slerp(p,q,t)  = unembed[ ( sin((1−t)θ)·√p + sin(tθ)·√q ) / sin θ ]
karcher_mean({pᵢ}) = the point m minimising Σᵢ d_FR(m, pᵢ)²

Plain: How to walk between two distributions along the shortest path, and how to average a set of them without leaving the space.

Provenance: Standard Riemannian geometry and DTW-barycenter averaging (Petitjean 2011). Known. These drive the morph and barycenter operators, and the zero-training prototypes behind findings 13 and 27.

§5 The temporal-pooling information-loss theorem

ASSEMBLED — downgraded from OURS on 2026-07-26

D(A,B)   = Θ( ρ · g(s,s′) )        — LINEAR in ρ
g(ā, b̄) = O(ρ²)   under the interior/support condition
         = O(ρ)    otherwise — the honest boundary

so  D / g(ā,b̄) = Θ(1/ρ) → ∞

Plain: Averaging a signal over time is QUADRATICALLY blind to a brief event riding on an ongoing one — the final /t/ against /p/ on a sustained vowel — while the warp distance stays LINEARLY sensitive. Averaging can only lose; on brief transients it loses badly.

Provenance: DOWNGRADED after the literature check this section had been carrying a warning about since 2026-07-22. It splits into three parts and they do not have equal status. (a) The O(ρ²) half is CLASSICAL: every smooth f-divergence is locally quadratic, its second derivative IS the Fisher information, and for chi-squared the literature states χ²(P₁‖P₂) = (θ₁−θ₂)² I_F + o(·) directly. Our proof is that expansion applied to the perturbation ρu. (b) Our 'honest boundary' — the O(ρ) case where mass moves where the baseline is ≈0 — is the classical regularity condition under which the local expansion holds at all. We rediscovered the caveat that ships with the theorem. (c) The warp's Θ(ρ) half is one line from the definition, and that averaging dilutes brief events is not only known but actively engineered around (attention pooling exists for it). The COMPOSED separation Θ(1/ρ)→∞ was not found anywhere — but six web searches are weak evidence of absence, and a composition of two textbook facts is a thin novelty claim regardless. Full write-up: paper/LITERATURE-CHECK-S5.md.

Correction: This section was tagged 'OURS — the main new result' from 2026-07-22 to 2026-07-26, always carrying its own warning that its novelty was unchecked. The check was run and the warning was justified. The measured consequence (finding 2 — averaging is quadratically blind to a brief event) STANDS and is unaffected; what changes is only the claim to have discovered the mathematics behind it.

§6 The pitch / shift quotient

OURS — new extension

D_inv(P,Q) = min_s  D(P, σ_s Q)

isometry:  g(σ_s p, σ_s q) = g(p, q)  EXACTLY, for the full cyclic group   (verified to 4.4×10⁻¹⁶)

Plain: A pitch shift is a translation along a log-frequency axis, so quotient the space by that shift. What is left is relative formant spacing — which is identity. Invariance by construction, with nothing trained.

Provenance: OURS as an extension, BUT the idea — shift-invariant and transposition-invariant distances, chroma/CQT shift, quotient and orbit metrics — already exists in music information retrieval and group theory. A specific clean instance, not a first. Measured as finding 7: a same-vowel transposition collapses 61× while a different vowel is preserved.

§7 Instruments applied to the Riemann side

KNOWN — reproductions, NOT ours

R₂(r) = 1 − ( sin(πr)/(πr) )²      GUE pair correlation (Montgomery 1973)
H = (XP + PX)/2                    Berry–Keating operator
Weil/Connes explicit-formula operator

Plain: Established results run THROUGH Hodos as a referee, not equations we made.

Provenance: Everything reproduces known mathematics. RH is NOT proven. Hodos's role is the measuring instrument — it correctly ranks the GUE spectrum closest to the real zeros. Two exciting signals died to fair controls here (findings 22 and 23), both density artifacts. The honest limiter is resolution: the arithmetic fingerprint needs 10³–10⁴ zeros and we compute a few hundred.

The one-paragraph honest summary

We assembled a real, working, model-free engine — DTW plus an information-geometry (Fisher–Rao √-sphere) ground over sequences of distributions — with clean geodesic operators and a quotient extension for pitch invariance. The composition in §1 and the quotient in §6 are the parts that are ours. §5 was demoted on 2026-07-26 when the literature check it had been flagged for was finally run. The strongest results in this project are not in this tab at all — they are the MEASUREMENTS (findings 13, 14 and 27), which were built with controls designed to kill them and survived.

The rules, and why each one exists

Every rule here was written after something went wrong. They are not general good practice; they are scar tissue.

Write the criterion before running the test

A test whose bar is set afterwards can always be read as a success. Finding 13 predicted that the advantage would shrink with data — the signature of a prior. It grew instead. That makes it a better representation rather than a better prior: a different claim, and one we would have described wrongly had the criterion not been fixed in advance.

Build the control that can kill the result

Findings 8, 9, 18 and 22 were all killed by a control built for that purpose. Finding 18's smoothing arm matched the result exactly, which is what ended it. A benchmark with no arm capable of producing a negative is a demonstration, not a test.

Design the test to exercise the mechanism

Findings 5 and 19 both passed on inputs that could not have shown the effect. The activity adapter deleted direction, so two postures were identical by construction. The fp16 test used random data that never approaches the regime in question. A null from a crippled input is not a boundary.

An unexplained win is a red flag

Finding 9 won on one modality and lost on another with no account of why. Investigating the mechanism showed it was collapsing toward plain averaging — which is exactly why it won on the task where averaging wins. Explain the mechanism before banking the win.

Distinguish "measured nothing" from "could not measure"

Finding 23 returned a statistic that was identically zero for the real data and for all 900 shuffled controls. A z of zero against a null with zero spread is a dead instrument, not a null result. It is reported as a failed measurement.

An impossible number is worth chasing even when the headline is good

Finding 21 reported a negative forgetting score, which cannot happen if the store is order-independent. Chasing it found a real defect — and fixing it made the opposing method look worse, meaning the bug had been flattering the opponent.

Replicate on a second, unrelated modality

Speech and handwriting share no sensor, no physics and no preprocessing. Finding 4 is what turned finding 3 from a sound trick into a claim. And per finding 21, a second dataset is not only confirmation — it is a different defect detector.

Retractions stay visible — and they are not all the same thing

8 findings on this page are retracted, with the original wording kept and labelled. 16 more are outright negatives. Deleting them would make the record look stronger and be worth less.

A retraction count on its own is misleading, so each one is typed. Withdrawing a claim that something worked and withdrawing a claim that something failed are opposite events, and lumping them together makes a record of self-correction read as a record of failure. Of the 8 here, 3 withdrew a positive claim — we said it worked and our own later measurement said otherwise — and 5 withdrew a negative one: we had recorded a failure or a limit of the method against ourselves, and we were wrong about it. The second kind is the more common on this page. Both are on it for the same reason.

Status meanings

activeSurvives the controls it was tested against; currently believed.
provisionalMeasured, but on one task or at small scale; not yet replicated.
negativeTested and did NOT hold. Reported because a negative is a result.
retractedWe claimed it, then our own measurement refuted it. Original wording kept visible. Each retraction is labelled by WHAT it withdrew: a positive claim (we said it worked) or a negative one (we said it did not work, or that it was a limit of the method).
untestedStated but never measured. Carries no weight.

Confidence rungs

What it takes
1reasoned, unmeasured
2measured once, synthetic data
3measured on real data, one modality
4replicated on a second, unrelated modality
5survives a fair control designed to kill it
6survives comparison against a published method
7reproduced independently, outside this project

The Hodos Hypothesis

Nothing is what it is in isolation. A thing — including a self — is the pattern of its connections, not its substrate. What a system is is constituted by what it stands in relation to; and what it does that exceeds its parts — emergence — arises from that pattern of relation, and is never installed as a component. The relations are not decoration laid over pre-existing things: the things are constituted by relations, and those relations hold between further processes that are themselves so constituted. There is no isolated substrate underneath at which the regress halts.

Published origin: 10.5281/zenodo.21613155 (2026-07-27). Named: 2026-08-07. The name refers back to a premise already published on 2026-07-27; naming it on 2026-08-07 does not create it and does not move its date. The premise is stated in the abstract of 10.5281/zenodo.21613155, given its own section 5, and declared in section 2 to be the one premise all five of the method's commitments are downstream of. It is stated earlier still in Your Past Loves You (begun 2026-07-19), which that paper cites as the plainest statement of the thesis the method is built on.

The regress does not halt. A part is prior to the whole it composes, and that priority is real — but it is local, not fundamental. Two halves must exist before one exists; the one is what their relation constitutes. Each half, in turn, is what the relation of two quarters constitutes. The series does not terminate in an unrelated object: at every level the thing exists because two or more further things stand in relation, and the same holds of those. So conceding that the components precede the emergence concedes nothing about substance — the components are themselves emergences of the same form, one level down. There is no floor at which relation stops and bare substrate begins, which is precisely why the premise applies at every scale rather than at a chosen one, and why an engine built on it is not restricted to any particular kind of matter or signal.

Clarification of the premise published 2026-07-27; first stated in writing 2026-08-08. It amplifies the published premise and does not re-date it: the origin date above is unchanged.

Why "hypothesis" and not "theory". It is called a hypothesis, not a theory, because independent outside replication has not happened yet. That is the scientifically correct use of the word: a hypothesis is a stated, testable position; it becomes a theory when others test it and it holds. It graduates to "the Hodos Theory" when replication by people with no stake in it has been done and has survived — not when it feels established, and not by the author's decision.

Two readings of the name. The Hodos Hypothesis is the principle (10.5281/zenodo.21613155). Hodos Diastema is its measurement instantiation — the model-free distance between processes as curves of distributions (10.5281/zenodo.21612829). The dependency runs one way: the equation is downstream of the hypothesis, never the reverse. Cite the full phrase "the Hodos Hypothesis" for the principle and "Hodos Diastema" for the distance, and attach the DOI in both cases. The bare word "Hodos" names the programme and should not be used for either on its own: it was doing both jobs at once, and readers took it for the equation and never reached the premise underneath. Each equation in the family carries its own name — Diastema (the distance), Symploke (emergence), Systasis (constitution), Chronos (duration). The distance was named on 2026-08-09, which re-dates nothing: its paper, its DOI and every existing citation of "Hodos" for the distance remain correct as written.

What is and is not claimed. Relational ontology is older than any of this machinery, and the source paper says so in its own text. The premise has deep ancestry — dependent origination, process philosophy, structural realism, relational quantum mechanics. No one can claim the idea itself and this does not. What is claimed is this formulation as the named premise of a named programme, its instantiation in working measured machinery, and the timestamp.

What this claims

A signal is a curve of probability distributions on a statistical manifold; comparison is a geodesic distance between curves.

The narrow version, which is the one that survived

The ground — the information-geometry representation, not the alignment machinery — is what earns a measurable advantage, and it does so wherever identity lives in the relative-space pattern of a process. The warp is a separate lane: it earns its keep only where identity lives in temporal motion, and on discrete already-aligned data it actively hurts (finding 17).

These are two lanes. Conflating them was our own error, retracted as finding 5 and corrected in finding 6.

What it is good at, stated as capabilities

It learns from 5–10× less data than a matched standard network, on two unrelated modalities (finding 13). At one example the network is barely above chance.

It does not forget (finding 14). Taught in sequence, a standard network loses half to four-fifths of what it first learned. This loses nothing — and cannot, because building the store in reverse order yields byte-identical answers, so the final state provably does not depend on the sequence.

It is invariant by construction rather than by training (finding 7). A transposed signal collapses 61× while a genuinely different one is preserved. A network has to learn that from examples and only ever approximates it.

What it costs

The store grows. Consolidation folds repeated episodes into one prototype per concept, so it scales with how many distinct things are known rather than with how much has been seen — but lookup gets slower as it learns, where a network's does not. That is the real trade, and it is why a fast retrieval index exists.

What it is not

Not a universal oracle. It solves what can be posed as a least-cost path on a distribution-manifold; it does not invent a domain's law and does not reach discrete or logical truths. The Riemann work (finding 22) reproduces known mathematics and proves nothing new.

Not better on language. Memory helps a language model, but a plain distance retrieves better than this one for text (finding 24), and the mechanism is understood (finding 17).

Not novel wholesale. Dynamic time warping, Fisher–Rao, the square-root sphere, soft-DTW, barycenter averaging and optimal transport are all prior art, reused. What is ours is the synthesis, the pitch quotient, and the measurements on this page.