Comparing processes as curves of distributions
Line of work
Field measured on
· Work applying this engine in a specific domain keeps its own record, its own numbering and its own evidence ladder. Nothing on this page is a clinical claim, and the lane below is not cleared or approved by any regulator: Cordthym — cardiac rhythm. A result in one line is not a claim about another.
foundations (38) — The distance itself, its operators, and what they were measured against. The results the rest of the record stands on.
architecture (5) — Memory-based learning: the retrieve-compose-decode loop, composition, and the coverage law. An architecture claim, not an application result.
cardiac rhythm (5) — Cardiac rhythm application of the architecture, on public patient recordings. Application results, NOT architecture claims, and no clinical claim.
machine telemetry (11) — Prognostics and anomaly attribution on machine sensor data. Method and negatives, not a predictor and not a product.
The hypothesis is the premise; each equation it generates carries its own name, and the name is Hodos plus one word. This matters for reading the record below: the findings were written before the equations were named, so they speak of "the distance", "the ground" or "the surplus". Those are the objects named here.
Hodos Diastema — How far apart are two processes?
The model-free distance between processes as curves of distributions. It is the equation nearly every finding on this page is measured with — where a finding says "the distance", "D", or "the Fisher–Rao ground", this is what it means. 10.5281/zenodo.21612829
Hodos Symploke — What does a relationship create?
The emergence surplus: the joint against the product of the marginals. Cleared all four criteria fixed before it ran. 10.5281/zenodo.21850666
Hodos Systasis — What are the parts made of?
Relational — and the 'lossy' half is WITHDRAWN. It genuinely is relational: destroying the shared coordinate system collapses it. The earlier claim that it loses to the window's own values was withdrawn on 2026-08-15/16 by our own measurement, not by argument — the paired leave-one-out gap is 4 items of 200, then 4 to 12 of 600, sign-stable but below what this design can resolve, and a gap the design cannot resolve is withdrawn rather than kept as a null. 10.5281/zenodo.21850666
Hodos Chronos — How much time has a process lived?
Clears four of its five pre-registered bars. The fifth, sampling invariance, did not clear at the registered configuration - 0.759 against a bar of 0.05 - and the paper says so on the same page that announces the other four. 10.5281/zenodo.21861429
The same representation is applied in every field below. Whether it helps
is a separate question, measured case by case — the fields where it did not are on this page
with the same weight as the fields where it did. Generated from manifest.json,
so it cannot disagree with the record underneath it. Click a row to filter.
| Field | Status mix | What the data is | |
|---|---|---|---|
| mathematics | 7 | 4 active 3 provisional | Identities and proofs. No dataset — these either hold or they do not. |
| speech | 15 | 12 active 1 provisional 1 negative 1 retracted | Real recorded spoken digits, six speakers. The across-speaker split is the hard one: the same word through a voice the system has never heard. |
| handwriting | 12 | 8 active 3 negative 1 retracted | Real pen motion from a graphics tablet. No audio, no shared sensor with speech — which is what makes it a replication rather than a second look. |
| text | 8 | 3 active 2 negative 3 retracted | Character-level English. Discrete and already aligned, which is the regime the method is measured to be WEAKEST in. |
| human activity | 3 | 1 active 1 provisional 1 retracted | Body-worn accelerometer windows — sitting, standing, walking. |
| light | 2 | 1 negative 1 retracted | Emission spectra of materials. |
| sonar | 3 | 2 negative 1 retracted | Sonar range profiles and Doppler returns. |
| turbulence | 1 | 1 negative | Turbulent versus calm flow. |
| cardiac rhythm | 5 | 3 active 2 negative | Inter-beat intervals from public patient recordings, per patient. |
| jet engine | 9 | 3 active 5 negative 1 retracted | NASA C-MAPSS turbofan run-to-failure. Simulated, one engine family — it is not flight telemetry and a result here is not a claim about one. |
| synthetic vehicles | 2 | 1 active 1 negative | Multi-channel systems built deliberately to break the machinery — mixed units, awkward distributions, planted faults with a known answer. |
| zeta zeros | 2 | 2 negative | The Riemann zeta zeros and candidate operator spectra. An honest appendix: everything here reproduces known mathematics and proves nothing new. |
| spacecraft telemetry | 4 | 1 active 1 provisional 2 negative | Real ESA mission telemetry (ESA Anomaly Dataset, CC BY 3.0 IGO), with channel-level ground truth and real, unequal subsystems. |
| synthetic pairs | 2 | 1 active 1 negative | Constructed pairs of processes with a known coupling, used where the right answer has to be known in advance. |
| astronomy | 5 | 3 active 1 provisional 1 retracted | Pantheon+/SH0ES: 1,701 public type-Ia supernovae, twelve measured columns per object. The catalogue ships inside the instrument that reads it. No cosmological claim is made on this data - the clock reads relations; its pre-registered cosmic-time bars did not clear at the registered configuration, and its own face says so. |
Counts and the table below are generated from manifest.json, so they
cannot disagree with the record. Newest finding: 2026-08-12.
| Finding | Field | Status | Rung | Date | |
|---|---|---|---|---|---|
| 1 | A process is literally a curve on a spherefoundations | mathematics | active | 5 | 2026-07-21 |
| 2 | Time-averaging is quadratically blind to a brief eventfoundations | mathematics | provisional | 2 | 2026-07-21 |
| 3 | The information-geometry ground beats Euclidean on real speech across voicesfoundations | speech | active | 3 | 2026-07-23 |
| 4 | The same win replicates on handwriting, with no audio at allfoundations | handwriting | active | 4 | 2026-07-23 |
| 5 | The method fails where identity is static — the predicted boundaryfoundations | human activity | retracted | 3 | 2026-07-23 |
| 6 | The corrected boundary: two lanes, not onefoundations | human activity | active | 3 | 2026-07-24 |
| 7 | Pitch invariance by construction, not by trainingfoundations | speech | active | 5 | 2026-07-22 |
| 8 | The distinctive ground does NOT generalise across light, echo and turbulencefoundations | light, sonar, turbulence | negative | 5 | 2026-07-22 |
| 9 | The min-mean distance 'win' was a padding degeneracyfoundations | light, sonar | retracted | 5 | 2026-07-22 |
| 10 | Doppler / velocity measurement is not an edgefoundations | sonar | negative | 5 | 2026-07-22 |
| 11 | Against a published method: the field's warp is better, our ground still winsfoundations | speech, handwriting | active | 6 | 2026-07-23 |
| 12 | The win decomposes: trajectory > ground > warpfoundations | speech, handwriting | active | 4 | 2026-07-24 |
| 13 | The geometry needs 5–10× less data than a standard networkfoundations | speech, handwriting | active | 4 | 2026-07-25 |
| 14 | It does not forget — and cannot, by constructionfoundations | speech, handwriting | active | 5 | 2026-07-26 |
| 15 | A compressed latent representation does NOT helpfoundations | text | negative | 5 | 2026-07-25 |
| 16 | Anything between the landed point and the answer must start as a pass-throughfoundations | text | active | 5 | 2026-07-25 |
| 17 | On discrete aligned data the warp HURTS and the ground is neutralfoundations | text | retracted | 5 | 2026-07-25 |
| 18 | Hierarchy does not beat a flat representationfoundations | speech, handwriting | negative | 5 | 2026-07-25 |
| 19 | The geometry is not fp16-safefoundations | text | retracted | 1 | 2026-07-24 |
| 20 | Measurement error: a growing label space looked like memory decayingfoundations | speech, handwriting | retracted | 2 | 2026-07-26 |
| 21 | An impossible value exposed a flaw that only a second dataset could revealfoundations | handwriting | active | 5 | 2026-07-26 |
| 22 | The Riemann work reproduces known mathematics and proves nothing newfoundations | zeta zeros | negative | 5 | 2026-07-22 |
| 23 | Character spacing is not Riemann level repulsionfoundations | zeta zeros, text | negative | 5 | 2026-07-25 |
| 24 | Memory helps a language model, but the geometry is not whyfoundations | text | retracted | 5 | 2026-07-25 |
| 25 | The Fisher-Rao ground IS a better retrieval key than L2 -- once characters have a learned metricfoundations | text | active | 5 | 2026-07-26 |
| 26 | The learned character metric encodes linguistic class, and it is not frequencyfoundations | text | active | 3 | 2026-07-26 |
| 27 | The equation predicts the next step better than a trained network, with no trainingfoundations | speech, handwriting | active | 5 | 2026-07-24 |
| 28 | The closed loop as a learner: retrieval and instant learning hold; the decode leg does not closearchitecture | speech | active | 3 | 2026-08-03 |
| 29 | Order-invariance holds exactly through the composed retrieval patharchitecture | speech | active | 3 | 2026-08-03 |
| 30 | Composition is task-dependent: blending helps the vote and hurts the artifact; the decode gap is structuralarchitecture | speech | active | 3 | 2026-08-04 |
| 31 | The personal condition: the loop closes end-to-end as a function of coverage, with zero trainingarchitecture | speech | active | 3 | 2026-08-04 |
| 32 | The decider is coverage-dependent, and blended evaluators suppress decode scoresarchitecture | speech | active | 3 | 2026-08-04 |
| 33 | The coverage law does NOT transfer to heartbeats: direction survives, magnitude does notcardiac rhythm | cardiac rhythm | negative | 3 | 2026-08-05 |
| 34 | Same organ, other side of the boundary: the loop works on cardiac RHYTHM, per patientcardiac rhythm | cardiac rhythm | active | 3 | 2026-08-05 |
| 35 | Measured against the field: below the published state of the art, and per-patient memory adds little on averagecardiac rhythm | cardiac rhythm | negative | 3 | 2026-08-05 |
| 36 | The gap to the published field was data, not a ceiling: 0.983 at 64 examples per state (pure windows only — see finding 37 for the harder protocol)cardiac rhythm | cardiac rhythm | active | 3 | 2026-08-05 |
| 37 | The excluded hard cases, scored: 0.966 on the continuous stream, and onset is caught on the first windowcardiac rhythm | cardiac rhythm | active | 3 | 2026-08-05 |
| 38 | Matched pairs on jet engines: something beyond present distance predicts which unit fails sooner — but the differential operators are not what reads itmachine telemetry | jet engine | negative | 4 | 2026-08-05 |
| 39 | Per-condition normalisation makes multi-regime fleets measurable at all — and a count caught the defect that would have hidden itmachine telemetry | jet engine | active | 4 | 2026-08-05 |
| 40 | The excursion replicated pre-registered on two unseen fleets — and is retired anyway, because part of it was a modelling choicemachine telemetry | jet engine | retracted | 4 | 2026-08-07 |
| 41 | Attribution on real engines: the null control holds on hardware, and the binary question turns out to be the wrong onemachine telemetry | jet engine | active | 3 | 2026-08-07 |
| 42 | Ordering measured against the right null: nine-tenths of the apparent concentration was the representation, not the faultmachine telemetry | jet engine | negative | 3 | 2026-08-08 |
| 43 | Mixed units are not the blocker they were assumed to be — plain z-scoring beats both distributional standardisers on channels built to break itmachine telemetry | synthetic vehicles | negative | 2 | 2026-08-07 |
| 44 | The layered architecture was refuted too broadly: it was the pooled statistic, not the hierarchymachine telemetry | synthetic vehicles | active | 2 | 2026-08-07 |
| 45 | First contact with real spacecraft telemetry: the obvious instrument alarms on everythingmachine telemetry | spacecraft telemetry | negative | 3 | 2026-08-08 |
| 46 | With the channel's own normal variability as the null, it works — and names the right channels at 4.5× the base rate against real ground truthmachine telemetry | spacecraft telemetry | active | 3 | 2026-08-08 |
| 47 | The layered architecture does NOT transfer to real, unequal subsystems — flat wins on 40 events and loses on nonemachine telemetry | spacecraft telemetry | negative | 3 | 2026-08-08 |
| 48 | Rare-but-normal events do alarm less than anomalies — by a registered margin, and still far too often to be usefulmachine telemetry | spacecraft telemetry | provisional | 3 | 2026-08-08 |
| 49 | The premise made falsifiable: destroy the relations and identity goes on handwriting, but not on a jet enginefoundations | handwriting, jet engine | negative | 3 | 2026-08-08 |
| 50 | Hodos Symploke: a quantity for what a relationship creates, with the control built into the equationfoundations | synthetic pairs, handwriting, jet engine | active | 3 | 2026-08-08 |
| 51 | Hodos Systasis: frames derived from relations really are relational — and lose to simply using the windowfoundations | handwriting, jet engine | negative | 3 | 2026-08-08 |
| 52 | Hodos Chronos: intrinsic time is well posed on smooth processes and diverges on real sensor datafoundations | synthetic pairs, jet engine | negative | 2 | 2026-08-08 |
| 53 | Hodos Chronos rebuilt: four of five pre-registered bars clear, and the obstruction that stopped thirteen attempts was an artifact of our own estimatorfoundations | speech, human activity, mathematics | provisional | 2 | 2026-08-09 |
| 54 | A clock made of relations, pointed at 1,701 supernovae - and its cosmic-time bars did not clear at the registered configurationfoundations | astronomy | active | 3 | 2026-08-10 |
| 55 | The shipped warp was not a distance: a thing's distance from itself was negative, and grew with its lengthfoundations | astronomy, mathematics | active | 3 | 2026-08-11 |
| 56 | RETRACTED: 'the whole converts' - the result described a catalogue that no longer existed, and does not reproduce on the current onefoundations | astronomy | retracted | 3 | 2026-08-11 |
| 57 | Chronos is oriented, and nothing had ever checked the direction: 63 of 66 relations read differently reversed - and the headline survives either wayfoundations | astronomy, mathematics | active | 3 | 2026-08-12 |
| 58 | The 'unexplained' detection band was the plant's own fault: |sin| halves the period, and the corrected rule is exact on 58 of 58 cellsfoundations | mathematics | active | 2 | 2026-08-12 |
| 59 | A second level of time cannot be resolved on this catalogue: at nine level-2 frames the statistic's own calibration is wider than it claimsfoundations | astronomy, mathematics | provisional | 2 | 2026-08-12 |
Claim
If constitution has no floor, whatever a pairing produces can itself pair - a surplus timeline standing in its own relation would be a SECOND level of time. The registered test runs the nested-surplus gates at the clock's own shape: 1,700 objects give 67 surplus frames give about nine level-2 frames, the only window arithmetic allows. The calibration gate failed there: on synthetic noise the score must sit inside the width it claims (|z| < 2), and it did not - so the ruler's own markings are wider than stated at this depth, no reading taken with it can be placed, and the run discarded itself before scoring any real data. A resolution verdict: the question needs a longer series, and this catalogue cannot ask it.
Evidence
N0 on registered seeds: nested z of -0.66 / +2.98 / +0.82 against a |z| < 2 calibration. Mechanism measured over 12 seeds: the null is over-dispersed, sd 1.27 against a calibrated 1.0. The ledger row records the discard.
Scope / limits
The gate is a CALIBRATION FLOOR, never a claim that anything is unrelated - under the premise nothing is, and a higher level would stand in relation to everything by construction. Failing it says the instrument cannot certify itself at this depth; it neither finds nor refutes a higher time. The premise's own derivation expects nesting to add nothing if constitution is already pairwise - and an unaskable question confirms nothing either way. The nested object is pairs of pairs throughout: one pair's surplus standing in one relation with a third part.
Controls that could have killed it
Claim
Which planted periods each window can detect had resisted a formula - several in-band periods scored zero and an out-of-band one scored - and was recorded as unexplained rather than quoted. Measured on a grid of planted arms: the planted cycle uses |sin|, which doubles the frequency, so the arm's true period is HALF its parameter, and the standing formula had been scoring the wrong period (14/36 against the measured band). The corrected rule - some multiple of the effective period fits between the minimum return separation and the window count, AND the departure gate is satisfiable on the cell's own scales - reproduces the band exactly: 36/36 on the measurement grid and 22/22 on fresh cells, with the arithmetic half of every prediction printed before any distance was computed.
Evidence
probe_detectable_band.py: every detected cell's return lags cluster at integer multiples of period/2 (period 22 detected at lags 33/44/55; period 30 at 45). probe_detectable_band_confirm.py: 22/22 fresh cells; a +/-1-frame lag tolerance FAILED on 8 cells whose clusters are +/-2 wide - the tolerance was the miscalibration, and the failure is reported rather than dropped.
Scope / limits
Instrument calibration on planted arms only; no real data enters. A not-detected cell is a statement about the instrument at that window and period, never about data. Consequence applied the same day: the calibration opened one honest extra window (W=7) for the conversion re-measurement above.
Controls that could have killed it
Claim
Chronos's estimator is an oriented function of its two arguments, and the shipped clock had read every relation in one direction chosen by nothing but the order the parts sat in a dictionary. Measured both ways: tau(a,b) differs from tau(b,a) on 63 of 66 relations, median relative difference 0.4485 against a 0.1828 same-direction seed floor - 2.45x, separable. The whole's total reads 113.978 one way and 106.318 the other (6.7% apart), but the share of time living between the things is 73.9% vs 74.2% - the claim that matters is direction-independent. Treatment settled from the framework's own canon: the two readings are two sides of ONE relation, never summed and never maxed - 66 relations carrying 132 numbers, not 132 relations.
Evidence
probe_is_time_directional.py: 63/66 differ, 2.45x the seed floor. run_direction_readings.py (full restatement, both directions, one seed): 61 relations resolved both ways, 2 one way only (each reading a real number one direction and exactly 0 the other), 3 neither. The per-relation convert/cycle classification is direction-blind to the last digit - max absolute difference 0.00e+00 both scores, zero verdicts flip - measured, not assumed from the joint's permutation symmetry.
Scope / limits
One catalogue, one seed for the both-directions restatement (the seed floor arm is what makes the asymmetry reportable). A relation live in one direction and exactly 0 in the other is a resolution verdict about this instrument, never a one-way relation - and never a horizon detection; it is a lead.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW WITHDRAWN: 'The whole converts, moderately: the state path scores 1.9134 against phase surrogates near 0.99 (min ratio 1.852 over a 1.5 bar) with zero returns - it converts and does not cycle.' Measured 2026-08-11 00:5x, matching a model distinction stated before the run; withdrawn the same day.
Evidence
Registered run: 1.9134, surrogates 0.9912/1.0331/0.9885, min ratio 1.852, returns 0. Withdrawal: five ledger rows marking each whole-level result not current, readable rather than deleted. Re-measurement sweeping W=5/6 (all usable windows, one planted period in band for all): conversion 0.890/0.948 against surrogates ~1.0 - ratios 0.877/0.942 against the 1.5 bar. Extension at W=7 after an instrument calibration opened it: 0.946, min ratio 0.943 - same verdict at the strongest window the path can carry.
Scope / limits
The retraction is about the CATALOGUE CHANGE, not an error in the registered run: the result was correct for the data it was computed on, and that data is no longer what the instrument reads. The small-window null is weak evidence - a 5-7 frame window gives the warp little to align - and transfers to nothing.
Controls that could have killed it
Why this was wrong
A data-cleaning fix landed between the run and the day's end: a not-measured marker (-9) in one of the twelve parts had been read as a value, and removing those rows shrank the WHOLE's path from 1,701 objects / 67 frames to 712 / 26 - the five whole-level results had all been scored on a path that no longer exists. At 26 frames the pre-registered window cannot be asked at all (15 windows against the 25 the test requires), and the re-measurement at every usable smaller window - W=5, 6, and later 7 - reads no visible conversion (ratios 0.877 / 0.942 / 0.943 against the 1.5 bar). The retraction is about the catalogue change, not an arithmetic error: the original number was correct for the data it was computed on, and that data is not what the instrument reads now.
Claim
The clock shipped with raw soft-DTW as its warp. Raw soft-DTW is the undebiased Cuturi-Blondel object: the softmin subtracts a gamma*log term at every cell, so the bias accumulates with path length and self-distance goes negative. Measured on the project's own state path: D(A,A) = -0.469, and -1.235 for a longer A. One of the 66 relations displayed a negative 'moved' value on the shipped face. Fixed to the debiased soft-DTW divergence (exactly 0 on identity); tau was unaffected because tau is Chronos and only the travel figure used Diastema.
Evidence
check_warp_is_a_distance.py: soft_dtw D(A,A) = -0.469 / -1.235; soft_dtw_divergence and dtw both 0.00000 on identity. After the fix, travel spans 3.79-329.53 with zero negatives.
Scope / limits
The engine itself was untouched - its own source names the divergence as the proper non-negative form and dtw as the master distance; the defect was the instrument's configuration choice. Every earlier Diastema-based figure in this lane was computed with the biased warp; the ledger is append-only and records them as they were.
Controls that could have killed it
Claim
Chronos and Diastema run over every pairwise relation of the twelve measured columns of the Pantheon+/SH0ES catalogue: 12 parts, 66 relations, 85 dials, packaged as a one-file desktop instrument. Its pre-registered cosmic-time bars did not clear at the registered configuration and the face says so: this is a clock OF the relations in one catalogue; whether it reads the universe's age is what those bars could not resolve at this depth. What it reads that survives its controls: most of the whole's accumulated time lives BETWEEN the things rather than inside any of them - roughly 74%, a share that later proved direction-independent.
Evidence
C1 epoch ordering 1.164 against a 1.5 bar - did not clear. C3v2 rate agreement +0.094 against 0.30 - did not clear, and the rise loses to its own surrogates. A4 sampling 0.147 against 0.05 - did not clear, a live confound quoted beside every trend on this data. What passes: the six-part K6 sub-run's resolvable edges land on known astrophysics nobody told it about (host mass vs brightness residual is the published mass step).
Scope / limits
One public catalogue, one epoch ordering, seeds fixed. Frames are windows of 96 objects, not equal spans of time - every per-frame trend is quoted with that units caveat. Nothing here is a cosmological claim.
Controls that could have killed it
Claim
Duration for a system of interacting parts, accumulated from the relations between the parts rather than along a sampling grid. Thirteen constructions traced a single trade-off — a clock could be sure about the direction of time or quiet when nothing happened, never both — and that trade-off was about to be written up as structural. It was not structural. It was a one-sided nonlinearity in our own estimator: `max(eps - null, 0)` clips a quantity resting on a floor, and a clipped quantity rises abruptly and decays back for no reason but the clip, which reads as an arrow of time that is not in the data. The amount must keep the clip; the arrow must not. The fourteenth version reads the amount off the clipped curve and the direction off the unclipped one, and clears four of the five criteria fixed before any of them ran.
Evidence
Bars registered in advance, scored on the shared harness. A0 time-asymmetry (speech): 2.020 / 2.044 / 2.122 / 2.133 across 4 seeds, every p < 1e-12, against a 1.5 bar. A1 stillness (277 human-activity runs, rest vs motion): 2.278 / 2.137 / 2.218 / 2.209 / 2.252 across 5 seeds, every p < 1e-33, against a 2.0 bar. A2 relation-sensitivity: 2.722 at p 7.5e-07 where a magnitude-only measure is arithmetically blind at 0.958. A3 exactness, on parts independent by construction: z = -0.190, i.e. indistinguishable from its own null. A4 sampling invariance did not clear at the registered configuration: 0.759 against a 0.05 bar.
Scope / limits
One construction over the six-part relational graph, scored on speech, human activity and synthetic independent parts. A0 cannot be scored on the activity data at all — it is time-reversible, so the arrow fires at chance there — and A1 needs a dataset with rest in it, so no single dataset carries all the bars. Nothing here has been through a second pair of eyes or an independent reproduction.
Controls that could have killed it
Claim
Every other quantity here takes the time axis as given. Defining duration instead as distance travelled through the process geometry gives a process its own clock: a system that sits still accrues no time however long the wall clock runs. The construction is defined by nothing but successive relations between states — but it is only a clock where the process is smooth. On noisy data the reading depends on how often you looked.
Evidence
Relative change in total duration when the sampling rate is halved: smooth curve 0.003, random walk 0.325, real turbofan sensor trajectory 0.500 — rising to 0.022 / 0.690 / 0.884 at stride 8. On real telemetry, halving the sampling rate roughly halves the measured duration. A still process accrues exactly 0.000.
Scope / limits
One ground metric, one estimator, synthetic curves plus one real sensor trajectory. It says nothing about processes observed finely enough to be smooth at the scale of interest.
Controls that could have killed it
Claim
Every domain this engine touches gets its frame from a hand-built adapter, so an engineer picks the axis each time. Building the frame instead from nothing but relations — locating each window by its normalised similarity to a set of reference windows, so no bin means a frequency or a sensor — produces a genuine relational encoding and a worse one. On handwriting it beat the hand-built adapter (0.960 against 0.935) and still lost to a frame built from the window's own absolute values, referencing nothing at all (0.980).
Evidence
Handwriting, 20 classes, 200 items: adapter 0.935 / relational 0.960 / absolute 0.980 / per-item references 0.215. Jet engine, healthy against mid-life, 100 items: adapter 0.800 / relational 0.580 / absolute 0.590, with the relational arm failing its validity bar against a 0.500 floor. One setting across both domains, never tuned per domain.
Scope / limits
One construction of a relational frame, two domains, one downstream engine. It bounds this construction, not the idea that representations can be derived.
Controls that could have killed it
Claim
What two processes do together, compared against what the same two would do if they did not interact, is a computable quantity in the same geometry the distance uses. At each instant the joint distribution of the pair is compared with the outer product of its own marginals; the surplus is read through the engine's master distance. If the parts genuinely do not interact the two are identical and the surplus vanishes — the control is a term in the equation rather than an arm bolted beside it.
Evidence
Four criteria fixed before running. Independent parts sit inside the quantity's own shuffle null (z = +0.06, p 0.46). Sweeping coupling from 0 to 0.95 gives z = +0.1 / +3.4 / +9.6 / +24.9 / +39.5 — monotone. On y = x squared, where Pearson correlation reads -0.095, it fires at z = +18.8, so it is not a correlation proxy. And on real data it reproduces a split measured by a different method before this equation existed: mean surplus z = +8.08 across handwriting part-pairs against +1.32 across jet-engine sensor pairs, a gap of +6.75 against a pre-registered bar of 5.0.
Scope / limits
Pairwise by construction; a K-part total carries a higher-order remainder this does not compute. The estimator needs many more samples than the joint has cells — a first version at 8 bins over a 16-sample window scored an INDEPENDENT pair at 1.29 against a coupled pair's 1.98, because a joint that sparse cannot resemble a smooth outer product whatever the parts are doing. Reported throughout as a z against its own null, never as a raw surplus: a finite window makes the quantity positive even for independent parts.
Controls that could have killed it
Claim
The principle this record sits under says a thing is the pattern of its connections rather than its substrate. Stated as an empirical prediction it can fail, and it does — on one of two modalities. Three surgeries were applied to the same data: destroy temporal ORDER, destroy CROSS-PART COUPLING while leaving every part's own values byte-identical, and destroy the SUBSTRATE with per-part monotone unit changes while leaving every relation intact. On handwriting the prediction lands hard. On jet-engine telemetry it does not: the parts must be POOLED, but they need not INTERACT.
Evidence
HANDWRITING (20 classes, 200 items, majority floor 0.050): intact 0.935. Destroying cross-part coupling — every part keeping its exact values and the representation's time-average exactly unchanged — collapses it to 0.315 (−0.620, p 0.0001). Destroying order alone: 0.400 (−0.535, p 0.0001). Replacing the substrate with per-part monotone transforms that change every number, unit and distributional shape: 0.885, a difference of exactly 10 items in 200. Non-interacting controls, each part classified in ISOLATION and merged assuming no interaction: 0.565 (vote) and 0.755 (independent likelihoods multiplied) — the intact method beats the better of them by +0.180, p 0.0001. JET ENGINE at 55% of life (120 items, floor 0.500): intact 0.775, coupling destroyed 0.758 — a difference of +0.017 at p 0.42, so the cross-part relation carries essentially nothing there — while the non-interacting controls reach only 0.608 and 0.567, so the intact method still beats them by +0.167 (p 0.0059).
Scope / limits
Two modalities, one representation each, leave-one-out nearest-neighbour. This tests the premise's PREDICTION on these data; it does not establish the general ontological claim, and it is not a promotion of the principle to a theory — that requires replication by people with no stake in it. Neither dataset supports the premise on all five pre-registered criteria, so no rung is earned for it.
Controls that could have killed it
Claim
82 of this mission's 200 annotated events are rare but NORMAL, and the dataset's authors state that not alarming on them is of high practical importance. Measured: the flat decider alarms on 0.975 of anomalies and 0.817 of rare-nominal events — a difference of +0.158, clearing a pre-registered 0.15 bar at p = 0.0005. The direction is real and registered. The absolute number is not usable.
Evidence
118 anomalies against 82 rare-nominal events, same instrument, same held-out-validated false-alarm rate of 0.025. Flat Bonferroni and FDR both give +0.158 at p = 0.0005 on a label-shuffle permutation. The layered decider goes the WRONG WAY, −0.133 at p = 0.98, and fails.
Scope / limits
One mission, one annotation scheme, and the events are labelled by ESA operators rather than by us. Being 'rare nominal' is a judgement about operational meaning, not a property of the signal, so a method reading the signal alone has no principled route to the distinction beyond magnitude.
Controls that could have killed it
Claim
Every previous test of the subsystem-then-channel hierarchy used four invented subsystems of five channels each. This mission's subsystems are real and unequal — 42, 17, 11 and 6 channels. There the layered test is strictly worse than testing every channel flat: it detects 0.636 of anomalies against 0.975, and the paired comparison is one-sided.
Evidence
Paired McNemar on the same 118 anomaly events: 40 events the flat test catches and the layered one misses, ZERO the other way, p < 0.0001. Attribution follows the same order — layered recall 0.217 at 3.19× base against flat 0.429 at 4.54×. Mechanism measured rather than argued: mean subsystems reaching stage one is 1.271 of 4 on anomalies against 0.020 on held-out nominal windows, so the stage-one gate is specific but far too conservative on the large subsystem.
Scope / limits
One mission, one subsystem partition. This does not refute hierarchy in general — it refutes it for THIS correction on THESE group sizes.
Controls that could have killed it
Claim
Replacing the null with an empirical, per-channel one — how far does THIS channel move between two normal stretches, measured over 1,600 event-free windows of its own history — turns the instrument of finding 45 into a working one. It holds its false-alarm rate on held-out event-free windows, detects 97.5% of annotated anomalies, and identifies which channels were involved at four and a half times the base rate.
Evidence
False alarm 0.025 (flat Bonferroni), 0.028 (FDR), 0.018 (two-stage) on 400 HELD-OUT event-free windows that played no part in building the null. Detection 0.975 of 118 annotated anomalies for both flat deciders. Attribution against the dataset's own channel-level labels: flat Bonferroni recall 0.429 at precision 0.591 — 4.54× the 0.130 base rate — with a channel-label shuffle p of 0.0010; FDR recall 0.562 at precision 0.528 (4.06× base). 76 channels, 200 annotated events, empirical null of 1,600 windows per channel with the resolution checked against the corrected bar before scoring.
Scope / limits
One mission. Windows are resampled to a fixed 256-point grid, so transients shorter than the grid are not resolved. Recall of 0.43 means the majority of labelled channels are still missed; this identifies SOME of the right channels well above chance, not all of them. Real flight telemetry, but one vehicle and one annotation scheme.
Controls that could have killed it
Claim
Asking whether a channel's distribution over a window differs from the window before it — a two-sample permutation test, the textbook construction — flags at least one channel on EVERY event-free window of real mission telemetry. False-alarm rate 1.000. It is not broken; it is correct and useless.
Evidence
150 event-free control windows drawn from stretches with no annotated event anywhere in them or in their reference, 76 channels, 2,000 permutations per channel, permutation count checked against the corrected threshold it must resolve (0.000500 against a 0.000658 bar, margin ×1.32). All three deciders — flat Bonferroni, flat FDR, two-stage — flagged 150 of 150. The pre-registered gate declared every other arm of that run VOID rather than failed, and they are reported nowhere.
Scope / limits
One mission, real flight telemetry, 5.5-hour windows on a fixed resampling grid.
Controls that could have killed it
Claim
Testing a subsystem by pooling its channels' distributions was measured at 0.03 detection against a flat test's 0.50, and the conclusion drawn was that the layered subsystem-then-channel architecture is refuted. That was too broad. Asking instead whether a subsystem CONTAINS a diverging channel — a maximum over its channels rather than a pooled statistic — recovers it: 0.67 detection at zero false alarms and 4.4 times cheaper.
Evidence
4 subsystems × 5 channels, 30 seeds, criteria fixed before the run. Two-stage MAX 20/30 = 0.67 [0.49, 0.81] with 0/30 false alarms, against flat Bonferroni 0.50 [0.33, 0.67], flat FDR 0.53, and the pooled two-stage version it replaces at 0.03 [0.01, 0.17]. The registered prediction landed: mean subsystems flagged on a fault vehicle goes from 0.033 to 0.867 — stage one now fires at all, which is the whole finding. Cost 219,330 permutation-divergences against 961,200.
Scope / limits
Synthetic vehicles with EQUAL, INVENTED subsystems of five channels each, one faulted channel, all channels standard normal. On real hardware a raw maximum is dominated by the noisiest channel unless it is standardised first — and real subsystems are not equal sizes, which changes the null a maximum is scored against.
Controls that could have killed it
Claim
For months the record carried an untested assumption: that pressures, temperatures, voltages and rates must each become a distribution before this geometry applies. Measured, it is unsupported. On ten channels chosen to be awkward — lognormal flow, heavy-tailed bus voltage, a valve pinned at its rail, an integer counter — plain per-channel z-scoring wins by a wide margin at equal false alarm.
Evidence
At a 20-frame baseline, detection 0.69 for z-score against 0.28 (median/MAD) and 0.27 (probability integral transform), thresholds fitted on held-out no-fault seeds so the comparison is at equal false alarm. Across baseline lengths 20 to 200 z-score runs 0.89 → 0.97 and never once fails its false-alarm gate. A saturation prediction written before the run landed exactly: the PIT is bounded by its baseline sample, so it goes flat at 0.10–0.12 from drift 3 upward while z-score climbs 0.48 → 0.94.
Scope / limits
Synthetic channels designed to be awkward, not real telemetry. The claim is 'not the bottleneck on data built to break it', NOT 'will never be the bottleneck'. z-score handled them at 0.89–0.97, so they may be milder than real mixed-unit hardware.
Controls that could have killed it
Claim
Asked which sensor leads on a degrading engine, the method looks strikingly consistent — until it is scored against how consistent it is on HEALTHY engines instead of against a uniform ideal. Against the healthy population's own concentration the honest gap is +0.102, not the 0.95 the weaker comparison implied. What does survive is narrower and real: WHICH sensor leads differs between the two fault modes, even though neither arm concentrates.
Evidence
Validity arm passes on 100 held-out healthy engines with 12 sensors reaching at least 5% of top-1. Primary arm: healthy top-1 entropy 2.509, fault 2.407, gap +0.102 against a pre-registered ≥0.40 bar, p = 0.0105 — FAIL. Discrimination arm: total-variation distance between the two fault modes' leading-sensor distributions 0.280, p = 0.0330 against a ≥0.25 bar — PASS. A third arm on a single marker sensor fails at p = 0.1231.
Scope / limits
Simulated data, two fleets, one engine family. Rankings are read from the cached attribution state rather than recomputed, so this run and the run it corrects are comparable on representation. The p of 0.0105 is NOT to be rounded into significance — the criterion failed as registered and the effect size misses by four times regardless.
Controls that could have killed it
Claim
The per-channel permutation null — built after an earlier version named channels on pure noise — holds on real degrading hardware: on engines that have NOT degraded the module names a sensor on 4 of 100 and 1 of 100. But on degraded engines it names 89–92% of all sensors, and that is CORRECT rather than broken, which makes the binary 'is this channel significant' question useless on coupled machinery.
Evidence
400 engine-runs (100 fault + 100 healthy on each of two fleets), 15 sensors common to both, permutation count sized from the corrected threshold it has to resolve (401 against a Bonferroni bar of 0.00333, margin ×1.34, checked before running). On unit 1, 14 of the 15 sensors have genuinely moved more than 1 sd from their own baseline — NRf +8.2, T50 +7.5, Ps30 +6.0, Nf +5.8, P30 −5.7. Across all 200 degraded engines the naming rate is 0.89 and 0.92.
Scope / limits
Simulated run-to-failure data, one engine family, no per-sensor ground truth — the dataset labels fault MODES, not sensors. The series is truncated to the healthy reference plus the scored window, so nothing here speaks to WHEN divergence started.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW WITHDRAWN: 'How far above its current position an engine had already been — one number, the excursion — predicts which of two matched engines fails sooner, and it replicates pre-registered on two unseen fleets at 0.721 and 0.764, p < 0.0001, against a control blind at exactly 0.500.'
Evidence
The replication is real and its numbers stand. What was missing was a test of whether the quantity survives the arbitrary choice inside it.
Scope / limits
RETIRED as a claim 2026-08-07. The measurement is not disowned; the interpretation is.
Controls that could have killed it
Why this was wrong
The excursion moves with a knob nobody has justified. Across blind configurations of the baseline it spans 0.533 to 0.795 on one fleet — a range that contains both 'strong result' and 'nothing'. The sensitivity was then localised to the REFERENCE distribution, and no baseline-free formulation recovers the signal: 0.535 and 0.604 against a 0.65 bar, with the radius forms sitting near chance at 0.51–0.57. So stability and signal turned out to live in different places — the form that is stable carries nothing, and the form that carries something is unstable. A pre-registered replication and an unresolved dependence on a modelling choice travel together or neither does. ⚠ This does NOT prove the effect is pure artefact: a radius upper-bounds a signed overshoot, so the failure of the baseline-free version proves only that no baseline-free version has been found. Those are different statements and merging them would be a second error.
Claim
On fleets that operate in six different flight conditions, treating the vehicle as one regime scores at chance. Giving each engine its own baseline PER CONDITION, read from operational settings alone, recovers the measurement — and the excursion replicates on two fleets that had never been touched.
Evidence
Control blind at 0.500 EXACTLY on 4 of 4 fleets (482 engines, 6 conditions, 2 fault modes). The naive single-regime treatment collapses to 0.516 and 0.509 — chance — against 0.721 and 0.764 for the per-condition version, permutation p < 0.0001 on both. The margin over the dynamics operators is +0.048 on FD002 against a pre-registered +0.10 bar, so 'decisively beats the operators' is NOT claimed and was recorded as a fail.
Scope / limits
Simulated run-to-failure data, one engine family. Baselines are set from an engine's first 8 flights in each condition; nothing is borrowed from another engine or a population model.
Controls that could have killed it
Claim
Given two engines equally far from their own baseline at the same age, one of which will fail much sooner, the differential-operator family built for exactly this question (velocity, acceleration, curvature as a three-vector) does NOT tell them apart on a fleet it has never seen. A single path statistic does.
Evidence
45 disjoint matched pairs, same cycle, distance-from-own-baseline matched to 0.091%; sooner-failing engine has mean remaining life 76 cycles against 149. On the held-out fleet FD003 — 100 engines untouched by any earlier choice — the dynamics three-vector scores 0.594 with a sign-flip permutation p = 0.2675, which is chance. Dynamics minus the path baseline is −0.156. A curvature direction FIXED IN ADVANCE scored 0.562 held-out against 0.737 on the fleet it was first spotted in. What did carry the signal is one number off the distance curve: how far above its current position the engine had already been.
Scope / limits
Simulated run-to-failure data from one engine family (NASA C-MAPSS). This is not flight telemetry and a result here is not a claim about spacecraft. One window size; the ordering between arms flips across the window sweep.
Controls that could have killed it
Claim
Findings 34 and 36 excluded windows straddling a rhythm change — the onsets and offsets a monitor exists to catch — which made their numbers non-comparable to published detectors evaluated on continuous streams. Scoring every window with nothing excluded, a mixed window labelled by its majority class: accuracy 0.966 across 9 patients, inside the published 95-98% band on the harder protocol. Transition windows do cost accuracy (0.902 on transitions vs 0.969 on pure windows) but they do not move the stream result out of band. Separately, onset is not slow: across 14 AF episodes the loop's first AF call comes on the FIRST window of the episode in every case (median lag 0 windows, 90th percentile 0).
Evidence
Per-patient stores of 64 pure windows per state (finding 36's configuration); queries are the remainder of that patient's stream IN ORDER with nothing filtered, up to 150 windows each. Per-patient stream accuracy ranged 0.873-1.000, with the weakest patient the one carrying by far the most transitions (28 of 150 windows). Onset lag measured only on true AF runs of at least 3 windows.
Scope / limits
Per-patient only; no cross-patient claim. One database, seed 0, one window length (64 intervals), majority-label convention for mixed windows. Published detectors differ in episode-level scoring conventions, so this is a fair-protocol comparison rather than an identical-protocol one. 14 episodes is a small sample for the onset claim.
Controls that could have killed it
Claim
Finding 35 recorded the loop BELOW the published 95-98% band at 0.935. Two follow-ups locate why, and it is not the method. First, a tune-on-dev / judge-on-held-out sweep over window length, sub-window and bin count moved held-out accuracy by exactly nothing (0.940 for both the swept-best and the hand-picked default, while the dev-set best rose 0.925 -> 0.945) — the hyperparameters are not the limit, and the dev gain was overfitting. Second, extending the store past the arbitrary 16 examples per state: accuracy 0.935 (16) -> 0.967 (32) -> 0.983 (64), still rising, with the plain-L2 control trailing at every level (0.898/0.942/0.947). The published comparison was between a tuned literature using full training data and a loop given 16 examples; matched for data, the loop reaches the top of that band.
Evidence
Coverage extension on 9 patients (one lacks enough windows at 64), 40 queries each, per-patient stores, same representation and splits as finding 34. Geometry vs plain L2: +3.7 / +2.5 / +3.6 points at 16 / 32 / 64. Tuning experiment: 11 configurations swept on 5 dev patients, verdict measured only on 5 patients never seen by the sweep; dev range 0.900-0.945, held-out difference between chosen and default 0.000.
Scope / limits
IMPORTANT AND UNFLATTERING: windows whose label is not pure are EXCLUDED, so rhythm-transition windows — onset and offset, the genuinely hard part — are not scored here, while published detectors are usually evaluated on continuous streams that include them. The comparison is therefore indicative, not like-for-like, and the number should not be quoted as beating those methods. Per-patient only; no cross-patient claim. 64 examples per state is roughly an hour of that patient's own labelled rhythm. One database, seed 0.
Controls that could have killed it
Claim
Two honest comparisons, both unfavourable. (1) Published RR-interval AF detectors report 95-98% accuracy on this database (sensitivity ~96-97%, specificity ~97-98%); the untuned loop of finding 34 reaches 93.5%, so on raw accuracy this approach is BEHIND the field, not ahead of it. (2) With store size held equal, a store of the patient's OWN windows beats a store built from other patients' windows by only +2.7 points accuracy and +3.7 points specificity on average across 10 patients — both far under the pre-registered 10-point bars. The headline personalisation pitch does not survive its own test.
Evidence
Personal vs population stores, 16 windows per class in BOTH arms so the comparison isolates WHOSE memories rather than how many, identical queries and splits per patient. Mean gains: accuracy +0.027, specificity +0.037, sensitivity -0.003. Four of ten patients showed exactly zero gain because the population store already scored 1.000 on them; three showed small negative gains. State-of-the-art figures are from the published literature on the same database, not re-measured here.
Scope / limits
One database, one window size, one representation, seed 0, per-patient evaluation. The loop is untuned — no window/bin/representation search was run — while published methods are tuned, so the accuracy comparison flatters the field somewhat; that does not change the direction of the result.
Controls that could have killed it
Claim
Finding 33 failed on single-beat shape. Asked instead about RHYTHM — atrial fibrillation, which is defined by irregularity over time and therefore sits on the temporal-evolution side of the two-lane boundary — the same loop works: per-patient recognition of AF vs that patient's own normal rhythm reaches 0.935 at 16 stored windows per state (chance 0.50), and the decode leg lands in-class 0.89-0.92 at every coverage level — the leg that never rose at all on beat morphology. The geometry beats a plain-L2 control on identical windows at all four coverage levels (+3.0/+5.5/+4.3/+3.7 points).
Evidence
10 patients from a public AF database, each containing both their own AF and their own non-AF periods, so the store holds that patient's own examples of each state. 64-interval windows, window-disjoint store/judge/query splits, transition windows excluded (labels must be pure), unblended 1-NN judge, 60 queries per patient. Recognition by coverage (2/4/8/16 windows per state): 0.895/0.888/0.915/0.935. Plain L2 on identical windows: 0.865/0.833/0.872/0.898. Decode: 0.892/0.915/0.907/0.895. BOTH pre-registered criteria FAILED AS WRITTEN and are reported as such: A1 required monotonic rise AND >= 0.85 — the level was cleared comfortably but a 0.7-point dip between the first two coverage points breaks monotonicity; A2 required the geometry to beat L2 by >= 5 points at top coverage — the margin there is 3.7, though the geometry leads at every level and by 5.5 at one of them.
Scope / limits
Per-patient only. No cross-patient claim is made or tested — that is the transfer problem finding 31 leaves open, in a new domain. One database, one window size, seed 0, deltas-of-intervals representation. Two states, chance 0.50.
Controls that could have killed it
Claim
Asked of a medical signal — one cardiac patient's heartbeats, 5 rhythm classes — the coverage law fails both pre-registered bars. Recognition does rise with coverage (0.370 -> 0.428 -> 0.517 -> 0.517 over stores of 1/2/4/8 examples per class, against a 0.200 chance floor) but plateaus at roughly half the bar of 0.85, and the decode leg never rises at all (0.407/0.450/0.417/0.467). The direction of the law transfers; its magnitude is modality-dependent.
Evidence
3 seeds per coverage point, beat-disjoint store/query/judge splits, unblended 1-NN judge, the same protocol and criteria as the speech confirmation. Deciders track each other closely (vote 0.370/0.411/0.472/0.511 vs prototype 0.370/0.422/0.511/0.483) — no regime flip is visible here, unlike speech.
Scope / limits
One dataset, one patient, one encoder (the magnitude-weighted slope-bin representation already used for this data). This bounds the coverage law; it does not bound memory-based learning in general, and it says nothing about medical signals whose identity IS carried by temporal evolution.
Controls that could have killed it
Claim
Which decision rule the loop should use depends on the memory regime. Across speakers (sparse, shifted), classification by consolidated barycentre prototypes beats the distance-weighted k-NN vote by 9.7 points mean over 3 seeds (0.672 vs 0.575; per-seed gains +10.5/+3.5/+15.0) — the loop as first measured in finding 28 was mis-wired. At dense personal coverage the ordering REVERSES: the raw vote wins (0.917/0.858/0.925) over prototypes (0.833/0.850/0.833), because the retrieved neighbourhood is the same speaker's own takes. Neither rule is 'the' decider; the loop should choose by regime.
Evidence
Same protocol and seeds as findings 28/31. Additionally: the decode evaluation itself is sensitive to finding 30's blending result — judging single-speaker responses against mixed-speaker barycentre prototypes suppressed the measured in-class rate (0.750 at t4) versus a 1-NN judge over raw clips (0.800 on the identical responses, registered before running as the final judge variant). Blending drifts artifacts off every real speaker's manifold, and that applies to yardsticks exactly as it applies to outputs.
Scope / limits
One dataset, one modality, 3 seeds. The regime boundary (where the ordering flips between prototype and vote) is not yet mapped — only its two endpoints are measured.
Controls that could have killed it
Claim
When the store contains the speaker's own examples of each word — the personalization condition the architecture is for — the full retrieve-and-read-out loop crosses both pre-registered bars, on 3 seeds with full query sets: recognition 0.900 mean (0.917/0.858/0.925; bar 0.85) and decoded responses in-class 0.872 mean (0.867/0.850/0.900; bar 0.80, unblended 1-NN judge). Both legs rise monotonically with coverage (recognition 0.575 -> 0.800 -> ~0.9; decode 0.400 -> 0.700 -> ~0.87 as coverage goes none -> 2 -> 4 examples per voice-word pair), and every improvement costs exactly one append — no gradient step exists anywhere in the system.
Evidence
3 seeds x 120 queries each, store = 4 takes per (speaker, digit), queries = held-out takes of the same pairs, yardstick = a third disjoint slice. Mechanism measured, not argued: 0.964 of retrievals return a memory of the SAME speaker (per-seed 0.967/0.967/0.958) — the representation entangles who with what, which is simultaneously why the across-speaker condition fails (finding 30: right word, wrong voice) and why the personal condition works. The coverage curve's earlier points come from the seed-0 probes at sparse and 2-take coverage.
Scope / limits
One dataset (spoken digits), one modality; the across-speaker condition remains open (finding 30) and identity-preserving quotients are the proposed route, not a result. Decode bar was originally hit exactly (0.800) on a 40-query subsample; the full 120-query sets confirm above it.
Controls that could have killed it
Claim
Three pre-registered follow-ups to finding 28's decode gap. (1) Blending hurts generation monotonically: reading out the single nearest memory verbatim lands in-class 0.475; the barycentre of the vote-winning cohort, 0.400; the consolidated all-members class prototype, 0.325 — the more memories are averaged into the artifact, the worse it gets. (2) The gap is structural, not data-starved: the best strategy swept at 5/10/20 examples per class gives 0.475/0.425/0.575 — no monotone rise, nowhere near the registered 0.70 bar. (3) At fixed shortlist width, index recall@5 degrades as the store grows (width 32: 0.980 -> 0.897 -> 0.797 over stores 50/100/200), yet the index arm's task accuracy is at or above the exact path at every width on the larger stores — the prefilter's misses are net-positive for the vote.
Evidence
Decode arms at 5 shots, 40 queries, identical yardstick as finding 28's L5 (held-out-speaker prototypes): nearest 0.475 / raw cohort 0.400 / consolidated 0.325. Shot sweep of the best arm: 0.475 (store 50) / 0.425 (100) / 0.575 (200). Index scaling: exact vote accuracy saturates across speakers at 0.483/0.500/0.483 while the store grows 4x; operating width (first width reproducing exact answers) 32 at store 50, 16 at store 100, none at store 200 — where every tested width scores ABOVE exact (0.550/0.517/0.517/0.500 vs 0.483), so the registered two-sided match criterion fails in the direction of improvement. Registered sublinear-width criterion: FAIL as stated.
Scope / limits
Single seed, one across-speaker split, 40-query decode subsample and 60-query index subsample — at 60 queries one answer = 1.7 points, so the 1-point match criterion effectively demands identical answers. Follow-up probes; any headline use gets the full 3-seed, 200-query treatment first.
Controls that could have killed it
Claim
Building the identical 50-entry store all at once versus one class at a time produces identical per-class vote accuracy on every one of 10 classes — the end state provably does not depend on arrival order, through the full distance-weighted cohort vote, not just 1-NN recall.
Evidence
Per-class accuracies batch vs incremental: 0.350/0.450/0.300/0.150/0.550/0.900/0.400/0.950/0.800/0.550 — equal to the third decimal in all 10 cases. This is the non-degenerate control for finding 28's failed retention criterion: the first class's fall from 1.000 (alone in the store) to 0.350 (nine rivals present) is entirely rival arrival — task difficulty growing — and 0% forgetting.
Scope / limits
Deterministic retrieval over an append-only store makes this near-structural; the measurement confirms the composed vote path introduces no order dependence either. One dataset, one split, seed 0.
Controls that could have killed it
Claim
The full retrieve-compose-decode loop, run end-to-end with pre-registered criteria on across-speaker spoken digits (5 shots/class): a distance-weighted 5-NN vote beats the single nearest memory by +2.7 points (positive on 3/3 seeds); a random-cohort control collapses to 0.175, so the vote is carried by which memories are retrieved; every class is recognisable the instant its five examples are appended, with zero gradient steps. The generation leg is the measured frontier: response trajectories composed from the vote-winning cohort land in-class only 0.400 of the time — below the vote's own accuracy.
Evidence
Vote 0.575 vs 1-NN 0.548 (mean of seeds 0/1/2; smallest seed gain +0.5pt — the effect is real but modest). Random cohorts weighted by their real distances: 0.175. Instant learning: several classes at 1.000 immediately after append. Composed decode: 16/40 in-class against held-out-speaker prototypes. Two additional criteria FAILED AS REGISTERED and are kept: a retention criterion whose baseline was degenerate (measured the first class when it was the only class — a one-label classifier scores 1.000 for free; the drop it 'detected' is label-space growth, not damage — see finding 29 for the non-degenerate control), and an index-match criterion missed by literal floating-point epsilon (the index changed exactly 2 answers in 200 at shortlist 32, the registered 1-point boundary).
Scope / limits
One dataset, one across-speaker split, 3 seeds for the composition arms and 1 for the rest. Recognition through the loop is the claim; generation through the loop is explicitly NOT closed at this shot count and split.
Controls that could have killed it
Claim
Continuing the curve along its geodesic — a closed-form step with zero learned parameters — predicts the next frame of a real signal 3-16x more accurately than a matched dense network that was trained to do it.
Evidence
Fisher-Rao error, lower is better. AUDIO (FSDD cochlea, F=32): geodesic continuation 0.0054 vs fair dense 0.0860 vs naive repeat-last 0.0866. PEN (handwriting, F=16): 0.1796 vs 0.3863 vs 0.3903. Paired: audio 646 wins / 2 losses, p=4e-190; pen 591/129, p=2e-71. The LEARNED geometry model beats fair dense on audio (p=2e-122) but only TIES on pen (397/353, p=0.12) — reported because it qualifies the headline: learning helps complex dynamics, not smooth real signals.
Scope / limits
Next-step prediction on continuous real sequences, two unrelated modalities. This is one-step continuation, NOT generation over long horizons, and not language.
Controls that could have killed it
Claim
Trained only to predict the next character, the model places characters on the simplex so that Fisher-Rao distance tracks linguistic class -- case pairs, vowels, sentence terminators, whitespace -- and this is not explained by how often the characters occur.
Evidence
Off-diagonal Fisher-Rao spread 0.5050 (min 0.1298, max 0.6349) against exactly 0.0000 for one-hot. Correlation with frequency is nil: Spearman +0.009, ~0.0% of rank variance. After residualising distance on the frequency gap, within-class pairs are closer than cross-class pairs by 0.0226; a 10,000-shuffle permutation null over class labels gives z=+4.46, p=0.00010. Tightest classes are the terminators . ? ! (-0.256) and separators , ; : (-0.212).
Scope / limits
One trained checkpoint (HODOS_embed, 30k steps, val ppl 5.29) on TinyShakespeare. Learned embeddings clustering by character class is a KNOWN property of neural language models and is not claimed as novel here. It is recorded because it is the precondition that makes finding 25 interpretable: without it there is no metric for the Fisher-Rao ground to act on.
Controls that could have killed it
Claim
Given character positions that carry a real metric, retrieving past contexts under the Hodos process distance predicts the next character better than retrieving under plain L2 over the identical frames. The only difference between the two arms is the distance function.
Evidence
Pooled over 5 seeds, 1000 held-out positions, each seed tuning its own temperature and blend weight on its own separate half: Hodos beats L2 by +0.0970 nats, p=2.0e-05. The warp HELPS by 0.0674 nats (p=2.5e-03) and the ground helps by 0.0296 nats (p=0.024) -- both reversing the retracted findings 17 and 24. The same experiment run in the old one-hot space reads -0.0249 nats, p=0.25: indistinguishable from zero, as a constant ground must be.
Scope / limits
k-NN character language modelling on TinyShakespeare against a unigram term. The learned metric is the 4,225-parameter embedding table from the trained HODOS_embed language model, so this shows the geometry works ON a learned metric -- it does not show the metric can be obtained without learning one.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'Retrieval is what helps; the geometric distance is not better than a flat one for this.'
Evidence
Memory improves by +0.105 nats, p=2.7e-06. Plain L2 retrieval beats Hodos retrieval by 0.121 nats, p=7.35e-04, and this survives per-arm temperature tuning.
Scope / limits
RETRACTED 2026-07-26. The geometry was never given a metric to work with.
Controls that could have killed it
Why this was wrong
Both of these were measured in a space where the Fisher-Rao ground is a CONSTANT FUNCTION. Contexts were built as label-smoothed one-hot rows, and on the actual vocabulary (V=65, eps=0.03) the Fisher-Rao distance between every pair of distinct characters is identical: min 2.9971, mean 2.9971, max 2.9971 -- spread exactly 0.0000. In that space Fisher-Rao and L2 induce the SAME ranking, so 'the ground ties L2' was an algebraic identity, not a measurement, and the warp was the only thing left that could move. The experiment could not have produced any other answer. Re-measured in a non-degenerate metric (finding 25) both conclusions reverse: the ground beats L2 and the warp HELPS. The H1 half of this finding -- that memory helps a language model at all -- survives and is carried into finding 25. Only the control's verdict on the geometry is withdrawn.
Claim
The apparent 'empty band' in character separations is an artifact of one-hot encoding, not a repulsion signature.
Evidence
One-hot characters have exactly ONE distinct pairwise distance — all 2080 pairs at 1.498555, spread 0.00e+00. Arbitrary points on the SAME simplex give 1334 distinct distances, so the constraint is the encoding.
Scope / limits
The structural question is genuinely shared — both ask what a distribution of separations looks like near zero. But level repulsion is a statement about a DENSE 1-D spectrum, and 65 points in 65 dimensions is the sparsest possible arrangement.
Controls that could have killed it
Claim
The distance works as a REFEREE — it correctly ranks a quantum-chaotic drum closest to the real Riemann zeros — but every attempt to use it to find a load-bearing arithmetic operator returned a null.
Evidence
Referee ranks GUE 0.0435 closest, GOE 0.0695 second, noise far. Reproduces Montgomery–Odlyzko (known). The relationship-kernel looked load-bearing at 4–6σ against random integers — but a density-matched control scored BETTER (0.062 vs 0.128), so the signal was a density-envelope artifact.
Scope / limits
RH is not proven, not approached. The real limiter is resolution: an arithmetic fingerprint needs ~10³–10⁴ zeros and we have hundreds.
Controls that could have killed it
Claim
A negative forgetting score — sequential beating joint — is impossible if the store is order-invariant, and chasing that impossibility found a real defect in the comparison.
Evidence
Handwriting reported hodos forgetting −0.083 on one seed. Cause: an incomplete final chunk is dropped from the task list, but the joint comparison arms still iterated ALL classes — so joint faced a 17-way problem while sequential faced 16-way.
Scope / limits
Speech divides evenly into pairs, so nothing was orphaned there and the flaw could ONLY surface on a second modality.
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'the geometry forgets too — task-1 accuracy drops 0.500 as later tasks arrive.'
Evidence
The raw drop, which is real but means something else.
Scope / limits
RETRACTED same day. This is a flaw in the METRIC, not in either system.
Controls that could have killed it
Why this was wrong
In class-incremental learning the LABEL SPACE GROWS: task 1 is a 2-way choice when first learned and a 10-way choice at the end. Comparing those two numbers makes the task getting harder look like memory degrading. Measuring against the joint-trained ceiling makes the growing label space cancel, leaving only the cost of having learned in sequence — which for this system is exactly zero. The file's own docstring warned about this trap before the first version was written that way anyway.
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'sph_log takes arccos of an inner product very close to 1.0; fp16 resolution there is coarser than the guard clamp, so every Log map collapses to zero and the geometry dies silently.'
Evidence
None — this was reasoned, never measured. Rung 1.
Scope / limits
RETRACTED 2026-07-24 by direct measurement on a T4.
Controls that could have killed it
Why this was wrong
Both halves were wrong. The premise is BACKWARDS: characters are label-smoothed one-hots, so two DIFFERENT consecutive characters are nearly ORTHOGONAL (real-text mean inner product 0.096), nowhere near 1.0. The fp16 danger band contains 0.0000 of real pairs and zero pairs lose real motion. And naive autocast never reaches the geometry anyway — the manifold tensors come out float32. The 'safety' code written to guard this imaginary hazard then introduced the only real bug the run found.
Claim
Lifting a trajectory into nested levels gives no advantage over the flat version on either modality.
Evidence
Speech: flat 0.920, lifted 0.935 — but the SMOOTHING control also 0.935, identical. Handwriting: lifted 0.870 < flat 0.877.
Scope / limits
FSDD utterances are 24 frames, so a lift yields ~5 — shallow. A null here means 'no benefit on short sequences', NOT 'no benefit'. The case hierarchy exists for is long sequences.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'For character contexts, the Fisher-Rao ground ties plain L2 exactly, and the entire performance gap is the alignment machinery.'
Evidence
Per-arm tuned: full Hodos 3.1117, ground-without-warp 2.9949, L2 2.9903. Ground vs L2: −0.0046, p=0.586 — a dead tie. The warp costs +0.1168, the whole gap.
Scope / limits
RETRACTED 2026-07-26. The tie was forced by the representation, not observed.
Controls that could have killed it
Why this was wrong
Both of these were measured in a space where the Fisher-Rao ground is a CONSTANT FUNCTION. Contexts were built as label-smoothed one-hot rows, and on the actual vocabulary (V=65, eps=0.03) the Fisher-Rao distance between every pair of distinct characters is identical: min 2.9971, mean 2.9971, max 2.9971 -- spread exactly 0.0000. In that space Fisher-Rao and L2 induce the SAME ranking, so 'the ground ties L2' was an algebraic identity, not a measurement, and the warp was the only thing left that could move. The experiment could not have produced any other answer. Re-measured in a non-degenerate metric (finding 25) both conclusions reverse: the ground beats L2 and the warp HELPS.
Claim
A randomly-initialised readout layer costs ~21 perplexity — a 2.5× degradation — regardless of its form. Initialised as a pass-through, the cost vanishes entirely.
Evidence
At identical dimension and parameters: no readout 13.528, random-init readout 34.318, identity-init readout 13.187. Both readout FORMS cost the same, so it is not which readout but having one.
Scope / limits
Applies to anything built downstream of a landed point — the memory, reasoning and decoder organs all sit there.
Controls that could have killed it
Claim
Squeezing characters onto a lower-dimensional concept simplex is monotonically worse, at matched parameters, with a working readout.
Evidence
13.545 (65 dims) → 16.089 (32) → 19.418 (16) → 23.113 (8) perplexity. No sweet spot.
Scope / limits
Character language modelling, small local config.
Controls that could have killed it
Claim
Taught classes in sequence, a standard network loses half to four-fifths of its first-task accuracy. This system loses nothing, because learning is appending to a store rather than overwriting weights.
Evidence
Forgetting measured against the joint-trained ceiling so the growing label space cancels. Speech: net +0.500, hodos +0.000. Handwriting: net +0.801, hodos +0.000. Per-seed hodos [0,0,0] on both. Learning INCREMENTALLY beats the network trained on EVERYTHING AT ONCE: 0.544 vs 0.422 (speech), 0.845 vs 0.702 (handwriting).
Scope / limits
Two modalities, 3 seeds each. Handwriting uses a random per-class holdout (no speaker axis), a weaker split than the across-speaker speech version.
Controls that could have killed it
Claim
Class-barycenter prototypes with NO training reach a standard method's 50-example accuracy using 5 examples (speech) or 10 (handwriting).
Evidence
Speech across-speaker: hodos 0.416/0.539/0.654/0.715/0.746 at 1/2/5/10/50 shots vs GRU 0.189/0.291/0.417/0.520/0.624. Handwriting: hodos 0.730→0.894 vs GRU 0.403→0.828. At ONE example the network is near chance (0.189 vs 0.10).
Scope / limits
Two modalities, 5 seeds, 200 test items, full shot sweep to 50.
Controls that could have killed it
Claim
The advantage over mean-pooling splits cleanly into three ingredients in a fixed order: using the trajectory at all, then the Fisher–Rao ground, then the alignment.
Evidence
FSDD across-speaker 5-shot, 3 seeds: FR+warp 0.657 / FR+nowarp 0.637 / L2raw+warp 0.602 / L2raw+nowarp 0.573 / mean-pool 0.483. Trajectory +9, ground +5–6, warp +2–3. Handwriting: same ordering, +25 / +3.4 / +1.8.
Scope / limits
Same ordering on a sound task AND a non-sound task, so it is not a speech artifact.
Controls that could have killed it
Claim
soft-DTW beats our hard-DTW warp — but the Fisher–Rao ground beats Euclidean under BOTH warps on BOTH modalities. The transferable contribution is the ground, and it composes with the field's better alignment.
Evidence
Handwriting z=2.75, errors 0.045→0.025 for soft-DTW over hard-DTW.
Scope / limits
Recommended distance is therefore soft-DTW + Fisher–Rao — adopt the field's warp, keep our geometry.
Controls that could have killed it
Claim
The conceptual link (Doppler is a time-warp, so the method should measure velocity) is real and elegant, but empirically purpose-built DSP measures it better.
Evidence
Matched filter is perfect (0.000 clean); Hodos ~0.13. On the velocity PROFILE — supposedly the structural edge — the Hodos warp-slope is too noisy (0.14–0.20) and loses even to a FLAT single-α baseline (0.07–0.12).
Scope / limits
Two fair benchmarks. Stopped after two rather than trying a third variant.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'min-MEAN + Fisher–Rao beats standard DTW on light by +0.21 (0.61→0.82), and it survives 6→12/class.'
Evidence
The apparent win.
Scope / limits
RETRACTED — the metric tracked one condition better than the other with no mechanism to explain it.
Controls that could have killed it
Why this was wrong
min-mean degenerates into long, padded, cheap-repeat paths (path padding 0.00 for standard DTW vs 0.61 for min-mean) which makes it behave MORE like plain averaging (correlation with mean-pool 0.779 → 0.845). Light is precisely the task where averaging wins — so it 'won' for the wrong reason, and the same collapse HURT on echo where timing is the signal. The object is Marzal–Vidal's Normalized Edit Distance (1993), proven non-metric. RETIRED.
Claim
The Fisher–Rao ground shows no edge over plain L2 on light spectra, sonar range-profiles or turbulence. It ties; it does not beat.
Evidence
1-NN LOO, 3 seeds, ground ablation, noise swept to break the ceiling. At 18/class: light 0.45 vs 0.44, echo 0.73 vs 0.75, turbulence 0.62 vs 0.68. The ground earns its keep in 0 of 3 modalities.
Scope / limits
Recognition only. Does not test barycenter, morph, pooling or the physics reproduction.
Controls that could have killed it
Claim
Quotienting the process space by the global bin-shift group gives a distance between ORBITS, making recognition invariant to transposition exactly — not approximately, and with no examples required.
Evidence
Same vowel shifted 6 bins: distance collapses 61× (0.4611→0.0075) while a DIFFERENT vowel is preserved (0.2911) — a 39× separation. The invariant recogniser REPAIRS a real base-engine misclassification.
Scope / limits
Exact-zero only for an exact array roll; independently synthesised transpositions carry edge-tail truncation (D_inv ≈ 7.5e-3, not 1e-16). Pad-mode invariance is approximate.
Controls that could have killed it
Claim
The GROUND earns its keep wherever identity lives in the relative-space pattern — INCLUDING static posture. The WARP earns its keep only where identity lives in temporal motion. These are separate lanes; the earlier single-axis 'motion vs static' framing conflated them.
Evidence
With a representation that preserves relative space (distribution over TOTAL acceleration orientation), subject-disjoint sit-vs-stand = 0.758 against chance 0.500. Fisher–Rao ground contributes +5.8 points over Euclidean on a task with NO motion; the warp contributes 0, correctly.
Scope / limits
One dataset (UCI HAR). The two-lane structure is supported by findings 3, 4 and 17 jointly. PRIOR ART, added 2026-07-26: the same demarcation is already stated in the pooling literature — orderless pooling suffices for stationary or weakly nonstationary signals with little discriminative temporal structure, while order-aware pooling is needed where the cues are transient or phase-dependent. This finding independently REPRODUCES a known boundary; it does not establish a new one. It stays active because it is measured and it is true.
Controls that could have killed it
Claim
ORIGINAL WORDING, NOW KNOWN TO BE WRONG: 'Fisher–Rao does not beat Euclidean on human-activity windows because activity identity in a short window is static — exactly what the principle predicts.'
Evidence
HAR: averaging won, 0.24 vs 0.26–0.27.
Scope / limits
RETRACTED 2026-07-24. The null was OUR OWN BROKEN INPUT, not a property of the method.
Why this was wrong
The adapter fed gravity-removed acceleration reduced to a MAGNITUDE — a representation with direction deleted. Sitting and standing are both 'at rest, no direction', so under that input they were identical BY CONSTRUCTION. The method was not failing on static identity; it was succeeding on an input from which the answer had already been removed. A null produced by a crippled representation is not a boundary.
Claim
The identical claim holds on pen motion — a completely different modality — and more strongly than on speech.
Evidence
UCI Character Trajectories, 2858 real handwritten characters, 4000 triples: paired McNemar z≈4.95 (Fisher–Rao fixes 49 errors, breaks 10).
Scope / limits
Single-writer temporal recognition, not an across-body handwriting split. Not a published leaderboard.
Controls that could have killed it
Claim
Recognising the same spoken word through a DIFFERENT human voice, the Fisher–Rao ground reduces error ~17% vs Euclidean.
Evidence
FSDD, 2000 triples, fresh seed: 0.1500±0.0080 vs 0.1805±0.0086 (~2.6σ).
Scope / limits
Word-level ABX on FSDD, not the published phone-level ZeroSpeech leaderboard. Within-speaker is a TIE — the N=400 within-speaker 'win' did NOT replicate and was a small-sample mirage.
Controls that could have killed it
Claim
Averaging a signal over time loses a short event at O(ρ²) while the curve distance loses it at Θ(ρ) — so pooling destroys exactly what a process comparison keeps.
Evidence
Log-log slopes 1.90 vs 1.00.
Scope / limits
Holds under an interior/support condition (corrected on review, 2026-07-22). A novelty check found the underlying effect is KNOWN prior art in detection theory and the DTW-averaging literature — the exponent statement may be a small novel instance at most. Do NOT claim as a new theorem.
Controls that could have killed it
Claim
The map p ↦ √p sends the simplex onto the unit sphere, under which Fisher–Rao becomes ordinary great-circle geometry. A signal over time is therefore a curve on a sphere, and comparing signals is comparing curves.
Evidence
2·geodesic == fisher_rao to 1e-9; chi² ≈ ¼·d_FR² with correlation 0.987 on real audio frames.
Scope / limits
Mathematical identity, not an empirical claim. Prior art (information geometry); reused, not invented.
Every equation the project rests on, with an honest label on each: ASSEMBLED from known parts, KNOWN and merely reused, or OURS. Read ASSEMBLED carefully — it is not a lesser category. Nearly every method in this field is a composition, and §1, the Hodos distance itself, is ASSEMBLED: the contribution is which ground it is composed with, and that composition is not something we found elsewhere. What the labels buy is that nothing here can be dismissed by pointing at a part of it.
A signal becomes a process P = (pⁱ, …, pᵀ), each frame a probability distribution over F bins. Δ^{F−1} is the probability simplex. BC(p,q) = Σᵇ √(pᵇ qᵇ) is the Bhattacharyya coefficient.
ASSEMBLED — the composition IS the contribution
D(P,Q) = ( min_{π ∈ Π_β} Σ_{(i,j)∈π} g(p⁽ⁱ⁾, q⁽ʲ⁾) ) / |π★|Plain: Line the two signals up in time so the total per-step difference is smallest, then divide by the length of that alignment.
Provenance: Dynamic Time Warping (Sakoe–Chiba 1978) composed with an information-geometry ground g. The warp is standard and reused; the ground is classical; THE COMPOSITION IS THE CONTRIBUTION — an elastic distance whose local cost is the exact Fisher–Rao geodesic between distributions, with closed-form geodesic operators on top. A literature check on 2026-07-26 found the neighbouring families (OTW, TiOT, TAOT, Riemannian Time Warping) but not this composition as a working engine with these operators. D is an alignment-tolerant divergence, NOT a metric; ordinary DTW fails the triangle inequality.
Correction: An earlier version of this sheet presented the minimum-MEAN alignment as 'the distinctive object'. That is RETIRED: minimising the mean directly is the Normalized Edit Distance (Marzal–Vidal 1993), whose optimum pads the path with cheap repeats and collapses toward plain averaging. Its one apparent benchmark win was that padding artifact (finding 9).
KNOWN
chi2(p,q) = ½ Σᵇ (pᵇ − qᵇ)² / (pᵇ + qᵇ) fisher_rao(p,q) = 2·arccos( BC(p,q) ) hellinger(p,q) = √( 1 − BC(p,q) ) wasserstein1 = Σᵇ | CDF_p(b) − CDF_q(b) |
Plain: Five interchangeable ways to measure the difference between two single frames. All closed-form, none learned.
Provenance: All established — Le Cam / Topsøe, Rao 1945, Amari, Bhattacharyya, optimal transport. We reuse them; we do not claim them. Our only move here is making the ground pluggable in one engine.
KNOWN, load-bearing framing
φ(p) = √p maps the simplex onto the positive orthant of S^{F−1}
‖√p‖² = Σᵇ pᵇ = 1
geodesic(p,q) = arccos( BC(p,q) )
d_FR(p,q) = 2·arccos( BC(p,q) )Plain: Taking the square root of a distribution puts it on a sphere, and the distance between distributions becomes the plain great-circle angle. So a signal over time is a curve on a sphere, and comparing signals is comparing curves.
Provenance: Classical information geometry (Rao 1945; Amari). Known — used as the exact geometric home of the engine. It is what makes the closed-form geodesics and means below possible.
KNOWN math, our operators
slerp(p,q,t) = unembed[ ( sin((1−t)θ)·√p + sin(tθ)·√q ) / sin θ ]
karcher_mean({pᵢ}) = the point m minimising Σᵢ d_FR(m, pᵢ)²Plain: How to walk between two distributions along the shortest path, and how to average a set of them without leaving the space.
Provenance: Standard Riemannian geometry and DTW-barycenter averaging (Petitjean 2011). Known. These drive the morph and barycenter operators, and the zero-training prototypes behind findings 13 and 27.
ASSEMBLED — downgraded from OURS on 2026-07-26
D(A,B) = Θ( ρ · g(s,s′) ) — LINEAR in ρ
g(ā, b̄) = O(ρ²) under the interior/support condition
= O(ρ) otherwise — the honest boundary
so D / g(ā,b̄) = Θ(1/ρ) → ∞Plain: Averaging a signal over time is QUADRATICALLY blind to a brief event riding on an ongoing one — the final /t/ against /p/ on a sustained vowel — while the warp distance stays LINEARLY sensitive. Averaging can only lose; on brief transients it loses badly.
Provenance: DOWNGRADED after the literature check this section had been carrying a warning about since 2026-07-22. It splits into three parts and they do not have equal status. (a) The O(ρ²) half is CLASSICAL: every smooth f-divergence is locally quadratic, its second derivative IS the Fisher information, and for chi-squared the literature states χ²(P₁‖P₂) = (θ₁−θ₂)² I_F + o(·) directly. Our proof is that expansion applied to the perturbation ρu. (b) Our 'honest boundary' — the O(ρ) case where mass moves where the baseline is ≈0 — is the classical regularity condition under which the local expansion holds at all. We rediscovered the caveat that ships with the theorem. (c) The warp's Θ(ρ) half is one line from the definition, and that averaging dilutes brief events is not only known but actively engineered around (attention pooling exists for it). The COMPOSED separation Θ(1/ρ)→∞ was not found anywhere — but six web searches are weak evidence of absence, and a composition of two textbook facts is a thin novelty claim regardless. Full write-up: paper/LITERATURE-CHECK-S5.md.
Correction: This section was tagged 'OURS — the main new result' from 2026-07-22 to 2026-07-26, always carrying its own warning that its novelty was unchecked. The check was run and the warning was justified. The measured consequence (finding 2 — averaging is quadratically blind to a brief event) STANDS and is unaffected; what changes is only the claim to have discovered the mathematics behind it.
OURS — new extension
D_inv(P,Q) = min_s D(P, σ_s Q) isometry: g(σ_s p, σ_s q) = g(p, q) EXACTLY, for the full cyclic group (verified to 4.4×10⁻¹⁶)
Plain: A pitch shift is a translation along a log-frequency axis, so quotient the space by that shift. What is left is relative formant spacing — which is identity. Invariance by construction, with nothing trained.
Provenance: OURS as an extension, BUT the idea — shift-invariant and transposition-invariant distances, chroma/CQT shift, quotient and orbit metrics — already exists in music information retrieval and group theory. A specific clean instance, not a first. Measured as finding 7: a same-vowel transposition collapses 61× while a different vowel is preserved.
KNOWN — reproductions, NOT ours
R₂(r) = 1 − ( sin(πr)/(πr) )² GUE pair correlation (Montgomery 1973) H = (XP + PX)/2 Berry–Keating operator Weil/Connes explicit-formula operator
Plain: Established results run THROUGH Hodos as a referee, not equations we made.
Provenance: Everything reproduces known mathematics. RH is NOT proven. Hodos's role is the measuring instrument — it correctly ranks the GUE spectrum closest to the real zeros. Two exciting signals died to fair controls here (findings 22 and 23), both density artifacts. The honest limiter is resolution: the arithmetic fingerprint needs 10³–10⁴ zeros and we compute a few hundred.
We assembled a real, working, model-free engine — DTW plus an information-geometry (Fisher–Rao √-sphere) ground over sequences of distributions — with clean geodesic operators and a quotient extension for pitch invariance. The composition in §1 and the quotient in §6 are the parts that are ours. §5 was demoted on 2026-07-26 when the literature check it had been flagged for was finally run. The strongest results in this project are not in this tab at all — they are the MEASUREMENTS (findings 13, 14 and 27), which were built with controls designed to kill them and survived.
Every rule here was written after something went wrong. They are not general good practice; they are scar tissue.
A test whose bar is set afterwards can always be read as a success. Finding 13 predicted that the advantage would shrink with data — the signature of a prior. It grew instead. That makes it a better representation rather than a better prior: a different claim, and one we would have described wrongly had the criterion not been fixed in advance.
Findings 8, 9, 18 and 22 were all killed by a control built for that purpose. Finding 18's smoothing arm matched the result exactly, which is what ended it. A benchmark with no arm capable of producing a negative is a demonstration, not a test.
Findings 5 and 19 both passed on inputs that could not have shown the effect. The activity adapter deleted direction, so two postures were identical by construction. The fp16 test used random data that never approaches the regime in question. A null from a crippled input is not a boundary.
Finding 9 won on one modality and lost on another with no account of why. Investigating the mechanism showed it was collapsing toward plain averaging — which is exactly why it won on the task where averaging wins. Explain the mechanism before banking the win.
Finding 23 returned a statistic that was identically zero for the real data and for all 900 shuffled controls. A z of zero against a null with zero spread is a dead instrument, not a null result. It is reported as a failed measurement.
Finding 21 reported a negative forgetting score, which cannot happen if the store is order-independent. Chasing it found a real defect — and fixing it made the opposing method look worse, meaning the bug had been flattering the opponent.
Speech and handwriting share no sensor, no physics and no preprocessing. Finding 4 is what turned finding 3 from a sound trick into a claim. And per finding 21, a second dataset is not only confirmation — it is a different defect detector.
8 findings on this page are retracted, with the original wording kept and labelled. 16 more are outright negatives. Deleting them would make the record look stronger and be worth less.
A retraction count on its own is misleading, so each one is typed. Withdrawing a claim that something worked and withdrawing a claim that something failed are opposite events, and lumping them together makes a record of self-correction read as a record of failure. Of the 8 here, 3 withdrew a positive claim — we said it worked and our own later measurement said otherwise — and 5 withdrew a negative one: we had recorded a failure or a limit of the method against ourselves, and we were wrong about it. The second kind is the more common on this page. Both are on it for the same reason.
| active | Survives the controls it was tested against; currently believed. |
| provisional | Measured, but on one task or at small scale; not yet replicated. |
| negative | Tested and did NOT hold. Reported because a negative is a result. |
| retracted | We claimed it, then our own measurement refuted it. Original wording kept visible. Each retraction is labelled by WHAT it withdrew: a positive claim (we said it worked) or a negative one (we said it did not work, or that it was a limit of the method). |
| untested | Stated but never measured. Carries no weight. |
| What it takes | |
|---|---|
| 1 | reasoned, unmeasured |
| 2 | measured once, synthetic data |
| 3 | measured on real data, one modality |
| 4 | replicated on a second, unrelated modality |
| 5 | survives a fair control designed to kill it |
| 6 | survives comparison against a published method |
| 7 | reproduced independently, outside this project |
Nothing is what it is in isolation. A thing — including a self — is the pattern of its connections, not its substrate. What a system is is constituted by what it stands in relation to; and what it does that exceeds its parts — emergence — arises from that pattern of relation, and is never installed as a component. The relations are not decoration laid over pre-existing things: the things are constituted by relations, and those relations hold between further processes that are themselves so constituted. There is no isolated substrate underneath at which the regress halts.
Published origin: 10.5281/zenodo.21613155 (2026-07-27). Named: 2026-08-07. The name refers back to a premise already published on 2026-07-27; naming it on 2026-08-07 does not create it and does not move its date. The premise is stated in the abstract of 10.5281/zenodo.21613155, given its own section 5, and declared in section 2 to be the one premise all five of the method's commitments are downstream of. It is stated earlier still in Your Past Loves You (begun 2026-07-19), which that paper cites as the plainest statement of the thesis the method is built on.
The regress does not halt. A part is prior to the whole it composes, and that priority is real — but it is local, not fundamental. Two halves must exist before one exists; the one is what their relation constitutes. Each half, in turn, is what the relation of two quarters constitutes. The series does not terminate in an unrelated object: at every level the thing exists because two or more further things stand in relation, and the same holds of those. So conceding that the components precede the emergence concedes nothing about substance — the components are themselves emergences of the same form, one level down. There is no floor at which relation stops and bare substrate begins, which is precisely why the premise applies at every scale rather than at a chosen one, and why an engine built on it is not restricted to any particular kind of matter or signal.
Why "hypothesis" and not "theory". It is called a hypothesis, not a theory, because independent outside replication has not happened yet. That is the scientifically correct use of the word: a hypothesis is a stated, testable position; it becomes a theory when others test it and it holds. It graduates to "the Hodos Theory" when replication by people with no stake in it has been done and has survived — not when it feels established, and not by the author's decision.
Two readings of the name. The Hodos Hypothesis is the principle (10.5281/zenodo.21613155). Hodos Diastema is its measurement instantiation — the model-free distance between processes as curves of distributions (10.5281/zenodo.21612829). The dependency runs one way: the equation is downstream of the hypothesis, never the reverse. Cite the full phrase "the Hodos Hypothesis" for the principle and "Hodos Diastema" for the distance, and attach the DOI in both cases. The bare word "Hodos" names the programme and should not be used for either on its own: it was doing both jobs at once, and readers took it for the equation and never reached the premise underneath. Each equation in the family carries its own name — Diastema (the distance), Symploke (emergence), Systasis (constitution), Chronos (duration). The distance was named on 2026-08-09, which re-dates nothing: its paper, its DOI and every existing citation of "Hodos" for the distance remain correct as written.
What is and is not claimed. Relational ontology is older than any of this machinery, and the source paper says so in its own text. The premise has deep ancestry — dependent origination, process philosophy, structural realism, relational quantum mechanics. No one can claim the idea itself and this does not. What is claimed is this formulation as the named premise of a named programme, its instantiation in working measured machinery, and the timestamp.
A signal is a curve of probability distributions on a statistical manifold; comparison is a geodesic distance between curves.
The ground — the information-geometry representation, not the alignment machinery — is what earns a measurable advantage, and it does so wherever identity lives in the relative-space pattern of a process. The warp is a separate lane: it earns its keep only where identity lives in temporal motion, and on discrete already-aligned data it actively hurts (finding 17).
These are two lanes. Conflating them was our own error, retracted as finding 5 and corrected in finding 6.
It learns from 5–10× less data than a matched standard network, on two unrelated modalities (finding 13). At one example the network is barely above chance.
It does not forget (finding 14). Taught in sequence, a standard network loses half to four-fifths of what it first learned. This loses nothing — and cannot, because building the store in reverse order yields byte-identical answers, so the final state provably does not depend on the sequence.
It is invariant by construction rather than by training (finding 7). A transposed signal collapses 61× while a genuinely different one is preserved. A network has to learn that from examples and only ever approximates it.
The store grows. Consolidation folds repeated episodes into one prototype per concept, so it scales with how many distinct things are known rather than with how much has been seen — but lookup gets slower as it learns, where a network's does not. That is the real trade, and it is why a fast retrieval index exists.
Not a universal oracle. It solves what can be posed as a least-cost path on a distribution-manifold; it does not invent a domain's law and does not reach discrete or logical truths. The Riemann work (finding 22) reproduces known mathematics and proves nothing new.
Not better on language. Memory helps a language model, but a plain distance retrieves better than this one for text (finding 24), and the mechanism is understood (finding 17).
Not novel wholesale. Dynamic time warping, Fisher–Rao, the square-root sphere, soft-DTW, barycenter averaging and optimal transport are all prior art, reused. What is ours is the synthesis, the pitch quotient, and the measurements on this page.