N NEUROHELIX
CWSF 2027 · COMPUTATIONAL NEUROSCIENCE

Shared biology.
Changing traits.
What can be detected?

NEUROHELIX measures when the effect of biological pathways on neurodevelopmental trait trajectories could be detected, and when it could not.

RESEARCH IN PROGRESS · Simulation + human-data testing
WHAT THE PROJECT FOUNDVersion 3
  • 7.3%
    Lower prediction error, in simulation.

    Planted pathway values improved on trait history alone. The size of this number is set by the simulator, not measured in people.

  • 4.5→7.1%
    False detections rise with noisy measurements.

    With no interaction planted, the detection rule still fired. A rate of 2.5% is expected by chance.

  • 0/3
    Primary human tests significant.

    Gene panels in blood did not differ between cases and controls in autism (two cohorts) or OCD.

NEW IN VERSION 3Detection limits, a measurement-noise control, and the OCD and ADHD tests are now reported.See what the method can and cannot detect ↗
15 curated genes4 biological systems300 simulated individuals344 blood samples · three cohorts
01 / THE FRAMEWORK

From connections
to predictions.

NEUROHELIX brings biological curation, a small knowledge graph, dynamical modeling, and machine learning into one interpretable research workflow.

01

Map the biology

Connect curated genes to pathways, four core systems, and three conditions.

Knowledge graph
02

Model the dynamics

Express trait persistence, pathway influences, and environmental variation mathematically.

HELIX-N / HELIX-M
03

Generate trajectories

Create virtual individuals with distinct pathway profiles and 36 observations each.

Controlled simulation
04

Test added value

Compare traits-only, pathway-enhanced, and shuffled-pathway predictors, then measure when the test fails.

HELIX-P
The question is incremental value.

Current traits already predict the next observation well. The experiment asks whether explicit pathway information adds useful signal beyond that strong baseline, and how large an effect and how much data that takes.

02 / BIOLOGICAL FOUNDATION

One network.
Four shared systems.

The directed knowledge graph organizes selected biological relationships. Explore each system to see its curated genes and condition links.

CURATED SYSTEM–CONDITION LINKS36 NODES · 39 EDGES

Simplified view of the system and condition layers. Links represent the project’s curated graph; missing links do not establish biological absence.

Inspect the full gene → pathway mapping +
GeneCurated pathwayCore system
The graph names the four pathway dimensions of the simulator and defines the gene panels tested in human data. It has no numerical role in the simulator and is not itself evidence of causal importance. ADGRL3 and DLGAP3 are the current names of the genes earlier called LPHN3 and SAPAP3.
03 / THE MATHEMATICAL ENGINE

Interpretable by design.

The controlled generator models three continuous synthetic traits: attention difficulty, compulsivity severity, and social communication difficulty.

OFFICIAL LINEAR GENERATOR v4.3
Ti(t+1) = c + A Ti(t) + Bi Pi + Ei(t)

c + ATBaseline and current-trait effects

BᵢPᵢIndividual pathway influence

Eᵢ(t)Environmental / process noise

The nonlinear experiments add one term: three trait × pathway products (attention × monoamine, compulsivity × glutamatergic, social communication × synaptic), scaled together by an interaction strength s. At s = 0 the generator is exactly the linear one above.

A controlled synthetic dataset

300virtual individuals
36observations each
35transitions each
7 → 3predictors → targets

Five computational profiles: synaptic, glutamate, monoamine, proteostasis, and balanced. These profiles are modeling categories, not diagnoses.

Every parameter was chosen for this study. None was estimated from human data. Timesteps are ordered observations, not calibrated calendar ages.

Explicit assumptions. Visible checks.

  • Stable transition matrixSpectral radius 0.92, below 1.
  • Interior neutral equilibrium3.5 for each trait; baseline intercept 0.28.
  • No clipping in the primary run0.0000% at the 0–10 trait boundaries.
  • Fixed pathway profilesConstant within each trajectory; not measured biology.
MODEL EXPLORER

See how pathway profiles
change the trajectory.

Choose a profile to view the deterministic mean trajectory under the official base matrices.

Illustration computed for this website: starts at [3.5, 3.5, 3.5], uses a 0.45 pathway shift, and omits noise and individual matrix variation. It is not a held-out prediction or a clinical forecast.

Attention difficultyCompulsivity severitySocial communication difficulty
04 / SYNTHETIC PREDICTION RESULTS

Pathway features add
signal in simulation.

In the linear experiment, adding the planted pathway values improves next-step prediction. Shuffling pathway profiles between individuals removes the benefit.

MAE REDUCTION7.3%

95% interval: 6.0–8.6%

RMSE REDUCTION7.3%

95% interval: 6.1–8.4%

REPEATED HOLDOUTS50/50

HELIX-P ahead on MAE, RMSE, and R²

Three models. The same held-out individuals.

240 training / 60 testing individuals · 8,400 / 2,100 rows

ModelInputsMAE ↓RMSE ↓R² ↑
Traits-only baseline30.07750.09700.9693
HELIX-P · traits + pathways70.07190.09000.9736
Decoy · traits + shuffled pathways70.07760.09720.9692

MAE = mean absolute error. RMSE = root mean squared error. R² is averaged over traits with variance weights; it is not a percentage accuracy score.

What makes the test stronger

  • Split by individualEntire trajectories stay together to avoid row-level train/test leakage.
  • Subject-cluster uncertainty2,000 bootstrap resamples of held-out individuals; 95% intervals.
  • A matched-input decoySeven inputs, but shuffled biology: MAE 0.0776, no better than baseline.
  • Repeated checksFive-fold grouped cross-validation and 50 grouped holdouts test robustness.

Two things to keep in mind

  • The size of the gain is set by the simulatorNoise puts a floor of about 0.071 on MAE. How far the traits-only model sits above that floor depends on how strong the pathway effects were made.
  • The gain is concentrated early28.5% at the first step, 2.4% at the last. Pathway effects accumulate in the traits, so later traits already carry most of what the pathway values add.
  • One control is unfinishedThe run with the pathway effect removed has been generated, but its prediction result has not been reported yet.

Repeated holdout splits reuse the same 300 individuals and are descriptive robustness evidence, not 50 independent replications.

WHAT THIS RESULT ESTABLISHES

The pipeline detects pathway information that was deliberately built into a simulation.

It does not establish predictive validity in people, causal biological mechanisms, or clinical usefulness. The generator contains pathway effects by design, and its parameters are not fitted to patient data.

Earlier experiments & model versions +
ExperimentBaseline MAEHELIX-P MAEReported reduction
Early HELIX-P linear0.16200.15981.36%
Linear comparison v4.10.16070.15950.79%
Official v5.0 / generator v4.30.07750.07197.30%

Generator settings and analysis versions differ. These values are separate experiments, not a like-for-like performance trend.

05 / DETECTION LIMITS · NEW

How strong. How many.
How clean.

A second experiment plants trait × pathway interactions and asks when they can be detected beyond a model that already contains the pathways. The answer depends on effect strength, sample size, and how cleanly the pathway values are measured.

300 INDIVIDUALS · s = 1.28/10

datasets with a detection

900 INDIVIDUALS · s = 0.710/10

datasets with a detection

NOISY VALUES · RELIABILITY 0.43/10

down from 10/10, at 300 individuals and s = 1.2

More individuals and stronger effects are detected more often.

Datasets with a detection, out of 10 independent simulated datasets per cell

350 datasets
Detection count by sample size (rows) and interaction strength s (columns)

A dataset counts as a detection when the 95% interval for the corrected gain lies above zero. The corrected gain compares the true interaction terms with decoy terms built from another individual’s pathway profile. With ten datasets per cell each count is imprecise (8 of 10 has a 95% interval of 49–94%), so the table shows trends, not thresholds. s is a simulator setting with no clinical unit.

THE MAIN CAUTION

With nothing planted, the detection rule still fired, and more often when the pathway values were noisy.

Across 1,000 simulated datasets with no interaction, the rule fired more often than the 2.5% expected by chance. A detected interaction is therefore not proof of a nonlinear mechanism.

Common axis: 0 to 10% of datasets. Reliability is the share of the observed variance that is true signal.

WHY IT HAPPENS

Two causes were identified.

  • Pathway values that come in clustersWhen true values come from a mixture of profiles, the best estimate of a true value from a noisy one is a curve. Products of traits and noisy pathway values let a linear model approximate that curve.
  • One model fitted across all time stepsSimulated traits drift toward a level set by the pathways. Fitting separate models per time step removed most of the effect; what remained was inconclusive.

Both are instances of a known problem: controlling for a variable that is measured with error. The follow-up controls were designed after the pattern was seen, and fixed before they were run.

WHAT ELSE NOISE DOES

The benefit of pathways shrinks fast.

Linear gain kept at reliability 0.862%
Linear gain kept at reliability 0.422%

The loss is steeper than the reliability itself, because traits already carry part of the pathway information and noise removes a larger share of the rest. Measurement-error theory, applied before the experiment, predicted 63% and 24%.

WHAT THIS RESULT ESTABLISHES

An interaction model beating an additive model is not sufficient evidence that the underlying dynamics are nonlinear.

The interaction model was also told which three products to use. These experiments concern detecting an interaction of known form; discovering its form was not tested. The artifact is small, roughly 6% of the gain from a planted interaction at s = 1.0, and matters most where planted effects are weakest.

06 / HUMAN-DATA TESTING

Real samples.
An honest result.

A retrospective analysis of public blood gene-expression data asks whether the project’s gene panels differ between cases and controls. It is a cross-sectional test and does not address prediction over time.

NO PRIMARY TEST SIGNIFICANT

The gene panels did not show a statistically significant case–control difference in autism (two cohorts) or OCD.

Every interval includes zero, and the two autism cohorts differ in the direction of the estimate. These tests could detect only moderate to large differences, so they make large effects unlikely and say little about small ones.

AUTISM · COHORT 1GPL570
99 blood samples

66 ASD · 33 controls · all male
Synaptic panel, 6 genes

Permutation p0.27
AUTISM · COHORT 2GPL6244
185 blood samples

103 ASD · 82 controls · mixed sex
Synaptic panel, 6 genes

Permutation p0.61
OCDGPL19612
60 blood samples

30 OCD · 30 controls
Glutamatergic panel, 3 genes

Permutation p0.56

The autism cohorts come from GSE18123 and are analyzed separately; one sample appearing on both platforms was removed from the second cohort. The OCD cohort is GSE78104. These are existing public data, not people recruited by NEUROHELIX.

No interval excludes zero.

Cases minus controls · standardized difference in the panel score (Hedges’ g) · 95% intervals

Primary panels
Primary gene-panel effects with uncertainty and detectable effect sizesAutism cohort 1: 0.24, interval minus 0.21 to 0.73, detectable at 0.61. Autism cohort 2: minus 0.07, interval minus 0.37 to 0.22, detectable at 0.46. OCD: 0.15, interval minus 0.34 to 0.63, detectable at 0.75. All intervals include zero.
Primary results
CohortMean difference95% bootstrap intervalHedges’ gPermutation pRandom-panel p
Autism · cohort 1+0.112−0.099 to +0.343+0.240.270.44
Autism · cohort 2−0.038−0.193 to +0.113−0.070.610.78
OCD+0.092−0.210 to +0.390+0.150.560.52

10,000 label permutations and 10,000 stratified bootstrap resamples per cohort. The random-panel p compares each panel with 10,000 random gene panels of the same size. P-values are uncorrected; with the smallest at 0.27, no multiple-testing correction can make any test significant. In the chart, intervals are the bootstrap intervals rescaled to standardized units.

WHAT COULD THESE TESTS DETECT?

Only moderate to large differences.

A known shift was added to the panel genes of the cases in each real dataset, and the test rerun many times.

False-positive rate with no shift (nominal 5%)4.7–5.5%
Detectable with 80% power · autism cohort 10.61
Detectable with 80% power · autism cohort 20.46
Detectable with 80% power · OCD0.75

Detectable values are standardized differences in the combined panel score, shown as triangles in the chart. The audit assumes every panel gene shifts equally; an effect in one or two genes would be harder to detect.

WERE THE GENES MEASURABLE IN BLOOD?

Several barely are.

ADGRL3 · both autism platformsbelow 10th percentile
SLC1A1 · OCD, one of three panel genesbelow 1st percentile
ADHD panel · genes detected in most samples1 of 5

Percentile is the rank of a gene’s median expression among all genes measured on the platform. A microarray returns a number for every probe, so a value is not evidence that a gene is expressed. This check was made after the tests were complete.

EXPLORATORY

ADHD: not a confirmatory test.

The ADHD dataset (GSE159104, whole-blood RNA-seq, 23 cases and 21 controls) is reported as exploratory. Four of the five monoamine-panel genes were detected in at most 9 of the 44 samples, so the panel score is not a meaningful measure in blood. Under a stricter measurability rule the panel cannot be evaluated at all, and an exploratory result had been seen before the reporting rule was settled.

EXPLORATORY

A longitudinal pilot, without pathways.

In a public dataset with two time points (OpenNeuro ds002424, 48 participants), adding eight cognitive scores to a model of later ADHD symptoms did not help: mean absolute error was 5.64 without them and 5.81 with them. The analysis was corrected after its first outcome was seen. It tests the evaluation procedure on real data and does not involve genes or pathways.

WHAT THE ANALYSIS DOES

The same panels, a new test.

  1. Keep the biology fixed.Gene panels come from the project’s knowledge graph without change. SHANK3 belongs to two systems. No gene was substituted and no weights were fitted.
  2. Build blood-expression scores.Apply log₂(x + 1), collapse probes by their median, standardize each gene, then average genes within each panel. Preprocessing uses no diagnosis labels.
  3. Compare predefined groups.One primary panel per condition: synaptic for autism, glutamatergic for OCD, monoamine for ADHD. Panels were not retuned after any result.
SENSITIVITY ANALYSES

Planned variations changed nothing.

Autism 1 · adjusted for agep = 0.36
Autism 2 · adjusted for sex and agep = 0.61
Autism 2 · males onlyp = 0.38
OCD · adjusted for age and sexp = 0.57
OCD · matched pairsp = 0.56

Restricting to autistic disorder only gave p = 0.26 in cohort 1 and p = 0.97 in cohort 2.

Interpretation, limitations & what remains open +

Blood is a proxy. These are gene-expression panel scores in peripheral blood, not direct measurements of brain pathway function, inherited variants, or changing behavioral traits.

A null result has a scope. The tests do not support these particular scores in these cohorts. They do not establish that the underlying pathways have no role in these conditions.

Documented, not preregistered. The analysis plans are stored as hashed files. A hash shows a file has not changed since it was hashed; it cannot show that every decision was made before the data were seen, and the plans were not deposited with an independent registry.

Technical variation remains relevant. The source autism study used batch correction that cannot be reproduced from the deposited data. This analysis uses the deposited intensities and is not an exact reproduction of that study.

Limited power. The tests reach 80% power only for standardized differences of 0.46 to 0.75. Smaller effects would usually have been missed.

Prediction is a separate goal. The 7.3% improvement is a simulation result. Longitudinal human traits and measured biological features are still needed to evaluate forecasting in people.

07 / CURRENT PROGRESS

A working method.
Known limits.

The project connects a controlled prediction benchmark, a measurement of that benchmark’s detection limits, and a retrospective human-data test. It does not show that pathways predict trait trajectories in people.

IMPLEMENTED & DOCUMENTED

Biology → graph → simulation

15 genes, 14 pathways, four core systems; a mathematical model; synthetic trajectory generation.

RESULTS REPORTED

Controlled linear prediction

Baseline and HELIX-P comparison, shuffled-pathway control, grouped validation, and cluster-bootstrap uncertainty.

RESULTS REPORTED

Interaction detection and sample size

Sixteen interaction strengths in one dataset, and a sweep over five sample sizes and seven strengths with ten datasets per condition.

RESULTS REPORTED

Measurement-noise controls

Noise reduces both the benefit of pathways and the detection of interactions, and raises the false-detection rate. Two follow-up controls traced part of the cause.

HUMAN-DATA ANALYSIS COMPLETE

Autism and OCD panel tests

Three cohorts, none significant; a sensitivity audit and a gene-measurability check. The ADHD analysis is exploratory.

PARTIALLY VERIFIED

Zero-biology control

Matched primary/control dataset generation is confirmed. The final predictive comparison with the pathway effect removed remains unreported.

NEXT RESEARCH PRIORITIES

Close the open control.
Then test measured trajectories.

  1. Report the zero-biology comparisonRerun the control so the prediction result with no pathway effect is recorded alongside the main result.
  2. Add runs to the sample-size sweepTen datasets per condition show trends. More are needed before any threshold can be stated.
  3. Test one longitudinal cohort with measured biologyThree or more time points, biology in a form with established links to these traits, and a published baseline to compare against.
  4. Strengthen biological provenanceAdd a citation for every gene assignment; distinguish curated associations from inferred mechanisms.
Limitations & open methodological questions +

Two different evidence types. Prediction gains come from simulation. The human study measures cross-sectional blood-expression scores; it does not validate trait forecasts or clinical use.

The model matches the simulator. Linear models were fitted to data from a linear rule, and the interaction model was given the correct products. Performance when those assumptions fail was not tested.

One-step target. Reported metrics evaluate the next observation; they do not measure forecasting several steps ahead.

A simplified biology. Four pathway values that never change, three traits, and a small unweighted graph of fifteen genes chosen for this project.