Map the biology
Connect curated genes to pathways, four core systems, and three conditions.
Knowledge graphNEUROHELIX measures when the effect of biological pathways on neurodevelopmental trait trajectories could be detected, and when it could not.
Planted pathway values improved on trait history alone. The size of this number is set by the simulator, not measured in people.
With no interaction planted, the detection rule still fired. A rate of 2.5% is expected by chance.
Gene panels in blood did not differ between cases and controls in autism (two cohorts) or OCD.
NEUROHELIX brings biological curation, a small knowledge graph, dynamical modeling, and machine learning into one interpretable research workflow.
Connect curated genes to pathways, four core systems, and three conditions.
Knowledge graphExpress trait persistence, pathway influences, and environmental variation mathematically.
HELIX-N / HELIX-MCreate virtual individuals with distinct pathway profiles and 36 observations each.
Controlled simulationCompare traits-only, pathway-enhanced, and shuffled-pathway predictors, then measure when the test fails.
HELIX-PCurrent traits already predict the next observation well. The experiment asks whether explicit pathway information adds useful signal beyond that strong baseline, and how large an effect and how much data that takes.
The directed knowledge graph organizes selected biological relationships. Explore each system to see its curated genes and condition links.
Simplified view of the system and condition layers. Links represent the project’s curated graph; missing links do not establish biological absence.
| Gene | Curated pathway | Core system |
|---|
The controlled generator models three continuous synthetic traits: attention difficulty, compulsivity severity, and social communication difficulty.
c + ATBaseline and current-trait effects
BᵢPᵢIndividual pathway influence
Eᵢ(t)Environmental / process noise
The nonlinear experiments add one term: three trait × pathway products (attention × monoamine, compulsivity × glutamatergic, social communication × synaptic), scaled together by an interaction strength s. At s = 0 the generator is exactly the linear one above.
Five computational profiles: synaptic, glutamate, monoamine, proteostasis, and balanced. These profiles are modeling categories, not diagnoses.
Every parameter was chosen for this study. None was estimated from human data. Timesteps are ordered observations, not calibrated calendar ages.
Choose a profile to view the deterministic mean trajectory under the official base matrices.
Illustration computed for this website: starts at [3.5, 3.5, 3.5], uses a 0.45 pathway shift, and omits noise and individual matrix variation. It is not a held-out prediction or a clinical forecast.
In the linear experiment, adding the planted pathway values improves next-step prediction. Shuffling pathway profiles between individuals removes the benefit.
95% interval: 6.0–8.6%
95% interval: 6.1–8.4%
HELIX-P ahead on MAE, RMSE, and R²
240 training / 60 testing individuals · 8,400 / 2,100 rows
| Model | Inputs | MAE ↓ | RMSE ↓ | R² ↑ |
|---|---|---|---|---|
| Traits-only baseline | 3 | 0.0775 | 0.0970 | 0.9693 |
| HELIX-P · traits + pathways | 7 | 0.0719 | 0.0900 | 0.9736 |
| Decoy · traits + shuffled pathways | 7 | 0.0776 | 0.0972 | 0.9692 |
MAE = mean absolute error. RMSE = root mean squared error. R² is averaged over traits with variance weights; it is not a percentage accuracy score.
Repeated holdout splits reuse the same 300 individuals and are descriptive robustness evidence, not 50 independent replications.
It does not establish predictive validity in people, causal biological mechanisms, or clinical usefulness. The generator contains pathway effects by design, and its parameters are not fitted to patient data.
| Experiment | Baseline MAE | HELIX-P MAE | Reported reduction |
|---|---|---|---|
| Early HELIX-P linear | 0.1620 | 0.1598 | 1.36% |
| Linear comparison v4.1 | 0.1607 | 0.1595 | 0.79% |
| Official v5.0 / generator v4.3 | 0.0775 | 0.0719 | 7.30% |
Generator settings and analysis versions differ. These values are separate experiments, not a like-for-like performance trend.
A second experiment plants trait × pathway interactions and asks when they can be detected beyond a model that already contains the pathways. The answer depends on effect strength, sample size, and how cleanly the pathway values are measured.
datasets with a detection
datasets with a detection
down from 10/10, at 300 individuals and s = 1.2
Datasets with a detection, out of 10 independent simulated datasets per cell
A dataset counts as a detection when the 95% interval for the corrected gain lies above zero. The corrected gain compares the true interaction terms with decoy terms built from another individual’s pathway profile. With ten datasets per cell each count is imprecise (8 of 10 has a 95% interval of 49–94%), so the table shows trends, not thresholds. s is a simulator setting with no clinical unit.
Across 1,000 simulated datasets with no interaction, the rule fired more often than the 2.5% expected by chance. A detected interaction is therefore not proof of a nonlinear mechanism.
Common axis: 0 to 10% of datasets. Reliability is the share of the observed variance that is true signal.
Both are instances of a known problem: controlling for a variable that is measured with error. The follow-up controls were designed after the pattern was seen, and fixed before they were run.
The loss is steeper than the reliability itself, because traits already carry part of the pathway information and noise removes a larger share of the rest. Measurement-error theory, applied before the experiment, predicted 63% and 24%.
The interaction model was also told which three products to use. These experiments concern detecting an interaction of known form; discovering its form was not tested. The artifact is small, roughly 6% of the gain from a planted interaction at s = 1.0, and matters most where planted effects are weakest.
A retrospective analysis of public blood gene-expression data asks whether the project’s gene panels differ between cases and controls. It is a cross-sectional test and does not address prediction over time.
Every interval includes zero, and the two autism cohorts differ in the direction of the estimate. These tests could detect only moderate to large differences, so they make large effects unlikely and say little about small ones.
66 ASD · 33 controls · all male
Synaptic panel, 6 genes
103 ASD · 82 controls · mixed sex
Synaptic panel, 6 genes
30 OCD · 30 controls
Glutamatergic panel, 3 genes
The autism cohorts come from GSE18123 and are analyzed separately; one sample appearing on both platforms was removed from the second cohort. The OCD cohort is GSE78104. These are existing public data, not people recruited by NEUROHELIX.
Cases minus controls · standardized difference in the panel score (Hedges’ g) · 95% intervals
| Cohort | Mean difference | 95% bootstrap interval | Hedges’ g | Permutation p | Random-panel p |
|---|---|---|---|---|---|
| Autism · cohort 1 | +0.112 | −0.099 to +0.343 | +0.24 | 0.27 | 0.44 |
| Autism · cohort 2 | −0.038 | −0.193 to +0.113 | −0.07 | 0.61 | 0.78 |
| OCD | +0.092 | −0.210 to +0.390 | +0.15 | 0.56 | 0.52 |
10,000 label permutations and 10,000 stratified bootstrap resamples per cohort. The random-panel p compares each panel with 10,000 random gene panels of the same size. P-values are uncorrected; with the smallest at 0.27, no multiple-testing correction can make any test significant. In the chart, intervals are the bootstrap intervals rescaled to standardized units.
A known shift was added to the panel genes of the cases in each real dataset, and the test rerun many times.
Detectable values are standardized differences in the combined panel score, shown as triangles in the chart. The audit assumes every panel gene shifts equally; an effect in one or two genes would be harder to detect.
Percentile is the rank of a gene’s median expression among all genes measured on the platform. A microarray returns a number for every probe, so a value is not evidence that a gene is expressed. This check was made after the tests were complete.
The ADHD dataset (GSE159104, whole-blood RNA-seq, 23 cases and 21 controls) is reported as exploratory. Four of the five monoamine-panel genes were detected in at most 9 of the 44 samples, so the panel score is not a meaningful measure in blood. Under a stricter measurability rule the panel cannot be evaluated at all, and an exploratory result had been seen before the reporting rule was settled.
In a public dataset with two time points (OpenNeuro ds002424, 48 participants), adding eight cognitive scores to a model of later ADHD symptoms did not help: mean absolute error was 5.64 without them and 5.81 with them. The analysis was corrected after its first outcome was seen. It tests the evaluation procedure on real data and does not involve genes or pathways.
Restricting to autistic disorder only gave p = 0.26 in cohort 1 and p = 0.97 in cohort 2.
Blood is a proxy. These are gene-expression panel scores in peripheral blood, not direct measurements of brain pathway function, inherited variants, or changing behavioral traits.
A null result has a scope. The tests do not support these particular scores in these cohorts. They do not establish that the underlying pathways have no role in these conditions.
Documented, not preregistered. The analysis plans are stored as hashed files. A hash shows a file has not changed since it was hashed; it cannot show that every decision was made before the data were seen, and the plans were not deposited with an independent registry.
Technical variation remains relevant. The source autism study used batch correction that cannot be reproduced from the deposited data. This analysis uses the deposited intensities and is not an exact reproduction of that study.
Limited power. The tests reach 80% power only for standardized differences of 0.46 to 0.75. Smaller effects would usually have been missed.
Prediction is a separate goal. The 7.3% improvement is a simulation result. Longitudinal human traits and measured biological features are still needed to evaluate forecasting in people.
The project connects a controlled prediction benchmark, a measurement of that benchmark’s detection limits, and a retrospective human-data test. It does not show that pathways predict trait trajectories in people.
15 genes, 14 pathways, four core systems; a mathematical model; synthetic trajectory generation.
Baseline and HELIX-P comparison, shuffled-pathway control, grouped validation, and cluster-bootstrap uncertainty.
Sixteen interaction strengths in one dataset, and a sweep over five sample sizes and seven strengths with ten datasets per condition.
Noise reduces both the benefit of pathways and the detection of interactions, and raises the false-detection rate. Two follow-up controls traced part of the cause.
Three cohorts, none significant; a sensitivity audit and a gene-measurability check. The ADHD analysis is exploratory.
Matched primary/control dataset generation is confirmed. The final predictive comparison with the pathway effect removed remains unreported.
Two different evidence types. Prediction gains come from simulation. The human study measures cross-sectional blood-expression scores; it does not validate trait forecasts or clinical use.
The model matches the simulator. Linear models were fitted to data from a linear rule, and the interaction model was given the correct products. Performance when those assumptions fail was not tested.
One-step target. Reported metrics evaluate the next observation; they do not measure forecasting several steps ahead.
A simplified biology. Four pathway values that never change, three traits, and a small unweighted graph of fifteen genes chosen for this project.