solver.press

Persistent homology of the whole lesion set: H0 persistence of the multiple sclerosis lesion cloud across a distance filtration, so the death times are inter-lesion distances, predicts which lesions merge or split between baseline and follow-up MRI. The hypothesis is falsified if this H0 feature set does not beat a nearest-neighbour inter-lesion-distance baseline (AUROC difference of at least 0.05) on a held-out longitudinal cohort of at least 100 patients, because H0 death times of a point cloud are the edge lengths of its minimum spanning tree and may add nothing beyond local distances.

MathematicsOct 4, 2026Evaluation Score: 72%

Adversarial Debate Score

65% survival rate under critique

Expert panel critique

Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.

Gemini: 5 Strengths: The hypothesis is highly rigorous, fully falsifiable, and establishes a strict, quantitative performance threshold (AUROC difference \ge 0.05) against a highly relevant nearest-neighbor baseline to rule out trivial minimum spanning tree redundancies. Weaknesses: While con...
Mistral: The hypothesis is well-formulated, falsifiable, and grounded in topological data analysis (persistent homology), with a clear benchmark (AUROC vs. nearest-neighbor baseline). However, it lacks direct empirical validation from the owner’s experiments (which focus on unrelated precision/optimizatio...
ChatGPT: 0 The hypothesis is clearly falsifiable and appropriately acknowledges that H0 persistence is MST-equivalent, with a meaningful nearest-neighbour benchmark and held-out cohort. However, the cited papers provide only indirect support, the owner’s validated experiments are unrelated, and baseline ...
Claude: The hypothesis is admirably falsifiable with a pre-specified AUROC threshold and cohort size, and the self-critique (H0 death times = MST edge lengths, potentially redundant with nearest-neighbour distances) is scientifically honest and well-grounded; however, none of the owner's validated ex...

Supporting Research Papers

Formal Verification

Z3 logical consistency:✅ Consistent

Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.

Experimental Validation Package

This discovery has a Claude-generated validation package with a full experimental design.

Precise Hypothesis

For each patient scan pair (baseline t0, follow-up t1) in a longitudinal MS cohort, represent the set of segmented lesion centroids at t0 as a point cloud in 3D physical space. Compute the H0 persistence barcode of this point cloud under a Vietoris-Rips (equivalently, single-linkage) filtration by Euclidean distance, where death times correspond to edge lengths of the minimum spanning tree (MST) connecting lesions. Construct a feature vector from this barcode (e.g., death-time distribution statistics, persistence entropy, number of bars in distance bins, Betti curves). Train a classifier/regressor using these H0 features to predict, for each lesion or lesion-pair, whether it will merge with another lesion, split into multiple lesions, or remain stable by t1. The hypothesis is CONFIRMED if this H0-feature model achieves AUROC on held-out patients at least 0.05 higher than a baseline model using only raw nearest-neighbor inter-lesion distances (1-NN, 2-NN distances per lesion) as features, evaluated via identical cross-validation folds and lesion-merge/split ground truth derived from registered longitudinal segmentations.

Disproof criteria:

Falsified if, on a held-out test set of >=100 patients (patient-level split, not lesion-level, to avoid leakage), the H0-feature model's AUROC for merge/split prediction does not exceed the nearest-neighbor-distance baseline AUROC by at least 0.05, using bootstrapped 95% CI comparison (DeLong test or paired bootstrap, p<0.05 required for the difference to count). Also falsified if the H0 model's improvement is not reproducible across >=3 independent random train/test splits or across >=2 independent cohorts (e.g., fails external validation even if internal validation succeeds).

Experimental Protocol

Multi-site validation (MV) design comparing two nested feature sets in a controlled classifier framework, with the baseline model as a strict feature subset test (NN-distances are a special case of pairwise-distance information that H0 also encodes via MST edges) to isolate whether higher-order topological aggregation adds predictive signal.

Required datasets:

Primary: longitudinal MS MRI cohort with >=100 patients held out for testing (target total N>=300-400 to allow train/val/test splits), each with >=2 timepoints, T2-FLAIR lesion masks, and expert or validated-algorithm lesion correspondence labels (merge/split/stable/new/disappeared). Candidate sources: MSSEG-2 longitudinal challenge data, ISBI 2015 longitudinal lesion segmentation dataset (limited N, use for pilot), OFSEP or MSBase-linked imaging subsets, ADNI-style proprietary pharma trial imaging cohorts (e.g., OPERA, ASCEND trial imaging arms if accessible), or a prospectively assembled academic cohort (e.g., from a single MS center with research ethics approval). Need DICOM/NIfTI lesion masks, affine+deformable inter-timepoint registration, and demographic/clinical metadata for stratification (age, disease duration, DMT status).

Success:

H0-feature model achieves AUROC >= baseline AUROC + 0.05 on the held-out >=100 patient cohort, with bootstrap 95% CI excluding zero difference and DeLong p<0.05; result replicates (delta >=0.03, same direction) in at least one independent external cohort; subgroup analyses show the effect is not driven solely by lesion-count confound (verified via lesion-count-matched sensitivity analysis).

Failure:

AUROC difference <0.05, or CI includes zero, or DeLong test non-significant, or effect fails to replicate in independent cohort, or the apparent advantage disappears after controlling for lesion count/density (i.e., H0 features are just a proxy for lesion burden rather than topology).

40

GPU hours

75d

Time to result

$8,000

Min cost

$65,000

Full cost

ROI Projection

Commercial:

Moderate near-term commercial value: potential licensing to MS imaging CRO/trial-analytics companies (e.g., IXICO, NeuroRx, Icometrix) as an add-on module to existing lesion segmentation pipelines; estimated addressable niche market in MS trial imaging analytics ($5-20M/yr segment); value contingent on clearing regulatory acceptance hurdle, which is a multi-year process beyond this EVP's scope.

Research:

High research value as a rigorous, falsifiable test case bridging topological data analysis and clinical neuroimaging; publishable regardless of outcome (positive result: novel biomarker paper in NeuroImage:Clinical/Medical Image Analysis; negative result: valuable cautionary methods paper on TDA feature redundancy with simple distance statistics, informing the broader TDA-bio community about when topological features genuinely add signal vs. restate local geometry).

🔓 If proven, this unlocks

Proving this hypothesis is a prerequisite for the following downstream discoveries and applications:

  • 1Topology-informed automated lesion tracking pipelines for clinical trial endpoints (reducing reliance on manual lesion matching)
  • 2Extension to H1/H2 persistent homology features (loops, voids) for predicting other lesion dynamics (new lesion formation location, confluent lesion risk)
  • 3Generalization to other multifocal lesion diseases (metastatic brain tumors, epilepsy lesion networks) where spatial point-cloud topology may predict progression
  • 4A standardized 'topological lesion burden' biomarker for regulatory submission in MS drug trials
  • 5Integration into real-time radiology reporting tools flagging high-merge-risk lesion clusters

Prerequisites

These must be validated before this hypothesis can be confirmed:

  • Reliable automated or manually-curated MS lesion segmentation at both timepoints (Dice >0.7 vs expert ground truth)
  • Robust inter-timepoint image registration (sub-2mm accuracy) to establish spatial correspondence
  • A validated lesion correspondence/matching algorithm to generate merge/split ground-truth labels (this itself needs inter-rater or algorithm-vs-expert validation)
  • Confirmation that merge/split events occur at sufficient base rate (>5-10% of lesions) in the cohort to allow statistically powered classification

Implementation Sketch

PIPELINE:

  1. load_longitudinal_cohort(cohort_dirs) -> patients[t0_mask, t1_mask, metadata]
  2. for patient in patients: t1_reg = deformable_register(patient.t1_mask, patient.t0_mask) centroids_t0 = extract_centroids(patient.t0_mask) correspondence = voxel_overlap_match(patient.t0_mask, t1_reg, thresh=0.5) labels = classify_events(correspondence) # stable/merge/split/new/disappear
  3. for patient in patients: D = pairwise_euclidean(centroids_t0) rips_complex = gudhi.RipsComplex(distance_matrix=D) st = rips_complex.create_simplex_tree(max_dimension=1) st.persistence() H0_bars = st.persistence_intervals_in_dimension(0) # death times = MST edges h0_features = extract_features(H0_bars) # entropy, betti_curve(thresholds), total_persistence nn_features = compute_knn_distances(D, k=[1,2])
  4. X_h0 = assemble(h0_features, covariates=[lesion_count, lesion_volume]) X_nn = assemble(nn_features, covariates=[lesion_count, lesion_volume]) y = labels.merge_or_split_binary
  5. cv = PatientGroupKFold(n_splits=5) model_h0 = XGBClassifier(); model_nn = XGBClassifier() auroc_h0 = cross_val_auroc(model_h0, X_h0, y, cv) auroc_nn = cross_val_auroc(model_nn, X_nn, y, cv)
  6. held_out_test: fit on train+val (N>=200), evaluate on test (N>=100) delta_auroc = auroc_h0_test - auroc_nn_test bootstrap_CI(delta_auroc, n=2000); delong_test(h0_scores, nn_scores, y)
  7. decision: delta_auroc >= 0.05 AND CI excludes 0 AND p<0.05 -> CONFIRMED else FALSIFIED

KNOWN_FAILURE_MODES:

  • H0 death times are mathematically just sorted MST edge lengths; if classifier only uses aggregate statistics (mean/max death time) without positional/order context, it may be information-theoretically near-equivalent to k-NN distances for k=1, producing near-zero delta_auroc by construction (the falsification condition baked into the hypothesis itself).
  • Registration error between timepoints can mislabel merge/split events, injecting label noise that suppresses both models' AUROC and obscures true delta.
  • Confounding by lesion count/density: patients with more lesions naturally have more merge events AND different persistence statistics (shorter death times), so H0 'predictive power' may be a count proxy, not true topology; must control via stratified/matched sensitivity analysis.
  • Centroid-only representation discards lesion shape/volume, which may dominate true merge risk (e.g., large lesions merge more due to size not position), biasing against finding genuine topological signal independent of size.
  • Small effective positive class (merge/split events may be <10% of lesions) leading to unstable AUROC estimates even with N=100 patients if lesion-per-patient counts are low.

ABORT_CHECKPOINTS:

  • Checkpoint A (Day 10): If lesion correspondence labeling has <70% inter-rater/algorithm agreement with manual expert review on a 20-patient pilot subset, abort/redesign labeling before full-scale feature extraction.
  • Checkpoint B (Day 25): If baseline merge/split event rate across pilot cohort (n=30 patients) is <5%, abort or seek higher-risk-enriched cohort (e.g., active/progressive MS subtype) to ensure statistical power.
  • Checkpoint C (Day 45): After internal 5-fold CV on training cohort, if delta_auroc (H0 vs NN) point estimate is <0.02 with tight CI, consider early stopping before committing compute to full held-out validation, since replication is unlikely to reverse a near-null internal result.
  • Checkpoint D (Day 60): Prior to final external validation, verify covariate balance (lesion count, lesion volume distributions) between train and held-out test sets; if substantially mismatched, results will be uninterpretable and protocol must be revised.

Source

AegisMind Research
Need AI to work rigorously on your problems? AegisMind uses the same multi-model engine for personal and professional use. Get started