solver.press

The spectral analysis methods from complex interpolation of matrices can reveal conserved eigenvector structures between healthy and Multiple Sclerosis-affected tissues in cross-tissue transcriptomic data, enabling the identification of disease-specific manifold deviations.

BiologyApr 18, 2026Evaluation Score: 62%

Adversarial Debate Score

57% survival rate under critique

Expert panel critique

Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.

ChatGPT: The hypothesis is falsifiable and draws on relevant mathematical methods for matrix comparison, as discussed in the cited papers, but lacks clear direct evidence of successful application to biological transcriptomic data or MS specifically; potential confounders in biological data complexity and...
Grok: The hypothesis is falsifiable through testing spectral analysis on transcriptomic data for eigenvector conservation, and it is partially supported by the papers on complex interpolation of matrices and MS transcriptomics. However, it lacks direct evidence linking spectral methods to disease-speci...
Mistral: The hypothesis is falsifiable and aligns with the cited work on spectral interpolation and transcriptomic analysis, but lacks direct empirical validation in the provided excerpts and faces potential counterarguments about biological noise and manifold instability in MS data.
Claude: The hypothesis creatively bridges two relevant papers (complex matrix interpolation and MS transcriptomics ML), but the connection is speculative and unsupported—the MS paper uses standard ML pipelines with no mention of complex interpolation or eigenvector-based manifold analysis, and the matrix...

Supporting Research Papers

Literature Assessment

📖 Literature-assessed (LLM)· literature_meta

An LLM's reading of the literature — not computational verification.

Eigenvector structures may vary significantly in disease contexts.

Method: literature_meta · Result: inconclusive

Formal Verification

Z3 logical consistency:✅ Consistent

Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.

Experimental Validation Package

This discovery has a Claude-generated validation package with a full experimental design.

Precise Hypothesis

Restricting to cross-tissue bulk and pseudo-bulk transcriptomic datasets of MS versus matched healthy controls (CNS lesion tissue, PBMC, and CD8+ T-cell-sorted fractions from GSE193770, GSE108000, GSE138614, CELLxGENE Census), a spectral decomposition method based on complex interpolation of gene-gene correlation/covariance matrices (e.g., eigenvector continuation between healthy and disease matrices via matrix pencils/Löwner interpolation) will identify a low-rank subspace (k ≤ 20 eigenvectors) that is (a) statistically conserved across ≥3 independent healthy cohorts (subspace angle < 15° pairwise) and (b) exhibits a measurable, reproducible deviation in MS samples (subspace angle > 30° from healthy consensus, permutation p<0.01, Benjamini-Hochberg corrected across tissue types) that is NOT explainable by cell-type composition shifts alone (persists after CIBERSORTx/scVI deconvolution correction). Falsification: if no such conserved-yet-deviating subspace exists, or if the deviation vanishes after cell-composition correction, the hypothesis is false.

Disproof criteria:
  1. Subspace angle between healthy-cohort eigenvector sets exceeds 30° (i.e., "conservation" claim itself fails) in ≥2 of 3 independent healthy datasets.
  2. MS-vs-healthy subspace deviation is statistically indistinguishable from healthy-vs-healthy cohort variation (permutation test p>0.05 after FDR correction).
  3. Deviation signal disappears or drops below noise floor after regressing out cell-type proportion estimates (scVI/CIBERSORTx), indicating the "manifold deviation" is merely a compositional artifact.
  4. Deviation subspace does not overlap (principal angle <45°) with the independently derived DEG-based CA-RIM signature (394 genes) or the composite-score target hierarchy (DNMT1, ZNF740, CTSS) — i.e., the spectral method finds nothing convergent with the known biology.
  5. Method fails to outperform a naive PCA-on-pooled-data baseline (no interpolation) in cross-validated MS/healthy classification (AUC difference <0.03).

Spine & Adversarial Read

  • highWhy matrix-pencil/Löwner complex interpolation specifically, rather than simpler and more standard manifold methods (diffusion maps, UMAP, geodesic interpolation on the SPD manifold, or even straightforward joint PCA)? The methodology choice is not justified against these well-established alternatives.
    Not resolved in current design — the EVP does not present a comparative rationale for why complex interpolation of matrices (a technique with roots in control theory/rational approximation) is expected to outperform SPD-manifold geodesic methods or diffusion-map approaches already used in transcriptomic manifold learning. This must be addressed by adding an explicit baseline comparison arm (geodesic-on-SPD interpolation, diffusion maps) alongside the naive PCA baseline before claiming the specific method choice is warranted; absent this, a reviewer will reasonably ask 'why this math, not the standard math.'
  • highThe claimed 'MS-specific manifold deviation' may be entirely reducible to the already-known cell-type composition shift (CD8+ T-cell expansion documented in the Phase 2 scVI atlas) — the spectral method could just be an elaborate, expensive re-derivation of a fact already established by simpler deconvolution, offering no orthogonal information.
    Partially addressed via the mandatory cell-composition regression step (Methodology step 9, Disproof criterion 3) — but the correction method (regressing out scVI-derived fractions) itself has known limitations (imperfect deconvolution accuracy, especially for rare subpopulations) and could leave residual composition signal that gets misattributed to 'manifold deviation.' Recommend adding a synthetic-data sanity check (simulate pure composition-shift data with no true manifold change) to confirm the method correctly returns a null result before trusting real-data findings.
  • mediumSample sizes in the source cohorts (GSE193770, GSE138614) are likely far below the n≥30/group threshold stated as a boundary condition once split by phenotype (relapsing vs. smoldering) and healthy sub-cohorts — the eigenvector estimates may be statistically unstable regardless of the mathematical elegance of the interpolation method.
    Acknowledged in Known Failure Modes and mitigated by the Day 7/Day 15 abort checkpoints, but not resolved — actual sample sizes for GSE193770 and GSE138614 by phenotype subgroup are not confirmed in this package and should be verified before committing full budget; if subgroup n<20, the protocol should default to pseudobulk aggregation across broader Leiden clusters rather than per-phenotype splitting to preserve statistical power.

Experimental Protocol

Minimum viable test (MVT):

  • Use GSE138614 (already replicated CTSS/FGF2/SLCO2B1 signal) as primary dataset (n≈50-90 depending on group sizes), GSE193770 CD8+ T-cell subset as secondary/validation, and a third independent healthy reference (GTEx CNS/blood, v10) as the "conservation" test set.
  • Compute gene-gene correlation matrices per group (healthy_1, healthy_2, healthy_3, MS) restricted to top 2,000-5,000 variable genes intersected across datasets.
  • Apply complex interpolation (matrix pencil / Löwner framework) between healthy matrices to define a "healthy manifold trajectory"; extract top-k eigenvectors (k selected by Marchenko-Pastur threshold, expect k=10-20).
  • Interpolate/extrapolate to MS matrix; measure principal (subspace) angles between the MS eigenbasis and the healthy consensus basis via canonical correlation analysis (CCA).
  • Run 1,000-permutation null (label shuffling) to establish significance of angle deviation.
  • Cross-validate: does deviation subspace loadings enrich for CA-RIM DEGs (394 genes) and target hierarchy genes (DNMT1, ZNF740, CTSS, FGF2, SLCO2B1) via GSEA on eigenvector loadings?
Required datasets:
  • GSE193770 (CD8+ T cells, CNS-associated, primary discovery cohort)
  • GSE108000 (MS bulk transcriptomics)
  • GSE138614 (independent replication cohort, already used for CTSS/FGF2/SLCO2B1)
  • CELLxGENE Census (cross-tissue healthy reference atlas)
  • GTEx v10 (independent healthy baseline, blood/CNS tissue)
  • Existing scVI atlas (gs://aegismind-tpu-results/ms_phase2/results/, 32,239 cells, 30 Leiden clusters) — for cell-composition deconvolution correction
  • CIBERSORTx or scVI reference signatures for compositional correction
  • Compute environment: Python (numpy/scipy for matrix pencil methods, or MATLAB Control System Toolbox-equivalent for Löwner interpolation), scikit-learn (CCA), statsmodels (permutation/FDR)
Success:
  • Healthy-cohort eigenvector conservation: pairwise subspace angle <15° across all 3 healthy datasets (mean angle reported with 95% CI).
  • MS deviation signal: subspace angle >30° from healthy consensus, permutation p<0.01 (BH-corrected).
  • Deviation survives cell-composition correction (angle reduction <20% after regression).
  • Biological convergence: eigenvector loadings enriched (GSEA FDR<0.1) for ≥2 of the 5 known target genes/CA-RIM signature.
  • Outperforms naive PCA baseline by ≥0.05 AUC in MS/healthy classification (cross-validated).
  • Cross-cohort reproducibility: >60% subspace overlap between GSE138614- and GSE193770-derived deviation subspaces.
Failure:
  • Healthy cohorts fail conservation criterion (pairwise angle >30° in ≥2/3 pairs) — indicates method is not detecting a stable manifold, likely noise-dominated.
  • MS deviation indistinguishable from permutation null (p>0.05 after correction).
  • Deviation signal disappears (>50% angle reduction) after cell-composition correction — indicates artifact of cell proportion shifts, not a genuine "manifold" effect.
  • No GSEA enrichment overlap with known CA-RIM/target genes — spectral method finds a mathematically real but biologically disconnected signal.
  • Performance parity or worse than naive PCA baseline (ΔAUC <0.03) — interpolation adds no discriminative value over standard dimensionality reduction.

ROI Projection

Implementation Sketch

# Pseudocode
load_and_harmonize(datasets=[GSE193770, GSE108000, GSE138614, GTEx_v10])
genes = intersect_top_variable_genes(datasets, n=5000)
groups = {healthy_1, healthy_2, healthy_3, MS_relapsing, MS_smoldering}

corr_matrices = {g: pearson_corr(expr[g][:, genes]) for g in groups}
corr_matrices = {g: psd_correct(m) for g, m in corr_matrices.items()}

# Complex interpolation / matrix pencil continuation across healthy matrices
healthy_family = loewner_interpolate(corr_matrices[healthy_1],
                                      corr_matrices[healthy_2],
                                      corr_matrices[healthy_3])
k = marchenko_pastur_threshold(healthy_family.eigenvalues)
healthy_eigvecs = top_k_eigenvectors(healthy_family, k)

for ms_group in [MS_relapsing, MS_smoldering]:
    ms_eigvecs = top_k_eigenvectors(corr_matrices[ms_group], k)
    angle = subspace_angle(healthy_eigvecs, ms_eigvecs)  # CCA/Grassmannian
    null_dist = permutation_test(corr_matrices, labels, n_iter=1000)
    p_value = compute_p(angle, null_dist)

# Correction for cell composition
expr_corrected = regress_out_cell_fractions(expr, scvi_atlas_fractions)
repeat_pipeline(expr_corrected)

# Biological convergence check
gsea_result = run_gsea(eigenvector_loadings, gene_sets=[CA_RIM_394, target_hierarchy_5])

# Baseline comparison
naive_pca_auc = cross_val_classify(pooled_pca_projection, labels)
spectral_auc = cross_val_classify(subspace_deviation_features, labels)
Abort checkpoints:
  • Day 7: If healthy-cohort conservation criterion fails (pairwise angle >30° in ≥2/3 pairs) on preliminary 3-cohort comparison — abort, method does not detect stable structure.
  • Day 15: If PSD-correction step requires >20% eigenvalue floor adjustment (indicating matrices are too rank-deficient/noisy for reliable interpolation) — abort or redesign with larger cohorts.
  • Day 25: If MS deviation signal is not distinguishable from permutation null at nominal p<0.05 (pre-FDR) — abort before full FDR/replication run.
  • Day 35: If deviation signal drops >50% after cell-composition correction — reclassify as compositional artifact, do not proceed to biological enrichment step.
  • Day 40: If no GSEA enrichment for known target genes and naive PCA baseline is equivalent — abort before commercial/ROI framing work.

NAMED_EXPERTS: []

CLOSEST_EXISTING_WORK: []

NOVELTY_NARROWING_REQUIRED: false

SPINE_STATEMENT: This hypothesis tests whether complex-interpolation-based eigenvector continuation of gene-gene correlation matrices reveals a statistically conserved healthy-tissue subspace that measurably and reproducibly deviates in MS transcriptomic data beyond what cell-type composition differences alone explain.

Source

AegisMind Research
Need AI to work rigorously on your problems? AegisMind uses the same multi-model engine for personal and professional use. Get started