The spectral analysis methods from complex interpolation of matrices can reveal conserved eigenvector structures between healthy and Multiple Sclerosis-affected tissues in cross-tissue transcriptomic data, enabling the identification of disease-specific manifold deviations.
Adversarial Debate Score
57% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Supporting Research Papers
- Machine Learning for analysis of Multiple Sclerosis cross-tissue bulk and single-cell transcriptomics data
Multiple Sclerosis (MS) is a chronic autoimmune disease of the central nervous system whose molecular mechanisms remain incompletely understood. In this study, we developed an end-to-end machine learn...
- Complex Interpolation of Matrices with an application to Multi-Manifold Learning
Given two symmetric positive-definite matrices A, B \in \mathbb{R}^{n \times n}, we study the spectral properties of the interpolation A^{1-x} B^x for 0 \leq x \leq 1. The presence of `common structur...
- Matrix Product States for Modulated Symmetries: SPT, LSM, and Beyond
Matrix product states (MPS) provide a powerful framework for characterizing one-dimensional symmetry-protected topological (SPT) phases of matter and for formulating Lieb-Schultz-Mattis (LSM)-type con...
Literature Assessment
An LLM's reading of the literature — not computational verification.
Eigenvector structures may vary significantly in disease contexts.
Method: literature_meta · Result: inconclusive
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.
This discovery has a Claude-generated validation package with a full experimental design.
Precise Hypothesis
Restricting to cross-tissue bulk and pseudo-bulk transcriptomic datasets of MS versus matched healthy controls (CNS lesion tissue, PBMC, and CD8+ T-cell-sorted fractions from GSE193770, GSE108000, GSE138614, CELLxGENE Census), a spectral decomposition method based on complex interpolation of gene-gene correlation/covariance matrices (e.g., eigenvector continuation between healthy and disease matrices via matrix pencils/Löwner interpolation) will identify a low-rank subspace (k ≤ 20 eigenvectors) that is (a) statistically conserved across ≥3 independent healthy cohorts (subspace angle < 15° pairwise) and (b) exhibits a measurable, reproducible deviation in MS samples (subspace angle > 30° from healthy consensus, permutation p<0.01, Benjamini-Hochberg corrected across tissue types) that is NOT explainable by cell-type composition shifts alone (persists after CIBERSORTx/scVI deconvolution correction). Falsification: if no such conserved-yet-deviating subspace exists, or if the deviation vanishes after cell-composition correction, the hypothesis is false.
- Subspace angle between healthy-cohort eigenvector sets exceeds 30° (i.e., "conservation" claim itself fails) in ≥2 of 3 independent healthy datasets.
- MS-vs-healthy subspace deviation is statistically indistinguishable from healthy-vs-healthy cohort variation (permutation test p>0.05 after FDR correction).
- Deviation signal disappears or drops below noise floor after regressing out cell-type proportion estimates (scVI/CIBERSORTx), indicating the "manifold deviation" is merely a compositional artifact.
- Deviation subspace does not overlap (principal angle <45°) with the independently derived DEG-based CA-RIM signature (394 genes) or the composite-score target hierarchy (DNMT1, ZNF740, CTSS) — i.e., the spectral method finds nothing convergent with the known biology.
- Method fails to outperform a naive PCA-on-pooled-data baseline (no interpolation) in cross-validated MS/healthy classification (AUC difference <0.03).
Spine & Adversarial Read
- highWhy matrix-pencil/Löwner complex interpolation specifically, rather than simpler and more standard manifold methods (diffusion maps, UMAP, geodesic interpolation on the SPD manifold, or even straightforward joint PCA)? The methodology choice is not justified against these well-established alternatives.Not resolved in current design — the EVP does not present a comparative rationale for why complex interpolation of matrices (a technique with roots in control theory/rational approximation) is expected to outperform SPD-manifold geodesic methods or diffusion-map approaches already used in transcriptomic manifold learning. This must be addressed by adding an explicit baseline comparison arm (geodesic-on-SPD interpolation, diffusion maps) alongside the naive PCA baseline before claiming the specific method choice is warranted; absent this, a reviewer will reasonably ask 'why this math, not the standard math.'
- highThe claimed 'MS-specific manifold deviation' may be entirely reducible to the already-known cell-type composition shift (CD8+ T-cell expansion documented in the Phase 2 scVI atlas) — the spectral method could just be an elaborate, expensive re-derivation of a fact already established by simpler deconvolution, offering no orthogonal information.Partially addressed via the mandatory cell-composition regression step (Methodology step 9, Disproof criterion 3) — but the correction method (regressing out scVI-derived fractions) itself has known limitations (imperfect deconvolution accuracy, especially for rare subpopulations) and could leave residual composition signal that gets misattributed to 'manifold deviation.' Recommend adding a synthetic-data sanity check (simulate pure composition-shift data with no true manifold change) to confirm the method correctly returns a null result before trusting real-data findings.
- mediumSample sizes in the source cohorts (GSE193770, GSE138614) are likely far below the n≥30/group threshold stated as a boundary condition once split by phenotype (relapsing vs. smoldering) and healthy sub-cohorts — the eigenvector estimates may be statistically unstable regardless of the mathematical elegance of the interpolation method.Acknowledged in Known Failure Modes and mitigated by the Day 7/Day 15 abort checkpoints, but not resolved — actual sample sizes for GSE193770 and GSE138614 by phenotype subgroup are not confirmed in this package and should be verified before committing full budget; if subgroup n<20, the protocol should default to pseudobulk aggregation across broader Leiden clusters rather than per-phenotype splitting to preserve statistical power.
Experimental Protocol
Minimum viable test (MVT):
- Use GSE138614 (already replicated CTSS/FGF2/SLCO2B1 signal) as primary dataset (n≈50-90 depending on group sizes), GSE193770 CD8+ T-cell subset as secondary/validation, and a third independent healthy reference (GTEx CNS/blood, v10) as the "conservation" test set.
- Compute gene-gene correlation matrices per group (healthy_1, healthy_2, healthy_3, MS) restricted to top 2,000-5,000 variable genes intersected across datasets.
- Apply complex interpolation (matrix pencil / Löwner framework) between healthy matrices to define a "healthy manifold trajectory"; extract top-k eigenvectors (k selected by Marchenko-Pastur threshold, expect k=10-20).
- Interpolate/extrapolate to MS matrix; measure principal (subspace) angles between the MS eigenbasis and the healthy consensus basis via canonical correlation analysis (CCA).
- Run 1,000-permutation null (label shuffling) to establish significance of angle deviation.
- Cross-validate: does deviation subspace loadings enrich for CA-RIM DEGs (394 genes) and target hierarchy genes (DNMT1, ZNF740, CTSS, FGF2, SLCO2B1) via GSEA on eigenvector loadings?
- GSE193770 (CD8+ T cells, CNS-associated, primary discovery cohort)
- GSE108000 (MS bulk transcriptomics)
- GSE138614 (independent replication cohort, already used for CTSS/FGF2/SLCO2B1)
- CELLxGENE Census (cross-tissue healthy reference atlas)
- GTEx v10 (independent healthy baseline, blood/CNS tissue)
- Existing scVI atlas (gs://aegismind-tpu-results/ms_phase2/results/, 32,239 cells, 30 Leiden clusters) — for cell-composition deconvolution correction
- CIBERSORTx or scVI reference signatures for compositional correction
- Compute environment: Python (numpy/scipy for matrix pencil methods, or MATLAB Control System Toolbox-equivalent for Löwner interpolation), scikit-learn (CCA), statsmodels (permutation/FDR)
- Healthy-cohort eigenvector conservation: pairwise subspace angle <15° across all 3 healthy datasets (mean angle reported with 95% CI).
- MS deviation signal: subspace angle >30° from healthy consensus, permutation p<0.01 (BH-corrected).
- Deviation survives cell-composition correction (angle reduction <20% after regression).
- Biological convergence: eigenvector loadings enriched (GSEA FDR<0.1) for ≥2 of the 5 known target genes/CA-RIM signature.
- Outperforms naive PCA baseline by ≥0.05 AUC in MS/healthy classification (cross-validated).
- Cross-cohort reproducibility: >60% subspace overlap between GSE138614- and GSE193770-derived deviation subspaces.
- Healthy cohorts fail conservation criterion (pairwise angle >30° in ≥2/3 pairs) — indicates method is not detecting a stable manifold, likely noise-dominated.
- MS deviation indistinguishable from permutation null (p>0.05 after correction).
- Deviation signal disappears (>50% angle reduction) after cell-composition correction — indicates artifact of cell proportion shifts, not a genuine "manifold" effect.
- No GSEA enrichment overlap with known CA-RIM/target genes — spectral method finds a mathematically real but biologically disconnected signal.
- Performance parity or worse than naive PCA baseline (ΔAUC <0.03) — interpolation adds no discriminative value over standard dimensionality reduction.
ROI Projection
Implementation Sketch
# Pseudocode load_and_harmonize(datasets=[GSE193770, GSE108000, GSE138614, GTEx_v10]) genes = intersect_top_variable_genes(datasets, n=5000) groups = {healthy_1, healthy_2, healthy_3, MS_relapsing, MS_smoldering} corr_matrices = {g: pearson_corr(expr[g][:, genes]) for g in groups} corr_matrices = {g: psd_correct(m) for g, m in corr_matrices.items()} # Complex interpolation / matrix pencil continuation across healthy matrices healthy_family = loewner_interpolate(corr_matrices[healthy_1], corr_matrices[healthy_2], corr_matrices[healthy_3]) k = marchenko_pastur_threshold(healthy_family.eigenvalues) healthy_eigvecs = top_k_eigenvectors(healthy_family, k) for ms_group in [MS_relapsing, MS_smoldering]: ms_eigvecs = top_k_eigenvectors(corr_matrices[ms_group], k) angle = subspace_angle(healthy_eigvecs, ms_eigvecs) # CCA/Grassmannian null_dist = permutation_test(corr_matrices, labels, n_iter=1000) p_value = compute_p(angle, null_dist) # Correction for cell composition expr_corrected = regress_out_cell_fractions(expr, scvi_atlas_fractions) repeat_pipeline(expr_corrected) # Biological convergence check gsea_result = run_gsea(eigenvector_loadings, gene_sets=[CA_RIM_394, target_hierarchy_5]) # Baseline comparison naive_pca_auc = cross_val_classify(pooled_pca_projection, labels) spectral_auc = cross_val_classify(subspace_deviation_features, labels)
- Day 7: If healthy-cohort conservation criterion fails (pairwise angle >30° in ≥2/3 pairs) on preliminary 3-cohort comparison — abort, method does not detect stable structure.
- Day 15: If PSD-correction step requires >20% eigenvalue floor adjustment (indicating matrices are too rank-deficient/noisy for reliable interpolation) — abort or redesign with larger cohorts.
- Day 25: If MS deviation signal is not distinguishable from permutation null at nominal p<0.05 (pre-FDR) — abort before full FDR/replication run.
- Day 35: If deviation signal drops >50% after cell-composition correction — reclassify as compositional artifact, do not proceed to biological enrichment step.
- Day 40: If no GSEA enrichment for known target genes and naive PCA baseline is equivalent — abort before commercial/ROI framing work.
NAMED_EXPERTS: []
CLOSEST_EXISTING_WORK: []
NOVELTY_NARROWING_REQUIRED: false
SPINE_STATEMENT: This hypothesis tests whether complex-interpolation-based eigenvector continuation of gene-gene correlation matrices reveals a statistically conserved healthy-tissue subspace that measurably and reproducibly deviates in MS transcriptomic data beyond what cell-type composition differences alone explain.