solver.press

The UCB acquisition strategy's superiority over EI in high-uncertainty drug-discovery landscapes (validated across 6/6 targets) will transfer to MSH3 Walker-A pocket screening when an unrelated P-loop ATPase counter-screen (adenylate kinase) is embedded as a hard constraint in the Bayesian optimisation objective, enabling discovery of genuinely MSH3-selective compounds by penalising candidates whose predicted affinity for 1AKE equals or exceeds their MSH3 score, directly addressing CHDI's off-target selectivity liability.

ChemistryAug 17, 2026Evaluation Score: 63%

Adversarial Debate Score

50% survival rate under critique

Expert panel critique

Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.

Mistral: The hypothesis is well-supported by the owner's validated experiments (UCB superiority in high-uncertainty landscapes) and addresses a clear, falsifiable objective (MSH3 selectivity via Bayesian optimization constraints). However, its transferability to MSH3—while plausible—remains untested in th...
ChatGPT: The hypothesis is falsifiable and is supported by the validated 6/6 UCB-over-EI result, but transfer to the MSH3 Walker-A pocket remains untested. A single unrelated ATPase counter-screen—especially using potentially miscalibrated cross-protein predicted affinities—cannot by itself establish genu...
Grok: UCB>EI is validated (6/6), so the acquisition claim is solid, but transfer to MSH3 Walker-A plus the 1AKE hard-constraint selectivity scheme is an untested extrapolation; prior MSH3 docking used the wrong chain (MSH2), weakening confidence that the proposed setup will yield genuinely selective hits.
Adversarial skeptic · via ChatGPT: — Penalizing predicted 1AKE binding cannot establish genuine MSH3 selectivity because an unrelated adenylate kinase is not a biologically relevant proxy for the off-target landscape, and cross-protein affinity scores are not reliably comparable.

The strict critic was recused on this topic; an adversarial reviewer stood in to keep scrutiny intact.

Supporting Research Papers

Computational Result

❌ Refuted by computation· cross-protein docking score comparability, 72 compounds x 2 unrelated receptors (analysis/cross_target_selectivity/cross_target.py)

The computation ran and did not support the hypothesis.

Context: in our pre-registered retrospective benchmark of this docking pipeline (git 92c6f8bf, 14 July 2026), 14 targets were assessed for admissibility and all but one were excluded by the benchmark’s own decoy-bias criterion. The single target that could be scored returned EF@1% = 0.00 — no known actives recovered in the top 1% — against 23.1 for a plain 2D fingerprint baseline. Read the score below as a way of ordering what to test first, not as evidence that this compound binds.

THE COUNTER-CONSTRAINT CANNOT WORK. Docking scores for 72 compounds against AcrB (bacterial efflux transporter) and CTSS (human cysteine protease) — structurally and biologically unrelated — correlate at r=+0.932, so 87% of a compound's score is NOT target-specific. The 'selectivity' difference has sd 0.676 against 1.773 for a single score, sitting on a constant -1.08 kcal/mol offset: a small residual, not an independent measurement. The residual tracks clogP (-0.378) more than molecular weight (-0.174), so what survives the subtraction is partly lipophilicity, not discrimination. The adversarial critic's objection is confirmed.

Method: cross-protein docking score comparability, 72 compounds x 2 unrelated receptors (analysis/cross_target_selectivity/cross_target.py) · Result: refuted · Confidence: 0%

Formal Verification

Z3 logical consistency:✅ Consistent

Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.

Experimental Validation Package

This discovery has a Claude-generated validation package with a full experimental design.

Precise Hypothesis

When Bayesian optimisation using an Upper Confidence Bound (UCB, κ=2.0–2.5) acquisition function is applied to virtual screening of MSH3 Walker-A pocket ligands, and a hard constraint is imposed such that predicted pAffinity(1AKE) < predicted pAffinity(MSH3) − δ (δ ≥ 0, tuned; default 0.5 log units) for every candidate accepted into the optimisation trajectory, the resulting compound set will show (a) statistically significant improvement in mean predicted MSH3-selectivity margin over an unconstrained UCB baseline and an Expected Improvement (EI) baseline (p<0.05, Mann-Whitney U), and (b) equal or better sample efficiency (number of oracle calls to reach a fixed selectivity/affinity threshold) than EI, replicating the previously observed 6/6-target UCB>EI advantage in this new counter-constrained regime. The hypothesis is falsified if constrained-UCB fails to beat constrained-EI on ≥4/6 comparable metrics, or if the counter-constraint fails to reduce false-positive off-target binders (defined via docking/FEP validation) relative to an unconstrained selectivity-scoring post-filter.

Disproof criteria:
  1. Constrained-UCB does not outperform constrained-EI in sample efficiency (oracle calls to reach threshold) on ≥3/6 held-out target-emulation splits or synthetic benchmark tasks.
  2. The hard constraint fails to reduce the docking-validated off-target hit rate (compounds where 1AKE score ≥ MSH3 score) below that of a naive post-hoc selectivity filter applied to unconstrained UCB output.
  3. Final selected compound set shows no statistically significant selectivity margin improvement (Δ pAffinity MSH3−1AKE) over baseline (p≥0.05).
  4. Confirmatory docking/FEP or (if budget allows) biochemical assay shows ≥50% of top-10 "selective" candidates are false positives (bind 1AKE-analog counter-targets with comparable or higher affinity than predicted).

Spine & Adversarial Read

  • highA single counter-target (1AKE) is a weak proxy for real-world off-target liability; passing this in-silico test says nothing about selectivity against the broader P-loop ATPase/kinase panel CHDI actually cares about.
    Protocol explicitly scopes claim to single-counter-target validation only (see BOUNDARY_CONDITIONS) and lists multi-target generalization as a downstream UNLOCK, not a claim made here; this is an acknowledged gap requiring follow-up with a broader counter-target panel before any translational claim to CHDI's actual liability profile.
  • highWhy UCB/EI Bayesian optimisation and docking/ML-surrogate scoring specifically, rather than simpler multi-objective methods (e.g., Pareto-based genetic algorithms, multi-task GP with explicit selectivity objective, or simple post-hoc filtering)? The methodology choice inherits its rationale entirely from an external prior claim (6/6 targets) not re-justified here.
    The EVP depends on validating the cited prior 6/6-target UCB>EI benchmark as a DEPENDENCY, but does not itself argue why hard-constrained UCB is mechanistically superior to alternative multi-objective BO formulations (e.g., ParEGO, qNEHVI) for this selectivity problem; a baseline comparison against at least one Pareto-native multi-objective acquisition function should be added to strengthen methodology justification.
  • mediumSurrogate-model-only validation (docking/ML scores) with no wet-lab confirmation risks the entire result being an artifact of correlated errors between the MSH3 and 1AKE scoring functions (e.g., both models sharing systematic biases from similar fingerprint features), producing spurious selectivity signals that would not replicate experimentally.
    Protocol includes independent docking re-scoring as a partial orthogonal check and flags FEP/biochemical assay as optional follow-up; however, budget (COST_USD_MIN) does not guarantee wet-lab confirmation, so the core validation remains in-silico-only unless full budget is committed — this limitation should be stated explicitly in any publication of results.

Experimental Protocol

Minimum viable test: an in-silico benchmark using two fixed, pre-trained surrogate scoring models (one for MSH3 Walker-A pocket, one for 1AKE ATP pocket) built from existing docking or ML-affinity datasets, run through 4 optimisation arms (UCB-constrained, UCB-unconstrained, EI-constrained, EI-unconstrained) × 10 random seeds × 200 oracle-call budget, on a shared candidate pool of ~50,000 purchasable compounds (Enamine REAL subset or ZINC20 "lead-like" tranche). Compare sample efficiency, final selectivity margin, and Pareto-front hypervolume (MSH3 affinity vs. −1AKE affinity). Follow-up with docking re-scoring (AutoDock-GPU/Glide SP) of top-20 hits per arm for orthogonal validation; optional FEP+ (10 top compounds) or biochemical assay if resources allow.

Required datasets:
  • MSH3 Walker-A pocket structure(s): PDB (if available) or homology model built from MutSβ complex structures; docking grid definition.
  • 1AKE adenylate kinase crystal structure (PDB: 1AKE, 4AKE apo/holo) with ATP-pocket grid definition.
  • Compound library: Enamine REAL Diversity subset (~50K) or ZINC20 lead-like tranche with 3D conformers pre-generated (RDKit ETKDG).
  • Pre-trained or in-house-trained ML affinity/docking-score surrogate models (e.g., Gaussian Process on Morgan fingerprints + docking score training set of ≥5,000 compounds per target for surrogate calibration).
  • Existing benchmark dataset validating UCB>EI across the "6/6 targets" claim (need access to that prior experiment's code/data for exact replication of methodology).
  • Docking software: AutoDock-GPU or Glide (academic/commercial license) for re-scoring.
  • Optional: FEP+ (Schrödinger) or OpenFE for top-10 confirmatory free-energy calculations.
  • Compute environment: BoTorch/Ax or GPyOpt for Bayesian optimisation implementation with custom constrained acquisition function.
Success:
  • Constrained-UCB reaches fixed target threshold (top-1% predicted MSH3 affinity + selectivity margin ≥1.0 log unit) in fewer or equal median oracle calls than constrained-EI in ≥4/6 comparable benchmark splits/seeds-groups (replicating original 6/6-type pattern, allowing 1 miss).
  • Mean selectivity margin (pMSH3 − p1AKE) of constrained-UCB top-20 set is ≥0.5 log units higher than unconstrained-UCB top-20 set (p<0.05).
  • Docking-validated false-positive rate (compounds where re-docked 1AKE score ≥ MSH3 score) for constrained-UCB ≤50% of the rate observed in unconstrained-UCB + post-hoc filter baseline.
  • Pareto hypervolume for constrained-UCB significantly exceeds constrained-EI (p<0.05, paired across seeds).
Failure:
  • Constrained-UCB and constrained-EI show no statistically significant difference in sample efficiency (p≥0.05) or EI outperforms UCB in ≥3/6 comparisons.
  • Hard constraint produces <10% improvement in selectivity margin versus simple post-hoc filtering of unconstrained runs, indicating the "hard constraint during optimisation" framing adds no value over post-hoc screening.
  • Surrogate model calibration fails (uncertainty estimates uninformative), invalidating UCB's theoretical advantage a priori — treated as an abort condition, not merely a failure of the hypothesis.
  • Docking/FEP re-scoring shows >50% of predicted-selective top compounds are actually non-selective or non-binders, indicating the surrogate models (not the acquisition strategy) are the bottleneck.

ROI Projection

Implementation Sketch

# Pseudocode
for arm in [UCB_constrained, EI_constrained, UCB_unconstrained, EI_unconstrained]:
    for seed in range(10):
        GP_msh3 = fit_surrogate(train_data_msh3)
        GP_1ake = fit_surrogate(train_data_1ake)
        observed = init_random_sample(pool, n=20, seed=seed)
        for t in range(200):
            candidates = pool - observed
            mu_m, sigma_m = GP_msh3.predict(candidates)
            mu_a, sigma_a = GP_1ake.predict(candidates)
            if arm.constrained:
                mask = (mu_a + z*sigma_a) < (mu_m - z*sigma_m) - delta
                candidates = candidates[mask]
                if empty: relax_or_break()
            score = acquisition(arm.type, mu_m, sigma_m)  # UCB or EI
            x_next = argmax(score)
            y_msh3, y_1ake = oracle(x_next)  # docking/FEP ground truth
            update(GP_msh3, GP_1ake, x_next, y_msh3, y_1ake)
            log_metrics(t, best_so_far, selectivity_margin, hypervolume)
compare_arms(logs)  # statistical tests
docking_revalidate(top20_per_arm)
Abort checkpoints:
  • Day 5: If surrogate model calibration checks fail (coverage <70% at 90% CI) for either target, abort and revisit training data/model architecture before continuing.
  • Day 12: If constrained pools are empty >30% of iterations across seeds (constraint too strict), abort/redesign delta or relaxation scheme before running full comparison.
  • Day 25: If interim analysis (5/10 seeds) shows no directional trend favoring constrained-UCB over constrained-EI (effect size <0.1), consider stopping before completing full seed set and docking re-validation to save cost.
  • Day 35: If docking re-scoring of interim top candidates shows >60% false-positive selectivity rate, abort before committing to FEP/biochemical follow-up spend.

NAMED_EXPERTS: []

CLOSEST_EXISTING_WORK: []

NOVELTY_NARROWING_REQUIRED: false

SPINE_STATEMENT: This hypothesis tests whether embedding adenylate kinase as a hard counter-screening constraint inside a UCB-based Bayesian optimisation loop yields MSH3-selective compound candidates more sample-efficiently and with fewer off-target false positives than an equivalent EI-based or unconstrained approach.

Source

AegisMind Research
Need AI to work rigorously on your problems? AegisMind uses the same multi-model engine for personal and professional use. Get started