On the SARS-CoV-2 Mpro retrospective benchmark (753 ChEMBL compounds: 257 measured actives and 496 measured inactives at pChEMBL<5, receptor PDB 7VU6), AutoDock Vina's below-chance ranking (AUROC 0.427) is a SCORING failure and not a pose-SAMPLING failure. Therefore: substituting DiffDock-L pose generation while retaining Vina scoring will change AUROC by less than 0.05, whereas rescoring poses with the gnina CNN scoring function will raise AUROC by at least 0.10 over the 0.427 baseline. Neither arm will reach the 0.763 AUROC obtained from seven physicochemical descriptors alone on the same 753 compounds.
Adversarial Debate Score
65% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Supporting Research Papers
- Measuring the Re-executability of Published Molecular Docking Claims
Published molecular docking scores depend on the receptor, ligand, software, search box, seed, and preparation choices; a paper reporting only the score has published a number with unknowable provenan...
- BCover: An Electronic Structure-Based Scoring Suite for Reaction-Aware Covalent Docking
Covalent virtual screening requires ranking compounds according to both noncovalent recognition and their ability to adopt a reaction-competent geometry with an appropriately reactive warhead. Here, w...
- Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disagree without indicating which prediction to trust. Consensus s...
- Reducing cross-sample prediction churn in scientific machine learning
Scientific machine learning reports predictive performance. It does not report whether the same prediction would survive a different draw of training data. Across 9 chemistry benchmarks, two classifie...
- Accelerated descriptor-free path sampling for protein-ligand binding kinetics
The kinetics of protein-ligand binding systems are increasingly recognized as a key determinant of drug efficacy, yet remain far harder to compute than binding affinities. Existing kinetics methods ei...
Computational Result
A structure-prediction score, not evidence of binding. No wet-lab assay has been run.
Context: in our pre-registered retrospective benchmark of this docking pipeline (git 92c6f8bf, 14 July 2026), 14 targets were assessed for admissibility and all but one were excluded by the benchmark’s own decoy-bias criterion. The single target that could be scored returned EF@1% = 0.00 — no known actives recovered in the top 1% — against 23.1 for a plain 2D fingerprint baseline. Read the score below as a way of ordering what to test first, not as evidence that this compound binds.
SUPPORTED, in a narrower form than first published. TWO INTERVENTIONS CLEAR THE FLOOR AND BOTH CONCERN WHAT IS COMPUTED FROM THE POSES. Rescoring identical poses with a CNN moved AUROC 0.3992 -> 0.5389 (+0.1397 [+0.0913,+0.1867]). Reading the whole pose ensemble rather than the top pose gives +0.055 [+0.043,+0.068] on the repaired receptor, every one of 40 cross-validation fold seeds above the floor. WHAT DOES NOT MATTER IS HOW THE POSES ARE FOUND: eight-fold search effort -0.019 [-0.032,-0.006] with an unchanged top ten, and supplying the correct binding site to Boltz-2 -0.013 [-0.026,+0.000]. Both sit below the 0.039 protocol floor, and the exhaustiveness-4/32 pair IS that floor, so eight times the compute is by construction indistinguishable from running the same protocol twice. THE SCORING GAIN STILL DOES NOT SURVIVE DECOMPOSITION: regressing seven free physicochemical descriptors out of both scores leaves gnina's residual at 0.4994 against plain Vina's 0.5187 -- the CNN's advantage is compound-intrinsic, not target recognition. THERE ARE TWO FLOORS, NOT ONE: 0.020 [0.011,0.029] for two scoring functions on fixed poses, 0.039 [0.024,0.056] for two runs of a whole protocol.
Method: Six single-factor interventions on a 751-compound SARS-CoV-2 Mpro panel (PDB 7VU6), each pre-registered before its data existed, read against two measured resolution floors. AutoDock Vina 1.2.5, gnina 1.3, DiffDock-L, Boltz-2 2.2.1. analysis/docking_value, analysis/pose_ensemble, analysis/three_arm_docking, analysis/boltz2 · Result: supported · Confidence: 0%
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.
This discovery has a Claude-generated validation package with a full experimental design.
Precise Hypothesis
On the fixed retrospective benchmark of 753 ChEMBL-annotated compounds (257 actives, pChEMBL≥5; 496 inactives, pChEMBL<5) docked against SARS-CoV-2 Mpro (PDB 7VU6, single rigid receptor conformation), three arms will be compared holding the compound set and receptor fixed: (A) AutoDock Vina pose generation + Vina scoring (baseline, expected AUROC 0.427); (B) DiffDock-L pose generation + Vina rescoring of the top DiffDock-L pose per ligand; (C) AutoDock Vina pose generation + gnina CNN rescoring of the top Vina pose per ligand. Claim 1 (sampling is not the bottleneck): AUROC(B) − AUROC(A) < 0.05 in absolute value. Claim 2 (scoring is the bottleneck): AUROC(C) − AUROC(A) ≥ 0.10. Claim 3 (ceiling): AUROC(B) < 0.763 and AUROC(C) < 0.763, where 0.763 is the AUROC of a logistic regression on seven physicochemical descriptors (MW, logP, HBD, HBA, TPSA, rotatable bonds, formal charge) alone on the same 753 compounds.
- If AUROC(B) − AUROC(A) ≥ 0.05, sampling contributes materially and Claim 1 is disproved.
- If AUROC(C) − AUROC(A) < 0.10, CNN rescoring does not recover claimed performance and Claim 2 is disproved.
- If either AUROC(B) or AUROC(C) ≥ 0.763, the "scoring-not-sampling, but still capped below descriptors" narrative is disproved.
- If baseline Vina AUROC on re-run deviates from 0.427 by more than ±0.03 (implementation/parameter mismatch), original benchmark is not reproduced and the entire comparison is invalid until resolved.
- If bootstrap 95% CI on any AUROC difference crosses zero (for the <0.05 claim) or crosses 0.10 (for the ≥0.10 claim), result is statistically inconclusive rather than a clean proof/disproof.
Spine & Adversarial Read
- highThe design conflates 'pose sampling quality' with 'pose selection/confidence ranking' — DiffDock-L's top-1 pose is chosen by its own confidence model, which is itself a learned scoring function; so Arm B is not a pure test of sampling, it already has scoring contamination baked into pose selection.Partially resolved by the RMSD-to-crystal validation step (Methodology step 9), which can show whether DiffDock-L's sampled pose ensemble contains near-native poses even if top-1 selection is wrong — but the EVP does not fully decouple 'DiffDock-L sampling' from 'DiffDock-L confidence scoring' as an independent ablation (e.g. re-scoring all 10 DiffDock-L-sampled poses with Vina, not just the top-1). This should be added as an additional arm (B') for a fully clean test.
- mediumWhy these specific methods (Vina, DiffDock-L, gnina) and this specific benchmark (753 ChEMBL compounds, 7VU6) rather than other established scoring functions (e.g. AutoDock4, Glide SP/XP) or other sampling methods (e.g. Glide, GOLD)? The generalizability of 'scoring not sampling' conclusions rests entirely on this one sampling method and one CNN scorer being representative of their respective classes.Not resolved in this EVP — methodology choice is justified only by availability/reproducibility (both tools are open-source and widely used), not by a principled argument that DiffDock-L and gnina are representative of 'best-in-class sampling' and 'best-in-class scoring' respectively. A stronger version would include at least one additional sampler (e.g. Glide or GOLD) and one additional ML scorer (e.g. RTMScore) as robustness arms; this is flagged as a scope limitation of the minimum-viable design.
- mediumThe 0.763 physicochemical-descriptor AUROC ceiling may reflect a benchmark construction artifact (e.g. actives systematically differing from inactives in molecular weight or logP due to how ChEMBL assay data was filtered) rather than a genuine 'ligand efficiency ceiling' — if so, all four arms are being compared against a metric that is itself confounded, undermining the interpretive claim that no docking method should be expected to beat it.Acknowledged but not resolved. The protocol should additionally report the individual discriminative power of each of the 7 descriptors and check for a simple confound (e.g. MW alone achieving high AUROC), which would indicate the benchmark itself over-separates actives/inactives by trivial physicochemical property rather than true binding-relevant chemistry — this analysis is not currently in the methodology and should be added before treating 0.763 as a meaningful ceiling.
Experimental Protocol
Minimum viable test: single-receptor, single-conformation, 753-compound benchmark, 3 arms (A/B/C) plus descriptor-only baseline (D, already established at 0.763 — replicate to confirm), n=1 run per arm with fixed random seeds, bootstrap resampling (2,000 iterations) for confidence intervals on AUROC differences. Full validation adds: 3 independent seeds per stochastic arm (DiffDock-L, gnina CNN ensemble), a second receptor conformation (alternate Mpro PDB, e.g. 6LU7) as a robustness check, and pose-quality RMSD-to-crystal validation for the top-N compounds with available co-crystal structures.
- 753-compound ChEMBL benchmark set (SMILES, pChEMBL values, active/inactive labels at pChEMBL=5 threshold) — must be reconstructed/confirmed from ChEMBL API + original benchmark curation notes.
- PDB 7VU6 (primary receptor) and PDB 6LU7 (secondary, robustness) — prepared with consistent protonation (PDB2PQR/Reduce), same binding-site definition (grid box centered on catalytic Cys145/His41).
- AutoDock Vina v1.2.x (deterministic mode, exhaustiveness=32, seed fixed).
- DiffDock-L (public checkpoint, default inference settings, 10 poses/ligand, confidence-model top-1 selection).
- gnina (CNN scoring, "default" and "dense" ensemble checkpoints) for rescoring.
- RDKit (descriptor calculation for 7-feature physicochemical baseline) + scikit-learn logistic regression.
- Compute environment: single-node GPU (for DiffDock-L and gnina CNN inference), CPU cluster for Vina docking (embarrassingly parallel per-ligand).
- Claim 1 supported: |AUROC(B) − AUROC(A)| < 0.05, with 95% CI upper bound < 0.07.
- Claim 2 supported: AUROC(C) − AUROC(A) ≥ 0.10, with 95% CI lower bound > 0.05.
- Claim 3 supported: AUROC(B) and AUROC(C) both < 0.763 (95% CI upper bound < 0.75).
- Baseline reproduction: re-run Vina AUROC within ±0.03 of 0.427.
- Robustness: directionally consistent results (same claim outcomes) on second receptor conformation (6LU7).
- DiffDock-L + Vina scoring improves AUROC by ≥0.05 → sampling is a nontrivial contributor, hypothesis fails.
- gnina CNN rescoring improves AUROC by <0.10 → scoring-fix claim overstated, hypothesis fails.
- Either rescoring arm meets or exceeds 0.763 → "capped below descriptors" claim fails.
- Baseline Vina AUROC does not reproduce 0.427 (±0.03) → benchmark/environment mismatch invalidates comparison; must resolve before interpreting other arms.
- Results flip sign or claim status on second receptor (6LU7) → conclusion is receptor-conformation-specific, not general.
ROI Projection
Implementation Sketch
# Pseudocode compounds = load_chembl_benchmark(n=753) # SMILES, pChEMBL, label receptor_7vu6 = prepare_receptor("7VU6.pdb") receptor_6lu7 = prepare_receptor("6LU7.pdb") # robustness arm def arm_A(compounds, receptor): poses = vina_dock(compounds, receptor, exhaustiveness=32, seed=42) scores = [p.top_vina_score for p in poses] return poses, auroc(scores, labels) def arm_B(compounds, receptor, vina_poses): diff_poses = diffdock_L_infer(compounds, receptor, n_samples=10) top_poses = select_top1_by_confidence(diff_poses) rescored = vina_score_only(top_poses, receptor) # no re-docking return auroc(rescored, labels) def arm_C(vina_poses, receptor): cnn_scores = gnina_rescore(vina_poses, receptor, ensemble="default") return auroc(cnn_scores, labels) def arm_D(compounds): descriptors = rdkit_descriptors(compounds, features=7) model = LogisticRegression(cv=5).fit(descriptors, labels) return auroc(model.cv_predict(), labels) for receptor in [receptor_7vu6, receptor_6lu7]: poses_A, auroc_A = arm_A(compounds, receptor) auroc_B = arm_B(compounds, receptor, poses_A) auroc_C = arm_C(poses_A, receptor) auroc_D = arm_D(compounds) # receptor-independent, compute once bootstrap_ci([auroc_A, auroc_B, auroc_C, auroc_D], n=2000) evaluate_claims(auroc_A, auroc_B, auroc_C, auroc_D)
- Checkpoint 1 (Day 2): If Vina baseline re-run does not reproduce AUROC 0.427 ± 0.03 — abort and debug environment/parameters before proceeding to Arms B/C.
- Checkpoint 2 (Day 5): If DiffDock-L fails to converge or produces >20% invalid poses (steric clashes, out-of-box) — abort Arm B pending pose-filtering fix.
- Checkpoint 3 (Day 8): If gnina CNN rescoring shows suspiciously perfect performance (AUROC > 0.95) — pause and investigate training-data leakage before reporting Claim 2 as supported.
- Checkpoint 4 (Day 10): If bootstrap CIs for all three claims are wide and inconclusive (span >0.15) after 2,000 iterations — consider this an underpowered result requiring either a larger/alternate compound set or explicit reporting as inconclusive rather than forcing a pass/fail call.
NAMED_EXPERTS: []
CLOSEST_EXISTING_WORK: []
NOVELTY_NARROWING_REQUIRED: false
SPINE_STATEMENT: This hypothesis tests whether AutoDock Vina's poor ranking performance on the SARS-CoV-2 Mpro benchmark is caused by its scoring function rather than its pose-sampling algorithm.