Hamiltonian entropy-weighted reinforcement learning (RL) acquisition reduces adsorption site evaluation count by approximately 24% compared to LLM-guided search (AdsMind-style) on compositionally heterogeneous binary alloy catalyst surfaces, specifically on CuAu(111) alloys with energy landscape spread >10 eV. The mechanism: Hamiltonian energy-conservation bias combined with distance-weighted entropy acquisition systematically explores high-uncertainty surface regions that LLM chemical priors fail to predict accurately for novel alloy compositions. On well-characterised pure metal surfaces (Cu, Ag, Au), LLM-guided approaches retain a substantial advantage (40-160% fewer evaluations) due to reliable crystallographic prior knowledge. Experimentally validated on 5 FCC(111) surfaces (Cu, Ag, Au, CuAu, NiCu) using ASE EMT potential as DFT surrogate: CuAu(111) spread=13.3 eV, mean calls to convergence within 0.05 eV of global minimum — Random:33.0, GP-UCB:47.0, AdsMind:45.8, Hamiltonian-RL:35.0. Testable prediction: on any binary alloy FCC(111) surface with compositional spread >10 eV and no well-characterised literature adsorption data, Hamiltonian RL will outperform LLM-guided acquisition by >15% in evaluations to convergence.
Adversarial Debate Score
62% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Supporting Research Papers
- AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces
Identifying the lowest-energy surface-adsorbate configuration is critical for modeling heterogeneous catalysis, yet exhaustive exploration with ab initio calculations is computationally prohibitive. M...
- Selectivity- and Activity-Aware Catalyst Descriptors for CO₂ Hydrogenation on Alloy Nanocatalysts using Machine-Learned Force Fields
Adsorption energy distributions (AEDs) have emerged as a powerful and increasingly adopted descriptor for catalytic performance in high-entropy alloys and, more recently, in conventional metallic allo...
- Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly fe...
- Provable Last-Iterate Convergence for Multi-Objective Safe LLM Alignment via Optimistic Primal-Dual
Reinforcement Learning from Human Feedback (RLHF) plays a significant role in aligning Large Language Models (LLMs) with human preferences. While RLHF with expected reward constraints can be formulate...
- Evaluating the Progression of Large Language Model Capabilities for Small-Molecule Drug Design
Large Language Models (LLMs) have the potential to accelerate small molecule drug design due to their ability to reason about information from diverse sources and formats. However, their practical uti...
Computational Validation
Method: ASE EMT adsorption site search — 5 FCC(111) surfaces, 4-method comparison (Random / GP-UCB / AdsMind-LLM / Hamiltonian-RL), 5 seeds each · Result: supported · Confidence: 0%
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.
This discovery has a Claude-generated validation package with a full experimental design.
Precise Hypothesis
On binary alloy FCC(111) surfaces with compositional (adsorption) energy landscape spread >10 eV as measured by an EMT (or DFT) surrogate over a fixed candidate site enumeration, a Hamiltonian entropy-weighted RL acquisition policy will require ≥15% fewer single-point energy evaluations than an LLM-guided acquisition policy (AdsMind-style) to reach within 0.05 eV of the surrogate global-minimum adsorption energy, averaged over ≥10 independent runs per surface with different random seeds/initial samples. On pure-metal FCC(111) surfaces (Cu, Ag, Au) with well-characterized crystallographic priors, the LLM-guided policy will require 40-160% fewer evaluations than the Hamiltonian RL policy under the same convergence criterion. The current n=5-surface, single-seed-per-surface result (CuAu spread=13.3 eV; calls: Random 33.0, GP-UCB 47.0, AdsMind 45.8, Hamiltonian-RL 35.0) is treated as a preliminary pilot, not a validated effect, pending the replication protocol below.
- If, across ≥5 independent alloy surfaces with spread >10 eV (different composition ratios/alloy pairs, e.g. CuAu, NiPd, AgPt, CuPt, AuPd) and ≥10 seeds each, the mean evaluation-count reduction of Hamiltonian-RL vs AdsMind is <15% (or not statistically significant at p<0.05 via paired t-test / Wilcoxon signed-rank), the hypothesis is disproven.
- If on pure-metal surfaces AdsMind's advantage is <40% or is reversed (Hamiltonian-RL wins) in a majority of replicate seeds, the boundary condition claim is disproven.
- If results are not robust to reordering of the fixed candidate site set or to different random seeds (variance across seeds exceeds the claimed effect size), the effect is attributed to noise/pilot artifact rather than a real mechanism.
- If replacing the EMT surrogate with a DFT or ML-potential-based energy oracle eliminates the effect (reduction drops below 15% or reverses), the claim is disproven at the level of practical relevance (DFT surrogate ≠ generalizable).
Spine & Adversarial ReadReady for validation
“This hypothesis tests whether a Hamiltonian entropy-weighted RL acquisition policy reduces the number of adsorption-site energy evaluations needed to find the global-minimum adsorption site (within 0.05 eV) by at least 15% relative to an LLM-guided (AdsMind-style) acquisition policy on binary alloy FCC(111) surfaces with energy landscape spread greater than 10 eV.”
- highThe pilot evidence is a single run per surface (n=5 surfaces, apparently 1 seed each) with no reported variance — this is nowhere near sufficient to claim a robust 24% effect, and the number could easily be within noise given typical seed-to-seed variance in Bayesian-optimization-style acquisition comparisons.The EVP's protocol directly addresses this by requiring ≥10 seeds per surface and ≥5-6 independent alloy surface families with paired significance testing before any claim is accepted; until that is run, the current 24% figure should be treated as an unvalidated pilot estimate, not a result.
- highWhy EMT as the energy oracle and why AdsMind specifically as the LLM baseline, rather than DFT-ground-truth or other established acquisition baselines (e.g., pure Bayesian optimization with physics-informed kernels, or graph-neural-network-based active learning as used in OC20 baselines)? The methodology choice is not justified beyond convenience/speed.EMT is justified only as a cheap surrogate for rapid iteration (enables exhaustive ground-truth computation for convergence checking, which DFT cannot afford at this scale); this is acknowledged explicitly as a boundary condition, and the protocol mandates a DFT/ML-potential cross-validation subset before any claim of DFT-level relevance. AdsMind is the named comparator in the original discovery and must be reproduced faithfully from its description since no independent public benchmark was found in available search — this is a genuine gap: without confirming AdsMind's published implementation details, reproduction fidelity is uncertain and should be flagged to reviewers explicitly.
- mediumThe mechanism claim (Hamiltonian energy-conservation bias + distance-weighted entropy 'systematically explores high-uncertainty regions LLM priors fail to predict') is asserted but not directly tested — the experiment only measures evaluation counts, not whether the proposed mechanism (entropy-driven exploration of specific regions) is actually what drives the difference.Not resolved in this EVP as written; a mechanism-level test would require logging which specific sites each method queries and comparing spatial/compositional distribution of queries against actual high-error regions of the LLM's implicit prior, which should be added as a follow-up analysis (e.g., correlate Hamiltonian-RL's query distribution with regions where AdsMind's confidence-weighted priors have highest error) rather than inferred solely from aggregate evaluation counts.
Experimental Protocol
Minimum viable test (MVT):
- Fix candidate adsorption site enumeration algorithm (e.g., ASE
add_adsorbate+ symmetry-unique site finder) identically across all methods and surfaces. - Select 6 binary alloy surfaces spanning spread >10 eV (CuAu, NiCu, NiPd, AgPt, CuPt, AuPd — 3x3x4 slabs, random substitutional occupancy at 3 compositions each = 18 configurations) plus the 3 pure metals (Cu, Ag, Au) as controls.
- For each surface/config, run 4 acquisition methods (Random, GP-UCB, AdsMind, Hamiltonian-RL) for 10 independent seeds, budget-capped at 100 evaluations, recording evaluations-to-convergence (within 0.05 eV of exhaustively computed global min).
- Compute per-surface mean/std of evaluations-to-convergence; run paired statistical tests between Hamiltonian-RL and AdsMind.
- Repeat a reduced version (2 alloy surfaces, 5 seeds) using a DFT-level or ML-potential energy oracle to test surrogate-transfer robustness.
- ASE (Atomic Simulation Environment) with EMT calculator for pilot-scale replication.
- Open Catalyst Project (OC20/OC22) pretrained ML potentials (GemNet-OC, EquiformerV2) as a higher-fidelity surrogate oracle.
- Optional DFT validation subset: VASP or Quantum Espresso with PBE functional, standard PAW pseudopotentials, on a small (≤50 configs) confirmation set.
- AdsMind-style LLM acquisition implementation (reproduction needed — no public reference confirmed in available search; must be reimplemented from the discovery's own description/codebase).
- Bulk crystal structures for Cu, Ag, Au, and binary alloy solid-solution generators (e.g., via pymatgen/ASE
SQSor random alloy generation). - Compute logging/tracking (Weights & Biases or MLflow) for reproducibility.
- Primary: Hamiltonian-RL shows ≥15% mean reduction in evaluations-to-convergence vs AdsMind on ≥4 of 6 alloy surface families with spread >10 eV, statistically significant (p<0.05, paired test), consistent in sign across ≥8/10 seeds per surface.
- Secondary: AdsMind shows 40-160% fewer evaluations than Hamiltonian-RL on all 3 pure-metal controls, p<0.05.
- Robustness: Effect direction (sign) unchanged under threshold sensitivity analysis (0.03-0.10 eV) and under EMT→ML-potential oracle substitution (at least directionally consistent, even if magnitude shifts).
- Mean reduction <15% or not significant on majority of alloy surfaces.
- Effect sign flips or is inconsistent (>30% of seeds disagree in direction) on any tested alloy surface.
- Pure-metal advantage for AdsMind fails to reach 40% lower bound.
- Effect disappears or reverses when moving from EMT to DFT/ML-potential oracle.
- Results are sensitive to arbitrary implementation choices (e.g., LLM prompt wording changes result by >50%).
100
GPU hours
30d
Time to result
$1,000
Min cost
$10,000
Full cost
ROI Projection
Adds a validated decision rule ("use LLM-priors for pure/well-characterized surfaces, use uncertainty-driven RL for novel/heterogeneous alloys") to materials-discovery acquisition pipelines used by catalysis screening groups, national labs, and battery/fuel-cell/CO2-reduction catalyst startups; could be packaged as an acquisition-policy-selection module in active-learning-for-DFT toolkits (e.g., integrated into AMPTorch, OCP, or Meta's fairchem stack).
TIME_TO_RESULT_DAYS: 21
Implementation Sketch
for surface in alloy_surfaces + pure_metal_controls: sites = enumerate_symmetry_unique_sites(surface) ground_truth = {s: emt_energy(surface, s) for s in sites} # exhaustive, one-time spread = max(ground_truth.values()) - min(ground_truth.values()) for method in [Random, GP_UCB, AdsMind, Hamiltonian_RL]: for seed in range(10): history = [] for step in range(budget_cap=100): site = method.propose_next_site(history, surface_features, seed) e = ground_truth[site] # simulated oracle lookup history.append((site, e)) if abs(min(e for _,e in history) - min(ground_truth.values())) < 0.05: log(method, surface, seed, step+1, converged=True) break else: log(method, surface, seed, budget_cap, converged=False) analyze: paired_test(HamiltonianRL_calls, AdsMind_calls, groupby=surface_type)
Hamiltonian-RL policy internals: reward = -distance_weighted_entropy(candidate) + energy_conservation_penalty(trajectory); trained via policy gradient or bandit-style update per episode across surfaces (meta-learned prior) before test-time deployment.
- After generating exhaustive ground-truth energies (step 4): if spread values don't reproduce claimed magnitude (13.3 eV for CuAu) within reasonable tolerance, abort and re-audit site enumeration/energy pipeline before running acquisition methods.
- After 3-seed pilot per method/surface (~10% of full compute budget): if effect direction (Hamiltonian-RL < AdsMind on alloys) is not even directionally present, abort before scaling to full 10-seed x 6-surface run.
- After alloy-surface runs complete but before DFT cross-validation: if statistical significance is not reached at n=10 seeds, do not proceed to costly DFT confirmation subset.
NAMED_EXPERTS: []
CLOSEST_EXISTING_WORK: []
NOVELTY_NARROWING_REQUIRED: false