Integrating Satisfiability Modulo Theories (SMT) solvers with large language models to enforce strict physical and chemical stoichiometry constraints reduces hallucinated compositions in generative inorganic crystal discovery.
Adversarial Debate Score
75% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Related patents (prior art)
This hypothesis overlaps subject matter covered by existing third-party patents. It is published as research, not as a patentable claim of ours.
Supporting Research Papers
- General-purpose LLMs as Constrained Crystal Composition Generators
The targeted discovery of inorganic materials remains challenging due to the vastness of compositional design spaces and the high cost of exhaustive screening. Task-specific generative artificial inte...
- Molecular Geometry Understanding Has Unintendedly Emerged in Frontier Large Language Models
Large language models (LLMs) have already shown strong capabilities in solving complex chemical problems expressed in natural language. Yet, many tasks performed by chemists require understanding of 3...
- Formalize, Don't Optimize: The Heuristic Trap in LLM-Generated Combinatorial Solvers
Large Language Models (LLMs) struggle to solve complex combinatorial problems through direct reasoning, so recent neuro-symbolic systems increasingly use them to synthesize executable solvers. A centr...
- Neuro-Symbolic Compliance: Integrating LLMS and SMT Solvers for Automated Financial Legal Analysis
Financial regulations are increasingly complex, hindering automated compliance-especially the maintenance of logical consistency with minimal human oversight. We introduce a Neuro-Symbolic Compliance ...
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.
This discovery has a Claude-generated validation package with a full experimental design.
Precise Hypothesis
Constraining an LLM's generative sampling process for inorganic crystal compositions with an SMT solver (encoding charge-neutrality, oxidation-state feasibility, stoichiometric ratio validity, and ionic radius/Pauling-rule constraints) reduces the rate of chemically invalid ("hallucinated") compositions from a measured baseline (unconstrained LLM) by a statistically significant margin, without reducing the rate of novel, synthesizable, property-relevant compositions among the valid outputs. Specifically: for a fixed LLM (e.g., a fine-tuned GPT-family or Llama-family model prompted/fine-tuned for crystal composition generation), the SMT-constrained pipeline will achieve ≥90% chemical validity (charge-balance + oxidation-state consistency, verified against ICSD/Materials Project ground truth and DFT-checked subsample) versus a baseline unconstrained validity rate of ≤60%, while retaining ≥70% of the unconstrained model's diversity (measured via unique composition count and Shannon entropy over element combinations) and ≥50% overlap with DFT-confirmed formation-energy-negative (stable/metastable, Ehull ≤ 50 meV/atom) structures relative to unconstrained sampling.
- SMT-constrained pipeline validity rate is not statistically significantly better than unconstrained baseline (p≥0.05, two-proportion z-test, n≥500 generated compositions per arm).
- SMT constraints reduce novel/diverse valid compositions by >50% relative to baseline (i.e., solver over-constrains and merely regurgitates known compounds), measured by unique composition count and Shannon diversity index.
- DFT-verified stability rate (Ehull ≤ 50 meV/atom) of SMT-passed compositions is not significantly higher than DFT-verified stability rate of a size-matched random valid subsample from baseline.
- Solver integration adds >10x wall-clock latency making the pipeline impractical (>10s/candidate at scale) without commensurate accuracy gain.
- Effect fails to replicate across ≥2 independent LLM backbones (e.g., GPT-4-class vs. Llama-3-class) or ≥2 independent held-out test sets (ICSD held-out split vs. OQMD).
Spine & Adversarial ReadReady for validation
“This hypothesis tests whether integrating an SMT solver as a hard-constraint validator/guide within an LLM's crystal-composition generation loop significantly increases the rate of chemically valid (charge-balanced, oxidation-state-consistent) inorganic compositions relative to unconstrained LLM generation, without collapsing compositional diversity or novelty.”
- highCharge neutrality and oxidation-state matching are necessary but nowhere near sufficient conditions for a real, synthesizable crystal — many charge-balanced formulas are thermodynamically unstable or structurally impossible, so 'SMT-valid' is a weak proxy for 'not hallucinated.' The success criteria risk optimizing a metric (rule-based validity) that is easy to satisfy but doesn't guarantee the actual scientific goal (real, useful materials).Partially addressed by requiring DFT/surrogate Ehull confirmation as a secondary success criterion (≥1.5x stability enrichment), but the EVP does not require full structural validation (space group, coordination geometry) — this is an acknowledged gap; a stronger version would add a structure-prediction step (e.g., CDVAE- or AIRSS-based structure search) on SMT-valid compositions before claiming 'non-hallucination.'
- highWhy SMT/Z3 specifically rather than simpler and cheaper alternatives — e.g., a rule-based Python validator, a differentiable soft-constraint loss during fine-tuning, or grammar-constrained decoding with a context-free grammar encoding valid stoichiometries? The methodology does not justify why the added complexity and potential latency cost of a full SMT solver is necessary versus these lighter-weight baselines.This is a genuine methodology justification gap. The EVP should be revised to include a mandatory baseline arm using a simple rule-based/regex validator with identical oxidation-state tables (no solver) to isolate what marginal benefit the SMT formalism (as opposed to just 'having any explicit validity check') actually provides. Without this ablation, a skeptical reviewer could argue the result demonstrates 'constraint-checking helps' rather than 'SMT-solving specifically helps.'
- mediumThe train/test split strategy and novelty metric are vulnerable to LLM memorization: modern LLMs (especially GPT-4-class) may have seen Materials Project and ICSD data during pretraining, making 'novel valid composition' claims unverifiable — the model could be regurgitating memorized compositions that merely pass the SMT filter, not generating genuinely new chemistry.Not resolved in current protocol. Mitigation would require using a post-training-cutoff held-out set (compositions/structures published or deposited after the LLM's known training cutoff date) and/or restricting evaluation to open, license-clear checkpoints with documented training corpora (favoring open Llama-3 over closed GPT-4-class API for this reason) — this should be added as an explicit protocol amendment before the study is considered publication-ready.
Experimental Protocol
Minimum viable test (MVT):
- Fine-tune or prompt-engineer one open LLM (Llama-3-8B or similar) on a composition-generation task using Materials Project (MP) training split (~80%).
- Generate n=1,000 candidate compositions unconstrained (baseline arm) and n=1,000 with SMT-solver-in-the-loop constrained decoding (Z3 or CVC5 encoding charge balance + oxidation state tables) (treatment arm), using identical seeds/prompts.
- Score validity: (a) rule-based — charge neutrality + oxidation state lookup against Materials Project ICSD; (b) DFT-based — random subsample of n=100 per arm relaxed via VASP/Quantum ESPRESSO or fast surrogate (M3GNet/CHGNet) to compute Ehull.
- Compare validity %, diversity metrics, and DFT-confirmed stability rate between arms with pre-registered statistical tests.
- Replicate on second LLM backbone and second held-out dataset (OQMD) to test generalization.
- Materials Project (MP) full snapshot (~154,000 structures, formation energies, oxidation states) — primary training/validation set.
- ICSD (Inorganic Crystal Structure Database) subset for held-out ground-truth validity checks (institutional license required).
- OQMD (Open Quantum Materials Database, ~1M+ DFT-computed entries) — secondary held-out generalization test.
- Oxidation state reference tables (Pymatgen
icsd_oxidation_states, Bartel et al. 2020 oxidation-state statistics). - Pauling electronegativity / ionic radius tables (Shannon radii).
- Pre-trained LLM checkpoints: Llama-3-8B-Instruct, GPT-4-class API access (or open equivalent), optionally CrystalLLM / MatBERT-style domain-adapted baselines if available.
- Fast surrogate DFT: M3GNet or CHGNet (pre-trained universal potentials) for large-scale Ehull screening; true DFT (VASP/QE) for confirmatory subsample.
- SMT solver: Z3 (Microsoft) or CVC5, with Pymatgen for chemistry-rule encoding.
- Primary: SMT-constrained validity rate ≥90% vs. baseline ≤60%, difference significant at p<0.01 (two-proportion z-test, n≥1,000/arm).
- Secondary: DFT-confirmed stability rate (Ehull≤50 meV/atom) among SMT-valid compositions ≥1.5x the rate in a matched random valid baseline subsample.
- Diversity retention: unique composition count in treatment arm ≥70% of baseline arm's unique count; Shannon entropy difference <15%.
- Generalization: effect replicates (validity gain ≥20 percentage points, p<0.05) across ≥2 LLM backbones and ≥2 datasets.
- Practicality: mean solver latency per candidate <1s; total pipeline throughput ≥500 compositions/hour on single GPU+CPU node.
- Validity gain <10 percentage points or not statistically significant (p≥0.05).
- Diversity collapse: unique composition count drops >50% relative to baseline, indicating solver forces trivial/repetitive outputs.
- No DFT stability enrichment: SMT-valid compositions show equivalent or worse Ehull distribution vs. random valid baseline.
- Failure to generalize: effect present in only 1 of 2 backbone/dataset combinations.
- Solver intractability: >20% of candidates cause SMT timeout (>5s) or solver non-termination at realistic constraint complexity.
480
GPU hours
75d
Time to result
$18,000
Min cost
$95,000
Full cost
ROI Projection
High cross-sector applicability: battery materials (solid electrolytes, cathode/anode chemistries), semiconductor and photovoltaic materials discovery, catalysis (single-atom/alloy catalysts), and thermoelectric materials all rely on generative composition screening pipelines vulnerable to hallucination waste. A validated SMT-LLM constraint framework is licensable as a plug-in module for existing materials-discovery platforms (e.g., Citrine Informatics, Materials Project workflows, battery/semiconductor R&D at national labs and industry: Toyota Research Institute, Samsung SDI, Panasonic, IBM Research). The neurosymbolic pattern (SMT + LLM for domain-constrained generation) is also transferable to adjacent domains (drug-like molecule generation with valence/synthon constraints, protein sequence design with biophysical constraints), broadening commercial reach beyond materials science into pharma/biotech tooling.
🔓 If proven, this unlocks
Proving this hypothesis is a prerequisite for the following downstream discoveries and applications:
- 1property-conditioned-crystal-generation-with-SMT-guardrails
- 2SMT-LLM-hybrid-pipelines-for-organic-synthesis-planning
- 3automated-DFT-triage-pipeline-for-generative-materials-discovery
- 4cross-domain-symbolic-constraint-LLM-frameworks-drug-design
Implementation Sketch
# Pseudocode: SMT-Constrained Crystal Composition Generation def build_smt_constraints(elements, oxidation_states_db): solver = z3.Solver() counts = {e: z3.Int(f"count_{e}") for e in elements} ox_states = {e: z3.Int(f"ox_{e}") for e in elements} for e in elements: solver.add(counts[e] >= 0) solver.add(z3.Or([ox_states[e] == v for v in oxidation_states_db[e]])) # charge neutrality solver.add(z3.Sum([counts[e] * ox_states[e] for e in elements]) == 0) # stoichiometry bounds, radius-ratio rules appended similarly return solver, counts, ox_states def llm_generate_candidate(prompt, model, temperature=0.8): return model.generate(prompt, temperature=temperature) # returns composition string def constrained_generation_loop(model, prompt, max_retries=5): for attempt in range(max_retries): candidate = llm_generate_candidate(prompt, model) elements, ratios = parse_composition(candidate) solver, counts, ox_states = build_smt_constraints(elements, OXIDATION_DB) # fix candidate ratios as additional constraints, check satisfiability solver.push() for e in elements: solver.add(counts[e] == ratios[e]) if solver.check() == z3.sat: return candidate, "valid" solver.pop() # optional: ask solver for nearest satisfiable assignment (MaxSMT/soft constraints) prompt = augment_prompt_with_feedback(prompt, candidate, reason="charge_imbalance") return None, "rejected_after_max_retries" def run_pipeline(n_candidates, model, dataset_split): results = [] for i in range(n_candidates): candidate, status = constrained_generation_loop(model, base_prompt(dataset_split)) results.append((candidate, status)) return results def evaluate(results, ground_truth_db, dft_surrogate): validity_rate = fraction_valid(results, ground_truth_db) subsample = random.sample([r for r in results if r[1]=="valid"], 100) ehull_scores = [dft_surrogate.relax_and_score(c) for c, _ in subsample] diversity = shannon_entropy(unique_element_sets(results)) return validity_rate, ehull_scores, diversity
- Day 15 (after constraint schema + solver built): if solver fails to encode >5 canonical known-valid compounds correctly (false negatives on ground truth), halt and revise oxidation-state/radius-ratio schema before proceeding.
- Day 30 (after baseline + treatment generation on backbone 1): if validity gain <10 percentage points, halt full-scale run and diagnose (schema too permissive? LLM already highly valid at baseline?) before investing in DFT confirmation and second-backbone replication.
- Day 45 (after DFT/surrogate subsample scoring): if no stability enrichment signal (Ehull distributions statistically indistinguishable, p≥0.2), halt before committing to true-DFT confirmatory runs (most expensive compute step).
- Day 60 (after first replication attempt): if effect fails to replicate on second backbone/dataset, halt and downgrade claim to "backbone-specific" rather than general finding; do not proceed to publication-scale write-up without explicit scope narrowing.