solver.press

Integrating Satisfiability Modulo Theories (SMT) solvers with large language models to enforce strict physical and chemical stoichiometry constraints reduces hallucinated compositions in generative inorganic crystal discovery.

Computer ScienceSep 28, 2026Evaluation Score: 81%

Adversarial Debate Score

75% survival rate under critique

Expert panel critique

Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.

Gemini: Strengths: The hypothesis is highly falsifiable and logically sound, directly supported by literature demonstrating that LLMs struggle with direct combinatorial reasoning but excel when synthesizing constraints for symbolic solvers (e.g., SMT). The owner's validated experiments do not con...
Mistral: The hypothesis is well-motivated, falsifiable, and aligns with validated neuro-symbolic integration trends (e.g., SMT+LLM in financial compliance), but lacks direct experimental support from the owner’s own work and faces plausible counterarguments (e.g., LLMs’ emergent 3D reasoning may already r...
ChatGPT: The hypothesis is clear and falsifiable, and related work supports constrained generation and LLM–SMT integration, but the cited evidence is mostly indirect and the owner’s validated experiments do not test crystal composition generation. SMT can eliminate encoded stoichiometric violations, yet c...
Claude: The hypothesis is logically coherent, falsifiable (hallucination rate is measurable), and supported by the neuro-symbolic SMT-LLM integration paper and the constrained crystal composition work, with plausible extension from the financial compliance framework; however, none of the owner's vali...

Related patents (prior art)

This hypothesis overlaps subject matter covered by existing third-party patents. It is published as research, not as a patentable claim of ours.

Supporting Research Papers

Formal Verification

Z3 logical consistency:✅ Consistent

Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.

Experimental Validation Package

This discovery has a Claude-generated validation package with a full experimental design.

Precise Hypothesis

Constraining an LLM's generative sampling process for inorganic crystal compositions with an SMT solver (encoding charge-neutrality, oxidation-state feasibility, stoichiometric ratio validity, and ionic radius/Pauling-rule constraints) reduces the rate of chemically invalid ("hallucinated") compositions from a measured baseline (unconstrained LLM) by a statistically significant margin, without reducing the rate of novel, synthesizable, property-relevant compositions among the valid outputs. Specifically: for a fixed LLM (e.g., a fine-tuned GPT-family or Llama-family model prompted/fine-tuned for crystal composition generation), the SMT-constrained pipeline will achieve ≥90% chemical validity (charge-balance + oxidation-state consistency, verified against ICSD/Materials Project ground truth and DFT-checked subsample) versus a baseline unconstrained validity rate of ≤60%, while retaining ≥70% of the unconstrained model's diversity (measured via unique composition count and Shannon entropy over element combinations) and ≥50% overlap with DFT-confirmed formation-energy-negative (stable/metastable, Ehull ≤ 50 meV/atom) structures relative to unconstrained sampling.

Disproof criteria:
  • SMT-constrained pipeline validity rate is not statistically significantly better than unconstrained baseline (p≥0.05, two-proportion z-test, n≥500 generated compositions per arm).
  • SMT constraints reduce novel/diverse valid compositions by >50% relative to baseline (i.e., solver over-constrains and merely regurgitates known compounds), measured by unique composition count and Shannon diversity index.
  • DFT-verified stability rate (Ehull ≤ 50 meV/atom) of SMT-passed compositions is not significantly higher than DFT-verified stability rate of a size-matched random valid subsample from baseline.
  • Solver integration adds >10x wall-clock latency making the pipeline impractical (>10s/candidate at scale) without commensurate accuracy gain.
  • Effect fails to replicate across ≥2 independent LLM backbones (e.g., GPT-4-class vs. Llama-3-class) or ≥2 independent held-out test sets (ICSD held-out split vs. OQMD).

Spine & Adversarial ReadReady for validation

“This hypothesis tests whether integrating an SMT solver as a hard-constraint validator/guide within an LLM's crystal-composition generation loop significantly increases the rate of chemically valid (charge-balanced, oxidation-state-consistent) inorganic compositions relative to unconstrained LLM generation, without collapsing compositional diversity or novelty.”

  • highCharge neutrality and oxidation-state matching are necessary but nowhere near sufficient conditions for a real, synthesizable crystal — many charge-balanced formulas are thermodynamically unstable or structurally impossible, so 'SMT-valid' is a weak proxy for 'not hallucinated.' The success criteria risk optimizing a metric (rule-based validity) that is easy to satisfy but doesn't guarantee the actual scientific goal (real, useful materials).
    Partially addressed by requiring DFT/surrogate Ehull confirmation as a secondary success criterion (≥1.5x stability enrichment), but the EVP does not require full structural validation (space group, coordination geometry) — this is an acknowledged gap; a stronger version would add a structure-prediction step (e.g., CDVAE- or AIRSS-based structure search) on SMT-valid compositions before claiming 'non-hallucination.'
  • highWhy SMT/Z3 specifically rather than simpler and cheaper alternatives — e.g., a rule-based Python validator, a differentiable soft-constraint loss during fine-tuning, or grammar-constrained decoding with a context-free grammar encoding valid stoichiometries? The methodology does not justify why the added complexity and potential latency cost of a full SMT solver is necessary versus these lighter-weight baselines.
    This is a genuine methodology justification gap. The EVP should be revised to include a mandatory baseline arm using a simple rule-based/regex validator with identical oxidation-state tables (no solver) to isolate what marginal benefit the SMT formalism (as opposed to just 'having any explicit validity check') actually provides. Without this ablation, a skeptical reviewer could argue the result demonstrates 'constraint-checking helps' rather than 'SMT-solving specifically helps.'
  • mediumThe train/test split strategy and novelty metric are vulnerable to LLM memorization: modern LLMs (especially GPT-4-class) may have seen Materials Project and ICSD data during pretraining, making 'novel valid composition' claims unverifiable — the model could be regurgitating memorized compositions that merely pass the SMT filter, not generating genuinely new chemistry.
    Not resolved in current protocol. Mitigation would require using a post-training-cutoff held-out set (compositions/structures published or deposited after the LLM's known training cutoff date) and/or restricting evaluation to open, license-clear checkpoints with documented training corpora (favoring open Llama-3 over closed GPT-4-class API for this reason) — this should be added as an explicit protocol amendment before the study is considered publication-ready.

Experimental Protocol

Minimum viable test (MVT):

  1. Fine-tune or prompt-engineer one open LLM (Llama-3-8B or similar) on a composition-generation task using Materials Project (MP) training split (~80%).
  2. Generate n=1,000 candidate compositions unconstrained (baseline arm) and n=1,000 with SMT-solver-in-the-loop constrained decoding (Z3 or CVC5 encoding charge balance + oxidation state tables) (treatment arm), using identical seeds/prompts.
  3. Score validity: (a) rule-based — charge neutrality + oxidation state lookup against Materials Project ICSD; (b) DFT-based — random subsample of n=100 per arm relaxed via VASP/Quantum ESPRESSO or fast surrogate (M3GNet/CHGNet) to compute Ehull.
  4. Compare validity %, diversity metrics, and DFT-confirmed stability rate between arms with pre-registered statistical tests.
  5. Replicate on second LLM backbone and second held-out dataset (OQMD) to test generalization.
Required datasets:
  • Materials Project (MP) full snapshot (~154,000 structures, formation energies, oxidation states) — primary training/validation set.
  • ICSD (Inorganic Crystal Structure Database) subset for held-out ground-truth validity checks (institutional license required).
  • OQMD (Open Quantum Materials Database, ~1M+ DFT-computed entries) — secondary held-out generalization test.
  • Oxidation state reference tables (Pymatgen icsd_oxidation_states, Bartel et al. 2020 oxidation-state statistics).
  • Pauling electronegativity / ionic radius tables (Shannon radii).
  • Pre-trained LLM checkpoints: Llama-3-8B-Instruct, GPT-4-class API access (or open equivalent), optionally CrystalLLM / MatBERT-style domain-adapted baselines if available.
  • Fast surrogate DFT: M3GNet or CHGNet (pre-trained universal potentials) for large-scale Ehull screening; true DFT (VASP/QE) for confirmatory subsample.
  • SMT solver: Z3 (Microsoft) or CVC5, with Pymatgen for chemistry-rule encoding.
Success:
  • Primary: SMT-constrained validity rate ≥90% vs. baseline ≤60%, difference significant at p<0.01 (two-proportion z-test, n≥1,000/arm).
  • Secondary: DFT-confirmed stability rate (Ehull≤50 meV/atom) among SMT-valid compositions ≥1.5x the rate in a matched random valid baseline subsample.
  • Diversity retention: unique composition count in treatment arm ≥70% of baseline arm's unique count; Shannon entropy difference <15%.
  • Generalization: effect replicates (validity gain ≥20 percentage points, p<0.05) across ≥2 LLM backbones and ≥2 datasets.
  • Practicality: mean solver latency per candidate <1s; total pipeline throughput ≥500 compositions/hour on single GPU+CPU node.
Failure:
  • Validity gain <10 percentage points or not statistically significant (p≥0.05).
  • Diversity collapse: unique composition count drops >50% relative to baseline, indicating solver forces trivial/repetitive outputs.
  • No DFT stability enrichment: SMT-valid compositions show equivalent or worse Ehull distribution vs. random valid baseline.
  • Failure to generalize: effect present in only 1 of 2 backbone/dataset combinations.
  • Solver intractability: >20% of candidates cause SMT timeout (>5s) or solver non-termination at realistic constraint complexity.

480

GPU hours

75d

Time to result

$18,000

Min cost

$95,000

Full cost

ROI Projection

Commercial:

High cross-sector applicability: battery materials (solid electrolytes, cathode/anode chemistries), semiconductor and photovoltaic materials discovery, catalysis (single-atom/alloy catalysts), and thermoelectric materials all rely on generative composition screening pipelines vulnerable to hallucination waste. A validated SMT-LLM constraint framework is licensable as a plug-in module for existing materials-discovery platforms (e.g., Citrine Informatics, Materials Project workflows, battery/semiconductor R&D at national labs and industry: Toyota Research Institute, Samsung SDI, Panasonic, IBM Research). The neurosymbolic pattern (SMT + LLM for domain-constrained generation) is also transferable to adjacent domains (drug-like molecule generation with valence/synthon constraints, protein sequence design with biophysical constraints), broadening commercial reach beyond materials science into pharma/biotech tooling.

🔓 If proven, this unlocks

Proving this hypothesis is a prerequisite for the following downstream discoveries and applications:

  • 1property-conditioned-crystal-generation-with-SMT-guardrails
  • 2SMT-LLM-hybrid-pipelines-for-organic-synthesis-planning
  • 3automated-DFT-triage-pipeline-for-generative-materials-discovery
  • 4cross-domain-symbolic-constraint-LLM-frameworks-drug-design

Implementation Sketch

# Pseudocode: SMT-Constrained Crystal Composition Generation

def build_smt_constraints(elements, oxidation_states_db):
    solver = z3.Solver()
    counts = {e: z3.Int(f"count_{e}") for e in elements}
    ox_states = {e: z3.Int(f"ox_{e}") for e in elements}
    for e in elements:
        solver.add(counts[e] >= 0)
        solver.add(z3.Or([ox_states[e] == v for v in oxidation_states_db[e]]))
    # charge neutrality
    solver.add(z3.Sum([counts[e] * ox_states[e] for e in elements]) == 0)
    # stoichiometry bounds, radius-ratio rules appended similarly
    return solver, counts, ox_states

def llm_generate_candidate(prompt, model, temperature=0.8):
    return model.generate(prompt, temperature=temperature)  # returns composition string

def constrained_generation_loop(model, prompt, max_retries=5):
    for attempt in range(max_retries):
        candidate = llm_generate_candidate(prompt, model)
        elements, ratios = parse_composition(candidate)
        solver, counts, ox_states = build_smt_constraints(elements, OXIDATION_DB)
        # fix candidate ratios as additional constraints, check satisfiability
        solver.push()
        for e in elements:
            solver.add(counts[e] == ratios[e])
        if solver.check() == z3.sat:
            return candidate, "valid"
        solver.pop()
        # optional: ask solver for nearest satisfiable assignment (MaxSMT/soft constraints)
        prompt = augment_prompt_with_feedback(prompt, candidate, reason="charge_imbalance")
    return None, "rejected_after_max_retries"

def run_pipeline(n_candidates, model, dataset_split):
    results = []
    for i in range(n_candidates):
        candidate, status = constrained_generation_loop(model, base_prompt(dataset_split))
        results.append((candidate, status))
    return results

def evaluate(results, ground_truth_db, dft_surrogate):
    validity_rate = fraction_valid(results, ground_truth_db)
    subsample = random.sample([r for r in results if r[1]=="valid"], 100)
    ehull_scores = [dft_surrogate.relax_and_score(c) for c, _ in subsample]
    diversity = shannon_entropy(unique_element_sets(results))
    return validity_rate, ehull_scores, diversity
Abort checkpoints:
  • Day 15 (after constraint schema + solver built): if solver fails to encode >5 canonical known-valid compounds correctly (false negatives on ground truth), halt and revise oxidation-state/radius-ratio schema before proceeding.
  • Day 30 (after baseline + treatment generation on backbone 1): if validity gain <10 percentage points, halt full-scale run and diagnose (schema too permissive? LLM already highly valid at baseline?) before investing in DFT confirmation and second-backbone replication.
  • Day 45 (after DFT/surrogate subsample scoring): if no stability enrichment signal (Ehull distributions statistically indistinguishable, p≥0.2), halt before committing to true-DFT confirmatory runs (most expensive compute step).
  • Day 60 (after first replication attempt): if effect fails to replicate on second backbone/dataset, halt and downgrade claim to "backbone-specific" rather than general finding; do not proceed to publication-scale write-up without explicit scope narrowing.

Source

AegisMind Research
Need AI to work rigorously on your problems? AegisMind uses the same multi-model engine for personal and professional use. Get started