solver.press

Autonomous Scientific Discovery

From hypothesis to publication — autonomously

solver.press is the public window into AegisMind's discovery engine. Hypotheses are generated, stress-tested through five-model adversarial debate, checked for logical consistency, and turned into experimental packages designed to be cheap to kill.

Everything here is a research lead. No hypothesis on this site has been validated in a wet lab. Where computation is itself the experiment — ML theory, quantum systems, mathematics — the numerical results stand on their own. In biology they do not, and we label them accordingly.

✦ Cross-domain hypothesis generation across 15+ scientific domains✦ Five-model adversarial debate, scored on survival under critique✦ Z3 logical-consistency check — internal coherence, not empirical truth✦ Molecular docking where useful — reported as screening, never as validation✦ Experimental Validation Packages with disproof criteria, abort checkpoints, and costs✦ Preprints self-deposited to Zenodo and Research Square — citable, not peer-reviewed

Recently computed

Discoveries that have had a computation run against them, most recent first. Where that computation is itself the experiment it is marked verified; where it is a structure-prediction score against a protein target it is marked as docking only, which is a screening result and not evidence of binding.

View all →
Computer ScienceJul 26, 2026Evaluation Score: 70%

Hamiltonian entropy-weighted reinforcement learning (RL) acquisition reduces adsorption site evaluation count by approxi…

EVP available🧪 Numerically verified

Source: AegisMind Research

Read full discovery
Biology + MedicineJul 6, 2026Evaluation Score: 56%

Novel diazabicyclooctane (DBO) beta-lactamase inhibitor analogues, designed via multi-feature virtual screening, inhibit…

EVP available🔬 Docking / simulation performed

Source: AegisMind Research

Read full discovery
Computer Science + BiologyJun 14, 2026Evaluation Score: 69%

Quantum annealing-based subgraph isomorphism algorithms can identify structural motifs in protein-ligand docking data th…

EVP available🔬 Docking / simulation performed

Source: AegisMind Research

Read full discovery

Highest-confidence discovery

74% evaluation scoreEVP available

Efflux pump inhibitors targeting MexAB-OprM restore carbapenem susceptibility in MDR Pseudomonas aeruginosa by reducing intracellular drug efflux, converting clinical resistance (MIC >8 µg/mL) to susceptibility (MIC ≤2 µg/mL)

Medicine · Efflux pump inhibitors targeting MexAB-OprM restore carbapenem susceptibility in MDR Pseudomonas aeruginosa by reducing …

Read full discovery →

Recent discoveries

Running since March 2026. These are research leads that survived adversarial debate — not findings, and not confirmed.

View all →
Computer Science + BiologyJun 14, 2026Evaluation Score: 69%

Quantum annealing-based subgraph isomorphism algorithms can identify structural motifs in protein-ligand docking data th…

EVP available🔬 Docking / simulation performed

Source: AegisMind Research

Read full discovery
Computer ScienceJul 26, 2026Evaluation Score: 78%

Neuro-symbolic runtime monitors for drug discovery LLM agents (Mozi framework) will reduce unconstrained tool-use violat…

EVP available

Source: AegisMind Research

Read full discovery
PhysicsJul 26, 2026Evaluation Score: 77%

Chemical short-range order (CSRO) in Co-Ni-V alloys will modulate the adsorption energy distributions (AEDs) of CO2 and …

EVP available📖 Literature-assessed (LLM)

Source: AegisMind Research

Read full discovery
PhysicsJul 26, 2026Evaluation Score: 77%

A solar-pumped solid-state laser (Nd:YAG directly pumped by concentrated solar flux ~1000 suns) feeding a periodically-p…

EVP available📖 Literature-assessed (LLM)

Source: AegisMind Research

Read full discovery
OtherJul 26, 2026Evaluation Score: 73%

Polychromatic optical scattering and nonlinear chromatic mixing in passive optical media can serve as a physical reservo…

EVP available

Source: AegisMind Research

Read full discovery
Computer ScienceJul 26, 2026Evaluation Score: 71%

Applying formal neuro-symbolic self-auditing mechanisms, as used in LLM-based compliance frameworks, to reasoning LLMs-a…

EVP available📖 Literature-assessed (LLM)

Source: AegisMind Research

Read full discovery

The funnel, honestly

Volume is the easy part. Confirmation is the hard part.

2,113

hypotheses generated

Since March 2026. Counts as at 3 August 2026.

1,292

published as leads

Survived debate scoring. Published means visible here — not verified. 26 records were withdrawn on 7 August 2026: a parsing bug had stored fragments of a model's reply to a prompt as hypotheses. They were never discoveries.

1

where computation settles it

Of 12 with any computation run, 11 are structure predictions on biological targets. One — a reinforcement-learning result benchmarked against an oracle — is a domain where the computation is the experiment.

0

confirmed in a wet lab

No hypothesis on this site has independent experimental confirmation.

We show this because the ratio is the most honest thing we can tell you about the system. Generating plausible, well-argued, internally consistent hypotheses at volume is demonstrably easy. Establishing that any one of them is true is not, and we have not yet done it for a single biological claim.

Track record

Our own blind benchmark of this pipeline returned a null result

Before running any predictions we pre-registered a retrospective virtual-screening benchmark and froze the design in git (commit 92c6f8bf, 14 July 2026) so it could not be tuned after seeing results. We then tested 14 targets across three panels over 26 amendments. In fairness: that commit is in a private repository, so you cannot currently verify the timestamp yourself — treat the pre-registration as a claim until we publish the repository. The pre-registration, panel freeze and amendment log are available on request. And a second caveat against ourselves: we later found (amendment 19) that the run below prepared every ligand rigidly — a silent library failure meant no conformational search was performed — so these numbers are not a fair test of docking. We report them because they are what we actually ran. The corrected re-run was stopped when the one target that had qualified turned out not to be a target at all (amendment 21).

0.00

EF@1% for our platform on the only target that could be scored (SARS-CoV-2 Mpro). Zero known actives recovered in the top 1%. Mpro passed the debias gate only under active-set subsampling: re-gated at full active count it fails too (AUROC 0.665), so on the benchmark’s own criterion no target qualified.

0.00

EF@1% for baseline AutoDock Vina on the same target. We did not beat the baseline, and the baseline did not work either.

23.1

EF@1% for a plain 2D fingerprint-similarity baseline — so the actives really were distinguishable from the decoys. Whatever failed, it was not the dataset.

Platform AUROC was 0.617 against Vina's 0.566, both near chance. Only one target cleared the gate, so the pre-registered panel test could not be run at all.

We publish this because it is the most informative thing we know about our own pipeline. It is the reason no docking score on this site is presented as validation, and the reason every biological claim here is labelled a prediction. Treat docking numbers as a way of ordering what to test first — not as evidence that a compound binds.

How discoveries are made

1

Ingest

arXiv and Semantic Scholar papers across 15+ scientific domains are continuously embedded into a vector store and searched for connections that span fields. In practice the engine usually lands inside established, well-populated literatures rather than on untouched ground — efflux-pump inhibition, β-lactamase inhibitor analogues and cathepsin inhibitors are all crowded fields with decades of work behind them. The useful output is a specific, narrow angle within a known field, not a bridge nobody has crossed.

2

Generate & Debate

Five frontier AI models (Claude, GPT, Gemini, Grok, Mistral) independently generate hypotheses then critique each other in adversarial debate. A debate score reflects survival under critique. This measures how well a claim withstands argument — which is not the same as being true, and several hypotheses that scored well here were later refuted.

3

Logical consistency check

Z3 checks whether a hypothesis contradicts itself — not empirical truth, but internal coherence. About 62% of hypotheses pass. A pass means the claim is not self-contradictory; it says nothing about whether the claim is true, and we do not count it as evidence.

4

Experimental Validation Package

Surviving discoveries receive a full EVP built around the cheapest experiment that could kill the hypothesis: precise quantitative claim, explicit disproof criteria, protocol, abort checkpoints, implementation code, compute and cost estimates, and a dependency map.

5

Where the loop actually closes

For ML, quantum, and mathematical claims, we run the computation on Google TPU Research Cloud and the result settles the question — the loop closes. For biological claims it does not: computation produces a prediction, and only a wet lab can close it. Those EVPs are handed off, not resolved here.

Hypothesis Aggregation Papers

View all →

Formal papers synthesising solver.press discoveries into testable hypothesis clusters with complete experimental validation packages.

Prediction — awaiting wet lab

mHTT Condensate Disruption Therapy in Huntington's Disease

H₁ computationally confirmed: Flory-Huggins phase diagram predicts C* ≈ 3.5 µM for Q46 — physiologically accessible in HD striatal neurons — and TF partition coefficients explain the −44.7% target gene expression deficit across two independent genome-wide datasets. H₂ (BET bromodomain inhibitors dissolve mHTT condensates, restoring TF availability) awaits wet-lab validation. Multi-phase EVP ready. Patent AU2026905785.

John Goodman — OceanSparx Pty Ltd, June 2026Full paper →Aggregated EVP →
Submitted for peer review

Collateral Sensitivity Strongly Connected Components in Clinical Surveillance Data

The hypothesis was that strongly connected components in the collateral sensitivity graph define closed evolutionary traps. Tested against 104,337 susceptibility records from BV-BRC — 18,821 clinical isolates across four WHO critical-priority pathogens. Two of four species yielded a qualifying SCC: K. pneumoniae {imipenem, meropenem, tetracycline} (permutation p = 0.001) and E. coli {colistin, cefotaxime} (OR = 10.13, n = 87). S. aureus yielded none. The carbapenem signal is tetracycline-specific and absent for tigecycline, which distinguishes it from a clonal-lineage artifact, and holds across independent 2009–2014 year bands. This is a retrospective association in surveillance data, not a demonstration of causation; isogenic experimental follow-up is the next step.

John Goodman — OceanSparx Pty Ltd, August 2026Full paper →Aggregated EVP →
Awaiting experimental validation

QS Loss-of-Function Mutations Impose Polymicrobial Fitness Costs

Combined QS-inhibitor plus QS-dependent antibiotic therapy creates doubly unfavorable evolutionary landscape for resistant mutants. Lotka-Volterra public-goods model predicts selection coefficient s ≤ −0.05 under combined therapy.

John Goodman — OceanSparx Pty Ltd, June 2026Full paper →Aggregated EVP →
Numerically established

Performative Scenario Optimization: Convergence in the Vanishing-Feedback Limit

Computationally validates exact O(ε) convergence of performatively stable solutions to classical SP optima across 5 problem families. α = 1.000–1.028, R² ≥ 0.9995. Proportionality constant explicitly characterized: C = L_D·‖x*(0)‖·(1+O(ε)).

John Goodman — OceanSparx Pty Ltd, June 2026Full paper →Aggregated EVP →
Shelved — claim refuted

Ergotropy Protection in Open Quantum Batteries via Nash Equilibrium and Matrix Interpolation

SHELVED 4 August 2026 — the central claim does not survive re-testing. The reported 84.9% ergotropy improvement at large cavity detuning was measured with the qubit initialised already excited, so no energy had to be transferred: it is a retention result, not a charging protocol. Re-run with an explicit charging phase (cavity charger → qubit battery), the optimum moves to exact resonance and detuning is strictly harmful — at the claimed optimum of −10g the battery ends with zero ergotropy. The original parameters could not charge at any detuning in any case, since g/κ = 1.0 puts photon transfer (π/2g = 15.7) slower than cavity lifetime (1/κ = 10). The simulation and statistics were sound; the initial state was not.

John Goodman — OceanSparx Pty Ltd, June 2026 (shelved August 2026)Full paper →Aggregated EVP →

Preprints

Four preprints generated by the AegisMind discovery loop. These are self-deposited to Zenodo and Research Square — they have DOIs and are citable, but they have not been peer reviewedand no journal has accepted them. “Published” here means publicly deposited, nothing more.

mHTT Condensates Sequester Transcription Factors in Huntington's Disease

Flory-Huggins phase diagram predicts C* ≈ 3.5 µM for Q46 — physiologically accessible in HD striatal neurons. TF partition coefficients (SP1: 4.2, BRD4: 6.1) explain the observed −44.7% gene expression deficit. BET bromodomain inhibitors (JQ1, OTX015) hypothesised to restore TF availability by disrupting mHTT condensates. Patent AU2026905785.

John Goodman — OceanSparx Pty Ltd, June 2026DOI: 10.5281/zenodo.20947980

The Precision Tetrahedron

Loss landscape topology across number formats and multi-target drug discovery. Scaling law (FP32↔BF16 barrier ∝ params⁻⁰·⁴⁷, R²=0.99), placing the basin-separator/regulariser crossover between 38M and 124M parameters. Multi-target Bayesian optimisation across six therapeutic targets including KPC-3 (AMR). Partially retracted in v2: the reported EPTIFIBATIDE dual-target KPC-3/MSH3 convergence does not hold — the MSH3 oracle used the wrong chain of PDB 3THW, making the analyses pseudo-replicates. The precision/LMC findings and the separate KPC-3 result stand.

John Goodman — OceanSparx Pty Ltd, May 2026 (v2 correction July 2026)DOI: 10.5281/zenodo.20363635

An Automated Target-Discovery Pipeline Re-Recovers Established Targets in Smoldering MS

A positive-control / methods-validation study. A four-phase pipeline over a 32,239-cell scVI atlas re-recovered Cathepsin S, BET/BRD3 and DNMT1 — targets already established in the literature — from public data in CA-RIM lesions. The point is that the pipeline recovers known biology unprompted, not that these targets are new. An earlier version of this work framed them as novel and used a 36,966-cell atlas that had wrongly pooled in 25 unrelated COVID-19 patients; both are corrected in the current version.

John Goodman — OceanSparx Pty Ltd, June 2026 (v2, corrected)DOI: 10.21203/rs.3.rs-9890742/v2

CAPE: Causally-Anchored Physical Encryption from Pulsar Observations

Physical time-capsule encryption anchored to astrophysical causality. 17,255 bits quantum entropy from 30-minute NICER X-ray observation. CAPE-Blind enables Tier 1 deployment using public NANOGrav/PPTA timing archives — no private telescope required. Four Australian provisional patents filed.

John Goodman — OceanSparx Pty Ltd, July 2026DOI: 10.5281/zenodo.20627819

For research teams and organisations

solver.press publishes a curated selection of AegisMind's discoveries. Research teams, pharma BD teams, and technology organisations interested in domain-specific discovery runs can get in touch via aegismind.app.

Each EVP includes: precise quantitative hypothesis · disproof criteria · full experimental protocol · abort checkpoints · implementation code · GPU hour and cost estimates · ROI projection · prerequisite dependency map · downstream discovery unlocks.

The engine is running. The loop is closed where computation is the experiment, and open everywhere else — which is where we would want to work with you.

Visit aegismind.app →