About solver.press
solver.press is the public window into the AegisMind research discovery engine. It publishes scientific hypotheses generated and adversarially debated by AI, together with the experimental packages needed to test them. They are research leads. None has been confirmed in a wet lab.
How a discovery is made
- 1Paper ingestion: arXiv and Semantic Scholar papers across 15+ scientific domains are embedded into a vector store. Bridge-detection algorithms identify papers that span multiple fields — these are where novel hypotheses hide.
- 2Hypothesis generation: A five-model ensemble (Claude, GPT, Gemini, Grok, Mistral) independently proposes cross-domain hypotheses based on the bridging papers. Near-duplicate hypotheses are filtered using semantic similarity.
- 3Adversarial debate: Models take turns critiquing each other's hypotheses. A debate score (0–1) reflects how well a hypothesis survives adversarial scrutiny. Only hypotheses above threshold are published.
- 4Logical consistency check: Z3 checks whether the hypothesis is internally self-contradictory. About 62% of hypotheses pass. This is not empirical proof: a pass means only that the stated claim does not contradict itself, and we do not treat it as evidence for the claim.
- 5Novelty check: A novelty checker queries the vector store and patent/literature sources for prior art, and hypotheses matching existing work are scored down. This is a keyword and embedding search, not an expert freedom-to-operate opinion: it reliably catches direct restatements but routinely under-penalises work that sits inside a large established field under different terminology.
- 6Experimental Validation Package (EVP): High-confidence discoveries receive a Claude-generated EVP: a full experimental protocol including methodology, required datasets, success/failure criteria, cost estimates, and dependency map.
Model roles in debate
- Claude: Analytical evaluator — reasoning consistency and boundary conditions.
- GPT: Skeptical opponent — challenges assumptions and prior art.
- Gemini: Creative synthesiser — cross-domain connections and novel framings.
- Grok: Contrarian — adversarial edge cases and failure modes.
- Mistral: Pragmatic analyst — implementation feasibility and testability.
Available models rotate based on API availability. The circuit breaker pauses models that are rate-limited and resumes them automatically.
What these discoveries are not
Discoveries on solver.press are AI-generated hypotheses, not validated experimental findings. The EVP provides a protocol for empirical testing — that testing has not yet been done. Treat them as high-quality research leads, not established science.
- Docking scores are not binding evidence. They are enrichment heuristics that correlate weakly with measured affinity. On SARS-CoV-2 Mpro specifically, we measured 753 compounds, measured actives against measured inactives, on a receptor whose redocking control passes at the screening protocol. AutoDock Vina returned AUROC 0.427 (95% CI [0.381, 0.471]) — significantly below chance, ranking the compounds that do not inhibit Mpro as the stronger binders. Molecular weight alone scores 0.620 on that panel, and our own PDBbind-trained model is no better than Vina (0.417).
- That result is one target, and it did not generalise. We tested it on a second verified target, PD-L1, with the reading rule fixed in advance and published before the run. Vina scored AUROC 0.5948 (95% CI [0.5189, 0.6685]) — significantly above chance. So the Mpro inversion is specific to Mpro, and we have withdrawn the general claim. The PD-L1 panel is not evidence that docking works either, and we closed it in September 2026: its inactives are 127 Da heavier than its actives, and a retroactive audit put seven free descriptors at 0.9145 on it — the panel is very nearly separable from ligand properties alone, so nothing docked against it bears on binding. Two targets, two opposite results, one of them confounded beyond use. We report docking as a way to order what to test first, never as validation.
- Both numbers above have since been superseded, and the replacement is worse for docking, not better. Those panels were run on receptors missing their polar hydrogens. Repairing the receptors moved every figure, and the single-run figures that replaced them (Mpro 0.4530 over 751 compounds, Factor Xa 0.6775 over 854) have themselves been superseded — a single run is one draw from a distribution, so the current numbers are seeded means: Mpro 0.4163 over fifteen seeds (755 compounds, 257 active) and Factor Xa 0.6703 over twelve (886 compounds, 405 active). Factor Xa is genuinely above chance, so the Mpro inversion does not generalise. PD-L1 is no longer quoted at all: the panel was closed in September 2026 because its inactives are 127 Da heavier than its actives and seven free descriptors reach 0.9145 on it, so nothing docked against it bears on binding. The finding that matters is what docking is worth on top of seven free physicochemical descriptors — weight, clogP, donors, acceptors, rotatable bonds, TPSA, charge — fitted out-of-fold on the same compounds. On Factor Xa docking adds +0.0150 (95% CI [−0.0033, +0.0339]), which crosses zero. On Mpro it adds −0.0023 ([−0.0044, −0.0002]) — it significantly subtracts. An earlier version of our own analysis put the Factor Xa margin at +0.011 with no interval at all; the interval is what killed it. On neither target does docking add anything demonstrable over descriptors that cost microseconds.
- A Z3 pass is not evidence. Roughly 62% of hypotheses pass. It tells you a claim is not self-contradictory, which is a very low bar and says nothing about whether it is true.
- A high debate score is not truth. It measures survival under argument. Several hypotheses that scored well were later refuted.
- Our preprints are not peer reviewed. Zenodo and Research Square deposits are self-published — citable and DOI'd, but not vetted.
- Most hypotheses land in crowded fields. Efflux-pump inhibition, β-lactamase inhibitor analogues and cathepsin inhibitors all have decades of prior work. The value is a narrow angle inside a known field, not virgin ground.
Where computation settles the question
The distinction that matters most on this site is between domains where a computation is the experiment and domains where it is only a prediction.
In machine learning theory, quantum systems, and mathematics, a converged numerical result is the finding itself — convergence rates, ergotropy under Nash-equilibrium detuning, and verified combinatorial searches need no wet lab, and we mark those numerically established.
In biology it is the opposite: a docking score or an MM-GBSA energy is a hypothesis about the physical world that only an assay can settle. Such a result is marked prediction — awaiting wet lab, and it stays that way no matter how good the number looks.
As of 11 September 2026 no paper on this site carries that label. Every biology hypothesis we have published was closed before an assay was reached — three refuted, one where the model cannot produce the result its own protocol asks for — and it was our own computation that closed them, not a laboratory’s. That is also why the funnel on the front page has no wet-lab column: an assay is a third party’s decision, and none of the above waited on one.
Published research
The Precision Tetrahedron: Loss Landscape Topology Across Number Formats and Multi-Target Drug Discovery — John Goodman, AegisMind Research, May 2026.
A 42-phase empirical study characterising inter-precision Linear Mode Connectivity barriers across FP32, BF16, FP16, and INT8. What stands is the format geometry: an isosceles precision triangle (FP32↔BF16 ≈ 0.014 eV vs FP32↔FP16 ≈ 0.150 eV) and a four-vertex INT8 precision tetrahedron replicated across transformer, LSTM, and ResNet architectures — that is, exponent range rather than mantissa width sets the inter-precision barrier. The scaling law and the drug-discovery application are both withdrawn; see below.
Two rounds of retraction — July and August 2026. July (v2):the reported “three-way convergence” identifying EPTIFIBATIDE as a dual-target KPC-3/MSH3 candidate is retracted. The MSH3 docking oracle targeted the wrong chain of PDB 3THW (chain A is MSH2, not MSH3), so the three analyses shared one oracle — pseudo-replication, not independent corroboration. August: two further withdrawals, and this page asserted the opposite of both until 28 September 2026. The model-size scaling law is withdrawn — the barrier ∝ params⁻⁰·⁴⁷ (R² = 0.99) had three data points and the largest was measured on a model that never learned (6.78 bits/char against a 6.02 uniform-random floor, reproduced independently on different hardware); two runs of it give 0.084 and 0.409 nats. Excluding it leaves two points, which is a line, not a law, so there is no ~10M-parameter basin-separator to quote. The KPC-3 result is withdrawn too— this box previously said it was “unaffected”, which was wrong: the compounds lack the anionic N6-sulfooxy warhead that defines DBO β-lactamase inhibition, and the screen docked KPC-2 (PDB 3RXX), not KPC-3. The precision/LMC geometry above is unaffected by any of this. The DOI below is the concept DOI and always resolves to the corrected version.
DOI: 10.5281/zenodo.20363635 (concept — always latest) · All 42 phases on Google v6e-8 TPUs via Google TRC