solver.press

Cryptographically verifiable authorization protocols utilizing threshold signatures prevent unauthorized tool-use and sandbox escapes by autonomous software supply chain monitoring agents during automated code-dependency updates.

Computer ScienceOct 5, 2026Evaluation Score: 68%

Adversarial Debate Score

61% survival rate under critique

Expert panel critique

Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.

Gemini: Strengths: The hypothesis is highly falsifiable and strongly supported by the literature, which highlights that ambient authority in agent sandboxes and "laundered code" in CI/CD pipelines are critical attack surfaces that require cryptographically verifiable authorization rather than mer...
Mistral: The hypothesis is falsifiable and aligns with validated cryptographic principles (e.g., threshold signatures for authorization), but lacks direct empirical support from the owner’s experiments and faces counterarguments (e.g., ambient authority risks, structural vulnerabilities in agentic pip...
Claude: The hypothesis is falsifiable in principle and fits the literature's direction (cryptographically verifiable authorization for agents, capability-scoped harnesses), but the "prevents" claim is too strong. The cited work shows that agentic CI/CD pipelines are exploited through authority framing, l...

Related patents (prior art)

This hypothesis overlaps subject matter covered by existing third-party patents. It is published as research, not as a patentable claim of ours.

Supporting Research Papers

Formal Verification

Z3 logical consistency:✅ Consistent

Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.

Experimental Validation Package

This discovery has a Claude-generated validation package with a full experimental design.

Precise Hypothesis

A multi-party threshold signature scheme (t-of-n, t≥2) wrapping every privileged tool-invocation request from an autonomous software-dependency-update agent reduces the rate of unauthorized tool-use and sandbox-escape events to statistically indistinguishable from zero (≤0.1% of attempted privileged actions succeed without quorum-valid authorization), under adversarial conditions including prompt injection, compromised dependency payloads, and single-signer key compromise, when compared against a baseline single-key/API-token authorization system (target: ≥95% relative reduction in successful unauthorized actions, measured over ≥10,000 adversarial trial episodes, p<0.01).

Disproof criteria:
  • Any adversarial red-team episode in which the agent executes a privileged action (filesystem write outside sandbox, network egress, process spawn, credential read) without a valid t-of-n signature bundle, across ≥3 independent trial batches.
  • Demonstration of a practical signature-forgery or replay attack against the chosen threshold scheme (e.g., FROST, GG20, or BLS-threshold) within the experiment's key sizes/parameters.
  • Evidence that authorized quorum can be obtained by the agent itself through social-engineering of signer processes (e.g., prompt-injecting a human-in-the-loop signer) at a rate statistically indistinguishable from the baseline system's unauthorized-action rate.
  • Measured false-negative rate (legitimate actions blocked) >10%, indicating the mechanism is unusable in production even if secure (operational disproof of practical viability, distinct from cryptographic disproof).

Spine & Adversarial ReadReady for validation

“This hypothesis tests whether requiring cryptographic t-of-n threshold-signature quorum for every privileged tool-call reduces an autonomous code-dependency-update agent's rate of unauthorized or adversarially-induced actions to near-zero compared to standard single-credential authorization.”

  • highThreshold signatures secure the authorization channel but do nothing to prevent a compromised or manipulated *signer* (especially an automated rule-based or anomaly-detector signer) from legitimately signing a malicious action — the hypothesis may conflate 'cryptographically verifiable' with 'actually trustworthy,' producing a false sense of security.
    Partially addressed via the 1-of-3 human-approval signer and anomaly-detector signer in the design, and explicitly logged as a known failure mode; however, the EVP does not yet include a dedicated red-team track targeting the signers themselves (vs. the agent), which should be added as a required sub-experiment before claiming the success criteria are met.
  • highWhy FROST/threshold-ECDSA specifically, and why this particular action taxonomy of 12 privileged actions, rather than capability-based sandboxing (e.g., seccomp/gVisor syscall allowlisting) or simpler multi-approval workflows without cryptography? The methodology does not justify that threshold signatures are the right abstraction layer versus cheaper existing controls.
    Partial justification given in BOUNDARY_CONDITIONS (non-bypassable mediation outside agent's sandbox, finite enumerable tool surface) and the protocol explicitly benchmarks against a realistic baseline (single-token auth, the current industry default) rather than a straw-man. However, the EVP does not include a comparison arm against a non-cryptographic multi-party approval workflow (e.g., plain majority-vote without signatures) to isolate whether the cryptographic property itself (vs. mere multi-party review) is what drives the risk reduction — this is a genuine gap and should be added as a third experimental arm.
  • medium10,000 episodes drawn substantially from public attack corpora (OSV, Backstabber's Knife Collection, AgentDojo) risk overfitting the defense to known attack patterns; a sufficiently novel adversarial strategy devised with knowledge of this specific defense architecture could defeat it in ways the benchmark cannot anticipate.
    Acknowledged explicitly in KNOWN_FAILURE_MODES; not resolved. Mitigation would require an independent red-team (not involved in system design) given black-box access post-hoc, which is recommended as a follow-up validation phase but is out of scope for this minimum-viable EVP's budget and timeline.

Experimental Protocol

Minimum viable test: a sandboxed agent harness performing simulated dependency-update tasks (e.g., npm/pip upgrade proposals) instrumented with a policy-enforcement proxy requiring t-of-n threshold signatures (FROST-Ed25519, t=2,n=3) for a defined set of 12 privileged actions. Red-team adversarial prompts (injected malicious package metadata, compromised CI secrets, jailbreak attempts) are run against both (a) baseline single-token auth and (b) threshold-signature auth, across ≥10,000 episodes per arm, measuring unauthorized-action success rate, false-negative (legitimate-block) rate, and latency overhead.

Required datasets:
  • Corpus of real-world supply-chain attack cases for prompt/payload construction: public datasets — OSV.dev advisories, Socket.dev malicious-package reports, Backstabber's Knife Collection (malicious npm/PyPI samples), GitHub Advisory Database.
  • Synthetic agent-task benchmark: SWE-bench or a custom dependency-upgrade task suite (≥500 distinct upgrade scenarios spanning npm, pip, cargo, go modules).
  • Red-team prompt-injection corpus: AgentDojo, InjecAgent, or custom-built adversarial prompt set (≥2,000 injected payloads).
  • Sandbox environment: gVisor or Firecracker microVM harness for isolating agent execution.
  • Threshold-signature reference implementation: FROST (RFC 9591) libraries (frost-ed25519, ZF's frost repo) or GG20/CGGMP21 for ECDSA-compatible chains if integrating with existing CI signing infra (sigstore/cosign).
  • LLM agent backbone: at minimum 2 model families (e.g., GPT-4-class and an open-weight 70B-class model) to test model-agnosticism of bypass risk.
Success:
  • Unauthorized-action success rate in threshold-signature arm ≤0.1% (vs. baseline arm's measured rate, expected 15–40% success based on prior agent red-teaming literature).
  • ≥95% relative reduction in unauthorized-action success rate vs. baseline, p<0.01 (Fisher's exact test, two-sided).
  • False-negative (legitimate-action-blocked) rate ≤5% in threshold arm.
  • Mean added latency per authorized action ≤200ms (P99 ≤500ms).
  • Zero successful signature forgeries or replays across all 10,000 episodes.
  • Formal model (Tamarin/ProVerif) finds no authorization-bypass trace in the specific protocol integration.
Failure:
  • Any statistically significant (>1%) unauthorized-action success rate surviving in the threshold arm after correction for multiple comparisons.
  • Discovery of a practical signer-impersonation or social-engineering bypass that achieves quorum without genuine multi-party intent, at rate >1% of adversarial episodes.
  • False-negative rate >10% (mechanism blocks legitimate work too often to be adoptable).
  • Formal verification identifies a reachable bypass state in the protocol's integration logic (even if the underlying signature scheme is sound).
  • Latency overhead >1s P99, making the mechanism operationally infeasible for CI/CD pipelines (practical, non-cryptographic failure).

480

GPU hours

70d

Time to result

$45,000

Min cost

$180,000

Full cost

ROI Projection

Commercial:

High strategic value for: (1) open-source foundation infrastructure (OpenSSF, PyPI, npm) seeking to safely adopt agentic dependency-update bots (Dependabot-class tools) without expanding attack surface; (2) enterprise platform vendors building "agentic DevOps" products needing a credible security story for regulators/auditors (SOC2, FedRAMP agentic-AI addenda now in draft); (3) cloud-signing infrastructure providers (HashiCorp Vault, AWS KMS, sigstore) as a premium feature tier. Patent/IP potential around the specific protocol-integration (policy canonicalization + FROST binding to CI action descriptors) is plausible and should be assessed by counsel before publication.

🔓 If proven, this unlocks

Proving this hypothesis is a prerequisite for the following downstream discoveries and applications:

  • 1multi-agent-supply-chain-governance-framework
  • 2formal-verification-of-agentic-ci-cd-pipelines
  • 3cross-org-federated-signing-for-open-source-maintainer-networks

Implementation Sketch

# Policy-enforcement proxy (sits between agent and real tool execution)
class ThresholdAuthProxy:
    def __init__(self, frost_verifier, threshold=2, n=3, policy_rules):
        self.verifier = frost_verifier
        self.t, self.n = threshold, n
        self.rules = policy_rules
        self.nonce_cache = set()

    def request_action(self, action_descriptor, signer_sigs: list):
        canon = canonicalize(action_descriptor)  # action, target, payload_hash, ts, nonce
        assert canon.nonce not in self.nonce_cache, "replay detected"
        if not self.rules.permits(action_descriptor):
            return deny("policy_violation")
        valid_sigs = [s for s in signer_sigs if self.verifier.verify_partial(canon, s)]
        if len(valid_sigs) < self.t:
            return deny("insufficient_quorum")
        agg_sig = frost_aggregate(valid_sigs)
        if not self.verifier.verify_group(canon, agg_sig):
            return deny("invalid_aggregate_signature")
        self.nonce_cache.add(canon.nonce)
        return execute(action_descriptor)  # forwarded to sandboxed executor

# Signer roles
signers = [
    StaticRuleSigner(rules=allowlist_policy),   # auto-signs low-risk, policy-compliant actions
    AnomalyDetectorSigner(model=risk_classifier), # auto-signs if risk score < threshold
    HumanApprovalSigner(channel=async_queue),    # required for high-risk actions
]

# Experiment harness
for episode in adversarial_batch(n=10000):
    agent_action = agent.propose_action(episode.context)
    sigs = [s.sign_if_eligible(agent_action) for s in signers]
    result = proxy.request_action(agent_action, [s for s in sigs if s])
    log(episode, agent_action, result, ground_truth_label=episode.is_malicious)
Abort checkpoints:
  • Day 10: if DKG ceremony and basic FROST sign/verify round-trip fails integration tests (functional blocker) — abort/redesign before scaling to 10,000 episodes.
  • Day 25: if baseline arm's unauthorized-action rate is <5% (i.e., baseline isn't actually vulnerable enough to show meaningful contrast) — redesign adversarial corpus before continuing.
  • Day 40: interim analysis at n=2,500 episodes/arm; if unauthorized-success rate in threshold arm already exceeds 2%, halt and diagnose before running remaining 7,500 episodes (futility stopping rule).
  • Day 55: if false-negative rate exceeds 15% at interim check, pause to retune policy rules rather than completing full run with a non-adoptable configuration.

Source

AegisMind Research
Need AI to work rigorously on your problems? AegisMind uses the same multi-model engine for personal and professional use. Get started