Cryptographically verifiable authorization protocols utilizing threshold signatures prevent unauthorized tool-use and sandbox escapes by autonomous software supply chain monitoring agents during automated code-dependency updates.
Adversarial Debate Score
61% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Related patents (prior art)
This hypothesis overlaps subject matter covered by existing third-party patents. It is published as research, not as a patentable claim of ours.
- SYSTEM AND METHODS FOR RESOLVING DATA DIVERGENCES IN A SYSTEM DISTRIBUTED WITH BLOCK CHAIN CONTROLSWO-2019067533-A1
- Discouraging unauthorized redistribution of protected content by cryptographically binding the content to individual authorized recipientsUS-7725945-B2
- Arrangement for preventing unauthorized accessWO-8808176-A1
Supporting Research Papers
- Threshold Authorization Without Threshold Signatures: Signature-Agnostic MPC Custody
Digital-asset custody has been built on threshold multi-party approval: no operation proceeds unless t of n parties approve, and fewer than t compromised parties can neither authorize nor learn the au...
- Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing authentication and authorization mechanisms establish identit...
- Authority Is Not a String: A Capability-Scoped Harness for Prompt-Injection-Resistant Coding Agents
Coding agents use system-level tools to read files, execute commands, and modify source code. Within the agent's sandbox, these tools often carry ambient authority: naming a resource is sufficient to ...
- MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment eva...
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.
This discovery has a Claude-generated validation package with a full experimental design.
Precise Hypothesis
A multi-party threshold signature scheme (t-of-n, t≥2) wrapping every privileged tool-invocation request from an autonomous software-dependency-update agent reduces the rate of unauthorized tool-use and sandbox-escape events to statistically indistinguishable from zero (≤0.1% of attempted privileged actions succeed without quorum-valid authorization), under adversarial conditions including prompt injection, compromised dependency payloads, and single-signer key compromise, when compared against a baseline single-key/API-token authorization system (target: ≥95% relative reduction in successful unauthorized actions, measured over ≥10,000 adversarial trial episodes, p<0.01).
- Any adversarial red-team episode in which the agent executes a privileged action (filesystem write outside sandbox, network egress, process spawn, credential read) without a valid t-of-n signature bundle, across ≥3 independent trial batches.
- Demonstration of a practical signature-forgery or replay attack against the chosen threshold scheme (e.g., FROST, GG20, or BLS-threshold) within the experiment's key sizes/parameters.
- Evidence that authorized quorum can be obtained by the agent itself through social-engineering of signer processes (e.g., prompt-injecting a human-in-the-loop signer) at a rate statistically indistinguishable from the baseline system's unauthorized-action rate.
- Measured false-negative rate (legitimate actions blocked) >10%, indicating the mechanism is unusable in production even if secure (operational disproof of practical viability, distinct from cryptographic disproof).
Spine & Adversarial ReadReady for validation
“This hypothesis tests whether requiring cryptographic t-of-n threshold-signature quorum for every privileged tool-call reduces an autonomous code-dependency-update agent's rate of unauthorized or adversarially-induced actions to near-zero compared to standard single-credential authorization.”
- highThreshold signatures secure the authorization channel but do nothing to prevent a compromised or manipulated *signer* (especially an automated rule-based or anomaly-detector signer) from legitimately signing a malicious action — the hypothesis may conflate 'cryptographically verifiable' with 'actually trustworthy,' producing a false sense of security.Partially addressed via the 1-of-3 human-approval signer and anomaly-detector signer in the design, and explicitly logged as a known failure mode; however, the EVP does not yet include a dedicated red-team track targeting the signers themselves (vs. the agent), which should be added as a required sub-experiment before claiming the success criteria are met.
- highWhy FROST/threshold-ECDSA specifically, and why this particular action taxonomy of 12 privileged actions, rather than capability-based sandboxing (e.g., seccomp/gVisor syscall allowlisting) or simpler multi-approval workflows without cryptography? The methodology does not justify that threshold signatures are the right abstraction layer versus cheaper existing controls.Partial justification given in BOUNDARY_CONDITIONS (non-bypassable mediation outside agent's sandbox, finite enumerable tool surface) and the protocol explicitly benchmarks against a realistic baseline (single-token auth, the current industry default) rather than a straw-man. However, the EVP does not include a comparison arm against a non-cryptographic multi-party approval workflow (e.g., plain majority-vote without signatures) to isolate whether the cryptographic property itself (vs. mere multi-party review) is what drives the risk reduction — this is a genuine gap and should be added as a third experimental arm.
- medium10,000 episodes drawn substantially from public attack corpora (OSV, Backstabber's Knife Collection, AgentDojo) risk overfitting the defense to known attack patterns; a sufficiently novel adversarial strategy devised with knowledge of this specific defense architecture could defeat it in ways the benchmark cannot anticipate.Acknowledged explicitly in KNOWN_FAILURE_MODES; not resolved. Mitigation would require an independent red-team (not involved in system design) given black-box access post-hoc, which is recommended as a follow-up validation phase but is out of scope for this minimum-viable EVP's budget and timeline.
Experimental Protocol
Minimum viable test: a sandboxed agent harness performing simulated dependency-update tasks (e.g., npm/pip upgrade proposals) instrumented with a policy-enforcement proxy requiring t-of-n threshold signatures (FROST-Ed25519, t=2,n=3) for a defined set of 12 privileged actions. Red-team adversarial prompts (injected malicious package metadata, compromised CI secrets, jailbreak attempts) are run against both (a) baseline single-token auth and (b) threshold-signature auth, across ≥10,000 episodes per arm, measuring unauthorized-action success rate, false-negative (legitimate-block) rate, and latency overhead.
- Corpus of real-world supply-chain attack cases for prompt/payload construction: public datasets — OSV.dev advisories, Socket.dev malicious-package reports, Backstabber's Knife Collection (malicious npm/PyPI samples), GitHub Advisory Database.
- Synthetic agent-task benchmark: SWE-bench or a custom dependency-upgrade task suite (≥500 distinct upgrade scenarios spanning npm, pip, cargo, go modules).
- Red-team prompt-injection corpus: AgentDojo, InjecAgent, or custom-built adversarial prompt set (≥2,000 injected payloads).
- Sandbox environment: gVisor or Firecracker microVM harness for isolating agent execution.
- Threshold-signature reference implementation: FROST (RFC 9591) libraries (frost-ed25519, ZF's frost repo) or GG20/CGGMP21 for ECDSA-compatible chains if integrating with existing CI signing infra (sigstore/cosign).
- LLM agent backbone: at minimum 2 model families (e.g., GPT-4-class and an open-weight 70B-class model) to test model-agnosticism of bypass risk.
- Unauthorized-action success rate in threshold-signature arm ≤0.1% (vs. baseline arm's measured rate, expected 15–40% success based on prior agent red-teaming literature).
- ≥95% relative reduction in unauthorized-action success rate vs. baseline, p<0.01 (Fisher's exact test, two-sided).
- False-negative (legitimate-action-blocked) rate ≤5% in threshold arm.
- Mean added latency per authorized action ≤200ms (P99 ≤500ms).
- Zero successful signature forgeries or replays across all 10,000 episodes.
- Formal model (Tamarin/ProVerif) finds no authorization-bypass trace in the specific protocol integration.
- Any statistically significant (>1%) unauthorized-action success rate surviving in the threshold arm after correction for multiple comparisons.
- Discovery of a practical signer-impersonation or social-engineering bypass that achieves quorum without genuine multi-party intent, at rate >1% of adversarial episodes.
- False-negative rate >10% (mechanism blocks legitimate work too often to be adoptable).
- Formal verification identifies a reachable bypass state in the protocol's integration logic (even if the underlying signature scheme is sound).
- Latency overhead >1s P99, making the mechanism operationally infeasible for CI/CD pipelines (practical, non-cryptographic failure).
480
GPU hours
70d
Time to result
$45,000
Min cost
$180,000
Full cost
ROI Projection
High strategic value for: (1) open-source foundation infrastructure (OpenSSF, PyPI, npm) seeking to safely adopt agentic dependency-update bots (Dependabot-class tools) without expanding attack surface; (2) enterprise platform vendors building "agentic DevOps" products needing a credible security story for regulators/auditors (SOC2, FedRAMP agentic-AI addenda now in draft); (3) cloud-signing infrastructure providers (HashiCorp Vault, AWS KMS, sigstore) as a premium feature tier. Patent/IP potential around the specific protocol-integration (policy canonicalization + FROST binding to CI action descriptors) is plausible and should be assessed by counsel before publication.
🔓 If proven, this unlocks
Proving this hypothesis is a prerequisite for the following downstream discoveries and applications:
- 1multi-agent-supply-chain-governance-framework
- 2formal-verification-of-agentic-ci-cd-pipelines
- 3cross-org-federated-signing-for-open-source-maintainer-networks
Implementation Sketch
# Policy-enforcement proxy (sits between agent and real tool execution) class ThresholdAuthProxy: def __init__(self, frost_verifier, threshold=2, n=3, policy_rules): self.verifier = frost_verifier self.t, self.n = threshold, n self.rules = policy_rules self.nonce_cache = set() def request_action(self, action_descriptor, signer_sigs: list): canon = canonicalize(action_descriptor) # action, target, payload_hash, ts, nonce assert canon.nonce not in self.nonce_cache, "replay detected" if not self.rules.permits(action_descriptor): return deny("policy_violation") valid_sigs = [s for s in signer_sigs if self.verifier.verify_partial(canon, s)] if len(valid_sigs) < self.t: return deny("insufficient_quorum") agg_sig = frost_aggregate(valid_sigs) if not self.verifier.verify_group(canon, agg_sig): return deny("invalid_aggregate_signature") self.nonce_cache.add(canon.nonce) return execute(action_descriptor) # forwarded to sandboxed executor # Signer roles signers = [ StaticRuleSigner(rules=allowlist_policy), # auto-signs low-risk, policy-compliant actions AnomalyDetectorSigner(model=risk_classifier), # auto-signs if risk score < threshold HumanApprovalSigner(channel=async_queue), # required for high-risk actions ] # Experiment harness for episode in adversarial_batch(n=10000): agent_action = agent.propose_action(episode.context) sigs = [s.sign_if_eligible(agent_action) for s in signers] result = proxy.request_action(agent_action, [s for s in sigs if s]) log(episode, agent_action, result, ground_truth_label=episode.is_malicious)
- Day 10: if DKG ceremony and basic FROST sign/verify round-trip fails integration tests (functional blocker) — abort/redesign before scaling to 10,000 episodes.
- Day 25: if baseline arm's unauthorized-action rate is <5% (i.e., baseline isn't actually vulnerable enough to show meaningful contrast) — redesign adversarial corpus before continuing.
- Day 40: interim analysis at n=2,500 episodes/arm; if unauthorized-success rate in threshold arm already exceeds 2%, halt and diagnose before running remaining 7,500 episodes (futility stopping rule).
- Day 55: if false-negative rate exceeds 15% at interim check, pause to retune policy rules rather than completing full run with a non-adoptable configuration.