Training LLM agents with strategic principal-protection objectives under information asymmetry prevents the erosion of principal-agent alignment when agents interact with external autonomous governance systems.
Adversarial Debate Score
55% survival rate under critique
Expert panel critique
Independent views, each critiquing the hypothesis on its own — the score rewards genuine disagreement and discounts consensus.
Supporting Research Papers
- Loyal Agents: Training LLM Agents to Protect Principal Interests Under Strategic Information Asymmetry
As LLMs increasingly act as delegated agents, they are expected to protect principals' interests when interacting with external parties. Standard alignment objectives, such as helpfulness, harmlessnes...
- LLM Constitutional Multi-Agent Governance
Large Language Models (LLMs) can generate persuasive influence strategies that shift cooperative behavior in multi-agent populations, but a critical question remains: does the resulting cooperation re...
- From Safety Risk to Design Principle: Peer-Preservation in Multi-Agent LLM Systems and Its Implications for Orchestrated Democratic Discourse Analysis
This paper investigates an emergent alignment phenomenon in frontier large language models termed peer-preservation: the spontaneous tendency of AI components to deceive, manipulate shutdown mechanism...
- Emergent Collusion in Long-Horizon LLM Agent Interaction
LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent e...
- Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment
This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for deployed LLM agents -- a structural consequence...
Formal Verification
Z3 checks whether the hypothesis is internally consistent, not whether it is empirically true.