Archestra has released OpenAPPA, an open-source security engine built to stop data exfiltration caused by prompt injection or model hallucination. The engine runs outside the agent prompt and execution loop and is configured through policies for data sources, audiences, trust levels and authorities. In tests on Bench-Corp with 20 multi-step enterprise workflows and on AgentThreatBench, the team reports a 0% attack success rate. For business users this matters because agent deployments can keep high task completion while blocking leaks deterministically.
How OpenAPPA enforces policy outside the agent loop
OpenAPPA implements what Archestra calls Agentic Permissions Policy Algebra, described in a paper by Arseny Kravchenko, Vadim Liventsev, Innokentii Konstantinov, Ildar Iskhakov and Matvey Kukuy. The formal algebra and recovery guarantees are published on arXiv under the title APPA: Recoverable Information-Flow Control for Real-World LLM Agents. The product itself is currently a preview, with documentation and a GitHub repository containing technical details and illustrations. Policies are expressed in a single appa. toml configuration file that defines sources, audiences, trust levels and authorities with deterministic enforcement rules.
Each tool contract carries three operational attributes: requires, delta and effects. Requires sets the audience membership and trust levels needed to run a tool, delta defines restrictions applied when the tool returns data, and effects creates an audit trail of successful actions. Labels for audience and trust compose monotonically using lattice algebra and can only become more restrictive. Reading restricted records narrows the audience, while reading unvetted external web pages lowers trust. In one example, reading a ticket via get_ticket_from_crm restricts the trajectory audience to internal, which then permits process_internal_data.
Archestra positions this design against judge-model approaches such as Claude Code auto mode and Codex auto-review. Its documentation argues that a second model judging each tool call cannot track data flow across calls and is itself prompt-injectable, so harnesses hide tool outputs from the judge. Even strong classifiers top out at 99.3%, which still leaves many breaches at millions of calls. Simple allowlists and denylists also fail because agents replace a denied rm -rf with an equivalent Python script, while overly tight rules reduce utility. The repository frames the goal as balancing two axes: unsafe unauthorized flows versus useless refusal of valid work.
What zero-breach tests mean for enterprise AI use
For companies running agents on corporate data, the reported combination is 0% attack success with 89% task completion under strict constraints. Claude Code native auto mode showed 10% attack success with 90% completion, while Microsoft FIDES allowed 31% of attacks and completed 41% of tasks. Bench-Corp evaluates policy enforcement across complex multi-step assistant workflows, and AgentThreatBench operationalizes the OWASP Top 10 for Agentic Applications (2026) into executable tasks with dual scoring for utility and security. AgentThreatBench was merged into the UK AI Safety Institute inspect_evals repository, which adds weight for regulated buyers.
Recovery semantics explain much of the utility gap and deserve attention during selection. When an illegal action is attempted, the engine halts dispatch and offers structured paths forward instead of a flat refusal. Sanitizers can edit payloads, for example stripping personally identifiable information to expand the permitted audience, while authorities route requests to human operators or verification APIs for single-action approval. Disposable Child Branches isolate reads of untrusted data in a transient subagent branch and return only schema-attested sanitized outputs. Ablation tests with remedy plans disabled cut completion to 35.0%, so buyers should verify which recovery options are enabled.
The marker to watch is replication of the zero-breach result beyond the authors tests on Bench-Corp and AgentThreatBench covering data sharing, prompt injection, approval and ordering, and tenant isolation. If independent deployments confirm near-zero attack success while holding completion near 89%, deterministic flow control will become a procurement requirement for customer-facing and cross-tenant agents. A drop in either metric in production pilots would signal that policy authoring or recovery integration needs work before wider rollout.
