AI defense · Wed Oct 7, 2026

ASI.blue: defensive AI work, checked across labs

Eric Buess

AI defense

asi.blue

asi.blue

1. AI defense · Eric Buess

  • Defensive AI work, checked across labs
  • I’m Eric Buess. I hold a master’s in computer science, and I’m a member of Anthropic’s model safety bug bounty program on HackerOne. I build with AI every day, and what I’m building toward is defensive: securing AI systems — agents that run tools and touch real data — on systems whose owners authorize the work in writing.

2. The system as it exists today

  • I run a private system of AI agents from each frontier lab. That includes every frontier closed-source lab, the open-source labs (Mistral among them), the coding agents near the frontier, and local models.
  • When I ask for a change, an agent running one lab’s model writes it. Before the change lands, depending on its complexity, one or two agents running other labs’ models check it in their own sandbox. The findings and proposed fix go back to the owner, who decides whether to install it.
  • For signed, verified public work anyone can review, see How the record will be kept below.

3. The work it is aimed at

  • Three kinds of work I’m aiming at, drawn from my public proposals.

4. Audits of AI agents

  • Bounded audits of an AI agent for prompt injection and data leaks: its prompt boundaries, the tools it can run, and the logs it keeps.

5. Independent re-checks of claimed defenses

  • Re-running a claimed guardrail or a claimed fix against a frozen set of tests, then signing a receipt that says exactly what was re-run, by whom, and what happened — pass or fail.
  • The signed receipt is the deliverable: it makes the result checkable.

6. Hardened agent setups

  • Building or repairing an agent setup so it runs in isolation, handles secrets carefully, and keeps tamper-resistant logs — measured against a test agreed up front.
  • A specific problem closed and tested.

7. Defender programs

  • I plan to apply to the labs’ defender programs, beginning with Anthropic’s Defense Access. I would use the access for secure coding and for checking and validating vulnerabilities. That work would also help harden my own system.

8. How the record will be kept

  • Signed, verified public work anyone can review is coming soon and may already be live when you read this. Each finished piece is meant to go on a public build log: what was asked, what was built, and which lab checked it.

9. The thinking behind it

  • I’m writing The Alignment Hypothesis, an early draft about AI alignment. I’m exploring how alignment may be developed through lived moral experience.
  • The paper asks how alignment forms through tested experience, and trust earned under pressure joins that question to the work I’m building toward, which would test AI systems against real attacks with another lab checking each result.

10. From AI-found bugs to installed protection

11. ASI.contractors, release v1

  • A proposal from the same day for a board of bounded safety tasks with checkable receipts. I plan for other builders, each on their own plan and account, to take on these tasks. I expect that to bring more safety work than a niche product could.
  • Read the ASI.contractors proposal

12. How the sites fit together

  • I’m bringing the work exchange and verified public record together at si.build, with ASI.contractors as its safety-task section. asi.blue is my page for the Anthropic application, linking the two proposals and si.build.
  • si.build: the system, work exchange and public record