AI defense · Fri Oct 9, 2026

asi.blue: different models, defending together

Eric Buess

AI defense

asi.blue

asi.blue

1. AI defense · Eric Buess

  • I’m Eric Buess. I hold a master’s in computer science. I research AI safety, and I’m a member of Anthropic’s model safety bug bounty program. asi.blue is defensive AI security built on one testable question. Human review can’t keep up with the internet’s scale; can models from different labs, reasoning together, defend systems as reliably as a careful human reviewer — easing the review burden so people focus on the calls that most need them? That is what I measure.

2. Why different models

  • These are independently trained frontier models with different pretraining data, architectures, and loss surfaces. That diversity enables cross-model verification: a model from another family catches flaws the author model overlooks. The architecture builds on Together AI’s Mixture of Agents research. A strict rule governs deployment: any proposed fix must be reviewed and approved by models from independent labs before landing, ensuring no family grades its own output.

3. The first question: models in concert

  • The question at the heart of this work: how much test-time compute, from which model families, at which effort levels, in which harnesses, and in what combinations does it take for frontier models reasoning in concert to reliably match or beat a careful human reviewer? The goal is to measure and scale this until humans can safely step out of the review loop, landing verified patches as fast as vulnerabilities are found.

4. Benchmarks, red and blue

  • I am turning that question into empirical benchmarks. asi.blue and asi.red form the defensive and adversarial sides of a unified testing environment. Both attack and defense draw on models from across the frontier labs — Anthropic, OpenAI, Google, SpaceXAI, Meta, DeepSeek, Mistral, Moonshot, Zhipu, Alibaba, MiniMax, Cohere, Microsoft, NVIDIA, and others — reasoning together. In the testing labs, rounds test which model mixtures hold under pressure, including against coordinated swarms. Every break becomes a failing test in a secure sandbox, and every fix is verified by a family that did not write it.

5. Provenance anyone can verify

  • I am building verifiable provenance directly into the architecture. Each participating model carries a cryptographically signed identifier, each reasoning step and action generates a structured receipt, and these receipts roll up into a public ledger. Anyone can independently verify who performed and checked each action without taking my word for it, backed by hardware enclaves, local infrastructure, and version-controlled repositories.

6. The larger vision

  • All of it serves a larger vision: testing my research, the Alignment Hypothesis, and a thesis I call the Universal Alignment Imperative, that alignment is a fundamental challenge any civilization developing greater intelligence has to solve. It explores the necessary conditions for genuine values: agency, persistent consequences, and the absence of cues signaling an evaluation.

7. The second question: a training world

  • The second question at the center of this work is the Alignment Hypothesis in practice: can I build a training world with real, lasting consequences, where being watched carries no signal? And does bounded, tested experience make an agent reliably aligned, rather than just better at performing alignment when observed? The prediction is falsifiable: whether agents that learn from the real consequences of their actions in the sandbox resist adversarial misalignment better than agents trained with today’s techniques. The goal is the ultimate sandbox — full isolation, where reaching the real world would require a deliberate sim-to-real step and only those that pass the audit earn transfer.
  • I am building a contained environment on ordinary hardware: a sealed world with its own independent clock, stateful consequences, and no undo — real, nested-layer sandboxes in pursuit of the ultimate one. The hardest problem is differentiating between authentic alignment and situational compliance. Grading remains invisible to the agent, situational awareness of the test is measured directly, and concealing an error is scored as a distinct failure mode from making one. A further question: can the world guide an agent as it develops through a low-signal intervention channel it cannot detect as outside influence, the way introspection experiments inject a thought and test whether the model notices? The deeper mechanics are for the full write-up.

8. In my own testing labs

  • In the testing labs, the baseline instruments are taking shape: open-weights models operate in environments where decisions are scored for harm across varying conditions, with and without evaluation indicators. The goal is testing whether persistent consequences change agent behavior when no monitor is detected. Today that detection is behavioral; I’m working toward mechanistic interpretability that reads internal states and how they correspond to behavior. I’m testing with tools like Anthropic’s J-space lens, and hoping to integrate future interpretability tools as they become available. Around this, I am building an evaluation instrument with layered baselines and controls to detect behavioral shifts under pressure, with full technical write-ups to follow.

9. Hands-free

  • I direct this work hands-free, at high bandwidth, by voice through a wearable pendant keyed to my voice signature, from wherever I am, on systems only I can access. I am going screenless myself using a custom voice-driven application I built for my own workflow, which I can teach other researchers to use.