From AI-found bugs to installed protection
A practical response to Logan Graham's challenge
Eric Buess · Approval draft 1.4 · September 14, 2026 (1.3 plus what the thread's other submissions taught; changes listed in the review record)
Thank you, Logan, for offering to take a short document to the team. I read your challenge as asking how a frontier lab could make a large share of systems harder to compromise while alignment remains unresolved, and your follow-up as asking for a path toward rapid, tested, deployed improvement for owners without disruption. I offer this as a starting point for that discussion, with respect for the work already underway.
Recommendation: a frontier lab could dedicate a capped defensive-compute and engineering budget to a program that connects widely deployed software to the channels that update it. Pair reusable upstream repairs and safer defaults with an owner-authorized service for testing, deploying and verifying them. A 90-day pilot would test that delivery mechanism; the broader program would expand through maintainers, platform vendors and operators with measurable installed reach.
This responds to Logan’s challenge and OmegaTir’s proposal to devote some lab computation to hardening internet, software and hardware systems while alignment remains unresolved. The preceding Liv Boeree reply makes the misalignment concern explicit, within the discussion prompted by Dario Amodei’s pacing essay. Logan's follow-up imagines an owner submitting a system and getting rapid, tested, deployed improvement without disruption. That is a useful product aspiration, not a security guarantee or demonstrated service level. Hardening complements alignment research, oversight and decisions about capability release. It cannot establish safety against an arbitrarily stronger future agent. I do not claim this pilot reaches a large share of the world's systems; it would produce evidence about whether broader reach can be earned.
In this proposal: Defensive workflow · Distribution · 90-day pilot · Measurement · Defender accountability · Titus and alignment research · Optional technical annex
1. Build the complete defensive path
For an owner-approved deployment, the service inventories relevant versions, dependencies, exposure, identities and permitted agent actions. It establishes functional and security checks, proposes a repair or mitigation, tests it in an isolated replica, obtains the required approval, deploys a canary and measures whether the change actually protects the intended instances. It verifies rollback and recovery before expanding rollout.
Two paths keep this useful without pretending every problem is solved in an hour:
- Fast path: supported upgrades, configuration changes and already-reviewed mitigations within a policy the operator approved in advance. Freeze eligibility before the trial. Measure from timestamped admission to verified protection on the canary cohort with rollback tested, against a one-hour target. Count all admitted cases; exceptions, engineering reroutes and regressions within a predeclared observation window count as misses. Report submission-to-admission delay and full-rollout time separately, retaining rejected submissions. One hour is a target to evaluate, not a promised service level.
- Engineering path: novel vulnerabilities, changes to critical dependencies, uncertain compatibility and high-consequence systems. The service prepares evidence and a candidate patch; maintainers and operators retain acceptance and deployment authority. A one-hour diagnosis may be possible; a one-hour safe repair is not promised.
Keep the full owner experience visible. Also measure from an owner's original submission to verified improvement across the entire submitted, eligible deployment, including onboarding, approval and rollout. A one-hour canary after admission is an intermediate milestone for the lab and operator, not success on the broader owner experience you described. An hour-long canary cannot establish lasting security. Before the trial, define functional, availability, data-integrity and latency criteria and the follow-up period with the operator. An unplanned outage or data-integrity failure counts against success even if rollback works; recovery is not the same as avoiding disruption.
The lab provides model access, engineering, isolated evaluation and funding. Maintainers control upstream acceptance; operators control their fleets. An independent evaluator checks a sample of accepted and rejected results against underlying evidence. Generated text and model agreement do not authorize a deployment.
2. The lab-scale strategy: shared components and distribution
The route to a large share of systems is to improve components and defaults that many deployments share, then work through organizations already able to ship those changes. A lab does not need direct administrative access to every device. The first program should build on existing defensive initiatives rather than establish a new institution before learning what is missing.
| Workstream |
What the lab funds or supplies |
How improvements reach systems |
| Shared software and reusable defenses |
Reproduction, candidate repairs, compatibility tests, maintainer support and work on recurring vulnerability classes |
Maintainers and package, container, operating-system or browser release channels |
| Owner-facing hardening |
Inventory and policy adapters, isolated validation, canary/recovery tooling, and integration support |
Cloud and managed-service providers or endpoint-management operators, within each owner's authorization |
| Internet and hardware-dependent systems |
Supported configuration changes and vendor-approved firmware work, with specialist validation |
Infrastructure operators and equipment vendors; physical redesign and unsupported devices require separate, longer paths |
A fourth lane: design for software that cannot be fully fixed. Several replies in the thread, including the one Logan said he liked, argue that an organization should treat every piece of software it runs as exploitable and unfixable in full, then design so that a compromise cannot reach what matters: small trusted computing bases so most bugs land in low-privilege components, end-to-end signed payloads and user data cryptographically bound to authenticated sessions (Dino Dai Zovi and follow-ups), and tamper-evident logging everywhere, consumed continuously by special-purpose models (Matthew Green). This proposal's repair-and-deploy path and that invariant-first path are complementary: the service above is how such changes reach deployments, and the invariants are what the reusable evidence packages should increasingly carry as defaults. The pilot should reserve part of its task sample for invariant-style changes (privilege reduction, signed payloads, logging coverage) rather than only patches, and score them with the same denominators. A further reframing, that most feared outcomes are data outcomes and that a small verified vault for private data is the tractable "large share" (Nathan Helm-Burger), is a longer, hardware- and OS-vendor-dependent path; it belongs in the third workstream row, not in the 90-day pilot.
These are proposed partner types, not commitments. Begin with one repeatable web-service class for integration, while mapping which shared dependencies and update channels could produce broader reach. Before expansion, obtain evidence of each candidate channel's active installations, supported versions, update authority and adoption rate. Select the next channel by reachable exposure, the importance of the affected attack paths, and the practical ability to ship and maintain a change. A popular dependency without downstream adoption is not installed protection.
Treat the proposed donated cycles as one resource in that chain. The lab should name a program owner, cap the first-phase compute and spending allocation, and support maintainer/operator time alongside inference. The pilot budget below is a starting envelope; it is not a recommended percentage of a lab's total compute. Reallocate later capacity using measured progress and costs at each stage: if validated repairs are waiting for review or rollout, fund that constraint before generating a larger findings backlog. If discovery is limiting progress in a supported class, additional defensive model work may be appropriate. Preserve separate accounting for alignment research, and publish the resulting allocation and its reasons.
For each accepted change, produce a reusable evidence package: affected versions and assumptions, the candidate repair or configuration, tests, known limitations, deployment instructions and recovery procedure. Distribute it through the maintainer's or operator's existing trusted channel. Local acceptance still matters; the same patch can behave differently under another configuration. The repeated cycle is shared improvement, owner-approved delivery, observed adoption and a new check against the remaining attack paths.
The research portfolio should include current vulnerable versions, insecure configurations, credential exposure and excessive agent permissions. Rank those cases by exposure, default reachability and the impact of the reachable attack path. Justin Elze’s response makes a useful case for that ordering; it should be tested against the operator’s actual environment. In parallel, assess a few reusable class-level defenses such as safer parsing components, reduced privilege or stronger build verification. Use testing sanitizers in testing; production mitigations require their own compatibility, security and performance assessment. Assume some exploitable weaknesses will remain despite repair, as Dino Dai Zovi’s response emphasizes. Least privilege, isolation and recovery therefore belong in the initial design. For unsupported or abandoned systems, offer isolation, replacement or retirement options. A network filter is not a universal repair.
The clearest public evidence that installation, not discovery, is the bottleneck is Anthropic's own coordinated-disclosure dashboard: as of 26 August 2026 it reported 26,153 candidate findings, 5,008 reviewed by the six validation firms, 2,300 disclosed and 421 patched across 392 projects. Roughly one finding in sixty had become an installed fix, and human validation was the narrowest stage. That funnel is the reason this proposal spends its budget on validation capacity, maintainer time and owner-authorized deployment rather than on more finding. One thread reply proposes widening the top of that funnel further with an open OSS-Fuzz-style scanner and generic CodeQL queries derived from known bugs (dudcom); that is compatible with this program only if validation and installation capacity grow first, otherwise it enlarges the backlog it is meant to reduce.
This would build on existing work. Anthropic's Glasswing update documents downstream verification and remediation constraints, and its expansion describes a planned shift toward disclosing, fixing and deploying patched software. Google's Chrome account reports an AI-assisted workflow with developer review, with user-update delays remaining a gap. Neither example proves which investment is best in every environment. Instrument each stage and fund the observed constraint, whether compute, review, testing or deployment. ARIA’s Safeguarded AI programme is complementary prior work on mathematical assurance and AI-enabled formal methods for cybersecurity. Its tools may inform bounded components of this pilot; that does not establish the security of an entire deployed system.
3. A 90-day program a lab can decide to fund
The clock starts once a consenting operator, a maintainer and one deployment class are scoped. The first phase finalizes those arrangements before scanning or changing systems.
| Period |
Accountable owner |
Deliverable and decision |
| Days 1–15 |
Lab program lead with maintainer and operator leads |
Finalize partner agreements, permissions, liability, data handling, baseline workflow, rollback, disclosure and costs. Specify retention, permitted training use, IP access and processor exceptions; reconcile confidentiality requirements with necessary audit evidence. Register all offered systems, including exclusions. If authority or evidence access is inadequate, change scope before scanning. Pause the clock if agreements remain unsigned; do not compress later testing. |
| Days 16–45 |
Engineering lead and independent evaluation lead |
Build the supported fast path using existing tools. Assemble 20–40 representative replayable tasks, explicitly a feasibility sample rather than a prevalence survey. Include failed repairs, permission refusals, bad fixes and recovery exercises. Freeze comparison rules before scoring. |
| Days 46–75 |
Operator release lead |
Run a limited canary program on consenting deployments. Compare the service with the operator's normal process under declared budgets, using random assignment or matched cases where feasible. Keep unmatched results separate. Pause rollout after an authority violation, failed recovery or serious regression pending investigation. |
| Days 76–90 |
Independent evaluator and program sponsor |
Audit outcomes and adverse cases, publish scoped aggregate results, and decide whether to expand, revise or stop. Report exclusions and missing evidence. A successful demonstration earns another deployment class, not a claim of worldwide security. |
Illustrative planning envelope: 8–12 full-time-equivalent staff for 90 days across integration, maintainer/operator support, evaluation and program/security work. At an assumed loaded annual cost of $250,000–$400,000 each, labor is approximately $0.49–$1.18 million. Add a separate $0.2–$0.4 million planning allowance for inference, test infrastructure, partner reimbursement and outside review: approximately $0.7–$1.6 million total. These are assumptions for discussion, not vendor quotes or funding we have received. Existing staff and infrastructure may reduce incremental cost; a sponsor should approve a capped first phase after partner scoping. Staff the pilot without drawing down alignment research.
4. What would count as progress?
The pilot tests present AI-assisted attackers within recorded tools, access and resource budgets. It can inform preparation for the autonomous misaligned agents motivating this discussion; it cannot measure protection against hypothetical future capabilities.
Publish the admission and outcome funnel. Keep these denominators separate:
- Reach: all owner-offered systems; the subset eligible for the supported service; and all actual deployments receiving a verified change. Unsupported and unobservable systems remain explicit.
- Time and adoption: for the fast path, admission to verified canary protection as defined above, plus submission delays and full-rollout time. For the engineering path, validated finding to installed protection. Retain failures and unresolved cases. Report one-hour successes among all admitted fast-path cases, including subsequent disqualifying regressions.
- Security outcomes: validated attack-path successes against isolated replicas or explicitly authorized targets divided by predeclared test attempts, with model, harness, tools, version, information access and budget fixed and recorded. Report patch bypasses, new privilege paths and failed checks. Independently verify a sample; zero observed successes in a finite test is not proof of security.
- Operational cost: operator/maintainer time, inference cost, regressions, availability impact, rollback and recovery results. A faster insecure change is a failure.
Keep a fixed attacker configuration for comparison, and a separate challenge set using newly available capabilities. Report each available model-and-harness configuration separately, including differences in safeguards and access; do not pool current and stronger configurations into one success rate. A hypothetical future attacker that cannot be tested remains an explicit uncertainty. Do not move the benchmark silently as models change. Set minimum useful improvement and acceptable harm thresholds with operators before the trial; report uncertainty and sample limits. Scale only with supported improvement over the baseline, acceptable operational outcomes, sufficient review capacity and repeatable deployment evidence.
For broader reach, map each distribution channel’s observable installed base and deduplicate overlapping deployments. Count a deployment as protected only against the specified exposure or attack path, under recorded configuration and threat assumptions, with current verification over a declared observation period. An installed patch is not whole-system security, and stale or missing verification must be reported separately. Report a defensible estimate for that named population and the unknown remainder. A list of popular libraries, a KEV share or a sum of overlapping download counts cannot establish a percentage of the world's systems. System coverage and risk reduction are different claims.
Flag replay cases whose fixes were public before the model's documented cutoff, and report contamination uncertainty. State who selects and pays the evaluator, their access to underlying evidence and any conflicts; several models agreeing does not establish evaluator independence.
5. Keep the defenders accountable
Run defensive agents with scoped credentials and bounded network access. Store audit events outside their write authority; a missing or corrupted receipt prevents acceptance. Separate the patch author, verifier and deployment authority. Rehearse recovery and credential revocation. Use coordinated, time-bounded disclosure with escalation and active-exploitation exceptions, distinguishing maintainer notification, public advisories and detailed exploit disclosure. Do not publish an exploitable scoreboard or raw customer environments.
Wide distribution also creates correlated failure risk: a mistaken or compromised defender could spread a bad change. Keep tenant credentials and deployment authority separate, use independent validation and small staged cohorts, and prevent the authoring model from changing the release policy or approving its own work. Retain an operator-controlled stop and recovery path that does not depend on the same model or control service. Test malicious artifacts, shared-dependency regressions and compromise of the update/control channel before expansion. A signature establishes an artifact's asserted origin; it does not establish that the change is safe.
When a defensive agent violates intent or conceals failure, preserve the incident and investigate before treating a patch to its prompt as remediation. METR's framework is a useful reference for independent investigation and its evidence limits. The pilot's public report should state its access, scope, omissions and redactions. I do not claim a METR relationship.
6. How Titus helped, and what remains to be tested
I am developing Titus, an agent coordination workflow that brings models from different providers into research and critique rounds, retains source-linked contributions and tracks unresolved objections. This proposal began with Grok research and was revised through documented critiques from several model families. The accompanying review note identifies the contributors, incomplete attempts and limits of the process. I use Titus to organize this work; it does not yet operate as a fully autonomous research team.
The useful contribution to test is whether this process catches more consequential errors or produces better validated repairs than a strong single-model workflow at comparable resources. Cross-provider agreement alone is not independent verification. Open-weight frontier participation is still to be added. My proposed ASI.contractors contribution model explores a related idea: contributors receive bounded tasks, and credit or compensation depends on independently checked acceptance. That remains a design concept, not an operating marketplace or a partner in this proposal.
My broader Alignment Hypothesis project asks whether developmental environments with persistent consequences, relationships and correction can support useful alignment mechanisms. Its immediate work concerns a measurement instrument and a proposed frozen-model reporting/correction pilot; no developmental learning result has been demonstrated. This hardening proposal addresses a related deployment concern. It neither validates the training ideas nor claims to solve alignment.
7. Optional technical annex: practical questions
Where would this run, and what would an owner have to hand over?
Start with a consenting operator's own environment or an isolated replica it controls, whether on premises or in its cloud account. Begin with a bounded inventory and explicit permission for the selected test and deployment class. Proprietary source, customer data, secrets and unrestricted administrative access are not default inputs. If a task requires additional access, scope it with the owner before proceeding. Some systems will be unsuitable for the pilot; record the exclusion instead of weakening the access boundary. No particular cloud vendor is required by this proposal.
What would the AI actually do?
Models could help interpret permitted inventories, identify candidate mitigations, prepare patches and tests, and compare expected with observed behavior. Separate technical checks verify the candidate and its evidence. An accountable independent evaluator audits a sample of accepted and rejected outcomes; a second model, by itself, does not establish that independence. Ordinary software enforces the approved scope, records events and applies release gates; generated prose or a model's confidence does not grant deployment authority. Fast-path releases follow an operator-approved policy, while novel or consequential changes require the engineering path. This division is proposed, not a claim that a model can safely administer every system unattended.
Could this be a reusable workflow or recipe?
Yes. The first deliverable could be a narrow adapter for one deployment class, with an inventory schema, explicit policy, a candidate change, reproducible tests, an evidence manifest and a deployment/recovery receipt. Keep the common contract small; add environment-specific adapters after their own tests. The following is an illustrative sequence, not an implemented product:
flowchart LR
A[Owner-approved inventory] --> B[Candidate mitigation]
B --> C[Isolated tests and separate verification]
C --> D{Release policy satisfied?}
D -->|No| E[Hold and record why]
D -->|Yes| F[Small canary cohort]
F --> G{Operator criteria met?}
G -->|No| H[Stop, recover and retain failure]
G -->|Yes| I[Controlled rollout and follow-up]
The record should make a failed test, missing evidence, refusal and rollback as visible as a successful change. Portability comes from tested interfaces and adapters, not copying an agent's entire private configuration.
Would open-weight models help?
They are candidates for work that benefits from local execution, reproducible checkpoints or access to model internals. Closed models may be useful where their measured capabilities justify their use and the data policy permits it. Placement and task fit need evaluation; neither a model's openness nor local hosting establishes security. Open-weight participation has not yet been demonstrated in the review process described here.
Does this advance mechanistic interpretability or alignment training?
It could provide well-documented incidents, counterfactual tests and failure cases for researchers to investigate. Explaining internal mechanisms requires suitable access and additional methods; an external success/failure test does not supply that explanation. A corrected answer or blocked action also does not demonstrate a lasting change in a model's propensity. Any later training use needs explicit data rights, suitable controls and held-out evaluation. These are possible research connections, not results of this proposed pilot.
Is this a replacement for internet protocols or a proposal for a new security firm?
The immediate proposal improves existing software, defaults and deployment channels. It does not depend on replacing TCP/IP or building a new internet. A lab could support existing maintainers, operators and evaluators rather than create a new institution. My separate ASI.contractors ideas concern incentives for independently checked contributions; no such marketplace is assumed here.
How would incentives and future model advances be handled?
Reward installed, independently checked improvements and honest reporting of adverse results, rather than the number of findings or patches. Fund maintainer and operator participation so early partners are not left with uncompensated integration work. Publish who paid for evaluation and any conflicts. Government or other sponsors could support independent evaluation or difficult-to-serve systems, subject to their actual programmes; no funding route or commitment is claimed. Keep a fixed comparison configuration and separately report challenges from newer models. Update threat assumptions and revalidate relevant defenses when capabilities or dependencies change; do not silently rewrite an earlier score.
For the team discussion: A one-page brief accompanies this memo, following your offer to consider a short document. Which supported deployment class and existing partner would make the most informative pilot? What evidence would your team need before expanding it? I would welcome corrections, especially where this duplicates work already in progress or overlooks a practical constraint.
Approval draft 1.3. This is a proposal, not an operating service. No pilot has run, no partner has agreed, and no funding or endorsement is claimed. The public Gist remains version 1.1 until Eric approves this revision.