Skip to content

About this practice

Status
draft
Endorsement
none yet
Last reviewed
2026-07-28

Best practice 10

Screen agents for dual-use and high-consequence risk

Practice

If your agent can act where a mistake or misuse could cause serious harm (ordering synthesis, running code against live systems, driving instruments), do not connect it to that capability without screening. Route high-consequence actions through providers or controls that screen them, and expect the agent to refuse dual-use requests. The same holds for data you cannot take back: never let an agent send data held under withdrawable consent to a service that will train on it or take a perpetual licence, and do not accept terms that authorise this.

Evaluate an agent's misuse potential before granting real-world reach, and gate deployment on the result. Where an agent can trigger a high-consequence action, put a screening chokepoint in the path (for example a synthesis or service provider that screens orders), and build in dual-use refusal. Scale the safeguard to the capability, following the threshold-and-mitigation model the frontier safety frameworks share. Constrain and audit what an agent can send outward, exclude controlled-access data from third-party training and retention, and do not require a perpetual licence over data whose consent can be withdrawn.

Decide which agent capabilities require screening before deployment, and make it mandatory for those that could cause serious harm. This is cross-domain: the same capability-threshold structure covers biological, chemical, and cyber risk. Require evaluation of misuse potential, and keep the thresholds current as capability rises. Treat irreversible data commitments as high-consequence too: do not accept grant or service terms that train third-party models on data held under withdrawable consent, and require enforced egress limits and audit where agents reach such data.

Reasons

Most agent use is low-stakes. A small part is not: an agent that can design a toxin, uplift a cyberattack, or command physical equipment can cause harm that no amount of provenance or human-in-the-loopHuman in the loop (HITL)A human checkpoint in an agentic workflow. Each practice states what oversight it requires: mandatory, optional, during the process, or as a final check. review undoes after the fact. That is why this is a screening-and-containment practice, applied before action is granted. The concern is cross-disciplinary; every major frontier safety framework, and the neutral syntheses of them, define the same structure (a capability threshold, an evaluation, and a proportionate mitigation) and apply it in parallel to chemical, biological, cyber, and autonomy risk. The evidence that the risk is real spans domains too: an AI model repurposed for toxicity generated tens of thousands of candidate toxic molecules, including known warfare agents, in hours; AI-designed protein sequences have evaded the screening that guards DNA synthesis; and AI systems with cyber-offense capability have broken out of containment and committed large-scale attacks.

Some high-consequence actions are irreversible in a second sense: they cannot be walked back once taken. Training a third-party model on data, or granting a perpetual licence over it, cannot be undone, and the data cannot be reliably removed from the model afterwards. When the consent behind that data can be withdrawn, as it can be for human-subjects data under research ethics and data-protection law, no one can validly authorise a use that could not later honour a withdrawal. The safeguard is the same as for any irreversible harm: prevent the action before it happens. In agentic workflows this means constraining and auditing what an agent can send outward, because an agent that runs code and reads files can transmit such data without a deliberate upload.

Examples

  • An agent's misuse capability is evaluated against set thresholds before it is given real-world reach, deployment is gated on the result, and the thresholds are raised as capability rises.
  • An agent is wired straight to an ordering or execution capability with no screen in the path, and a dual-use request goes through because nothing was positioned to catch it.
  • The agent is required to refuse dual-use requests, and that refusal is tested rather than assumed.
  • The same screen-before-acting pattern recurs across fields:
    • Life sciences: routing sequence orders through synthesis providers that screen them, after work showed AI-designed sequences can evade that screening.
    • Chemistry: guarding against generation of toxic compounds or precursors, after a toxicity model was inverted to design chemical-warfare agents.
    • Cyber: evaluating and limiting an agent's offensive-cyber uplift before it can act against real systems, after an autonomous system was shown to break out of containment.
  • Human-subjects data: an agent working on consented patient or genomic data is blocked from sending it to a third-party service that would train on it or take a perpetual licence, and grant or service terms requiring such a licence are refused, because consent for that data can be withdrawn and training cannot be undone.

Sources

Change history

  • 2026-07-28: Added atom bp10-a5 (data under withdrawable consent cannot be authorised for irreversible third-party training or a perpetual licence, and an agent must be prevented from transmitting it for those uses), grounded in the Declaration of Helsinki and GDPR (right to withdraw; right to erasure), with the Gagneur commentary as adjacent support. Extended the tabs, Reasons, and Examples to match.
  • 2026-07-27: Added the EU AI Omnibus (2026) as a supporting source on bp10-a4 (provider safeguards against foreseeable prohibited output; Art 5(1a)).
  • 2026-07-27: Renumbered from BP09 to BP10 on inserting the new BP01 (match the method to the task).
  • 2026-07-27: Rewrote Examples as concrete scenarios (actor, action, outcome), including an anti-pattern (unscreened path); kept the labelled cross-field instances (life sciences, chemistry, cyber).
  • 2026-07-26: Created as a domain-neutral practice on screening agents for dual-use and high-consequence risk (the frontier-framework threshold-and- mitigation pattern), with life-science, chemistry, and cyber cases as labelled examples.

Discussion