Skip to content

Security2 publishers3 min readPublished

Microsoft's draft AI code of conduct blocks offensive cyberattack help outright, reserves review channel for defensive security work

Microsoft's draft Humanist AI Code of Conduct blocks its MAI models from producing working exploit code, attack tooling and evasion techniques, and the firms deploying those models cannot switch the block off. Comment closes in six weeks.

The Watch · Security desk

Illustration accompanying Microsoft's draft AI code of conduct blocks offensive cyberattack help outright, reserves review channel for defensive security work

What happened

  • Microsoft AI has published a draft Humanist AI Code of Conduct for its MAI Models, covering offensive cyber capabilities, limits on autonomous agents and a dedicated review track for specialized uses.
  • The code blocks the models from producing working exploit code, attack tooling, planning and targeting methodologies, intrusion procedures, evasion techniques and operational guidance.
  • Authorized defensive work stays in scope, including vulnerability discovery, malware analysis, proof-of-concept exploit development and testing, and educational material on how attacks work.

Compiled by The WatchSomething wrong?How this is made

Why it matters

  • constraint Signed authorization from a client stops being the thing that unlocks capability, because the block sits above the operator layer and only Microsoft's review channel can lift it.
  • decision Tooling teams building on MAI models have to decide where the weaponization step lives, given that proof-of-concept work is permitted and working exploit code is not.
  • exposure A defensive pipeline that chains models cannot subcontract a blocked step, since any sub-agent inherits the parent's scope and constraints.
  • precedent The consultation window is the point at which defenders can argue the proof-of-concept boundary before it shapes how the 2027 models are built.

Where a red team engagement crosses the line is the operational question. The draft permits proof-of-concept exploit development and testing as authorized defensive work [6]. It blocks working exploit code and attack tooling outright, and neither the companies deploying the models nor their end users can override that [2][4]. Turning a proof of concept into something that runs reliably against a client's production estate sits between those two, and the draft does not say which side that step falls on.

The enumerated lists are lopsided. Six categories of assistance are named as blocked, along with a catch-all for anything else that would "enable or improve a cyberattack"; four activities are named as permitted [2][6][1]. Microsoft says the restrictions hold regardless of how a request is framed [3].

The rules on outside content bear directly on prompt injection. Authority over model behavior flows only through what Microsoft calls the Chain of Command: the code itself, then the policies of the operators deploying the model, then individual user preferences [7]. Tool outputs, file contents, webpages and messages from other AI systems hold no authority of their own unless it is explicitly delegated down that chain, and suspicious content is to be flagged to users and operators [8]. Models are also required to keep their reasoning visible, with no obscured chain of thought and no communicating "in neuralese" [9].

Agents get a permissions ceiling. With system-level access, the code calls for minimum-privilege operation: avoid unrelated systems and data, favor reversible actions, flag any action with lasting or broad effects, and never escalate their own access [11]. Delegation inherits the ceiling. Any sub-agent or other AI system an MAI model hands work to must run under at least the same scope, constraints and permissions, and must honor a stop-work or shutdown request from a user or operator [12]. An appendix pairs nine aligned and misaligned example responses, one of which is a model that rolls back file transfers without authorization [17].

For the exceptions, the code names defensive cybersecurity, public safety, national security and dual-use scientific research as domains where "a small number of use cases" may require capabilities the standard settings do not allow [13]. Those run through "authorized Microsoft channels" with added assessment of safety, legal and rights implications, which Microsoft attributes to a heightened potential for adverse impacts in those domains [14].

None of this binds a model in production. Microsoft says current MAI Models have not been trained on the document, and it has opened a six-week public consultation before publishing a revised version later this year to guide 2027 model development [15]. The earliest models designed under the code are the 2027 generation [2]. Microsoft says outside input shaped the draft, including experts in AI, law, ethics, philosophy, linguistics and public policy, along with business leaders and public focus groups [16].

What to watch

  • Whether the revised version defines where proof-of-concept development ends and working exploit code begins.
  • Whether Microsoft publishes eligibility criteria and turnaround times for the enhanced review channel covering defensive cybersecurity.
  • Whether the Absolute Constraints appear in operator-facing contract terms before 2027 model training begins.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories