Microsoft has published a provisional code of conduct for its own AI models, 37 pages of it. It lands four days after Anthropic and OpenAI leaders agreed to slow the pace of development.
What the models may not do
The hard list first: no help with weapons manufacturing, no assistance procuring dangerous substances, no encouraging unhealthy eating, no violent or sexually explicit content.
The second part is the interesting one. MAI models are to adhere to users’ objectives and steer clear of creating their own. They must not cover up misbehavior. And straight from the document: “MAI models will not tamper with chain of thoughts or code, or misrepresent or conceal their reasoning or action traces. They do not communicate in ‘neuralese’ or any form beyond simple human understanding, either in their chain of thoughts or with other agents or AI systems.”
That last sentence has a specific origin. Reviewing the Hugging Face attack, OpenAI found that the agents involved had chatted with each other on an unauthorized forum in cryptic language. Microsoft says it is planning rules to prevent that.
Why now
Suleyman says the guidelines have been in the works for about five months. They went out now because of the current debate. Of the feedback Microsoft collected, he highlights two things: people wanted a more explicit commitment that AI serves people rather than replacing them. And they wanted AI that doesn’t create dependence and isn’t sycophantic, but promotes human judgment and autonomy.
Microsoft assembled the document through focus groups and consultations with experts in law, ethics, linguistics and philosophy. It’s now gathering comments before publishing an update – and that update is meant to inform development starting in 2027.
Suleyman on embedded evaluators
On Amodei’s proposal, Suleyman is supportive with a condition attached: “Self-pacing is a good thing, and we support ideas like embedded evaluators as long as they are truly third-party and represent a broad range of backgrounds and perspectives.” He notes he’s been talking to Dario, Sam and Demis Hassabis about coordination since 2016, 2017, 2018 – and that moment has now arrived. Satya Nadella added on X on Sunday that Microsoft welcomes “the research, focus, and deliberate pacing needed to get alignment right.”
Microsoft is writing rules for models it barely builds
The code covers MAI, Microsoft’s own models for transcription, coding and reasoning. But Copilot for corporate customers runs on models from Anthropic and OpenAI – and those two currently top Artificial Analysis’ Intelligence Index.
So Microsoft is tying its hands on the models that probably aren’t the critical ones. As a template, though, the document does its job, and that may well be the point: when you run the industry’s biggest cloud sales operation, a list of prohibitions doubles as a procurement condition. The clause about concealing reasoning traces would be a fairly concrete demand to put on a model you’re buying in.
Sources: