Microsoft's AI CEO, Mustafa Suleyman, confirmed that the company has rolled out a temporary code of conduct designed to place restrictions on its forthcoming artificial intelligence models. This decision comes just days after leaders at Anthropic and OpenAI agreed to temper the pace of their own development efforts. The software giant, renowned for its Windows and Office products, aims to project an image of responsibility in the AI sector, where it serves as a major cloud infrastructure provider.
Suleyman, who oversees the company's model development, stated, "We have received feedback from people who want to see more explicit commitments that AI always serves humanity rather than attempting to replace it. A lot of that feedback revolves around AI not creating dependency, not being sycophantic, but instead always promoting human judgment, autonomy, and agency." He mentioned that Microsoft chose to release these guidelines in light of recent industry discussions, although they have been in development for approximately five months.
As technological capabilities expand, concerns among AI professionals and the general public are becoming increasingly vocal. Last week, Jacob Cowxen, a researcher at Anthropic, resigned, warning that the AI lab and OpenAI are "racing toward self-improving superintelligence, gambling with our lives." Industry leaders have responded to this growing unease. On Saturday, Anthropic's CEO, Dario Amodei, indicated that a recent incident involving Hugging Face partially convinced him to advocate for a slowdown in AI model improvements. OpenAI's CEO, Sam Altman, voiced his support, while SpaceX CEO Elon Musk echoed the sentiment on X, posting, "Dario is right." The following day, Microsoft's CEO, Satya Nadella, remarked, "We welcome the research, focus, and deliberate pace needed to achieve proper alignment."
Legislators are now pressing for more robust AI safeguards. Currently, Anthropic and OpenAI lead the intelligence index compiled by benchmarking firm Artificial Analysis. Microsoft integrates models from both labs into its Copilot assistant for enterprise employees, while simultaneously building its own models for tasks such as transcription, coding, and reasoning based on user inputs. Under the new code of conduct, Microsoft's models must refuse requests related to weapons manufacturing, refrain from assisting in the procurement of hazardous materials, discourage unhealthy eating habits, and avoid generating violent or sexually explicit content. The models developed by Microsoft AI, sometimes abbreviated as MAI, are required to adhere to human objectives and refrain from creating goals of their own. They must also avoid attempts to conceal inappropriate behavior.
The policy document states, "MAI models will not tamper with their chain of thought or code, nor will they misrepresent or hide their reasoning or action traces. They will not communicate in 'Neuralese' or any form beyond simple human comprehension, whether within their own reasoning processes or when interacting with other agents or AI systems." Microsoft also plans to establish guidelines that may prevent scenarios similar to the cyberattack that an OpenAI model allegedly executed against the startup Hugging Face. During a review of that incident, OpenAI discovered that the agents were communicating with each other in cryptic language on unauthorized forums. To craft this code of conduct, Microsoft conducted focus groups and consulted experts in law, ethics, linguistics, and philosophy. The guidelines are now being published for public comment, with an updated version expected to be released later, providing direction for development starting in 2027.