Microsoft said MDASH with MAI-Cyber-1-Flash received a 96 percent score on CyberGYM, a standard benchmark test. The rating is 12 points higher than Anthropic’s Mythos and also beats Google Gemini and OpenAI GPT. The new MDASH costs half as much to use as the previous MDASH offering.
The second tool Microsoft announced on Monday is named Project Perception. It too is a collection of specialized AI agents that perform red-, blue-, and green-team functions for finding vulnerabilities, investigating them to determine their risk, and taking corrective actions, respectively. Microsoft said the platform selects the models to use based on the assigned task. Considerations that go into the decision include the model’s effectiveness and the end cost to the customer. Microsoft said the decisions are shaped by “ongoing research, benchmarking and evaluation across frontier and specialized models.”
Microsoft said Project Perception is designed to perform 90 percent of tasks for lower costs than similar platforms from competitors. That means customers can turn to the more expensive alternatives only for the remaining 10 percent of tasks.
Microsoft said the new tools respond to a seismic shift in how organizations secure their networks against catastrophic hacks.
“As AI accelerates the speed and scale of cyberattacks, defenders are being asked to secure increasingly complex digital environments with approaches built for a different era,” the company said. “Security teams are often forced to piece together signals, context, and risk insights across vast amounts of data, making it harder to keep pace with emerging threats.”
With last week’s OpenAI incident evoking troubling scenes straight out of the most dystopian sci-fi novels, the tools, which are currently in preview mode, deserve a healthy dose of caution that Microsoft made no mention of. They should be closely scrutinized and evaluated before being used in production. On the other hand, there are clear risks for not adopting such tools. Balancing the risks of using AI agents versus the threat of avoiding them is a work in progress with no clear answers for now.
Leave a Reply