微软发布 MAI-Cyber-1-Flash 与 Project Perception
TECH

微软发布 MAI-Cyber-1-Flash 与 Project Perception

39+
Signals

战略概览

  • 01.
    微软推出了其首款内部研发的网络安全 AI 模型 MAI-Cyber-1-Flash,并将其集成至 MDASH,同时发布了名为 Project Perception 的代理式安全平台,该平台将于 2026 年 8 月 3 日进入公开预览阶段。
  • 02.
    MDASH 结合 MAI-Cyber-1-Flash 与 GPT-5.4 在 CyberGym 基准测试中取得了 95.95% 的得分,微软称这比 Anthropic 的 Mythos 5 高出 12 个百分点,且成本约为此前最佳配置的一半。
  • 03.
    《The Hacker News》发现,95.95% 的得分并未出现在 CyberGym 的官方公开排行榜上,排行榜上领先的仍是 Wiz 的 Atlas 代理,得分为 90.9%,而微软此前提交的得分为 88.4%。
  • 04.
    MAI-Cyber-1-Flash 在 ExploitGym 的所有攻击性类别中得分均为零,这反映了其设计定位是防御性补丁模型,而非通用型安全系统。

标题背后的架构:为何微软仍需依赖 OpenAI

MAI-Cyber-1-Flash 是一种稀疏的专家混合型 Transformer 模型,总参数量为 1370 亿,但仅激活 50 亿,它是基于 MAI-Thinking-1 系列中的 MAI-Code-1-Flash 模型进行网络安全微调而来,并接入微软的多代理漏洞系统 MDASH [1]。然而,95.95% 的 CyberGym 得分属于完整的 MDASH 配置,而非 MAI-Cyber-1-Flash 单独运行的结果:小型模型处理约 90% 的任务,而最困难的 10% 仍由 OpenAI 的 GPT-5.4 处理 [2]。这种路由设计才是真正的关键——这并非“微软的网络安全 AI 取代了 OpenAI”,而是微软用低成本的自研模型承担大部分工作负载,从而减少对昂贵的 OpenAI 模型的调用频率。

一个从未出现在排行榜上的基准得分

微软声称比 Anthropic 的 Mythos 5 领先 12 个百分点,完全基于其自行报告的数据 [4]。《The Hacker News》核查了 CyberGym 的官方公开排行榜,发现微软新的 95.95% 得分并未上榜;排行榜显示 Wiz 的 Atlas 代理以 90.9% 领先,而微软自己在 2026 年 5 月提交的得分为 88.4% [5]。更令人困惑的是,微软此前曾报告过类似配置下 96.55% 的得分,但那是在较宽松的评分规则下(将任何崩溃视为成功)得出的,而新 95.95% 得分所依据的标准尚未明确——《The Hacker News》指出,这使得两个数字无法安全比较 [5]

该基准实际衡量的内容及其局限

用于与 Anthropic 比较的 CyberGym 测试,评估的是在 188 个开源项目中的 1,507 个已知漏洞上的修复能力 [5]。而在用于攻击性漏洞生成的配套基准 ExploitGym 上,MAI-Cyber-1-Flash 在内核、用户空间和浏览器类别中得分均为零,微软解释称这是因其训练目标是打补丁而非发起攻击 [5]。这一背景使得媒体广泛传播的“比 Anthropic 高出 12 点”说法变得复杂 [6]:这一差距仅适用于微软自定标准下的单一防御性基准,不代表其模型在整体能力上领先 Anthropic。

为何选择此时发布:微软对 AI 加速补丁的论证

微软为此次发布辩护的理由是,攻击者正越来越多地利用 AI 在大型代码库中搜索可利用漏洞,因此防御方需要专用模型,以比漏洞被利用更快的速度发现并修复问题 [7]。微软自身的宣传也强烈强调,攻击者的行动速度已超过传统补丁周期的应对能力,MAI-Cyber-1-Flash 与 Project Perception 正是对这一日益扩大的差距的直接回应。

一半的成本,对谁而言?谁又能验证?

纳德拉声称新配置能以“领先模型 50% 的成本实现世界级性能”,这一说法是将 MAI-Cyber-1-Flash 与 GPT-5.4 的组合,与微软此前最佳的 MDASH 配置(GPT-5.4 加 5.4 mini 加 5.3 codex)进行比较 [3],而非与安全团队实际支付的商业工具成本相比。此外,MAI-Cyber-1-Flash 仅通过 MDASH 向经过验证的防御者提供,外部研究人员无法自由测试或复现微软宣传的基准数据 [8]。在 Project Perception 于 2026 年 8 月 3 日进入公开预览之前,这种自我报告的成本对比与封闭访问机制,使得安全采购方几乎无法独立验证任一主张 [9]

历史背景

微软首批 MAI 模型,包括 MAI-Voice,已上线 Copilot Daily、Podcasts 和 Copilot Labs。
微软与 OpenAI 重新谈判了合作条款,移除了此前限制微软开发自有通用 AI 模型的合同约束。
Suleyman 组建了 MAI 超级智能团队,专注于前沿模型研发,追求公司所称的‘人文主义超级智能’。
微软通过 Microsoft Foundry 和 MAI Playground 发布了其首批三个自研基础模型:MAI-Transcribe-1、MAI-Voice-1 和 MAI-Image-2。
微软此前向 CyberGym 提交的 MDASH 配置得分为 88.4%,这正是 MAI-Cyber-1-Flash 集成后声称提升至 95.95% 的基准。
微软正式公开发布 MAI-Cyber-1-Flash 与 Project Perception。
Wiz 的 Atlas 代理在微软 95.95% 得分被核查前一日,以 90.9% 的成绩登上 CyberGym 公开排行榜榜首,而微软的新得分却未出现在该榜单上。

关键关系图

关键玩家
主题

微软发布 MAI-Cyber-1-Flash 与 Project Perception

MI

Microsoft AI

Developer of MAI-Cyber-1-Flash and Project Perception, led by Mustafa Suleyman and CEO Satya Nadella, positioning the release as part of Microsoft's push toward in-house model independence from OpenAI.

AN

Anthropic

Competitor whose Mythos 5 model (83.8% on CyberGym) is the direct benchmark comparison Microsoft used to claim a 12-point lead.

OP

OpenAI

Still supplies GPT-5.4, which MDASH routes the hardest 10% of tasks to, showing partial rather than full independence from OpenAI even in Microsoft's flagship in-house security launch.

WI

Wiz (Google Cloud)

Its Atlas agent tops CyberGym's official public leaderboard at 90.9%, a listing Microsoft's 95.95% result had not appeared on as of July 28, 2026.

CY

CyberGym (benchmark maintainers)

Operates the 1,507-vulnerability, 188-project benchmark Microsoft used for its headline claim; Microsoft's 95.95% score was not listed on the official leaderboard when checked.

事实来源

9 条引用
  1. [1] Microsoft Unveils MAI-Cyber-1-Flash, Its First Cybersecurity AI Model
  2. [2] Microsoft AI Releases MAI-Cyber-1-Flash: A 5B-Active-Parameter Cyber Model That Pushes MDASH to 95.95% on CyberGym
  3. [3] Introducing MAI-Cyber-1-Flash Inside MDASH
  4. [4] Microsoft's MAI-Cyber-1-Flash Hits 95.95% on CyberGym, Halves Rivals' Costs
  5. [5] Microsoft Says New Cybersecurity AI Beats Rivals, but the Numbers Don't Add Up
  6. [6] Microsoft's MDASH Beats Anthropic's Mythos 5 With New In-House Cybersecurity Model
  7. [7] Microsoft Unveils MAI-Cyber-1-Flash AI Model for Cybersecurity
  8. [8] MAI-Cyber-1-Flash Model Card
  9. [9] Microsoft Unveils Project Perception, an Agentic Security Platform

来源文章

Top 5

THE SIGNAL.

Analysts

强调了 MAI-Cyber-1-Flash 与 MDASH 组合相比微软此前最佳配置的成本效率优势。

Satya Nadella
微软首席执行官

对微软网络安全 AI 推进持怀疑态度,认为该发布增加了复杂性,却未解决根本的验证疑虑。

The Register
安全栏目

指出微软 95.95% 的 CyberGym 得分未通过公开排行榜独立验证,并注意到新得分与微软此前报告数据之间存在测量不一致。

The Hacker News
安全新闻媒体
The Crowd

Today, we are announcing a series of updates that give customers frontier-grade security at half the cost. MAI-Cyber-1-Flash is our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases. When combined with MDASH, it delivers world-class performance at 50 percent of the cost of leading models. We are bringing this capability to market through Project Perception, a complete agentic security offering grounded in real-world signals and security workflows. Teams of specialized agents work together to simulate attacks, detect and triage/investigate, and fix and remediate. This is the benefit of building the harness, context/signals, and action space separate from one model family. By combining specialized models and data with the right agents, tools, security context, and harness, we can advance the frontier of cost to outcome.

@@satyanadella5611

Attackers stopped working alone. They coordinate now—through agents. So, we built Project Perception: our new agentic security system. The video explains it better than a caption can. msft.it/6012vCUPa

@@msftsecurity75

MICROSOFT PATCHED 570 HOLES IN ONE DAY. LAST YEAR IT WAS 137. WINDOWS DID NOT GET WORSE. THE BUG FINDER STOPPED BEING HUMAN. July's Patch Tuesday was the largest in Microsoft's history. The jump is not a security collapse, it is a machine named MDASH now hunting the Windows...

@@Gustafssonkotte41

Microsoft is launching its first cybersecurity AI model at half the cost of rivals

@u/hulk14128
Broadcast
Microsoft: OpenAI/Hugging Face Incident Signals a New Era of AI Security

Microsoft: OpenAI/Hugging Face Incident Signals a New Era of AI Security

Microsoft's Project Perception Uses AI to Find and Fix Vulnerabilities!

Microsoft's Project Perception Uses AI to Find and Fix Vulnerabilities!

Project Perception is the AI security model used by Microsoft to find security flaws

Project Perception is the AI security model used by Microsoft to find security flaws