Clarity · Edition

The Board Room

Sunday, May 24, 202636 sources · 9 min read

The Signal

AISI confirmed this week that Anthropic's Mythos became the first AI model to achieve

Simultaneously, TrustedSec demonstrated that AI reduces commercial EDR reverse engineering from weeks to days across all five major products tested, and exploit weaponization windows have collapsed to 4 hours.

Key intelligence

  1. 01

    AI Cyber Capability Crosses Full-Takeover Threshold

    AISI confirms Mythos cleared both simulated attack ranges (first model ever). EDR products now transparent via AI in days, not weeks. Exploit weaponization compressed to 4 hours. OpenAI's Daybreak launched with 8 major security vendors. The defensive model assumed obscurity bought time. That time is gone.

  2. 02

    Execution Layer War: Who Owns Where Agents Act

    SAP (€100M fund + Knowledge Graph), ServiceNow (MCP-based Action Fabric), Apple (agent App Store gatekeeper), and Google (Gemini Intelligence on 3B+ devices) all moved this week to claim the layer where AI agents commit writes. MCP is emerging as de facto standard. The platform that hosts the agent captures the margin.

  3. 03

    AI Infrastructure Financialized: Compute Locked, Power Platformized

    xAI leasing 220K GPUs (45% of Colossus) to Anthropic signals compute is now a financial instrument. Fervo Energy IPO at $10B+ (33% pop) validates power as platform. Microsoft's $100B OpenAI commitment is the floor for frontier play. Nebius at 4:1 demand ratio. Capacity for 2027 is being pre-sold now.

  4. 04

    AI Liability Regime Being Written — In Courts, Not Congress

    a16z published the most comprehensive AI liability lobbying blueprint while courts are actively deciding cases that could impose developer liability for user misuse. ODNI vs Commerce battle over pre-release model evaluation will resolve in quarters. Open-source AI strategy is directly threatened by developer-liability outcomes. The 18-month window to influence defaults is open.

  5. 05

    Enterprise AI Governance Vacuum: Budgets Blown, Foundations Missing

    ServiceNow blew its full-year Anthropic budget by May. 85% of organizations spending millions on agentic AI lack adequate data foundations. Duolingo walked back blanket AI mandate after discovering 20% 'slop tax.' The pattern is clear: adoption velocity is outrunning governance, cost controls, and quality management.

Deep dives

  1. 01

    AI Cyber: From 81% Hack Rate to Full Network Takeover in One Week

    The Capability Jump

    Tuesday's briefing put AI offensive capability at 81% success on individual tasks. This week's data is a category change, not another tick on the curve. AISI confirmed Anthropic's Mythos as the first model to clear both simulated attack ranges end-to-end. Full network takeover, not persistence and not lateral movement. OpenAI's GPT-5.5-cyber completed one of the two. The trend line in AI cyber task completion was already doubling every few months. These results broke above that trend.

    The security model of the defensive stack was built on the premise that the cost of understanding the agent exceeded the value of bypassing it for most adversaries. That premise is no longer true for a growing share of the threat population.

    EDR Is Now a Glass Box

    TrustedSec ran LLMs against five commercial EDR products and found all five share the same architectural pattern: YARA-style rules, behavioral logic, allowlists, prefilters, Lua scripting engines readable after a single decryption pass, and local ML classifiers. Work that took a skilled reverse engineer weeks now takes days with AI assistance. The entire endpoint security category has been running on security-through-obscurity. The obscurity just left the building.

    The Exploit Window Has Collapsed

    PraisonAI was actively exploited 4 hours after disclosure. Microsoft's MDASH system found 16 exploitable flaws in a single Patch Tuesday cycle using multi-model AI analysis. An 18-year undetected RCE in NGINX's rewrite module is the other half of the story: foundational infrastructure harbors defects nobody audited. AI-powered vulnerability discovery, with Mozilla finding 271 real bugs in Firefox using custom harnesses, is compounding the problem from both sides at once.


    The Market Response

    OpenAI's Daybreak launched with CrowdStrike, Palo Alto Networks, Cisco, Cloudflare, Oracle, Zscaler, Akamai, and Fortinet as partners. Congress is holding closed-door Mythos demos, with access routed through NSA over CISA. That ordering prioritizes offensive and intelligence operations over civilian defense, which means the private sector is on its own for several years. Security equities are up 20% YTD. The market is starting to price this in and is almost certainly underpricing it.

    The Foxconn Compound

    Nitrogen ransomware exfiltrated 8TB of confidential designs from Apple, Intel, Google, and Nvidia through a single contract manufacturer. Separately, a Raspberry Pi honeypot dressed as an AI stack was indexed by Shodan in 3 hours and absorbed 113,000+ attacks in a month. A reasonable skeptic would say these are two different stories. They are not. Supply chain custody risk and AI infrastructure exposure are the same asset class under two custody regimes, and both are being probed at machine speed.

    What to do

    1. Commission red team exercise targeting your EDR with AI-assisted reverse engineering within 30 days

      NowAll five major EDR products share identical patterns — your detection gap is larger than your last pen test suggested
    2. Rewrite patch SLAs to 72-hour maximum for internet-facing critical vulnerabilities by end of Q3

      This sprint4-hour exploit windows mean current multi-day SLAs are the exposure window, not the response window
    3. Audit all AI infrastructure tooling (LiteLLM, Ollama, model registries) for security posture — many deployed without review

      NowCISA KEV additions confirm AI routing infrastructure is being actively exploited in the wild
    4. Evaluate AI-powered defensive vulnerability scanning for your own codebase before adversaries operationalize the same capability

      This sprintMozilla's 271 bugs with custom harness vs. curl's 1 bug with generic scanning proves the moat is in orchestration, not model access
  2. 02

    The Execution Layer War: Four Platforms Claim Where Agents Act

    The Collision

    Four of the largest platforms in enterprise software moved in the same week to own the surface where AI agents commit writes, which is the execution layer between intelligence and action. SAP, ServiceNow, Apple, and Google did not coordinate. They did not need to. When four incumbents announce the same thing in the same quarter, the announcement is the market telling you the UI-centric era is closing. The consequence for any technology leader is that the product story written eighteen months ago does not survive contact with this quarter.

    Agents that act across finance, HR, IT, and procurement need one authoritative place to reconcile state. Two authoritative places is zero authoritative places.

    The Architectural Split

    ServiceNow adopted MCP (Model Context Protocol) servers as the standard for its headless Action Fabric, which is what an incumbent does when it wants the rest of the ecosystem to settle the protocol fight on its preferred terms. SAP is playing the opposite hand, building a vertically integrated Knowledge Graph behind a €100M fund whose entire purpose is to make SAP's own agents contextually superior inside SAP's own data universe. A reasonable skeptic would say one of these companies has read the room and the other has not. The reasonable skeptic is being too tidy. These are competing theories of where value sits in the agent economy, open interoperability vs. data-moat integration, and both have worked before in adjacent fights.

    Platform Gatekeepers Move

    Apple is inserting itself at the agent layer on iOS, specifically at the case where agents 'spin up smaller apps on the spot after Apple has already approved the parent app.' That sentence is the entire App Store rent-extraction model translated into the agent era, which is to say it is not new behavior, just a new surface. Google's Gemini Intelligence ships this summer across more than three billion devices, which converts 97%+ market share in key markets into a default that competitors will spend years trying to dislodge. The app stops being the product, and the agent becomes the interface.

    Where Value Migrates

    LayerOld ModelNew Model
    InterfaceApp UI (you own)Agent surface (platform owns)
    OrchestrationIntegration middlewareMCP / Knowledge Graph
    System of RecordWhere work happensDatabase agents query
    PricingPer-seatPer-action / per-outcome

    The a16z TAM Claim

    Andreessen Horowitz staked a public position this week that $150B+ of GTM value migrates from the traditional CRM toward the AI orchestration layer. The Lemkin numbers are the actual argument: 80% fewer human seats and 83% higher total spend with 20+ agents in the loop. Seat-based pricing is structurally breaking at the same time per-customer revenue is going up, which is the pattern that usually takes pricing committees four quarters to admit is real. The CRM stops being where the work happens. It becomes where the work is recorded.

    Vercel's production telemetry already shows agentic workloads at 59% of all token volume, which is not a forecast, it is a Tuesday. Products without agent-compatible APIs get bypassed the way sites outside the default search index got bypassed in 2005, and most of those sites spent the next three years convinced they were a special case.

    What to do

    1. Conduct an agent-readiness audit of your product architecture by end of Q3 — can third-party AI agents discover, invoke, and orchestrate your workflows without a human UI?

      This sprint59% of token volume is already agentic; products that aren't agent-addressable are being routed around now, not in 2027
    2. Evaluate MCP as a strategic standard for your platform roadmap within 60 days

      This sprintServiceNow's adoption of MCP as the Action Fabric communication layer is pulling enterprise ecosystem toward that protocol
    3. Model per-action/per-outcome pricing scenarios and pilot with 3-5 customers this quarter

      This quarterSeat-based revenue is directly threatened when agents replace human seat consumption — model this before your largest customers ask
    4. Map your Apple iOS agent distribution exposure before WWDC and model fee/approval structure into unit economics

      This sprintApple is building governance that prevents agents from routing around the 30% tax — consumer-facing agent roadmaps need this priced in
  3. 03

    Compute Gets Financialized: xAI Concedes, Power Becomes Platform

    The xAI Signal

    Elon Musk, who publicly called Anthropic "misanthropic and evil," has agreed to lease them 220,000 GPUs — 45% of Colossus 1. The financial logic has beaten the competitive logic. Grok never achieved meaningful traction and trails open-source models in developer surveys, and the lease revenue almost certainly exceeds what Grok could generate from the same silicon. A reasonable skeptic would call this a one-off. The skeptic should explain why one of five frontier labs is willing to financialize its infrastructure at all. The population of viable frontier labs is contracting, and the excess is moving onto the lease market.

    GPU supply has become a financial instrument first and a strategic moat second. That is good news for infrastructure buyers who want optionality and bad news for anyone who told a board last year that vertical integration of compute was the durable advantage.

    Power as Platform Business

    Fervo Energy's IPO at a $10B+ valuation, with shares jumping 33% above an already-raised price target, prices AI power supply as a platform business rather than a utility commodity. The more useful number is Google's option for 3 gigawatts against only 658MW currently contracted. At roughly 50MW per large data center, 3GW is sixty-plus facilities from a single supplier. Power contracts signed this year set competitive position in 2028 to 2030.

    The Lockup Accelerates

    Court filings put Microsoft's commitment at $100B+ to OpenAI by June 2026, against $30B in direct revenue. OpenAI separately committed $20B to Cerebras, underwriting its $56B IPO, which popped 70% on day one. Nebius reports 4+ customers per GPU while growing 684% year-over-year. The supply curve for frontier silicon is being negotiated in private, in blocks of $10B and up.


    What This Means for Enterprise Buyers

    The optionality most infrastructure plans were quietly relying on, the assumption that capacity would be available somewhere at some price when the workload showed up, is the line item being deleted. The xAI lease and the broader financialization of GPUs could meaningfully alter compute economics for enterprises over the next 12 to 18 months as excess infrastructure hits the secondary market. The two trends sit in tension: long-term capacity is being locked up by hyperscalers, while short-term supply may improve as failed frontier players become landlords. The decision this quarter is which side of that tension a buyer underwrites.

    Infrastructure Investment Hierarchy

    1. Frontier compute, locked via bilateral commitments such as OpenAI/Cerebras and Microsoft/OpenAI
    2. Power supply, being platformized via Fervo/Google at 3GW
    3. Neocloud capacity, with 4:1 demand and 684% growth at Nebius
    4. Secondary market, emerging via the xAI/Colossus lease

    What to do

    1. Assess whether the GPU financialization trend creates procurement advantages — explore secondary-market and lease options within 90 days

      This quarterxAI leasing 45% of Colossus signals a new supply channel for enterprise buyers at potentially better economics than hyperscaler premiums
    2. Secure long-term power supply agreements or partnerships for any planned data center expansion this quarter

      This quarterFervo's $10B+ validation means the market has priced power as strategic — contracts signed now set 2028-2030 position
    3. Model your AI compute needs for 2027-2028 and determine whether multi-year commitments are warranted at current prices

      This quarterThe marginal buyer arriving in 2027 will get compute but not 2025 terms — the window for favorable long-horizon procurement is closing
    4. Evaluate Cerebras and alternative silicon as a strategic hedge against GPU monopoly pricing

      Watch$56B market validation means credible alternatives to the default GPU stack now exist — diversification is no longer theoretical
  4. 04

    AI Liability Architecture: Courts Are Deciding Faster Than Congress

    The Blueprint

    a16z has published what is, on any honest reading, the most comprehensive AI liability lobbying blueprint the industry has produced. The headline proposals are user-liability defaults and damages caps. The subtext is that the venture class has decided the legal architecture of the next decade is worth $115.5M in political capital right now, which makes a16z the largest disclosed donor of the 2026 midterm cycle. Marc Andreessen sits on the White House tech council.

    The Tempo Mismatch

    Courts are deciding cases now. Congress is debating. The likely sequence is precedent-setting rulings landing before any comprehensive federal framework, producing a patchwork of judicial standards that subsequent legislation has to work around rather than replace. Active litigation against general-purpose AI tools could impose substantial penalties on developers for user misuse before a legislative framework exists at all.

    The competitive moat for any serious operator in this space for the next five years will be the quality of the audit trail, the defensibility of the evaluation process, and the contractual allocation of residual risk with upstream vendors.

    The Open-Source Threat

    If developer liability for downstream use becomes the standard, the economic logic of releasing an open-source model stops working. No rational actor open-sources a model that generates unbounded liability for every downstream application. Product strategies that quietly assume continued access to open weights, which describes most of them, carry an unpriced dependency on regulatory outcomes that the P&L does not show.

    The ODNI vs. Commerce Battle

    Inside the White House, the intelligence community is proposing a center inside ODNI for evaluating new AI models before release. It is a licensing regime in everything but name. Commerce's alternative, voluntary agreements through CAISI, preserves speed-to-market but lacks enforcement. CAISI published and then retracted voluntary testing agreements with Google, Microsoft, and xAI in the same week. The outcome determines whether frontier AI companies operate under pre-release gating or a voluntary framework.

    Two Scenarios, Different Businesses

    DimensionCommerce-Led (Voluntary)IC-Led (Pre-Release Gate)
    Release timelinesUnchangedExtended by months
    Compliance costDisclosure + export controlsClassified obligations
    Market structureFavors speed/startupsFavors deep-pocket incumbents
    Open sourceViable with safe harborsPotentially uninsurable

    What to do

    1. Commission legal exposure audit of AI products against three competing liability frameworks (absolute, safe harbor, user-liability presumption) by end of Q3

      This quarterCourts are setting precedent now — the outcome determines whether your AI products face negligence or strict liability
    2. Begin building audit-ready AI governance infrastructure — model cards, safety testing documentation, incident reporting — that would satisfy proposed safe harbor requirements

      This quarterSafe harbors reward companies that built governance proactively; retrofitting after rules harden costs 3-5x more
    3. Evaluate your open-source AI dependencies and develop contingency plans for a world where open-source model availability contracts

      This quarterDeveloper-liability regimes make open-sourcing an uninsurable risk — most AI strategies have unpriced dependency on continued open-weight access
    4. Engage in the federal legislative process through industry coalitions before the 18-month window closes

      This quarterCompanies not in the room will live under rules written by companies that are — a16z's $115.5M makes the stakes clear

From the editor's desk

Stories

  • Update: Anthropic tripled ARR from $9B to $30B+ in four months while xAI leased it 220K GPUs — the frontier lab hierarchy from 6 months ago no longer exists

  • ServiceNow blew its entire annual Anthropic budget by May — enterprise AI cost governance is structurally broken without SLAs, telemetry, or predictable pricing from model providers

  • Lovable dissolved its growth management layer and replaced it with autonomous 'High-Impact IC' roles — former VPs report 90% time on building vs. coordination, and the model is expanding after 5 months

  • Duolingo walked back blanket AI mandate after quantifying ~20% 'slop tax' on AI-generated content — performative adoption ≠ productive adoption

  • Abridge raised at $5.3B on 80M+ clinical conversations, positioning as 'clinical intelligence layer' above EHRs — the wedge-and-expand playbook working in a $4T market

  • Sigstore provenance forgery now confirmed — TeamPCP/Shai-Hulud extracts OIDC tokens from CI/CD runner memory, compromised TanStack, UiPath, and Mistral AI packages

  • Microsoft MDASH deploying 100+ coordinated AI agents against its own codebase, found 16 exploitable flaws in a single Patch Tuesday — self-audit at machine speed is the new baseline

  • Data center opposition hitting permitting risk: 4,000 complaints against one 9GW facility, states considering outright bans — compute scarcity premium rises for pre-permitted sites

The Bottom Line

AI offensive capability crossed from 'can hack individual systems' to 'full autonomous network takeover' this week, while the platform layer where AI agents act is being claimed simultaneously by SAP, ServiceNow, Apple, and Google — and the firms best positioned to weather both shifts are those that treat compute as a financializable asset (xAI just proved it is) and governance as a product (courts are writing liability defaults right now, not Congress). The two decisions that compound from here: architect your product to be the thing agents orchestrate through rather than the thing they bypass, and compress your security posture against adversaries that now see through your EDR in days rather than months.