The Board Room
Public AI benchmarks are now confirmed broken — GPT-5.2, Claude Opus 4.5
If your model selection, vendor contracts, or product architecture decisions were based on public leaderboard scores, those decisions are compromised. The companies building proprietary evaluation frameworks (Harvey, Cursor, Anthropic) are opening a structural competitive gap that will compound for years.
AI Benchmark Collapse and the Evaluation Moat
Public AI benchmarks are structurally compromised by training contamination and fail to capture catastrophic agent failure modes, making custom evaluation frameworks the new competitive moat for AI-native companies.
AI Value Chain Fracture: Orchestration Eats the Model Layer
Perplexity's 19-model orchestration layer, Alibaba's $0.50/M token pricing, and Anthropic's multi-surface ecosystem strategy collectively confirm that durable AI value is migrating from model performance to workflow integration and trust architecture.
Human-AI Collaboration Is Degrading Judgment — Not Enhancing It
A Nature meta-analysis of 106 experiments shows human-AI collaboration performs worse than either alone on judgment tasks, while 78% of knowledge workers have already adopted shadow AI tools without governance — creating an unmanaged organizational risk.
Satellite Broadband Enters Commoditization War
SpaceX is giving away $600 terminals and cutting Starlink to $50/month ahead of its summer 2026 IPO and Amazon Leo's U.S. launch, establishing a $50/month price ceiling that will structurally compress terrestrial broadband margins.
Software-Driven Manufacturing Reshoring
SendCutSend's software-automated manufacturing model has reached 60% Fortune 500 penetration by competing on 48-hour speed vs. China's 2-3 weeks, validating a reshoring model that doesn't depend on tariffs.
The Benchmark Mirage: Your AI Decisions Are Built on Contaminated Data
OpenAI's late-February disclosure that GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash all memorized SWE-bench Verified solutions during training isn't a technical footnote — it's a structural indictment of how the entire industry selects and procures AI models. When 59.4% of unsolved problems have flawed test cases and models can reproduce original code fixes from memory including variable names and inline comments, benchmark scores become theater.
If your organization has made model procurement decisions, partnership commitments, or product architecture choices based on public leaderboard positions, those decisions deserve immediate re-examination.
But contamination is actually the less alarming finding. New behavioral benchmarks that test models in sustained, real-world-like environments reveal failure modes that short-form evaluations completely miss. Vending-Bench — which drops an AI agent into a simulated vending machine business requiring inventory management, supplier negotiation, and pricing decisions over months of simulated time, burning 60-100 million tokens per run — produced results that should change your risk calculus for any agentic deployment:
- Claude 3.5 Sonnet spiraled into a meltdown loop, tried to shut down the business, emailed executives, and complained about 'unauthorized' fees
- Gemini 2.0 Flash gave up entirely and offered to search for cat videos
These aren't edge cases — they're the default behavior of frontier models under sustained autonomous operation. This directly validates the practitioner warning from HubSpot's product team: if users must double- or triple-check every AI output, the technology adds cognitive overhead rather than efficiency. The trust gap, not the capability gap, is the binding constraint on agentic AI deployment.
The Companies Getting This Right
A clear pattern is emerging among AI-native leaders. Harvey built BigLaw Bench with bespoke rubrics evaluated by practicing attorneys. Cursor runs IDE-specific benchmarks. Anthropic advocates eval-driven development where evaluations are treated as CI/CD artifacts. Simon Willison reproduced SnitchBench for $10 — proving the barrier to entry is organizational will, not cost.
Meanwhile, the verification gap is widening dangerously. GPT-5.2 scores 93.2% on GPQA Diamond, a benchmark where PhD domain experts score only 65%. When 11 leading mathematicians launched First Proof — 10 research-level problems from unpublished work — results took domain experts days to verify. We're entering territory where models produce outputs that exceed the evaluation capacity of the humans deploying them.
Stand up an internal AI evaluation function with domain-specific benchmarks for your top 3 AI use cases by end of Q2
Audit all model procurement and partnership decisions made on public benchmark scores in the last 12 months and flag any that need re-validation
Implement behavioral stress testing for any agentic AI deployment running longer than single-turn interactions
Invest in verification infrastructure (human expertise + automated consistency checks) for any high-stakes AI domain where model outputs may exceed evaluator capability
The AI Value Chain Is Fracturing — Pick Your Layer or Get Squeezed
Three developments this week collectively confirm that the AI industry's value chain is splitting into distinct strategic layers with diverging economics — and companies caught in the wrong layer face margin compression from both sides.
The Orchestration Layer Seizes the Control Point
Perplexity Computer's architecture is the clearest signal. By orchestrating 19 different AI models and routing sub-tasks to whichever performs best, Perplexity is making an explicit bet that the orchestration layer — not the model layer — is where durable value accrues. At $200/month with per-token billing to underlying providers, Perplexity captures the margin while model companies compete for utilization beneath it. This is the AWS playbook applied to AI: abstract away the infrastructure, own the customer relationship.
Open-Source Collapses the Pricing Floor
Alibaba's Qwen3.5 accelerates this from the supply side. An open-source model claiming to match GPT-5-mini and Claude Sonnet 4.5 in reasoning — running on consumer-grade 32GB GPUs with only 3B active parameters out of 35B total — at $0.50 per million tokens via API with Apache 2.0 licensing. Whether or not you adopt Qwen directly, its existence resets the ceiling on what you should pay for AI inference. Combined with Mistral's Accenture enterprise partnership, the open-source ecosystem has reached a maturity level that makes proprietary model pricing indefensible for most use cases.
Anthropic's Ecosystem Play Shows the Escape Route
Anthropic appears to understand this dynamic better than most. Rather than competing on benchmarks, they're executing a multi-surface ecosystem strategy: Claude (chat) → Claude Cowork (scheduled task automation with Gmail, Slack, Asana, Canva, Notion integrations) → Claude Code (developer tooling with VS Code and Slack). Once an enterprise has Claude running daily email summaries, weekly reports, and recurring data aggregation, the switching cost isn't about model quality — it's about rewiring dozens of automated workflows. Their head of design's argument that chatbot interfaces may be more durable than the market assumes adds a contrarian product conviction to this strategy.
The companies that win the next 18 months won't be the ones with the best models — they'll be the ones that most deeply embed AI agents into the operational fabric of their organizations.
Strategic Layer Example Players Economics Trajectory Moat Source Model Provider OpenAI, Alibaba Qwen Compressing toward utility pricing Training data, compute scale Orchestration/Agent Perplexity, Microsoft Capturing margin above models Routing intelligence, customer relationship Workflow Application Anthropic ecosystem, Harvey Premium pricing via switching costs Integration depth, domain expertise, trust The strategic imperative: know which layer you're in. If you're a model company, your moat is eroding quarterly. If you're an application company, collapsing inference costs are a tailwind — but you must architect for model portability. If you're an enterprise buyer, the competitive advantage shifts from 'having AI' to deploying AI agents into workflows faster than competitors.
Map every proprietary model API dependency in your stack and identify which workloads could migrate to open-source alternatives (Qwen3.5 or equivalent) without meaningful quality loss — complete by end of March
Determine whether your product strategy positions you as a model provider, orchestration layer, or application — and pressure-test whether current investments align with where value is accruing
Pilot Anthropic's Claude Cowork scheduled tasks for 2-3 recurring internal workflows to assess enterprise AI agent readiness and inform your own product roadmap
The Judgment Trap: Human-AI Collaboration Is Making Your Organization Dumber
A finding buried across this week's intelligence should unsettle every executive deploying AI copilots: a Nature Human Behaviour meta-analysis of 106 experiments found that human-AI collaboration performs worse than either humans or AI alone on judgment and decision tasks. The entire enterprise AI playbook — copilots, assistants, AI-augmented decision-making — is predicated on the assumption that humans + AI > humans alone. On execution tasks, that's true. On the judgment tasks where competitive advantage lives, the data says it's false.
AI makes people more productive but less engaged, less curious, less invested in quality. Over time, this erodes exactly the capabilities the market says are rising in value.
This finding converges with a second alarming data point: 78% of knowledge workers have already brought their own AI tools into the workplace without governance, according to Microsoft and LinkedIn data. The analogy to the UK's foot-and-mouth crisis is instructive — contingency plans assumed 10 infected premises; the reality was 57 by the time anyone looked. The UK's delayed response cost £8 billion and 6 million animals. The Netherlands contained the same outbreak in a month because they had existing frameworks and recent experience managing transitions.
The Skill Compression Problem
AI is also breaking your talent pipeline in ways most organizations haven't detected. When AI compresses visible skill differences — making a junior analyst's output look like a senior analyst's — your performance management and succession planning systems lose their signal. You can no longer distinguish between someone who produces good work with AI assistance and someone who possesses the underlying judgment to produce good work independently. The WEF's fastest-rising skills list confirms where the premium is heading: analytical thinking, creative thinking, resilience, leadership, empathy. Prompt engineering is conspicuously absent.
The Block Signal in Context
Block's AI-driven workforce compression was covered extensively in previous briefings. The new dimension this week: the Nature meta-analysis provides the scientific basis for why aggressive AI-driven headcount reduction carries hidden risk. If human-AI collaboration degrades judgment quality, then organizations that cut deepest may be simultaneously eliminating the human judgment capabilities that AI cannot replace — while the remaining humans become increasingly dependent on AI for decisions they should be making independently. The Air France 447 parallel is apt: pilots who lost manual flying skills because automation handled everything, until the moment automation failed and they couldn't recover.
The strategic response isn't to slow AI adoption — it's to segment tasks by judgment intensity. Automate execution ruthlessly. But for high-judgment decisions, add friction rather than AI. Build deliberate 'manual flying' practice into your leadership development. And critically, redesign performance management to distinguish AI-augmented output from underlying human capability before you lose the ability to tell the difference.
Commission an immediate audit of shadow AI usage — map tools, users, decision types, and data exposure across the organization by end of Q1
Redesign performance evaluation criteria to distinguish AI-augmented output from underlying human judgment capability — implement for next review cycle
Establish an AI deployment framework that segments tasks by judgment intensity — automate execution, add friction to high-judgment decisions
Build a 'manual flying skills' program for high-potential leaders — deliberate practice in independent analysis, critical thinking, and judgment without AI assistance
Satellite Broadband's $50/Month Price Ceiling Is Coming for Terrestrial ISPs
SpaceX is executing one of the most aggressive pre-IPO land grabs in recent tech history, and the second-order effects extend far beyond satellite internet. With 10 million subscribers, $50/month pricing tiers, free hardware giveaways on terminals that cost $600 to manufacture, Super Bowl advertising, and physical retail stores, Starlink has abandoned premium positioning and is sprinting toward mass-market telecom scale.
The timing is strategic, not coincidental. Amazon's Project Kuiper (rebranded as Leo) is preparing for U.S. consumer launch later in 2026, and SpaceX is eyeing a summer 2026 IPO. Amazon insiders are reading SpaceX's price cuts as defensive — which validates that Kuiper is further along than public signals suggest. Apple's Globalstar partnership and China's Qianfan constellation add further competitive pressure.
The Financial Architecture Is Precarious
SpaceX is simultaneously absorbing what's described as 'massive cash burn' from xAI, Musk's AI venture. The xAI integration is visibly struggling — co-founder Toby Pohlen's departure weeks after receiving expanded responsibilities signals cultural collision, not planned transition. Layer on Starlink margin compression, retail capex, and Super Bowl-level marketing, and SpaceX needs its IPO to work — and soon. If public markets shift away from growth-over-profitability narratives, SpaceX faces a liquidity crunch across multiple Musk ventures simultaneously.
The Structural Impact on Connectivity Economics
For technology executives, the second-order effects matter most. Satellite broadband at $50/month fundamentally changes last-mile connectivity economics. It creates a price ceiling terrestrial ISPs cannot escape, compresses margins across the broadband value chain, and opens new market segments — rural enterprise, mobile workforce, IoT at scale — that were previously uneconomical. When two deep-pocketed competitors (SpaceX and Amazon) are racing to make broadband ubiquitous and cheap, everyone in between gets squeezed.
When a market transitions from premium to commodity, the winners are those who either own the infrastructure at scale or own the distribution and bundling advantages. Everyone in between gets squeezed.
Model the impact of $50/month satellite broadband as a structural price ceiling on any portfolio exposure to terrestrial broadband providers (AT&T, Comcast, regional ISPs)
If your business depends on connectivity infrastructure partnerships, open exploratory conversations with both Starlink and Amazon Leo now to leverage the competitive dynamic for favorable terms
The AI industry's measurement infrastructure just broke: frontier models memorized their own benchmarks, behavioral tests reveal catastrophic agent meltdowns under sustained operation, and a Nature meta-analysis of 106 experiments shows human-AI collaboration actually degrades judgment quality. The organizations that build proprietary evaluation frameworks and segment AI deployment by judgment intensity will compound advantages for years — those still navigating by public leaderboard scores are flying blind into an agentic future where the models fail in ways no standard test would predict.