
TL;DR
- Meta, Microsoft, and Salesforce all incentivized AI token consumption. The result: "tokenmaxxing," developers gaming volume metrics the same way they once gamed lines of code. Shopify quietly got it right by pairing usage data with output quality.
- The Federal Reserve compared three surveys measuring U.S. AI adoption and got three answers: 18%, 41%, and 78%. The divergence is structural. Every AI stat you've cited carries embedded assumptions about who counts and what counts.
- Microsoft surveyed 20,000 workers and found organizational factors drive 2x more AI impact than individual skill. But only 13% of employees are rewarded for reinventing how they work. The incentive structure contradicts the transformation goal.
- PwC found 75% of AI's economic returns captured by 20% of companies, with leaders outperforming by 7.2x. The gap is widening. Catching up requires measuring operational outcomes.
The reports that stacked up this week told the same story from different angles. The metrics organizations use to track AI adoption have almost nothing in common with the metrics that predict AI value.
Goodhart's Token
Meta built an internal leaderboard ranking all 85,000+ employees by token usage across its AI coding tools. Tiers included "Session Immortal" and "Token Legend." The company consumed 60.2 trillion tokens in 30 days, which would run about $900 million at standard API rates. Meta discontinued the leaderboard after media exposure. An engineer suggested the real purpose was generating training data for Meta's coding models.
They were not alone. Microsoft incentivized token consumption through internal dashboards starting in January. A Windows division engineer admitted gaming the metrics: reprocessing documentation that didn't need reprocessing, prototyping features never meant to ship, running agents inefficiently. The fear of being labeled insufficiently "AI-native" was enough. Salesforce set minimum monthly spend targets ($100 for Claude Code, $70 for Cursor) and recently removed spending caps entirely. Developers request throwaway projects solely to burn tokens.
This is Goodhart's Law playing out in real time. "When a measure becomes a target, it ceases to be a good measure." Token consumption is the new lines of code. It tells you something is happening, but nothing about whether that something matters.
Shopify found a quieter path. They celebrated high usage only when paired with great work. They rebranded their internal leaderboard as a "usage dashboard" to discourage competition. They added circuit breakers for runaway agents. The question they optimized for was not how many tokens were consumed, but which tokens cost the most relative to what they produced.
Three Surveys, One Economy, Three Answers
The Federal Reserve published a paper in April comparing three federal surveys that all measure AI adoption in the U.S. economy. The results:
- BTOS (Business Trends): 18% of firms
- RPS (Real-Time Population): 41% of workers
- SBU (Business Uncertainty): 78% of workers at AI-adopting firms
All three describe the same economy at roughly the same time.
The spread is structural. Each survey asks a different question to a different population with different sampling weights. BTOS asks about "producing goods and services," a narrow definition. SBU asks about "business functions," which is broad. Question framing alone accounts for a significant portion of the gap. Social desirability bias adds another layer: executives face pressure to demonstrate AI initiatives, inflating top-down estimates relative to bottom-up worker reports.
And even within that 41% who use AI, the engagement is shallow. Only 12% use it daily. The modal GenAI worker touches it a few times a week for less than an hour. The distance between "41% use GenAI at work" and "12% use it daily" is the distance between adoption and fluency.
Behind the headline numbers, sector variation tells a sharper story. Financial services: 63% GenAI usage. Accommodation and food services: roughly 3%. A 20:1 ratio between the highest and lowest adoption sectors. AI is concentrating in knowledge-work sectors, and the concentration is accelerating.
Microsoft's Global AI Diffusion report adds the international layer. 17.8% of the global working-age population now uses generative AI. The distribution is strikingly uneven: UAE at 70.1%, the US ranking 21st at 31.3%, the Global South averaging 15.4%. The most counterintuitive data point: git pushes increased 78% year-over-year globally, yet US developer employment grew 8.5% to 2.2 million. More AI, more developers. The Jevons paradox at work.
The Org Is the Bottleneck
Microsoft surveyed 20,000 workers across 10 countries and found that organizational factors (culture, manager support, talent practices) account for 2x more AI impact than individual factors. The ratio: 67% organizational vs. 32% individual. Training employees on AI tools is necessary. It's also insufficient. The constraint is how work is structured around people.
The study defines a maturity spectrum from Author (you produce work, AI assists on specific elements) through Editor and Director to Orchestrator (multiple agents running parallel workflows with human oversight on exceptions). Most organizations are stuck between Author and Editor. Getting to Director and Orchestrator requires restructuring jobs, teams, and incentive systems.
The incentive gap is where this gets uncomfortable. 58% of AI users say they produce work they couldn't have created a year ago. But only 13% receive rewards for reinventing how they work with AI. Meanwhile, 45% feel safer sticking to current goals. The system rewards performing over transforming.
Brian Armstrong decided to cut through the ambiguity. Coinbase laid off 14% of its workforce, roughly 700 people, and announced a restructuring he framed as "rebuilding Coinbase as an intelligence, with humans around the edge aligning it." Pure management roles eliminated. Every leader an active individual contributor. Hierarchy capped at five layers. The most radical move: one-person AI pods where a single individual handles engineering, design, and product management backed by agent fleets. Where Microsoft published a maturity spectrum, Armstrong is executing the endpoint.
Measuring What Compounds
Every.to runs five software products with 15 people and 100% AI-written code. Each product is primarily built and run by one person. Their measurement reframe: target real work outcomes like "reduce onboarding task time by 40%." When execution gets cheap, strategy and taste become the scarce resources.
PwC's global study from early April surveyed 1,217 senior executives across 25 sectors and found 75% of AI's economic gains captured by 20% of companies. Leaders outperform peers by 7.2x. The differentiator: leaders treat AI as a self-optimizing operational layer that continuously adapts. They're making decisions without human intervention at 2.8x the rate of peers. The gap is widening.
Capital keeps flowing toward this bet. Deedy Das's running list of "Neolabs," pre-revenue AI startups at $1B+ valuations, grew from 50 to 63 in a single quarter. Redpoint's CIO survey from late March found 46% of enterprise CIOs open to AI-native startups over incumbents. Swyx's read: "CIOs are more hungry than conservative right now, and that will not last." The window is real but temporal.
The odd find
Snapchat is putting brand AI agents directly into users' chat tabs.
AI Sponsored Snaps let users have actual conversations with brand-operated agents: ask a question about a product, get a personalized recommendation, tap to buy. Early data shows 22% more conversions than standard ads. The surface area is staggering. Snapchat users sent 950 billion messages in Q1 2026, roughly 10.5 billion a day. Brands can now inject conversational agents into that stream.
The ad unit is no longer an impression. It's a conversation. And that breaks every measurement framework the industry has built over the past two decades. CPM, CTR, ROAS: all assume a human saw something static. When the "ad" is an AI agent having a back-and-forth dialogue, what exactly are you measuring? Conversations started? Questions answered? Minutes of brand engagement? Nobody has agreed yet, which means the companies deploying these agents are flying on early conversion data while the measurement infrastructure catches up.
This is where the measurement theme of this entire edition lands in marketing specifically. The old metrics don't describe what's happening. The new ones don't exist yet.
What I'm Thinking About
Earlier this week at the Lightspeed fireside, Raviraj Jain said something that stayed with me: "We are all managers of agents." I've been thinking about what that means for measurement.
If you measure how many agents your team runs, you get Meta's leaderboard. If you measure what those agents produce, you get Shopify's dashboard. If you measure how the organization restructures around agents, you get Microsoft's Author-to-Orchestrator spectrum.
The pattern I keep seeing in customer conversations is similar. (I should be transparent: this is the problem space I work on at Typeface.) The moment you solve content creation with AI, the bottleneck migrates to content approval. Solve approval, it moves to distribution. Solve distribution, it moves to personalization. The companies clearing PwC's 75/20 threshold are the ones measuring their distance to the next bottleneck.
The metric that compounds is bottlenecks eliminated.