
TL;DR
- Karpathy declared vibe coding over and replaced it with "agentic engineering." But the bigger reveal was Software 3.0: his entire MenuGen app became unnecessary when a single prompt to Gemini did the same thing directly in the pixels. Programming is now prompting.
- Todd Saunders extended this to vertical SaaS: when AI can learn any domain, companies split into "rails companies" (infrastructure) or features on someone else's harness. "There is no third outcome."
- While the paradigm shifts, the productivity numbers everyone cited were wrong. Karpathy cited METR and Stanford: the real median gain is 10-15%, not 60%. It takes 30-100 hours of deliberate practice before the tools consistently pay off.
- ActivTrak found only 3% of workers are in the productivity sweet spot (7-10% of time in AI tools). 83% of workers in a UC Berkeley study said AI increased their workload. The gap between adoption and effective use is 19x.
- Expedia bet its Gen Z strategy on a livestreamer. 25% of Gen Z bypasses travel sites entirely. Expedia built a content-to-commerce pipeline with IShowSpeed (150M followers): watch the trip, book the trip. The booking app isn't the product. The content surface is. Same logic as MenuGen, different industry.
Karpathy vibe-coded a restaurant menu app. Full stack: OCR pipeline, Vercel deployment, image generation. A real product. Then he saw the Software 3.0 version: take a photo of the menu, give it to Gemini, say "overlay images onto the food items." The model rendered results directly into the pixels.
"All of my MenuGen is spurious. That app shouldn't exist."
That sentence landed differently than most conference quotes. A builder looked at something he built and realized the entire category of effort was unnecessary. Not because the app was bad. Because the paradigm shifted underneath it.
The paradigm
At Sequoia's AI Ascent, one year after coining "vibe coding," Karpathy killed his own term. His replacement, "agentic engineering," gets the headlines. But the deeper shift is the one he almost buried in the middle of the talk.
Software 1.0: writing code. Software 2.0: arranging datasets, training neural nets. Software 3.0: prompting. The context window is the lever. The LLM is the interpreter. Programming is now prompting.
This isn't code getting faster to write. It's a new category of information processing that wasn't possible before. LLM knowledge bases are another example Karpathy gave: "This is not even a program. There was no code that would create a knowledge base based on a bunch of facts."
Vibe coding raised the floor (everyone can build software). Agentic engineering preserves the ceiling (quality, security, reliability). They are different things, not sequential stages. The developer keeps taste, judgment, and specification. Agents handle implementation.
What this eats
If Software 3.0 means the context window replaces the codebase, entire categories of applications become unnecessary. Karpathy's MenuGen is the small version. The large version is playing out across industries.
Todd Saunders extended the logic to vertical SaaS. When AI can be trained on any domain quickly, proprietary domain knowledge stops being a durable moat. Companies split into "rails companies" (owning payments, identity, compliance, data as infrastructure) or "domain-only companies" that become features on someone else's harness. "There is no third outcome."
Ben Lang published YC's emerging playbook for AI-native companies: "token-max, not headcount-max" as the operating metric. Software factories where repos have no handwritten code, just specs and test harnesses where agents iterate until tests pass. Every action must produce an artifact AI can learn from to make the company "queryable."
In February, Azeem Azhar and Nathan Warren provided the macro data anchoring this: monthly AI revenue grew from $772M (January 2024) to $13.8B (December 2025). 18x in 24 months. Realized revenue, not projected TAM.
The capital is following the paradigm. The question is whether organizations can.
Meanwhile, on the ground
They mostly can't. Not yet.
Gallup crossed 50% this week. Half of all U.S. employees now use AI at work. Morgan Stanley, in the same week, found that regular AI usage jumped 13% while confidence in using the technology fell 18%. The people who use AI most are the least confident they understand it.
Karpathy quantified why. At the same Sequoia talk, he cited the research everyone should have been reading: the METR study found experienced developers on their own codebases worked 19% slower with AI tools. Stanford measured median productivity gains at 10-15%, not the 60% that went viral. And it takes 30-100 hours of deliberate practice before the tools consistently pay off.
ActivTrak analyzed real-time productivity data and found the sweet spot at 7-10% of time in AI tools. Only 3% of workers are there. 57% spend less than 1%. An eight-month UC Berkeley ethnographic study found that 83% of workers said AI increased their workload: expanded task scope, dissolved boundaries between work and rest, parallel processing until the aggregate effect was exhaustion.
Korn Ferry found the leadership gap underneath: 52% of organizations plan autonomous AI agents, but only 22% believe their leaders can manage human-AI teams. The paradigm is shifting. The organizations aren't ready.
Four voices described different facets of the same gap. In an internal memo that leaked last April, Shopify CEO Tobi Lutke made AI usage a structural condition of employment: "Reflexive AI usage is now a baseline expectation." Gokul Rajaram predicted product design will cease to exist as an independent function by end of 2026. Jaya Gupta named the structural mechanism: experience is now a tax. And Aaron Levie predicted a consulting onslaught because "there is no shortcut to change management for the enterprise."
The odd find
Software 3.0 isn't just eating apps. It's eating how people discover things.
25% of Gen Z bypasses online travel agencies entirely. They decide where to travel on TikTok and YouTube, not on brand domains. Expedia's response: a year-long partnership with IShowSpeed (150 million followers), structured as a full content-to-commerce pipeline.
They built a TikTok account (@Exspeedia_) for real-time clips from Speed's travels, and a custom microsite (Exspeedia.com) where viewers book the exact flights, hotels, and activities featured in the content. The gap between inspiration and transaction collapses to zero.
This is Karpathy's MenuGen logic in a different domain. The booking app isn't the product. The content surface is. The traditional interface, a search box on a travel site, becomes as unnecessary as Karpathy's OCR pipeline once you realize the user already decided what they wanted before they ever opened the app.
What I'm thinking about after this week's reading
Karpathy looked at his own app and saw it was unnecessary. That's the cleanest version of where we are. The paradigm is shifting from building software to prompting intelligence, and the shift is moving faster than most organizations can absorb it.
The 30-100 hour practice threshold is the most useful number in all of this. It connects the sweet spot (3%), the confidence paradox (usage up, confidence down), and the leadership gap in a single frame: the technology moved from 1.0 to 3.0 while the org chart is still debating 2.0.
The organizations that will cross the gap are the ones that budget for the learning curve instead of expecting the tools to work on contact. The paradigm rewards practice, not enthusiasm.