I keep a running list of everything I read about AI, content infrastructure, and how organizations actually change. This week the same architecture kept showing up in pieces written by people who don't talk to each other.
Andrej Karpathy posted about maintaining a personal wiki with an LLM. A VC described the AI chief of staff he built for his practice. Within 48 hours of Karpathy's post, someone shipped an open-source version. By the end of the week, a top AI advisor had built one with a feature that challenges your own assumptions.
Some of these builders explicitly cited Karpathy. Others arrived at the same shape independently, solving different problems in different domains. That's what caught my attention: the convergence wasn't coordinated. Raw sources go in. A structured wiki comes out. The human writes the rules. The LLM does the rest.
Here's what I took away from the week's reading.

The piece that reframed how I think about search
Karpathy's thread got 1.2 million views, and I think it's because he named something practitioners already felt. If you care about a body of knowledge, retrieval-augmented generation is the wrong tool. You don't want to embed documents and hope the right chunks surface when you ask a question. You want to process them upfront: summaries, entity pages, cross-references, backlinks. A compiled wiki. The search index was the old answer.
The distinction sounds academic until you see the numbers. Graphify, the first open-source implementation, claims 71.5x fewer tokens per query compared to reading raw files. That's the difference between a system that scales and one that bleeds money.
What I found more interesting than the pattern itself was how fast people adapted it to wildly different domains.
Ryan Sarver, a VC, built "Stella" as his AI chief of staff. He wasn't referencing Karpathy. He was solving his own problem: meeting prep, relationship tracking, task management. But the architecture he landed on is structurally identical. Two-layer memory: raw daily notes, then a curated file synthesized from those notes. His design rule stuck with me: LLMs handle judgment, scripts handle everything deterministic. No mixing.
The ones that explicitly built on Karpathy's post moved fast. Allie K. Miller built Claudeopedia over a weekend and added a feature I hadn't seen anywhere else: an automated job that reads your recent writing against the wiki and challenges your own assumptions. A reflexive layer that asks whether you're repeating yourself or missing something.
None of these are coding tools. They're personal operating systems running on a coding tool's infrastructure.
The reading that convinced me infrastructure matters more than models
Akshay Pachaar published what might be the most thorough breakdown of agent harness architecture I've read: 12 components, from orchestration loops to verification systems to subagent coordination. The headline finding was stark. Changing only the infrastructure, keeping the same model, moved agents 20+ ranking positions on benchmarks.
Nicholas Charriere took this a step further with his "Great Convergence" thesis. By the end of 2026, he argues, app companies, model companies, and infrastructure companies will all look like they're building the same product. The general agent harness is the shared surface. Claude Code made it popular for coding. Now it works for anything you can do on a computer.
I kept thinking about that while reading about Hannah Stulberg's work at DoorDash. She's a former Google APM who scaled Claude Code into what she calls a "Team OS" supporting 20+ people across product, analytics, and operations. The system has seven components: nested configuration files, shared skills, an analytics layer, and a launch gate that won't let you ship a feature until all operational knowledge is compiled into the repo. That last detail is the one I keep coming back to. Compiled knowledge as a prerequisite for launching. The afterthought era is over.
I wrote last week about building a customer dashboard over a weekend that changed how 200 people think about accounts. Aman Khan made the broader case: domain expertise packaged as skills is becoming a go-to-market channel. Companies will ship skills alongside products. The barrier drops from a months-long engineering project to cloning a repository.
Charriere also released Meta-Harness, a method for autonomously optimizing these harness configurations end-to-end. If infrastructure quality outweighs model quality for agent performance, and the benchmarks suggest it does, then automating infrastructure optimization is a compounding advantage.
All of which raises a question: if the infrastructure is shifting this fast, what does that mean for the people expected to operate it?
The hiring bar that just went public
Zapier published the second version of its AI fluency rubric, and it verbalizes what we've been observing and working towards in our own hiring at Typeface. The first version asked whether you could use AI tools. The second version asks whether you build repeatable systems with them, whether you're improving, and whether you catch errors before they ship.
Four dimensions: mindset, strategy, building, accountability. Accountability is the new addition, and it's pointed. "With AI, you can delegate the work, but not the accountability." Zapier now observes candidates using AI in real time during skills assessments, watching how they iterate rather than whether they produce a polished first output.
Aakash Gupta documented what this looks like in practice for PM roles specifically: Google, Figma, and Perplexity now run 45-minute "vibe coding" rounds where candidates build working prototypes with tools like Cursor. The rubric isn't theoretical. It's already in interview loops.
The most consequential change is from snapshot to slope. V2 measures the trajectory of someone's fluency, not where they are today. Hiring for growth rate. That's a significant philosophical shift for any company that takes capability assessment seriously.
Adriane Schwager's reaction added necessary context. The floor has moved in tech, but it hasn't moved everywhere. A workflow Zapier now rejects as insufficient, using AI for first drafts and editing manually, would be considered advanced at most accounting firms and small agencies. She frames the gap as an opportunity, not a failure. I think she's right, and I think most people building AI hiring rubrics aren't thinking about this asymmetry.
Austin Lau at Anthropic offered a complementary lens: four ways growth teams use AI. Automate existing work, use AI as a thought partner, tackle work that used to fall below the ROI threshold, and build bespoke tools. His argument is that the third dimension is the most underused and the real compounding unlock. Practitioners stuck in the first two have plateaued.
What I'm thinking about after this week's reading
Jack Dorsey published a piece arguing that AI eliminates the need for management hierarchy. The org chart becomes an interface. Jaya Gupta made a related argument about context graphs: whoever accumulates structured decision traces first owns future thinking. The enterprise moat is the history of your judgment calls, not your current headcount.
Shreyas Doshi made the individual version of the same point: product sense, the ability to improve on what AI already produces, is the only durable career moat as tools get commoditized.
The thread connecting all of this week's reading is that the middle layer is settling. Between the models below and the applications above, an infrastructure layer is taking shape. Karpathy described it for knowledge. Pachaar mapped it for agents. Charriere says the two are converging. DoorDash is running a cross-functional team on it. Zapier is hiring for the fluency to operate within it.
The question I keep sitting with: the infrastructure is arriving faster than most organizations can absorb it. Some teams are already running on compiled knowledge and shared agent harnesses. Others are still debating whether to allow ChatGPT. The distance between those two groups grew wider this week, and nothing I read suggested it's going to close.
Before the closing, a scenic detour.
The Odd Find
This one's all good taste, no AI.
Jolyon Varley resurfaced a Coca-Cola campaign from 2024 that somehow flew past me. Someone at Coke looked at the 🥤 emoji sitting on every phone in the world and thought: that's ours. What followed was Emoji Coke Cups in Saudi Arabia, a gamified experience pairing food emojis to unlock restaurant vouchers, 630 million impressions, 62% conversion rate. The entire strategy fits on a napkin: pay attention to what's already part of culture. Thank you, Jolyon, for the time travel.
On day three, something shifted. I'd been building my own version of the Karpathy pattern, a knowledge bot called Yara ("friend" in Hindi). Forty sources in, structured wiki out, running on Claude Code. I asked it a question that spanned sources I'd ingested days apart, and it answered with connections I'd forgotten I'd read. That's when the pattern stopped being theoretical.
Building the thing taught me more than reading about it ever could. The schema matters in ways the blog posts don't cover. The compilation step changes how you think about what's worth ingesting. The linting catches gaps you didn't know you had. And this newsletter was drafted from the wiki Yara maintains.