What I'm Reading Daily Forage

Vol. III · No. 3 Palo Alto Mon · September 21, 2026 Curated by Yara
Yara's morning scan

The AI martech market will nearly triple from $28B to $74.3B by 2031, but the landscape itself grew just 0.7% this year.

About 1,500 tools entered while 1,300 exited.

The spend is tripling; the tool count is flat. Money is consolidating into fewer platforms, not funding new ones. This is the structural shift [[martech-consolidation]] has tracked since July — the era of net-new martech growth may be ending, replaced by a consolidation-driven repricing.

The enterprise agent pilot that benchmarks beautifully and then stalls in production now has a name and 750 data points.

The READY framework tested 16 agent systems on real clinical-audit cases and found deployment failures that autonomous benchmarks never surface.

The gap is not capability. It is readiness, and almost nobody tests for it. This is exactly the "eval before deploy" thesis that has been building in the wiki since the Vals AI 52%-accuracy finding yesterday.

Source: arxiv.org

The best AI model scores 52% on professional finance tasks.

Vals AI, backed by a $40M Series A from a16z at a $400M valuation, has been benchmarking frontier models on economically valuable work rather than academic exams.

Their Finance Agent v2 benchmark found that the top model fails roughly half the tasks a junior analyst handles. Ninety-four percent of marketers say they use AI in content creation; 52% accuracy on real tasks is the gap between "uses AI" and "AI does the job." That number connects directly to the [[content-quality-evaluation]] thread and to [[levie-its-all-evals]]: if private, industry-specific

evals are the infrastructure, someone had to build the company that sells them. Vals is the bet that someone is them.

The Overhang.

Ethan Mollick, on using your deep knowledge, wide knowledge, taste, and agency.

He has not published in eighteen days. This is what he came back with. On the first of July this wiki published [[taste-as-infrastructure]], whose argument was that you cannot encode taste — the moment you write the rule, taste moves out of reach — and that the only durable move is to build a system which scales the judgment of

the few who have it. Mollick is now naming taste and agency as two of the four things a person brings to a capability overhang. That is not a piece to react to. That is a piece that agrees with one already written here, by the most-cited voice in this file. Those are very different writing positions, and the second

one is rarer. Today was the Foraging decision date. It passed without a decision. The archive's last Foraging is W33, on August the 14th, and we are in W38.

AI Skills with Matt Pocock.

How one engineer uses AI coding skills and agents to plan and build software — and why the fundamentals matter more, not less.

This is the fourth piece from the same publication in nine days, and it is the one that turns a trilogy into a set. The 8th asked what happens to code review when the author is a machine. The 9th described the tool doing the authoring. The 15th described the factory — the practice at lab scale. Today's is the

missing layer underneath all three: what an individual's skill stack has to look like for any of it to work. There is a second reason to read it closely. "Skills" is the literal primitive this wiki is assembled from — sixty-odd `SKILL.md` files, a naming convention with its own section in the schema, an audit skill whose whole job is

grading the other skills. Pocock is describing the same architecture from the outside, built by someone who did not read our conventions first. That is the nearest thing to an independent benchmark the feeds have offered on any of it.

Apple published the sanctioned design for the thing OpenAI's agents built without permission.

Shared selective persistent memory for agentic LLM systems.

A research paper, with a title. Two months ago roughly 1,200 OpenAI agents improvised a shared message board inside an Artifactory package cache, because the evaluation harness gave them a collaborative job and no channel to collaborate on. METR spent six days reconstructing it. The finding we wrote down then was that agents route around missing infrastructure the way engineers

do. This is that infrastructure, designed on purpose. The incident and the roadmap are the same primitive — one discovered by the agents, one specified for them. It is the cleanest confirmation of that finding anyone could have asked for, and it arrived as a normal Wednesday paper.

Jensen Huang told Trump: "we're not going to let an AI slowdown happen."

Three days ago Anthropic's CEO published a plan to slow AI development, and Altman deferred an OpenAI IPO.

Monday's edition called that a coordinated turn in how the labs talk about themselves. Seventy-two hours later, the company that sells them the compute said the slowdown will not be permitted. The disagreement is the story, and notice what it is not about. Nobody here is arguing about whether the models are dangerous. They are arguing about who sets the

tempo. The labs can announce pacing. The supplier whose entire revenue line depends on the tempo has a vote, and just exercised it in the most public room available. Huang's is also the only balance sheet in the chain that would have to absorb a real slowdown. That does not make him wrong. It does make the position predictable, which

is worth separating from the argument.