Here's something I keep hearing in conversations with enterprise marketing leaders: "We've tried AI content tools. They're fine. But they don't really fit into how we work."
That word, "fine," is doing a lot of heavy lifting. It means the generation quality is acceptable. It means the novelty has worn off. And it means the hard problems, the ones that actually determine whether AI creates leverage or just creates more content, are still unsolved.
We've spent the last few years working on these problems. Some of what we've learned feels clear. Some of it we're still figuring out. This piece is an attempt to map the territory as I see it right now, not as a finished argument, but as a framework I keep coming back to.
What the First Wave Got Right
The first generation of AI content tools brought the cost of content creation down by an order of magnitude. They forced a useful reckoning inside enterprise marketing teams about what "good" actually means and who decides. And they normalized the idea that AI belongs in the content workflow at all. Three years ago, most enterprise marketing leaders were still treating AI as a compliance risk or a novelty. The tools that shipped in 2023 and 2024 made the category real.
But here's what they mostly didn't do: they didn't build the infrastructure that makes generated content useful at enterprise scale. In our latest research at Typeface, 95% of marketing leaders report rising content demand, but only 14% feel completely confident they can keep pace. That 81-point gap isn't about generation. It's about everything around it.
The Four Gaps That Actually Matter

Orchestration Across Workflows
Enterprise content production is not a linear process. It involves brand guidelines that need to be enforced, regional variations that need to be managed, approval workflows that need to be respected, and channel-specific formatting that needs to be applied. Most AI content tools treat generation as a standalone event. They generate something, and then your existing processes take over. The problem is that your existing processes weren't built to handle volume that's 10x what they used to handle. The seams show immediately. In our research, 67% of marketers say their brands miss key opportunities because of slow or outdated content review processes. At large enterprises, 71% need more than a day just to approve quick-turn content. The orchestration gap isn't abstract. It's measurable.
Real orchestration means the AI system understands the full workflow, not just the generation step. It knows what happens before content is created (briefing, context-setting, persona selection) and what happens after (review, approval, distribution, performance tracking). When those pieces are connected, the leverage compounds. When they're not, you have a faster typewriter.
Knowledge Integration
The question I hear most often from enterprise buyers isn't whether AI can write a product description. It's whether AI can write one that reflects their brand's specific voice, their customer's specific segment profile, and the regulatory requirements specific to their market. Those things live in systems: a brand guideline document, a CRM, a product catalog, a compliance database.
Most AI content tools don't have access to any of that. They generate against a generic prompt. The result sounds like everyone else's content, because it is. Enterprise buyers figured this out quickly. In our research, 87% of marketers say AI-generated content feels generic. "AI-generated" became shorthand for "sounds like it could be from any company," which is a problem if your differentiation is supposed to come from your brand.
Solving this requires structured knowledge integration: a system that can ingest, organize, and reason against a company's specific context. Not a one-time upload of a style guide. A living, queryable layer of institutional knowledge that feeds into every content decision.
Quality Evaluation
This is the one I find hardest to talk about with confidence, because I think the industry, ours included, is still working through it.
Here's the version of the conversation I keep having. Enterprise marketing leader says: "We tried AI content tools and the quality wasn't good enough." I ask: how were you measuring quality? A pause. Then: "We could just tell."
That's not a knock on enterprise marketers. It's a knock on the tools for not giving them better instruments. When you're producing five pieces of content a month, subjective judgment is a reasonable quality control mechanism. When you're producing five hundred, it isn't.
The honest truth is that taking something fundamentally subjective and making it objective enough to evaluate at scale is one of the hardest problems in this space. Quality evaluation means assessing output against defined dimensions: brand voice fidelity, persona accuracy, claim accuracy, tone consistency, regulatory compliance. But defining those dimensions with enough precision to evaluate programmatically? That work is painstaking.
Here's what it actually looks like. You're working with a customer. The content scores well on brand voice. It's accurate. It follows the guidelines. And then someone on the marketing team says: "It's on brand, but it's kind of... uninteresting." So you push toward something more edgy. And guess what? Now it's not on brand. You start asking harder questions. How much does brand voice actually constrain what's possible? Where's the line between "consistent" and "boring"? How do you encode the difference between "safely on brand" and "compellingly on brand"?
These are real questions, and sometimes the marketers themselves don't have the answers yet. They're discovering what they want through the process of reacting to what the system produces. That means quality evaluation isn't a configuration step you do once. It's an ongoing conversation between the system and the people using it, where the definition of "good" keeps evolving.
There is no easy button here. Anyone who tells you they've fully solved quality evaluation for AI content is either working on a narrower problem than you think, or oversimplifying. The teams making the most progress are the ones treating it as ongoing, collaborative work with their customers rather than a feature they shipped.
Closed-Loop Measurement
Content without measurement is an expense. Content with measurement is an investment. The difference, in practice, is whether the system can tell you which content worked, why it worked, and what to do differently next time.
Most AI content tools stop at generation. Some offer basic engagement metrics. Very few close the loop in a way that actually changes what gets generated next. The result is that teams are generating more content than ever, and they know less about what's working than they did before, because the volume makes it impossible to reason about manually.
Closed-loop measurement means connecting content performance data back to the generation system, so that what gets created next is informed by what worked last time. It means building feedback cycles that get faster and more precise over time. This is conceptually straightforward and operationally very difficult. The data pipelines, the attribution models, the feedback signals that are clean enough to actually learn from: all of that is still being built across the industry.
Why This Is a Category Shift, Not a Feature Gap
I've heard the objection: couldn't existing tools just add these capabilities? Couldn't an orchestration layer be bolted onto a generation tool?
Sometimes yes. But mostly no, and here's why.
Building a generation tool and building a Content Operating System require fundamentally different architectural decisions. A generation tool is built around a model and a prompt interface. A Content OS is built around data flows: how knowledge enters the system, how workflows are represented, how performance data feeds back in, how approvals are tracked. Those decisions get made early and shape everything that comes after.
A tool that starts as a generation layer and tries to add orchestration on top will struggle in the same way a spreadsheet struggles when you try to use it as a database. You can go pretty far, but you're working against the grain of what the thing was built to do.
The category shift matters because it changes what enterprise buyers should be evaluating. The right question isn't "does this tool generate good content?" The right question is "does this system support how my organization actually produces content at scale, and does it get better over time?"
Those are different questions. They have different answers.
Why the OS Framing Matters Now
Here's the thing about the four gaps I described above: they're hard enough in a world where the channels stay the same. But the channels aren't staying the same.
Most marketing teams today operate in silos. There's a web team, a social team, an ads team, an email team. Each has its own tools, its own processes, its own content pipeline. The first instinct with AI is to make each silo more efficient: faster web copy, faster ad variants, faster email campaigns. That's useful. It's also thinking too small.
The real opportunity is making marketing truly multi-channel in a way it's never been before. And the urgency comes from the fact that the channels themselves are transforming. It's unlikely that web doesn't change fundamentally. Search is already changing. Ads will evolve. And increasingly, there will be a version of every channel designed for humans and one designed for agents. Those are fundamentally different. Content that works for a person browsing a website is not the same as content that works for an agent evaluating options on their behalf. The channels as we know them today may not look the same in two or three years.
That's why the OS framing matters. A point tool that generates content for today's channels is useful until the channels shift. An operating system is built to adapt. It's the difference between an application and a platform.
What does that OS layer need to support? The four gaps map directly: orchestration that can rewire when workflows change, knowledge integration that persists across channels even as channels transform, quality evaluation that evolves with the work, and measurement that closes the loop regardless of where content lands. Those aren't just nice-to-haves for today's marketing stack. They're the foundation for a marketing stack that can absorb whatever comes next.
A system that's built only to fit into today's workflows is building for a world that's already changing. The companies that build for adaptability, not just efficiency, are the ones that will be ready when the ground shifts. And it will shift.
That's also where this gets genuinely hard. Because you can't transform how a marketing organization works and do it fast at the same time. The technology can move quickly. The organizational change required to actually use it? That's a different challenge entirely. And it's the one that determines whether any of the infrastructure actually delivers on its promise. I'll have a lot more to say about that in my next post.
I should be transparent: this is exactly the problem space I work on at Typeface. That gives me conviction that the category is real, and also means you should read this with that context in mind. I'm not a neutral observer. I'm someone who has chosen to spend years building in this space because I believe the infrastructure layer is what determines whether AI content actually works at enterprise scale.
What This Means for Enterprise Buyers
If you're evaluating AI content platforms, a few questions worth pressing on:
How does your system maintain and update brand knowledge? Not "do you have a style guide upload," but: how does that knowledge stay current, how does it propagate across outputs, and how do you handle conflicts between brand standards and what a model naturally wants to generate?
What does your quality evaluation look like, and how much of it requires working directly with your team to calibrate? Be skeptical of anyone who makes this sound easy. The best answers here will be honest about how much iteration the process requires.
How does content performance feed back into future content decisions? Is that a manual process or a systematic one?
And maybe most importantly: how does your system adapt when the channels change? Is this built to fit today's marketing stack, or to support the next version of it?
The vendors who answer these questions with specificity, including specificity about what's still hard, are the ones building infrastructure. The ones who pivot to model benchmarks or output samples are still thinking about the generation layer.
Conclusion
The generation problem is largely solved. What's not solved is everything around it: the knowledge that goes in, the workflows that shape it, the evaluation that ensures it, the measurement that closes the loop. And underneath all of that, the channels themselves are in flux, which means the infrastructure has to be built for adaptability, not just for today's workflows.
Even when the infrastructure is right, there's a deeper challenge: the organizational transformation required to actually use it. In our research, 48% of marketing leaders cite cultural resistance and change management as a top barrier to scaling AI. The technology is the easier part. The harder part is people and process. That's the subject of my next piece.