Quality Evaluation — The Hardest Problem in AI Content
A conversation I keep having with many customers and users of Typeface.
Enterprise marketing leader: "We tried AI content tools. The quality wasn't good enough."
Me: "How were you measuring quality?"
A pause. Then: "We could just tell."
When you're producing five pieces of content a month, subjective judgment works fine. When you're producing five hundred, it doesn't.
So you try to make it objective. You define dimensions: brand voice fidelity, persona accuracy, tone consistency, regulatory compliance. You build scoring systems.
And then someone on the marketing team says: "It's on brand, but it's kind of... uninteresting."
You push toward something more edgy. Now it scores lower on brand. How much does brand voice actually constrain what's possible? How do you encode the difference between "safely on brand" and "compellingly on brand"?
Sometimes the marketers themselves don't have the answers yet. They're discovering what they want through the process of reacting to what the system produces. The most AI-savvy marketers get this and treat it as a process: iterate till you get it dialed in.
Quality evaluation is the hardest problem in AI content right now, and there is no easy button. Anyone claiming they've fully solved it is either working on a narrower problem than you think, or oversimplifying.