***Learn more about CAPED (Complexity-Aware Planning, Estimation, & Delivery)***

“Claude, Make Me a Backlog”

In recent conversations with two different clients, they lamented backlogs full of AI slop. Piles of detailed epics and stories no one really chose or even fully understands.

Does AI have a place in backlog management? If so, what’s the right way to use these tools?

The incentives that make this happen

Many POs operate under a set of incentives that make AI-generated backlogs seem like the secret to productivity.

The real measure of PO productivity would be, how much value flows through your team? A development team is a relatively fixed, rather large investment (salaries and benefits, tools and infrastructure, that ever-growing token budget). A good PO maximizes the return on that investment for their org.

That’s hard to measure. It’s a lagging indicator, so even if you can measure it, there’s a long delay between the work and the results. For big multi-team programs, the disconnect between one PO’s work and a valuable result feels almost insurmountable.

So, people look for proxies. What easy-to-measure leading indicators show a PO is doing their job well?

The easiest and most obvious choice? A full and detailed backlog. The PO who manages a long list of detailed backlog items must be doing their job, right? … Right?

Georgetown computer science professor and author Cal Newport says (Monday Advice, June 1, 2026),

So what happens when you set up a work environment in which visible activity is rewarded, then you give everyone a machine that can automate those efforts, making them essentially free? Well, what’s going to happen is work will become a mad performative dash, a button mashing. Who can turn out more slop quicker than the next person?

Adding AI tools in an environment that already rewards performative work just exacerbates the existing problem. It’s logical for a PO to take advantage of these tools in this incentive structure, but it’s counterproductive for the org and for the PO’s own career.

Why it’s counterproductive for your org and for your career

When people use AI tools to generate a backlog, they’re behaving as if the constraint on the flow of value is having enough ideas and enough detail for the backlog. But that was never the constraint. There are always more things you could do than things you have capacity to actually do.

No one ever comes to me and says, “Richard, I’ve got all this capacity. If only I could come up with something to do with it.” It’s always, “If only we had more resources!”

The job of the PM or PO is to choose, out of all the many things we could spend our limited capacity on, which ones actually matter for our business and what order is best to do them in. Then, it’s to communicate those in a way that aligns everyone on why they matter, what success looks like, and therefore, what’s in scope and what’s out.

Reasonable vs Correct

Delegating this work to an LLM isn’t going to produce the best outcome. AI doesn’t know what your product should do. It can’t. Newport explains:

What you’re going to get out of the LLM is a reasonable sounding plan…because that’s what it does. it’s a story completer.

Reasonable sounding plans get you in trouble. You need correct plans. And the way that humans actually make plans is we do a few things.

One, we test a bunch of possibilities internally to see what makes the most sense. And two, we have some sort of notion of correctness and a world model we can use to evaluate plans to see, does this actually do what I wanted to do? … [The LLM] doesn’t have a world model to test it against… It has no ability to do future simulation of possible outcomes. And so we just get a story that sounds like a good plan, but often has issues along the way.

This distinction between reasonable outputs and correct outputs explains why AI tools are better at software code than prose. Reasonable and correct (1) correlate more tightly in a constrained context like code and (2) correct is testable, often in an automated way.

“Correct” in a backlog is a judgment about the future and about tradeoffs. Reasonable and wrong backlogs are quite common.

As AI tools get faster and cheaper, the cost of producing a reasonable (but likely wrong) backlog approaches zero. The cost of a human PO isn’t going down. If you don’t add human perspective and judgment, you’re replaceable.

How to add value as a PM or PO using AI as a tool rather than your replacement

On the other hand, a good PM or PO who brings all their human capability to the table can still add a lot of value by pointing a team (and their AI collaborators) in the right direction. Even if AI is making development faster and cheaper in many orgs, it’s still expensive to go in the wrong direction. And there are ways to use AI to make you better at that job rather than reasonable yet worse.

The principle here is this: Apply the Theory of Constraints. Keep the parts of the job that benefit from human judgment. Use tools to improve your judgment and free up more capacity for that part of the job.

Here are some recent examples of how we’ve seen PMs and POs apply AI tools to enhance their effectiveness without replacing human judgment…

  1. Process data about customer behavior and needs. Product people are often constrained on customer evidence that would improve their judgment. AI is good at transcribing interviews, summarizing them and pulling out key quotes. It’s good at analyzing a database of support tickets to find patterns. It can make sense of Google Analytics or other product instrumentation to see how users actually behave.
  2. Generate solution prototypes for customer tests. Once a problem is validated and you’re doing solution testing, AI tools can be used to generate prototypes at various fidelity levels to use to test customer preferences and behavior. These can range from low-fidelity paper prototype equivalents to clickable, realistic mini-applications. (Beware the risks associated with too-good prototypes, though.)
  3. Improve your functional tests. We’ve had great results generating and improving Cucumber scenarios using Claude Code and the Gherkin guidance I created based on my BDD book. Regression testing can take a lot of a PO’s time, and poorly written automated tests can become a maintainability headache. But a good suite of PO-readable automated tests is quite valuable.
  4. Do more comprehensive analysis, at the right time. Once you’re in the complicated domain, backlog generation involves a lot of analysis. “What are all the variations here that we should consider building?” AI tools are good at enumerating variations. But be careful not to let them make the calls about what’s in scope and what’s top priority. Claude can generate the list of possibilities. The PO should select and rank the ones that belong in the backlog.

All these examples use AI to enhance human effectiveness, exploiting and elevating the constraint, to put it in Theory of Constraints terms.

If AI is a big line in your 2027 plan, make sure you’re pointing it at your real constraint. Otherwise, you’re just producing reasonable, wrong work faster.

We help organizations look at the whole system (strategy and priorities, the incentives around them, how work gets shaped) to see where AI will increase the flow of value and where the constraint is human judgment about what to build. When it’s judgment, we help your product people get better at it.

Schedule a call to talk it through. We’ll help you unpack the challenge, ask useful questions, and identify some ways forward.

Last updated