Shipping AI features
The PM job that barely existed three years ago — specifying, testing and shipping a feature whose output you can't fully predict. Evals, failure modes, fallback UX, and the artifacts that make you the person who decides when it's ready.
Every other category on this site is about using AI to do your existing job faster. This one is about the job itself changing.
When your feature is a deterministic form, “done” is a checklist and QA can run it. When your feature calls a model, the same input produces a different output on Tuesday than it did on Monday, and the old acceptance criteria quietly stop meaning anything. Somebody has to decide what “good enough to ship” means, encode it so it can be re-checked on every change, and own the answer when it regresses. At most companies nobody owns that yet — engineering treats it as a product question and product treats it as a technical one.
That gap is the opportunity. The artifact that closes it is an eval set: a collection of real inputs with a defined view of what a good response looks like, run automatically. It is a product document that happens to execute. Writing it requires knowing what customers actually need, which is the one thing an engineer can’t do for you — which is exactly why it’s the most defensible thing a PM can own right now.
The tools here are engineer-shaped and don’t apologise for it. That’s the tax on being early.
Ranked — 2 tools
Best first · our pick highlightedpromptfoo
Our pickA PM who wants to turn 'the AI answer feels worse this week' into a number, without asking engineering for anything.
Claude Code
A PM who wants to stop writing tickets for small internal things and just build them — and who can tolerate a terminal.