Guide · step 4 of 4
Your AI agent is live. What do you look at now?
Conversation counts tell you nothing. The four things worth reading every week: what it could not serve, what it claimed without doing, which tools failed, and what people kept asking for.
Half the value of an agent is in what it answered well. The other half — the half that tells you what to build next — is in what it could not serve, and it appears in no successful conversation.
What it could not serve
The only list that tells you what to add, written by the people who wanted it. It reads as three piles, because it is three different jobs: the thing is missing (an entry to write), the means to find it is missing (a tool to wire up), or it is not your business (a decision, not work).
What it claimed without doing
"That is passed on" with no tool ever called. It only shows by cross-checking the text against the real calls, and it appears in no satisfaction metric — the conversation went fine, it is just that nobody will call back.
Which tools failed
A calendar unavailable, a webhook timing out, a search returning nothing usable. Distinct from the rest: these are outages, they get repaired rather than decided.
What people asked for again
A single request is an anecdote. The same request from several people over several weeks is a roadmap — derived from real demand, at zero research cost.
A customer asking for something you do not sell has just told you what they would have bought. No market study gives you that, and nobody ever filled in the survey you did not send.
Frequently asked questions
Why is the conversation count not enough?
Because it goes up when things go well and when they go badly. Two hundred conversations, a quarter of which leave without an answer, is a bad month that looks like a good one. The number that matters is not how many people talked, it is how many left with what they came for.
How do I know these requests are real and not invented by the model?
By not depending on it alone. A search that came back empty is a fact written into the call result: it records itself without any model having to want to. The agent qualifies what only it knows — what it had tried — and a separate reviewer goes over what neither of them saw. The three correct each other.
A week with no unmet demand at all — is that a bad sign?
No, it is the normal case. Most conversations find what they came for, and a register where everything is logged is worth no more than an empty one. What should worry you is the opposite: a tool that always finds something to report is measuring its own eagerness to please.
What do you actually do with it?
You write the missing entries, wire up the missing tools, and explicitly dismiss what is not your business. That third move matters as much as the other two: without it the same lines come back forever and you stop opening the list.