← All articles
EngineeringOperationsAgents

The Agent Said the Data Was Complete. The Tool Call Said Otherwise.

A confident, wrong answer is the hardest kind to catch, because nothing about it looks broken. Here is how to read an agent run back to the step that produced it, using pagination as the worked example.

Key takeaways

  • A run that completes without errors can still be wrong, and that is the run nobody opens.
  • Most surprising answers come from what a tool returned, not from the model’s reasoning about it. Open the call before rewriting the prompt.
  • An LLM hides a partial result better than a traditional pipeline does: the answer is the usual length, in the usual shape, with real values in it.
  • You cannot reconstruct a non-deterministic path afterwards. Either the run was recorded or the investigation starts with a guess.

“The agent gave me the wrong answer.” It is the least actionable bug report in production AI, and the most common one.

It is also the point where a lot of teams reach for the prompt. That is usually the wrong end. The prompt is the part you can see; the run is the part that explains what happened.

What follows is one worked example, using a healthcare agent asked to reconstruct a patient’s hospital care pathway. The failure in it is not exotic. It is pagination, which has been breaking software for forty years. What is new is how convincingly a language model can paper over it.

The failure that does not look like one

Most agent failures announce themselves. A tool times out, a schema check fails, a request 500s, and something in your stack goes red.

The expensive ones do not. The agent answers. The answer is well structured, specific, full of real values, and wrong. Nothing in it looks like an error, so nothing catches it except a person who happens to know better.

Here is a concrete one. An agent is asked to reconstruct a patient’s hospital care pathway and to say whether the returned results are complete. It produces a clean chronological answer, admission by admission, and states that the results appear complete.

The tool it called had already said the opposite. The answer never mentioned it.

You cannot see that from the reply. You can see it immediately from the run.

Start from the conversation, not a reproduction

The instinct when someone says “the agent gave me the wrong answer” is to try to make it happen again. With a non-deterministic agent that is a bad first move: you are gambling on reproducing a path the model may not take twice, and you burn tokens doing it.

Start where the complaint is. Find the user, open their conversation, and read what they were actually given.

Conversations listed by end user, with the answer the agent returned shown alongside.
The complaint has a conversation behind it, and the conversation has a run behind it.

This is the answer under investigation. It is confident and well organised, which is exactly why nobody questioned it.

Read the run before reading the answer

Open the run that produced it. Before looking at any step in detail, the header already narrows things down.

A completed run: 15.773 seconds, 29,432 tokens, function calling mode, two iterations, one tool used twice.
Completed in 15.773s across two iterations, 29,432 tokens, one tool called twice.

Four things are worth reading off that header before anything else:

  • Completed, not failed
  • Two iterations
  • One tool, called twice
  • 688 then 28.7K tokens
What the run header tells you before you open a single step.

The run completed. No step errored. That rules out the whole class of problems people usually look for first, and it points at the other class: every step did what it was told, and the result was still wrong.

The token split is the tell. The first model call spent 688 tokens deciding what to do. The final one spent 28,744 writing the answer up. Almost all of the work was presentation, built on top of whatever the two tool calls returned. So the question becomes: what did they return?

Open the tool call

This is the step most debugging never reaches, because the final answer gives you no reason to suspect it.

The query_healthcare_data call: the arguments the agent generated, and the records that came back.
What the agent asked for, and what the tool handed back.

The arguments are fine. The agent asked for the right patient and the right resource:

Input · what the agent generated
{
  "subject_id": "10014354",
  "resource": "transfers"
}

The response is where it falls apart. Alongside the records, the tool returns its own account of how much it gave you:

Output · the part that matters
"total":       3903,
"returned":    100,
"limit":       100,
"offset":      0,
"complete":    false,
"truncated":   true,
"next_offset": 100

The tool was explicit. There are 3,903 records. It returned 100. It is not complete, it is truncated, and here is where to carry on from.

The agent wrote its summary from the first hundred and told the user the results looked complete. Every date in that answer is real. The conclusion drawn from them is not.

Why a good model still gets this wrong

It is tempting to call this a model failure. It is more useful to call it an interface failure, because that is the part you control.

The completeness flags arrive as two fields in the middle of a large JSON payload, surrounded by exactly the records the model was asked to summarise. Nothing in the prompt told it that truncated outranks the content. Nothing stopped it from answering. A model optimising for a fluent, complete-looking reply will give you one.

Tool returns 100 of 3,903
truncated: true, next_offset: 100
Model summarises what it has
Confident, partial answer
The flags were present at every step. Nothing was built to act on them.

There are three honest fixes, and they are not alternatives so much as layers:

  • Page in the tool, not the prompt. If a caller asks for a patient’s transfers, the tool should follow next_offset until it is done, and only then return. Correctness belongs below the model wherever it can live there.
  • Make truncation loud. If the tool must return partial data, say so somewhere the model cannot treat as background: a status line ahead of the records rather than a flag buried after them.
  • Make the agent state its evidence. The original request asked it to say whether the results were complete. Requiring it to answer that from the flags, as a separate claim, turns a silent assumption into something reviewable.

None of those are reachable until someone opens the tool call. That is the whole argument for keeping the run.

The same read, for deterministic runs

Autonomous agents get the attention here because their path varies. Workflows have the opposite problem: the path is fixed, so when the output is wrong the question is which step changed the data.

A completed workflow run: trigger, custom function, HTTP request and end, each with its own duration and its input and output payloads.
Four steps, 1.641s. Each one opens on its own input and output.

Same discipline, fewer unknowns. The trigger shows what arrived. Each step shows what it received and what it passed on. The step that broke the shape of the data is the one where the input looks right and the output does not.

Worth noticing in that run: the same truncated and total: 3903 appear in the workflow’s payload too. It is a property of the data source, not of the agent. Anything built on that tool inherits it.

What to take from this

The specific bug is a pagination bug, and pagination bugs are old news. What is new is how well an LLM hides one.

A traditional pipeline that reads 100 of 3,903 rows produces something visibly short. An agent produces a fluent answer of the usual length, in the usual shape, with real values in it. The failure is invisible at the only layer most teams look at.

A confident answer is not evidence of a correct one. The run is.

Three things worth carrying into whatever you are building:

  • Completed is not correct. A run with no errors still deserves a look when the output is disputed, and that is precisely the run nobody checks.
  • Debug from the tool boundary outwards. Most surprising answers come from what a tool returned, not from the model’s reasoning about it. Open the call before you rewrite the prompt.
  • Keep the runs you did not think you needed. You cannot reconstruct a non-deterministic path after the fact. Either it was recorded or the investigation starts with a guess.

If you want to see this on your own agents, observability is part of the same platform that runs them, so there is nothing to instrument. Start free with $10 in credit and open the first run you disagree with.

XPECTRUM / ENGINEERINGMore from Xpectrum ↗