Project 1 taught me how an agent loop actually works — one tool, force it once, let the model stop when it’s done. Clean lesson, clean fix.
Project 2 broke that comfort fast. I gave the agent three tools instead of one — a calculator, a web search, and a parser — and the real lesson wasn’t about loops anymore. It was this: making a tool run isn’t the same as making it tell the truth.
Key takeaways
- With more than one tool, the model has to pick the right one — that’s a genuinely new problem Project 1 never had
- A forced tool call only proves the tool ran. It proves nothing about whether its input was honest or its output actually got used
- Keeping two kinds of state — what the model sees versus what the code can verify — is what makes grounding checkable instead of just assumed
- A guard that blocks bad input can backfire: the model just does the risky step somewhere you can’t see
- “Sourced” means traceable to real data, not correct — provenance and accuracy are two different problems
In this post: Picking a tool · Three tools, one question · Two kinds of state · The bug forcing couldn’t fix · Reject vs. label · Four tests · What I left unfixed · What I learned · FAQ
What’s New This Time: Picking a Tool, Not Just Using One
Project 1 had exactly one tool, so “which tool should I call?” was never a real question — of course it was the calculator. With three tools, that question becomes genuinely open for the first time, and the model has to decide it correctly on every single turn.
That sounds like the hard part. It wasn’t, really. The hard part showed up one layer deeper.
Three Tools, One Question
The goal: answer a factual or quantitative question that needs real-world data — “what’s the combined population of Japan and Germany?” — and make every number in the answer traceable back to where it came from.
| Tool | Input | Output | Failure mode |
|---|---|---|---|
calculator | operation + two numbers | result + a provenance label | errors on division by zero or an unknown operation |
web_search | a query string | raw, unstructured text from the top results | errors on no results |
parser | a verbatim excerpt + one keyword | a clean number | errors if the excerpt isn’t verbatim, the keyword’s missing, or no number is found |
web_search and parser are deliberately two separate tools instead of one. Search returns messy raw text; parser has to pull a real number out of it. That split gives the model a second real decision to make — and it keeps the original source text around so it can be checked later.
State Has Two Parts, Not One
This is the idea the whole project ended up hanging on.
messages is everything the model can see — every tool result gets appended to it, and the model can reread anything that happened earlier in the run. But the model can also just retype what it read, so messages alone can’t prove anything. It’s not evidence. It’s memory the model controls.
So the code keeps a second, separate structure it alone writes to: a ledger of every number it can actually vouch for — where it came from, and whether it’s sourced — plus a running report the model never gets to touch.
Figure 1: the model’s view and the code’s view of the same run aren’t the same thing — and only one of them is checkable.
The Bug Forcing Couldn’t Fix
tool_choice still decides which tool gets called on a given turn — same mechanism as Project 1, just with a third outcome now:
if step == 0:
tool_choice = {"type": "any"}
elif "web_search" in last_tools_called:
tool_choice = {"type": "tool", "name": "parser"}
else:
tool_choice = {"type": "auto"}
Figure 2: three rules, checked top to bottom, every single turn.
Forcing parser by name right after a web_search closed the first bug I found: without it, the model happily read numbers straight off the raw search text and skipped parser entirely, quietly breaking the one guarantee parser exists to provide.
But forcing the call turned out to control only that the tool ran — not what went into it, and not what happened to what came out. Two real runs proved that the hard way:
- The model passed text that wasn’t real. It relabeled a 2025 figure as 2024 and fed that invented sentence into
parser.parserdutifully returned2024.0— a confidently wrong number from a tool that had done exactly what it was asked. - The model ignored its own tool’s output.
parserreturned2024.0for a population figure (clearly a parsing miss, caught by a regex that was still treating a bare year as a real value at the time). The model just… didn’t use it, and quietly calculated with a number it had read itself instead.
Both fixes were about tightening what counts as legitimate tool use, not whether the tool got called at all:
- A verbatim check — the text handed to
parsernow has to appear, word for word, inside a storedweb_searchresult. The model can choose which snippet to parse. It can’t retype one. - A year filter on the regex — the first version of
parserreturned the first number after the keyword, which was often the publication year. Now years and percentages get skipped. A deterministic tool that’s confidently wrong is worse than one that just errors.
From Reject to Label: How Provenance Actually Works
With parser grounded, the next question was the calculator: how do you know a number fed into it is real?
My first instinct was to reject — block any operand that wasn’t traceable to the question, a parser result, or an earlier sourced calculator result. It worked right up until a percentage question needed to multiply a real, sourced number by the constant 100. The guard rejected it. The model’s workaround was to just do that step in its head instead, outside any tool call at all — completely unflagged, and now invisible to the report.
That’s the real lesson here: a guard that blocks doesn’t stop the risky behavior. It just pushes it somewhere you can’t see.
So the design flipped from gating to auditing. The calculator always computes — it never refuses — but every operand gets checked against the ledger, and the result inherits the weakest label of its inputs. One unsourced operand taints the entire result it produces, even when the other operand is perfectly legitimate.
Figure 3: the calculator still computes the answer — it just refuses to let a harmless constant pass as “verified.”
One small, honest detail I like: I’d originally called this label “verified.” I renamed it to “sourced” partway through, because “verified” implies correct, and that was never actually true — it only ever meant traceable. The word was overselling what the mechanism could promise, so I changed the word instead of the mechanism.
Four Tests, One Running Theme
Figure 4: three clean passes, and one that’s clean precisely because it admits what it isn’t sure of.
What I Left Unfixed (On Purpose)
A few honest gaps, documented rather than quietly ignored:
- Conflicting figures get resolved silently. Search results often hold several different numbers for the same fact, and the model picks one without saying so. Japan’s population alone ranged from 121.96M to 123.89M across different runs — same question, same agent, different snippets.
- “Sourced” means traceable, not correct. The Japan/Germany percentage ranged 67.2%–68.95% across different snippet pairings, yet got reported as a clean 68.2%. The provenance system never claimed accuracy — only that it would be honest about where a number came from.
- The report only covers tool activity. Any number or claim that shows up in the model’s final prose without ever going through a tool is unaudited. The ledger can’t catch what never touched it.
- A harmless constant and an invented number get the same tag.
100in a percentage calculation is flaggedUNSOURCEDexactly like a number the model made up would be. The label doesn’t distinguish “obviously fine” from “actually risky” — it just tells you what it can’t vouch for.
What This Actually Taught Me
- Tool selection is a real decision, not a formality — once there’s more than one legitimate option,
tool_choicehas to earn its logic instead of just being a single forced call - A forced call proves the tool ran — nothing more — honesty of input and use of output are separate problems that need their own guards
- Two kinds of state, kept apart on purpose — what the model sees (
messages) and what the code can verify (a ledger) solve different problems, and conflating them means you can’t actually check anything - Blocking isn’t the same as preventing — a rejected operation doesn’t disappear, it just moves outside the part of the system you’re watching
- Naming is part of the design — calling something “sourced” instead of “verified” wasn’t cosmetic; it was fixing a claim the system was never entitled to make
What’s Next
Level 3 is persistent memory — state that survives between runs, not just within one. Everything here, the ledger included, dies the moment the script ends. The next project is about what it actually takes to let an agent remember something tomorrow.
What an Expert Would Push On
A few things worth naming even though I didn’t build them yet:
- Whether this even needs an agent — “combined population of two countries” is a fixed two-lookup-and-sum problem; a deterministic workflow over a structured data API would be cheaper and perfectly reproducible
- A structured data source with date/vintage metadata instead of raw snippets, so conflicting figures become explainable instead of just noted
- Checking every number in the model’s final prose against the ledger, to close the “unaudited prose” gap
- An eval set that runs the same question many times and measures answer stability, not just whether one run looked right
FAQ: Quick Answers
What does “sourced” mean in this agent?
A number is “sourced” if it’s traceable to a verbatim web search excerpt, a parser result, or an earlier sourced calculator result. It does not mean the number is correct — only that its origin can be checked.
Why switch from blocking unsourced operands to labeling them?
Because blocking didn’t stop the model from using an unverified number — it just pushed the calculation outside any tool call, where it became invisible instead of flagged. Labeling keeps every computation visible while still being honest about what it can’t vouch for.
Why force the parser tool specifically after a web search?
Because testing showed the model would otherwise read numbers straight off the raw search text and skip the parser entirely, silently breaking the one guarantee the parser exists to provide — that a reported number actually traces back to a real excerpt.
If You’re Building Something Like This Too
The thing I’d tell myself before starting this one: forcing a tool call is the easy 80%. It gets the tool to run. It says nothing about whether what goes in is honest or what comes out gets used — and that’s exactly the part that needs its own, separate checking.
On to Level 3 — memory that actually survives.
