Project 1 taught me how an agent loop actually works — one tool, force it once, let the model stop when it’s done. Clean lesson, clean fix.

Project 2 broke that comfort fast. I gave the agent three tools instead of one — a calculator, a web search, and a parser — and the real lesson wasn’t about loops anymore. It was this: making a tool run isn’t the same as making it tell the truth.

Key takeaways

  • With more than one tool, the model has to pick the right one — that’s a genuinely new problem Project 1 never had
  • A forced tool call only proves the tool ran. It proves nothing about whether its input was honest or its output actually got used
  • Keeping two kinds of state — what the model sees versus what the code can verify — is what makes grounding checkable instead of just assumed
  • A guard that blocks bad input can backfire: the model just does the risky step somewhere you can’t see
  • “Sourced” means traceable to real data, not correct — provenance and accuracy are two different problems

In this post: Picking a tool · Three tools, one question · Two kinds of state · The bug forcing couldn’t fix · Reject vs. label · Four tests · What I left unfixed · What I learned · FAQ


What’s New This Time: Picking a Tool, Not Just Using One

Project 1 had exactly one tool, so “which tool should I call?” was never a real question — of course it was the calculator. With three tools, that question becomes genuinely open for the first time, and the model has to decide it correctly on every single turn.

That sounds like the hard part. It wasn’t, really. The hard part showed up one layer deeper.

Three Tools, One Question

The goal: answer a factual or quantitative question that needs real-world data — “what’s the combined population of Japan and Germany?” — and make every number in the answer traceable back to where it came from.

ToolInputOutputFailure mode
calculatoroperation + two numbersresult + a provenance labelerrors on division by zero or an unknown operation
web_searcha query stringraw, unstructured text from the top resultserrors on no results
parsera verbatim excerpt + one keyworda clean numbererrors if the excerpt isn’t verbatim, the keyword’s missing, or no number is found

web_search and parser are deliberately two separate tools instead of one. Search returns messy raw text; parser has to pull a real number out of it. That split gives the model a second real decision to make — and it keeps the original source text around so it can be checked later.

State Has Two Parts, Not One

This is the idea the whole project ended up hanging on.

messages is everything the model can see — every tool result gets appended to it, and the model can reread anything that happened earlier in the run. But the model can also just retype what it read, so messages alone can’t prove anything. It’s not evidence. It’s memory the model controls.

So the code keeps a second, separate structure it alone writes to: a ledger of every number it can actually vouch for — where it came from, and whether it’s sourced — plus a running report the model never gets to touch.

Two kinds of state in one runLeft panel: messages, what the model sees — every tool result is appended and the model can reread it, but can also retype it, so it isn’t provable on its own. Right panel: the ledger, what the code can verify — raw search text, sourced numbers, and report lines the model can’t edit, which is what actually gets checked.MESSAGESwhat the model sees• every tool_result gets appended• model can reread earlier results• but can also retype what it reads→ not provable on its ownLEDGERwhat the code can verify• raw search text, stored verbatim• sourced numbers (parser, calculator)• report lines the model can’t edit→ this is what gets checkedBoth reset at the end of the run — this is State, not Memory.

Figure 1: the model’s view and the code’s view of the same run aren’t the same thing — and only one of them is checkable.

The Bug Forcing Couldn’t Fix

tool_choice still decides which tool gets called on a given turn — same mechanism as Project 1, just with a third outcome now:

if step == 0:
    tool_choice = {"type": "any"}
elif "web_search" in last_tools_called:
    tool_choice = {"type": "tool", "name": "parser"}
else:
    tool_choice = {"type": "auto"}

How tool_choice is decided each turnA three-rule cascade. If it is turn zero, tool_choice is any, so some tool must run. Otherwise, if web_search was the last tool called, tool_choice is forced to parser so the number gets grounded. Otherwise tool_choice is auto, so the model can choose freely or stop.Is this turn 0?yesanymust call a tool — model picks whichnoWas web_search thelast tool called?yesforced: parserground the number that was just foundnootherwiseautomodel chooses freely, including stopping

Figure 2: three rules, checked top to bottom, every single turn.

Forcing parser by name right after a web_search closed the first bug I found: without it, the model happily read numbers straight off the raw search text and skipped parser entirely, quietly breaking the one guarantee parser exists to provide.

But forcing the call turned out to control only that the tool ran — not what went into it, and not what happened to what came out. Two real runs proved that the hard way:

  • The model passed text that wasn’t real. It relabeled a 2025 figure as 2024 and fed that invented sentence into parser. parser dutifully returned 2024.0 — a confidently wrong number from a tool that had done exactly what it was asked.
  • The model ignored its own tool’s output. parser returned 2024.0 for a population figure (clearly a parsing miss, caught by a regex that was still treating a bare year as a real value at the time). The model just… didn’t use it, and quietly calculated with a number it had read itself instead.

Both fixes were about tightening what counts as legitimate tool use, not whether the tool got called at all:

  • A verbatim check — the text handed to parser now has to appear, word for word, inside a stored web_search result. The model can choose which snippet to parse. It can’t retype one.
  • A year filter on the regex — the first version of parser returned the first number after the keyword, which was often the publication year. Now years and percentages get skipped. A deterministic tool that’s confidently wrong is worse than one that just errors.

From Reject to Label: How Provenance Actually Works

With parser grounded, the next question was the calculator: how do you know a number fed into it is real?

My first instinct was to reject — block any operand that wasn’t traceable to the question, a parser result, or an earlier sourced calculator result. It worked right up until a percentage question needed to multiply a real, sourced number by the constant 100. The guard rejected it. The model’s workaround was to just do that step in its head instead, outside any tool call at all — completely unflagged, and now invisible to the report.

That’s the real lesson here: a guard that blocks doesn’t stop the risky behavior. It just pushes it somewhere you can’t see.

So the design flipped from gating to auditing. The calculator always computes — it never refuses — but every operand gets checked against the ledger, and the result inherits the weakest label of its inputs. One unsourced operand taints the entire result it produces, even when the other operand is perfectly legitimate.

Provenance taint, traced through a real calculationStep one divides two parser-sourced populations and the result is sourced. Step two multiplies that sourced result by the constant 100, which is not in the ledger, so the result is labeled unsourced even though it is still computed. The final report line prints the number with that label rather than hiding it.STEP 1 · sourceddivide(germany, japan)both operands from parser→ sourcedSTEP 2 · constantmultiply(result, 100)100 isn’t in the ledger→ UNSOURCEDREPORT LINE68.2%[UNSOURCED constant: 100]printed, not hidden

Figure 3: the calculator still computes the answer — it just refuses to let a harmless constant pass as “verified.”

One small, honest detail I like: I’d originally called this label “verified.” I renamed it to “sourced” partway through, because “verified” implies correct, and that was never actually true — it only ever meant traceable. The word was overselling what the mechanism could promise, so I changed the word instead of the mechanism.

Four Tests, One Running Theme

Four tests on the finished research agentTest one, the combined population of Japan and Germany, ran two searches and two parses then a sourced addition, correct. Test two, forty five times twelve, went straight to the calculator with no wasted search. Test three, the population of Japan alone, stopped after search and parse with no needless calculator call. Test four, Germany’s percentage of Japan’s population, divided two sourced numbers correctly but flagged the multiply-by-100 constant as unsourced rather than hiding it.TEST 1 · combined lookup“Combined population: Japan + Germany?”2× web_search → 2× parser → calculator add✓207,183,674 — sourcedTEST 2 · efficiency“What’s 45 times 12?”straight to calculator — no wasted search✓540, correct, efficientTEST 3 · no overreach“What’s the population of Japan?”search → parse → stop, no needless calculator call✓correct, minimal stepsTEST 4 · honest labeling“Germany’s % of Japan’s population?”divide sourced, ×100 flagged UNSOURCED⚠68.2%, constant flagged

Figure 4: three clean passes, and one that’s clean precisely because it admits what it isn’t sure of.

What I Left Unfixed (On Purpose)

A few honest gaps, documented rather than quietly ignored:

  • Conflicting figures get resolved silently. Search results often hold several different numbers for the same fact, and the model picks one without saying so. Japan’s population alone ranged from 121.96M to 123.89M across different runs — same question, same agent, different snippets.
  • “Sourced” means traceable, not correct. The Japan/Germany percentage ranged 67.2%–68.95% across different snippet pairings, yet got reported as a clean 68.2%. The provenance system never claimed accuracy — only that it would be honest about where a number came from.
  • The report only covers tool activity. Any number or claim that shows up in the model’s final prose without ever going through a tool is unaudited. The ledger can’t catch what never touched it.
  • A harmless constant and an invented number get the same tag. 100 in a percentage calculation is flagged UNSOURCED exactly like a number the model made up would be. The label doesn’t distinguish “obviously fine” from “actually risky” — it just tells you what it can’t vouch for.

What This Actually Taught Me

  • Tool selection is a real decision, not a formality — once there’s more than one legitimate option, tool_choice has to earn its logic instead of just being a single forced call
  • A forced call proves the tool ran — nothing more — honesty of input and use of output are separate problems that need their own guards
  • Two kinds of state, kept apart on purpose — what the model sees (messages) and what the code can verify (a ledger) solve different problems, and conflating them means you can’t actually check anything
  • Blocking isn’t the same as preventing — a rejected operation doesn’t disappear, it just moves outside the part of the system you’re watching
  • Naming is part of the design — calling something “sourced” instead of “verified” wasn’t cosmetic; it was fixing a claim the system was never entitled to make

What’s Next

Level 3 is persistent memory — state that survives between runs, not just within one. Everything here, the ledger included, dies the moment the script ends. The next project is about what it actually takes to let an agent remember something tomorrow.

What an Expert Would Push On

A few things worth naming even though I didn’t build them yet:

  • Whether this even needs an agent — “combined population of two countries” is a fixed two-lookup-and-sum problem; a deterministic workflow over a structured data API would be cheaper and perfectly reproducible
  • A structured data source with date/vintage metadata instead of raw snippets, so conflicting figures become explainable instead of just noted
  • Checking every number in the model’s final prose against the ledger, to close the “unaudited prose” gap
  • An eval set that runs the same question many times and measures answer stability, not just whether one run looked right

FAQ: Quick Answers

What does “sourced” mean in this agent?

A number is “sourced” if it’s traceable to a verbatim web search excerpt, a parser result, or an earlier sourced calculator result. It does not mean the number is correct — only that its origin can be checked.

Why switch from blocking unsourced operands to labeling them?

Because blocking didn’t stop the model from using an unverified number — it just pushed the calculation outside any tool call, where it became invisible instead of flagged. Labeling keeps every computation visible while still being honest about what it can’t vouch for.

Why force the parser tool specifically after a web search?

Because testing showed the model would otherwise read numbers straight off the raw search text and skip the parser entirely, silently breaking the one guarantee the parser exists to provide — that a reported number actually traces back to a real excerpt.


If You’re Building Something Like This Too

The thing I’d tell myself before starting this one: forcing a tool call is the easy 80%. It gets the tool to run. It says nothing about whether what goes in is honest or what comes out gets used — and that’s exactly the part that needs its own, separate checking.

On to Level 3 — memory that actually survives.