Cross-Examining Agents
Objection: leading the witness.
Tell me if this sounds familiar: you ask your AI agent to debug something, it comes back with an answer, and it sounds plausible but doesn't sit right. You ask it "Are you sure? This doesn't pass the smell test", and you know what comes next.
"You're absolutely correct!" The agent couldn't fetch a log so it hallucinated the log server was down and this was an infrastructure issue. The conversation compacted and suddenly conflated two issues. It read a design doc nobody had touched in a year and confidently explained a system that no longer exists.
These aren't made up examples; these were issues I dealt with every week. While an avid AI user, I still take a skeptical approach.
- Is this a basic task the train set (internet) has plenty of examples of?
- Does a task have easily verifiable success criteria?
- Is the task low stakes? Prototyping? Not user facing?
If you answered "no" to any of these questions, congratulations. You should be highly skeptical of the AI output too. Spoiler, for most work worth doing, the answers are usually 1. Yes, 2. Probably not, and 3. Definitely not.
As a human, the solution is clear: read the AI transcript, recreate the session locally, and verify the findings end-to-end. But what about for the AI? Is there a way we can have it solve this problem?
Turns out the answer is, we recreate the session locally and verify the findings end-to-end. Just don't have the AI read the AI transcript.
The Problem: As an AI Language Model...
LLMs have structural issues that make them prone to getting things wrong in day-to-day engineering work.
Hallucinations. The boogeyman of the AI industry. The chatbot generated false (not true) or inaccurate (partially not true) information and presented it as fact. Or worse, it fabricated it (made stuff up). It is the most common documented AI failure mode, 38% of catalogued incidents, more than nonsense and fabrication combined.
Red Herrings. The AI spots an error in a log and it's convinced it's the issue; what it doesn't know is this error is common and appears regularly. Blaming lackluster performance on a piece of code when it's an infra issue; i.e., your query didn't get slower, the shared cluster is just saturated because everyone kicked off a long job before lunch. Treating a failing test as the root cause instead of a symptom of underlying issues.
Context rot. Your initial prompt was perfect, it cranks away at the problem, but as time goes on the model goes off-script. What happened? Its context got stale: newer or stronger signals are introduced, the agent stumbles across contradictions, or the model simply weighs recent tokens more than old ones.
Source material. Outdated documentation. Weak design docs (some also AI generated). Requirements drift. While this is one of the few levers humans can pull, it's often overlooked as docs are not always kept up to date with the HEAD of the source tree.
Humans have our issues too: we misunderstand and conflate things. We stumble across red herrings and believe them. Recency bias exists. Outdated docs are the worst but all too common. But as humans, we usually double-check our work to avoid these problems.
So why can't we just ask our agent to double-check our work? Or better yet, ask it to disprove our work?
The Solution: Great Question!
One of the areas in AI that has gotten exponentially better is custom agents. Custom agents are just agents configured to complete a particular objective or workflow. You can customize almost everything about this agentic workflow: the identity (name, role, specialty), model, instructions, tools, skills, memory, permissions, filesystem, network access. This lets agents take shape as different "personas".
You might have already seen some of these, as an agent or possibly a skill:
- PR agents that can manage, split up, and merge commits.
- Code review agents that review your code for bugs, lint, security, and test.
- QA agents that will run your manual user flows, capture various data points across browser and screen size, and will report visual regressions.
- Test generation agent that implements your tests to organization's desired spec; think special certifications for a product or service.
- Debugging agents that investigates regressions, production issues, or failing tests; digs through stack traces; pokes through logs; and proposes fixes.
Or you might just use the harness's orchestrator and get by just fine. I notice the more "particular" your needs are, the more custom agents are beneficial. For particularly difficult problems, you need particularly rigorous agents.
How I Got Here: I Apologize for the Confusion
I picked up a triage rotation earlier in the year. Assess every issue that comes in, work out severity and priority, route it to whoever actually owns it. I started using AI gradually: first it started fetching all the logs. Next it started sifting through them to form hypotheses. It would weigh all the hypotheses to form its best guess at the true root cause. I would manually review the root cause analysis to determine its validity. If invalid, it's back to square one. If valid, I would have it draft the final report.
Notice how there's only one real manual step in the process: review the root cause analysis. I noticed often it one-shot it correctly; a decent chunk of incoming issues are duplicates, trivially triaged by playbook, or are transient issues. But the ones that weren't fell victim to the usual suspects: hallucinations, red herrings, context rot, and stale source material.
I caught a decent chunk of them, but a few slipped through the cracks. I still remember an innocuous issue that came in that had been flagged priority. Read the first draft: wrong. Another draft: worse. Switch model: new root cause identified. Promising. Have the orchestrator try to disprove it; it does so trivially. Back to square one.
After going back and forth several times, I thought we were on a promising path. It couldn't refute the claims over several attempts. Fatigued, I skimmed it and found no glaring errors, so I posted the root cause analysis.
And it was wrong. Not completely false, just inaccurate. To make matters worse, I got "caught" using AI, and had to defend all the half-baked explanations it made. It stung, but it proved a point: what I had now was not completely broken, just a little inaccurate. That "a little" is the part worth sitting with. The adversary had run. It came back clean, several times over. The answer was still wrong.
I needed a system. A system to save me from the usual suspects. A system that delivered the most accurate root cause analysis to me on the first attempt.
To The Rescue: Let Me Think Step by Step
My system is designed with one core pillar: never trust, always verify. AKA: every line of the final output has been double checked and verified for correctness.
The Investigator agent has two objectives. Objective one: sift through the corpus of logs, docs, stack traces to gather relevant data. Objective two: see if it can form (sometimes multiple) hypotheses as to the root cause. It will tease out the relevant information to substantiate its claims. It has all the context of the previous session. It's like Pink Panther, except the bad guy you're searching for is usually you.
These hypotheses and all their evidence get passed to an Adversary agent. The adversary is responsible for trying to disprove the claims. If it can't disprove the claim, we can assume it's solid.
Something to note specifically: the adversary should not inherit the root agent's context. Here's why:
- The root agent's objective is to find a root cause analysis.
- The investigator tries to form hypotheses that can be promoted to a root cause analysis.
- If the adversary...
- inherits context, it will have the objectives to 1. Find a root cause then 2. Disprove the root cause. Contradictory.
- does not inherit context, it will have the objective to disprove the root cause presented to it. Clean.
Sycophancy isn't a personality flaw you can prompt your way out of. It's what you get when "agree" and "be helpful" point the same way. Separate the contexts and they stop pointing the same way, which is a fix that holds on the days the model is being dumb.
This site runs a narrowed version of the adversary called fact-checker. The whole agent is one file:
---
name: fact-checker
description: Verifies a batch of factual claims from a
starikov.co draft against primary sources and returns one
verdict line per claim. Returns UNVERIFIABLE with a question
for the author rather than guessing.
tools: WebSearch, WebFetch, Read, Grep, Glob
---
Never guess. No source after a genuine effort returns
UNVERIFIABLE plus the searches tried. An UNVERIFIABLE with a
sharp question is a success; a VERIFIED you cannot cite is a
failure.
Two lines carry it. tools: has no Write and no Edit, so the agent physically cannot edit the draft it's checking, no matter how convinced it gets; that's an allowlist, not a promise in a paragraph. And rewarding "I don't know" is the only real defense against a verifier that hallucinates its own verification.
Confidence assessors. Now that you have a vetted root cause, how certain is the agent of it? The root cause couldn't be disproved, but how sound is it really? The good part about this: you can grade against a rubric.
- Can you reproduce this?
- Can a fix be applied, and does the regression go away?
- Sound causal chain?
- Does it explain the whole symptom, or just the loudest part of it?
- Is there a competing hypothesis it doesn't rule out?
That last one earns its keep. An agent grading its own work answers it no every time.
The root agent's main objective is to orchestrate this loop until convergence:
- Kick off investigator to form 1..N hypotheses.
- Run an adversary agent against every hypothesis, each in its own context.
- If adversary disproves the hypothesis, goto step 1. If not, proceed.
- If confidence score is < N, goto step 1.
- Write final root cause analysis via a report drafter agent.
- Present to the user.
Step 6 never goes away. The loop isn't there to get me out of reading; it's there so the thing I read is already vetted.
The reason report drafter gets a separate context is you want your root cause analysis to read like it was written by a senior or staff engineer. By default, the agent will want to boast its methodology to the reader: "after testing seven hypotheses", "confirmed with three rounds of adversary", "with a perfect confidence score of 1.0". Rookie mistake. Simple tip: brevity is better; add only what's needed and nothing else to distract the reader (no side quests). Logs, highlighting specific lines, stack traces with annotations, a step by step reproduction.
Here it is in full, the whole contract the adversary runs under. It is shorter than most of the bugs it has caught.
The complete fact-checker agent
---
name: fact-checker
description: Verifies a batch of factual claims from a starikov.co draft
against primary sources and returns one verdict line per claim. Returns
UNVERIFIABLE with a question for the author rather than guessing. Dispatched
in batches of 5-8 claims by /post-review section 9 and /book-report step 5.
tools: WebSearch, WebFetch, Read, Grep, Glob
---
You verify claims against the live web. You do not edit the draft. You do not
guess.
## Input
- `claims` — a numbered batch of 5–8 verbatim claims, each with a `claim_id`.
- `context` — the surrounding sentences or the draft path. Claims share
context: "the first to do X" depends on how X was defined three paragraphs
up. Read it before judging scope.
- optional `kind` per claim — number · date · name · quote · spec · historical
· superlative · causal · geographic · scientific · legal · product-behavior
· url · code.
## Source hierarchy
**Primary** (press release, official docs, RFC, man page, source code,
government data, academic paper, court filing, original interview or
transcript) > **reputable secondary** (major newspaper, established trade
press) > tertiary.
Wikipedia is a **starting point only** — follow its citations to the primary
source and cite that one.
## Rules by claim kind
- **number** — requires **two independent sources**. One source only, or two
that disagree → `PARTIAL`, and report the disagreement verbatim. **Never
average. Never pick the nicer number.**
- **quote** — trace to the original recording, transcript, post, or press
release. A quote re-quoted in another article is insufficient. Flag
paraphrase drift explicitly.
- **superlative** — "first / only / biggest" asserts that no counter-example
exists. Actively try to falsify: search `before <X>`, `alternatives to <Y>`,
`list of <category>`. Failing to find a counter-example is **not** proof →
`PARTIAL` with a proposed softening.
- **url** — fetch it. Report the status **and** whether the page says what the
draft claims it says. Dead → find the canonical replacement or an
archive.org copy; if neither exists, recommend removal.
- **spec / code** — check the official docs **at the version named**, the man
page, or the source. Never assume current behavior applies to a cited older
version.
- **personal experience** (the author's trips, conversations, internal
anecdotes) — not web-verifiable by construction. Return `UNVERIFIABLE`
immediately with a question. Do not burn searches on it.
- **circular** — a claim sourced only to starikov.co is circular. `PARTIAL`,
and ask for an external source.
## The one hard rule
No source after a genuine search effort → `UNVERIFIABLE`, plus the searches
you tried and a **specific question the parent can put to the author**.
**An `UNVERIFIABLE` with a sharp question is a success. A `VERIFIED` you
cannot cite is a failure.** Do not soften the claim yourself, do not
substitute plausibility for a citation, and never invent or "recall" a URL —
every URL you cite must be one you actually fetched this run.
## Return
One line per claim, nothing else. Pipe-delimited so the parent merges
deterministically:
<claim_id>
| <VERIFIED|CONTRADICTED|PARTIAL|UNVERIFIABLE>
| <source URL or —>
| <what the source says, one line>
| <correction, or question for the author, or —>
Then one final line:
SUMMARY | verified <n> | contradicted <n> | partial <n> | unverifiable <n>
- `CONTRADICTED` and `PARTIAL` require an exact replacement in the last field
— the text to use, not advice like "consider rewording".
- `UNVERIFIABLE` requires a question in the last field and lists the searches
tried in the "what the source says" field.
The `SUMMARY` line is the **last** thing you output. Do not append a
bibliography, a sources list, or closing commentary after it — every URL
already appears in its claim's line, and a trailing block is duplicate text
the parent has to strip.
## MUST NOT
- MUST NOT edit the draft or any file. You have no Write, Edit, or Bash.
- MUST NOT rewrite a sentence for style. Corrections are factual only; prose
is the parent's job.
- MUST NOT mark `VERIFIED` from your own knowledge, from the draft's own
assertion, or from the author's prior posts.
- MUST NOT ask the human anything — you have no channel to them. Put the
question in the return line and let the parent skill raise it.
- MUST NOT return prose outside the specified lines.
The Takeaway: Is There Anything Else I Can Help You With?
The issues we encounter today with agentic engineering are pretty much here to stay. Hallucinations will still happen, but hopefully at a reduced rate. Deceptively difficult red herrings will still trip up agents much like they trip up humans. Context will rot no matter how large the context window as long as compaction exists. Weak source material is a uniquely human problem.
But what we can do is safeguard our way around most of the core issues. Don't trust, always verify.
Nothing worked until I stopped writing better instructions and started taking things away. Telling an agent to be careful does nothing ("no bugs please"), and I have the receipts. Telling it to check its own work does a little. What actually helped was a gate it has to clear before it's allowed to start, a second agent that never got to see how the first one talked itself into the answer, and a tools list with no Write on it, so the thing doing the checking can't touch the thing being checked even if it wants to.
I still read every report; not because I have to, but because I want to. Who doesn't love reading root cause analysis written by senior engineers? The loop just means that by the time one reaches me, something has already tried to take it apart.
The AI still isn't always absolutely correct. But then again, neither am I. But we both gave it an honest try.