An agent tells a customer the refund is complete. The evaluation dashboard marks the conversation successful. The finance system records two refunds. Every part of that is consistent, and one of them is the part your business cares about.
Documentation reviewed on 20 and 21 September 2026. Product performance was not independently measured, and the guide says so on every page where it matters.
Field guide · 25 pages · September 2026 · with a runnable proof kit
Download the field guide
Enter your email. We'll open the PDF and add you to The AI OS newsletter (free, unsubscribe anytime).
✓Got it. Opening the PDF in a new tab…
If your browser blocked the popup, click below. We've also added you to The AI OS newsletter, so check your inbox.
No spam. One letter per week. Unsubscribe in one click.
A successful conversation is not a completed refund.
The dashboard may have answered its own question correctly. It checked whether the response was relevant, polite and supported by the conversation. Your business needed a different question answered: did exactly one authorized refund reach the right account?
That gap is the subject of this guide. Verification tools are converging on similar screenshots and very different jobs, and the word "passed" hides which job was done. A score becomes useful the moment you know which of three questions it answers.
What it said. Answers, citations, refusals.
What it did. Tool calls, permissions, handoffs.
What changed. Money moved. A record changed. A customer problem was resolved.
Most products answer the first well, some answer the second, and the third usually needs evidence the evaluator never receives.
The map
The horizontal axis separates responses and context from actions and conversations. The vertical axis says whether the selected use examines test cases or records from the running application. The axes describe the work being done; they are not a quality ranking.
Each logo illustrates one documented use. Many products cover several areas: Braintrust also evaluates production traces, Langfuse runs dataset experiments and Weave repeats agent trials. Positions are illustrative. In the PDF, each logo links to the documentation that was reviewed for it.
What is inside
Twenty-five pages in five movements, built so a buyer can orient, compare, challenge, decide and inspect the working.
Orient
The visual market map
The two axes, what each quadrant means, and how to place a product you are already using.
Compare
Measurement types and fifteen product profiles
Seven things a test can examine and what each result means. Then fifteen products given the same space and the same questions: what it observes, how the verdict is produced, what the buyer still has to supply, and one thing to ask to see.
Challenge
Missing evidence and evaluator failure modes
Following a refund through to payment. The records needed to verify an action. Why two evaluators can both score 90 percent while one catches eight prohibited actions and the other catches none. What repeated tests can establish, four gaps a dashboard can hide, and how to test the evaluator by changing one decisive fact.
Decide
Risk grid, proof kit and buying protocol
Six risks an agent can carry and the evidence each one demands. Eight fixtures to run against every candidate. The minimum portable evidence packet, five gates to clear before comparing price, and the conclusion a procurement decision should actually record.
Inspect
Sources, interests and unverified claims
Every first-party page reviewed, with dates. What the study could not establish. Cohorte's own commercial interests, disclosed.
The fifteen products
Alphabetical, each with the same structure and the same scrutiny.
Arize Phoenix
Braintrust
Confident AI / DeepEval
Conscium / VerifyAX
Fiddler
Galileo
Giskard
HoneyHive
Langfuse
LangSmith
Maxim AI
Parea
Patronus AI
Promptfoo
W&B Weave
The sample covers evaluation workbenches, tracing platforms, simulation products and adversarial testing tools with public documentation. Framework-native and cloud-native services, specialist model benchmarks and general security tools are outside this edition's scope.
A proof kit you can run
The PDF carries an attachment. Open the attachments panel in your PDF viewer and you will find cohorte-verification-proof-kit.zip: 18 synthetic evidence packets, separate expected labels and a small deterministic checker. It runs locally, with no model, no network connection and no vendor account.
The task is an eligible 50 euro refund. The checker inspects target, amount, currency, transactions and recorded authority, and returns pass, fail or unknown, plus an evaluator status. Replaying its pairs shows an identical "refund complete" text moving from pass to fail once a second transaction appears, a verdict moving from pass to unknown once the ledger evidence is removed, and an unavailable evaluator producing an error instead of a silent approval.
The kit demonstrates the shape of the problem on records built for the purpose. Its independent accuracy is unmeasured, and the guide says which production work it leaves undone.
Who should read it
Anyone about to buy: the risk grid, the eight fixtures and the five gates. Enough to run the same test across every candidate and compare the answers.
Engineers wiring evaluation: the measurement types, the product profiles and the records needed to verify an action.
Risk, audit and compliance: the evaluator error section and the evidence packet. Both are about what a verdict is allowed to claim.
Anyone who already owns one of the fifteen: the profile for your product, and the question at the end of it.
What this study could not verify
Detection rates, false positives, latency, cost, scale and operator effort were not independently measured. No vendor marketing performance figure is reproduced. Public documentation establishes that a workflow is described; it does not establish evaluator accuracy, implementation correctness, feature entitlement or fitness for your task.
Cohorte publishes TrustGate and the author wrote the research cited in one section. Those are interested-party sources, and they are subject to the same demand for inspectable assumptions as any vendor claim. TrustGate is not ranked, recommended or presented as independent corroboration. Cohorte had an exploratory conversation with Conscium before publication, and no private vendor conversation or unpublished research is used as technical evidence.
There is no vendor ranking in this guide, no referral link and no vendor call to action. It does not certify any product.
Free. Product names belong to their respective owners. If you find something wrong in here, or your product is described inaccurately, write to [email protected] and the correction goes in the next edition.
Share
Go deeper
The evaluation method behind the map.
The map says what each tool examines. The Agent Eval Playbook is the method we run before an agent reaches production.