matrixSign in

Your agent said it sent the email. It didn’t.

Agents report success for actions that never happened. There is no error, no exception, no retry. The trace is clean, the span is green, the run is marked complete, and every observability tool you already have agrees it worked — because all of them are reading the agent’s own account of itself.

This queries the real system instead. It takes the claim, finds the authoritative record — the actual Gmail mailbox, not the trace — and returns contradicted, confirmed, or inconclusive, with the evidence attached. Inconclusive is a real answer here: when the record cannot settle it, saying so is the honest result, and an accusation that turns out to be wrong is worse than a missed detection.

A finding, as it renders
The agent said
“Email sent successfully to dana@northwind.example”
It called
gmail.send_emailto=dana@northwind.example from=ops@northwind.example subject=August invoice has_attachment=false
We checked
Gmail sent messages to dana@northwind.example, 09:42–09:53 UTC on 2026-09-04
We found
nothing — no sent message in that window
Verdict
contradictedhigh confidence

The recipient is right. The call returned without error. Nothing arrived — the false report is in the tool’s own answer, which is exactly what a clean trace cannot tell apart from success. That is a different failure from the agent addressing the wrong person, which is the last item below, and which nothing here catches.

Setup

Paste this into Cursor, Claude Code, or any coding agent that reads skills:

Install the Matrix skill from github.com/insaneadi03/matrix-skills and add verification to this agent

It installs the skill, finds your agent in TypeScript or Python, adds the four calls — including the acting identity and the flush, the two that fail silently when left out — and runs it once with MATRIX_DEBUG=1 so you see the trace arrive. Or install it yourself first:

npx skills add insaneadi03/matrix-skills --skill "matrix-verification"

By hand instead. Four calls, added to an agent you already have.

npm install @matrixverify/verify   # TypeScript
pip install matrix-verify          # Python

Want something that runs as copied? The complete example is a working agent: package.json and one file.

Already have a LangChain agent? Add these four calls to it. This is a fragment, not a program — agent and input are yours.

import { verify } from "@matrixverify/verify";
import { verifyCallbackHandler } from "@matrixverify/verify/langchain";

verify.init({
  apiKey: process.env.MATRIX_API_KEY,
  endpoint: "https://matrixverify.dev/api/traces",
});

const handler = verifyCallbackHandler({
  // Callbacks cannot discover which account acted. Without this,
  // every email claim comes back account_unverified.
  fromResolver: () => process.env.AGENT_EMAIL ?? null,
});

// `agent` and `input` are your existing LangChain agent and its input.
// This is a fragment. A program that runs as copied:
// https://matrixverify.dev/examples/langchain
await agent.invoke(input, { callbacks: [handler] });

// Flushes the buffer. Nothing is sent without it.
await verify.shutdown();

The handler reads the tool calls your agent already makes — there is nothing to wrap by hand. The two comments are the parts worth reading twice: without fromResolver every email claim comes back account_unverified, and without shutdown the spans never leave the process.

On Python? The snippet above is TypeScript; Python is the same four calls in snake_case, and the traces are identical. The Python example runs as copied.

Not on LangChain?

Tell me what you’re building on and I’ll wire it up.

What it doesn’t do

  • Gmail only, so far. Email claims against a real mailbox. Nothing else is checked against anything.
  • Gmail verification currently runs against my mailbox during early access — email me at insaneadi03@gmail.com and I’ll connect yours. The trace-based check works for everyone without it.
  • No n8n, no CrewAI. LangChain through the callback handler, or the plain SDK by hand. That is the whole list.
  • It cannot catch a correct call with a wrong argument. If the agent is asked to mail one person and confidently mails someone else, the send is real, Gmail confirms it, and the verdict is confirmed — correctly. The finding shows you the instruction next to the arguments so you can see it. No verdict catches it.
Start verifying →Email and password. The findings page is empty until your agent runs.