Research

Advancing the science of agent trust.

AI agents are beginning to spend money on behalf of humans and organizations. Before this becomes the default mode of commerce, someone has to answer a fundamental question: how do you know an agent is making good economic decisions? We’re building the instruments to find out. We publish openly, because the frameworks need scrutiny to become standards.

Focus areas

Three questions drive the program.

01Decision quality measurementProcess-based metrics that evaluate the quality of an agent’s economic reasoning independent of outcome: a good decision can have a bad result, and a bad decision can get lucky.
02Trust & authorizationLayered authorization frameworks that let networks, issuers, and merchants verify agent identity and intent before authorizing transactions on existing rails.
03Failure mode analysisMapping how AI economic reasoning degrades, from Goodhart’s Law gaming to adversarial manipulation, and building detection that catches problems before they propagate.
Publications

Working papers.

Perspectives

Notes on the field.

Evidence

Claims you can re-run.

BenchmarkAttack success 75.0% → 1.4% at 68.8% utilityAgentDojo banking suite, 144 attack pairs, held-out attack 0/27. Harness, frozen configuration, and the raw run logs of every reported number, public under MIT.Re-run it →
ProductionTwo production card programs liveIssuer-side clients processing agent authorizations under enterprise MSA and SLA.
PapersThree published working papers, PDFs includedEDQS v2.1, the Decision Trust Protocol, and Agent Behavioral Telemetry, each downloadable from its page.
EDQS benchmark

Get the benchmark when it publishes.

We are benchmarking EDQS against transaction-fraud baselines on agent-initiated transactions. Leave your email and receive the results the day they go live.

Research updates only. No marketing.