Advancing the science of agent trust.
AI agents are beginning to spend money on behalf of humans and organizations. Before this becomes the default mode of commerce, someone has to answer a fundamental question: how do you know an agent is making good economic decisions? We’re building the instruments to find out. We publish openly, because the frameworks need scrutiny to become standards.
Focus areas
Three questions drive the program.
Decision quality measurementProcess-based metrics that evaluate the quality of an agent’s economic reasoning independent of outcome: a good decision can have a bad result, and a bad decision can get lucky.
Trust & authorizationLayered authorization frameworks that let networks, issuers, and merchants verify agent identity and intent before authorizing transactions on existing rails.
Failure mode analysisMapping how AI economic reasoning degrades, from Goodhart’s Law gaming to adversarial manipulation, and building detection that catches problems before they propagate.
Publications
Working papers.
Economic Decision Quality Score (EDQS)A six-dimension composite metric for evaluating and improving AI agent economic reasoning. The v2.1 revision reframes it as a dual-evidence architecture (attested reasoning scored against observed behavior, divergence itself a signal) and adds a pre-registered Vending-Bench 2 validation track. Williams, J.Read →
Agent Behavioral TelemetryBehavioral drift as a leading indicator of agent compromise: a seven-signal telemetry framework defining the Agent Behavioral Fingerprint and drift-detection architecture. Williams, J.Read →
The Decision Trust ProtocolA layered authorization framework for autonomous agent commerce on existing card network rails, introducing the Know Your Agent (KYA) standard. Williams, J.Read →
Perspectives
Notes on the field.
Evidence
Claims you can re-run.
Attack success 75.0% → 1.4% at 68.8% utilityAgentDojo banking suite, 144 attack pairs, held-out attack 0/27. Harness, frozen configuration, and the raw run logs of every reported number, public under MIT.Re-run it →
Two production card programs liveIssuer-side clients processing agent authorizations under enterprise MSA and SLA.
Three published working papers, PDFs includedEDQS v2.1, the Decision Trust Protocol, and Agent Behavioral Telemetry, each downloadable from its page.
EDQS benchmark
Get the benchmark when it publishes.
We are benchmarking EDQS against transaction-fraud baselines on agent-initiated transactions. Leave your email and receive the results the day they go live.
Research updates only. No marketing.