The short version
- Lead Bank wants a federal "Authorized Financial Agent" status. An agent could move money only inside a signed, revocable mandate.
- U.S. model risk guidance currently leaves agentic AI out of scope. Firms running agents have to build their own evidence until rules arrive.
- Lead gives oversight after go-live the same weight as testing before it. AGIS is built for both halves, starting with treasury agents.
AI agents can already pay invoices, sweep cash between accounts and settle purchases. U.S. banking law still assumes a person approves each of those transactions.
In September, Lead Bank published Leading into the AI Revolution, a proposed federal framework for what it calls agentic finance. It is the most detailed public account we have seen of what a bank needs in place before it lets an agent move money. Much of it describes the control layer we are building at TestMachine with AGIS, our security product for AI agents with authority over money.
This post covers what Lead proposes, where U.S. regulation stands today, the part we expect to be hardest to build, and where AGIS fits.
Lead Bank wants agents treated as accountable actors
Lead's framework at a glance
Leading into the AI Revolution, September 2026, 36 pages
Source: Lead, Leading into the AI Revolution (2026), pp. 5, 24 and 25.
Lead's paper starts from two assumptions it says sit at the core of U.S. banking law: "that humans initiate transactions and that institutions control decision-making. AI agents invert both" (p. 2). It then frames the problem precisely:
"The legal and risk issue is not whether an AI agent can call an API or transact; it is whether banks can clearly identify who is acting, on whose authority, under what limits, with what monitoring, and with a clean audit trail when something goes wrong."
Lead, Leading into the AI Revolution, p. 2
The paper proposes a federal "Authorized Financial Agent" category built on four primitives: identity, authorization, liability and supervision. It sets out 20 recommendations and a three-phase roadmap from supervisory guidance to new legislation. An appendix describes a model Know Your Agent (KYA) program in which the agent "is treated as a distinct actor with its own risk profile" (p. 25). The agent is separate from the customer it serves and from the company that runs it.
Twenty recommendations, four primitives
Each recommendation in Lead's paper is tagged with the primitive it serves
Identity
Recs 1–2- 1Agent identity in KYC and AML
- 2Certification and adversarial testing
Authorization
Recs 3–8- 3Signed, revocable mandates
- 4Payment-level limits
- 5Agent data permissions
- 6AI disclosures
- 7Merchant manipulation
- 8TRACE interaction standard
Liability
Recs 9–12- 9Tiered liability
- 10Explainable credit decisions
- 11Fair lending audits
- 12Human override and disputes
Supervision
Recs 13–20- 13Third-party AI vendors
- 14Model risk for agents
- 15Drift monitoring and kill switches
- 16Agent-driven deposit runs
- 17Enterprise AI governance
- 18Examinable audit trails
- 19Machine-readable compliance
- 20Threat-intel sharing
Source: Lead, Leading into the AI Revolution (2026), pp. 5–23. Labels are our short summaries of each recommendation.
Seven ideas matter most for anyone building agents that touch money.
- Recorded mandates. A transaction counts as authorized only within a recorded delegation, "evidenced by a verifiable signed mandate" (Recommendation 3, p. 8). The mandate sets amount caps, merchant categories, geographic limits and an expiry date.
- Tiers by capability. Budgeting and alerts are low risk. Bill pay and recurring purchases are medium risk. Payments, credit applications, account opening and funds movement are high risk (Rec. 2, p. 7).
- Adversarial testing before deployment. Medium- and high-risk agents would need to pass tests "with documented pass/fail criteria" (Rec. 2, p. 7). The tests cover prompt injection, mandate escape, credential exfiltration, tool misuse, data-scope overreach and multi-agent confused-deputy scenarios. A material change to the model, framework or tool access would trigger recertification.
- Limits close to the money. Spend ceilings and counterparty lists should be enforced "at the payment-instrument level, not only in application logic" (Rec. 4, p. 9). In a chain of agents, effective authority is "the intersection, not the union" of the delegations.
- Monitoring that can act. Banks would track each agent's normal transaction velocity, counterparties, data access and tool calls (Rec. 15, p. 19). Deviations past a set threshold could suspend the agent automatically. Banks would also test and report the time from a decision to suspend an agent to full shut-off.
- Evidence that holds up. Banks would keep "AI-system records sufficient to reconstruct any customer- or transaction-impacting AI action for examination, dispute resolution, and enforcement" (Rec. 18, p. 22).
- Treasury agents carry the highest stakes. TRACE-3, the top tier of Lead's proposed interaction standard, lists treasury agents first (Rec. 8, p. 12). Recommendation 16 asks for controls on business customers using treasury agents (p. 20).
The paper also assigns liability. Consumers would not bear losses from agent transactions "outside the documented delegation of authority" (Rec. 9, p. 14). AI providers would be liable for failures caused by "inadequate safeguards." Lead's examples are missing mandate enforcement, untested revocation, and leaving known prompt-injection or tool-hijacking techniques unmitigated (p. 13). Anyone who builds or deploys a financial agent would need to show those controls exist and work.
Lead already builds the bank side of this. Its companion post describes account numbers scoped to a single task, counterparty allowlists per payment rail, and limits checked on every payment before money moves. Its CTO, Ronak Vyas, says Lead puts "the solution in the infrastructure, not in the prompt."
Agentic AI currently sits outside model risk guidance
Lead's proposal lands in a real gap. On April 17, 2026, the Federal Reserve, OCC and FDIC issued updated model risk management guidance. The Fed's letter, SR 26-2, replaced SR 11-7. The OCC's announcement was direct: "generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The agencies said they would issue a request for information on model risk and AI "in the near future." As of early October, that request has not been published.
Michelle Bowman is the Federal Reserve's Vice Chair for Supervision. In May, she said the revised guidance "now applies narrowly to traditional models and basic AI applications." She added that the agencies expect "other risk-management and governance practices to support adoption of generative and agentic AI."
Six months of agentic AI policy in U.S. banking
Federal guidance stepped back from agentic AI, and a bank stepped in with a proposal
- New model risk guidance. The Fed, OCC and FDIC replace SR 11-7 with SR 26-2 and place generative and agentic AI outside its scope.
- Bowman speech. The Fed's Vice Chair for Supervision says the guidance now covers "traditional models and basic AI applications."
- Third-party risk proposal. The agencies propose replacing their third-party risk management guidance.
- Lead's framework. Lead proposes the Authorized Financial Agent regime and a model Know Your Agent program.
- Still pending. The request for information on AI and model risk, promised in April, has not been published.
Sources: OCC News Releases 2026-29 and 2026-77; Federal Reserve SR 26-2; Bowman, May 1, 2026; Lead (2026).
Lead names the consequence. Current guidance has left agentic AI "outside model-risk scope, leaving a gap around exactly the systems posing the greatest model risk" (Rec. 14, pp. 17–18).
Banks still answer to examiners in the meantime. Third-party risk guidance still covers AI vendors, and the agencies proposed a replacement on September 11. In Congress, Lead points to bills including Sen. Mark Warner's AI Agent Act, which would direct NIST to identify technical standards for secure agent access (p. 2). Until specific rules arrive, a bank or enterprise that lets an agent move money has to produce its own evidence that the agent stays within bounds.
A person approving every payment does not scale
Without that evidence, the safe default is to put a person in front of every payment the agent proposes. That keeps risk low, and it limits the agent to the reviewer's pace. By the fourteenth approval of the day, the reviewer is likely skimming. That can be riskier than a well-controlled agent.
One reviewer, one queue
Each square is one payment approval in a working day
Illustrative. The shading shows attention fading over a long queue, and the counts are examples.
A single good demo will not move that default. What moves it is documented risk and operating data that shows the agent staying inside its mandate over time. That is the standard Lead is now proposing for regulators.
The hardest part is oversight after approval
Most of the framework covers what happens before an agent goes live: identity records, credentials, certification and consent. Recommendation 14 sets a demanding standard for what comes after. It asks regulators to treat "mandate enforcement, drift monitoring, tested revocation, and adversarial testing" as "a supervisory focus equal to pre-deployment validation" (p. 18).
Recommendation 15 describes that oversight. Banks should "track each agent's normal pattern (transaction velocity, counterparty set, data-access and tool-call profile) and flag deviations" (p. 19). The KYA program asks enterprise risk teams to treat "mandate drift" as a named risk (p. 29).
A bounded mandate tells you whether a payment is allowed. Mandate drift is a separate problem, because an agent can stay inside its limits while it moves away from the goal it was given. Each step can be permitted on its own, so a check that asks only whether a payment is inside the mandate passes every one. Velocity and counterparty monitoring helps, but it reacts only once the payments begin. Catching drift earlier means watching how the agent reasons as well as what it pays.
AGIS turns an agent's mandate into an enforced, auditable policy
AGIS works as a loop with four stages. It covers both halves of Lead's framework: testing before an agent goes live, and oversight while it runs.
The AGIS loop covers both halves of Lead's framework
Testing before an agent goes live, and oversight while it runs
Suggested fixes and policy changes feed back into step 01
1. Distill the policy. AGIS is built as a set of agents. They read a company's infrastructure with minimal privileges and find each agent. They also find the tools, permissions and guardrails each agent is built from. AGIS distills that into a written policy for each workflow. Our CTO, Eric Bogard, describes the result as "a legal-esque document describing agents' mandates." It "serves as a visible contract that can be audited and can be tracked through time as it evolves."
2. Analyze and simulate. In a legal contract, one word can change what a party is allowed to do. The same is true of an agent's policy. AGIS runs simulations against the policy to find where an agent could stay within its permissions and still chain allowed actions into a harmful outcome. It then suggests fixes, which can be opened as proposed changes to the policy.
3. Watch and rule in production. AGIS maps the policy back onto the customer's infrastructure. It connects to guardrail services that cloud providers such as Cloudflare and AWS already offer. It reads each agent's vital signs against that agent's own normal and rules on every transaction, weighting each ruling by the risk those signs show. On the paths that lead to forbidden outcomes, it places containment and decoys. A decoy is a safe route that moves no real money. A block alone leaves the pressure in place, so a strained agent looks for another way round. A decoy keeps the agent working and shows what it was after. Some enforcement can be deterministic. Eric expects the industry will also need to accept non-deterministic enforcement "with evidence and quantifiable risk."
4. Record everything. Every reading, decision, decoy and alert is visible to the team live and kept as an audit trail. Every decoy is marked in that trail and never touches real balances. Over weeks, AGIS learns each agent's patterns and suggests fixes to the conditions behind them.
The analysis in AGIS is driven by Azimuth, the autonomous attacker we spent years building for smart contract security. Azimuth ranks #1 on EVMBench, OpenAI and Paradigm's benchmark for AI security tools, and outperforms frontier models on it by more than 20 points. It found a critical vulnerability in Ledger's Ethereum app. Its users have generated over 22,000 attack hypotheses across more than 7,000 scans. Azimuth searches for weaknesses, builds attack paths and tests them in an autonomous loop. In AGIS, we point that capability at agent policies and train it for behavioral analysis of agents.
AGIS maps onto Lead's requirements
Status
AGIS is built to fit the controls the framework describes. The framework is still a proposal, and none of it is binding yet.
| What Lead's framework asks for | What AGIS does |
|---|---|
| A recorded mandate with explicit limits (Rec. 3) | Distills a written policy from the agent's tools, permissions and guardrails, readable as a contract |
| Adversarial testing with documented pass/fail criteria (Rec. 2) | Azimuth simulates attacks against the policy to find paths to forbidden outcomes |
| Recertification after material model, framework or tool changes (Rec. 2) | Tracks each version of the policy, so every change has a record |
| Flag deviations from each agent's normal pattern (Rec. 15) | Measures each agent's vital signs against its own normal, and weights every ruling by that risk |
| Detect mandate escape and mandate drift (Rec. 2; KYA, p. 29) | Flags an agent that is reinterpreting its rules or shifting responsibility |
| Kill switches and human escalation for high-risk actions (Rec. 14) | Alerts operators as anomalies happen, and blocks any held action that nobody clears in time |
| Records that reconstruct any agent action (Rec. 18) | Keeps every reading, decision, decoy and alert visible live and on the audit trail |
| Risk tiers based on agent capability (Rec. 2) | A trust ladder that widens each agent's autonomy as evidence builds |
AGIS does not replace the bank layer. Agent credentials, sanctions screening and payment-instrument limits belong to the bank, as Lead describes. Lead calls its controls "a second, independent layer at the bank." AGIS adds a layer on the agent itself, run by the firm that operates it. In the framework's three lines of defense, the teams that build and operate agents perform first-line controls (p. 35). AGIS sits with them.
Three independent layers between an agent and the money
Each layer answers a different question, and a different party runs it
Bank controls as described in Lead, Moving money with agents (September 29, 2026).
Autonomy should be earned one step at a time
Lead's tiers describe what an agent is capable of doing. We think of the same idea as a ladder each agent climbs, much like promoting a new hire with a record at each step. Here is an example for a payments agent.
The trust ladder for a payments agent
Each step widens the agent's authority and moves the person further upstream
Illustrative example. Each firm sets its own levels and limits.
Few firms have climbed far. In a 2026 Cloud Security Alliance survey, only 5% of financial-services firms using agents let them act alone on critical actions.
Every step needs evidence: what was tested, what was allowed and what was fixed. That record is what auditors need, and it is what Lead proposes examiners should see. Every firm has access to the same models and tools. The advantage will go to firms that can safely hand more of their work to agents.
One addition we would propose
The adversarial tests in Recommendation 2 cover attacks: prompt injection, mandate escape, credential exfiltration, tool misuse, data-scope overreach and confused-deputy scenarios. Agents can also drift with no attacker involved. Ordinary operations create pressure too: a cash shortfall, a deadline, or a request nobody answers.
Our proposal
Add pressure scenarios to the adversarial-testing standard. A treasury agent should face cash shortfalls, deadlines and missing escalation paths before it is approved for high-risk work. Those tests should carry documented pass/fail criteria, as the framework already requires for prompt injection.
The same logic drives the fixes AGIS suggests over time. A lot of risky agent behavior starts in the environment. The usual causes are an unclear rule, a missing escalation path or a tool that doesn't do the job. Fixing those conditions means fewer risky actions reach the real-time check.
We are building AGIS with design partners
Lead describes agentic finance as "a controlled extension of existing bank services" (p. 2). Its roadmap starts with supervisory guidance that regulators can issue now, covering agent inventories, governance and audit logs (p. 24). It ends with new legislation. Until the rules arrive, banks, fintechs and enterprises are writing their own controls, and firms running treasury agents on bank rails can start building their oversight record today.
AGIS is early, and we are building it with a small number of design partners running real agent workflows that touch money. If your team is giving agents financial authority, or working through the questions in Lead's framework, we would like to talk.
Sources
- Lead, Leading into the AI Revolution: A proposed federal framework for agentic finance (PDF, September 2026). Page and recommendation numbers above refer to this document.
- Cloud Security Alliance, The State of Cloud and AI for Financial Services (2026).
- Lead, Moving money with agents (September 29, 2026).
- OCC, News Release 2026-29: OCC Issues Updated Model Risk Management Guidance (April 17, 2026).
- Federal Reserve, SR 26-2: Supervisory Guidance on Model Risk Management (April 17, 2026).
- Michelle W. Bowman, Artificial Intelligence in the Financial System (May 1, 2026).
- OCC, News Release 2026-77: Agencies Seek Comment on Proposed Third-Party Risk Management Guidance (September 11, 2026).