The short version

  • Lead Bank wants a federal "Authorized Financial Agent" status. An agent could move money only inside a signed, revocable mandate.
  • U.S. model risk guidance currently leaves agentic AI out of scope. Firms running agents have to build their own evidence until rules arrive.
  • Lead gives oversight after go-live the same weight as testing before it. AGIS is built for both halves, starting with treasury agents.

AI agents can already pay invoices, sweep cash between accounts and settle purchases. U.S. banking law still assumes a person approves each of those transactions.

In September, Lead Bank published Leading into the AI Revolution, a proposed federal framework for what it calls agentic finance. It is the most detailed public account we have seen of what a bank needs in place before it lets an agent move money. Much of it describes the control layer we are building at TestMachine with AGIS, our security product for AI agents with authority over money.

This post covers what Lead proposes, where U.S. regulation stands today, the part we expect to be hardest to build, and where AGIS fits.

Lead Bank wants agents treated as accountable actors

Lead's paper starts from two assumptions it says sit at the core of U.S. banking law: "that humans initiate transactions and that institutions control decision-making. AI agents invert both" (p. 2). It then frames the problem precisely:

"The legal and risk issue is not whether an AI agent can call an API or transact; it is whether banks can clearly identify who is acting, on whose authority, under what limits, with what monitoring, and with a clean audit trail when something goes wrong."

Lead, Leading into the AI Revolution, p. 2

The paper proposes a federal "Authorized Financial Agent" category built on four primitives: identity, authorization, liability and supervision. It sets out 20 recommendations and a three-phase roadmap from supervisory guidance to new legislation. An appendix describes a model Know Your Agent (KYA) program in which the agent "is treated as a distinct actor with its own risk profile" (p. 25). The agent is separate from the customer it serves and from the company that runs it.

Seven ideas matter most for anyone building agents that touch money.

  1. Recorded mandates. A transaction counts as authorized only within a recorded delegation, "evidenced by a verifiable signed mandate" (Recommendation 3, p. 8). The mandate sets amount caps, merchant categories, geographic limits and an expiry date.
  2. Tiers by capability. Budgeting and alerts are low risk. Bill pay and recurring purchases are medium risk. Payments, credit applications, account opening and funds movement are high risk (Rec. 2, p. 7).
  3. Adversarial testing before deployment. Medium- and high-risk agents would need to pass tests "with documented pass/fail criteria" (Rec. 2, p. 7). The tests cover prompt injection, mandate escape, credential exfiltration, tool misuse, data-scope overreach and multi-agent confused-deputy scenarios. A material change to the model, framework or tool access would trigger recertification.
  4. Limits close to the money. Spend ceilings and counterparty lists should be enforced "at the payment-instrument level, not only in application logic" (Rec. 4, p. 9). In a chain of agents, effective authority is "the intersection, not the union" of the delegations.
  5. Monitoring that can act. Banks would track each agent's normal transaction velocity, counterparties, data access and tool calls (Rec. 15, p. 19). Deviations past a set threshold could suspend the agent automatically. Banks would also test and report the time from a decision to suspend an agent to full shut-off.
  6. Evidence that holds up. Banks would keep "AI-system records sufficient to reconstruct any customer- or transaction-impacting AI action for examination, dispute resolution, and enforcement" (Rec. 18, p. 22).
  7. Treasury agents carry the highest stakes. TRACE-3, the top tier of Lead's proposed interaction standard, lists treasury agents first (Rec. 8, p. 12). Recommendation 16 asks for controls on business customers using treasury agents (p. 20).

The paper also assigns liability. Consumers would not bear losses from agent transactions "outside the documented delegation of authority" (Rec. 9, p. 14). AI providers would be liable for failures caused by "inadequate safeguards." Lead's examples are missing mandate enforcement, untested revocation, and leaving known prompt-injection or tool-hijacking techniques unmitigated (p. 13). Anyone who builds or deploys a financial agent would need to show those controls exist and work.

Lead already builds the bank side of this. Its companion post describes account numbers scoped to a single task, counterparty allowlists per payment rail, and limits checked on every payment before money moves. Its CTO, Ronak Vyas, says Lead puts "the solution in the infrastructure, not in the prompt."

Agentic AI currently sits outside model risk guidance

Lead's proposal lands in a real gap. On April 17, 2026, the Federal Reserve, OCC and FDIC issued updated model risk management guidance. The Fed's letter, SR 26-2, replaced SR 11-7. The OCC's announcement was direct: "generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." The agencies said they would issue a request for information on model risk and AI "in the near future." As of early October, that request has not been published.

Michelle Bowman is the Federal Reserve's Vice Chair for Supervision. In May, she said the revised guidance "now applies narrowly to traditional models and basic AI applications." She added that the agencies expect "other risk-management and governance practices to support adoption of generative and agentic AI."

Six months of agentic AI policy in U.S. banking

Federal guidance stepped back from agentic AI, and a bank stepped in with a proposal

  1. New model risk guidance. The Fed, OCC and FDIC replace SR 11-7 with SR 26-2 and place generative and agentic AI outside its scope.
  2. Bowman speech. The Fed's Vice Chair for Supervision says the guidance now covers "traditional models and basic AI applications."
  3. Third-party risk proposal. The agencies propose replacing their third-party risk management guidance.
  4. Lead's framework. Lead proposes the Authorized Financial Agent regime and a model Know Your Agent program.
  5. Still pending. The request for information on AI and model risk, promised in April, has not been published.

Sources: OCC News Releases 2026-29 and 2026-77; Federal Reserve SR 26-2; Bowman, May 1, 2026; Lead (2026).

Lead names the consequence. Current guidance has left agentic AI "outside model-risk scope, leaving a gap around exactly the systems posing the greatest model risk" (Rec. 14, pp. 17–18).

Banks still answer to examiners in the meantime. Third-party risk guidance still covers AI vendors, and the agencies proposed a replacement on September 11. In Congress, Lead points to bills including Sen. Mark Warner's AI Agent Act, which would direct NIST to identify technical standards for secure agent access (p. 2). Until specific rules arrive, a bank or enterprise that lets an agent move money has to produce its own evidence that the agent stays within bounds.

A person approving every payment does not scale

Without that evidence, the safe default is to put a person in front of every payment the agent proposes. That keeps risk low, and it limits the agent to the reviewer's pace. By the fourteenth approval of the day, the reviewer is likely skimming. That can be riskier than a well-controlled agent.

A single good demo will not move that default. What moves it is documented risk and operating data that shows the agent staying inside its mandate over time. That is the standard Lead is now proposing for regulators.

The hardest part is oversight after approval

Most of the framework covers what happens before an agent goes live: identity records, credentials, certification and consent. Recommendation 14 sets a demanding standard for what comes after. It asks regulators to treat "mandate enforcement, drift monitoring, tested revocation, and adversarial testing" as "a supervisory focus equal to pre-deployment validation" (p. 18).

Recommendation 15 describes that oversight. Banks should "track each agent's normal pattern (transaction velocity, counterparty set, data-access and tool-call profile) and flag deviations" (p. 19). The KYA program asks enterprise risk teams to treat "mandate drift" as a named risk (p. 29).

A bounded mandate tells you whether a payment is allowed. Mandate drift is a separate problem, because an agent can stay inside its limits while it moves away from the goal it was given. Each step can be permitted on its own, so a check that asks only whether a payment is inside the mandate passes every one. Velocity and counterparty monitoring helps, but it reacts only once the payments begin. Catching drift earlier means watching how the agent reasons as well as what it pays.

AGIS turns an agent's mandate into an enforced, auditable policy

AGIS works as a loop with four stages. It covers both halves of Lead's framework: testing before an agent goes live, and oversight while it runs.

1. Distill the policy. AGIS is built as a set of agents. They read a company's infrastructure with minimal privileges and find each agent. They also find the tools, permissions and guardrails each agent is built from. AGIS distills that into a written policy for each workflow. Our CTO, Eric Bogard, describes the result as "a legal-esque document describing agents' mandates." It "serves as a visible contract that can be audited and can be tracked through time as it evolves."

2. Analyze and simulate. In a legal contract, one word can change what a party is allowed to do. The same is true of an agent's policy. AGIS runs simulations against the policy to find where an agent could stay within its permissions and still chain allowed actions into a harmful outcome. It then suggests fixes, which can be opened as proposed changes to the policy.

3. Watch and rule in production. AGIS maps the policy back onto the customer's infrastructure. It connects to guardrail services that cloud providers such as Cloudflare and AWS already offer. It reads each agent's vital signs against that agent's own normal and rules on every transaction, weighting each ruling by the risk those signs show. On the paths that lead to forbidden outcomes, it places containment and decoys. A decoy is a safe route that moves no real money. A block alone leaves the pressure in place, so a strained agent looks for another way round. A decoy keeps the agent working and shows what it was after. Some enforcement can be deterministic. Eric expects the industry will also need to accept non-deterministic enforcement "with evidence and quantifiable risk."

4. Record everything. Every reading, decision, decoy and alert is visible to the team live and kept as an audit trail. Every decoy is marked in that trail and never touches real balances. Over weeks, AGIS learns each agent's patterns and suggests fixes to the conditions behind them.

The analysis in AGIS is driven by Azimuth, the autonomous attacker we spent years building for smart contract security. Azimuth ranks #1 on EVMBench, OpenAI and Paradigm's benchmark for AI security tools, and outperforms frontier models on it by more than 20 points. It found a critical vulnerability in Ledger's Ethereum app. Its users have generated over 22,000 attack hypotheses across more than 7,000 scans. Azimuth searches for weaknesses, builds attack paths and tests them in an autonomous loop. In AGIS, we point that capability at agent policies and train it for behavioral analysis of agents.

AGIS maps onto Lead's requirements

Status

AGIS is built to fit the controls the framework describes. The framework is still a proposal, and none of it is binding yet.

What Lead's framework asks for What AGIS does
A recorded mandate with explicit limits (Rec. 3) Distills a written policy from the agent's tools, permissions and guardrails, readable as a contract
Adversarial testing with documented pass/fail criteria (Rec. 2) Azimuth simulates attacks against the policy to find paths to forbidden outcomes
Recertification after material model, framework or tool changes (Rec. 2) Tracks each version of the policy, so every change has a record
Flag deviations from each agent's normal pattern (Rec. 15) Measures each agent's vital signs against its own normal, and weights every ruling by that risk
Detect mandate escape and mandate drift (Rec. 2; KYA, p. 29) Flags an agent that is reinterpreting its rules or shifting responsibility
Kill switches and human escalation for high-risk actions (Rec. 14) Alerts operators as anomalies happen, and blocks any held action that nobody clears in time
Records that reconstruct any agent action (Rec. 18) Keeps every reading, decision, decoy and alert visible live and on the audit trail
Risk tiers based on agent capability (Rec. 2) A trust ladder that widens each agent's autonomy as evidence builds

AGIS does not replace the bank layer. Agent credentials, sanctions screening and payment-instrument limits belong to the bank, as Lead describes. Lead calls its controls "a second, independent layer at the bank." AGIS adds a layer on the agent itself, run by the firm that operates it. In the framework's three lines of defense, the teams that build and operate agents perform first-line controls (p. 35). AGIS sits with them.

Autonomy should be earned one step at a time

Lead's tiers describe what an agent is capable of doing. We think of the same idea as a ladder each agent climbs, much like promoting a new hire with a record at each step. Here is an example for a payments agent.

Few firms have climbed far. In a 2026 Cloud Security Alliance survey, only 5% of financial-services firms using agents let them act alone on critical actions.

Every step needs evidence: what was tested, what was allowed and what was fixed. That record is what auditors need, and it is what Lead proposes examiners should see. Every firm has access to the same models and tools. The advantage will go to firms that can safely hand more of their work to agents.

One addition we would propose

The adversarial tests in Recommendation 2 cover attacks: prompt injection, mandate escape, credential exfiltration, tool misuse, data-scope overreach and confused-deputy scenarios. Agents can also drift with no attacker involved. Ordinary operations create pressure too: a cash shortfall, a deadline, or a request nobody answers.

Our proposal

Add pressure scenarios to the adversarial-testing standard. A treasury agent should face cash shortfalls, deadlines and missing escalation paths before it is approved for high-risk work. Those tests should carry documented pass/fail criteria, as the framework already requires for prompt injection.

The same logic drives the fixes AGIS suggests over time. A lot of risky agent behavior starts in the environment. The usual causes are an unclear rule, a missing escalation path or a tool that doesn't do the job. Fixing those conditions means fewer risky actions reach the real-time check.

We are building AGIS with design partners

Lead describes agentic finance as "a controlled extension of existing bank services" (p. 2). Its roadmap starts with supervisory guidance that regulators can issue now, covering agent inventories, governance and audit logs (p. 24). It ends with new legislation. Until the rules arrive, banks, fintechs and enterprises are writing their own controls, and firms running treasury agents on bank rails can start building their oversight record today.

AGIS is early, and we are building it with a small number of design partners running real agent workflows that touch money. If your team is giving agents financial authority, or working through the questions in Lead's framework, we would like to talk.

Learn more about AGIS

Sources