

When I first started briefing portfolio managers on agentic systems in 2024, the typical reaction was polite scepticism. By the start of 2026, the same managers were asking me which agents had passed their internal model risk reviews and how quickly they could be wired into the order management system. The shift in tone has been faster than anything I have seen in two decades of working at the intersection of capital markets and software.
What changed was not the underlying language models. It was the recognition that finance, more than almost any other industry, runs on the patient assembly of evidence, the careful translation of policy into action, and the relentless documentation of every decision. Those are exactly the things that well-designed AI agents do at a speed and consistency no analyst team can match. The catch, of course, is that the regulators expect the same audit trails and the same accountability they have always demanded.
In this guide I want to walk through where AI agents in finance are already creating measurable value, where the failure modes lurk, and how I would approach a deployment if I were sitting in the chief operating officer’s chair at a global bank or asset manager today.
Finance has three properties that make it unusually well suited to agentic automation. The first is that almost every workflow already produces a structured paper trail, which means agents can be supervised, replayed, and audited without rebuilding the underlying systems. The second is that the cost of a single human analyst hour is high enough that even modest automation produces a defensible business case. The third is that the same workflow repeats thousands of times a day across hundreds of issuers, counterparties, or clients, which gives the agent enough surface area to learn and improve.
I often tell finance leaders that the question is no longer whether agents will appear in their stack but whether the agents will be built internally, embedded by their existing software vendors, or brought in by their fintech challengers. The answer in most large institutions is all three at once, which is why governance has to be the starting point rather than an afterthought.
“Treat every agent the same way you would treat a junior analyst with extraordinary speed and zero institutional memory. Supervise the work, retain the artefacts, and never let them sign anything that binds the firm.”
Equity research is the area where I have seen the cleanest early wins. A modern research agent can ingest an 8-K filing, reconcile the new numbers with the prior consensus model, draft a one-page update with supporting citations, and queue the revised model for an analyst to approve, all within minutes of the filing hitting EDGAR. The same agent can monitor a watchlist of two hundred names rather than the twenty an analyst can realistically follow.
The leading tools in this space are Hebbia, AlphaSense AI, and an increasing number of bespoke systems built on top of LangGraph or the OpenAI Agents SDK. The pattern is similar across vendors:
| Task | Human analyst time | Agent-assisted time |
| Earnings note for a covered name | 4 to 6 hours | 30 to 45 minutes |
| Initiation of coverage draft | 3 to 4 weeks | 4 to 6 days |
| Quarterly model refresh | 6 to 8 hours | 60 to 90 minutes |
| Thematic screen across 500 names | 2 to 3 days | 2 to 3 hours |
I would not yet trust an agent to publish unsupervised. I do trust it to do the spadework that lets a senior analyst spend the day on judgement rather than copy and paste.
On the trading desk, agents are showing up in two very different places. The first is signal generation, where systems like the much-discussed BridgewaterBot prototypes ingest news, alternative data, and macroeconomic releases to propose trade ideas with supporting evidence. The second is execution, where an agent sits next to the trader and helps choose venues, sizes orders, and watches for adverse selection.
I am cautious about over-claiming on the signal side. The history of systematic trading is littered with strategies that looked brilliant in backtests and humiliating in production. What I have seen work is a narrower pattern where the agent does not pick the trade but instead packages the evidence so that the human portfolio manager can decide in seconds rather than hours.
On execution, the wins are easier to measure. An execution copilot that reduces implementation shortfall by even a basis point on a multi-billion-dollar book pays for itself in weeks. The key is to keep the agent inside guardrails defined by the firm’s existing transaction cost analysis and to log every recommendation so the desk can review what would have happened if the human had not overridden it.
Know-your-customer and anti-money-laundering work is the operational backbone that every regulated firm has to maintain and the one almost nobody enjoys. It is also where I expect agents to deliver the largest cost savings over the next three years.
A modern KYC agent does the following:
The same agent, run continuously, becomes a perpetual KYC engine that re-screens clients whenever a watchlist updates or an ownership change is detected. I have seen large banks reduce the average time to onboard a corporate client from twenty-five days to seven, with quality metrics that satisfy the second line of defence.
The trap to avoid is treating the agent as the decision maker. Regulators expect a named human to own the risk rating, and the agent’s job is to make that human’s review as fast and as well-evidenced as possible.
Fraud detection has always been a machine learning problem. What agents add is the ability to investigate alerts rather than simply score them. The classic rule-based system fires hundreds of thousands of alerts a day, most of which a human team triages with a quick glance. An agent can do that glance at scale, pulling the customer history, the counterparty record, and any related cases into a single dossier within seconds.
I have seen card issuers cut their false-positive rate by forty per cent simply by routing every alert through an investigation agent that decides whether to auto-clear, escalate, or freeze. The cost saving is real, but the customer-experience gain is bigger. A legitimate customer whose card stops working at the supermarket is a customer who is one step closer to switching providers.
The agent does not replace the fraud analyst. It replaces the first thirty seconds of every investigation, which is the part that scales badly with volume.
In wealth management, the agentic copilot is starting to look like the genuinely useful junior associate that every adviser has wanted for years. It listens to the client meeting, drafts the follow-up, updates the financial plan, prepares the trade tickets, and queues the regulatory disclosures.
The Reg BI and fiduciary obligations in the United States, together with the suitability regimes in the United Kingdom and the European Union, mean that the adviser remains the decision maker. The copilot’s role is to remove the administrative burden that has historically forced advisers to spend more than half their week on paperwork rather than on clients.
I would watch the productivity numbers carefully here. Early deployments at large wirehouses are showing advisers handling between twenty and thirty per cent more households without a drop in client satisfaction scores. That number alone will rewrite the economics of the industry.
Every agent that touches a regulated workflow has to leave behind the same kind of evidence trail that a human would. In practice this means logging:
The European Union’s MiFID II regime, the SEC’s Reg BI, FINRA’s Rule 3110 supervision requirements, and the UK Senior Managers and Certification Regime all assume that a named human is responsible for the activity. An agent does not change that. It changes the volume of evidence the human has to absorb to make a decision.
I have started recommending that firms treat their agent logs as a regulated record from day one. That means retention policies that match the longest applicable regulatory window, which is often seven years, and access controls that match the firm’s existing books-and-records policy.
Most large financial institutions run their model risk management under SR 11-7 in the United States or its equivalents elsewhere. Agentic systems do not fit comfortably into that framework because the model is only one component of a much larger system that also includes retrieval, tools, prompts, and orchestration logic.
I have been pushing firms to expand their inventory to cover the agent as a system, not the model as a component. The questions the second line should be asking include:
“The single biggest mistake I see is firms deploying agents under the existing model risk policy without recognising that the system has new failure modes the policy was never written to address.”
A short, opinionated list of the vendors I think are doing the most interesting work in early 2026:
I am deliberately not naming the consumer-facing robo-advisers, because the agentic shift there is happening behind the scenes rather than in the customer-facing app.
The failure modes I worry about most are not the dramatic ones. They are the quiet ones that erode trust over time.
Each of these has a mitigation, but the mitigations only work if the firm has decided in advance who owns the risk and what the response looks like.
The firms that are getting the most value are the ones that have invested in a shared internal platform rather than letting every business unit roll its own. The platform typically provides:
This is essentially the same platform discipline that the firm already applies to its trading systems. The difference is that the platform serves dozens of agents rather than a handful of strategies, and it has to make it easy for business users to launch new agents without compromising the controls.
The talent market for AI agent engineers in finance is brutal. I have seen offers north of four hundred thousand US dollars in total compensation for senior engineers with both LLM and capital markets experience, and the supply is not catching up quickly.
My advice to firms is to buy the platform and the obvious horizontal tools, build the agents that touch the firm’s proprietary data and workflows, and partner with specialist vendors for the regulated edge cases like sanctions screening and surveillance.
| Layer | Recommended approach |
| Foundation models | Buy from frontier labs |
| Agent framework | Buy or open source |
| Retrieval and data | Build with the firm’s data team |
| Workflow agents | Build for proprietary cases |
| Compliance agents | Buy from specialist vendors |
| Monitoring and evaluation | Build internally |
If I were given a clean sheet and ninety days, here is roughly how I would spend them.
The point is not to ship something perfect. The point is to learn what the firm’s own governance, technology, and culture will tolerate before the stakes get higher.
By the end of 2026 I expect every tier-one bank and asset manager to have at least one production agent in a revenue-generating workflow. By 2028 I expect agents to be the default interface to the firm’s data for most analysts and bankers. By 2030 I expect the org chart of a typical finance function to look meaningfully different, with fewer associates and more agent supervisors.
The firms that win will be the ones that treat agents as colleagues with extraordinary speed and extraordinary blind spots, and that build the governance to amplify the strengths while containing the blind spots.
Brian Jagger is an AI Architect and Software Engineer with over 15+ years of experience in generative AI, AI-first software development, and digital accessibility. As the Co-founder & CTO of TechA11y and Founder of GuardRailz, he has built innovative AI solutions for businesses, education, and enterprise clients. Brian combines deep technical expertise with a creative background in film and media, helping professionals leverage AI to build impactful, scalable solutions.
QUICK FACTS
The agents themselves are not directly regulated in most jurisdictions, but every regulated activity they touch is. That means MiFID II, Reg BI, SR 11-7, the Senior Managers Regime, and the equivalent rules elsewhere all apply to the human and the firm using the agent.