← AiTimeline Home

Technology · AI in Finance

Agentic AI in Finance Timeline 2023–2026: How Wall Street Is Moving From Chatbots to AI Agents

📅 Updated 2 September 2026Goldman Sachs, Anthropic, Reuters, Bank of England, FSB, FINRANot investment advice
Advertisement

View as Web Story

In short

How Wall Street moved from chatbots to AI agents (2023-2026): Goldman-Anthropic Claude agents, the 51% bank pilot survey, and what regulators allow.

For the first phase of Wall Street’s generative-AI boom, the human still did almost everything important. The AI summarised a filing; a person read the summary. The AI drafted a memo; a person sent it. The AI suggested code; an engineer ran it. Agentic AI changes that sequence: instead of returning one answer, an agent can take a goal, find information, call software tools, move between applications, check its own intermediate results and complete several steps — increasingly deciding which step comes next. This agentic AI in finance timeline tracks how Wall Street and its regulators actually moved from chatbots to agents between 2023 and 2026. The headline is not that banks handed trading floors to robots. It is that they started giving AI real authority over workflows — and now have to decide where that authority stops.

Data last verified: 2 September 2026. This article separates what banks and AI vendors have officially announced, what regulators have published, what news organisations have reported, and what remains simulation or future risk. It contains no investment advice, no trading signals, no buy/sell or portfolio recommendations, and no instructions for autonomous live trading.

⚠️ What this article is — and is not. This is not a claim that autonomous AI has taken over trading, that Goldman Sachs lets AI agents run its trading floor, or that any major bank routinely lets an unsupervised language model select, size and execute live trades. Banks are deploying and piloting agentic systems across research, engineering, client vetting, operations, treasury and trading-related workflows while keeping tighter human and system controls around high-stakes financial decisions. Every case study below is labelled with its verification status.

🧠 AI Overview Summary

Between 2023 and 2026, Wall Street’s use of AI evolved from chatbots that answer questions to agents that complete multi-step workflows using approved data and tools. Enabling technology such as OpenAI’s June 2023 function calling and Microsoft’s AutoGen made structured tool use and multi-agent designs practical. Goldman Sachs rolled out a firmwide AI assistant in June 2025, detailed an AI-driven operating model in early 2026, and in February 2026 confirmed it had spent roughly six months with embedded Anthropic engineers building Claude-based agents for trade and transaction accounting, client due diligence and onboarding. In May 2026 Anthropic released ten finance-specific agent templates. By mid-2026, a KPMG survey cited by Reuters put 51% of surveyed banks as piloting AI agents, mostly as “digital coworkers” in research, wealth, treasury and operations. Regulators — FINRA, ESMA, the Federal Reserve, the Bank of England and the Financial Stability Board — treat fully autonomous AI trading as a future risk to study, not a documented mainstream practice; the FSB’s most immediate concern as of August 2026 is AI-amplified cyber risk.

📊 Agentic AI on Wall Street — status, 2 September 2026
🟢 Widespread
AI assistants across large banks
🟢 Deploying / piloting
Agentic workflows in operations & research
🟢 In use
Software-engineering agents in bank workflows
🔵 Development / rollout
Client onboarding & KYC agents
🟢 Active area
Research, treasury, wealth, trading support
🔴 Not established
Unsupervised LLM live trading as normal practice
🔴 Not established
Fully autonomous major-bank portfolio management
⚪ Research / simulation
Agent-to-agent autonomous markets (Project Logos)
51%
Banks in June 2026 KPMG survey piloting AI agents
Status labels defined below; every figure is sourced inline in the timeline.
⚡ Agentic AI in Finance — Quick Facts
Enabling milestoneOpenAI function calling, 13 Jun 2023
Multi-agent frameworkMicrosoft AutoGen, 2023
Goldman AI assistantFirmwide, Jun 2025 (~10,000 users)
Goldman + AnthropicConfirmed 6 Feb 2026
Anthropic finance agents10 templates, 5 May 2026
Banks piloting agents51% (KPMG survey, Jun 2026)
⚡ Quick Answers — AI Overview Ready

Agentic AI in Finance: Key Questions

What is agentic AI in finance?
AI systems that pursue a goal across multiple steps by planning, using software tools, accessing approved data and taking actions — rather than only generating an answer. In banking they are used for research, engineering, KYC preparation, reconciliation and client-service support.
Is Wall Street using AI agents?
Yes. Major institutions including Goldman Sachs, JPMorgan, Morgan Stanley, Citi, BNY and UBS are piloting or deploying agents across operations, wealth, client vetting, engineering, research and trading-related workflows, per Reuters reporting in July 2026.
Are AI agents executing trades autonomously?
Not as a documented mainstream practice. AI is used across trading-related research, analytics, surveillance and workflow automation. Bank of England analysis in 2026 indicates more autonomous AI in markets is mostly used for research, coding and surveillance, not fully autonomous trading.
What is bounded autonomy?
The model most banks actually use: an agent is connected to specific tools, specific data, specific permissions and specific workflows, with human approval gates around high-stakes actions such as moving funds, approving clients or executing trades.
📚 Key Takeaways

What actually changed between 2023 and 2026

  • The shift is from AI that answers to AI that completes workflows. Chatbot, copilot, tool-using assistant, agent, agentic workflow, multi-agent workflow, bounded autonomous system — Wall Street moved several rungs up that ladder, but not to the top.
  • Function calling (June 2023) made tool use structured and reliable — it did not invent AI agents. ReAct-style research and LangChain agent concepts predate it.
  • Agentic AI is not the same as multi-agent AI. One agent can use many tools; a multi-agent system splits work across specialised agents. Not every bank deployment uses multi-agent architecture.
  • Goldman Sachs is the clearest public case. Firmwide AI assistant (June 2025), then software-engineering agents, then an AI-driven operating model (early 2026), then Claude-based agents with Anthropic for accounting, due diligence and onboarding (February 2026).
  • Banks let AI write code before they let it control capital. Code can be tested, reviewed, sandboxed and rolled back; a live trade changes counterparty exposure and capital in ways that cannot simply be undone.
  • “AI in trading” usually means research, analytics, surveillance, pricing or execution support — not an LLM independently selecting a security, sizing a position, executing and managing risk.
  • Regulators separate three horizons: today’s concern is cyber and operational resilience; the emerging concern is client-facing and financial decision-making autonomy; the future systemic concern is correlated behaviour across many similar agents.
  • The FSB’s most immediate AI concern (August 2026) is cyber risk — frontier AI changing the speed, scale and economics of finding and exploiting vulnerabilities — not an imminent AI-driven flash crash.
  • Bounded autonomy is the operating model. The Wall Street question is no longer “can AI write the memo?” but “what are we willing to let it do after the memo?”

🔴 Reading the status labels used throughout this article

🟢 VERIFIED DEPLOYMENT — announced by the institution or vendor as in production use. 🔵 PILOT / DEVELOPMENT — confirmed as being built or tested, not yet general production. 🟡 REPORTED / PARTIALLY VERIFIED — from credible reporting or a survey, with details (sample, scope) not fully public. ⚪ CONCEPT / SIMULATION — a research project or simulated environment, not a live market system. 🔴 NOT YET ESTABLISHED — not a documented normal practice at major institutions; do not assume it.

Generative AI answers. Agentic AI acts. Wall Street’s question is no longer “can AI write the memo?” It is “what are we willing to let it do after the memo?”

Chatbot vs Agent vs Multi-Agent System

Three different things that get blurred together in coverage

A chatbot takes a question and returns an answer. An agent takes a goal, makes a plan, calls tools, takes actions and returns a result — and can maintain state across the steps. A multi-agent system divides that work across specialised agents, for example a research agent, a risk agent and a compliance agent coordinated by an orchestrator, usually with a human approval step before anything is actioned.

The important correction: agentic does not automatically mean multi-agent. A single agent using several tools is still an agentic system, and many bank deployments are exactly that. Multi-agent architecture is an option, not a requirement — and it introduces its own failure modes (agent disagreement, error propagation, authority confusion, feedback loops) rather than being automatically safer.

Chatbot — question → answer
Agent — goal → plan → tools → actions → result
Multi-agent — goal → research agent → risk agent → compliance agent → human approval → authorised system → audit log

Generative AI Assistant vs Agentic AI

CapabilityGenerative AI assistantAgentic AI
Main purposeGenerate or analyse informationComplete goals across steps
Typical interactionPrompt → responseGoal → plan → tools → actions
Tool useOptionalOften central
Memory / stateUsually limited task contextCan maintain workflow state
AutonomyLow to moderateBounded / moderate, depending on permissions
Human roleReviews outputSets boundaries, approvals and exceptions
Finance exampleSummarise a 10-KCollect documents, run checks, prepare workflow output
High-risk actionHuman executesUsually permission-gated
Audit requirementPrompt / output historyFull action, tool and approval trail

Autonomy is a spectrum, not a switch. The realistic picture is “bounded and permission-gated,” not “constant prompting” versus “humans only supervise.”

The Real Wall Street Model: Bounded Autonomy

Banks generally do not want an unconstrained AI model with unlimited access to money, markets, client accounts or regulated systems. Instead, agentic AI is connected to specific tools, specific data, specific permissions, specific workflows and human approval gates. The critical word is bounded.

An agent may be allowed to read documents, query databases, calculate values, draft forms and reconcile records. It is not necessarily allowed to move funds, send regulated communications, approve client onboarding, change risk limits or execute trades. The more authority an agent gets, the more important its permission boundary becomes.

Human-in-the-loop vs human-on-the-loop

Human-in-the-loop: the agent must pause and get approval before a critical action. Human-on-the-loop: the agent can operate inside pre-approved boundaries while a human monitors and can intervene. Full autonomy: minimal or no human approval. For regulated finance, many high-risk applications should stay bounded and approval-gated — and neither in-the-loop nor on-the-loop is mandatory for every agent; the right setting depends on the action’s blast radius.

Wall Street’s AI Autonomy Ladder

2026 adoption is moving up the ladder — but not all the way

Answer. Respond to a question. Read-only.
Draft. Produce a memo, note or first-pass model. Human sends.
Research. Gather documents and data across sources; synthesise.
Use tools. Call internal APIs, pricing services, databases.
Update workflow. Write to internal records or case files, within scope.
Trigger approved action. Start a workflow that a human or authorised system signs off.
Make a financial decision. Choose an allocation or client outcome inside limits. Highly restricted.
Control capital autonomously. Create and execute strategies with minimal oversight. Not a documented normal major-bank practice for LLM trading.

The question is not “can an AI agent do the job?” It is “how much authority should it be given?”

Where Is the Agent Allowed to Act?

Research — high access
Documentation — high access
Reconciliation / accounting — high access
Compliance preparation — controlled
Client decision — restricted
Portfolio decision — highly restricted
Live market execution — highest control

An AI does not need direct exchange access to create large productivity gains — it can automate everything around the final controlled action.

Diagram of Wall Street's AI autonomy ladder from answering questions to controlling capital, and where 2026 bank adoption sits on it

Goldman Sachs’ Agentic-AI Stack

The clearest public sequence from assistant to digital coworker

Jun 2025 — GS AI Assistant firmwide (~10,000 already using it)
2025 — agentic software-engineering tools enter developer workflows
Jan 2026 — One Goldman Sachs 3.0: AI-driven operating-model redesign
Feb 2026 — Claude-based agents with Anthropic: trade/transaction accounting, client due diligence, onboarding
May 2026 — broader finance-agent ecosystem (Anthropic templates)
Jul–Sep 2026 — agents used increasingly as “digital coworkers”

At the February 2026 report, Goldman was described as “in the early stages” of developing these agents and expecting to launch them “soon” — not as having deployed autonomous Claude agents across the bank. Goldman executives said they were surprised at how capable Claude was beyond coding, in accounting and compliance work that combines parsing large volumes of documents with applying rules and judgement. The bank also reported a roughly 30% reduction in the time to onboard new institutional clients using Claude’s reasoning. 🔵

Why code is easier than money

✅ Code

  • Can be tested before it runs
  • Can be reviewed by a human or another system
  • Runs in a sandbox first
  • Version controlled — can be rolled back

❌ A live trade

  • The market moves while you act
  • Execution changes counterparty exposure
  • Capital changes hands
  • Some effects cannot simply be undone

This is why banks let AI write code before they let it control capital — software engineering is a relatively controlled environment for agent autonomy.

A Note on the Model Version

Anthropic released Claude Opus 4.6 on 5 February 2026, improving agentic task performance, coding, long-horizon work and document and spreadsheet workflows, and some reporting connects that model to Goldman’s agent development. But Goldman’s confirmation refers broadly to Claude-based agents. Unless a first-party Goldman source names the exact production model, the bank’s agent strategy is better described model-agnostically. Model versions change quickly; the architecture is what lasts: approved data → model → tools → permissions → checks → human gate → audit trail.

Who Is Using Agents?

Based on July 2026 Reuters reporting and company disclosures; blanks mean no public evidence, not “no”

InstitutionReported agentic / AI focus areasAutonomous live trading
Goldman SachsAI assistant, software engineering, accounting, due diligence, onboarding, operating-model redesignNo public evidence
JPMorganInternal productivity, research, treasury and financial-services workflowsNo public evidence
Morgan StanleyAdviser-support agents, client interaction; humans kept in the loopNo public evidence
CitiEnterprise AI, agentic architecture, digital-assistant rolloutNo public evidence
BNYAI systems managed operationally like “digital employees” with defined identitiesNo public evidence
UBSDigital assistants with human oversightNo public evidence
Visa, AIGNamed by Anthropic among financial-services adopters of its agentsNot applicable / no public evidence

“Digital employees” is an operational management metaphor for identity, permissions and ownership — not a claim that an AI legally becomes an employee. No entry here should be read as evidence of autonomous trading.

“Trading” Is Not One Thing

When a source says AI is used in trading, it can mean any of: research, analytics, data processing, surveillance, execution support, pricing, workflow automation or algorithm development. It does not automatically mean an LLM independently selects a security, chooses a position, sets size, executes and manages risk. There are at least six distinct meanings, and collapsing them into “AI trading” is the single most common error in coverage:

  1. AI helps trading research — summarising filings, news, positioning.
  2. AI helps algorithm development — writing and testing execution code.
  3. AI adjusts existing models — tuning parameters a human owns.
  4. AI suggests orders — a human decides.
  5. AI manages a portfolio under constraints — inside hard limits and suitability rules.
  6. AI autonomously creates and executes strategies — the frontier, not documented as mainstream.

Algorithmic trading is not agentic AI

Algorithmic trading has existed for decades: a predefined rule or model receives market data and executes. An agentic system can potentially plan, reason, select tools, adapt its workflow and make sequences of decisions. Not every algorithm is an AI agent.

High-frequency trading is not an LLM agent

High-frequency trading prioritises latency, deterministic execution, market microstructure and specialised infrastructure. LLM agents add reasoning and tool orchestration — but also latency, non-determinism and model risk. It is technically misleading to say “LLM agents execute trades at zero latency.” The more accurate concern: machine-speed decisions can amplify risk if more autonomous AI is eventually connected directly to financial decision-making.

What Regulators Actually Worry About

Three horizons that should be kept separate

Today — cyber risk, operational resilience, governance, vendor and cloud concentration
Emerging — client-facing autonomy and AI in financial decision-making
Future systemic — correlated behaviour across many similar agents, direct portfolio autonomy, agent-to-agent markets
Body2026 action / positionStatus
FINRA (US)2026 Annual Regulatory Oversight Report treats AI agents as an emerging trend; flags autonomy and scope creep, authority beyond intended scope, auditability and transparency, data sensitivity, domain knowledge and incentive design — on top of hallucination, bias and privacy risks🟡 Guidance
ESMA (EU)26 February 2026 supervisory briefing on algorithmic trading under MiFID II: governance, testing, outsourcing and the interpretation of key concepts; existing trading rules still apply when AI is used🟢 Published
Federal Reserve (US)Examining agentic AI both as a user (applying AI to financial-stability analysis) and supervisor; AI named among top potential market shocks in the 2026 Financial Stability Report; US interagency work in April 2026 noted traditional model-risk guidance does not specifically encompass generative/agentic AI🟡 Reported
Bank of EnglandJuly 2026 Financial Stability Report: frontier AI raises financial-stability risk mainly via cyber and operational vulnerabilities and third-party concentration; analysis indicates more autonomous AI in markets is currently weighted toward research, coding and surveillance, not fully autonomous trading🟢 Published
Financial Stability Board31 August 2026 letter to G20: frontier AI’s effect on cyber risk is the most immediate concern for the financial system — it could alter the speed, scale and economics of cyber attacks, amplified by concentrated third-party providers🟢 Published
BIS Innovation Hub + BoE + BundesbankProject Logos: an agent-based simulation comparing rules-based and LLM-based asset managers in a simulated market, to study how agents allocate capital and whether they amplify correlated decision-making⚪ Simulation

🔮 Project Logos — regulators testing the future

Actual market today → simulated agent market → study collective behaviour → design future oversight. Project Logos is a simulation, not proof that central banks allow AI agents to manage real portfolios. Its research question is the systemic one: if many systems share common models, data, signals, cloud providers and optimisation objectives, could they react similarly to the same signal and amplify a market move? One bad AI agent is an operational risk; thousands of similar AI agents can become a systemic question.

On flash crashes: regulators are not claiming an AI-driven “Flash Crash 2.0” is expected tomorrow. The historical analogy is used to frame a research question about correlated automated behaviour — not to describe a documented agent-driven crash. And using AI does not erase existing regulation: rules on supervision, records, market conduct, best execution, risk controls and communications can still apply. Regulation here is technology-neutral.

The Agentic-AI Governance Problem: Who Authorised the Action?

Traditional software usually follows defined logic. Generative AI can produce unexpected output. Agentic AI adds another dimension: unexpected action. So the key governance question is not only “was the answer correct?” but “was the agent allowed to do that?

For every action, a bank may need to reconstruct which model acted, what data it saw, which tools it called, what permissions it had, what rule it applied, what intermediate output it generated, whether another agent reviewed it, whether a human approved it, and what happened afterward. That makes auditability, permissioning, identity, monitoring and action logs into core financial infrastructure.

Liability is not a simple choice

Responsibility for an agent’s mistake is not a clean pick between the bank, the model provider and a human supervisor. It can depend on the activity, jurisdiction, contract, regulatory obligation, delegation, supervision and product design. In many regulated contexts, an institution cannot simply outsource its responsibility to an AI vendor.

Traceability, not full explainability

Banks realistically need traceability — observable inputs, outputs, actions, tool calls, approvals and policy decisions — rather than a promise of full explanation of a model’s internal reasoning. Every action needs a receipt: input → plan → tool call → data accessed → policy decision → approval → action → result.

Illustrative control — not a real agent

Agent passport

Name: KYC-Agent-23. Role: document preparation. Can read: approved KYC systems. Can write: draft case file. Can approve client: no. Can move money: no. Can execute trade: no. Human owner: compliance team. Audit log: on. Expiry: set.

Agent-specific risks

Beyond generic GenAI risk

Prompt injection (an external document carrying a malicious instruction the agent then acts on), data leakage across client data, market data and material non-public information, and long multi-step workflows drifting out of scope. An AI agent acting for a professional does not get a free pass on insider-information rules or information barriers.

Multi-agent systems do not automatically add safety. A “risk agent” checking a “research agent” can add defence — but if both share the same model, data and prompt assumptions, they can fail together. Real safety diversity includes deterministic rules, traditional risk engines, independent data, hard position limits and human review, not just asking another language model.

Finance Agent Use-Case Maturity, 2026

Use case2026 maturityAutonomy risk
Document summarisationHighLow
Research preparationHighLow–Medium
Software engineeringHighMedium
Pitchbook draftingHighMedium
KYC preparationGrowingMedium
Reconciliation / accountingGrowingMedium
Compliance monitoringGrowingMedium–High
Client interactionPilot / growingHigh
Portfolio recommendationsControlled / pilotHigh
Autonomous portfolio decisionsLimited / researchVery High
Unsupervised live tradingNot established as mainstreamExtreme

Cost, Tokens and ROI

Agents do not automatically reduce operating cost. An agentic workflow can require many model calls, long contexts, multiple agents, tool execution, monitoring and human review. Goldman Sachs Research has said it expects agentic AI to materially increase token consumption, because agents think, check, retry, call tools, delegate and process more context. One chatbot handles one task; a multi-agent setup adds an orchestrator, a researcher, a reviewer, a risk check and retries. The right scorecard measures cycle-time reduction, manual touches, error rate, exception rate, human review time, compute cost, regulatory incidents and customer outcome — not “jobs eliminated” as the only metric.

Will AI Agents Replace Bankers?

The honest answer is not yes or no: tasks are likely to change before whole occupations disappear. Work most exposed includes document review, data gathering, routine coding, reconciliation, pitch materials and basic analysis. Human-heavy work — relationships, judgement, negotiation, risk ownership and regulatory accountability — remains central.

For decades, a junior banker learned by doing the slow work: reading the filings, checking the spreadsheet, updating the presentation, finding the number that did not reconcile. AI agents are very good candidates for exactly that kind of work. The productivity case is obvious. The training problem is not. If software completes the first five steps, the next generation of bankers may arrive at step six without having learned why the first five mattered.

Consider an analyst who asks an AI: “Prepare tomorrow’s client meeting.” A chatbot writes a briefing note. An agent could open the CRM, retrieve the client’s holdings, check recent earnings, review previous meeting notes, update approved analytics, find upcoming maturities, prepare the presentation and schedule a reminder. The human asked one question; the machine performed many tasks. That is the agentic shift.

Now change the instruction to “Manage the client’s portfolio.” The same convenience becomes a governance problem. What can the agent buy? How much? Using which data? Under which suitability rules? Who approved the strategy? Who stops it? That is why agentic finance will likely advance one permission boundary at a time.

Agentic AI in Finance: The Full Timeline (2023–2026)

Newest first. Each entry carries a verification status.

The Digital-Coworker Phase

September 2026Current status

What happened: Across large banks, AI assistants are widespread, agentic workflows are deploying or piloting in operations and research, software-engineering agents are in real use, and client onboarding and KYC agents are in development or rollout. Autonomous portfolio management and unsupervised LLM live trading are not established as normal major-bank practice.

Why it matters: The unresolved question is no longer whether an agent can do the task — it is how much authority comes next, and where it stops.

The most important Wall Street AI story of 2026 is not “the robots took over trading.” It is “banks started giving AI agents real authority — and now have to decide where that authority stops.”
🟢 Verified deploymentBounded autonomy

FSB: Frontier-AI Cyber Risk Is the Most Immediate Concern

31 August 2026Financial Stability Board

What happened: In a letter to G20 finance ministers and central-bank governors, FSB Chair Andrew Bailey said frontier AI’s potential impact on cyber risk is the most immediate concern for the financial system, because it can materially alter the speed, scale and economics of finding and exploiting vulnerabilities, and because critical third-party providers are highly concentrated.

Why it matters: The biggest current regulatory fear is not an AI trader causing a flash crash tomorrow. It is cyber and operational resilience — a different risk from future market autonomy, and one that should be kept separate.

🟢 PublishedCyber + operational resilience

Bank of England: Autonomous AI in Markets Is Still Mostly Lower-Risk Work

July 2026Financial Stability Report; Project Logos

What happened: The Bank of England’s July 2026 Financial Stability Report warned that frontier AI increases financial-stability risk mainly through cyber and operational vulnerabilities and dependence on a small number of cloud and AI providers. Its analysis indicated that more autonomous AI systems in markets are used primarily for research, coding, surveillance and other lower-risk operational functions, rather than fully autonomous trading. Separately, the BIS Innovation Hub London Centre, the Bank of England and the Deutsche Bundesbank set up Project Logos to observe LLM-based agents acting as portfolio managers in a simulated market.

Why it matters: This is the central distinction of the whole story — “AI in trading” today is mostly research and surveillance, and the portfolio-manager scenario is being studied in simulation, not run live.

🟢 Published⚪ Simulation (Project Logos)

Reuters: Banks Promote Agents From Research Aids to “Digital Coworkers”

13 July 2026Reuters; KPMG survey

What happened: Reuters reported major banks — Goldman Sachs, JPMorgan, Morgan Stanley, Citi, BNY and UBS — accelerating agentic-AI adoption across wealth management, client vetting, trading-related work, treasury and operations. A June KPMG survey cited by Reuters found 51% of surveyed banks were piloting AI agents. BNY was reported to be experimenting with AI systems managed operationally like “digital employees,” with defined identities and management structures.

Why it matters: The “51%” figure is a survey of surveyed banks piloting agents — not 51% of every bank worldwide, and not production deployment. It marks a shift in framing from research aid to coworker, with humans still in the loop for high-stakes decisions.

“Digital employees” is an organisational metaphor for identity and permissions — it does not mean an AI legally becomes an employee.
🟡 Reported / survey51% piloting

Anthropic Launches Ten Finance-Specific Agent Templates

5 May 2026Anthropic

What happened: Anthropic released ten agent templates for financial services: five research and client-coverage agents (pitch builder, meeting preparer, earnings reviewer, model builder, market researcher) and five finance and operations agents (valuation reviewer, general-ledger reconciler, month-end closer, statement auditor, KYC screener), shipped as plugins for Claude Cowork and Claude Code and as cookbooks for Claude Managed Agents. Anthropic identified financial-services customers including Goldman Sachs, Citi, Visa and AIG, and formed a services-led joint venture with Goldman Sachs, Blackstone and Hellman & Friedman.

Why it matters: Finance-specific agent workflows moved from custom experimentation toward repeatable products. Naming an institution as a customer does not mean it uses all ten templates.

🟢 Product launch10 templates

ESMA Supervisory Briefing on Algorithmic Trading

26 February 2026ESMA (EU)

What happened: ESMA issued a supervisory briefing to help national regulators supervise algorithmic trading under MiFID II, covering governance, testing, outsourcing and the interpretation of key concepts, including emerging AI use.

Why it matters: Using AI does not remove existing trading rules. Pre-trade controls, governance, testing and outsourcing obligations continue to apply.

🟢 PublishedTechnology-neutral

Goldman Sachs Confirms Claude-Based Agents With Anthropic

6 February 2026Goldman Sachs; Anthropic; CNBC / Reuters

What happened: Goldman confirmed it had spent roughly six months working with embedded Anthropic engineers to build Claude-based autonomous agents for internal workflows. Reuters-confirmed development areas: trade and transaction accounting, client due diligence, and client onboarding. Goldman said it was still in the early stages of developing these agents and planned to launch them soon; it reported a roughly 30% reduction in institutional-client onboarding time.

Why it matters: This is the central milestone of the story — and it is post-trade and operational work, not autonomous trading. February 2026 should not be rewritten as “Goldman deployed autonomous Claude agents across the bank.”

Goldman executives were reportedly surprised at how capable Claude was beyond coding, in accounting and compliance tasks that mix document parsing with rules and judgement.
🔵 Pilot / developmentAccounting + onboarding + due diligence

Anthropic Releases Claude Opus 4.6

5 February 2026Anthropic

What happened: Anthropic released Claude Opus 4.6, with improvements in agentic task performance, coding, long-horizon work, financial analysis, research and document and spreadsheet workflows.

Why it matters: Some reporting links the model to Goldman’s agent development, but Goldman’s public confirmation refers broadly to Claude-based agents. The bank’s strategy is best described model-agnostically — model versions change faster than architecture.

🟢 Model releaseDon’t hard-wire the model

Goldman’s One GS 3.0 Operating Model

January 2026Goldman Sachs

What happened: Around its Q4 2025 earnings call and 2026 annual meeting, Goldman detailed “One Goldman Sachs 3.0,” a multi-year, AI-driven operating-model transformation. Initial workstreams included client onboarding and KYC, enterprise risk management, vendor management, lending, regulatory reporting and sales enablement, with emphasis on auditability, data lineage, risk insights and workflow automation.

Why it matters: The redesign frames AI as an operating-model change, not a set of point tools — and starts in operations and control functions, not trading.

🔵 Operating-model redesignKYC, risk, lending, reporting

FINRA Flags AI Agents as an Emerging Risk

December 2025FINRA 2026 Annual Regulatory Oversight Report

What happened: FINRA’s 2026 report explicitly treated AI agents — systems that plan, decide and act without predefined rules — as an emerging trend, flagging autonomy and scope creep, authority exceeding intended scope, and auditability and transparency, alongside data sensitivity, domain knowledge and incentive design, plus existing GenAI risks of hallucination, bias and privacy.

Why it matters: A US self-regulatory body put agent-specific risks on member firms’ agenda before most retail-facing deployments existed. Traditional GenAI risks still apply on top.

🟡 GuidanceAutonomy, authority, auditability

Software-Engineering Agents Enter Bank Workflows

2025Goldman Sachs and peers

What happened: Banks explored and adopted agentic software-engineering tools, with Goldman among those testing autonomous coding assistants. Code is a relatively controlled setting for agent autonomy because it can be tested, reviewed, version-controlled, sandboxed and rolled back.

Why it matters: Engineering became the intermediary step between the copilot phase and higher-stakes finance workflows — banks let AI write code before letting it touch capital.

🟢 In useControlled environment

Goldman Sachs Rolls Out the GS AI Assistant Firmwide

23 June 2025Goldman Sachs internal memo

What happened: Goldman made its GS AI Assistant available across the firm, with about 10,000 employees already using it at rollout. Typical tasks: summarising documents, drafting content and performing data analysis.

Why it matters: This is the copilot phase — “help me understand this,” not “complete this approved workflow.” Large institutions, not just retail investors and junior analysts, were the adopters.

🟢 Verified deployment~10,000 users

The Copilot Era: Wall Street Learns to Talk to AI

2024Multiple institutions

What happened: Large banks scaled generative-AI assistants for document summarisation, research assistance, code generation, internal search and meeting preparation.

Why it matters: Adoption at scale by major institutions set the base for the agentic step — but the AI still mostly returned information for a human to act on.

🟢 WidespreadCopilot phase

Microsoft Publishes AutoGen for Multi-Agent Workflows

2023Microsoft Research

What happened: Microsoft released AutoGen, a framework for building applications with multiple conversing agents, popularising multi-agent orchestration among developers. LangGraph and other production-orchestration tools followed in 2024.

Why it matters: Multi-agent design became a practical pattern — but it is an option, not a requirement, and it adds coordination-failure modes of its own.

🟢 FrameworkMulti-agent != required

OpenAI Introduces Function Calling

13 June 2023OpenAI

What happened: OpenAI added function calling, letting models generate structured arguments for external functions and APIs: model → tool request → external system → result → model.

Why it matters: Function calling made tool integration more structured and reliable — an important enabling technology for modern agents. It did not invent AI agents; agent concepts and ReAct-style research predate it.

Do not say function calling launched in late 2023, or that OpenAI invented AI agents.
🟢 Enabling milestoneStructured tool use
2022
–23

ReAct and Early Agent Frameworks Gain Developer Attention

2022 – early 2023Research and open-source

What happened: ReAct-style “reason + act” research and early LangChain chains and tool-using agent concepts drew developer interest, establishing the pattern of an LLM planning steps and calling tools.

Why it matters: The conceptual groundwork for agents existed before function calling made it robust — the enabling technologies stacked up rather than arriving in one moment.

⚪ Research / open-source

📝 Update History

  • 2 September 2026 — Full rewrite: corrected “Wall Street gave AI the power to trade” framing to the verified “chatbots to agents, with bounded autonomy” picture; added FSB August 2026 cyber warning, BoE July 2026 FSR, Project Logos, Anthropic May 2026 templates, Goldman-Anthropic February 2026 detail, FINRA and ESMA 2026 actions.
  • 31 August 2026 — FSB frontier-AI cyber-risk letter to G20.
  • July 2026 — Reuters “digital coworker” reporting; KPMG 51% survey; Bank of England FSR and Project Logos.
  • 5 May 2026 — Anthropic ten finance-agent templates.
  • 6 February 2026 — Goldman Sachs confirms Claude-based agents with Anthropic.
  • January 2026 — One Goldman Sachs 3.0 operating model.
  • June 2025 — GS AI Assistant firmwide.
  • 13 June 2023 — OpenAI function calling.

People Also Ask

Did Wall Street give AI the power to trade?
Not in the sense of handing trading floors to autonomous machines. Banks are deploying and piloting agents across research, engineering, client vetting, operations, treasury and trading-related workflows, while keeping tighter human and system controls around high-stakes financial decisions such as executing trades or moving funds.
Is Goldman Sachs using Claude?
Yes. Goldman confirmed in February 2026 that its Anthropic collaboration is based on Claude, with agents in development for trade and transaction accounting, client due diligence and onboarding. That does not mean every Goldman AI workload runs on Claude.
What is a multi-agent trading system?
An architecture that splits work across specialised AI components — for example research, risk analysis, compliance and coordination. Such systems can be research projects, simulations or controlled enterprise workflows; they are not evidence that major banks have deployed fully autonomous LLM trading fleets.
Can AI cause a flash crash?
AI and algorithmic systems can in principle contribute to rapid market feedback loops. There is no evidence that current LLM agents have caused a new flash crash. Regulators are studying future risks from greater AI autonomy and correlated behaviour, using historical episodes only as an analogy.
Who is responsible if an AI agent makes a mistake at a bank?
There is no universal single answer. Responsibility depends on jurisdiction, activity, contracts, regulatory requirements and supervisory structure. Banks generally cannot assume that using an outside AI provider removes their own compliance responsibilities.

Frequently Asked Questions

What is agentic AI in finance?
AI systems that pursue a goal across multiple steps by planning, using tools, accessing approved data and taking actions, rather than simply generating an answer. In finance they are used for research, software engineering, KYC preparation, reconciliation and client-service support.
What is the difference between generative AI and agentic AI?
Generative AI primarily creates or analyses content in response to a prompt. Agentic AI can go further by planning a sequence of steps, calling tools and performing approved actions to complete a goal.
Is Wall Street using AI agents?
Yes. Major institutions are piloting or deploying AI agents across operations, wealth management, client vetting, engineering, research, trading-related workflows and treasury, per July 2026 Reuters reporting.
Is Goldman Sachs using agentic AI?
Yes. Goldman has publicly confirmed development and use of agentic AI. Its Anthropic work has included agents for trade and transaction accounting, client due diligence and onboarding, and it uses AI in software engineering and a broader operating-model transformation.
Is Goldman using Claude Opus 4.6?
Anthropic released Claude Opus 4.6 in February 2026 and some reporting connects it to Goldman’s agent development. Goldman’s public confirmation refers broadly to Claude-based agents, so its wider strategy should be described model-agnostically unless a specific workload and model pairing is directly confirmed.
Are AI agents executing trades autonomously on Wall Street?
Not as a documented mainstream practice. AI is used across trading-related research, analytics, surveillance and workflow automation. Bank of England analysis in 2026 indicates fully autonomous AI trading is not the dominant use of agentic systems at major institutions.
What is bounded autonomy?
The model most banks use: an agent is connected to specific tools, data, permissions and workflows, with human approval gates around high-stakes actions. The more authority an agent has, the more important its permission boundary becomes.
What is human-in-the-loop AI?
A system where a human must review or approve a decision before a critical action occurs.
What is human-on-the-loop AI?
A system allowed to operate within defined boundaries while a human supervises and can intervene.
Is agentic AI the same as multi-agent AI?
No. A single agent can use many tools and still be agentic. A multi-agent system uses multiple specialised agents. Multi-agent architecture is an option, not a requirement, and it adds its own coordination-failure modes.
Did OpenAI’s function calling invent AI agents?
No. Function calling, introduced on 13 June 2023, made tool integration more structured and reliable. Agent concepts, ReAct-style research and early LangChain agents predate it.
What is Microsoft AutoGen?
A framework Microsoft Research published in 2023 for building applications with multiple conversing AI agents, which popularised multi-agent orchestration among developers.
When did Goldman Sachs roll out its AI assistant?
Goldman made the GS AI Assistant available firmwide on 23 June 2025, with about 10,000 employees already using it at rollout for tasks such as document summarisation, drafting and data analysis.
What is One Goldman Sachs 3.0?
A multi-year, AI-driven operating-model transformation Goldman detailed around its Q4 2025 earnings and 2026 annual meeting, with initial workstreams in client onboarding and KYC, enterprise risk, vendor management, lending, regulatory reporting and sales enablement.
What did Goldman and Anthropic announce in February 2026?
Goldman confirmed roughly six months of work with embedded Anthropic engineers building Claude-based agents for trade and transaction accounting, client due diligence and onboarding. It said it was in the early stages and planned to launch the agents soon.
Does that mean Goldman deployed autonomous agents across the bank?
No. As of the February 2026 report, Goldman was still developing these agents, focused on operational and post-trade work, with launch described as upcoming rather than complete.
What are Anthropic’s finance agent templates?
Ten templates launched on 5 May 2026: pitch builder, meeting preparer, earnings reviewer, model builder, market researcher, valuation reviewer, general-ledger reconciler, month-end closer, statement auditor and KYC screener, shipped as plugins and cookbooks.
Which banks use Anthropic’s finance agents?
Anthropic named financial-services customers including Goldman Sachs, Citi, Visa and AIG. Being named as a customer does not mean an institution uses all ten templates.
What did the KPMG survey find?
A June 2026 KPMG survey cited by Reuters found 51% of surveyed banks were piloting AI agents. This is a survey of surveyed banks piloting, not 51% of every bank worldwide, and not production deployment.
What is algorithmic trading, and is it agentic AI?
Algorithmic trading has existed for decades: a predefined rule or model receives market data and executes. Agentic systems can plan, reason, select tools and adapt a workflow. Not every algorithm is an AI agent.
Is high-frequency trading done by LLM agents?
No. High-frequency trading prioritises latency, deterministic execution and specialised infrastructure. LLM agents add reasoning and tool orchestration but also latency, non-determinism and model risk.
Can AI agents cause a flash crash?
There is no evidence current LLM agents have caused one. Regulators study whether correlated AI systems could amplify market moves if more autonomous AI is connected to financial decision-making, using past episodes as an analogy.
What is Project Logos?
A research project by the BIS Innovation Hub London Centre, the Bank of England and the Deutsche Bundesbank that observes LLM-based agents acting as portfolio managers in a simulated market, to study how they allocate capital and whether they amplify correlated decisions. It is a simulation, not a live deployment.
What does the FSB say about AI risk?
In an August 2026 letter to the G20, FSB Chair Andrew Bailey said frontier AI’s potential impact on cyber risk is the most immediate concern for the financial system, because it can alter the speed, scale and economics of cyber attacks and third-party providers are highly concentrated.
What does FINRA say about AI agents?
FINRA’s 2026 Annual Regulatory Oversight Report treats AI agents as an emerging trend and flags risks including excess autonomy and scope creep, authority beyond intended scope, weak auditability and transparency, data sensitivity, domain knowledge and incentive design, on top of hallucination, bias and privacy risks.
Does using AI remove existing financial regulation?
No. Rules on supervision, records, market conduct, best execution, risk controls and communications can still apply. ESMA’s February 2026 briefing reflects this technology-neutral approach.
What are the biggest risks of agentic AI in banking?
Excess autonomy, incorrect permissions, hallucinations, sensitive-data exposure, poor auditability, cyber risk, third-party dependence, prompt injection and difficulty controlling long multi-step workflows.
Does a multi-agent design make a workflow safer?
Not automatically. A risk agent checking a research agent adds defence, but if both share the same model, data and assumptions they can fail together. Safety diversity includes deterministic rules, independent data, hard limits and human review.
Do AI agents get a free pass on insider information?
No. If a professional cannot use material non-public information improperly, an AI agent acting for them cannot make that restriction disappear. Banks’ information barriers may require equivalent or stronger data segmentation for agents.
Do agents always reduce cost?
No. Agentic workflows can require many model calls, long contexts, multiple agents, tool execution, monitoring and human review, which can increase compute cost. Goldman Sachs Research expects agentic AI to materially increase token consumption.
Will AI agents replace bankers?
Tasks are likely to change before whole occupations disappear. Document review, data gathering, routine coding, reconciliation, pitch materials and basic analysis are most exposed; relationships, judgement, negotiation and regulatory accountability remain human.
What is the junior-banker training problem?
Many repetitive tasks also function as training. If agents prepare models, summaries and first drafts, firms have to work out how junior staff still learn the underlying work — a major 2026 management question.
Is this article investment advice?
No. It is editorial coverage of how banks and regulators are adopting agentic AI. It contains no trading signals, no buy/sell or portfolio recommendations and no instructions for autonomous trading.

Related AiTimeline Coverage

⚠️ Editorial & Accuracy Note

This article separates company and vendor announcements (Goldman Sachs, Anthropic), regulator publications (FINRA, ESMA, the Federal Reserve, the Bank of England, the Financial Stability Board, the BIS Innovation Hub), and news reporting (Reuters, CNBC), and labels each with a verification status. Figures were verified as of 2 September 2026. In India, the Securities and Exchange Board of India regulates algorithmic and automated trading and has consulted on AI/ML use in markets; there is no public evidence of major Indian institutions running unsupervised autonomous LLM trading. Nothing here is investment advice, a trading signal, or a recommendation to buy, sell or allocate. “Pilot” is not “production,” “research” is not “execution,” and “trade accounting” is not “trading.”

Advertisement