Features Hub

When the AI System Can’t Be Held Accountable, Who is?

Wed 3 Jun 2026 | Ritesh Singhania

Humanoid robot holding a document against a digital background with currency symbols and data graphics, representing AI in financial analysis.

At the IMF spring meetings in Washington, finance ministers and central bank governors found themselves in emergency discussions about an AI model.

Anthropic’s Claude Mythos had autonomously identified exploitable vulnerabilities across every major operating system and browser – capabilities that Anthropic itself described as unprecedented in their ability to identify and exploit weaknesses at scale.

The Bank of England Governor, Andrew Bailey, has called it a serious challenge, while Canada’s Finance Minister compared it to an unknown threat with no clear boundaries – and the US Treasury urged major banks to test their systems before any public release.

What makes that threat so serious for financial institutions specifically is not only what adversaries could do with those capabilities – it’s that institutions will deploy them internally too. Mythos-level AI will be used inside regulated firms to automate decisions, execute tasks, and operate across systems at a scale no human oversight function can match. The governance frameworks that those firms are relying on were not built for that, and the consequences of a failure won’t be visible until significant harm has already occurred.

Legal Accountability Was Designed for Inspectable Systems

Senior managers are legally responsible for outcomes in their area – whether or not AI was involved in producing them. In the UK, that accountability is formalised through the Senior Managers and Certification Regime, where named individuals carry prescribed responsibilities for specific functions.

The core problem is that accountability was designed for a different kind of system – one where the person responsible could inspect the model, trace a decision, and validate an output. As the AI systems now being deployed are more complex and probabilistic, that direct interrogation becomes increasingly impossible. Legal accountability is outpacing meaningful oversight because generative and agentic AI doesn’t work that way.

Direct model inspection is increasingly no longer a realistic basis for oversight, either. The teams building and deploying AI systems are increasingly using AI to validate them – deploying models to evaluate the outputs of other models. Human oversight functions cannot match that at the same depth or pace, meaning independent challenge can no longer mean directly inspecting every model. It has to mean governing the standards and conditions under which automated validation is conducted – and maintaining the judgment to escalate when it fails.

From Model Oversight to Control Oversight

That is the shift we are experiencing – from model oversight to control oversight. Being accountable stops meaning “I signed off on this model” and starts meaning “I can show the control environment is sound – and I know when it isn’t” – governing data inputs, prompt structure, escalation triggers, and the points at which human judgement must intervene.

Zango’s research, drawing on interviews with 27 C-suite and senior leaders across UK and European financial institutions, found that in several firms, compliance and risk functions had limited visibility into which AI tools were even being used across the organisation. One practitioner put it plainly: “If I ask the question – show me everywhere AI is being used across this organisation – I wouldn’t be able to get an answer.”

That is a skills problem as much as a structural one. Compliance and risk professionals need to know what questions to ask of systems they often can’t directly inspect and recognise when automated validation isn’t sufficient. UK Government research has identified the key AI skills gaps in financial services as being in governance, ethics, and interpreting AI outputs – concentrated in compliance and legal teams. Our research corroborates that finding.

The Need for Sector-Specific Standards

What’s missing – and what this moment should accelerate – is an authoritative, sector-specific standard translating regulatory expectations into operational AI governance practice.

Through our work deploying AI agents for risk and compliance functions across financial institutions, we see the consequences of that gap directly – every firm approaching the same governance questions differently, with no shared baseline to work from. The US moved to fill this gap in February with a public-private Financial Services AI Risk Management Framework. GovTech Singapore has done equivalent work on agentic systems specifically. The UK and EU have not.

From Urgency to Structural Reform

The Mythos moment has concentrated minds at the highest level. That urgency is welcome – but the response can’t stop at patching vulnerabilities. That means moving oversight from models to control environments, and building the sector-specific implementation standard that the US and Singapore have already developed. Without it, Mythos-level capabilities won’t only pose external threats – they will create risks inside financial institutions that governance frameworks cannot address.

Experts featured:

Tags:

agentic AI AI Accountability AI Governance Standards AI Regulation Antropic
Send us a correction Send us a news tip

Subscribe for News in Your Inbox