>>
Technology>>
Artificial intelligence>>
How Sphinx Turned Multi-Agent ...Most compliance software tries to give analysts a faster answer. But false-positive rates on transaction alerts often sit between 90 and 95 percent, according to PwC analysis of AML alert engines.
The goal, says applied machine learning specialist Alexandre Berkovic, should be to have an auditable path from evidence to judgment: a decision that can be reviewed, challenged, and traced back to the information that produced it.
An expert in ML for financial crime and compliance, Berkovic believes a plausible output is not enough: "These figures of false positives can sometimes be even as high as 98 percent, and I realized something needed to be done to fix that."
However, his response is not to replace the analyst. Berkovic, who has in-depth knowledge in multi-agent systems that make onboarding, sanctions screening, and transaction-alert decisions auditable for supervised banks, has come up with a different solution.
He solved the problem by making multiple AI agents debate the case before a human ever sees a recommendation.
Berkovic explains: "By fine-tuning the large models, and designing multi-agent decision architectures, we can solve the hard problems of reliability, interpretability, and calibration. These arise when machine learning systems are placed in high-stakes production environments."
This is how he founded Sphinx, making it possible to apply multi-agent machine learning to financial crime and compliance work.
Working mainly with regulated banks and fintechs, the company runs on Berkovic's proprietary Interpretable Agentic Framework. The system has AI agents argue a case before a human reviewer sees the recommendation. That setup is built for onboarding, sanctions screening, and transaction alerting, with a decision trail auditors can follow.
Berkovic did not start by taking a general-purpose large language model and bolting an explanation on top. He started from a different question: what would a compliance decision look like if disagreement were built into the system? That question is what led to Sphinx's proprietary Interpretable Agentic Framework.
That thinking dates back to his training. He studied at Imperial College London, where he completed an MEng in Design Engineering with First Class Honors, then attended Massachusetts Institute of Technology (MIT) as a Master of Business Analytics candidate in 2022–23, with work tied to MIT Sloan and the Operations Research Center. At MIT he was already treating hard problems as adversarial and sequential, not static puzzles with a single neat answer.
Alexandre Jacquillat is Associate Professor of Operations Research and Statistics at MIT Sloan. He taught Berkovic in 15.093 and watched that approach show up in class. What stood out to him was not flashy technique. It was judgment: picking a method because it matched the shape of the problem. Jacquillat says: "What I valued in his course project was not the technique alone but the judgment behind it: he saw that the problem was adversarial and sequential rather than static, and he chose robust optimization because it fit the structure, not because it was available."
Berkovic's system splits a compliance decision across three roles: defender, prosecutor, and judge.
How does it work? The defender makes the best case that the customer or transaction is legitimate. The prosecutor looks for risk, contradictions, adverse information, or gaps that still need an answer. The judge weighs both sides and issues the final calibrated decision.
Berkovic identifies that system as follows: "The system I am most proud of is the proprietary Interpretable Agentic Framework we developed as a data science architecture."
Each agent receives the complete case record, including customer documents, identity attributes, and open-source intelligence. That matters because the debate is grounded in the same evidence, not separate or incomplete prompts. Competing interpretations can therefore confront the same facts before the final decision.
The architecture also changes what interpretability means. In a conventional machine-learning system, an explanation can be attached after the model has produced an output. Here, deliberation is part of the decision process itself. The roles created in the system then produce a sequence of claims, challenges, and adjudication that can be inspected alongside the final ruling.
For compliance teams facing large volumes of ambiguous alerts, this system can be essential. The structure becomes reasoned so an analyst can assess the evidence and a regulator can follow the path to the outcome.
Berkovic says: "When the defender and prosecutor must confront the same evidence, weak arguments become visible before the judge rules. It means the final decision becomes reviewable."
This approach also gives calibration a defined place in the architecture. The final agent weighs the competing cases against the institution's risk thresholds and available evidence. This builds a practical bridge between generative reasoning and a controlled compliance workflow.
Berkovic says: "A single model tends to collapse ambiguity into one answer. But a structured debate forces the system to surface the ambiguity and then resolve it under explicit roles, which is closer to how experienced compliance teams actually reason."
The deployment method is equally important. By starting from the tools and thresholds the institution already uses, the system inherits its risk appetite and operational constraints rather than imposing new ones. That is what allows production deployment in weeks rather than months.
Sphinx reports a 94 percent reduction in hallucinations relative to single-model systems. That figure is a company-reported comparison, but it supports the underlying argument: reliability can improve through architecture, evaluation, workflow design, and evidence tracking rather than post-hoc explanations alone.
Ian Gilligan, who holds CAMS and CAMS-FCI credentials, has seen the model in action. He led Axos Bank's 2025 review of sixteen AML AI vendors. He put Sphinx's multi-agent setup next to the rest of the field.
Gilligan says: "Every other system I reviewed was built to make an investigator faster. His system conducts the investigation. One agent argues for suspicion, a second argues against it, a third rules against the bank's own risk framework and writes a narrative a reviewer can treat as a work product. It's impressive, it makes a difference, and it works."
A prototype can show that two models disagree. But a production system has to fit inside a compliance team's existing way of working, without forcing a bank to rebuild its technology stack. Sphinx treats its agents as an extension of the team, not a replacement.
Berkovic explains the practical setup: "When the clients supply credentials, our agents interact directly with their existing tools."
The result is a deployment model closer to giving a new analyst system access than to running a traditional software integration project.
However, behind that setup is a harder data-science choice: "The system models how experienced analysts actually work," says Berkovic.
"Compliance judgment combines structured fields, documents, and external intelligence, as well as procedural rules, and tacit knowledge.
"Sphinx tries to encode those judgments into generative and multi-agent systems while preserving an observable trail of what happened."
Berkovic puts the method this way: "We treat the compliance process as an observable data-generating process, mine it for structure, and then reproduce it at machine speed while preserving full decision provenance."
That means the workflow now becomes a source of training and evaluation information: "Analysts generate evidence about which sources matter," says Berkovic.
"It involves looking at contradictions which cause escalation, and which missing documents block onboarding. You can also look at which patterns lead to a disposition. Reproducing those structures is different from automating one isolated task."
However, the operational burden is substantial. Berkovic says that over time, that volume desensitizes analysts and increases the risk of missing a true positive.
A second structural frustration can be the hiring constraint. The traditional way to scale compliance capacity is to add headcount, but experienced analysts are scarce and expensive.
But the benefit is that the system is built to handle repeatable investigation and evidence assembly, while human analysts keep the judgment calls and escalation decisions.
Berkovic says: "The goal is to remove the volume of routine analytical work that currently prevents analysts from applying judgment where it is most needed."
He believes his method sits at the infrastructure layer of financial crime prevention. Often, model error or system failure can expose a client to multi-million-dollar fines. In extreme cases, it can jeopardize an operating license.
Berkovic says: "That reality imposes continuous pressure on model validation, monitoring, and safeguards. The responsibility is structural, not episodic."
Deployment follows the same principle. Berkovic says: "We typically integrate into their systems within one to two weeks. A dedicated group of data scientists, engineers, and compliance analysts works directly with the client to replicate existing workflows as closely as possible. Once the system is live and performing to specification, the operational goal is for the client to no longer need to manage the underlying technology. It functions as a reliable, auditable extension of their compliance function."
Sphinx currently works with publicly regulated U.S. financial institutions that manage tens of billions of dollars in assets. Its fintech clients include Equals Money in the UK, Roundtable and Fipto in France, and a range of U.S. companies in embedded finance, banking-as-a-service, and cross-border payments. They also include credit unions and digital-asset infrastructure.
Sphinx's architecture also operates inside publicly regulated U.S. financial institutions supervised by the Office of the Comptroller of the Currency, the Federal Deposit Insurance Corporation, and the Federal Reserve.
Berkovic says: "Because the framework is already running in production inside entities supervised by the OCC, the FDIC, and the Federal Reserve, the benefits are operational rather than theoretical. They appear as reduced false positives, faster legitimate onboarding, cleaner audit trails, and the capacity to serve more customers without proportional increases in compliance headcount. Outcomes that simultaneously improve competitive position and supervisory standing within the U.S. financial system."
The resulting system is designed around failure analysis and predictable behavior while retaining the human compliance function as the authority for judgment and escalation. Debate is useful because structured disagreement can expose weak reasoning before a final decision enters an accountable process.
Berkovic believes clients need regulatory-grade results without sacrificing speed, cost control, or reliability. Even a bank with a solid compliance stack can still drown in case volume. Meanwhile, a fintech company may need to grow without hiring a matching army of analysts.
He wants his client experience to feel closer to a high-touch analytical consultancy engagement. This is why the Interpretable Agentic Framework focuses on two principal workflow areas.
Berkovic says: "Enhancing the analytical capacity of financial-crime and compliance teams on two primary fronts: customer onboarding and transaction alerting. Onboarding involves multi-source document verification, identity resolution, open-source intelligence gathering, and procedural checks. Alerting involves investigating potentially suspicious activity that may require regulatory reporting."
That distinction affects the deployment design. While onboarding requires rapid synthesis across heterogeneous evidence before activating an account, transaction alerting requires investigation of activity that may or may not warrant regulatory reporting.
Berkovic says: "The architecture can support both, while the evidence, workflow, and decision thresholds remain specific to each process."
Berkovic says U.S. financial institutions lose hundreds of billions of dollars annually to fraud and related activity. Nasdaq Verafin's 2026 Global Financial Crime Report supports this, estimating U.S. fraud losses at about $196 billion in 2025, including roughly $179 billion in bank fraud. That same year, the FBI's Internet Crime Complaint Center logged $20.9 billion in reported U.S. internet-crime losses, up from $16.6 billion in 2024.
He adds that by designing data science systems that reduce those losses, it is possible to accelerate legitimate customer onboarding and produce decisions that withstand regulatory scrutiny: "This is concretely valuable to the American financial system."
The figures speak for themselves when it comes to Sphinx's success rate. For Equals Money, it reports a 94 percent reduction in false positives and an increase in straight-through processing from 42 percent to 87 percent.
Sphinx also says average onboarding time at another client, Alviere, fell to 2.4 minutes, while case-review time at Equals Money dropped from 15 minutes to 2 minutes.
Berkovic explains: "At Alviere, average onboarding time fell to 2.4 minutes. At Equals Money, case review times dropped from fifteen minutes to two minutes, and the cycle from identifying missing documentation to contacting the customer moved from multiple days to minutes. We report sanctions-screening precision of 99.9 percent across our portfolio."
He adds another client realized "more than 200 hours of analyst time saved per month, and another achieved approximately 10x faster resolution times."
Jorgen Osio Norgaard, Chief Compliance and Risk Officer at Alviere and a Sphinx customer, says: "The numbers are public: 86.1 percent of cases closed with no human touch, 99.7 percent when you include correct escalations, a 98.7 percent false-positive detection rate, and seventeen days of investigative work returned to the team in a single month."
Sphinx has also reported millions of alerts processed and hundreds of thousands of cases across multiple countries.
Berkovic explains how they do it: "By systematically reducing low-value analytical work, we enable institutions to grow customer bases and transaction volumes without linear growth in staffing. We can also improve decision provenance. These cleaner, machine-generated audit trails strengthen relationships with partner banks, external auditors, and U.S. supervisors. It's a win-win."
Berkovic tends to show up in industry forums in the same way he runs technical work: directly, with little distance between himself and the problem.
His public appearances include the Fintech Summit in San Francisco, Imagination In Action, and a TEDx talk focused on managing AI agents rather than promoting software. In 2024, France's presidential palace invited him twice: first to a May convening of French AI talent at the Élysée, and again in September for preparatory talks ahead of the AI Action Summit. In 2026, he served as a judge for both the MCP Hackathon and RoboHacks at Y Combinator.
He is also a member of the Association of Certified Anti-Money Laundering Specialists (ACAMS). That membership started in 2025, and it matters because Sphinx operates in the same regulatory environment ACAMS was created to support.
Those descriptions match the production challenge in regulated AI: the work has to be technically sound, deployable, and usable by people who are not data scientists.
Production-grade compliance AI has to answer more than whether a model can classify a case. It must show what evidence it considered, how it handled competing interpretations, why it crossed a threshold, and how an analyst or regulator can reconstruct the decision.
Sphinx's architecture turns those requirements into a working process. The defender and prosecutor create competing cases; the judge produces the calibrated ruling; and the surrounding system preserves the evidence and provenance needed for review.
Berkovic says: "The architecture is deliberately designed for interpretability and is now running in production inside publicly regulated U.S. financial institutions supervised by the OCC, the FDIC, and the Federal Reserve." The intent is to support human judgment, not replace it.
The intended operating state is for the technology to become a reliable, auditable extension of the compliance function rather than something the client continues to manage.
Berkovic puts it this way: "As someone who helped convert research-grade generative and multi-agent machine learning into operational systems that regulated institutions can trust, systems that produce more consistent, better-calibrated decisions than purely manual processes and that strengthen the practical defense against financial crime at scale."
The operational meaning is straightforward. Debate becomes useful when it creates a record; reliability becomes measurable when it survives real workflows. Automation becomes credible when people can inspect and challenge the reasoning.
Comments