FinCrime Agent Benchmark
FCA Bench tests LLMs on real end to end anti-financial crime work, from standard operating procedure ingestion, to sanctions, PEP and adverse media screening and ending with alert review and adjudication.
3
500
9
Risk Categories
Test cases
AI Models
Introduction
Why We Built FCA Bench
Few spaces show as much promise for the application of AI as anti-money laundering (AML) and know-your-customer (KYC). The space is known for research intensive, repetitive work and 99% percent false positive rates. But the fact that one mistake can lead to million dollar penalties and reputational damage, as well as the need to explain AI to auditors and regulators means that finding effective, regulatory-friendly AI solutions is incredibly challenging.
Most AI benchmarks score models based on defined tasks with structured context. Numerous benchmarks already exist for legal, economic and coding scenarios, but AML/KYC doesn’t run on generic tasks. AML/KYC tools must be tuned to each financial institution’s unique risk exposure and risk appetite through an institution-specific standard operating procedure (SOP).
AI is only credible when every decision is consistent and ties back to institution’s own SOP. This is the standard expected by model risk management teams, auditors and examiners.
We created the FinCrime Agent Benchmark (FCA Bench) to measure:
Can AI accurately ingest an institution’s SOP?
Can AI identify relevant financial crime risk across sanctions, PEPs and adverse media?
Can AI accurately and consistently adjudicate financial crime risk?
How does a general purpose AI compare to a purpose-built agentic harness like Castellum.AI’s Arbiter?
Methodology
FCA Bench Structure
FCA Bench compares Castellum.AI's purpose-built LLM harness, Arbiter, against general AI models across a core AML/KYC workflow: Screening names for sanctions, PEPs and adverse media risk and adjudicating the alerts generated.
FCA Bench reports the difference between the two conditions. Each model was run through the benchmark six times, three times on it own and three times with the Castellum.AI harness, to ensure consistency.
The FCA Bench tests 500 cases composed of names and identifying details (DOB, ID and location) that are evaluated across both screening and adjudication. The structure was chosen to reflect real-world customer and transaction data.
Same model, two conditions
Each model is run through the benchmark twice — once on its own, once through the Castellum.AI harness.
9 models tested
The model on its own.
- Fixed prompt applied across each frontier model
- Given the financial institution's AML/KYC policy & procedure document directly
- Screens against the open web
- Returns an adjudication based on the sample policy and its screening results
The same model, inside the product.
- Maintained sanctions, PEP & adverse-media dataset, kept current
- Matching across transliteration, aliases & name collisions
- SOP-driven adjudication flow
- Returns a structured, cited decision
Full definitions and caveats are in the detailed methodology below.
Key Findings
A Purpose-Built Harness Improves LLM Accuracy by Over 50%
Standalone vs Castellum.AI's Harness
- Cases fully correct: every finding surfaced and every decision reached correctly for that subject.
- Decision consistency: the model reaches the same disposition on the same alert across three separate runs.
Unharnessed LLMs Present a Catastrophic Risk to Compliance Programs
Standalone LLMs reached incorrect conclusions between 69% (Muse) and 29% (Astra) of the time. While results like this for standalone LLMs are disappointing, it’s crucial to remember that in AML/KYC, even one error or missed true positive can lead to multi-million dollar fines and the dismissal of senior leadership. Quite simply, standalone LLMs are dangerous to rely upon for compliance work.
Standalone LLMs missed between 46% to 72% of true positives. Of the 152 true positives layered into the 500 test cases, standalone LLMs missed between 70 and 109 depending on the model.
Standalone LLMs could not effectively identify or adjudicate risk. The best standalone model, Astra, resolved only 71.1% of cases correctly (evaluated across 3 repeated runs). The worst, Muse, got just 30.8% right. This is the median standalone outcome, not an outlier, Gemma (53.6%), Kimi (50.6%) and Gemini (60.3%) all land well short of usable when not harnessed.
The weakest model, when harnessed, still beat the strongest standalone LLM by 28%. Muse achieved 99.2% accuracy with Castellum.AI's harness, beating the GPT 6 Astra (unharnessed) by 28% across the same cases (71.1%).
Only harnessed AI was consistently correct. Re-running each case three times, harnessed decisions agreed with themselves 97.2% of the time on average (95.1%–99.6%); standalone averaged 45.2% (31%–71%). An even more dangerous discovery is that standalone models didn't even surface the same alerts between runs, meaning that there was no ability to reproduce results and no ability to truly predict how an unharnessed AI would work over time. By contrast, an AI harnessed by Castellum.AI surfaced the same cases every run, with zero drift. In regulated workflows, having repeatable, explainable outcomes is a requirement for validation and regulator sign off.
Standalone LLMs were also unable to accurately query compliance data, creating an additional risk around tool use, on top of model drift, lack of consistency and inaccurate adjudications. Harnessed AI that used Castellum.AI’s risk dataset found 100% of expected hits. Standalone LLMs only found 61% of expected hits on average, (40.5%–81.1%), across sanctions, PEPs and adverse media, despite being given clear instructions and direct links to risk data sources.
Leaderboard
Screening Effectiveness
Here, FCA bench tested whether a standalone LLM could identify sanctions, PEPs and adverse media risk when pointed at direct and reputable sources (e.g. US Treasury sanctions lists, government PEP registries and reliable media).
The harnessed LLMs all used Castellum.AI’s purpose-built screening and matching algorithm as well as Castellum.AI's risk dataset (both of which are part of the harness).
Screening effectiveness
Can the AI identify the correct sanctions, PEPs and adverse media results?
Toggle between the different risk data types to see how LLMs perform across Adverse Media, PEPs, Sanctions and Overall.
Adjudication
Of the alerts the model raised, how often did it make the right call per the SOP?
Adjudication effectiveness
Of the alerts identified, how often did the AI reach the right decision against the SOP?
Detailed Results
Complete numbers behind the leaderboard. Every model, every stage and data type. Switch the view and sort any column.
Every model, every stage
The full numbers behind the leaderboard. Toggle between each model on its own, the same model through the Castellum.AI harness, and the uplift between them.
Sample Cases
Evaluating Financial Crime Risk
Given the same task to identify financial crime risk and adjudicate that risk against a prospective customer profile, the AI harnessed with Castellum.AI catches risk the standalone model misses and effectively dispositions the risk with clear rationales.
See it on a case
The Takeaway
The Harness Matters More than the AI Model
Foundational models on their own are not fit for AML/KYC. Standalone LLMs missed 55.7% of risk on average and agreed with its own prior decisions only 45.2% of the time, meaning that someone using standalone LLMs risks getting the wrong decision 55.7% of the time. A compliance program cannot be built on a coin flip.
AML/KYC is not a convenience or spray and pray use case like improving email outbound. Reducing analyst workload by half means nothing if the underlying accuracy doesn't hold. A tool that cuts alert review time by 50% while screening accurately only 80% of the time gets executives fired, not promoted. Unlike drafting an email or generating an image, a wrong answer here is a missed sanctions hit, a cleared bad actor, or a SAR that never gets filed. The bar to clear isn't "better than manual review." It's "defensible to an examiner" and the challenge, while often framed as a technical one, is actually a regulatory one.
The same models, run through Castellum.AI's harness, surfaced and adjudicated alerts consistently and accurately: 99.2%+ end-to-end accuracy and 97.2%+ decision consistency across every model tested, with zero drift in which alerts even surfaced from one run to the next. These metrics are based on the FAC Bench test and a baseline SOP. Live implementations with client-specific tuning reaches 100% accuracy and deterministic-level consistency.
An institution that chooses to build, not buy, must add a purpose built harness and take on the full regulatory burden of proving that its agent is reliable: SOPs correctly ingested and enforced, sanctions/PEP/adverse-media data kept current, and adjudication that produces the same defensible decision every time and under ongoing model risk management for as long as the agent stays in production.
Benchmark Test Details
FCA Benchmark Testing
Each model was run three times under two conditions (AI standalone or AI using Castellum.AI’s harness), using the same 500 case task set. Both conditions start with the ingestion of a baseline SOP and each condition then followed the same two-step flow: Screening for risk followed by adjudication of identified risk.
SOP Details and Prompting
The AI was provided a baseline SOP representative of a financial institution’s policies and procedures pertaining to screening and adjudication for sanctions, PEPs and adverse media risk in the course of KYC onboarding, ongoing monitoring (perpetual KYC or pKYC) and payment screening.
Standalone AI was provided risk-specific prompts incorporating the SOP into the screening and adjudication process. Included in the prompts were specific target instructions for financial crime risk sources (including pointing the AI directly at US OFAC sanctions).
Castellum.AI-harnessed AI utilized the company’s agentic screening and adjudication infrastructure.
500 Case Task Set
The FCA Benchmark used 500 cases split across three financial crime risk categories: Sanctions (160 cases), Politically Exposed Persons (180 cases) and Adverse Media (160 cases). Cases across each risk category incorporated exact name matches, fuzzy name matches (Characters within a name are altered) and no-match cases (e.g. Name should not generate an alert). Only names were used for screening, with supplementary identifying information (exact or mismatched DOB, IDs, locations, etc.) appended to the case for the adjudication process.
Screening Process
Standalone: Provided as text in the screening prompt, including explicit instructions on where to find the data. Standalone runs were given web search as a tool for this step; two models (GPT 5.5 and Muse Spark 1.2) were additionally run in a no-web-search baseline to isolate that tool's contribution.
Castellum.AI Harness: Enforced structurally, as search parameters within Castellum.AI’s product.
Adjudication Process
Standalone: Provided as natural-language instructions in the prompt, with no mechanism guaranteeing the model actually follows them case to case.
Castellum.AI Harness: Enforced through a decision matrix built from the SOP.
This distinction is the crux of the harness/no-harness comparison: the same institutional policy is either compiled into a ruleset the system enforces every time, or handed to the model as a suggestion it may or may not apply consistently.
Grading
FCA Bench scores both screening and adjudication to confirm whether the standalone or Castellum.AI harnessed AI generated a relevant alert and adjudicated the generated alert correctly.
Screening and adjudication pass rates were combined for a end-to-end score, where the AI passes if both screening and adjudication are correct, reflecting a real-world compliance use case where true positive risk must be identified and escalated.
Consistency was scored as agreement across the three repeated runs for each AI under each conditions for each case.
Testing flow
Frequently Asked Questions
Common Questions
-
How the same LLM performs on a key financial crime compliance workflow: screening and adjudicating financial crime risk across sanctions, PEPs and adverse media. FAC Bench measures how different LLMs perform on their own and against Castellum.AI’s agentic harness.
-
Castellum.AI brings a consistent adjudication flow to whatever model sits behind it along with a maintained sanctions, politically exposed persons (PEP) and adverse-media data layer. All nine models that the FAC Bench tested improved when paired with Castellum.AI’s harness. It supplies the data, matching, policy logic and controls that make adjudications accurate, consistent and defensible.
-
The FCA Bench is a point-in-time evaluation on a test set covering one SOP and 500 test cases. The standalone LLMs used a fixed prompt for screening and fixed prompt for adjudication. A fixed prompt was used to ensure fairness across LLMs and test each under the same conditions. Each standalone LLM was also configured to use built-in web browsing tool(s) to search and retrieve relevant risk data when pointed to specific risk data sources (e.g. US OFAC’s official sanctions list).
-
FCA Bench test results are current as of 16 September 2026. The benchmark is periodically updated to include new frontier models and validate scores for existing benchmarked models.
-
Yes. The test cases are built against Castellum.AI's financial-crime risk database of sanctions, PEPs and adverse media, which is updated in real time. Unharnessed LLMs frequently fail to identify relevant sanctions, PEPs and adverse media risk even when provided specific instructions and direction to reference official data sources like OFAC.
-
Directly. The Castellum.AI harness surfaces risk using current data and returns a structured disposition for each alert, creating a clear record of the evidence considered, the policy applied and how the decision was reached for model validation and audit review.
-
Aggregate results, detailed methodology and sanitized examples are available on request to approved parties. Reach out to contact@castellum.ai for more information.
Learn more about Castellum.AI’s harness
Connect with the team to see how our agentic harness improves accuracy, consistency and explainability across AML/KYC workflows.