Banking & FinancialAgentic AI · Governance & Safety

Trade surveillance: clearing the noise without missing the signal

Legacy lexicon rules generated thousands of alerts a day and almost no findings. An agentic triage layer that correlates trading and communications, drafts a rationale, and leaves every disposition to a human.

EngagementFixed-scope build
Timeline to v14–8 weeks
PatternRepresentative engagement
71%
false positives cleared with evidence
faster alert-to-disposition
Full
rationale on every case

The problem

A capital markets business ran surveillance for market abuse across trading activity and recorded communications. The system was rule- and lexicon-based, and it behaved the way those systems do: it flagged every message containing certain words and every trade fitting a coarse pattern, producing thousands of alerts a day of which the overwhelming majority were plainly benign on inspection.

The consequence is the dangerous part. Surveillance analysts working an unclearable queue develop a rhythm of rapid dismissal, and rapid dismissal is precisely how a genuine case gets closed in eight seconds. The firm's regulator was less interested in the alert count than in whether anyone was meaningfully reviewing it.

The hypothesis

Most surveillance triage is context assembly: what does this trader normally do, what was the market doing at that moment, does this message actually relate to that order, has this pattern been reviewed and dismissed before. That work is mechanical. If an agent assembled it and drafted a reasoned view, analysts could spend their day adjudicating rather than clearing.

The build

  • Context enrichment per alert — the trader's behavioral baseline, the order and execution timeline, contemporaneous market data and news, related alerts within a window, and prior dispositions for similar patterns.
  • Communications correlation — messages linked to the specific orders they plausibly relate to by instrument, timing, and counterparty, so a flagged phrase is read against what was actually happening rather than in isolation.
  • Language understanding over lexicons — replacing keyword triggers with models that read intent in context, which cuts the flagging of ordinary desk chatter that happens to contain a trigger word while catching euphemism the lexicon never had.
  • Drafted rationale, cited — the agent proposes benign, review, or escalate, and must cite the specific evidence. The reasoning is stored with the case permanently.
  • Human disposition, always — no alert is closed by the system without analyst sign-off, and escalation to compliance is a human act. Autonomy was granted per alert typology, only after demonstrated agreement, and never for the typologies that matter most.

Design choice that mattered: the asymmetry is not close. A false escalation costs an analyst twenty minutes; a false dismissal is a regulatory event. We tuned deliberately toward over-escalation and reported those thresholds to compliance in writing rather than burying them in a configuration file.

Rollout

Three months of shadow running, with the agent producing a view that analysts never saw until after they had recorded their own. That produced an agreement rate per typology and, more valuably, a clear map of where the system was reliable and where it was not. Insider-dealing typologies remained fully human throughout, by choice.

Results

The bulk of benign volume now arrives pre-cleared with the evidence attached, alert-to-disposition time fell several-fold, and analysts read fewer cases far more carefully. The firm can also now answer the question its regulator actually asks — show us your review process — with a documented rationale on every single case rather than a timestamp.

What we'd tell you

  • Shadow mode, per typology, before any autonomy. Aggregate accuracy hides the typologies you care about.
  • Write the threshold asymmetry down and share it with compliance. It is a policy decision, not a tuning parameter.
  • Store the reasoning permanently. A disposition without a rationale is indefensible three years later.
  • Some typologies should stay entirely human. Choosing that deliberately is a strength, not a gap.
← All success stories
Keep reading

More from the field.

Banking & Financial

Consumer lending, decided in hours

60%
faster time-to-decision
files per underwriter
Read the story
Banking & Financial

Early-warning credit risk

60 days
earlier risk signal
↓ roll rates
into later buckets
Read the story
Banking & Financial

The 6 a.m. advisor brief

advisor capacity for client time
6 a.m.
brief ready daily
Read the story
Get started

Bring us a hypothesis. Leave with a system.

Tell us what's eating your team's time. We'll give you an honest read on whether AI is the right tool — and if it is, a scoped v1 with a timeline and cost.