Banking & FinancialGenerative AI · Governance & Safety

Reviewing every recommendation instead of two percent of them

Advice surveillance that reads the recommendation, the client profile, and the rationale together — flagging what a compliance officer should look at, and never issuing a finding on its own.

EngagementFixed-scope build
Timeline to v14–8 weeks
PatternRepresentative engagement
100%
of advice reviewed, up from ~2%
more issues caught pre-complaint
0
automated adverse findings

The problem

A wealth and advice business reviewed suitability the way most do: a random sample of roughly two percent of recommendations, examined after the fact by a compliance team that could not possibly cover more. The sample caught problems eventually, usually after a client complaint had already been filed, and it gave the board very little assurance about the ninety-eight percent nobody read.

The hard part is that suitability is not a rule check. Whether a recommendation fits depends on the client's objectives, risk tolerance, time horizon, existing holdings, and circumstances — and on whether the documented rationale actually connects those things to the product recommended.

The hypothesis

Language models are good at exactly this kind of reading: comparing a recommendation and its stated rationale against a structured client profile and identifying where the reasoning does not hold together. Not to decide suitability — to direct expert attention to the cases most likely to be unsuitable, so that reviewers cover the population instead of a sample.

The build

  • Structured client profile assembly — objectives, risk tolerance, horizon, capacity for loss, existing portfolio, and recorded circumstances pulled into a consistent representation per client.
  • Recommendation and rationale reading — the advice document and the file notes read together, with the specific claims linking client need to product identified and extracted.
  • Coherence and concentration checks — deterministic rules for the things that should be deterministic (concentration limits, product permissions, fee disclosure presence) alongside model-based review of whether the rationale actually addresses the client's stated objectives.
  • Triage into a reviewer queue — cases ranked by concern with the specific tension quoted: a growth product recommended to a client whose file states capital preservation, a rationale that never mentions the horizon, a switch with no cost comparison.
  • No findings, ever — the system flags and cites. Every determination of unsuitability is made by a compliance officer, recorded under their name.

Design choice that mattered: we treated a missing rationale as a stronger signal than a questionable one. Advice that is poorly justified in the file is a documentation risk regardless of whether the product was right — and it is the thing that loses complaints and enforcement cases years later.

Rollout

We calibrated against historical cases with known outcomes, including upheld complaints, and measured whether the system would have surfaced them. Compliance reviewers then worked the ranked queue in parallel with the existing random sample for a full quarter, which gave both a precision measure and a direct comparison of what the sample had been missing. Fairness and consistency checks ran across advisers and client segments to ensure the flagging was not concentrating on particular groups for spurious reasons.

Results

Coverage went from a small sample to the full population, materially more issues were identified before they became complaints, and — the outcome compliance valued most — the firm could describe its advice oversight to its regulator as complete rather than sampled. Reviewer headcount did not change; what they read did.

What we'd tell you

  • Triage, never adjudicate. In regulated advice, the finding must belong to a named human.
  • Keep the deterministic checks deterministic. Concentration limits are code, not inference.
  • Calibrate against upheld complaints — they are your ground truth and you already have them.
  • Test the flagging itself for bias across advisers and client segments. Surveillance can quietly become unfair.
← All success stories
Keep reading

More from the field.

Banking & Financial

Consumer lending, decided in hours

60%
faster time-to-decision
files per underwriter
Read the story
Banking & Financial

Early-warning credit risk

60 days
earlier risk signal
↓ roll rates
into later buckets
Read the story
Banking & Financial

The 6 a.m. advisor brief

advisor capacity for client time
6 a.m.
brief ready daily
Read the story
Get started

Bring us a hypothesis. Leave with a system.

Tell us what's eating your team's time. We'll give you an honest read on whether AI is the right tool — and if it is, a scoped v1 with a timeline and cost.