Scalable Capital · Financial crime platform · Product case study

One client, four products, no one watching the sequence

An MVP for transaction monitoring across Broker, Wealth, Credit and Banking, and the plan to deliver it in six months with six data engineers.

Build the client-level record of behaviour first. Every later control — fixed rule or trained model — is written against it.

Today Scalable can answer “was this transaction unusual for this product?” It cannot answer “was this unusual for this client?”

Task 1The MVP
What this proposes to build, in one picture
26 weeks · six data engineers · three controls · go-live at week 14 · stages detailed in 2.1
WEEK 0210 141824 26 M0M1M2 M3M4M5 DiscoveryThe record The loopRestrict CompleteDecide First-party layering alerts only Credit and loan fraud restricts first Account takeover needs device data GO-LIVE · WEEK 14 1 Back-test 2 Shadow — nobody sees it 3 Alerts reach the queue — an analyst sees it 4 Restriction applied — the customer feels it

1.1The problem this solves, and what good looks like

The problem and the outcome the MVP is aiming at.

A layering sequence across four days and three product lines A client deposits ninety thousand euros into the settlement account, buys and sells an ETF in the broker, adds a second reference account and makes it primary, then withdraws eighty-eight thousand euros to that new account. Each product sees one ordinary event. Only the client record sees the sequence. DAY 1DAY 2 DAY 3DAY 4 AS EACH PRODUCT SEES IT nothing here is unusual Banking & settlement a deposit and a payout Broker two ordinary trades Reference accounts a customer preference €90,000 in from reference account A €88,000 out to reference account B Buy ETF €89,400 Sell ETF €88,900 Account B added own name, verified B set as primary payouts now go here AS THE CLIENT RECORD SEES IT this is the alert Client deposit → trade → reference account switch → withdraw 98% of the deposit round-tripped in four days, and the payout account changed mid-sequence
Figure 1  A first-party layering sequence. No product line sees more than one ordinary event. The pattern exists only where the products are joined at the client, which is the record this MVP builds. The same case reappears in Figure 5 as the alert an analyst opens.

A platform that scales with Scalable Capital

Scalable Capital started as a neobroker and has been widening its product suite ever since. Each new product is another touchpoint in the customer's financial life, and the banking licence added the ability to hold and lend money.

That growth is the problem. The more products a customer holds, the harder it becomes to see their financial crime risk whole. Risk is assessed per product, but it is carried by the customer.

So the platform unifies the products and their events at the client, and it does so in a way that scales as the suite keeps growing.

The gap that matters

Each product line sees one ordinary event. The signal lives in the aggregation across the customer, which no single product can see.

So: build the record. Products become inputs to one client-level view. Controls are written against it.

What good looks like

Two north stars, one per category. ★ 1 Precision for how well the system detects; ★ 2 Cycle time per alert for how well it runs. Everything else on this table is a supporting metric — it explains a north star moving, it does not replace it.

CategoryMeasureDefinitionWhy this one
System
performance
★ 1  PrecisionTrue positives over all alerts, measured per control and across the system.The primary number. It is what every other improvement shows up in.
Third-party requestsAnother institution contacts Scalable about a client it suspects of crime.A proxy for what we miss. A true positive here is crime inside the bank that we did not catch. Deliberately caveated, and never improved by taking in less.
Authority follow-upThe financial intelligence unit forwards a report to the police, and an investigation follows.The strongest external signal that a report was worth filing.
Operational
efficiency
★ 2  Cycle time per alertAlert raised to alert closed.Serves efficiency and the customer at once. See below.
Alerts per analyst per dayVolume against the agreed capacity ceiling.The hard limit on how strong the control set can be. Section 2.2 turns on this number.
Alerts per active clientTotal alerts divided by monthly active clients.Has to fall over time. This is the actual proof that the platform scales rather than grows headcount.

Two outcomes that matter and resist measurement

Cheaper to extend. No clean metric. Testable form: the twelfth product costs less to integrate than the fourth.

Innocent customers feel as little as possible. Only restriction severity is this team's — no front-end capacity here (A5). So: graded restrictions, not one blunt freeze. Harm measured as days under the most severe active restriction and money blocked over time.

For a true positive, time under restriction is the cost of doing the job. For a false positive it is pure harm. Cycle time is the customer-experience lever this team actually owns.

Compliance sets the risk appetite. Never traded against. See 2.5.

1.2What the MVP contains, and what it deliberately leaves out

Scope, and the reasoning behind what is excluded.

Every product goes into the record. Release 1 then writes controls for the three typologies that carry the most value, and leaves the rest absorbed but unwatched — so covering them later is a control to write, with no change to the data model.

FINANCIAL CRIME Money laundering First-party fraud Third-party fraud First-party layering Credit and loan fraud Account takeover Highest volume Biggest lever on the analyst queue Direct liability The only place the bank loses its own money Highest severity A customer robbed while we were watching
Built in release 1
Monitored, with alerts reaching a human
One typology per branch above
  • First-party layeringAbsorbs structuring and source-of-funds cases
  • Credit and loan fraudThe only place Scalable loses its own money
  • Account takeoverHighest severity per case
Three product surfaces
  • Bank and settlement account
  • Credit and loans
  • Securities transfer in and outA closed cash loop is not a closed value loop
And the platform underneath
  • The label loopAlert to case to typed outcome, stored in our own backend
  • The explanation on every alert
  • Vendor integrationAlerts, scores and outcomes written back to us
In the record, not monitored
Ingested and joined at the client. No control written against them yet. Adding one later costs a control, not a schema change.
  • Crypto
  • Derivatives
  • Savings plans
  • Wealth and managed portfoliosWhite-label included: a partner arrangement means two firms monitor, not one
  • Private equity
  • European Investor Exchange
Why absorb them at all
  • Every product offered needs a risk assessment, and products change into shapes that need watching later. The cost of being ready is one ingestion, paid once.
Out of scope
Each with the reason on the card
  • Sanctions and PEP screeningList matching against a name, checked at onboarding and rescreened periodically. Not behavioural, and a separate system.
  • Market abuse and insider dealingA different regulation, driven by order book and news data, and watched at venue level.
  • Third-party mule networksRemoved by the closed loop today — but the data model must not preclude them
  • Identity fraud at onboardingCaught at onboarding. This brief is about product usage, which is everything after it.
  • Investment scam, victim sideScalable cannot be the collection point; the remainder looks like account takeover
  • Anything customer-facingNo front-end capacity on this team (A5).
Assessed, not built

Two pieces of work inside the MVP produce a decision rather than a system: the build-versus-buy assessment on case management, and the time-boxed check on whether a model can be trained at all. Both are sized as analysis, and both are on the critical path because what they conclude sets the next release.

Figure 2  The middle column is the one that needs the picture. In prose, “absorbed but not monitored” reads as a hedge. Drawn beside the other two states it reads as what it is: the decision that keeps the twelfth product cheap.

What release 1 has to have produced to be finished

Release 1 is done when: three typologies are covered by controls that explain themselves at the right latency; every closed case carries a typed outcome label; both open questions in 2.3 are answered with a named next decision.

1.3How one client is watched across many products

Monitoring in a multi-product setting, and the holistic client view.

“Holistic client view” means three things here.

  1. Thresholds apply to the client, not the account. Small amounts through five products: invisible to five per-product controls, obvious to one client-level control.
  2. The sequences cross products. Every typology worth catching is a sequence, not a state, and no single product line contains one.
  3. One risk state per client, readable by any control. It does not decide whether an alert fires (see 1.4).
Product streams feeding one client record, which the controls read Five product event streams converge into a single client record. Three typology controls read from that record and raise alerts. Beneath, six signals that exist only when products are joined at the client. PRODUCT EVENT STREAMS ONE RECORD CONTROLS READ FROM IT Broker Banking & settlement Credit & loans Wealth Securities in / out THE CLIENT RECORD Thresholds apply here, not to an account. One client. Every product. One sequence. accepts a new product's events without a schema change Layering control fires on the client · overnight Credit fraud control fires on a payment · real time Account takeover control fires on a payment · real time SIGNALS THAT ONLY EXIST ONCE THE PRODUCTS ARE JOINED · Round-trip speed: money in, traded, straight back out · Value traded against value deposited · Everything sold, then the payout account changed · Loan drawn down, then paid out within hours · A deposit, then securities transferred away entirely · Money in that does not match the declared profile
Figure 3  Products are inputs. The client is where controls are written. None of the six signals in the lower band can be seen by a control that reads one product. The record is also built to take a sixth and a seventh stream without being reshaped, which is what the middle column of Figure 2 depends on.

1.4Where fixed rules are right, and where a model earns its place

Deterministic controls against adaptive ones.

Evidence decides between them. Both answer one question: at what point does risk become unacceptable and an alert has to fire.

Release 1 is entirely deterministic. Anything real-time that restricts a customer has to be a fixed rule: a false-positive cost measurable in advance, and a policy statement defensible to a regulator. A model has neither until it is trained on outcomes that do not exist yet.

The three deterministic controls, the performance floor, and what becomes adaptive later A table of three controls with their latency, the entity each fires on, the stage each launches at and its default restriction. Below it, the rule that every control holds a minimum precision. Below that, the label gate and the three adaptive steps that follow it. CONTROLLATENCY FIRES ONLAUNCHES AT DEFAULT RESTRICTION First-party layering money laundering Overnight batch is enough The client a pattern, not a payment Stage 3 alerts, no restriction None by configuration, not by design Credit and loan fraud first-party fraud Real time the loss happens now The outgoing payment the drawdown is only an input Stage 4 first to restrict Block the payment, freeze top of the queue Account takeover third-party fraud Real time needs device and IP The outgoing transaction event level, or we are a day late Stage 4 gated on A3 Freeze the account severity is extreme EVERY CONTROL LAUNCHES WITH A MINIMUM PRECISION IT HAS TO HOLD — COMPLIANCE SETS IT Below the minimum → the control moves back down a stage, is reworked, and re-enters at back-test. Rework does not recover it → the numbers go to Compliance with a proposal to remove it, and a statement of how the risk stays covered. GATE — ENOUGH CORRECTLY TYPED OUTCOME LABELS, AND DATA IN A USABLE SHAPE 1  Baselines €10,000 is enormous for one client and routine for another. This replaces the arbitrary number. 2  Lowering risk Rules add risk and cannot subtract it. “This fired, but she always does this.” The largest precision lever there is. 3  Triage and ranking Weighs every signal at once and orders the queue. The safe first place to put a model — nothing it decides reaches a customer.
Figure 4  Each adaptive step needs more label volume than the one before, which is why the label loop in Figure 7 is the first thing built and not the last. Stage numbers refer to the four-stage rollout in Figure 8.

1.5Making an alert explain itself

Explainability, for the people who have to act on it.

An alert has to say what it was measured against. At launch that is a risk pattern, not the client's own history — there is none worth comparing to until the record has run.

Every alert shows four things: what fired, the values, the threshold crossed, and how often this control is right. The last one calibrates the analyst — three in ten confirmed beats guessing.

ALERT 48213
Client C-90412  ·  raised 09:14 CET
Rule LAY-03  version 2.4
Financial crime
money laundering
first-party layering
What fired
Outgoing transfer of €100,000 to an account in Switzerland
This control treats transfers above €50,000 to Switzerland as risky.
This client sent €100,000.
31%
of alerts from this control were confirmed
true positives over the last 90 days
The sequence behind it
Day 1  €90,000 in from reference account A
Day 2  ETF bought, then sold — net −€500
Day 3  Reference account B added, then set as primary
Day 4  €88,000 out to reference account B
True positive False positive Send a request for information Escalate
These four buttons are the label loop. Nothing downstream in Figure 4 or Figure 7 exists without them.
Figure 5  One alert, as the analyst opens it — the same case as Figure 1. The rule version is pinned to the alert, which is what lets precision be measured over time and what an auditor needs in order to see which policy this control implements.

Three audiences, and only two of them are in scope

  1. Analyst — the card above.
  2. Auditor — the policy each control implements, and the version pinned to each alert. Without it, precision cannot be tracked across a change.
  3. Customer — an expectation and a timeline, never the reason. Telling them is tipping off, applied most strictly in Germany.

Rule for the first model: a model may never be the only reason a customer experiences something, unless its reasons can be rendered. Gradient boosting with per-feature attribution gives that. Ranking a queue changes nothing a customer feels — the safe place to start.

1.6Data, entities and the two dependencies that could break this

Data requirements, the entity model, and technical dependencies.

One product commitment: a stated latency per source, real time by default. Physical layout is an engineering decision.

Batch is a cost concession — taken where real-time retrieval is genuinely too expensive, and named where taken. A feed built for batch has to be rebuilt when tomorrow's control is real-time, and that rebuild is the integration cost Figure 2 promises to avoid.

Sources, entities, latency and controls, with the release-1 critical path and two dependencies marked Nine data sources feed the entities on the client record. Six sources are on the release-1 critical path. Entities are held to three latency commitments, which feed the three controls. Two sources carry assumptions that could break the plan: credit drawdown data and device and IP data. SOURCESENTITIES ON THE CLIENT RECORD LATENCY IT IS HELD TOCONTROLS Banking transactions Securities transfer in and out Credit application, drawdown A2 Reference account events Credential and profile changes Device, IP and session A3 Broker orders and trades Wealth portfolios Crypto, derivatives, plans, EIX on the release-1 critical pathabsorbed, nothing written against it yet client declared income, occupation, expected volume, wealth source accountper product line reference account own name, verified, primary transactiondirection, amount, currency, counterparty, time trade or order loanapplication, drawdown device and session alert·case events carry as much signal REAL TIME transactions, device and IP, credential changes, loan drawdown OVERNIGHT trades, reference events AT ONBOARDING a later release Credit and loan fraud fires on the outgoing payment Account takeover fires on the outgoing transaction First-party layering fires on the client Identity fraud out of scope for release 1 THE TWO PLACES THIS PLAN BREAKS FIRST A2 Credit application and drawdown data Assumed available. If it is not, credit and loan fraud leaves release 1 and waits for that integration — which also removes the first control that would restrict anyone. A3 Device, IP and session data Probably exists somewhere and has almost certainly never been a financial crime input at a brokerage. Both real-time controls need it. Checked in week one, not week ten.
Figure 6  Every typology worth catching is a sequence, not a state, so non-transaction events carry as much signal as payments: a reference account added and made primary, a password or phone number changed, a login from a new device, a loan drawn down, a portfolio sold in one go. Putting the two risky assumptions inside the architecture picture is deliberate — in a separate table nobody reads them against the thing they threaten.

One open question: does one person reliably map to one client record? Joint accounts, powers of attorney and children's accounts are where more than one person legitimately touches one record — the shape this has to get right without treating a spouse as an intruder.

1.7What happens to an alert, and why the outcome is the asset

How alerts feed into case handling and investigation.

De-risk first that a vendor can carry our label taxonomy, then decide build versus buy during discovery. The current state of case management is unknown, and assessing it is explicitly in scope for this MVP.

An investigator needs three things: the client's activity across every product, the device and login history with origin, and the full transaction history. Everything else is workflow.

flowchart TB
  A(["Created"]) --> B(["Ready for review"])
  subgraph CASE["ONE CASE — the first alert opens it, later alerts join it while it is open"]
    direction TB
    B --> C(["Waiting for customer"])
    C --> B
    B --> D(["Escalated — moves to the
second-line queue"]) B --> TP(["True positive"]) B --> FP(["False positive"]) TP --> NC(["Needs a check
by a line manager"]) FP --> NC end TP --> LS[["OUR INTERNAL FINANCIAL CRIME LABELS
typology mandatory on every outcome"]] FP --> LS NC --> LS LS --> PR(["Precision, per alert"]) LS --> AD(["Every adaptive step
in Figure 4"]) PR --> Q{"At or above the
minimum precision?"} Q -->|yes| K(["Keep it running"]) Q -->|no| R(["Back down a stage, rework,
re-enter at back-test"]) classDef st fill:#FFFFFF,stroke:#106E9E,stroke-width:1.5px,color:#0A4E71 classDef term fill:#E7F0F6,stroke:#106E9E,stroke-width:1.5px,color:#0A4E71 classDef store fill:#131820,stroke:#131820,color:#FFFFFF classDef risk fill:#F7E9E7,stroke:#A93226,stroke-width:1.5px,color:#A93226 classDef ad fill:#EFEAF7,stroke:#6C4AB0,stroke-width:1.5px,color:#6C4AB0 class A,B,C,D st class TP,FP,NC term class LS store class Q,R risk class PR,K st class AD ad
Figure 7  The typology is not a status. It is mandatory metadata on every alert, cascading through three levels — financial crime, then money laundering or first-party fraud or third-party fraud, then the specific typology. That is what makes an outcome trainable rather than merely closed. Escalation is a state in its own right and moves the alert to the second-line queue, which works the same set of states.

The case is the structure around the alerts, and it has a routing consequence

An alert opens a case. Later alerts on that client join it while it is open. Routing consequence: a money-laundering alert in waiting for customer, joined by a fraud alert that freezes the account, moves to the fraud team under a tighter service level — a blocked customer is a different clock.

The dependency I would escalate on day one

Discovery gates the case-management decision, and assessing it is in scope here. Discovery establishes what the existing tooling does, what the gap to the capability above is, and whether the market closes it faster than we can. The label loop cannot close until that decision is made, so it is made in M0, not deferred.

All alert and case data stays in our own backend, whichever way the assessment lands. A vendor owning the outcome history means swapping the vendor loses the asset the models train on.

1.8What would make this fail

Risks, edge cases and design challenges.

RiskHow it is handled, and how we would see it coming
Alert volume is more than operations can absorb at go-live. A backlog is itself a regulatory finding.Back-test every control on production-scale data before anything is switched on, and watch alerts per analyst per day against the agreed ceiling through shadow running. If the number breaches, typology scope is cut rather than precision degraded.
Precision does not reach the level Compliance set.Same instrument, before go-live. After go-live it is the minimum precision on each control, checked continuously, with rework or removal as the response.
Device and IP data does not exist in a usable form. Both real-time controls depend on it.The most likely technical surprise on this plan, so it is checked in week one of discovery. If the data has to be generated, account takeover leaves release 1 and the work starts immediately with the team that owns client-side telemetry.
The vendor's data model and typology structure do not fit ours, and we find out after signing.This one is not watched, it is closed. Walk the vendor through exactly what we intend to build and confirm they support it before the contract. A small ask early and an impossible one late.
Labels are captured badly, and the whole adaptive roadmap dies quietly a year later.First and second line have to own the typology taxonomy rather than receive it, and check labels inside their own alert quality assurance — confirming a fraud alert is actually fraud and not merely financial crime. Label quality is tracked as a number on the control scorecard.
Case management stays undecided and blocks the label loop.Escalated on day one with a request for a named owner and a decision date, not for a solution.
The feed becomes the real ceiling and the vendor takes the blame for it.An engine only detects what it is sent. If the data layer is incomplete or late the vendor underperforms, and we risk churning vendors over our own fault. Any vendor benchmark has to control for how complete the feed was.

Edge cases the first release has to survive

  • Securities transferred out. The cash loop is closed, the value loop is not. Assets move to another broker and become cash elsewhere.
  • Dividends, coupons and corporate actions. Inflow that never came from a reference account. A control assuming all inflow is a deposit fires on every dividend.
  • Joint accounts, powers of attorney, children's accounts. Legitimate third-party control that looks exactly like account takeover.
  • Migration. Baselines from zero make every long-standing client look new, therefore unusual. Load the maximum history available.

Three challenges worth naming

  1. Risk appetite is Compliance's. Precision targets, alert ceiling and default restrictions are decided by them. Needs a standing forum — hence 2.4.
  2. Tipping-off caps every customer-facing improvement. The most useful thing to tell a customer is the thing we may not say.
  3. Benchmarking an internal build means running both. Unless one runs silently, alert volume doubles. 2.2 resolves this at no cost.
Task 2How the delivery is driven

2.1Rolling out four components that are not four build tracks

The rollout approach for each platform component.

ComponentOur stanceHow it rolls out
Transaction monitoring coreBuild the feed, rent the engineOne source at a time, each reconciled against its own source system before it joins the client record. This is the critical path and it starts in week one.
Fraud signals and controlsWe specify them, the vendor authors themOne typology at a time, each moving through the four stages below.
Specialist integrationsCase by case, against a fixed schemaEach integration is scoped on its own merits. The schema defines the minimum data points every new product must supply, so onboarding one is a mapping exercise rather than a redesign.
Case management and workflowAssess in discovery, then decideThe brief asks for a position on it, so assessing it is our work. Discovery sizes the gap and the decision lands in M0, before anything is built against it.

A control is not switched on. It moves through four stages

Between one stage and the next, only two things change: who sees the alert, and what happens to the client.

  1. Back-test — run against our own history, inside the vendor's rule builder. Only we see it, and the client experiences nothing.
  2. Shadow — run against live traffic. Still only we see it.
  3. Alerts reach the analyst queue — worked and closed with a typed outcome. Nothing happens to the client except what an analyst decides.
  4. A restriction is applied when the alert fires, before a human looks. The only stage where the system acts on a customer by itself.

Restriction = account frozen, outgoing blocked, incoming blocked, or a request for information. Same unit as the harm measure in 1.1.

Stages 1 and 2 are a procurement requirement. A rule builder that cannot replay our history, or run a rule silently against live traffic, fails the evaluation. Into the week-1 walkthrough.

Back-test says what a rule would have done to yesterday. Shadow says what it does to today, and produces the live volume the queue is planned against.

Stages run backwards too. A misbehaving control moves down a stage, not off: harm removed, signal and volume kept. Stage 3 before stage 4 because shipping at stage 3 is reversible.

Six milestones over twenty-six weeks, with each control's rollout stages and the two gated dependencies A Gantt chart with four component lanes and one dependency lane across twenty-six weeks. The monitoring core is built feed by feed through week ten. Three controls each move through back-test, shadow, alerting, and restricting. A go-live line sits at week fourteen. Case management gates milestone two and device and IP data gates milestone four. M0M1 M2M3 M4M5 DiscoveryThe record The loopFirst restriction Real-time completeDecide wk 0–2wk 2–10 wk 10–14wk 14–18 wk 18–24wk 24–26 Transaction monitoring core build the feed, rent the engine feed by feed, each reconciled the next product lines Fraud signals and controls we specify them, the vendor authors them control specification and the vendor walkthrough First-party layering fires on the client  ·  overnight  ·  no restriction at launch 1 2 3 Credit and loan fraud fires on the outgoing payment  ·  real time  ·  blocks it and freezes the account 1 2 3 4 Account takeover fires on the outgoing transaction  ·  real time  ·  freezes the account 1 2 3 4 Specialist integrations not a build track the data half becomes a requirement on the core   ·   the product-specific views join the case-management assessment Case management assess, do not build decision owned by another team Device and IP data outside this team does it exist in a usable form? if not, account takeover leaves release 1 gates M2 gates M4 GO-LIVE · WEEK 14 1  Back-testour own history — only we see it 2  Shadowlive traffic — still only we see it 3  Alerts reach the queuenothing automatic reaches the client 4  A restriction is appliedwhen it fires, before a human looks minimum precision, measured from here on — below it the control moves down a stage gated on something this team does not control
Figure 8  Build order follows volume. The order in which controls start restricting people follows who loses the money. Layering is built first because it is assumed to carry the most volume, and volume is what consumes analyst capacity. Credit fraud restricts first, because it is the only control where the moment it fires is the moment the loss happens. The two dashed risers are the honest part of this plan: both sit outside this team, and each is marked so that a slip is traced to its cause instead of quietly absorbed.

Why specialist integrations is not a third build track

  1. The data half → the monitoring core. The schema accepts any product's events without being reshaped, and every data point is built for real-time retrieval.
  2. The user-facing half → the case-management assessment. A crypto panel or a securities order view is needed only when a case calls for it. It joins the assessment list.

How the restriction on each alert is chosen

The logic is one distinction. Fraud: act on the money. Credit fraud and account takeover default to blocking the payment and freezing the account, because funds are leaving now. Money laundering: do not. Freezing on a laundering suspicion risks tipping off the customer, so the layering control alerts without restricting. Any single alert can override its default.

Those rules are jurisdictional, so this is configuration and not architecture. Scalable will not stay in one market.

The six milestones, and what would move each date

Every date is attached to a condition. When the condition moves, the date moves visibly. A plan with no dates is evasion.

 Finished whenEndsHow firmWhat would move it
M0The seven assumptions are answered and the vendor walkthrough is done. Above all: does device and IP data exist, is credit drawdown data available, and what is the real gap in the existing case management.Week 2FirmNothing. It is our own work, and it is mostly asking questions rather than building.
M1The three critical product lines and the event sources are on the client record, at the latency the contract requires. Proved by replaying a case financial crime already investigated, end to end.Week 10FirmThe size of the existing data gap, which M0 measures.
M2The layering control is alerting into the queue. Alerts are closed with typed outcomes, and first and second line have checked the labels. Alerts per analyst per day are inside the agreed ceiling.Week 14Firm once M0 decidesThe case-management decision taken in M0 — extend what exists, or buy.
M3The credit fraud control is restricting. The first automatic mitigation in production.Week 18Firm once M2 lands
M4Account takeover is restricting. Real-time coverage is complete.Week 24GatedWhether device and IP data exists. If it has to be built, account takeover leaves release 1.
M5Whatever discovery could not answer: the real incidence mix, and whether a model can be trained at all. Each with a named next decision.Week 26EstimateWhether enough alerts have closed to measure precision per control.

Week 14, first alerts in production · week 18, first automatic restriction · week 24, release-1 scope complete · week 26, both open questions answered. Indicative, for six engineers, estimated without seeing the existing system — which is what M0 corrects.

Why the last milestone is small, and might be smaller

Measurement is not at the end. Each control is measured from its first closed outcome. Layering has ten weeks of evidence by week 24.

One week-2 question settles both: does the existing case history carry usable typology labels? Yes → incidence mix queryable at once, build order changes on the spot, feasibility check has a corpus. No → both fall to week 26.

Caveat: even answered, incidence is a prior. Those cases came from a different system firing different alerts, so anything the old queries missed carries no outcome. Enough to set a build order. Not enough to close the question.

2.2Where the ambition goes, and what is bought instead

Trade-offs between speed, control strength and technical ambition.

Spend ambition where reversing the decision would be expensive. Buy speed everywhere reversing it is cheap.

Put the ambition where the decision is irreversible. Buy, defer or keep simple everything that is not. A rule set is rewritten, a vendor is swapped — but a data layer built to the wrong latency has to be rebuilt, and the rebuild takes the label history with it.

Built
Bought, deferred, or kept deliberately simple
The one thing we build
The client-level data layer

Schema, entities, and a stated latency per source. Built wrong, it cannot be reversed — it has to be rebuilt, and the rebuild destroys the outcome history that every model depends on.

It is also the thing no vendor can sell us. No engine arrives knowing Scalable's Broker, Wealth, Credit and Banking activity joined at one client. Buying the engine does not buy the client view.

The monitoring engineRule builders are commoditised and several vendors are good.Reverse: swap it
Case managementAnother team owns it, generically, for settlement and operations too.Reverse: not ours to
Machine learningGated on labels that do not exist yet, not on appetite.Reverse: start it
Anything customer-facingNo front-end capacity on this team.Reverse: add people
Product-specific viewsOnly needed when a case calls for one.Reverse: build one
Seven product linesIn the record, nothing written against them.Reverse: write a control
Figure 9  Each card on the left names what would reverse the decision. That is the test being applied, and it is why the list is long and the right-hand side has one item on it.
What we give upWhat it buysWhat it costsWhat would reverse it
Building the engineTime to market, and something to benchmark an internal build againstOur controls are only as expressive as the vendor's rule builder, and their data model may not fit oursThe walkthrough shows the builder cannot express our controls or replay our history
Precision in release 1The label loop starts turning immediatelyAnalysts work more false positives in the first quarter than they will laterVolume breaches the analyst ceiling in back-test — then we cut scope, not precision
Machine learning in release 1No dependency on labels we do not haveThe largest precision lever, lowering risk on clients who always behave this way, sits unused for a releaseThe feasibility check passes both preconditions
Building case managementSix engineers stay on the feedWe depend on another team's roadmap for the thing that gates M2The assessment shows the gap cannot be bridged and the market offers nothing better
A restriction on the layering controlNo automatic customer harm on the highest-volume typologyMoney laundering is detected but never interruptedCompliance sets an appetite that says otherwise, or we enter a market where it is standard

What actually limits how strong the controls can be at launch

Analyst headcount, not detection capability. Breach the alerts-per-analyst ceiling and you get a backlog, and a backlog is itself a regulatory finding.

If the back-test says the control set is too large, we cut typology scope rather than degrade precision. Compliance has the final word.

Two trades I would refuse

  1. Shipping without typed outcomes. Kills the adaptive roadmap quietly.
  2. Going live without a back-test on production data. Turns a measurable risk into a surprise for operations.

Benchmarking an internal build costs nothing extra

The internal build runs at stage 2 while the vendor's control runs at stage 3 or 4. Same live traffic, no extra alerts in the queue, so the benchmark costs nothing extra.

Ambition is also limited by who is on the team

Six data engineers, no data scientists, no front-end engineers — which is why nothing customer-facing and no machine learning in release 1.

Do not ask for data scientists before there is evidence there are labels to train on. Borrow existing capability, run the check cheaply, let the result make the case.

2.3What is proved before go-live, and what can only be measured after

Validation before launch, what follows after, and how the system improves.

Before go-live, prove everything the customer can feel and everything that cannot be undone. After go-live, measure everything that can only be measured live.

Precision cannot be known before go-live, and the estimate is optimistic

Precision cannot be known before go-live. A back-test scores against outcomes from a different system firing different alerts. Cases it would have caught but the old queries never surfaced carry no outcome, so they cannot count against it.

The back-test is a capacity instrument, not a quality one. Volume: reliable. Recall: partial, by replaying known filed reports. Precision: an estimate. Plan the analyst ceiling conservatively against it.

Before go-live — week 14
Everything a customer can feel, and everything that cannot be undone
  1. The feed is correctEvery source reconciled against its own source system.
  2. The record can reconstruct a real caseReplay a case already investigated, end to end. Checks joins, event capture and latency at once.
  3. Volume fits the queueBack-test against the agreed ceiling.
  4. The loop closesOne real alert: created, queued, worked, closed with a typed outcome, stored in our own backend.
  5. Restrictions can be liftedApply and reverse every mitigation. Without the reversal path, every false positive becomes an engineering ticket.
  6. The explanation rendersAn analyst works shadow alerts and can say why each fired.
After go-live
Ordered by what unblocks what
  1. Real precision per controlFrom closed cases. The first true number about this system.
  2. The real incidence mixThe build order is reprioritised against it.
  3. Threshold tuningThe first thing that needs changing, and the cheapest.
  4. The model feasibility verdictIf labels fall short, label capture becomes the next release, not the model.
  5. Account takeover to a restrictionOnce device and IP data allows it.
Later
  1. The case-management decision executedWhichever way it went.
  2. The next product lines onto the recordThe first real test of the claim that this platform is cheap to extend.

Item 5 on the left is the one most teams miss, and the only one purely for the customer.

Three things make the system better, and they are not the same thing

Tuning. Adjust, replace or propose removing a control on its numbers — from day one, with Compliance approving any removal.

Compounding. Labels accumulate, so client history replaces arbitrary thresholds — each step needing more labels than the last.

Coverage. Neither of the other two finds a risk nobody wrote a control for, so a regular threat-modelling pass asks what surface nothing watches yet.

Every alert this system produces makes the next one better — but only because its outcome is typed and stored in a backend we own.

2.4Two forums, one gate, and a rule that stops it growing

Operating cadence, forums and artefacts.

Stakeholders are steered on milestones M0–M5, with exit conditions agreed before each one starts. Status is which condition is unmet and who owns it — not velocity. The engineering team's own rhythm stays with its manager.

The timeline turns on things we do not fully control — whether device and IP data exists, and a vendor procurement cycle. A milestone with an honest exit condition survives a slipped dependency; a sprint commitment does not.

A milestone can be re-scoped, never quietly re-dated.

Whoever bears the consequence of a milestone being wrongly declared finished is the one who judges whether it is finished — not the product manager, who has every incentive to close it.

First line judges whether the record can reconstruct a real case. Operations judges whether the queue can absorb the volume. Compliance judges whether a control may restrict a customer automatically.

Financial crime control forum
Monthly · longer once a quarter
Product · first-line team leads · second line · Compliance
Where a control's numbers are looked at and something is decided about them: tune it, replace it, or take the evidence for removing it to Compliance.
Once a quarter it runs longThe coverage map and the appetite sheet: the alerts-per-analyst ceiling, precision minimums, default restrictions, and any change to the typology taxonomy.
Standing document
Control scorecard
Volume, precision, cycle time and label quality, per control
Milestone review
When a milestone closes
Leadership · engineering manager · product · second line
Not a status meeting. It exists to close one milestone against its exit conditions and to take the decision that the next one depends on.
Standing document
Milestone one-pager
Exit conditions, their status, and the decision being asked for
New-product gate
Triggered by a launch, not a date
Owned by this team · run against any product launch
The one piece of process worth creating rather than inheriting. Three questions:
  1. What movement of value does this product create?
  2. Which typology does it touch?
  3. Is that covered, or explicitly accepted?
Its output lands in
The coverage map
Where the claim that this platform scales stops being a property of the schema and becomes a process
The rule that keeps this from growing. Every forum has one standing document as its agenda — no slides. Every meeting ends with a decision recorded in it. If no decision is due, the meeting does not happen.
Figure 10  Everything not shown here is a working rhythm rather than a forum: the weekly with the engineering team, and the vendor sync during integration. The one meeting where both lines of defence must be in the room is the control forum. Second line instructs first line and owns the policy; first line knows what is actually being seen. A control tuned with only one of them present gets rejected by the other, and the product manager's position on this platform is beside that relationship rather than inside either side of it.

Six artefacts, and not all of them need a meeting

ArtefactWhat it is for
Coverage mapProduct surface against typology, with coverage status. Makes “no gaps” auditable, and records what is covered by product design rather than by detection — which is how the MVP justifies what it leaves out.
Control specificationOne per control: typology label, the entity it fires on, latency, launch stage, default restriction, the policy it implements, and its version. Reference material rather than an agenda, and it doubles as the artefact an auditor needs.
Control scorecardThe tuning input, and the evidence behind any request to Compliance to remove a control.
Milestone one-pagerExit conditions and their status.
Assumption registerThe seven assumptions, each with an owner and a test. This is the M0 work plan — closing an assumption is a unit of progress, which is what makes discovery a milestone rather than a preliminary.
Decision logWhat was decided, by whom, and what would reverse it. Especially for decisions Compliance owns, so that a later change is a re-decision rather than an argument.

Label quality is produced by stakeholders, not supplied by the system. First and second line check typology labels in their own alert QA, which holds only if they own the taxonomy.

2.5What I would escalate in week one

Risks, open questions and decisions to escalate early.

Escalation is for what this team cannot resolve alone. Everything else is a task.

It escalates only if it cannot be resolved inside the team and gets more expensive the longer it stays open — and it goes up with a recommendation and a date, never as an open problem.

 WhatWhoThe askWhat a late answer costs
1Case managementLeadership, plus whoever runs the current toolingAccess and a decision slot in M0. We bring the recommendation.A conversation in week one; a milestone in month four, because the first control's alerts have nowhere to go.
2Does device and IP data exist in a form financial crime can use?Data engineering, and whoever owns client-side telemetryAnswer it. If the answer is no, start the work to generate it now.There is a project attached to the answer, on someone else's lead time. Late, and account takeover leaves release 1.
3Is one person guaranteed to be one client record?Onboarding and identityA yes or a no.Almost certainly fine. It earns a line because joint and children's accounts are exactly the shape client-level monitoring has to get right.

The decisions that are Compliance's, not ours

Risk appetite produces three numbers: the alerts-per-day ceiling, the minimum precision per control, and the default restriction per alert.

We arrive with a position on all three. The decision is theirs; the recommendation is ours.

The one risk we close before signing

Get agreement that the vendor walkthrough happens before the contract. A small ask early and an impossible one late.

Two positions I would simply state

  • History. Load the maximum available — baselines from nothing make every long-standing client look new.
  • The model feasibility check needs no new headcount.

No new assumptions

Task 2 adds no assumptions. Scope and delivery stand on the same seven, listed in the appendix.

Next

Week 1, two calls. Ask data engineering whether device and IP data exists in a form financial crime can use. Ask the case-management team for a named owner and a decision date. Everything in M0 runs behind those two answers.

Prepared for Scalable Capital · 1 September 2026
Figures 1 to 10 are original to this document. Every number in them is illustrative unless it is named as an assumption.
Assumptions A1 to A7 are listed in full in section 00, each with the test that would prove it wrong.

00Assumptions and open questions

Seven assumptions. Each has a falsifying test and a date.

Six of the seven are answered by week 2, by asking questions before anything is built.

 AssumptionWhat would prove it wrongAnswered
A1Layering carries the most volume, credit fraud the most direct loss, account takeover the highest severity per case at lower incidence.A query of the existing case history, if that history carries usable typology labels. If it does not, the first production measurement window.Week 2, or week 26
A2Credit application and drawdown data is available to financial crime.Data discovery. If it is not available, credit fraud waits for that integration and leaves release 1.Week 2
A3Device and IP data is not currently an input to financial crime detection.Data discovery. Both real-time controls depend on it.Week 2
A4The existing set of customer restrictions can be graded by severity: full freeze, then outgoing blocked, then incoming blocked, then a request for information only.A review of the current control set with operations.Week 2
A5This team has no front-end engineering capacity, so nothing customer-facing is in scope.Team composition confirmed at the start.Week 1
A6The existing in-house case management can be extended to the capability list in 1.7.The build-versus-buy assessment.Week 2
A7Financial crime has a right of access to every product's data, so a separate data source is never a reason to exclude a typology.Confirmation from second line and data governance.Week 2
Open questions I am not answering here

1. Can crypto be withdrawn to an external wallet? If yes, it only has to be monitorable like any other outbound value. 2. Current alert volume, financial crime headcount, current precision — every capacity number here is an assumption until these exist. 3. Portfolio-backed or unsecured loans? Changes the size of a loss, not the design.