One client, four products, no one watching the sequence
An MVP for transaction monitoring across Broker, Wealth, Credit and Banking, and the plan to deliver it in six months with six data engineers.
The first release is not a better rule engine. It is the client-level record of behaviour that every later control is written against, whether that control is a fixed rule or a trained model.
Today Scalable can answer “was this transaction unusual for this product?” It cannot answer “was this unusual for this client?”
00Assumptions and open questions
Stated up front, because every decision that follows rests on them. Each one has a test that would prove it wrong and a date by which it is answered.
Seven assumptions carry this proposal. Six of them are answered in the first two weeks by asking, not by building, which is why discovery is a milestone rather than a preliminary.
| Assumption | What would prove it wrong | Answered | |
|---|---|---|---|
| A1 | Layering carries the most volume, credit fraud the most direct loss, account takeover the highest severity per case at lower incidence. | A query of the existing case history, if that history carries usable typology labels. If it does not, the first production measurement window. | Week 2, or week 26 |
| A2 | Credit application and drawdown data is available to financial crime. | Data discovery. If it is not available, credit fraud waits for that integration and leaves release 1. | Week 2 |
| A3 | Device and IP data is not currently an input to financial crime detection. | Data discovery. Both real-time controls depend on it. | Week 2 |
| A4 | The existing set of customer restrictions can be graded by severity: full freeze, then outgoing blocked, then incoming blocked, then a request for information only. | A review of the current control set with operations. | Week 2 |
| A5 | This team has no front-end engineering capacity, so nothing customer-facing is in scope. | Team composition confirmed at the start. | Week 1 |
| A6 | The existing in-house case management can be extended to the capability list in 1.7. | The build-versus-buy assessment. | Week 2 |
| A7 | Financial crime has a right of access to every product's data, so a separate data source is never a reason to exclude a typology. | Confirmation from second line and data governance. | Week 2 |
Can crypto be withdrawn to an external wallet? If it can, the requirement is only that the withdrawal is monitorable like any other outbound movement of value. — What are the current alert volume, the financial crime operations headcount, and the current precision? Every capacity number in this document is an assumption until those three exist. — Are the loans portfolio-backed or unsecured? This does not change the design, but it changes what a realistic loss looks like.
1.1The problem this solves, and what good looks like
Required sub-point: the problem and the outcome the MVP is aiming at.
What forces this now is growth, not regulation
Every product a bank offers needs a risk assessment and has to be monitorable. That is a fixed requirement and it is not the interesting part of this problem.
The interesting part is that Scalable's product design has been doing most of the protective work, and it is running out. Money can only enter from an account in the client's own name at an EEA credit institution, only one of those accounts is primary at a time, and payouts go only to that primary account. Third-party deposits are rejected and returned. There are no business accounts. That closed loop removes whole categories of financial crime by construction — a mule network cannot use an account it cannot pay into.
Three things are ending that. The banking licence brought deposits and lending, so the bank now lends its own money and can lose it. The product line keeps widening, and each new product is another surface. And the client base is past a million, which means the current approach — SQL queries over data pipelines raising alerts — is at its ceiling for reasons of volume rather than quality.
The gap that matters
Figure 1 is the whole argument. Each product line sees a single ordinary event. Nobody is wrong. The pattern is only visible where the products are joined, and today they are not joined anywhere that a control can read.
So the first release is a record, not a rule engine. Products become inputs to one client-level view, and controls get written against that view.
What good looks like
Two measures lead, and they are the same mechanism seen from two sides. Better precision means fewer alerts for the same number of reports filed.
| Measure | Definition | Why this one |
|---|---|---|
| Precision | True positives over all alerts, measured per control and across the system. | The primary number. It is what every other improvement shows up in. |
| Third-party requests | Another institution contacts Scalable about a client it suspects of crime. | A proxy for what we miss. A true positive here is crime inside the bank that we did not catch. Deliberately caveated, and never improved by taking in less. |
| Authority follow-up | The financial intelligence unit forwards a report to the police, and an investigation follows. | The strongest external signal that a report was worth filing. |
| Cycle time per alert | Alert raised to alert closed. | Serves efficiency and the customer at once. See below. |
| Alerts per analyst per day | Volume against the agreed capacity ceiling. | The hard limit on how strong the control set can be. Section 2.2 turns on this number. |
| Alerts per active client | Total alerts divided by monthly active clients. | Has to fall over time. This is the actual proof that the platform scales rather than grows headcount. |
Two outcomes that matter and resist measurement
The platform has to get cheaper to extend. There is no clean metric for this. The testable form of the claim is that integrating the twelfth product costs less than the fourth did.
Customers who did nothing wrong should feel as little as possible. Two things drive that: the quality of the customer-facing experience, and how severe a restriction we apply. Only the second is this team's to control, because there is no front-end capacity here (A5). So the commitment is a graded set of restrictions rather than one blunt freeze, and the harm is measured as days under the most severe active restriction and money blocked over time. Restricting an empty account and restricting a funded one are not the same harm.
Compliance sets the risk appetite and that is never traded against. It is a boundary on the work rather than a goal of it, which is why it appears here as one sentence and again in 2.5 as a set of decisions that belong to someone else.
1.2What the MVP contains, and what it deliberately leaves out
Required sub-point: scope, and the reasoning behind what is excluded.
Scope has three states here, not two, and the middle one carries most of the judgment. A product can be present in the record without anything watching it. That is what keeps the platform cheap to extend: adding coverage later costs a control, not a change to the data model.
- First-party layeringAbsorbs structuring and inconsistent source of funds — the same shape
- Credit and loan fraudThe only place Scalable loses its own money
- Account takeoverHighest severity per case
- Bank and settlement account
- Credit and loans
- Securities transfer in and outA closed cash loop is not a closed value loop
- The label loopAlert to case to typed outcome, stored in our own backend
- The explanation on every alert
- Vendor integrationAlerts, scores and outcomes written back to us
- Crypto
- Derivatives
- Savings plans
- Wealth and managed portfoliosIncluding the white-label partnerships. Under AML rules a partner arrangement means two firms monitor, not one, so Scalable monitors regardless.
- Private equity
- European Investor Exchange
- Every product offered needs a risk assessment, and products change into shapes that need watching later. The cost of being ready is one ingestion, paid once.
- Sanctions and PEP screeningList matching against a name, checked at onboarding and rescreened periodically. Not behavioural, and a separate system.
- Market abuse and insider dealingA different regulation, driven by order book and news data, and watched at venue level.
- Third-party mule networksRemoved by the closed loop today. The data model must not preclude them, because a full banking proposition removes that protection.
- Identity fraud at onboardingCaught at onboarding. This brief is about product usage, which is everything after it.
- Investment scam, victim sideMoney can only arrive from the client's own account, so Scalable cannot be the collection point. The forced-liquidation remainder produces the same signal as account takeover.
- Anything customer-facingNo front-end capacity on this team (A5).
Two pieces of work inside the MVP produce a decision rather than a system: the build-versus-buy assessment on case management, and the time-boxed check on whether a model can be trained at all. Both are sized as analysis, and both are on the critical path because what they conclude sets the next release.
What release 1 has to have produced to be finished
Three typologies covered by controls that can explain themselves, running at the right latency. Every closed case carrying a typed outcome label. And the two open questions in 2.3 answered, each with a named next decision.
1.3How one client is watched across many products
Required sub-point: monitoring in a multi-product setting, and the holistic client view.
“Holistic client view” is easy to claim and worth nothing as a phrase. It means three specific things here.
- Thresholds apply to the client, not the account. A client moving small amounts through five products is invisible to five per-product controls and obvious to one client-level control.
- The sequences cross products. Every typology worth catching is a sequence rather than a state, and no single product line contains one.
- There is one risk state per client that any control can read — though as 1.4 explains, it does not decide whether an alert fires.
1.4Where fixed rules are right, and where a model earns its place
Required sub-point: deterministic controls against adaptive ones.
These are not competing philosophies. Both answer one question — at what precise point does risk become unacceptable and an alert has to fire — and the choice between them is decided by evidence rather than preference.
Release 1 is entirely deterministic, for one reason. Anything that runs in real time and restricts a customer has to be a fixed rule, because a fixed rule has a false-positive cost you can measure in advance and a policy statement you can defend to a regulator. A model has neither until it has been trained on outcomes that do not exist yet.
One thing a client risk rating must not do
An alert fires because a signal is risky, never because the client is marked high risk. A holistic rating mixes too many things together, and using it as a trigger means the same client keeps alerting for being the client they already were.
A model that grades risk over the client lifecycle is a different matter. It is an event source like any other, and a sharp jump in its score that is empirically tied to fraud is a legitimate reason to raise an alert.
1.5Making an alert explain itself
Required sub-point: explainability, for the people who have to act on it.
“Unusual” means nothing on its own. It only means something once you can see what it was measured against, and the honest answer at launch is that it was measured against a risk pattern rather than against the client's own history. There is no history worth comparing to until the record has been running.
So an alert shows what fired, the actual values, the threshold it crossed, and how often alerts from that control turn out to be real. The last one matters more than it looks: an analyst who knows that three in ten of these are confirmed arrives at the case calibrated instead of guessing.
Rule LAY-03 version 2.4
› money laundering
› first-party layering
This client sent €100,000.
true positives over the last 90 days
Day 2 ETF bought, then sold — net −€500
Day 3 Reference account B added, then set as primary
Day 4 €88,000 out to reference account B
Three audiences, and only two of them are in scope
The analyst needs what is on the card above. The auditor needs the policy each control implements and the version pinned to each alert, without which precision cannot be tracked across a change. The customer needs an expectation and a timeline and never the reason, because telling them the reason is the offence of tipping off, and Germany applies that most strictly of the markets Scalable operates in.
The customer side is out of scope for this release only because there is no front-end capacity on this team (A5). It is a resourcing limit and not a design position.
When a model does arrive, gradient boosting with per-feature attribution gives an alert-level explanation of what drove the score. That leads to one rule worth stating now, because it decides where the first model is allowed to go: a model may never be the only reason a customer experiences something, unless its reasons can be rendered. Ranking a queue changes nothing a customer can feel, which is why it is the safe place to start.
1.6Data, entities and the two dependencies that could break this
Required sub-point: data requirements, the entity model, and technical dependencies.
How the data is physically laid out is an engineering decision and I am not going to make it here. The product commitment is a different one: a stated latency for every source, and the default is real time.
Batch is a cost concession. It is taken where real-time retrieval is genuinely too expensive, and it is named where it is taken. The reason for that default is not that faster is better. It is that a feed built for batch because today's control runs overnight has to be rebuilt when tomorrow's control does not — and that rebuild is exactly the integration cost the middle column of Figure 2 promises to avoid.
One entity question sits underneath all of this and is worth asking out loud even though the answer is probably fine. Client-level monitoring is a fiction unless one person reliably maps to one client record. Onboarding and identity checks should guarantee that. Joint accounts, powers of attorney and children's accounts are the cases where more than one person legitimately touches one record, and those are exactly the shape that client-level monitoring has to handle without treating a spouse as an intruder.
1.7What happens to an alert, and why the outcome is the asset
Required sub-point: how alerts feed into case handling and investigation.
I do not know the current state of case management, and that unknown is one of the two things shaping this MVP. So the position here is to specify the capability that is needed, then run a build-versus-buy assessment against it rather than assume either answer.
The capability is not elaborate. An investigator needs a consolidated view of what the client does across every Scalable product, the device and login history including where the logins came from, and the full transaction history. Everything else is workflow.
flowchart TB
A(["Created"]) --> B(["Ready for review"])
subgraph CASE["ONE CASE — the first alert opens it, later alerts join it while it is open"]
direction TB
B --> C(["Waiting for customer"])
C --> B
B --> D(["Escalated — moves to the
second-line queue"])
B --> TP(["True positive"])
B --> FP(["False positive"])
TP --> NC(["Needs a check
by a line manager"])
FP --> NC
end
TP --> LS[["OUR OWN LABEL STORE
typology mandatory on every outcome"]]
FP --> LS
NC --> LS
LS --> PR(["Precision, per control"])
LS --> AD(["Every adaptive step
in Figure 4"])
PR --> Q{"At or above the
minimum precision?"}
Q -->|yes| K(["Keep it running"])
Q -->|no| R(["Back down a stage, rework,
re-enter at back-test"])
classDef st fill:#FFFFFF,stroke:#106E9E,stroke-width:1.5px,color:#0A4E71
classDef term fill:#E7F0F6,stroke:#106E9E,stroke-width:1.5px,color:#0A4E71
classDef store fill:#131820,stroke:#131820,color:#FFFFFF
classDef risk fill:#F7E9E7,stroke:#A93226,stroke-width:1.5px,color:#A93226
classDef ad fill:#EFEAF7,stroke:#6C4AB0,stroke-width:1.5px,color:#6C4AB0
class A,B,C,D st
class TP,FP,NC term
class LS store
class Q,R risk
class PR,K st
class AD ad
The case is the structure around the alerts, and it has a routing consequence
An alert opens a case. Further alerts on that client join it while it stays open. That is straightforward until two alerts of different kinds land in the same case: a money-laundering alert sitting in waiting for customer is joined by a fraud alert that freezes the account. The case now has to move to the fraud team under a tighter service level, because the customer is currently blocked and the clock on that is a different clock. The case also records what the system did automatically, what the customer did in response, and every restriction applied.
Case management is owned by a separate team, which also runs it for securities settlement and other operational processes, and it is on their roadmap rather than ours. The label loop cannot close without it, and everything adaptive depends on the label loop. In week one this costs a conversation. In month four it costs a milestone, because the first control's alerts have nowhere to go.
Whichever way the assessment lands, one thing does not move: all alert and case data stays in our own backend. If a vendor owns the outcome history, swapping the vendor means losing the asset the models are trained on, and the platform stops being ours.
1.8What would make this fail
Required sub-point: risks, edge cases and design challenges.
| Risk | How it is handled, and how we would see it coming |
|---|---|
| Alert volume is more than operations can absorb at go-live. A backlog is itself a regulatory finding. | Back-test every control on production-scale data before anything is switched on, and watch alerts per analyst per day against the agreed ceiling through shadow running. If the number breaches, typology scope is cut rather than precision degraded. |
| Precision does not reach the level Compliance set. | Same instrument, before go-live. After go-live it is the minimum precision on each control, checked continuously, with rework or removal as the response. |
| Device and IP data does not exist in a usable form. Both real-time controls depend on it. | The most likely technical surprise on this plan, so it is checked in week one of discovery. If the data has to be generated, account takeover leaves release 1 and the work starts immediately with the team that owns client-side telemetry. |
| The vendor's data model and typology structure do not fit ours, and we find out after signing. | This one is not watched, it is closed. Walk the vendor through exactly what we intend to build and confirm they support it before the contract. A small ask early and an impossible one late. |
| Labels are captured badly, and the whole adaptive roadmap dies quietly a year later. | First and second line have to own the typology taxonomy rather than receive it, and check labels inside their own alert quality assurance — confirming a fraud alert is actually fraud and not merely financial crime. Label quality is tracked as a number on the control scorecard. |
| Case management stays undecided and blocks the label loop. | Escalated on day one with a request for a named owner and a decision date, not for a solution. |
| The feed becomes the real ceiling and the vendor takes the blame for it. | An engine only detects what it is sent. If the data layer is incomplete or late the vendor underperforms, and we risk churning vendors over our own fault. Any vendor benchmark has to control for how complete the feed was. |
Edge cases the first release has to survive
- Securities transferred out. The cash loop is closed. The value loop is not. Assets can be moved to another broker and turned into cash somewhere else entirely, which is why securities transfer is one of the three critical surfaces rather than a later product.
- Dividends, coupons and corporate actions. Money arriving that never came from a reference account. A control that assumes all inflow is a deposit fires on every dividend payment.
- Joint accounts, powers of attorney, and accounts held for children. Legitimate third-party control that looks exactly like account takeover.
- Migration. Clients and history that predate the new record. If baselines start from zero, every long-standing client looks new and therefore unusual. We load the maximum history available, and the only real limit is data already deleted under retention rules.
Three challenges worth naming rather than hiding
- Risk appetite belongs to Compliance, not to this team. Precision targets, the alert ceiling and default restrictions are set jointly and decided by them. That needs a standing forum, which is why 2.4 exists.
- Tipping-off caps every customer-facing improvement. The customer experience outcome has to be bounded honestly rather than over-promised, because the most useful thing to tell a customer is the thing we are not allowed to say.
- Benchmarking an internal build against the vendor means running both. Unless one of them runs silently, alert volume doubles. Section 2.2 resolves this at no cost.
2.1Rolling out four components that are not four build tracks
Required sub-point: the rollout approach for each platform component.
| Component | Our stance | How it rolls out |
|---|---|---|
| Transaction monitoring core | Build the feed, rent the engine | One source at a time, each reconciled against its own source system before it joins the client record. This is the critical path and it starts in week one. |
| Fraud signals and controls | We specify them, the vendor authors them | One typology at a time, each moving through the four stages below. |
| Specialist integrations | Not a third build track | It splits, and each half becomes a requirement on a component we already own. |
| Case management and workflow | Assess, do not build | The decision comes before the work. Escalated on day one, because another team's roadmap gates the second milestone and that milestone gates everything after it. |
A control is not switched on. It moves through four stages
Between one stage and the next, only two things change: who sees the alert, and what happens to the client.
- Back-test — run against our own history, inside the vendor's rule builder. Only we see it, and the client experiences nothing.
- Shadow — run against live traffic. Still only we see it.
- Alerts reach the analyst queue — worked and closed with a typed outcome. Nothing happens to the client except what an analyst decides.
- A restriction is applied when the alert fires, before a human has looked at it. This is the only stage where the system acts on a customer by itself.
A restriction here means the account frozen, outgoing payments blocked, incoming payments blocked, or a request for information sent. The word is chosen to match the customer harm measure in 1.1 — days under the most severe active restriction — so the rollout and the harm are counted in the same unit.
The first two stages are a procurement requirement, not just a way of working. A rule builder that cannot replay our own history, and cannot run a rule silently against live traffic, fails the vendor evaluation whatever its typology library looks like. That goes into the walkthrough in week one.
Shadow still earns its place after a back-test, because they answer different questions. A back-test says what a rule would have done to yesterday. It does not say what it does to today. Shadow produces no labels — nobody works a shadow alert — but it produces live volume, and live volume is the number the analyst queue is planned against.
The stages also run backwards. A control misbehaving in production moves down a stage rather than being switched off. Moving down removes the harm and keeps the signal and the volume data. That is also why stage 3 comes before stage 4: shipping at stage 3 is a reversible act.
Why specialist integrations is not a third build track
It splits in two, and each half is a requirement on something we already own. The data half belongs to the monitoring core: we have to be able to watch any product, so the schema accepts any product's events without being reshaped, and every data point is built for real-time retrieval by default. The user-facing half belongs to the case-management assessment: a chain-analytics panel for crypto or an order view for securities is only needed when a case calls for it, and what is possible depends entirely on which case-management system we end up with. So it is not scoped as work. It is added to the list the build-versus-buy assessment is run against.
They named four components. Folding one into the other two is a judgment call, and it reads as discipline when it is argued and as an omission when it is not.
How the restriction on each alert is chosen
The typology carries the default, and for the two fraud typologies that default is always set: block the outgoing payment and freeze the account. Those are the critical cases, and a critical case should not depend on somebody remembering to configure it. Any individual alert can override the default. The typology sets what happens unless somebody decides otherwise; it never sets what is possible.
That distinction matters because rules on tipping off are jurisdictional. Freezing on a money-laundering suspicion is possible in the UK and is being piloted in France, and Scalable will not stay in one market. The layering control launches with no restriction because that is what we configure here today, not because the platform forbids it. Encoding a German rule as a property of the system is exactly the decision that makes the twelfth product expensive.
The six milestones, and what would move each date
Steering on milestones does not mean refusing to give dates. It means each date is attached to a condition, so that when the condition moves, the date moves visibly instead of quietly. A plan with no dates is not steering. It is evasion, and it is the first thing anyone will ask about.
| Finished when | Ends | How firm | What would move it | |
|---|---|---|---|---|
| M0 | The seven assumptions are answered and the vendor walkthrough is done. Above all: does device and IP data exist, is credit drawdown data available, and what is the real gap in the existing case management. | Week 2 | Firm | Nothing. It is our own work, and it is mostly asking questions rather than building. |
| M1 | The three critical product lines and the event sources are on the client record, at the latency the contract requires. Proved by replaying a case financial crime already investigated, end to end. | Week 10 | Firm | The size of the existing data gap, which M0 measures. |
| M2 | The layering control is alerting into the queue. Alerts are closed with typed outcomes, and first and second line have checked the labels. Alerts per analyst per day are inside the agreed ceiling. | Week 14 | Gated | Case management. The largest schedule risk on this plan, and it sits with another team. |
| M3 | The credit fraud control is restricting. The first automatic mitigation in production. | Week 18 | Firm once M2 lands | — |
| M4 | Account takeover is restricting. Real-time coverage is complete. | Week 24 | Gated | Whether device and IP data exists. If it has to be built, account takeover leaves release 1. |
| M5 | Whatever discovery could not answer: the real incidence mix, and whether a model can be trained at all. Each with a named next decision. | Week 26 | Estimate | Whether enough alerts have closed to measure precision per control. |
Indicative, for six engineers, and an outside estimate made without having seen the existing system — which is precisely what M0 exists to correct. The four dates a stakeholder actually wants are week 14, first alerts in production · week 18, first automatic restriction · week 24, release-1 scope complete · week 26, both open questions answered.
Measurement does not happen at the end. Each control is measured against its minimum precision from its first closed outcome, which for layering means ten weeks of evidence by the time week 24 arrives.
Two questions do need all three controls to have run, and they are both put to the existing case history first, in week two. Does that history carry usable typology labels? If it does, the real incidence mix can be queried immediately and the build order changes on the spot, and the model-feasibility check has a corpus to run against. If everything is bunched into one undifferentiated category, neither question can be answered until the label loop has been turning, and both fall to week 26.
One caveat I would state rather than hide: even when discovery answers the incidence question, it answers it as a prior, not as the final number. Those historical cases came from a different system firing different alerts, so anything the old queries never surfaced carries no outcome at all. That is enough to set a build order on. It is not enough to close the question.
2.2Where the ambition goes, and what is bought instead
Required sub-point: trade-offs between speed, control strength and technical ambition.
A rule set is cheap to reverse: rewrite it. A vendor is cheap to reverse if we own the data: swap it. Case management is not ours to reverse at all. But a data layer built to the wrong latency cannot be reversed, only rebuilt — and the rebuild takes the label history with it. So almost the whole ambition budget goes into one place.
Schema, entities, and a stated latency per source. Built wrong, it cannot be reversed — it has to be rebuilt, and the rebuild destroys the outcome history that every model depends on.
It is also the thing no vendor can sell us. No engine arrives knowing Scalable's Broker, Wealth, Credit and Banking activity joined at one client. Buying the engine does not buy the client view.
| What we give up | What it buys | What it costs | What would reverse it |
|---|---|---|---|
| Building the engine | Time to market, and something to benchmark an internal build against | Our controls are only as expressive as the vendor's rule builder, and their data model may not fit ours | The walkthrough shows the builder cannot express our controls or replay our history |
| Precision in release 1 | The label loop starts turning immediately | Analysts work more false positives in the first quarter than they will later | Volume breaches the analyst ceiling in back-test — then we cut scope, not precision |
| Machine learning in release 1 | No dependency on labels we do not have | The largest precision lever, lowering risk on clients who always behave this way, sits unused for a release | The feasibility check passes both preconditions |
| Building case management | Six engineers stay on the feed | We depend on another team's roadmap for the thing that gates M2 | The assessment shows the gap cannot be bridged and the market offers nothing better |
| A restriction on the layering control | No automatic customer harm on the highest-volume typology | Money laundering is detected but never interrupted | Compliance sets an appetite that says otherwise, or we enter a market where it is standard |
What actually limits how strong the controls can be at launch
Not detection capability. Analyst headcount. Release-1 strength is bounded by alerts per analyst per day. A control set strong enough to breach that ceiling produces a backlog, and a backlog is itself a regulatory finding, so the ceiling is a hard limit rather than a preference.
That is why back-testing is the first stage of every control rather than a checklist item before launch: the volume number has to exist before the scope decision, not after it. And the consequence, stated plainly — if the back-test says the control set is too large for the team, we cut typology scope rather than degrade precision. Fewer controls working properly beats three controls the queue cannot absorb. That call is made with Compliance, who have the final word, because the risk appetite is theirs.
Two trades I would refuse
- Shipping without typed outcomes. The cheapest-looking corner to cut, and it quietly kills the entire adaptive roadmap, because everything past M5 depends on label volume and label quality. Speed bought here is paid for with the second year.
- Going live without a back-test on production data. The volume risk and the precision risk have the same mitigation. Skipping it converts a measurable risk into a surprise that lands on operations.
Benchmarking an internal build costs nothing extra
Running an internal build alongside the vendor doubles alert volume unless one of them runs silently. It already does. The internal build runs at stage 2 while the vendor's control runs at stage 3 or 4 — same rollout stages, no new machinery, no extra alerts in the queue, and a direct comparison on identical live traffic. The benchmark stops being a separate project.
Ambition is also limited by who is on the team
Six data engineers, no data scientists, no front-end engineers. Two decisions already taken in Task 1 — nothing customer-facing, no machine learning in release 1 — are the same constraint seen twice.
What follows from that is a position rather than a request: we do not ask for data scientists before there is evidence there are labels for them to train on. The way to get that evidence is not a headcount case. It is to find existing machine-learning capability elsewhere in the company, run the feasibility check cheaply with it, and let a working first result make the argument. De-risk first, then ask.
2.3What is proved before go-live, and what can only be measured after
Required sub-point: validation before launch, what follows after, and how the system improves.
Precision cannot be known before go-live, and the estimate is optimistic
This is structural rather than a matter of care. A back-test scores a new control against historical outcomes produced by a different system firing different alerts. Cases the new control would have caught, but the old queries never surfaced, were never investigated and carry no outcome at all. So they cannot count against it.
A back-test therefore gives volume reliably, a partial recall signal — replay against known filed reports and ask whether the control would have fired — and precision only as an estimate. The back-test is a capacity instrument, not a quality one, and the alerts-per-analyst ceiling should be planned conservatively against it rather than exactly to the number it produces.
Before go-live — week 14
- The feed is correctEvery source reconciled against its own source system. A control validated on a wrong feed is validated on nothing.
- The record can reconstruct a real caseReplay a case financial crime already investigated, end to end. The best single test available — it checks the joins, the event capture and the latency at once.
- Volume fits the queueBack-test against the agreed ceiling. A backlog on day one is a regulatory finding.
- The loop closesOne real alert: created, queued, worked, closed with a typed outcome, stored in our own backend. If this does not work, nothing after M2 happens.
- Restrictions can be liftedApply and reverse every mitigation. If the reversal path was never built, lifting a block needs an engineer — and every false positive becomes an engineering ticket.
- The explanation rendersAn analyst works shadow alerts and can say why each one fired. Validated by an analyst using it, not by a screenshot.
After go-live
- Real precision per controlFrom closed cases. The first number about this system that has ever been true.
- The real incidence mixThe build order is reprioritised against it. If account takeover turns out not to be rare, it moves up.
- Threshold tuningThe first thing that needs changing, and the cheapest.
- The model feasibility verdictIf labels fall short, fixing label capture becomes the next release rather than the model.
- Account takeover to a restrictionOnce device and IP data allows it.
- The case-management decision executedWhichever way it went.
- The next product lines onto the recordThe first real test of the claim that this platform is cheap to extend.
Item 5 on the left is the one most teams miss, and it is the only item on that list that exists purely for the customer.
Three things make the system better, and they are not the same thing
Tuning. Review precision per control, then adjust it, replace it, or propose removing it. This works from day one and needs nothing but closed outcomes. Removing a control is not ours to decide — an underperforming control comes off only with explicit Compliance approval, given case by case and never in advance. What we own is the evidence that makes that conversation possible: precision and volume over time, plus a proposal for how the risk stays covered once the control is gone.
Compounding. Labels accumulate. Client history and peer comparison replace arbitrary thresholds. Lowering risk on a client who genuinely always behaves this way becomes possible, which is the largest precision lever available. Then a model that ranks the queue. Each step needs more label volume than the one before, which is why none of it starts until the loop has been turning for a while.
Coverage. Neither of the other two ever finds a risk nobody wrote a control for. So a regular threat-modelling pass asks what surface now exists that nothing watches: new products, new attack routes, and the places the closed loop already leaks — securities transferred out, crypto withdrawn to an external wallet, joint accounts, loan proceeds.
Improvement is not something the vendor supplies. It is a property of the label store we refused to let them own.
2.4Two forums, one gate, and a rule that stops it growing
Required sub-point: operating cadence, forums and artefacts.
The engineering team's iteration rhythm is the engineering manager's call and I would not touch it. Stakeholders are steered on milestones instead: M0 to M5, exit conditions agreed before each milestone starts, and status reported as which condition is not yet met and who owns the blocker — not as velocity or scope burned down.
This suits this initiative specifically rather than as a general preference. The three things that will actually set the timeline are all outside this team: whether device and IP data exists, another team's case-management roadmap, and a vendor procurement cycle. Committing to sprint scope against those produces theatre, and it costs trust the first time a dependency slips. A milestone with an honest exit condition survives a slipped dependency. A sprint commitment does not.
Two rules stop that from being only a label. Exit conditions are agreed when the milestone starts, not when it ends. And a milestone can be re-scoped but never quietly re-dated — if M4 slips, the one-pager names the unmet condition and who owns the blocker.
First line judges whether the record can reconstruct a real case, because they know what an investigation needs to see. Operations judges whether the queue can absorb the volume, because they are the ones who drown. Compliance judges whether a control is fit to restrict a customer automatically, because they carry that exposure.
- What movement of value does this product create?
- Which typology does it touch?
- Is that covered, or explicitly accepted?
Six artefacts, and not all of them need a meeting
| Artefact | What it is for |
|---|---|
| Coverage map | Product surface against typology, with coverage status. Makes “no gaps” auditable, and records what is covered by product design rather than by detection — which is how the MVP justifies what it leaves out. |
| Control specification | One per control: typology label, the entity it fires on, latency, launch stage, default restriction, the policy it implements, and its version. Reference material rather than an agenda, and it doubles as the artefact an auditor needs. |
| Control scorecard | The tuning input, and the evidence behind any request to Compliance to remove a control. |
| Milestone one-pager | Exit conditions and their status. |
| Assumption register | The seven assumptions, each with an owner and a test. This is the M0 work plan — closing an assumption is a unit of progress, which is what makes discovery a milestone rather than a preliminary. |
| Decision log | What was decided, by whom, and what would reverse it. Especially for decisions Compliance owns, so that a later change is a re-decision rather than an argument. |
Label quality is something stakeholders produce, not something the system provides. Task 1 commits first and second line to checking typology labels inside their own alert quality assurance. That only holds if they own the taxonomy rather than receive it. So the taxonomy is agreed in the quarterly session of the control forum, and label quality appears on the scorecard as a tracked number — visible to the people whose work produces it.
2.5What I would put in front of someone else, in week one
Required sub-point: risks, open questions and decisions to escalate early.
This is a different list from 1.8. That one asked which risks the system has to handle. This one asks what gets handed to somebody else, and when. A risk that can be resolved inside this team is not an escalation. It is a task.
Something belongs here if it cannot be resolved inside the team and the answer gets more expensive the longer it stays open. And it is escalated with a recommendation and a date, never as an open problem — handing someone a problem moves the work sideways, handing them a recommendation they can approve or reject moves it forward.
| What | Who | The ask | What a late answer costs | |
|---|---|---|---|---|
| 1 | Case management | The team that owns it, plus leadership | A named owner and a decision date. Not a solution. | In week one this costs a conversation. In month four it costs a milestone, because the first control's alerts have nowhere to go. |
| 2 | Does device and IP data exist in a form financial crime can use? | Data engineering, and whoever owns client-side telemetry | Answer it. If the answer is no, start the work to generate it now. | This is the one with a project attached to the answer, and that project sits with another team on its own lead time. Late, and account takeover has no path into release 1. |
| 3 | Is one person guaranteed to be one client record? | Onboarding and identity | A yes or a no. | Almost certainly fine, and raised for completeness. It earns a line because of joint accounts and children's accounts — legitimate cases where more than one person touches one record, which is exactly the shape client-level monitoring has to get right. |
The decisions that are Compliance's, not ours
Risk appetite produces three numbers: how many alerts a day the team can absorb, what precision each control has to hold, and what restriction each alert carries by default. Without them, the scope decision in 2.2 has nobody to arbitrate it and the minimum precision in Figure 4 has no value in it.
We arrive with a position on all three rather than a blank sheet. The decision is theirs. The recommendation is ours.
The one risk that gets closed rather than watched
The vendor's data model and typology structure may not fit ours, which would mean rework after signature. Every other risk on this page is monitored over time. This one is removed before commitment, by walking the vendor through exactly what we intend to build and confirming they support it. The escalation is procedural: get agreement that the walkthrough happens before the contract rather than after. A small ask early and an impossible one late.
Two positions I would state rather than escalate
- History. We load the maximum available. The only limit is data already deleted under retention rules. Baselines built from nothing make every long-standing client look new, and therefore unusual.
- The model feasibility check is not a headcount request. Find existing machine-learning capability elsewhere in the company, run the check cheaply, and let a working first result make the case for resourcing.
No new assumptions
Task 2 adds nothing to the register at the top of this page. The delivery plan rests on A1 for the build order, and on A2 and A3 for what release 1 can contain. The product scope and the delivery plan stand on the same seven assumptions, and the same seven tests would falsify either.