Backtest before you switch on: testing monitoring rule changes safely
Replay recent scored events through the proposed configuration, then read each decision that would change as well as the count. A backtest shows what a change would have done to the past; your written expectation is what tells you whether that is the right result.

Charles Archibong, Co-founder
· 5 min read

Key takeaways
- Write down the effect you expect before running a backtest, or any result will look acceptable.
- Read the individual decisions that flip; a small net change can hide large moves in both directions.
- History is shaped by your old decisions, so a backtest can understate what a looser rule would let through.
- Change one thing per test, publish it as a new version, and check live results against the forecast.
You know what a rule change will do before it goes live by replaying recent history through the proposed configuration and reading the decisions that would change. That is a backtest. It answers a narrow question well: on events we have already seen, what would this configuration have decided differently?
It does not tell you whether the change is right. That judgement comes from a written expectation made before the test, and from the individual events that flip, not the summary count. Teams that skip either step end up approving changes because the number looked small.
What does a backtest actually tell you?
A backtest takes stored events, scores each one again with the candidate rules, and compares the new decision with the one recorded at the time. The useful outputs are:
how many events would move between ALLOW, REVIEW and BLOCK;
which events move, and in which direction;
which rules fire more or less often under the new configuration.
In Myaza Transaction Monitoring, the Test changes step does this before you save. It replays recent scored events as a read-only historical test and shows how many decisions would change. It creates no alerts, investigations, customer updates or charges, so it is safe to run as often as you like. Backtests run against up to 1,000 stored events. The monitoring rules documentation covers the configuration it tests.
Why write the expected result first?
Without a stated expectation, any backtest result can be rationalised. A 3% drop in REVIEW decisions sounds like a win if you wanted fewer alerts and a loss if you wanted better coverage.
Write two or three sentences before running the test:
The change: "Raise the NGN single-transaction limit from ₦2,000,000 to ₦3,500,000."
The reason: "Most threshold alerts on salary-account customers in the last quarter were cleared as false positives."
The expected effect: "Fewer threshold REVIEW decisions on salary accounts. No change to BLOCK decisions. No confirmed-fraud event from the last quarter should move to ALLOW."
The figures are illustrative. The last sentence is the important one: it names the result that would make you reject the change.
How do you read the decisions that flip?
The net change hides the gross movement. A configuration that sends 30 events from REVIEW to ALLOW and 28 from ALLOW to REVIEW shows a net change of 2, yet it has changed 58 decisions.
Sort the flips into four groups and read a sample from each:
Movement | Question to ask |
|---|---|
REVIEW to ALLOW | Were any of these alerts confirmed as suspicious when they were worked? |
ALLOW to REVIEW | Do these look like the risk the change was meant to catch, or new noise? |
BLOCK to REVIEW or ALLOW | Is loosening a block intended? This group deserves the closest reading. |
ALLOW or REVIEW to BLOCK | Would blocking these have stopped legitimate customers? |
Cross-reference the first group with alert labels. If an event that moves to ALLOW had an alert confirmed as a true positive, the change would have removed a genuine detection. One of those is enough to stop and rethink.
For custom rules, also check that the rule fires at all. A condition written against a field that the integration does not send will show zero fires, which a quick glance can mistake for "no impact".
What can a backtest not tell you?
Four limits matter in practice.
History reflects your old decisions. If a past configuration blocked a transaction, the customer may have given up, tried a different amount, or moved elsewhere. The events that would have followed never happened. A backtest of a looser rule therefore sees fewer of the transactions it would now allow, and can understate the effect.
The window is finite. A replay of recent events may not include a quiet month, a seasonal peak, or a typology that appears a few times a year. If your business has a salary cycle or festive-season peaks, check whether the window covers one.
Labels lag. Recent alerts may not have been worked yet. An event with no label is not evidence of anything, so treat a flip on an unlabelled event as unknown rather than as a safe change.
Customer behaviour moves. Criminal behaviour adapts to controls, and legitimate behaviour shifts with prices and products. A backtest is a forecast, not a guarantee.
None of these limits is a reason to skip the test. They are reasons to watch live results after the change.
How do you run a change from proposal to production?
A simple sequence keeps changes controlled and auditable.
Propose one change. One rule, one limit or one weight. Bundled changes cannot be attributed.
Write the expected effect, including the result that would make you reject it.
Run the backtest in the environment where the change will apply.
Read the flips in the four groups above and cross-check confirmed alerts.
Get a second reviewer for any change that loosens a BLOCK or removes coverage of a named risk.
Publish as a new version. In Myaza, publishing a fraud rule creates an immutable version, so an earlier decision always points to the exact rule that produced it. That is what lets you answer "why was this transaction allowed in March?" a year later.
Compare live results with the forecast after a set period, and record whether the expectation held.
Permissions support the second-reviewer step. In Myaza, anyone with Risk Intelligence access can view the configuration, while editing needs the monitoring:manage permission, so you can separate the people who propose changes from those who apply them.
Does the same discipline apply to ongoing monitoring policies?
Yes, with one difference. Rule changes alter how individual events are scored. Continuous monitoring policies decide which customers are re-checked, how often, and when a change is material enough to raise an alert.
Myaza stores monitoring policies as immutable versions. Changing a policy's default frequency updates customers who follow the default, but it does not overwrite a customer with an explicit frequency override. Before changing a policy default, check how many customers follow it and how many have overrides, because the change will reach only the first group. The continuous monitoring documentation describes the policy and override fields.
The questions are the same as for rules: what do you expect to change, for whom, and what result would make you reverse it.
A checklist for every rule change
One change per test.
Expected effect written before the test, including a rejection condition.
Backtest run in the same environment as the change.
Flips read by direction, with confirmed alerts cross-checked.
Second reviewer for anything that loosens a block or removes coverage.
Published as a new version with the reason recorded.
Live results compared with the forecast on a set date.
If a change cannot be explained in those terms, it is not ready to go live.
Sources

Charles Archibong
Co-founder
Charles Archibong co-founded Myaza Trust. He writes about identity verification, financial technology, and the practical work of building trusted digital services.


