Insight Fraud & Risk Workflows

How to Test a New Fraud Signal Without Replacing Your Existing Fraud Engine

A practical method for testing one additional fraud signal against the current workflow without replacing the existing fraud engine or granting the candidate production authority.

QUICK ANSWER

A new fraud signal should not be adopted because it looks predictive in isolation. The useful question is whether adding that signal to a defined candidate changes a real fraud-review decision enough to justify the next controlled stage.

The lowest-risk way to answer that question is usually to preserve the existing fraud engine as the baseline, define exactly how the new signal would influence a candidate treatment, replay both on the same approved historical population, hold relevant review constraints constant, and compare the resulting reviewable evidence. If the candidate is not yet defined, signal discovery and candidate-building should be treated as a separate scope rather than silently mixed into validation.

START WITH THE EXISTING FRAUD ENGINE, NOT A REPLACEMENT PROJECT

Most established financial institutions already have some combination of fraud scores, rules, device intelligence, transaction monitoring, investigator queues, case-management systems and manual judgement. The operational question is often not whether to throw that stack away.

It is whether one additional source of context can improve a specific decision inside it.

Examples may include:

  • a new device-risk signal;
  • graph-derived relationship features;
  • behavioural context;
  • an external consortium or network signal;
  • merchant or beneficiary risk context;
  • a new scam indicator;
  • a credit-side or broader risk signal where the workflow and authority permit its use;
  • a new rule-derived feature;
  • an internal model output that previously was not available to the review workflow.

These are examples, not a recommendation that every institution should combine every signal. The value of a signal is workflow-specific.

THE FIRST QUESTION IS NOT “IS THE SIGNAL PREDICTIVE?”

A signal can correlate with fraud and still fail operationally.

It may arrive too late. It may be available only for a small population. It may duplicate information already embedded in the baseline. It may increase sensitivity while also creating too many false escalations. It may improve a global metric but fail to change the set of cases investigators can actually review. It may introduce data-quality, provenance, privacy or policy constraints that make deployment unattractive.

A stronger question is:

If we add this signal through a declared candidate treatment, what changes relative to the current baseline under the same relevant operational constraints?

That formulation connects the signal to a decision rather than treating it as an isolated feature-engineering exercise.

STEP 1 — DEFINE THE WORKFLOW AND DECISION

Choose one bounded workflow.

For example:

“Fraud review prioritisation for outbound account-to-account transfers.”

Then define the decision:

“Does adding signal X to the current prioritisation logic produce enough evidence to justify a non-customer-impacting shadow stage?”

The question should be specific enough that a positive, negative or inconclusive result changes a next action.

If no owner can explain what would happen after the result, the evaluation is not yet well bounded.

STEP 2 — DEFINE THE SIGNAL BEFORE DEFINING ITS VALUE

A signal should have an explicit contract.

At minimum, document:

  • source;
  • business meaning;
  • event timestamp;
  • availability timestamp;
  • coverage;
  • missing-value behaviour;
  • update frequency;
  • expected population;
  • permitted use;
  • version or source identity where relevant.

This matters because “the same signal” can mean different things if its generation logic, refresh cadence or source system changes.

A signal is not useful to an evaluation simply because it exists in a data warehouse. It must have been available at the point in time when the real decision would have been made.

STEP 3 — CHECK FOR TEMPORAL LEAKAGE

This is one of the easiest ways to create a misleading fraud result.

Suppose a signal incorporates a chargeback outcome, investigator disposition or post-event network relationship that became known only after the original transaction decision. Using that signal in historical replay can make the candidate appear stronger than it could have been in production.

The evaluation should therefore distinguish:

EVENT TIME

When the transaction or review decision occurred.

EVIDENCE TIME

When each signal or feature became available.

OUTCOME TIME

When the label or reviewed outcome became sufficiently mature to use for evaluation.

Only information that would have been available before the declared decision point should be allowed into the candidate path.

STEP 4 — DEFINE THE CANDIDATE TREATMENT

A signal by itself is not a workflow change.

The institution must define how the signal affects the existing process.

Examples:

  • add the signal as an input to a challenger score;
  • use it to re-rank the existing alert queue;
  • add a bounded rule when the signal crosses a declared threshold;
  • use it as additional context for a defined review segment;
  • use it to create a treatment score while leaving the baseline model untouched.

This distinction matters commercially and technically.

If the institution already has the candidate logic, the evaluation can focus on comparison and evidence.

If the institution only has a raw signal and asks AEGI or another evaluator to invent the best model, threshold and operational use, that is candidate-building work. It should be explicitly scoped rather than hidden inside “validation.”

STEP 5 — RECONSTRUCT THE CURRENT BASELINE

A new signal can only be evaluated relative to a credible baseline.

The baseline may be:

  • current model score;
  • current rule set;
  • current alert ranking;
  • current threshold;
  • current investigator routing logic;
  • a declared combination of these.

Record the baseline version and material configuration. If the baseline cannot be reconstructed on the historical population, the comparison may not answer the real operational question.

Do not quietly substitute a simplified baseline merely because it is easier to reproduce. If approximation is unavoidable, label it as a limitation.

STEP 6 — CHOOSE A HISTORICAL WINDOW AND OUTCOME BASIS

The historical window should represent the intended decision as closely as possible while preserving outcome maturity.

Questions include:

  • Does the period include seasonality relevant to the fraud pattern?
  • Are labels sufficiently mature?
  • Did policy or product changes materially alter the population?
  • Were all required signals available during the period?
  • Is the customer segment comparable to the intended future use?

An evaluation window is not just “the data we happen to have.” It defines the population to which the result directly applies.

STEP 7 — HOLD THE IMPORTANT CONSTRAINTS CONSTANT

If the fraud-review workflow is capacity constrained, compare the baseline and candidate at the same review capacity.

For example, if investigators can review only the highest-ranked K cases, evaluate what each treatment places inside those K slots.

Other constraints may include:

  • latency;
  • customer-contact limits;
  • manual review hours;
  • policy exclusions;
  • business hours;
  • segment-specific routing;
  • technical availability or fallback behaviour.

The principle is simple: do not allow the candidate to win by consuming resources that the baseline was not allowed to use.

STEP 8 — SELECT A PRIMARY METRIC AND GUARDRAILS

The primary metric should follow the decision.

For ranked review, examples include fraud-labelled cases surfaced at K, precision at K, recall at K or a customer-defined review-value metric.

Guardrails may include:

  • false escalations;
  • investigator effort;
  • latency;
  • customer friction;
  • segment stability;
  • missing-signal behaviour;
  • fallback rates;
  • calibration or score distribution changes where relevant.

Do not infer metrics that were not measured. Equal alert counts do not prove equal operational cost. Additional fraud-labelled cases do not automatically equal prevented fraud losses.

STEP 9 — ANALYSE WHO MOVED, NOT ONLY THE TOTAL

A useful signal evaluation looks beyond headline totals.

Compare:

  • cases reviewed by both baseline and candidate;
  • cases added by the candidate;
  • cases removed by the candidate;
  • positive and negative outcomes in each group;
  • important segments or fraud typologies;
  • missing-signal cases;
  • cases close to the review cutoff.

This can reveal whether the signal is adding genuinely new information or simply reshuffling already-obvious cases.

If the candidate improves the total but systematically drops a material risk segment, that may be a Refine or Stop outcome rather than Continue.

STEP 10 — QUANTIFY UNCERTAINTY AND LIMITATIONS

Fraud outcomes are noisy and often sparse. Small improvements near a review threshold can also be unstable.

Where appropriate, use confidence intervals, paired resampling or other declared uncertainty methods around the primary metric.

Record limitations such as:

  • incomplete labels;
  • short historical period;
  • population drift;
  • unresolved data lineage;
  • missing operational effort measures;
  • signal coverage gaps;
  • baseline reconstruction approximations;
  • policy changes during the sample period.

The goal is not to make the result look statistically sophisticated. It is to show how much confidence the institution should place in the conclusion.

STEP 11 — MAKE CONTINUE, REFINE AND STOP ALL VALID

A credible evaluation should not be designed to prove that the signal deserves deployment.

CONTINUE

The candidate produces enough evidence to justify the next controlled stage.

REFINE

The signal may be useful, but the candidate, protocol, evidence or guardrails need revision.

STOP

The candidate does not justify progression under the declared basis, the baseline is adequate, or the evidence is insufficient for the intended decision.

STOP can be valuable. It may prevent unnecessary integration work, vendor spend, model change or operational disruption.

STEP 12 — KEEP PRODUCTION AUTHORITY SEPARATE

A positive historical result is not authority to change the live fraud workflow.

The institution may still require model validation, architecture review, privacy/security review, policy approval, change-management controls, shadow testing, operational readiness or executive sign-off.

A bounded evaluation should make the next decision clearer without manufacturing authority it does not possess.

A WORKED CONCEPTUAL EXAMPLE

Assume the current fraud engine ranks eligible cases using Baseline B. The institution is considering Signal X, which is available before the review decision for 92% of the target population.

Candidate C uses the same baseline score but adds Signal X through a declared ranking treatment. The historical window and labels are frozen. The review team’s capacity K is held constant.

The evaluation compares:

Baseline B — Top-K review set

versus

Candidate C — Top-K review set

The report then examines:

  • positive cases surfaced in each Top-K set;
  • candidate-only and baseline-only cases;
  • signal-missing behaviour;
  • key risk segments;
  • confidence interval around the primary difference;
  • any operational guardrail breaches.

If Candidate C surfaces more relevant cases but fails on a required segment or workload constraint, the result may be REFINE rather than CONTINUE.

This example illustrates the method only. It is not a claim that Signal X will improve any real institution’s fraud performance.

HOW THIS RELATES TO AEGI SHIELD

AEGI Shield is designed to work alongside existing fraud and risk systems rather than require model replacement for the initial evaluation proposition.

For a bounded Controlled Evaluation, the institution provides a defined workflow, baseline/candidate basis, approved historical evidence, outcome semantics, review constraints and responsible owners. AEGI structures the agreed comparison, reproducibility and reviewable evidence within the declared scope.

Where the proposed change involves additional approved context, Shared Risk Context may be part of the treatment. It is not a requirement for every first evaluation.

The output supports a Continue, Refine or Stop decision. Customer-action and production-change authority remain with the institution.

WHEN THIS APPROACH IS NOT A FIT

Do not force a new-signal evaluation when:

  • no baseline can be reconstructed;
  • the signal was not available at decision time;
  • outcome labels are too immature;
  • the proposed use of the signal is undefined;
  • the signal requires an unapproved data path;
  • the workflow owner cannot define what decision the result will change;
  • the task is actually open-ended model development rather than evaluation.

In those cases, feasibility or candidate-design work should happen before a controlled comparison.

FREQUENTLY ASKED QUESTIONS

Should a new fraud signal be tested as a standalone predictor?

Usually not if the real decision is how it will change an existing workflow. The signal should be evaluated through a defined candidate treatment against the current baseline.

Why is evidence time important?

A signal that appears useful in historical data may be invalid for the intended workflow if it became available only after the original decision point. Event time, evidence time and outcome time should be separated explicitly.

Does testing a new signal require replacing the current fraud engine?

No. The current fraud engine can remain the baseline while the signal is added through a bounded challenger, rule or re-ranking treatment.

What if the institution has a signal but no candidate logic yet?

That is candidate-design work, not simply validation. It should be scoped separately so an open-ended modelling project is not hidden inside a controlled evaluation.

CONCLUSION

The safest way to test a new fraud signal is not to replace the existing engine and hope for a better result. Preserve the current workflow as the baseline, define how the new signal changes one candidate treatment, replay both on the same historical evidence, hold important constraints constant and record enough evidence to decide whether the change deserves progression.

That turns “we found an interesting signal” into a much more useful institutional question:

Does this signal create enough decision-relevant value, in this workflow and under these constraints, to move to the next controlled stage?

NEXT STEP

Bring one proposed change. If you already have a signal, candidate and current baseline, AEGI can first assess whether the workflow and historical evidence are sufficiently bounded for a Controlled Evaluation.

CLAIM BOUNDARY

This article is educational and does not establish that any specific signal will improve fraud detection, reduce fraud losses or satisfy institution-specific validation, security, privacy or regulatory requirements. AEGI customer value remains to be established within each institution’s own evaluation scope.

SOURCES

[1] IBM, “Why operational capacity is becoming the binding constraint in fraud detection,” 16 June 2026:

https://www.ibm.com/think/insights/fraud-detection-no-longer-limited-models-limited-by-operations

[2] Kim et al., “Champion-challenger analysis for credit card fraud detection: Hybrid ensemble and deep learning,” Expert Systems with Applications, 2019:

https://www.sciencedirect.com/science/article/pii/S0957417419302167

[3] MAS remarks on Annual Report 2024/2025, via BIS — AI governance, evaluation, testing and explainability:

https://www.bis.org/speeches/20250805-remarks-mas-annual-report-20242025

Related AEGI resources

NEXT STEP

Bring one proposed change.

Controlled Evaluation Bring One Question