Insight AEGI Shield

When Is an AI or Risk Workflow Ready for Controlled Evaluation?

A practical readiness checklist for deciding whether a banking AI, fraud or risk workflow can be evaluated as one bounded change rather than becoming an open-ended build project.

QUICK ANSWER

A workflow is ready for controlled evaluation when the institution can define one current baseline, one bounded proposed change, one usable evidence path, one material operating constraint and one real decision consequence.

If those elements are missing, the next step may be workflow design, data engineering, model development or governance scoping rather than evaluation.

AEGI Shield treats readiness as a control. It should not hide an open-ended build programme inside something called a “pilot”.

WHY READINESS MATTERS

Evaluation quality depends on the question being bounded before execution begins.

If the baseline changes during the test, the candidate keeps evolving, the labels are immature, the population is unclear or success criteria are chosen after results are seen, a polished report can still provide weak evidence.

The first purpose of readiness assessment is therefore to decide whether the proposed evaluation can answer a meaningful institutional question.

READINESS CONDITION 1: A REAL WORKFLOW

The institution should be able to identify the operational workflow being changed.

Examples include fraud-review prioritisation, account-opening risk review, credit-fraud escalation, case routing, a new signal entering an existing queue or an AI-assisted recommendation entering a controlled review step.

“Use AI for fraud” is not a bounded workflow.

READINESS CONDITION 2: A RECONSTRUCTABLE BASELINE

The baseline is the authorised reference point against which the candidate will be compared.

It can include:

  • a production model version;
  • a set of rules and thresholds;
  • a ranking treatment;
  • a manual review process;
  • a current combination of model and policy behaviour.

The baseline must be reconstructable well enough that the candidate is not given an artificial advantage.

READINESS CONDITION 3: ONE BOUNDED PROPOSED CHANGE

The candidate should be clear enough to freeze.

It might be:

  • a new signal;
  • a challenger model;
  • a threshold adjustment;
  • a ranking treatment;
  • Shared Credit-Fraud Risk Context;
  • an AI-assisted recommendation;
  • a policy-routing change.

If several unrelated changes are introduced at once, it becomes difficult to know which change produced the observed effect.

READINESS CONDITION 4: APPROVED HISTORICAL EVIDENCE

The institution needs an evidence path that can support the declared comparison.

This does not mean “all available data”. It means the minimum approved evidence necessary to reconstruct the baseline, run the candidate and measure the intended outcome without hindsight leakage.

If important candidate inputs were unavailable at the original decision time, the evaluation must account for that limitation.

READINESS CONDITION 5: OUTCOME SEMANTICS

The institution should know what the outcome label means and when it becomes mature enough to use.

For fraud, a label can depend on investigation, customer reporting, chargeback or other operational processes. For credit, relevant outcomes can mature over much longer periods.

An immature label can make a precise-looking metric misleading.

READINESS CONDITION 6: A BINDING OPERATING CONSTRAINT

Many banking workflows are not limited by model inference. They are limited by investigator capacity, manual review slots, customer-friction tolerance, latency, policy boundaries or available follow-up resources.

The evaluation should hold the material constraint consistently across the baseline and candidate.

For example, if fraud operations can review only the top 1% of applications, comparing two models at different review volumes does not answer the operational question.

READINESS CONDITION 7: A DECISION OWNER

Someone must own the decision the evaluation is intended to inform.

The owner may sit in fraud, risk, credit, product, model risk or another function. The key point is that the evaluation should not produce evidence with no institutional consumer.

A good readiness question is:

If this result is positive, who can authorise the next bounded step? If it is negative, who can decide to stop?

READINESS CONDITION 8: CONTINUE, REFINE AND STOP ARE ALL POSSIBLE

An evaluation is weaker when the organisation has already decided that the candidate must progress.

AEGI treats three outcomes as first-class:

  • Continue — evidence supports another controlled step;
  • Refine — the candidate, scope or evidence basis needs material revision;
  • Stop — evidence does not justify further progression.

This reduces confirmation bias.

READINESS CONDITION 9: AUTHORITY IS EXPLICIT

The evaluation should state what the candidate is allowed to do.

Historical replay should not mutate production. A shadow stage should not silently change customer outcomes. A recommendation should not become action authority simply because confidence is high.

Bank-owned authority boundaries need to be clear before execution.

READINESS CONDITION 10: MATERIAL LIMITATIONS CAN BE REPORTED

A credible evaluation must be allowed to report that the evidence is weak, incomplete or non-generalizable.

If the organisation expects every study to produce a success story, the evaluation process is not yet mature enough.

A SIMPLE GO / REFINE / NOT-YET CHECK

GO when the workflow, baseline, candidate, historical evidence, outcome, operating constraint and decision owner are all sufficiently defined.

REFINE when the question is useful but one or two elements need clearer definition or a narrower scope.

NOT YET when there is no valid baseline, no frozen candidate, no usable evidence path or no actual decision to make.

WHAT AEGI DOES AT THE READINESS STAGE

AEGI can help structure the question, identify whether the baseline and candidate can be compared fairly, define the evidence boundary and make the decision output explicit.

That is different from promising that every workflow is ready for AEGI.

A refusal to run a weak evaluation can itself be evidence of governance discipline.

FREQUENTLY ASKED QUESTIONS

What if we have a new signal but no candidate logic?

That is candidate-design work first. It should not be hidden inside a validation exercise.

What if we have a candidate but no stable baseline?

Baseline reconstruction becomes the first task. Without it, comparative evidence will be weak.

What if the labels are incomplete?

The evaluation can wait, use a different declared outcome, or report insufficient evidence.

What if the team only wants a technical benchmark?

That can still be useful, but it should be labelled as technical or mechanism evidence rather than institution-specific workflow value.

NEXT STEP

If you can describe the workflow, current baseline, proposed change and decision that needs evidence, AEGI can first assess whether the question is sufficiently bounded for a Controlled Workflow Evaluation.

Bring One Workflow →

CLAIM BOUNDARY

Readiness for controlled evaluation is not production approval, regulatory approval, formal model validation or security approval. Those remain separate institution-owned gates.

RELATED AEGI RESOURCES

NEXT STEP

Bring one proposed change.

Controlled Evaluation Bring One Question