Insight AI Change Evaluation

When Should an AI Candidate Be Stopped Rather Than Promoted?

Why Stop should be a normal, reviewable outcome in financial-services AI evaluation — not a failed project.

QUICK ANSWER

An AI candidate should be stopped when the available evidence does not justify further progression under the institution’s declared decision criteria. That can happen because the candidate produces no material improvement, breaches an important guardrail, relies on invalid or unavailable evidence, creates excessive operational burden, cannot be reproduced, changes risk in an uncontrolled way, or simply fails to outperform an adequate baseline.

A credible evaluation process must allow Stop. If every test is designed to produce a “go” decision, the process is not independently evaluating the candidate; it is only documenting a predetermined promotion path.

WHY TEAMS STRUGGLE TO STOP

AI projects accumulate momentum. Teams invest time, budget and reputation in a candidate. A new model may be technically interesting, a new vendor may have executive sponsorship, or a pilot may have been announced publicly. Those factors can create a subtle bias toward progression even when the evidence is weak.

This is especially dangerous in regulated or high-impact workflows because the cost of moving an inadequately supported change forward can appear only later — during integration, investigation workload, customer friction, control review or production monitoring.

The purpose of staged evaluation is to reduce that uncertainty before commitment increases.

STOP CONDITION 1: NO MATERIAL DECISION VALUE

A candidate may improve a technical metric but not enough to change the operational decision. A fraud-ranking model that improves a global score but leaves the Top-K review set almost unchanged may not justify implementation effort. A new signal may add complexity without materially changing prioritisation. A model refresh may be statistically different but operationally irrelevant.

This is why materiality should be defined before seeing the result where practical.

The question is not “Did anything improve?” It is “Did enough improve to justify the next stage?”

STOP CONDITION 2: THE BASELINE IS ADEQUATE

The current system is not automatically a problem because it is old. A stable baseline with acceptable performance, known limitations and mature controls may be preferable to a candidate whose marginal benefit is small.

“Baseline adequate” is therefore a legitimate evaluation outcome.

This can create real value by preventing unnecessary implementation, integration, retraining, governance and change-management work.

STOP CONDITION 3: A MATERIAL GUARDRAIL FAILS

A candidate can improve the primary metric and still be unsuitable if it breaches a critical constraint.

Examples include:

  • excessive false escalations;
  • unacceptable customer friction;
  • materially higher investigator workload;
  • latency outside the workflow’s decision window;
  • unstable behaviour in a protected or high-risk subgroup;
  • security or data-handling requirements that cannot be satisfied;
  • dependence on evidence that is unavailable at the time of decision.

Primary metrics do not override guardrails automatically.

STOP CONDITION 4: THE EVIDENCE BASIS IS INVALID

A precise result can still be wrong if the evaluation used temporally leaked information, mismatched populations, inconsistent label definitions, a changed baseline, incomplete lineage or non-reproducible transformations.

In such cases, the correct result may be “not evaluable in the declared scope,” not a positive or negative performance judgement.

That distinction matters because an invalid test should not be converted into a commercial claim.

STOP CONDITION 5: THE CANDIDATE CHANGED DURING THE TEST

If the candidate’s model version, prompt, rule set, threshold, feature set or other material configuration changes during evaluation, the result may no longer refer to the same candidate.

The solution is not always to discard all work. The institution can determine whether the change is immaterial or whether a new evaluation basis is required. But evidence should remain bound to the version actually tested.

STOP CONDITION 6: THE OPERATIONAL COST IS TOO HIGH

A candidate may surface more useful cases but demand more investigator time, more specialist review, more customer contact or more infrastructure than the institution is willing to allocate.

Operational burden should be measured directly where it matters. Equal review slots, for example, do not necessarily mean equal investigator-hours.

If the candidate only works under resources the institution does not have, it is not ready for that workflow.

STOP CONDITION 7: THE NEXT STAGE CANNOT BE CONTROLLED

A candidate may look promising historically but require a live test that the institution cannot run safely, securely or reversibly. If there is no acceptable shadow path, no approved data path or no responsible owner for the next stage, progression may need to stop until those conditions exist.

A strong historical result does not create authority by itself.

REFINE IS DIFFERENT FROM STOP

Stop does not always mean abandon the idea permanently. Sometimes the right outcome is Refine.

Refine is appropriate when the problem remains worth solving but the current candidate or evidence basis is not sufficient. The team might need to narrow the scope, fix data timing, modify the candidate, change the metric, obtain more mature outcomes or define a safer next stage.

The distinction should be explicit:

Continue — evidence supports the next declared stage.

Refine — the question remains valid, but candidate or evidence needs work.

Stop — evidence does not justify progression on the current basis.

WHY STOP IS COMMERCIAL VALUE

A paid evaluation is not only valuable when a candidate succeeds. Avoiding a weak implementation can save engineering effort, integration cost, review time, vendor spend and governance burden.

This is why a fixed-scope evaluation should not be priced purely as a success fee tied to positive uplift. The decision value exists even when the correct answer is not to proceed.

HOW AEGI FRAMES STOP

AEGI Shield treats Continue, Refine and Stop as first-class outcomes of a bounded Controlled Evaluation. The objective is not to promote a candidate. It is to produce reviewable evidence for one decision.

Where the evidence shows that the baseline is adequate, the candidate is weak or the evaluation basis is insufficient, the result should say so directly. AEGI does not acquire production authority merely because it performed the evaluation.

FREQUENTLY ASKED QUESTIONS

Is Stop a failed pilot?

Not necessarily. If the evaluation correctly prevents an unsupported change from progressing, it has served its decision purpose.

What if executives already want the candidate deployed?

The evaluation should still record the declared criteria, evidence and limitations. Governance value is highest when the process can make disagreement visible rather than reverse-engineer a positive conclusion.

Can a stopped candidate be evaluated again?

Yes, if the candidate, evidence or decision context changes materially. The new evaluation should be treated as a new basis rather than silently reusing the old result.

Should every small guardrail breach cause Stop?

No. Materiality and remediation rules should be declared according to the workflow and institutional risk appetite.

CONCLUSION

The ability to stop is a sign of a mature AI evaluation process. It keeps technical enthusiasm separate from institutional authority and ensures that evidence, not project momentum, determines progression.

NEXT STEP

If your team has one proposed AI or risk-workflow change but is genuinely uncertain whether it should progress, AEGI can first assess whether the baseline, candidate and decision criteria are sufficiently bounded for a Controlled Evaluation.

CLAIM BOUNDARY

This article is educational and does not define regulatory or model-risk approval criteria. Stop, Refine and Continue thresholds are institution-specific and should reflect the applicable use case, policy, risk appetite and governance framework.

Related AEGI resources

NEXT STEP

Bring one proposed change.

Controlled Evaluation Bring One Question