Insight AEGI Shield

AEGI BAF-003 Evidence Note: What 494 → 544 at Fixed Review Capacity Shows — and Does Not Show

A public evidence note on AEGI's frozen BAF-003 Shared-Context evaluation: fixed capacity, measured queue-composition difference, reproducibility, and the limits of what can be claimed.

QUICK ANSWER

In AEGI’s frozen public-synthetic BAF-003 evaluation, the review capacity stayed fixed at 2,417 Top-1% review slots. The control placed 494 fraud-labelled applications in that review queue. The Shared-Context treatment placed 544. The net difference was +50, equivalent to +20.7 per 1,000 review slots. The reported 95% confidence interval for the normalised difference was +9.1 to +32.3 per 1,000, and the frozen rerun reproduced exactly.

The correct interpretation is narrow: under the declared public-synthetic evaluation, representation choice changed which fraud-labelled applications reached the same fixed review queue.

It is not production-bank performance proof, a fraud-loss-reduction claim, a universal model-superiority claim or evidence of GXS, MAS, Feedzai or bank endorsement.

WHY A FIXED REVIEW CAPACITY MATTERS

Fraud operations are frequently capacity-constrained. A model can produce millions of scores, but investigators can review only a limited number of cases.

If two approaches are compared using different queue sizes, an apparent performance improvement can simply come from reviewing more cases.

BAF-003 therefore keeps the review budget fixed. The question is not “how many positives can we find if we increase capacity?” It is:

At the same review capacity, does the treatment change which fraud-labelled applications reach the queue?

THE PUBLIC-SYNTHETIC SOURCE

The evaluation uses a public synthetic variant from the Bank Account Fraud (BAF) dataset suite released by Feedzai researchers. BAF was created as a privacy-preserving, large-scale synthetic tabular benchmark derived from the characteristics of a real-world bank-account-opening fraud setting.

The official source repository and paper are available from Feedzai’s public BAF materials. AEGI’s use of the dataset and the results described here are AEGI-conducted; source attribution does not imply Feedzai endorsement of AEGI or the AEGI evaluation design.

WHAT WAS HELD CONSTANT?

The purpose of the frozen comparison was to isolate a representation treatment rather than compare unrelated modelling pipelines.

The declared comparison held the learning procedure and review capacity fixed while comparing:

  • Control: the baseline representation used by the frozen evaluation;
  • Shared-Context treatment: the benchmark-specific treatment adding the declared shared-context representation.

The evaluation then measured the fraud-labelled applications appearing inside the same Top-1% review budget.

THE RESULT

The headline queue counts were:

  • Fixed review slots: 2,417
  • Control fraud-labelled applications in queue: 494
  • Shared-Context treatment fraud-labelled applications in queue: 544
  • Net difference: +50
  • Normalised difference: +20.7 per 1,000 review slots
  • Reported 95% CI: +9.1 to +32.3 per 1,000

The frozen rerun reproduced exactly under the declared evidence package.

WHAT THE RESULT SUPPORTS

BAF-003 supports four limited conclusions.

1. Representation can affect queue composition.

At the same review capacity, the treatment changed which labelled cases reached review.

2. The measured difference was not created by expanding the review budget.

The Top-1% queue size stayed fixed.

3. The evaluation is useful as a supporting representation signal.

It gives a concrete reason to test comparable hypotheses in a bank-controlled historical replay rather than relying only on architecture claims.

4. The frozen result is reproducible under its declared public evaluation package.

Reproducibility strengthens the evidence about what happened in the benchmark. It does not expand the benchmark’s scope.

WHAT THE RESULT DOES NOT SUPPORT

It does not prove production fraud-loss reduction.

The dataset is public synthetic, not a customer’s live production environment.

It does not prove universal model superiority.

The result is tied to the declared benchmark, treatment, split, capacity and learning procedure.

It does not prove that every bank will see the same effect.

Different institutions have different populations, labels, fraud typologies, operational constraints and model estates.

It does not authorise production deployment.

A positive offline result is one evidence stage. Production requires institution-owned integration, security, validation, governance and approval processes.

It does not prove that Shared Risk Context is always useful.

The value of additional context must be tested against the actual workflow. In some settings the existing baseline may already be sufficient.

WHY AEGI PUBLISHES THE LIMITS WITH THE NUMBER

A benchmark becomes misleading when the headline result travels farther than the conditions that produced it.

AEGI therefore treats the number, the scope and the limitation as one evidence object:

494 → 544 at the same 2,417 review slots, in the frozen public-synthetic BAF-003 evaluation.

Removing “fixed capacity,” “public-synthetic” or the declared comparison would change the meaning of the claim.

HOW THIS CONNECTS TO A CUSTOMER EVALUATION

BAF-003 is not intended to close the customer question. It is intended to justify asking it.

The next stage is institution-specific:

  1. define one workflow;
  2. reconstruct the current baseline;
  3. define the proposed treatment;
  4. freeze the historical cohort and material constraints;
  5. compare review, control and evidence value;
  6. decide Continue, Refine or Stop.

If a comparable effect does not survive the bank’s own historical replay, the public benchmark should not be used to override that result.

FREQUENTLY ASKED QUESTIONS

Does +50 mean 50 fraud cases were prevented?

No. It means 50 more fraud-labelled applications were present in the fixed review queue under the declared public-synthetic treatment. Prevention, investigation outcomes and financial loss are different questions.

Why use fraud-labelled applications rather than accuracy?

The evaluation is capacity-oriented. It asks what reaches a constrained review queue, which is closer to an operational prioritisation question than aggregate accuracy alone.

Does the confidence interval prove production significance?

No. It quantifies uncertainty for the declared benchmark statistic. It does not create external validity for a different institution or production environment.

Is Feedzai endorsing the AEGI result?

No endorsement is claimed. Feedzai is the source of the public BAF dataset suite; AEGI designed and executed the AEGI-specific evaluation.

NEXT STEP

The practical question is whether a comparable representation effect exists in a real institution’s bounded workflow under its own evidence and review constraints.

See the bank-controlled evaluation path →

CLAIM BOUNDARY

BAF-003 is supporting public-synthetic representation evidence. It is not customer validation, production performance proof, fraud-loss-reduction evidence, regulator approval, source-provider endorsement or a universal claim about Shared Risk Context.

SOURCES

RELATED AEGI RESOURCES

NEXT STEP

Bring one proposed change.

Controlled Evaluation Bring One Question