Heelius closes an L1 alert for $2.80 against a $32.90 loaded human cost

Model card · heelius-investigator-4.2

What it does, how it was evaluated, and where it fails

Published because a compliance team cannot adopt a model it cannot explain to a regulator — and because the limitations section is the part that matters.

Model architecture

A pipeline, not a prompt

A language model alone cannot pass a model-risk review. Heelius separates retrieval, graph inference, scoring and generation into four governed layers, so each one can be evaluated, versioned and challenged on its own terms.

RetrievalL1

Deterministic evidence layer

Typed connectors to your core, TM system, KYC store, card and wallet ledgers. Retrieval is deterministic and replayable — the same alert pulls the same records a year later, which is what makes a case defensible on examination.

Typed schemaReplayableNo free-text search
Graph MLL2

Entity resolution & network features

A learned linkage model resolves counterparties across devices, addresses, registrations and payment references, then extracts network features — circularity, fan-in/fan-out asymmetry, cluster co-registration, hop distance to known typologies.

Linkage model3-hop expansionTypology proximity
ReasoningL3

Weighted disposition model

Findings are scored as aggravating or mitigating against your institution's own thresholds, not a vendor default. The score is calibrated, so a stated 75% confidence means the disposition holds roughly 75 times out of 100 on held-out reviewed cases.

CalibratedPolicy-conditionedAuditable weights
GenerationL4

Constrained narrative synthesis

Language generation is confined to composing retrieved facts. A verification pass re-checks each assertion against source records and rejects any sentence that cannot be grounded — the failure mode is a blank section, never an invented one.

Citation-boundVerifier passNo free assertion

Evaluation · held-out adjudicated set

n = 41,200 · refreshed quarterly · champion/challenger

Precision on escalation
0.94
of escalations, share a human also escalated
Recall on true suspicion
0.991
of human-confirmed suspicion the agent surfaced
False-negative rate
0.9%
held-out set of 41,200 adjudicated alerts
Calibration error
0.021
expected calibration error, 10-bin
Narrative groundedness
99.8%
assertions traceable to a retrieved record
Median latency
155s
intake to released narrative

Every model version ships with a model card, a documented evaluation protocol and a challenger held in parallel. Drift on population stability, disposition mix and QA concurrence is monitored per typology and per corridor, with automatic reversion to human routing when a segment moves outside its control band.

Intended use

Heelius investigates alerts raised by a transaction-monitoring system and produces a recommended disposition with a supporting narrative. It is a decision-support system. The filing determination remains with the institution and its MLRO, and the console records who made it.

It is designed for L1 and L2 alert investigation. It is not designed to tune monitoring thresholds, to score customers at onboarding, or to make a decision to exit a relationship.

Evaluation protocol

The held-out set is 41,200 alerts adjudicated by human analysts across seven institutions and eleven jurisdictions, stratified by typology, channel and disposition. It is refreshed quarterly and never used for model selection — a separate development set carries that load.

Recall on true suspicion is the governing metric: the share of human-confirmed suspicious activity the agent also surfaced. Precision on escalation is reported alongside it, because an agent that escalates everything is trivially safe and commercially useless.

Narrative quality is scored on groundedness — the proportion of asserted facts that resolve to a retrieved record — and on a blind rubric applied by former examiners who do not know which narratives were machine-drafted.

Known limitations

Retrieval quality bounds everything downstream. Where a customer file is thin, stale or contradictory, the agent's confidence falls and the case routes to a human — but it cannot manufacture evidence the institution does not hold.

Novel typologies are, by construction, under-represented in the evaluation set. New corridors and products start under human routing and are released to auto-close only after a segment-level concurrence review.

Adverse media in low-resource languages has thinner source coverage. The agent reports coverage explicitly rather than treating an empty result as a clean one.

The agent does not resolve sanctions alias ambiguity. Where identifiers diverge but a name matches strongly, it escalates — always. That behaviour is a policy gate, not a score, and it is not tunable downward.

Model risk governance

Aligned to SR 11-7 and the equivalent EU and UK expectations. Each release ships a model card, a documented evaluation, a change log and a validation pack sized for an independent review function.

A challenger model runs in shadow on live traffic. Promotion requires non-inferior recall on the held-out set and MLRO sign-off, recorded in the console's governance view.

Drift & monitoring

Population stability, disposition mix, confidence distribution and QA concurrence are monitored per typology and per corridor. A segment that leaves its control band reverts to human routing automatically and raises a governance event.

Reviewer overturns are treated as a first-class signal. Each one is reviewed for whether it reflects a model error, a policy ambiguity or a genuine disagreement, and the classification is recorded.

Data handling

Customer data is never used to train or fine-tune shared models. Institution-specific adaptation happens through configuration, policy packs and reviewer feedback held inside that institution's tenant.

Evidence retrieved during an investigation is retained in the case file for the institution's stated retention period and is deleted on the same schedule as the underlying case.