Model architecture
A pipeline, not a prompt
A language model alone cannot pass a model-risk review. Heelius separates retrieval, graph inference, scoring and generation into four governed layers, so each one can be evaluated, versioned and challenged on its own terms.
RetrievalL1
Deterministic evidence layer
Typed connectors to your core, TM system, KYC store, card and wallet ledgers. Retrieval is deterministic and replayable — the same alert pulls the same records a year later, which is what makes a case defensible on examination.
Typed schemaReplayableNo free-text search
Graph MLL2
Entity resolution & network features
A learned linkage model resolves counterparties across devices, addresses, registrations and payment references, then extracts network features — circularity, fan-in/fan-out asymmetry, cluster co-registration, hop distance to known typologies.
Linkage model3-hop expansionTypology proximity
ReasoningL3
Weighted disposition model
Findings are scored as aggravating or mitigating against your institution's own thresholds, not a vendor default. The score is calibrated, so a stated 75% confidence means the disposition holds roughly 75 times out of 100 on held-out reviewed cases.
CalibratedPolicy-conditionedAuditable weights
GenerationL4
Constrained narrative synthesis
Language generation is confined to composing retrieved facts. A verification pass re-checks each assertion against source records and rejects any sentence that cannot be grounded — the failure mode is a blank section, never an invented one.
Citation-boundVerifier passNo free assertion
Evaluation · held-out adjudicated set
n = 41,200 · refreshed quarterly · champion/challenger
- Precision on escalation
- 0.94
- of escalations, share a human also escalated
- Recall on true suspicion
- 0.991
- of human-confirmed suspicion the agent surfaced
- False-negative rate
- 0.9%
- held-out set of 41,200 adjudicated alerts
- Calibration error
- 0.021
- expected calibration error, 10-bin
- Narrative groundedness
- 99.8%
- assertions traceable to a retrieved record
- Median latency
- 155s
- intake to released narrative
Every model version ships with a model card, a documented evaluation protocol and a challenger held in parallel. Drift on population stability, disposition mix and QA concurrence is monitored per typology and per corridor, with automatic reversion to human routing when a segment moves outside its control band.