Model Training and Testing
Understand how Athia payment models are trained, evaluated, served, experimented and promoted.
On this page
Athia treats a trained model as a candidate, not a production decision. Data quality, offline evaluation, artifact compatibility, serving readiness and controlled live evidence must all pass before a model receives meaningful traffic.
Two systems, one lifecycle#
Athia separates offline model development from online model operations.
| Responsibility | Data and training pipelines | Athia platform |
|---|---|---|
| Prepare training data | Query governed Snowflake data, validate it and build model-specific features | Track the dataset and training context exposed with a model version |
| Train and evaluate | Train candidates, run temporal or cohort-aware evaluation and package artifacts | Register model versions and evaluation evidence |
| Test compatibility | Validate feature schemas, encoders, metadata and artifact contracts | Check loadability, health and serving readiness |
| Run experiments | Produce candidate artifacts and expected metrics | Assign shadow or controlled traffic, collect outcomes and compare variants |
| Promote and operate | Reproduce training and retain evidence | Promote a winner, hot-reload or deploy it, monitor it and roll back when required |
The training code and Snowflake-facing pipelines live in DATA-Athena-Snowflake. Registry, experiment, serving and runtime controls live in athena-platform. This boundary keeps access to training data separate from the services that make online predictions.
From governed data to a production model#
Payment models supported by the lifecycle#
The same lifecycle can be used for several decision families, depending on account configuration and availability:
Some decision families use a single predictor; others train several related outputs or combine model evidence with eligibility and safety rules.
What is tested#
1. Data gates
Before training starts, the pipeline checks that the source is usable: freshness, minimum volume, required fields, label availability, class balance and feature distributions. Time-aware splits and leakage checks prevent future information from entering past examples.
2. Offline model quality
Evaluation is specific to the decision. It can include discrimination, precision and recall, calibration, ranking quality, error by cohort and the expected business effect. Temporal holdouts show whether a candidate generalizes beyond the period that trained it.
Replay and policy-aware tests compare model decisions with historical or live-observed outcomes where the data supports a valid comparison. No single metric is sufficient for promotion.
3. Artifact and feature contracts
The model package includes the model artifact plus metadata required to reproduce its inputs and outputs. Tests validate feature order and types, encoders, supported model identifiers, output schema and compatibility with the serving runtime.
4. Serving readiness
A registered version must load successfully, expose health and readiness, respond to representative requests and meet latency and error expectations. Readiness is separate from model quality: a strong candidate that cannot be served safely does not advance.
5. Live evidence
Shadow mode records what the model would have decided without changing the transaction. Controlled experiments then assign an agreed share of eligible traffic to a candidate and compare it with a baseline. Guardrails watch both payment outcomes and operational health.
Promotion gates#
| Gate | Evidence required |
|---|---|
| Data | Valid, sufficiently fresh and representative training input |
| Offline evaluation | Model-specific thresholds plus acceptable cohort behavior and calibration |
| Contract | Reproducible artifacts, metadata, feature schema and load test |
| Serving | Healthy deployment with acceptable latency and error behavior |
| Experiment | Controlled evidence against the active baseline and no breached guardrail |
| Approval | Authorized promotion decision with a retained audit trail |
A promoted model remains reversible. Traffic weights can be reduced, the prior model can be restored, and shadow mode can be used again while a new candidate is investigated.
What you see in Athia#
The model operations area brings together: