“It probably flags everything.”
A legitimate feature with a larger importance score than the leaked one is audited on every run, and must come back clear. It does.
Hindsight release audit · run 20260728T175344-81df34
· case credit-default-leaked-payment-event
· verdict confirmed
· decision BLOCK
Lending · The hard case
The same bank, but the defect is realistic: only about one application in seven has post-decision data attached, so the model barely looks suspicious.
The record this page argues from
Every number above is read from this. Nothing is computed in the browser.
Loading the record...
What is being audited
synthetic datasetModel
credit_default_v1_leaky
The version being considered for release
Its id in DataHub
urn:li:mlModel:(urn:li:dataPlatform:mlflow,hindsight.credit_default_v1_leaky,PROD)
Feature examined
days_since_last_payment
built by the pipeline
feature_pipeline_leaky — one column out of everything the model was shown
Where the feature came from
| Step | Table | Column | Status |
|---|---|---|---|
| 1 | payment_events_after_decision |
payment_recorded_at |
rows appear after the decision |
| 2 | feature_pipeline_leaky |
days_since_last_payment |
knowable at decision time |
| 3 | credit_default_v1_leaky |
— | knowable at decision time |
In plain English
It scored solid in testing. When we rebuilt it using only information that existed at decision time, it dropped to solid. About 24% of its apparent skill came from seeing the future.
Imagine a student who aces a practice exam. Impressive - until you notice the answer sheet was sitting on the desk. The score was real; the ability was not. Retake the exam without the answers and you learn what they actually know.
How Hindsight knew
Nobody told Hindsight which feature was suspicious. It read
DataHub — the catalog that records every
table, column and pipeline in the company, and
which column was built from which. Following
that map backwards from the model showed one feature drawing on
payment_events_after_decision — a table whose rows
only appear after the moment the loan is approved or declined.
Running on metadata recorded from a real DataHub instance. Connect DataHub to publish the finding back into the catalog.
What actually happened
The problem, on a calendar
Everything left of the red line was knowable when the decision had to be made. Everything right of it did not exist yet. Watch the bottom row cross the line.
the decision itself
all of it existed before the decision
stops at the decision - cleared for release
still collecting data 31 days after the decision
| Asset | Column | Availability | Status |
|---|---|---|---|
applications_at_decision_time |
prediction_time |
ends at the cutoff | Source |
customer_history_point_in_time |
prior_delinquencies |
ends at the cutoff | Source |
feature_pipeline_safe |
prior_delinquencies |
ends at the cutoff | Clear |
feature_pipeline_leaky |
days_since_last_payment |
extends 31 days past the cutoff | Violation |
How certain is this?
Either one is enough on its own. Hindsight reports which fired rather than blending them, so you can see exactly what settled it.
Why you can believe this
“It probably flags everything.”
A legitimate feature with a larger importance score than the leaked one is audited on every run, and must come back clear. It does.
“It can only ever say no.”
One scenario is expected to pass, and does — the same model after the repair. A gate with no “yes” is not a gate.
“The numbers were tuned to look good.”
The data generator is frozen and its seed committed. Whatever it produces is published — including a confirmation policy that was missed by 5.1 points, disclosed rather than hidden.
“An LLM decided this.”
It cannot. A language model may explain the evidence; reaching a verdict runs through deterministic code with no model in the path.
Why nobody noticed
This is what leakage actually looks like in the wild. Nobody notices a model that scores a little better than expected; they ship it. The statistical test alone is too weak here - it is the query itself that gives the defect away.
What it costs
Obvious leaks get caught eventually. This is the kind that reaches production and quietly underperforms for a year.
What to do now
Pre-production release audit
Every claim above traces to a measurement below. Nothing here is inferred by a language model.
Release decision
BLOCK Deterministic verdict: confirmedWhat the agent actually did
Every DataHub call, SQL check and validation step, in order. Entries marked LIVE hit a real DataHub instance during this request; RECORDED entries replay responses captured from a real instance so this page works without Docker.
Directional column lineage
The timeline above the fold shows this on a calendar. Here is the exact finding from the transformation itself:
The post-outcome source is referenced, but no directional available_at <= prediction_time cutoff was found. The transformation does not reference available_at.
Calibration regression fixture ยท constant across scenarios
This controlled regression pair is deliberately held constant across audits: it proves the gate does not confuse importance with illegal availability.
The longer bar is the safe one. Hindsight blocks the shorter feature and clears the longer, because importance measures whether a feature is useful — never whether the information was allowed to exist yet. A detector built on ablation flags exactly the wrong one here.
| Feature | Ablation delta | Verdict |
|---|---|---|
days_since_last_payment |
0.190616 | confirmed |
prior_delinquencies |
0.234114 | clear_for_release |
Counterfactual test
24.5%
of the apparent feature advantage disappeared
AUC 1.000000 is expected: this synthetic planted leak is total by construction. Real leakage can be subtler, and the generator is frozen rather than tuned to look realistic.
False-positive defence
The legitimate prior_delinquencies feature remains predictive because it existed
before the decision. It must come back clear on every run.
Inspectable execution
Smallest safe repair
payment.available_at <= application.prediction_time
The proposal is verified independently. Hindsight never merges or applies pipeline code automatically.
Human approval boundary
Dry-run is the default. Explicit approval writes the field tag, structured verdict, linked audit Document and active incident — then re-reads every one to prove it persisted.
Write-back is disabled on the public demo. Against a live instance this form publishes four records once a person ticks the approval box: a tag on the offending column, the verdict as a structured property, an audit document, and an open incident. Each one is read back afterwards, and a write that cannot be re-read is treated as a failure.
Clone the repository and run uv run hindsight serve against your own
DataHub to exercise it. The example record it produces is in
examples/audit_document.md.