Overview / Runs / 20260728T175344-81df34
Recorded from DataHub Core

Lending · The hard case

Will this applicant repay the loan?

The same bank, but the defect is realistic: only about one application in seven has post-decision data attached, so the model barely looks suspicious.

Download

What is being audited

synthetic dataset

Model

credit_default_v1_leaky

The version being considered for release

Its id in DataHub urn:li:mlModel:(urn:li:dataPlatform:mlflow,hindsight.credit_default_v1_leaky,PROD)

Feature examined

days_since_last_payment

built by the pipeline feature_pipeline_leaky — one column out of everything the model was shown

Decision made 10 Jan 2026, 09:00 UTC Nothing after this was knowable
Feature's data arrived 10 Feb 2026, 09:00 UTC The feature reached past the moment the decision was made.

Where the feature came from

  1. Source table payment_events_after_decision payment_recorded_at rows appear after the decision
  2. Feature feature_pipeline_leaky days_since_last_payment
  3. Model credit_default_v1_leaky
Table view of this path
StepTable ColumnStatus
1 payment_events_after_decision payment_recorded_at rows appear after the decision
2 feature_pipeline_leaky days_since_last_payment knowable at decision time
3 credit_default_v1_leaky knowable at decision time

Records re-tested 4,000 Audited 28 Jul 2026, 17:53 UTC Run 20260728T175344-81df34

In plain English

This model was cheating.

It scored solid in testing. When we rebuilt it using only information that existed at decision time, it dropped to solid. About 24% of its apparent skill came from seeing the future.

BLOCK confirmed 24% of the skill was borrowed

The simplest way to think about it

Imagine a student who aces a practice exam. Impressive - until you notice the answer sheet was sitting on the desk. The score was real; the ability was not. Retake the exam without the answers and you learn what they actually know.

The exam
Will this applicant repay the loan?
The answer sheet
The same payment history problem - but it only reaches a minority of records, so the model's score barely moves. It looks like an ordinary, slightly-good model.
Retaking it fairly
Rebuild the data as it looked on the day, then re-test

How Hindsight knew

Nobody told Hindsight which feature was suspicious. It read DataHub — the catalog that records every table, column and pipeline in the company, and which column was built from which. Following that map backwards from the model showed one feature drawing on payment_events_after_decision — a table whose rows only appear after the moment the loan is approved or declined.

Running on metadata recorded from a real DataHub instance. Connect DataHub to publish the finding back into the catalog.

What actually happened

  1. Day 0 Thousands of applications are decided. Everything the model may know stops here.
  2. Later Only some customers make a payment that gets joined back in. Roughly one record in seven is contaminated - not all of them.
  3. Testing The model scores 0.88 instead of 0.83. A small, plausible improvement. Nobody questions it.
  4. Audit The statistical test does not fire, but the query proves it anyway. This is why there are two independent routes to a verdict.

The problem, on a calendar

One fact arrived after the decision

Everything left of the red line was knowable when the decision had to be made. Everything right of it did not exist yet. Watch the bottom row cross the line.

applications_at_decision_time prediction_time

the decision itself

customer_history_point_in_time prior_delinquencies

all of it existed before the decision

feature_pipeline_safe prior_delinquencies

stops at the decision - cleared for release

feature_pipeline_leaky days_since_last_payment
+31 days

still collecting data 31 days after the decision

Stays before the decision Feature under audit Reaches past the decision
Timeline of data availability relative to the prediction cutoff on 10 Jan 2026. The table below repeats every value.
Table view of the timeline
Asset Column Availability Status
applications_at_decision_time prediction_time ends at the cutoff Source
customer_history_point_in_time prior_delinquencies ends at the cutoff Source
feature_pipeline_safe prior_delinquencies ends at the cutoff Clear
feature_pipeline_leaky days_since_last_payment extends 31 days past the cutoff Violation

How certain is this?

We read the code that built it The code that builds the feature reads from a table whose rows only appear after the decision, with nothing stopping it. That is proof, not an estimate — no retraining needed.
We rebuilt it and re-tested Performance fell when future data was removed, but not far enough to cross the configured confirmation policy. This route is supporting evidence only.

Either one is enough on its own. Hindsight reports which fired rather than blending them, so you can see exactly what settled it.

Why you can believe this

Four ways this tool could be fooling you, and what stops each

“It probably flags everything.”

A legitimate feature with a larger importance score than the leaked one is audited on every run, and must come back clear. It does.

“It can only ever say no.”

One scenario is expected to pass, and does — the same model after the repair. A gate with no “yes” is not a gate.

“The numbers were tuned to look good.”

The data generator is frozen and its seed committed. Whatever it produces is published — including a confirmation policy that was missed by 5.1 points, disclosed rather than hidden.

“An LLM decided this.”

It cannot. A language model may explain the evidence; reaching a verdict runs through deterministic code with no model in the path.

Why nobody noticed

This is what leakage actually looks like in the wild. Nobody notices a model that scores a little better than expected; they ship it. The statistical test alone is too weak here - it is the query itself that gives the defect away.

What it costs

Obvious leaks get caught eventually. This is the kind that reaches production and quietly underperforms for a year.

What to do now

  1. Do not release this model version.
  2. The same one-line repair. The size of the leak does not change the fix.
  3. Re-run the audit; the score you get then is the real one.

Pre-production release audit

The evidence behind that conclusion

Every claim above traces to a measurement below. Nothing here is inferred by a language model.

Release decision

BLOCK Deterministic verdict: confirmed
Release decision BLOCK verdict: confirmed
Records tested 4,000 4,000 excluded post-cutoff
Advantage lost 24.5% under point-in-time rebuild
Audit runtime 0.133s credit-default-leaked-payment-event

What the agent actually did

Backend activity

11 operations

Every DataHub call, SQL check and validation step, in order. Entries marked LIVE hit a real DataHub instance during this request; RECORDED entries replay responses captured from a real instance so this page works without Docker.

Directional column lineage

One feature reached across the decision

Temporal violation

The timeline above the fold shows this on a calendar. Here is the exact finding from the transformation itself:

The post-outcome source is referenced, but no directional available_at <= prediction_time cutoff was found. The transformation does not reference available_at.

Calibration regression fixture ยท constant across scenarios

Importance gets this exactly backwards

This controlled regression pair is deliberately held constant across audits: it proves the gate does not confuse importance with illegal availability.

Planted leaked feature post-outcome feature
0.21
Legitimate control pre-cutoff control
0.24

The longer bar is the safe one. Hindsight blocks the shorter feature and clears the longer, because importance measures whether a feature is useful — never whether the information was allowed to exist yet. A detector built on ablation flags exactly the wrong one here.

Table view of the comparison
FeatureAblation deltaVerdict
days_since_last_payment 0.190616 confirmed
prior_delinquencies 0.234114 clear_for_release

Counterfactual test

Point-in-time reconstruction

Supporting evidence
Observed training AUC0.880428
Honest reconstructed AUC0.833788

24.5%

of the apparent feature advantage disappeared

AUC 1.000000 is expected: this synthetic planted leak is total by construction. Real leakage can be subtler, and the generator is frozen rather than tuned to look realistic.

False-positive defence

Strong signal, still safe

Clear

The legitimate prior_delinquencies feature remains predictive because it existed before the decision. It must come back clear on every run.

Observed and reconstructed AUC0.923926
Ablation delta0.234114
Advantage retained100%
Verdictclear_for_release

Inspectable execution

Evidence trace

  1. 01 Transformation Verification violation
  2. 02 Point In Time Reconstruction failed
  3. 03 Deterministic Verdict confirmed
  4. 04 Safe Control clear for release
  5. 05 Writeback awaiting human approval

Smallest safe repair

Cut off future knowledge

SQL verified
payment.available_at <= application.prediction_time

The proposal is verified independently. Hindsight never merges or applies pipeline code automatically.

Human approval boundary

Publish evidence to DataHub

Dry-run is the default. Explicit approval writes the field tag, structured verdict, linked audit Document and active incident — then re-reads every one to prove it persisted.

Write-back is disabled on the public demo. Against a live instance this form publishes four records once a person ticks the approval box: a tag on the offending column, the verdict as a structured property, an audit document, and an open incident. Each one is read back afterwards, and a write that cannot be re-read is treated as a failure.

Clone the repository and run uv run hindsight serve against your own DataHub to exercise it. The example record it produces is in examples/audit_document.md.