Overview / Runs / 20260729T050141-dc73ad
Recorded from DataHub Core

Lending · Loan approval

Will this applicant repay the loan?

A bank wants to predict whether someone will repay a loan, so it can decide who to approve.

Download

What is being audited

synthetic dataset

Model

credit_default_v1_leaky

The version being considered for release

Its id in DataHub urn:li:mlModel:(urn:li:dataPlatform:mlflow,hindsight.credit_default_v1_leaky,PROD)

Feature examined

days_since_last_payment

built by the pipeline feature_pipeline_leaky — one column out of everything the model was shown

Decision made 10 Jan 2026, 09:00 UTC Nothing after this was knowable
Feature's data arrived 10 Feb 2026, 09:00 UTC The feature reached past the moment the decision was made.

Where the feature came from

  1. Source table payment_events_after_decision payment_recorded_at rows appear after the decision
  2. Feature feature_pipeline_leaky days_since_last_payment
  3. Model credit_default_v1_leaky
Table view of this path
StepTable ColumnStatus
1 payment_events_after_decision payment_recorded_at rows appear after the decision
2 feature_pipeline_leaky days_since_last_payment knowable at decision time
3 credit_default_v1_leaky knowable at decision time

Records re-tested 4,000 Audited 29 Jul 2026, 05:01 UTC Run 20260729T050141-dc73ad

In plain English

This model was cheating.

It scored perfect in testing. When we rebuilt it using only information that existed at decision time, it dropped to solid. About 55% of its apparent skill came from seeing the future.

BLOCK confirmed 55% of the skill was borrowed

The simplest way to think about it

Imagine a student who aces a practice exam. Impressive - until you notice the answer sheet was sitting on the desk. The score was real; the ability was not. Retake the exam without the answers and you learn what they actually know.

The exam
Will this applicant repay the loan?
The answer sheet
The model was given how long it had been since the customer's last payment - but it counted payments made AFTER the loan was already approved.
Retaking it fairly
Rebuild the data as it looked on the day, then re-test

How Hindsight knew

Nobody told Hindsight which feature was suspicious. It read DataHub — the catalog that records every table, column and pipeline in the company, and which column was built from which. Following that map backwards from the model showed one feature drawing on payment_events_after_decision — a table whose rows only appear after the moment the loan is approved or declined.

Running on metadata recorded from a real DataHub instance. Connect DataHub to publish the finding back into the catalog.

What actually happened

  1. Day 0 Maya applies for a loan. The bank must decide today. Everything the model is allowed to know stops here.
  2. Day 31 Maya makes her first payment. This fact did not exist when the decision was made.
  3. Later Someone gathers examples to teach the model, and attaches every payment on record. The model can now see Day 31 while pretending to sit on Day 0.
  4. Testing The model looks flawless. Of course it does. It is reading the answer.

The problem, on a calendar

One fact arrived after the decision

Everything left of the red line was knowable when the decision had to be made. Everything right of it did not exist yet. Watch the bottom row cross the line.

applications_at_decision_time prediction_time

the decision itself

customer_history_point_in_time prior_delinquencies

all of it existed before the decision

feature_pipeline_safe prior_delinquencies

stops at the decision - cleared for release

feature_pipeline_leaky days_since_last_payment
+31 days

still collecting data 31 days after the decision

Stays before the decision Feature under audit Reaches past the decision
Timeline of data availability relative to the prediction cutoff on 10 Jan 2026. The table below repeats every value.
Table view of the timeline
Asset Column Availability Status
applications_at_decision_time prediction_time ends at the cutoff Source
customer_history_point_in_time prior_delinquencies ends at the cutoff Source
feature_pipeline_safe prior_delinquencies ends at the cutoff Clear
feature_pipeline_leaky days_since_last_payment extends 31 days past the cutoff Violation

How certain is this?

We read the code that built it The code that builds the feature reads from a table whose rows only appear after the decision, with nothing stopping it. That is proof, not an estimate — no retraining needed.
We rebuilt it and re-tested With the future removed, the model lost enough of its edge to cross the configured confirmation policy.

Either one is enough on its own. Hindsight reports which fired rather than blending them, so you can see exactly what settled it.

Why you can believe this

Four ways this tool could be fooling you, and what stops each

“It probably flags everything.”

A legitimate feature with a larger importance score than the leaked one is audited on every run, and must come back clear. It does.

“It can only ever say no.”

One scenario is expected to pass, and does — the same model after the repair. A gate with no “yes” is not a gate.

“The numbers were tuned to look good.”

The data generator is frozen and its seed committed. Whatever it produces is published — including a confirmation policy that was missed by 5.1 points, disclosed rather than hidden.

“An LLM decided this.”

It cannot. A language model may explain the evidence; reaching a verdict runs through deterministic code with no model in the path.

Why nobody noticed

In testing, every application already has months of payment history attached, because the test data was collected later. On the day of a real decision, none of it exists yet.

What it costs

The bank approves people it should have declined and declines people it should have approved - and only finds out months later, one default at a time.

What to do now

  1. Do not release this model version.
  2. Rebuild the feature so it only counts payments recorded before the decision.
  3. Re-run the audit; the score you get then is the real one.

Pre-production release audit

The evidence behind that conclusion

Every claim above traces to a measurement below. Nothing here is inferred by a language model.

Release decision

BLOCK Deterministic verdict: confirmed
Release decision BLOCK verdict: confirmed
Records tested 4,000 4,000 excluded post-cutoff
Advantage lost 55.1% under point-in-time rebuild
Audit runtime 0.444s credit-default-leaked-payment-event

What the agent actually did

Backend activity

11 operations

Every DataHub call, SQL check and validation step, in order. Entries marked LIVE hit a real DataHub instance during this request; RECORDED entries replay responses captured from a real instance so this page works without Docker.

Directional column lineage

One feature reached across the decision

Temporal violation

The timeline above the fold shows this on a calendar. Here is the exact finding from the transformation itself:

The post-outcome source is referenced, but no directional available_at <= prediction_time cutoff was found. The transformation does not reference available_at.

Calibration regression fixture ยท constant across scenarios

Importance gets this exactly backwards

This controlled regression pair is deliberately held constant across audits: it proves the gate does not confuse importance with illegal availability.

Planted leaked feature post-outcome feature
0.21
Legitimate control pre-cutoff control
0.24

The longer bar is the safe one. Hindsight blocks the shorter feature and clears the longer, because importance measures whether a feature is useful — never whether the information was allowed to exist yet. A detector built on ablation flags exactly the wrong one here.

Table view of the comparison
FeatureAblation deltaVerdict
days_since_last_payment 0.301712 confirmed
prior_delinquencies 0.226554 clear_for_release

Counterfactual test

Point-in-time reconstruction

Confirmation route fired
Observed training AUC1.000000
Honest reconstructed AUC0.833630

55.1%

of the apparent feature advantage disappeared

AUC 1.000000 is expected: this synthetic planted leak is total by construction. Real leakage can be subtler, and the generator is frozen rather than tuned to look realistic.

False-positive defence

Strong signal, still safe

Clear

The legitimate prior_delinquencies feature remains predictive because it existed before the decision. It must come back clear on every run.

Observed and reconstructed AUC0.924842
Ablation delta0.226554
Advantage retained100%
Verdictclear_for_release

Inspectable execution

Evidence trace

  1. 01 Transformation Verification violation
  2. 02 Point In Time Reconstruction passed
  3. 03 Deterministic Verdict confirmed
  4. 04 Safe Control clear for release
  5. 05 Writeback awaiting human approval

Smallest safe repair

Cut off future knowledge

SQL verified
payment.available_at <= application.prediction_time

The proposal is verified independently. Hindsight never merges or applies pipeline code automatically.

Human approval boundary

Publish evidence to DataHub

Dry-run is the default. Explicit approval writes the field tag, structured verdict, linked audit Document and active incident — then re-reads every one to prove it persisted.

Write-back is disabled on the public demo. Against a live instance this form publishes four records once a person ticks the approval box: a tag on the offending column, the verdict as a structured property, an audit document, and an open incident. Each one is read back afterwards, and a write that cannot be re-read is treated as a failure.

Clone the repository and run uv run hindsight serve against your own DataHub to exercise it. The example record it produces is in examples/audit_document.md.