JournalMachine Learning

Your Classifier Is Measuring Your Sampling, Not the World

A ticket classifier at 94% offline and 71% live was not implemented wrong. The test set was drawn from the same snapshot as the training set.

21 Jun 20267 min readEvaluation · Drift · MLOps

A ticket classifier hits 94% on a held-out split and 71% in its first month live. Nobody made a mistake. The held-out set was carved at random from the same snapshot as the training set, so it inherited every bias that snapshot contained — and inherited them in exactly the proportions that flatter the model.

Random splits assume a world that has stopped

Real inputs drift, and they drift in ways a random split structurally cannot capture:

  • Vocabulary shifts as products get renamed and features get deprecated.
  • Class balance moves with the season, the release cycle, and the incident calendar.
  • Users discover the tool and start writing for a machine, changing the input distribution because the model exists.

That last one is the interesting failure. Deploying the model changes the data the model sees. No offline split anticipates it.

Split on time instead

# optimistic: leaks future vocabulary and label balance backwards
train, test = train_test_split(df, test_size=0.2, random_state=0)

# honest: train on the past, evaluate on the future
cutoff = df.timestamp.quantile(0.8)
train = df[df.timestamp <= cutoff]
test  = df[df.timestamp >  cutoff]
A one-line change that buys an honest number

If accuracy falls sharply under a temporal split, that lower number is the one to plan around. It is not a worse measurement — it is a measurement of the thing you actually care about.

Abstention beats accuracy

The operational fix is rarely another two points offline. It is instrumenting confidence and routing the low-confidence tail to a human.

In a service desk, a classifier that knows when to abstain is worth more than one that is marginally more accurate and always certain.

Abstention also gives you a live drift signal for free: when the share of low-confidence predictions climbs, the world has moved before your accuracy metrics catch up.