The Slow Poison: Why Your AI Gets Worse Every Week
E1

The Slow Poison: Why Your AI Gets Worse Every Week

The Slow Poison: Why Your AI Gets Worse Every Week

This is Part 7 of our 16-week series on Context Degradation—the hidden failure modes that break AI systems before anyone notices. Zillow's $881 Million Lesson

In 2021, Zillow shut down its iBuying division and laid off 25% of its workforce.

The reason: their home pricing algorithm had systematically overvalued properties. Zillow bought houses at prices higher than they could sell them. They lost $881 million in a single quarter.

The algorithm wasn't always wrong. It was trained on years of housing data. It performed well in backtesting. It worked in early deployment.

Then the market shifted. And the algorithm didn't notice.

What Went Wrong

Zillow's Zestimate algorithm was trained on historical housing transactions. In a stable market, this works reasonably well—past sales predict future prices.

But 2021 wasn't stable:

Pandemic-driven relocations changed demand patterns Remote work shifted preferences toward different housing types Supply chain issues affected new construction Interest rate expectations created buying pressure Unprecedented price appreciation in some markets

The features that predicted prices in 2019 didn't predict prices in 2021. The relationships had shifted. The model was confident. The confidence was misplaced.

DRIFT: Reliability Decay Over Time

In our degradation taxonomy, DRIFT is specifically:

Declining ω (omega): Reliability decreasing over time Stable apparent performance: Until the gap becomes catastrophic

The signature of drift is that it's invisible until it's catastrophic. The model keeps producing outputs. The outputs look reasonable. But they're increasingly disconnected from reality.

Drift happens because the world changes and models don't:

Training data ages User behavior evolves Market conditions shift Regulations update Competitors adapt

Static models in dynamic worlds drift toward irrelevance.

The Two Stages of Drift

Drift isn't sudden. It's gradual—which makes it harder to detect.

Stage 1: Silent Degradation

The model continues performing within acceptable parameters on your monitoring metrics. But the relationship between predictions and reality is slowly decoupling.

You don't notice because:

Individual predictions still look plausible Aggregate metrics average out errors You're measuring what you measured at deployment The drift is too slow to trigger alerts

Stage 2: Catastrophic Visibility

At some point, degradation crosses a threshold. Errors compound. Losses accumulate. What was invisible becomes undeniable.

For Zillow, this happened when they realized they owned billions of dollars in overpriced inventory.

Why Standard Monitoring Misses Drift

Most ML monitoring focuses on:

Model metrics: Accuracy, precision, recall, F1 Infrastructure metrics: Latency, throughput, errors Feature drift: Statistical shifts in input features Concept drift: Changes in the target relationship

These help but have blind spots:

Metric lag: By the time accuracy drops measurably, you've already made many bad decisions.

Ground truth delay: For predictions about future events (home prices, loan defaults), you don't know you're wrong until the future arrives.

Threshold blindness: Gradual degradation doesn't trigger alerts designed for sudden failures.

Distribution blindness: Feature drift detection catches obvious shifts, not subtle changes in correlation structure.

Zillow's Specific Failure

Zillow had sophisticated monitoring. They had data science teams. They had executives asking questions.

What they lacked was a mechanism to detect reliability drift separate from prediction drift.

The model's predictions weren't obviously wrong. A house valued at $400K selling for $380K isn't a red flag in isolation. But systematic overvaluation of 5-10% across thousands of homes adds up.

The reliability of the model—its omega—was declining. But they were measuring accuracy on old data, not reliability in the current market.

What a Certificate Would Have Caught

A Context Quality Certificate tracks omega over time. Declining omega signals drift before it becomes catastrophic.

For Zillow, the certificate would have shown:

Omega trending downward: Model reliability decreasing over weeks/months Alpha-omega gap widening: Confidence staying high while reliability dropped Temporal anomaly: Recent predictions performing worse than older ones

These signals enable intervention:

Pause or slow down buying decisions Require additional verification for high-value properties Trigger model retraining or recalibration Adjust bidding margins to account for uncertainty

The key is continuous measurement of reliability, not just periodic retraining.

The Broader Pattern

Zillow's failure was expensive and public. But drift affects every deployed model:

Recommendation systems: User preferences evolve. Content catalogs change. Models trained on last year's behavior recommend for last year's users.

Fraud detection: Fraudsters adapt. What caught fraud in January doesn't catch fraud in December.

Credit scoring: E…

Read the full article →