Zillow Offers algorithm failure: $407.9m, and a court order says management overrode the model

Zillow Offers is the most cited AI failure in business. A company let an algorithm price houses, the algorithm got it wrong, and the company wrote off a fortune and fired a quarter of its staff.
I went to the primary sources for a working paper, and the story does not hold. Not because the numbers are wrong. Because the decisive act was not the model's.
What Zillow disclosed
On 2 November 2021 Zillow announced the wind-down. From its own filings: an inventory write-down of about $304 million in the third quarter, attributed to "purchasing homes at higher prices than current estimated future selling prices". The audited full-year 2021 write-down was $407.9 million, of which $211.0 million was on homes still held at year end. The Homes segment lost $881 million before income taxes across 2021, against $320 million in 2020. And a workforce reduction of approximately 25%.
Cite the percentage, not a headcount. Zillow never published an absolute number. One outlet reported 2,000, another 1,600 out of 6,400. Both cannot be right, and neither is Zillow's figure.
The arithmetic that was public the whole time
Here is the part that needed no inside knowledge.
On 25 February 2021 Zillow published a release titled "Zillow Starts Making Cash Offers For the Zestimate". The Zestimate became the opening cash offer in more than 20 markets. The chief operating officer, Jeremy Wacksman: "This exciting advancement demonstrates the confidence we have in the Zestimate."
Zillow also published its own error rates: a median error of 1.9% for on-market homes and 6.9% for off-market homes. Zillow Offers bought off-market homes.
A 6.9% median error means half of all offers were wrong by more than 6.9%. On a $400,000 house that is more than $27,600, against a target margin in the low single digits. At its advertised accuracy, with the market perfectly flat, the model's typical error was already larger than the entire business margin. And the Zestimate estimates value now, while the balance sheet was exposed three to six months forward.

That was on Zillow's own website. Anyone could have done that sum.
The comparison, and the correction it forced
There was a control group, which almost no AI failure has. Opendoor was in the same market, in the same quarter, with a larger book.
I got this comparison wrong on the first attempt, and how I got it wrong is more useful than the fixed version.
I had written that Opendoor recorded $39 million of inventory valuation adjustment across all of 2021 against Zillow's $304 million in a single quarter, and that the market therefore was not the variable. Three problems. The two figures are not the same kind of number: Opendoor's $39 million is a non-GAAP measure covering homes still held at period end, and its audited 2021 figure is $56 million. Opendoor did not record nothing in the quarter that killed Zillow, it recorded $31 million. And the market caught Opendoor twelve months later: $737 million of inventory valuation adjustments in 2022 and a net loss of $1,353 million, with $573 million and a $928 million loss in the third quarter alone.
Measured against the book each was marking, Opendoor at the end of 2022 was in a worse position than Zillow at the end of Q3 2021. Mike DelPrete, who produced the original pricing analysis, wrote a year later that "on a per home basis, each company incurred similar losses".

So the market was the variable. It just arrived at the two companies at different times.
I made that error in a specific direction: the one that produced the tidier sentence. That is the same direction every error in this literature travels, which is the whole subject of the paper this comes from.
What actually separates them
The two failures were not the same failure, and the difference is in a federal court order.
Ruling on the securities litigation against Opendoor in February 2024, the court distinguished the parallel case against Zillow, and in doing so described what Zillow had done. Zillow "sought to rapidly scale up its iBuying operation in an attempt to catch up to its competitors", and to do it "devised a plan that it called 'Project Ketchup.'" Under that plan, Zillow "applied systematic overlays to drive up offers well above the pricing indicated by its algorithm and pricing analysts."
The court's own restatement: "Zillow purposefully raised its offer prices well beyond those generated by its algorithm."
Read that again. The model produced a price. A human decision then raised it, on purpose, systematically, to win market share.
Zillow's model was not overruled by the market. It was overruled by management, upward, deliberately. Opendoor's later and larger loss is what the same downturn does to a book nobody overrode.
Why this matters more than the write-down
Every retelling of this case treats it as evidence about algorithms. It is evidence about the step after the algorithm.
The decisive moment was not the model's output. It was the conversion of that output into an offer the company could not take back, and the fact that a person could move the number on the way through with nothing to stop them.
That distinction changes what you would build to prevent it. If the model failed, you buy a better model. If the output was overridden on the way to committing capital, a better model changes nothing at all.

Two limits, because they are real. Opendoor's own account of its pricing cannot be used as evidence here: a court held it adequately pleaded that six of Opendoor's statements about its algorithm were "actionably misleading", and although the case settled with no admission, that contest disqualifies its self-description. And the size of an inventory write-down is itself set by the same pricing apparatus whose accuracy is the question. A small write-down is weak evidence of an accurate model, which is a general point about any AI system marking its own homework.
One claim in wide circulation should be named and dropped: that Zillow removed human review from its offers. I could find no primary or reported source. It traces to social posts and content marketing. It is a tidy explanation with nothing behind it, and the truth is nearly its opposite: the humans were very much in the loop, and they were the ones pushing the number up.
This is the second of several pieces drawn from a working paper I published this week, which traces nineteen headline AI failure statistics to their primary sources and works through seven adverse cases. The full paper, with every source stated and every claim checked, is here. The first piece, on where "80% of AI projects fail" actually comes from, is here.