katsuma onishi

Used Car Price Analysis, Take Two

Project · August 2026 · Python, statsmodels

Read the write-up ↗

My first linear regression project, done again properly. Same craigslist data, but de-duplicated, kept out to $125,000 instead of truncated at $44,500, and modeled in log-price — the transformation chosen by Box-Cox rather than by deleting the histogram's tail. Every assumption gets checked, every claim carries a confidence interval, and every headline number is measured on 14,391 listings the model never saw.

What improved

On identical held-out data, the 2025 specification's typical error is $3,032; the rebuilt model's is $1,835. Its 95% prediction intervals cover the true price 95.3% of the time — uncertainty that means what it says. The write-up re-scores the three autotrader cars from the original, including the Lexus IS F the old model missed by 31%, and is honest about what a brand-level model still can't see: the dataset's one real IS F, scored leave-one-out, misses its interval by ninety-two dollars — its badge is worth more than any regression on ad features can know.

What it taught me

The techniques from STATS 3A03 aren't decoration — each one changed a conclusion. Box-Cox picked the response scale, partial F-tests justified the depreciation curve's shape, robust standard errors kept the inference honest after Breusch-Pagan rejected constant variance, and out-of-sample interval coverage was the difference between quoting an error and standing behind one.