katsuma onishi

Analyzing and Predicting Used Car Prices, Take Two — The Technical Write-Up

project · august 2026 · python, statsmodels · back to the short version · source on github · the 2025 version

this is the full, assumption-by-assumption version of the write-up; the short version tells the same story in a five-minute read. in march 2025 i built my first linear regression project: a model that predicts what a used car should sell for, trained on craigslist listings. it worked, and i was proud of it. it also cut every corner i didn't yet know existed. this is the same project done again with the full toolkit from my mcmaster stats courses (stats 3a03, math 3mb3) — the transformation chosen by box-cox instead of by truncating the data, every assumption checked instead of assumed, every claim carrying a confidence interval, and every headline number measured on listings the model never saw.

listings modeled
72,120
de-duplicated, from 426,880 raw
r² (held-out, log price)
0.78
2025 version: 0.69, in-sample
median error
$1,835
2025 spec on same data: $3,032
95% interval coverage
95.3%
measured on 14,405 held-out cars

contents

  1. why do it again
  2. the data, cleaned honestly
  3. what the data looks like
  4. letting box-cox choose the transformation
  5. three models, two nested tests
  6. checking the assumptions
  7. how well it predicts
  8. what moves the price
  9. the depreciation curve
  10. six cars, including my dream car
  11. what i'd still improve
  12. reproduce it
  13. references

why do it again

the 2025 version got the big things right: real data, a sensible set of predictors, a train/test split, and honest writing about where it failed. but rereading it after finishing my degree, three things stood out.

2025 versionthis version
responseprice in dollarslog(price), chosen by box-cox
duplicate listingskeptremoved (171,428 rows)
luxury carsdeleted above $44,500kept to $125,000
classic carsmixed in (back to 1980)scoped out (1996+); they appreciate, not depreciate
diagnostics—residuals, q-q, scale-location, leverage, cook's distance, vif, breusch-pagan
inferencep-values from default standard errorsheteroskedasticity-robust (hc1) standard errors, intervals and tests
model choiceone modelthree nested models compared with formal tests, aic, and cross-validation
uncertaintyone mae number95% prediction intervals, coverage verified out of sample
headline numbersin-sampleheld-out test set + 5-fold cross-validation

the data, cleaned honestly

the data is the same kaggle snapshot of craigslist listings (426,880 rows, april–may 2021) scraped by austin reese. instead of one big filter, every cleaning decision is applied one step at a time and logged, so the whole funnel is auditable — the table below is generated straight from the pipeline. rows missing any of the duplicate-key fields are left alone at the first step: a duplicate you can't verify isn't a duplicate.

stepwhyrows keptremoved
raw listingskaggle export, april–may 2021426,880—
de-duplicate repostssame year/make/model/price/odometer posted in several cities is one car, not many255,452171,428
price $1,000–$125,000$0 means "call for price"; the $3.7 billion tundra is a typo. luxury cars stay in227,40828,044
odometer 100–400,000 mizero and ten-million-mile readings are data entry errors221,2666,142
model year 1996+older cars are collector cars; they appreciate, which is a different pricing regime206,88314,383
fuel gas/diesel/hybrid"other" is unlabeled; the few "electric" rows are mostly mislabeled gas cars196,8919,992
transmission automatic/manual"other" is unlabeled, not a category187,6209,271
title clean/rebuilt/salvagelien, missing and parts-only listings aren't straightforward sales183,6204,000
known body type, no busesunlabeled body types dropped; buses aren't passenger cars131,88251,738
no motorcyclesthe harley-davidsons still standing at this step are motorcycles in a car dataset131,85428
complete casesrows missing any modeling variable72,14759,707
brands with 100+ listingscan't estimate a brand effect from 14 alfa romeos, 7 ferraris and 6 aston martins72,12027

the biggest judgment call is complete cases: requiring cylinders, condition and drive costs 60k rows. the trade is deliberate — those variables earn their keep in the model, and the alternative (imputing or "not stated" levels) is on the future-work list.

the final sample: 72,120 listings, 34 brands, model years 1996–2022, prices $1,000 to $125,000. about 2.4% of these cars sit above the old $44,500 cap — the segment the 2025 model was never allowed to learn.

what the data looks like

asking prices are strongly right-skewed (skewness 2.46): median $9,000, mean $12,857, and a long tail of trucks and luxury cars. on a log scale the same distribution is nearly symmetric (skewness −0.03). that's the first hint that the model should think in percentages, not dollars.

two histograms: asking price is strongly right-skewed in dollars, and near-symmetric on a log scale
the same 72,120 asking prices, in dollars and in logs. skewness drops from 2.46 to −0.03.

price against age and against mileage, with price on a log scale. the median line falls roughly straight for the first decade — which on a log axis means a constant percentage drop per year — then flattens: old cars stop losing value once they're cheap.

hexbin plot of price versus vehicle age with the median falling from about 46 thousand dollars new to about 5 thousand at 20 years
price vs age at listing. each hexagon is a bundle of listings; the orange line is the median at each age.
hexbin plot of price versus odometer, median falling smoothly with mileage
price vs odometer. same shape: a steady percentage decline, flattening at high mileage.

brands sort themselves before any model is fit — porsche's median asking price is six times saturn's. (ram and gmc rank high not because trucks are luxurious but because their listings are trucks; the regression later separates the truck effect from the brand effect.)

dot plot of median asking price for 34 brands, from saturn around 4 thousand dollars to porsche around 24 thousand
median asking price by brand, with listing counts. raw medians — no model yet.

the 2025 version's correlation analysis, redone on the log scale it should have used: log(price) correlates with age at −0.63 and odometer at −0.53 (stronger than the −0.55 and −0.48 in dollar space), with cylinders a mild +0.24. age and odometer are themselves correlated (+0.55) — cars age and accumulate miles together — which is exactly why the regression, not the correlation table, gets the last word on who does what.

correlation matrix of price, log price, age, odometer and cylinders
pearson correlations. blue positive, red negative.

letting box-cox choose the transformation

the 2025 write-up ended with "i would like to apply a natural logarithm to the price variable". this version starts there, but doesn't take the log on faith. the box-cox procedure fits the full covariate model at every power transformation of the response — λ = 1 is untransformed dollars, λ = 0 is the log — and profiles the likelihood.

box-cox profile log-likelihood peaking at lambda 0.18, with lambda 0 marked
profile log-likelihood of the box-cox parameter, fit within the full model on the training split.

the mle is λ̂ = 0.18 with a razor-thin 95% interval of (0.175, 0.180) — with 57,715 observations, the likelihood can tell apart values of λ that are practically identical, so the interval excludes both 0 and 1. the honest reading follows the textbook advice: λ = 1 (dollars) is emphatically rejected; λ̂ sits near zero; and among nearby values, λ = 0 is the one with an interpretation — coefficients become percentage effects on price. i use the log, and let the diagnostics section judge whether it was enough.

three models, two nested tests

three specifications, each an answer to "what did the change buy?". m1 is the 2025 model refit on the improved sample; m2 changes the response to log(price) and adds drive wheels (a variable the 2025 version threw away); m3 adds a quadratic age term, because the depreciation curve visibly flattens. all three are fit by ols on the same 80% training split (seed 138, same as 2025), with the same 20% held out.

log(price) = β₀ + β₁·age + β₂·age² + β₃·odometer + β₄·cylinders
           + brand + condition + fuel + title + transmission + body type + drive + ε
modelresponseparamsr²adj. r²aic
m1 — the 2025 specdollars570.6700.6691,177,230
m2 — log price, + drivelog590.7660.76653,625
m3 — m2 + age²log600.7720.77252,129

r² is not comparable between the dollar and log rows (different response scales) — the out-of-sample comparison below is the fair one. aic is comparable only between m2 and m3.

the additions are tested formally, the way nested models should be. because breusch-pagan (next section) rejects constant error variance, both tests are run as wald tests on the heteroskedasticity-robust (hc1) covariance; the classical partial f-tests, which assume the rejected constant variance, agree on both counts.

multicollinearity, the thing correlation tables can't adjudicate, is checked with variance inflation factors on m2: the largest is 2.87 (4wd), and no term exceeds the usual concern threshold of 5 (full table in vif.csv). age and odometer are correlated, but not enough to destabilize the coefficients.

checking the assumptions

this is the section the 2025 version didn't have. the residual plots below are the before-and-after argument for the log transformation: in dollar space (m1, top row) the residuals fan out as fitted values grow — textbook heteroskedasticity — and the q-q plot bends hard away from normal in both tails. in log space (m3, bottom row) the residual band is level and the q-q plot hugs the line through the central 99%+ of the data, with modest heavy tails.

four diagnostic panels comparing the dollar model's funnel-shaped residuals and bent q-q plot with the log model's level residual band and near-straight q-q plot
residuals vs fitted and normal q-q, for the dollar model (top) and the log model (bottom). a 6,000-listing sample is plotted; the statistics use all 57,715.
scale-location plot with a level trend line, and residuals versus leverage with no influential points
final model: scale-location (orange line = rolling median spread) and residuals vs leverage.

the formal version of the same story:

how well it predicts

every number in this section is computed on the 14,405 held-out listings, with 5-fold cross-validation on the training split as a second opinion. predictions from the log models are converted back to dollars with duan's smearing estimator (a naive exp() would be biased low; the smearing factor here is 1.071).

held-out test setmaemedian aemapermser² (dollars)
m1 — the 2025 spec$4,265$3,03258.2%$6,5870.667
m2 — log + drive$3,266$1,86833.7%$5,8060.741
m3 — final$3,300$1,83533.3%$6,0230.722

cross-validation agrees: mae $4,236 / $3,215 / $3,250, mape 58.3% / 34.1% / 33.7% for m1 / m2 / m3.

the log transformation, not extra complexity, does the heavy lifting: on identical data the 2025 specification's typical (median) error is $3,032, and thinking in percentages cuts it to $1,835 — 39% smaller — while mape drops from 58% to 33%. between m2 and m3 the choice depends on the loss you care about: m3 is better in log space (aic 52,129 vs 53,625), on median error and on mape; m2 edges it on the mean-type dollar metrics (mae, rmse, dollar r²), which are dominated by the priciest cars. i take m3 as the final model — its errors are smaller for the typical car, and its quadratic term is what makes the depreciation curve below believable — and report the comparison so the trade-off is visible.

bar charts of mean and median absolute error for the three models on the held-out test set
held-out error, mean and median. the log models (m2, m3) cut the typical error by more than a third.
hexbin of predicted versus actual price for held-out listings, concentrated along the diagonal across two orders of magnitude
predicted vs actual for the held-out set, both axes log-scaled. the 2025 version's plot bent away from this line at both ends; modeling in logs straightens it.

the part i care most about: uncertainty that means what it says. the final model's 95% prediction intervals, built in log space and exponentiated, cover the true price for 95.3% of the 14,405 held-out cars — within a third of a point of nominal. coverage holds across the price range, from 93.4% in the cheapest quartile of predictions to 97.2% in the priciest. the intervals are honest and they are wide (the upper bound is typically ~4.4× the lower): a single regression on ad-listed features genuinely cannot pin a used car's price much tighter, and pretending otherwise is what a point forecast does.

what moves the price

because the response is log(price), each coefficient converts to a percentage effect on price, holding everything else fixed. all intervals below are 95%, from robust standard errors. these are associations in 2021 craigslist asking prices, not causal effects — a diesel badge can't be bolted onto a corolla to raise its value 82%.

dot and interval chart of percentage effects on price for the non-brand coefficients, from diesel plus 82 percent to condition salvage minus 40 percent
percentage effect on asking price with 95% confidence intervals. blue = raises price, red = lowers it.

brand effects, now measured against an otherwise identical ford rather than as raw medians: porsche commands +110%, lexus +55%, toyota +39% — while fiat (−24%) and mercury (−22%) sit at the bottom. the 2025 version's "price boosters" table mixed brand prestige with body type; holding body type fixed is what separates "porsches are expensive" from "trucks are expensive".

dot and interval chart of brand premiums versus ford, from fiat at minus 24 percent to porsche at plus 110 percent
price premium vs an otherwise identical ford, 95% confidence intervals. full table in coefs_final.csv.

the depreciation curve

the model's age terms, drawn as a curve for a concrete car (a 4-cylinder automatic toyota sedan, clean title, good condition, carrying the median mileage for its age): it sheds half its nearly-new value by year five, and by year twelve it's a $6,000 car. this is the shape the quadratic term buys — the 2025 model's straight line in dollar space couldn't fall fast early and flatten late at the same time.

predicted price of a typical toyota sedan falling from about 32 thousand dollars at age zero to under 5 thousand past age 15
predicted asking price by age for a typical gas toyota sedan, at the median mileage for each age.

six cars, including my dream car

the 2025 write-up tested the model on three cars i found on autotrader: a corolla, a rav4, and the lexus is f i still daydream about. those three are re-scored below — but scoring 2025 canadian listings with a model of the april–may 2021 us market mixes two questions: is the model good, and has the market moved? (it has.) so the fair test comes first: the same three archetypes scored against 2021 listings the model didn't fit. the corolla and rav4 come straight from the held-out test split. the dataset's only real is f landed on the training side of the split, so it's scored leave-one-out: the model is refit without that listing before predicting it — out of sample by construction (and with the largest cook's distance anywhere at 0.003, removing one row barely moves the fit anyway).

caractualpredicted95% intervalerror
2021 listings, out of sample — the in-market test
2015 toyota corolla le, 125k mi$9,990$12,357$5,478 – $24,293+24%
2017 toyota rav4, 27k mi$23,991$30,498$13,520 – $59,959+27%
2008 lexus is f, 53k mi (leave-one-out)$34,999$17,750$7,867 – $34,907−49%
2025 autotrader listings — four years of market drift
2015 toyota corolla, 82k mi$10,048$8,366$3,709 – $16,448−17%
2022 toyota rav4, 56k mi$20,944$26,523$11,757 – $52,145+27%
2009 lexus is f, 51k mi$29,332$12,421$5,505 – $24,427−58%

the 2025 version fed the corolla's odometer in kilometres into a model expecting miles; the re-scores above use the converted 82k miles.

dot and interval chart of six cars showing predictions, actual prices and 95 percent prediction intervals; four actuals fall inside their intervals, and both lexus is f listings sit above theirs — the 2008 car only just
predictions (blue) vs actual prices (orange) with 95% prediction intervals (gray bands).

the toyotas behave: both corollas and both rav4s fall inside their intervals, even across four years of market drift. the is f does not behave, twice — four of the six actual prices land inside their intervals, and both misses are the same car. the 2021 is f's $34,999 asking price sits ninety-two dollars above its interval's upper edge; the 2025 listing sits about $4,900 above its. and this time the model can't blame the data. the reason is visible in the model's own structure: to a manufacturer-level regression, an is f is "a 13-year-old 8-cylinder lexus sedan", i.e. a used ls 460. the 500-horsepower badge lives in the model column, which carries 9,650 distinct values and stayed out of this regression. the honest conclusion isn't that the car is mispriced — it's that a limited-production performance car's value is its model name, and a brand-level model has a measurable ceiling there. (the 2025 version "did better" on the is f — 31% off instead of 58% on the same autotrader listing — but by accident: its straight-line depreciation overvalued all old cars, which happened to flatter the one old car that actually holds its value.)

what i'd still improve

reproduce it

everything on this page — every number, table and figure, in both color schemes — is generated by one pipeline in the repository:

data/raw/vehicles.csv        # kaggle export (1.4 gb, see readme for link)
analysis/clean.py            # the attrition table above
analysis/model.py            # fits, box-cox, nested tests, vif, robust inference
analysis/diagnostics.py      # residual and influence figures (run model.py first)
analysis/validate.py         # held-out metrics, cv, interval coverage, the six cars
analysis/eda.py, results_figs.py, style.py
python analysis/run_all.py   # raw csv in, report assets out

the cleaned modeling table ships with the repo as data/cars_clean.parquet, so everything after the cleaning stage can be reproduced without the 1.4 gb download.

references