katsuma onishi

Analyzing and Predicting Used Car Prices, Take Two

project · august 2026 · python, statsmodels · full technical write-up · source on github · the 2025 version

what should a used car actually cost? in 2025 i built my first linear regression to answer that. this is the do-over, with everything i learned finishing my stats degree at mcmaster — same question, same craigslist data, a much more honest answer. the short story is below; the full technical write-up has every test and assumption for the curious.

listings modeled
72,120
cleaned from 426,880 raw ads
typical miss
$1,835
old approach: $3,032
price range hit rate
95.3%
on 14,405 cars it never saw
dream car verdict
missed
and that's the interesting part

the setup

the data is 426,880 craigslist car ads from spring 2021. raw craigslist is messy: the same car posted in five cities, $0 "call for price" ads, a $3.7 billion pickup truck. after removing 171,428 duplicate reposts and the junk, 72,120 real listings remain — and this time the expensive cars stay in. (the 2025 version deleted everything over $44,500 to make the data look nicer, which is exactly why it couldn't price anything interesting.) each listing has the asking price plus the basics: age, mileage, brand, body type, condition, title history, transmission, drivetrain.

cloud of listings showing price falling as cars age, from a median around 46 thousand dollars when new to about 5 thousand at 20 years
the whole story in one picture: cars lose value fast, then stop. each hexagon is a bundle of listings, darker = more of them; the orange line is the typical price at each age.

notice the drop isn't a straight line — a car loses a huge slice of its value early, then the curve flattens. that shape is why the model thinks in percentages, not dollars: knocking $2,000 off matters a lot on a $6,000 civic and barely registers on a $60,000 truck. (the formal tool for this decision is called box-cox, and it agrees — details in the technical write-up.)

the model, and how good it is

the model is a linear regression: it learns how much each feature — age, mileage, brand, body type, condition, title, transmission, drivetrain — adds or subtracts from the price, all at once. i trained it on 80% of the listings and graded it on the 14,405 it had never seen. no peeking.

predicted versus actual price for over 14 thousand unseen listings, hugging the diagonal line from 1 thousand to over 100 thousand dollars
on cars it never saw, predictions hug the "perfect" line — from $1,000 beaters to $100,000+ trucks. the typical miss is $1,835; the old approach missed by $3,032 on the exact same cars.

one more thing the 2025 version didn't have: every prediction comes with a price range, not just a number. the ranges are built to catch the true price 95% of the time — and on the unseen cars they caught it 95.3% of the time. an accuracy claim you can check is worth ten you can't.

what actually moves the price

because the model holds everything else equal, it can separate questions that raw averages mix up — like "are trucks expensive, or are truck brands expensive?" (answer: trucks. being a ram or gmc adds little once the model knows it's a truck.)

chart of price effects: diesel plus 82 percent, off-roader plus 72, truck plus 51, down through minus 12 percent per year of age, minus 17 for a rebuilt title and minus 40 for fair condition
blue raises the price, red lowers it. every effect is measured with everything else held equal — a diesel version of the same truck, a rebuilt-title version of the same corolla.

put the age effects together and you get the depreciation curve every car buyer feels: half the value gone by year five, a $6,000 car by year twelve.

predicted price of a typical toyota sedan falling from about 32 thousand dollars new to under 5 thousand past age 15, with the halfway point marked at year five
the model's depreciation curve for a typical toyota sedan. steep early, flat late — which is why "buy a three-year-old car" is real advice.

the road test: six real cars, including my dream car

numbers on a chart are one thing; real cars are the fun part. i scored six: a corolla, a rav4, and the lexus is f i daydream about — once each from 2021 listings the model never fit, and once each from the 2025 autotrader ads i tested in the original project.

six cars with predictions, actual prices and price ranges; the corollas and rav4s land inside their ranges, both lexus is f listings sit above theirs
blue dot = model's prediction, orange diamond = real asking price, gray band = the model's 95% range. the everyday cars behave. the is f doesn't — twice.

both corollas and both rav4s land inside their ranges. the is f escapes its range in both years — the 2021 car by just ninety-two dollars, the 2025 one by miles. and the model's excuse is actually the lesson: all it can see is "a 13-year-old 8-cylinder lexus sedan," so it prices a sensible used luxury car. it has no way to know the is f is a 500-horsepower collector's item, because that lives in the model name — a column with 9,650 different values this regression doesn't use. some cars are worth what their badge means, and no spreadsheet column carries that.

what i learned

want the rigor? the full technical write-up has the cleaning log, box-cox, the diagnostics, robust inference, and all fourteen figures — and everything regenerates from one pipeline in the repository. data: used cars dataset by austin reese. built on coursework from stats 3a03 and math 3mb3, mcmaster university.