Scraping dealer sites yourselfis cheap to start and expensive to keep
Writing a scraper for a few dealership websites is an afternoon. Keeping one running across the market is a standing engineering commitment: site platforms change templates, aggregators republish the same car, and nothing tells you a source went quiet. We scan 67,091 dealer websites daily to produce 19,349,581 live listings. Here is the honest comparison, including when you should build.
Should I scrape dealer websites myself or buy a listings feed?
Build if the job is small, bounded or constrained; buy if it is broad and ongoing. Scraping a handful of dealership sites for a one-off study is straightforward, and if a data-residency rule says the raw pages cannot leave your infrastructure, a third-party feed is not an option regardless of price. The calculation flips with scale and time. Producing 19,349,581 current listings means scanning 67,091 dealer websites every day, resolving 6,779,948 unique VINs out of duplicate postings across aggregators, and monitoring fill rates so a broken parser surfaces as a metric instead of a silent gap. That maintenance load never ends, and it does not shrink as your crawler matures - it grows with the number of sites you cover.
Hermes Data vs building it yourself
| Feature | Hermes Data | building it yourself |
|---|---|---|
| Scale of the crawl | 67,091 dealer websites scanned every day | Every site you want is a site you add and keep working |
| Platform churn | Parser changes are our problem, and they land before you notice | A template or platform change breaks extraction, often silently |
| Duplicate vehicles | 6,779,948 unique VINs resolved out of 19,349,581 postings | You write and tune VIN dedupe across aggregator republishing |
| Knowing it still works | Fill rates measured and published - 94.3% with a VIN | You build the monitoring that tells you a source went quiet |
| Cost shape | Per-call credits with no subscription - you pay for calls you make | Engineer time, fetch infrastructure and storage, every month |
| When it is the right call | Broad, current coverage you need working this week | A few sites, a one-off study, or a hard data-residency rule |
The number that decides it
One rooftop is a weekend. Ten is a sprint. 67,091 dealer websites, refreshed daily and producing 19,349,581 listings, is a system with an on-call rotation. Before you scope a build, write down how many sites you actually need current every morning - that figure settles the argument faster than any feature comparison.
Platform churn is the recurring cost
Dealership websites concentrate on a limited set of platform vendors - our own website_provider column is how we know - and those vendors ship template changes on their own schedule. Each change can quietly alter the markup an extractor depends on. The failure mode is not an exception in your logs, it is a field that starts arriving empty for one platform's worth of stores while everything looks green.
Dedupe and monitoring never finish
The same car is posted by the dealer and republished by aggregators, so raw counts overstate the market until you resolve them - we carry 6,779,948 unique VINs. Then you need fill-rate monitoring to catch degradation: we track and publish ours, at 94.3% of listings with a VIN and 96.8% with a photo. Both of these are permanent jobs, not launch tasks.
When you should build
Three cases, honestly. You need a handful of specific sites and nothing more. You are doing a one-off study where the data is thrown away afterwards. Or you have a data-residency or contractual requirement that the collection happen inside your own infrastructure. In any of those, build it - buying a broad feed would be paying for coverage you will not use.
Frequently asked questions
Should I build my own dealership inventory scraper or buy a listings feed?
It depends on breadth and duration. For a few dealership websites, or a study you run once and finish, building is reasonable and probably faster than procurement. For broad coverage that has to stay current, the work is ongoing: daily scanning across tens of thousands of sites, VIN deduplication across aggregators, and monitoring that catches a broken parser before your users do. Hermes Data scans 67,091 dealer websites daily and sells access per API call.
What usually breaks in a homegrown dealership scraper?
Three things, in order of how often we see them. Website platform vendors change templates, so extraction that worked last month returns empty fields. Aggregators republish the same vehicle, so counts inflate until VIN deduplication is right. And nothing announces a failure - a store that stops responding simply contributes zero rows, which looks identical to a store with no inventory unless you are monitoring fill rates per source.
When is building your own listings pipeline the right call?
When the scope is narrow enough that maintenance stays small, or when a rule forces it. A handful of named dealership sites, a one-time research project, or a data-residency requirement that raw collection stay inside your own infrastructure are all good reasons to build. Buying broad coverage for a narrow need is waste, and we would rather say so than pretend otherwise.
What does buying the data cost compared with building?
Listings meters each API call by request class rather than charging a subscription: keyed lookups, aggregates, full records and scans cost 1, 2, 5 and 10 millicredits against a credit of 1,000. Paging search results at 100 listings per call is roughly 10,000 listings per credit, and failed calls are not billed. Compare that against engineer time plus fetch and storage infrastructure carried every month, not against the cost of the first prototype.
More on vehicle listings
Try Hermes Data for $0.99
Self-serve access to the automotive data your business runs on. The $0.99 trial gives you 20 credits to spend across Contacts, Dealers, and Dealer Groups. Then auto-upgrade to Starter ($99/mo) — or cancel any time.
Contacts, Dealers, and Dealer Groups are all self-serve on one credit wallet — spend your trial credits on any of them. Vehicle Listings is live too, on its own API subscription (separate from the credit wallet). Explore Listings.