Used-Car VDP Photo Benchmark: Data + Replication
Quick answer: CarPixAI reviewed 108 live, public used-vehicle detail pages from 9 dealer websites across 8 states on July 28, 2026. The median listing scored 8 out of 12. This is a transparent observational benchmark of the sampled pages—not a nationally representative industry study and not evidence that a photo score causes sales.
A separate multi-platform replication then applied the same frozen rubric to a fixed 122-URL frame. It included 111 VDPs across 17 domains and seven platform families, with a median of 8 out of 12. The two releases remain separate and neither is industry-representative.
The strongest pattern in this sample was gallery quantity: 76.9% of listings exposed at least 20 unique vehicle images. Presentation quality was less consistent. Only 17.6% earned the full two points for the published hero image, while 20.4% earned full marks for repeatable, non-distracting backgrounds.
The public release includes the complete anonymized row-level dataset, aggregate calculations, rubric, data dictionary, methodology, limitations, and reproducible collection/scoring script. Dealer names, domains, listing URLs, VINs, stock numbers, titles, states, and source images are excluded from the public downloads.
Want a decision-focused review of your own set? Attach the photos you have (up to 12 at a time) to the free conversational Car Listing Photo Grader and ask about the lead image, missing coverage, consistency, reshoots, or eligible presentation fixes. Its answer is directional and does not turn these purposive samples into a market-representative score.
Download the anonymized CSV · Download row-level JSON · Download aggregate JSON · Download QA results · Download methodology · Download rubric · Download data dictionary
Independent v1.1 multi-platform replication
After v1 was frozen, a previously dispatched sourcing batch returned a broader fixed frame of 122 verified VDP URLs across 19 dealer domains, 11 states, and seven website-platform families. CarPixAI ran that pool as a separate replication rather than replacing favorable or unfavorable v1 rows after seeing the original results.
The replication included 111 VDPs across 17 domains. 11 fixed-frame candidates were excluded without replacement: 10 because robots policy remained unreachable and 1 because fewer than three vehicle images could be extracted. All seven platform families remained represented.
| Measure | Frozen v1 | Separate v1.1 replication |
|---|---|---|
| Included VDPs | 108 | 111 of 122 fixed candidates |
| Dealer domains | 9 | 17 |
| Website-platform families | 2 | 7 |
| Mean score, 0–12 | 7.79 | 8.14 |
| Median score, 0–12 | 8 | 8 |
| At least 20 detected images | 76.9% | 84.7% |
The replication mean was 0.35 points higher, while both medians were eight. Full-pass rates were higher in the replication for hero image, gallery completeness, condition proof, and background consistency, but lower for mobile crop and visible edit safety. These are unweighted descriptive differences—not platform effects—because dealer, vehicle, platform, and extraction composition all changed.
| Full-pass dimension | v1 | v1.1 replication | Difference |
|---|---|---|---|
| First hero image | 17.6% | 25.2% | +7.6 pp |
| Gallery completeness | 76.9% | 84.7% | +7.8 pp |
| Condition proof | 39.8% | 51.4% | +11.6 pp |
| Background consistency | 20.4% | 35.1% | +14.7 pp |
| Mobile crop safety | 50% | 46.8% | -3.2 pp |
| Visible edit safety | 50.9% | 46.8% | -4.1 pp |
Replication QA: Two rows from each platform family were independently rescored with a second model. Across 70 comparisons, exact agreement was 38.6%, 90% were within one point, and mean absolute difference was 0.71. This again shows that the visual scores are directional and reviewer-sensitive.
Download v1.1 CSV · Download v1.1 row JSON · Download v1.1 summary · Download v1.1 QA · Download v1.1 methodology · Download v1.1 data dictionary · Download descriptive comparison
Try it now
Try the same upload workflow with one inventory photo
Upload or select a real car photo, choose a cleaner background, enter your email, open the magic link, then process and download finished images from your dashboard. The free trial includes 5 photos with no credit card required.
What the sample contained
| Unit of analysis | One live public used-vehicle VDP as observed on the collection date |
|---|---|
| Completed sample | 108 VDPs across 9 dealer domains |
| Dealer mix | 54 independent-group or independent listings and 54 franchise listings |
| Geography | 8 states across Great Plains, Midwest, South Central, Southeast |
| Collection date | July 28, 2026 |
| Scoring | Five model-assisted visual dimensions plus one deterministic gallery-count dimension |
| Rubric version | v1.0-2026-07-28 |
| Representativeness | Purposive, balanced observational sample; not probability-weighted or nationally representative |
Descriptive findings from the sampled VDPs
Across the 108 sampled listings, the mean score was 7.79 and the median was 8 out of 12. Scores ranged from 3 to 12. These figures describe only this fixed sample and collection date.
| Scoring dimension | Mean, 0-2 | Full-pass listings | Full-pass rate |
|---|---|---|---|
| First hero image | 0.98 | 19 of 108 | 17.6% |
| Gallery completeness | 1.68 | 83 of 108 | 76.9% |
| Condition proof | 1.31 | 43 of 108 | 39.8% |
| Background consistency | 1.11 | 22 of 108 | 20.4% |
| Mobile crop safety | 1.34 | 54 of 108 | 50% |
| Visible edit safety | 1.36 | 55 of 108 | 50.9% |
Gallery count was the most common full pass: 83 listings had at least 20 unique detected vehicle images. Condition proof was less complete in the representative image sample: 43 listings showed enough breadth for a full score. This does not prove a missing proof image was absent from every unsampled gallery position.
Visual-score sensitivity: A stratified 12-row second-model QA pass rechecked 60 visual-dimension scores. Exact agreement was 40%, 85% were within one point on the 0–2 scale, and mean absolute difference was 0.75. The gallery-count metric is deterministic, but the visual percentages above are model-sensitive directional observations—not ground truth. The disagreement is published rather than hidden.
Most frequent published-hero issues
| Observable issue code | Listings | Share of sample |
|---|---|---|
| Text or branding overlay | 57 | 52.8% |
| Cluttered background | 42 | 38.9% |
| Visible editing artifact | 14 | 13% |
| Harsh glare | 4 | 3.7% |
| Tight crop | 3 | 2.8% |
| Placeholder image | 3 | 2.8% |
The explicit 12-point rubric
Every VDP receives zero, one, or two points in six dimensions. A zero indicates missing or unusable observable evidence, one indicates a usable result with material limitations, and two indicates a full pass under the published definition. The rubric was frozen before the full sample was scored.
| Dimension | 0 points | 1 point | 2 points |
|---|---|---|---|
| First hero image | Placeholder, heavily obstructed, severely cropped, blurry, or not useful as a vehicle hero | Usable but affected by clutter, glare, darkness, obstruction, weak angle, or tight crop | Sharp, readable whole-vehicle presentation with a useful angle, manageable background, and crop room |
| Gallery completeness | Fewer than 12 unique detected vehicle images | 12 to 19 unique detected vehicle images | At least 20 unique detected vehicle images |
| Condition proof | Representative images provide almost no condition evidence | Some interior, feature, odometer, wheel, cargo, or condition evidence appears | Interior plus at least two additional buyer-proof categories appear in the representative sample |
| Background consistency | Severe clutter, inconsistent settings, or repeated obstruction | Usable presentation with visible distractions or setting/framing variation | Repeatable, non-distracting exterior setting with coherent framing |
| Mobile crop safety | Important vehicle edges are already clipped or likely to disappear | Readable hero, but at least one edge is tight | Safe space remains around roof, bumpers, and tires for common card crops |
| Visible edit safety | Obvious distortion, concealment, implausible compositing, or misleading treatment | Uncertainty, aggressive overlays, or inconsistent visual cues that merit source review | No observable manipulation warning signs in the reviewed images |
The visible-edit score is deliberately narrow. A two does not prove an image is unedited, and it does not prove every vehicle detail matches an unavailable source photo. It only means the reviewed images showed no obvious warning sign under this rubric.
Methodology
- Source selection: Nine US dealer-owned websites were purposively selected to cover independent and franchise inventory across several regions. Aggregators and marketplaces were excluded.
- Robots check: The collection script requested each public robots.txt and continued only when the sitemap and representative VDP path were not disallowed for the general crawler group.
- VDP frame: Eligible URLs came from public dealer sitemaps, contained a VIN-shaped identifier, and were explicitly marked used or pre-owned in the path.
- Deterministic selection: Eligible URLs within each domain were ranked using a published SHA-256 rule and a versioned sampling salt. The first eligible rows were processed until the fixed domain quota was met.
- Inclusion: A page had to be live and yield at least three unique VIN-linked vehicle images. Failed, removed, blocked, placeholder-only, or insufficient-image pages were recorded internally and replaced by the next ranked page from the same domain.
- Image review: The hero plus images sampled from the beginning, middle, and end of the gallery were normalized to JPEG and reviewed at temperature zero using google/gemini-2.5-flash. The number of reviewed images is present in every public row.
- Deterministic count: Gallery completeness used the number of unique source image paths after thumbnail and resize variants were deduplicated.
- Missing data: No score was imputed. A source that could not satisfy inclusion criteria was excluded and replaced within its existing domain quota.
- Anonymization: Public rows retain an anonymous dealer ID, broad segment, region, URL hash, body style, evidence fields, and scores. They omit dealer/customer identities and all source identifiers that would expose a VIN, stock number, page title, URL, state, or image.
- QA: 12 rows were selected across all four score bands and re-reviewed with openai/gpt-4.1-mini. Exact and within-one-point agreement are published in the QA download. This is a second-model sensitivity check, not a human reliability study.
The complete executable procedure is in scripts/build-vdp-photo-benchmark.mjs. The downloadable methodology records the sampling salt, thresholds, reviewer configuration, replacement rules, and file lineage.
Anonymized representative examples
Raw dealer photos are not republished. The examples below are de-identified row summaries selected mechanically from the lowest, nearest-to-median, and highest total scores. They illustrate how the rubric behaves without naming or visually exposing a dealership, vehicle listing, customer, VIN, stock number, plate, or location.
Lower-scoring example: 3/12
Anonymous independent listing from the Southeast; suv; 3 unique detected images and 3 representative images visually reviewed. Hero 0/2, completeness 0/2, condition proof 0/2, background 0/2, mobile crop 2/2, and visible edit safety 1/2.
Recorded priority: Replace all placeholder images with actual photos of the vehicle.
Typical-score example: 8/12
Anonymous independent group listing from the Midwest; sedan; 34 unique detected images and 8 representative images visually reviewed. Hero 1/2, completeness 2/2, condition proof 1/2, background 1/2, mobile crop 1/2, and visible edit safety 2/2.
Recorded priority: Replace map and review images with actual vehicle photos.
Higher-scoring example: 12/12
Anonymous franchise listing from the Southeast; truck; 22 unique detected images and 8 representative images visually reviewed. Hero 2/2, completeness 2/2, condition proof 2/2, background 2/2, mobile crop 2/2, and visible edit safety 2/2.
Recorded priority: none
Limitations and scope boundaries
- Directional observational sample, not a probability sample of all US dealerships or all live used vehicles.
- Dealer-owned VDPs were sampled from nine public sitemaps across eight states using a deterministic URL hash rank.
- The sample is intentionally balanced between independent and franchise listings, so aggregate percentages are not market-share weighted.
- Most sampled websites used the same inventory website/image-delivery platform, which may cluster gallery and presentation behavior.
- Scores use the published hero plus representative images sampled from the beginning, middle, and end of each gallery; not every image received multimodal review.
- Gallery completeness is based on detected image count; sampled visual evidence may miss a proof photo elsewhere in the full gallery.
- Visible edit safety records observable warning signs only. It cannot prove an image is unedited or that every vehicle detail matches an unavailable source photo.
- Model-assisted scores can contain classification error. A documented stratified second-pass QA review is published separately; no formal human inter-rater study was performed.
- The benchmark does not measure or imply leads, appointments, sales, time-to-sale, price realization, or causal business impact.
- Listings change and may be removed after the collection date.
The sample was balanced for this audit rather than weighted to the actual US dealer population. That makes independent-versus-franchise coverage easier to inspect, but it means the overall percentages should not be projected to the industry. Website-platform concentration may also make gallery behavior more similar than it would be in a wider platform sample.
The v1.1 replication broadens platform coverage, but it is still purposive and unweighted: franchise rows outnumber independent rows, one platform contributes 41 of 111 included VDPs, and gallery extraction behaves differently across platforms. The replication reduces one v1 limitation without creating a representative industry sample.
This work evaluates observable photo merchandising only. It does not evaluate vehicle quality, dealer ethics, pricing, availability, customer experience, lead volume, appointments, sales, days-to-sale, or return on investment. It cannot establish that improving a score changes any commercial outcome.
How dealers can use the framework
Use the blank dealer photo audit template to apply the same six dimensions to your own inventory. Keep the sample rule stable, preserve the source date, and compare workflow changes over time without turning an internal score into an unsupported sales claim.
For the practical workflow behind the first image, use the VDP hero-image previewer and the first-nine VDP photo checklist. If a source photo is sharp and truthful but the environment is distracting, test a copy with the car background remover. Retake blurry, cropped, stale, or inaccurate source photos instead of trying to repair them with AI.
FAQ
Is this a representative automotive industry study?
No. The frozen v1 contains 108 purposively sampled public VDPs, and the separate v1.1 replication contains 111 of 122 fixed candidates. Neither is a probability sample, and neither should be generalized to every US dealership.
Can the findings be reproduced?
The public methodology, rubric, row-level calculations, URL hashes, sampling rule, and executable script are available. Live listings can change or disappear, and dealer URLs are withheld from public files to avoid naming or ranking stores, so exact visual replay requires the internal provenance manifest.
Did CarPixAI compare edited images with source photos?
No. The visible-edit dimension identifies observable warning signs in published images. It does not establish whether an image was edited or whether every detail matches an unavailable original.
Does a higher photo score cause more leads or sales?
No causal outcome was measured. The benchmark only describes observable photo practices on the collection date.
Frequently asked questions
What did the 2026 used-car VDP photo benchmark review?
The frozen v1 reviewed 108 live public used-vehicle detail pages from nine dealer websites across eight states. A separate v1.1 multi-platform replication then included 111 of 122 fixed candidates across 17 dealer domains, 11 states, and seven website-platform families. Both used the same six-dimension, 12-point rubric.
Is the CarPixAI VDP benchmark representative of the automotive industry?
No. Both releases are purposive observational samples, not probability samples or market-share-weighted estimates. The v1 and v1.1 percentages describe their separate sampled pages only and are not pooled.
Can the benchmark data be downloaded?
Yes. The report provides separate anonymized v1 and v1.1 row-level CSV and JSON files, aggregate summaries, QA sensitivity files, methodology, data dictionaries, the frozen rubric, and a descriptive cross-release comparison. Dealer domains, listing URLs, VINs, stock numbers, titles, exact states, and source photos are withheld from public files.
Does the visible edit safety score prove a photo was not edited?
No. The score records observable warning signs in the representative published images. A full score does not prove an image is unedited or that every vehicle detail matches an unavailable source photo.
Does a higher benchmark score cause more vehicle leads or sales?
No causal outcome was measured. The benchmark evaluates observable photo merchandising only and does not measure leads, appointments, sales, days-to-sale, price realization, or return on investment.
How should marketers cite the used car photo benchmark?
Cite it as a practical used-car photo audit method for dealers, not as a universal sales-lift statistic. State that the published percentages describe the sampled pages only and that no causal sales outcome was measured.
Ready to upgrade your listing photos?
Try CarPixAI free: 5 photos, no credit card required.
Try 5 photos free