# Methodology

DealerClaw collected paired initial HTML and browser-rendered observations on September 7, 2026. The completed sample has 486 distinct U.S. dealership websites and 2,430 analytic pages: one homepage, one vehicle results page and three distinct vehicle detail pages per website.

## Sampling and page selection

A 20-site pilot preceded the main method freeze. The original frozen candidate frame contained 1,801 distinct domains. Public dealer-group and association rosters supplied the frame. Their identifying entries remain in the private archive. This is a diversity sample with access limits, not a probability sample of U.S. dealerships.

Main rules were frozen at 13:42:48.188185 UTC on September 7. Regional targets were 125 each. Concentration caps were 40 websites per state, 20 per identified dealer group, 100 per manufacturer family and 200 per supported provider. The final sample covered 39 states. Main and replacement queues used the frozen rules. Bounded inventory discovery looked for new and used inventory. Fixed hash ordering selected eligible detail-page candidates, with both conditions represented where found. Extra discovery pages and later rechecks are not added to analytic denominators.

We stopped collection at 16:54:30.098 UTC after the final replacement queue. This was an operational decision, not a predeclared stopping time. Unattempted reserves remain. The completed sample fell 14 short of the target: 13 in the West and one in the South. See [sampling outcomes](/insights/dealership-structured-data-study-2026/assets/sampling.md).

## Measurement

Extraction covered JSON-LD, Microdata and RDFa, including nested entities and linked nodes, using Schema.org v30.0. Explicit declarations and inherited types were counted separately. Car counts under Vehicle and Product; AutoDealer counts under LocalBusiness and Organization. BreadcrumbList presence does not by itself establish an inventory list.

Field completeness is an explicit, nonempty property on a bound relevant entity. The primary vehicle is identified before counting its fields, so other vehicle recommendations do not fill its missing properties. Presence, completeness, technical validity and agreement with visible content are separate measurements. Unknown or uninterpretable extraction is not automatically absence. A valid block remains positive even when another block has a parsing failure.

Price comparisons distinguish MSRP, selling prices, fees, eligibility conditions and monthly payments. Inventory comparisons use displayed cards and corresponding vehicle identities or detail links. Pagination, sold cards and full-stock totals remain separate. Automatic agreement labels are candidates unless an audit explicitly confirms the relevant scope.

Provider attribution used supported publication credits. Asset hosts alone were insufficient. Only groups with at least 20 completed websites appear in the article comparison. Provider labels in this edition preserve those groups. Association does not establish responsibility for configuration or search performance.

## Verification and revisions

Original timestamped captures, source hashes, selected URLs, extraction output and collection/analysis versions are retained privately. Analysis corrections were documented after the freeze and did not rewrite the original captures. Direct AI-assisted review covered 132 pages from 44 websites, 95 selected field records and 70 rendering-category checks on 41 final pages, plus targeted reviews and fresh rechecks. Separate raw-source recounts checked primary-vehicle binding and headline counts. These were separate code paths, not independent human adjudication.

The September 8 new/used mileage split used the original captures. It added no websites and was not part of the frozen primary analysis. Thirteen supplementary source reviews checked examples across condition and mileage-presence categories. See [mileage methods](/insights/dealership-structured-data-study-2026/assets/mileage.md) and [verification notes](/insights/dealership-structured-data-study-2026/assets/verification.md).

## Denominators and limits

Every included website contributes five analytic pages. Detail-page results cluster three pages within each website, with further dealer-group and provider dependencies. No national confidence intervals or ranking, AI-citation, traffic or sales effects are reported. Completeness rates use evaluable bound subjects, with unknowns separate. The descriptive homepage increase is 162/486. The clean paired-rendering table uses 484 homepages because two are parse-limited. These are different counting bases.

Names and direct identifiers are withheld from this edition. Published summary tables support checking counts and percentages; raw-source replication requires the private archive. [Anonymization scope](/insights/dealership-structured-data-study-2026/assets/methodology.md#publication-scope)

## Relationship to earlier research

Earlier dealership studies use different scoring systems, access tests, page choices and outcomes. Their percentages are not directly comparable with this selected-VDP adoption rate. This study adds paired source states, explicit field definitions and scoped content verification without reusing a proprietary readiness score. Background research and standards are attributed in the [source ledger](/insights/dealership-structured-data-study-2026/assets/sources.md) and article.

## Publication scope

The public release includes the approved article, anonymous charts, summary tables, field definitions and verification notes. It excludes the review ZIP, spreadsheet, site and vehicle rows, original identities and source captures. Provider labels are consistent across the published article and charts. DealerClaw and cited research retain attribution.
