The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Commercial property-data brokers misreport and omit substantial transaction data: 1–2% of matched sales show large, systematic price errors and 12–15% of transactions are missing or misclassified, producing meaningful differences in estimated property‑tax regressivity.

Hidden Errors in Big Data: The Case of Property Records
Evelyn Smith, Emma Harvey, Jacob Goldin, Daniel E. Ho · July 30, 2026
arxiv descriptive high evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Evelyn Smith unresolved corpus identity
  2. Emma Harvey unresolved corpus identity
  3. Jacob Goldin unresolved corpus identity
  4. Daniel E. Ho unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Evelyn O. Smith provider ID
  2. Emma Harvey provider ID
  3. Jacob Goldin provider ID
  4. Daniel E. Ho provider ID
Auditing ATTOM and Cotality against Cook County administrative records (2018–2021) reveals 1–2% of matched sales have >5% price discrepancies—often exact multiples consistent with transfer‑tax imputation errors—and coverage/misclassification affecting roughly 12–15% of transactions, which materially alters estimates of property‑tax regressivity.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between brokers: more often than not, brokers make identical reporting errors for the same transactions. Third, to illustrate the significance of these errors, we measure their impact on estimates of property tax regressivity, finding that they drive significant wedges between estimates depending on the data source. These findings generalize to two other large counties in the U.S., and highlight the crucial importance of open administrative data and transparency from brokers regarding data provenance and lineage.

Summary

Main Finding

Brokered property datasets widely used in research and ML (ATTOM and Cotality) contain systematic coverage and imputation errors relative to county administrative “ground-truth” records. These errors (1) cause 1–2% of matched transactions to have sale prices differing by >5%—often by simple integer ratios (e.g., 2/3, 1/3, 2)—(2) produce coverage losses on the order of 12–15% of transactions (missing records, misclassified sale/property filters, duplicates), and (3) are highly correlated across brokers (errors often shared), producing meaningful biases in downstream economic estimates such as property‑tax regressivity. The problems generalize beyond Cook County (robustness checks in NYC and Philadelphia).

Key Points

  • Two failure modes:
    • Imputation errors: broker sale prices sometimes differ substantially from county deed prices (1–2% of matched sales; many discrepancies are exact/simple multiples).
    • Coverage errors: 12–15% of ground-truth arms‑length single‑family sales are not represented as such in brokered data due to missing records or misreported filters (e.g., marked as foreclosure, condo, multi‑parcel).
  • Quantitative highlights (Cook County, 2018–2021):
    • Ground-truth sales (after filters): ~153k transactions.
    • ATTOM matched ~134,427 transactions (≈87.8% of ground truth); Cotality matched ~129,606 (≈84.7%).
    • ATTOM: 1,976 matched transactions had sale-price discrepancies >5%; distribution of ratios of broker to ground-truth price is dominated by 2/3 (≈50% of discrepant cases), plus 2.0, 1/3, etc.
    • Sources of non-match (ATTOM): ~37% missing record, ~50% misreported filter, ~8.6% duplicates, ~4% fuzzy-match thresholds. (Cotality had more missing records and fewer misreported filters.)
    • Discrepancies sometimes amount to hundreds of thousands of dollars (conditional mean errors large for discrepant subset).
  • Mechanism: Evidence points to imputation via municipal transfer‑tax rates. If brokers impute sale price from reported transfer‑tax payments but use the wrong municipality’s rate or the wrong payer’s share, the imputed price will be an exact multiple of the true price (matching the simple ratios observed).
  • Shared errors: Many of the same transactions are misreported by both brokers, consistent with broker-to-broker data resale/propagation or common imputation routines.
  • Validation: Authors treated county administrative data as ground truth, and manually checked a random 2% sample of broker–county disagreements against deed records; county prices matched deed records in the overwhelming majority of sampled disagreements (Cook County accuracy ≈99.8–99.9% in sampled checks).
  • Downstream impact: These errors materially change estimates of property-tax assessment regressivity (the paper shows significant wedges in regressivity estimates depending on data source).

Data & Methods

  • Data sources:
    • Brokered datasets: ATTOM and Cotality.
    • Ground-truth: county administrative transaction and assessment data (Cook County), with robustness checks using New York City and Philadelphia.
  • Sample period: 2018–2021.
  • Filters applied (to create comparable samples): arms‑length transactions of single‑family homes; exclude sales < $10k, multi‑parcel sales, repeated transactions in the same year, and records missing address/date/price.
  • Matching procedure:
    • Matched records using address, latitude/longitude, and sale date (fuzzy matching thresholds applied; robustness tests for thresholds are reported).
    • Matched counts and unmatched-cause coding (manual classification of reasons for no-match).
  • Error classification:
    • Imputation/accuracy error: difference in sale price or assessed value between broker and county records; flagged >5% absolute difference.
    • Coverage/representativeness error: missing records, misreported sale/property characteristics leading to exclusion from filtered sample, duplicates.
  • Validation:
    • Manual verification of a random 2% sample of broker–county disagreements against deed records to confirm county record accuracy.
  • Analysis of mechanism:
    • Examined distribution of broker/ground-truth price ratios; identified spikes at simple rational multiples consistent with transfer‑tax imputation mistakes.
  • Impact analysis:
    • Measured effect on property-tax assessment regressivity estimates across data sources and years to quantify substantive consequences.

Implications for AI Economics

  • Data provenance and lineage matter for economic inference and for ML systems that use brokered administrative-like data. Errors in upstream broker processing (imputation, harmonization) propagate into models and policy metrics.
  • Multiple brokers does not guarantee robustness: because brokers often share sources and resale data, errors are frequently correlated across brokers; naive ensemble/“cross-check” with another broker may provide a false sense of security.
  • High‑stakes or policy-relevant analyses that rely on brokered property data should:
    • Prefer open administrative sources where available; treat broker data as convenience products requiring validation.
    • Implement routine validation and auditing steps: sample verification against deeds/administrative records, distributional checks, and search for simple-ratio patterns that indicate systematic imputation errors.
    • Document data provenance, cleaning, and imputation steps (both for reproducibility and for assessing bias).
    • Run sensitivity analyses (e.g., recompute key estimates dropping suspect transactions, or adding uncertainty bounds for imputed prices).
    • Demand greater transparency from brokers about lineage and imputation procedures; when contracting with brokers, require disclosure of methods and access to source identifiers.
  • For ML pipelines:
    • Incorporate data-quality checks (unit tests, anomaly detection for simple-ratio imputation artifacts), propagate uncertainty from input data to downstream model outputs, and prioritize retraining/updates when administrative corrections are published.
  • For research in AI economics:
    • Be cautious interpreting precise-sounding estimates from large brokered datasets: the “Big Data Paradox” can produce precise but biased estimates if upstream errors are overlooked.
    • When possible, publish code and matching procedures so others can audit sensitivity to data source and matching thresholds.
  • Policy: Governments and funders should invest in making administrative data more accessible and standardized to reduce reliance on opaque brokers and to improve the evidence base used by researchers, appraisers, and automated systems.

If you want, I can extract the paper’s key tables (match rates, ratio distributions, and match-failure breakdowns) into a compact CSV-style summary or produce a checklist for auditing brokered property datasets you or your team could run on other jurisdictions.

Assessment

Paper Typedescriptive Evidence Strengthhigh — The authors compare near-complete broker datasets to administrative ‘‘ground-truth’’ records for the universe of filtered single-family arms-length sales in Cook County (2018–2021), manually verify a random sample of disagreements against deed records, document systematic patterns (exact multiples consistent with transfer-tax imputation errors), quantify coverage gaps, and report robustness checks in two other large counties; sample sizes are large and the verification step supports the ground-truth assumption. Methods Rigorhigh — Careful record-level matching with fuzzy-address and date checks, explicit filtering rules, manual verification of disagreements against deeds, decomposition of match-failure causes, and examination of systematic error patterns (ratios consistent with transfer-tax imputation) indicate rigorous empirical work; limitations arise from focusing primarily on a small set of jurisdictions and from relying on county data as 'ground truth' (which they partially validate). SampleUniverse of arms-length single-family home transactions in Cook County, IL for 2018–2021 from three sources: Cook County administrative records (treated as ground truth), ATTOM, and Cotality; filters exclude sales < $10k, multi-parcel sales, properties with multiple transactions in a year, and records missing address/date/amount; resulting counts ~153k (ground truth), 152k (ATTOM), 148k (Cotality); matched subsets of ~134k (ATTOM) and ~130k (Cotality); random 2% manual verification of disputes against deed records; robustness checks using New York City and Philadelphia data reported in appendices. Themesgovernance inequality IdentificationDeterministic and fuzzy matching of broker records (ATTOM and Cotality) to county administrative transactions (Cook County open data) for arms-length single-family home sales 2018–2021, treating county records as ground truth and validating a random 2% of disagreements against deed records; robustness checks in two additional counties and sensitivity checks on matching thresholds. GeneralizabilityPrimary analysis concentrated on Cook County (with robustness checks in NYC and Philadelphia); results may differ in other U.S. counties or internationally., Focus restricted to arms-length single-family home sales; other property types (condos, multi-parcel, commercial) may exhibit different error profiles., Time period 2018–2021; broker practices, municipal transfer taxes, or broker updates after ingestion could change error patterns over time., Findings pertain to the two audited brokers (ATTOM, Cotality); other brokers or custom-aggregated datasets may behave differently., Matching and filtering rules (and thresholds) affect measured coverage; some unmatched cases may reflect matching limitations rather than broker omissions.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
For approximately 1–2% of matched Cook County transactions from 2018–2021, ATTOM-reported sale prices differ from ground-truth sale prices by more than 5%. Output Quality negative Accuracy of broker-reported property sale prices
Reading fidelity high
Study strength high
n=134427
1,976 transactions, approximately 1–2% of matched sales
0.3
For approximately 1–2% of matched Cook County transactions from 2018–2021, Cotality-reported sale prices differ from ground-truth sale prices by more than 5%. Output Quality negative Accuracy of broker-reported property sale prices
Reading fidelity high
Study strength high
n=129606
1,919 transactions, approximately 1–2% of matched sales
0.3
Brokered datasets exhibit coverage errors affecting roughly 12–15% of ground-truth transactions. Output Quality negative Coverage and representativeness of property transaction data
Reading fidelity high
Study strength high
n=153044
12–15% of ground-truth transactions
0.3
Nearly half of ATTOM match failures are attributable to broker misreporting of sale or property characteristics. Output Quality negative Transaction coverage resulting from reported sale and property characteristics
Reading fidelity high
Study strength high
n=18617
49.8% of ATTOM match failures
0.3
For ATTOM transactions with sale-price errors exceeding 5%, the conditional mean error was hundreds of thousands of dollars per transaction in each year from 2018 to 2021. Output Quality negative Magnitude of broker-reported sale-price error
Reading fidelity high
Study strength high
n=1976
$398,934 (2018); $166,999 (2019); $200,411 (2020); $162,269 (2021)
0.3
The predominant ATTOM sale-price discrepancies are exact ratios, especially two-thirds, one-third, and two times the ground-truth price. Output Quality negative Pattern and plausibility of sale-price misreporting
Reading fidelity high
Study strength high
n=1976
998 cases (50.51%) had a 2/3 ratio
0.3
The observed exact-ratio sale-price discrepancies are plausibly caused by brokers imputing prices from incorrect local transfer-tax rates or incorrect allocations of buyer and seller tax shares. Output Quality negative Validity of imputed property sale prices
Reading fidelity high
Study strength medium
n=75
0.18
Broker-derived estimates of property-tax assessment regressivity differ significantly depending on the data source because of broker data errors. Inequality negative Estimated regressivity of property-tax assessments
Reading fidelity high
Study strength medium
n=153044
0.18
County administrative sale-price data were highly accurate relative to underlying deed records in the authors’ validation sample. Output Quality positive Accuracy of county administrative sale-price records
Reading fidelity high
Study strength medium
n=74
99.8% overall accuracy for Cotality discrepancies; 99.9% overall accuracy for ATTOM discrepancies
0.18
Aggregate distributions of sale prices, assessed values, and assessment ratios appear comparable across ground-truth, ATTOM, and Cotality data despite substantial transaction-level discrepancies. Output Quality null_result Aggregate distributions of property prices, assessed values, and assessment ratios
Reading fidelity high
Study strength high
n=153044
0.3

Notes