0 cumulative citations
View corpus contextCommercial property-data brokers misreport and omit substantial transaction data: 1–2% of matched sales show large, systematic price errors and 12–15% of transactions are missing or misclassified, producing meaningful differences in estimated property‑tax regressivity.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Big data are the foundation for an increasing share of academic research and AI models deployed in both the public and private sectors, prompting substantial growth over time in reliance on brokered datasets. Brokered property records, which are ubiquitous in studies of gentrification, inequality, and the property tax in the U.S. and serve as inputs to property valuation models, are one notable example. In this paper, we audit two prominent brokered property datasets, finding errors in these data which bias key measures of economic inequality. First, we document that for 1-2% of matched sales in Cook County, IL, from 2018-2021, broker-provided sale prices differ from ground truth sale prices by more than 5%. Moreover, missing data and conceptual differences in the reporting of deed and property characteristics lead to coverage errors ranging from 12 to 15% of transactions. Second, we show that misreporting is highly consistent between brokers: more often than not, brokers make identical reporting errors for the same transactions. Third, to illustrate the significance of these errors, we measure their impact on estimates of property tax regressivity, finding that they drive significant wedges between estimates depending on the data source. These findings generalize to two other large counties in the U.S., and highlight the crucial importance of open administrative data and transparency from brokers regarding data provenance and lineage.
Summary
Main Finding
Brokered property datasets widely used in research and ML (ATTOM and Cotality) contain systematic coverage and imputation errors relative to county administrative “ground-truth” records. These errors (1) cause 1–2% of matched transactions to have sale prices differing by >5%—often by simple integer ratios (e.g., 2/3, 1/3, 2)—(2) produce coverage losses on the order of 12–15% of transactions (missing records, misclassified sale/property filters, duplicates), and (3) are highly correlated across brokers (errors often shared), producing meaningful biases in downstream economic estimates such as property‑tax regressivity. The problems generalize beyond Cook County (robustness checks in NYC and Philadelphia).
Key Points
- Two failure modes:
- Imputation errors: broker sale prices sometimes differ substantially from county deed prices (1–2% of matched sales; many discrepancies are exact/simple multiples).
- Coverage errors: 12–15% of ground-truth arms‑length single‑family sales are not represented as such in brokered data due to missing records or misreported filters (e.g., marked as foreclosure, condo, multi‑parcel).
- Quantitative highlights (Cook County, 2018–2021):
- Ground-truth sales (after filters): ~153k transactions.
- ATTOM matched ~134,427 transactions (≈87.8% of ground truth); Cotality matched ~129,606 (≈84.7%).
- ATTOM: 1,976 matched transactions had sale-price discrepancies >5%; distribution of ratios of broker to ground-truth price is dominated by 2/3 (≈50% of discrepant cases), plus 2.0, 1/3, etc.
- Sources of non-match (ATTOM): ~37% missing record, ~50% misreported filter, ~8.6% duplicates, ~4% fuzzy-match thresholds. (Cotality had more missing records and fewer misreported filters.)
- Discrepancies sometimes amount to hundreds of thousands of dollars (conditional mean errors large for discrepant subset).
- Mechanism: Evidence points to imputation via municipal transfer‑tax rates. If brokers impute sale price from reported transfer‑tax payments but use the wrong municipality’s rate or the wrong payer’s share, the imputed price will be an exact multiple of the true price (matching the simple ratios observed).
- Shared errors: Many of the same transactions are misreported by both brokers, consistent with broker-to-broker data resale/propagation or common imputation routines.
- Validation: Authors treated county administrative data as ground truth, and manually checked a random 2% sample of broker–county disagreements against deed records; county prices matched deed records in the overwhelming majority of sampled disagreements (Cook County accuracy ≈99.8–99.9% in sampled checks).
- Downstream impact: These errors materially change estimates of property-tax assessment regressivity (the paper shows significant wedges in regressivity estimates depending on data source).
Data & Methods
- Data sources:
- Brokered datasets: ATTOM and Cotality.
- Ground-truth: county administrative transaction and assessment data (Cook County), with robustness checks using New York City and Philadelphia.
- Sample period: 2018–2021.
- Filters applied (to create comparable samples): arms‑length transactions of single‑family homes; exclude sales < $10k, multi‑parcel sales, repeated transactions in the same year, and records missing address/date/price.
- Matching procedure:
- Matched records using address, latitude/longitude, and sale date (fuzzy matching thresholds applied; robustness tests for thresholds are reported).
- Matched counts and unmatched-cause coding (manual classification of reasons for no-match).
- Error classification:
- Imputation/accuracy error: difference in sale price or assessed value between broker and county records; flagged >5% absolute difference.
- Coverage/representativeness error: missing records, misreported sale/property characteristics leading to exclusion from filtered sample, duplicates.
- Validation:
- Manual verification of a random 2% sample of broker–county disagreements against deed records to confirm county record accuracy.
- Analysis of mechanism:
- Examined distribution of broker/ground-truth price ratios; identified spikes at simple rational multiples consistent with transfer‑tax imputation mistakes.
- Impact analysis:
- Measured effect on property-tax assessment regressivity estimates across data sources and years to quantify substantive consequences.
Implications for AI Economics
- Data provenance and lineage matter for economic inference and for ML systems that use brokered administrative-like data. Errors in upstream broker processing (imputation, harmonization) propagate into models and policy metrics.
- Multiple brokers does not guarantee robustness: because brokers often share sources and resale data, errors are frequently correlated across brokers; naive ensemble/“cross-check” with another broker may provide a false sense of security.
- High‑stakes or policy-relevant analyses that rely on brokered property data should:
- Prefer open administrative sources where available; treat broker data as convenience products requiring validation.
- Implement routine validation and auditing steps: sample verification against deeds/administrative records, distributional checks, and search for simple-ratio patterns that indicate systematic imputation errors.
- Document data provenance, cleaning, and imputation steps (both for reproducibility and for assessing bias).
- Run sensitivity analyses (e.g., recompute key estimates dropping suspect transactions, or adding uncertainty bounds for imputed prices).
- Demand greater transparency from brokers about lineage and imputation procedures; when contracting with brokers, require disclosure of methods and access to source identifiers.
- For ML pipelines:
- Incorporate data-quality checks (unit tests, anomaly detection for simple-ratio imputation artifacts), propagate uncertainty from input data to downstream model outputs, and prioritize retraining/updates when administrative corrections are published.
- For research in AI economics:
- Be cautious interpreting precise-sounding estimates from large brokered datasets: the “Big Data Paradox” can produce precise but biased estimates if upstream errors are overlooked.
- When possible, publish code and matching procedures so others can audit sensitivity to data source and matching thresholds.
- Policy: Governments and funders should invest in making administrative data more accessible and standardized to reduce reliance on opaque brokers and to improve the evidence base used by researchers, appraisers, and automated systems.
If you want, I can extract the paper’s key tables (match rates, ratio distributions, and match-failure breakdowns) into a compact CSV-style summary or produce a checklist for auditing brokered property datasets you or your team could run on other jurisdictions.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| For approximately 1–2% of matched Cook County transactions from 2018–2021, ATTOM-reported sale prices differ from ground-truth sale prices by more than 5%. Output Quality | negative | Accuracy of broker-reported property sale prices |
Reading fidelity
high
Study strength
high
|
n=134427
1,976 transactions, approximately 1–2% of matched sales
|
| For approximately 1–2% of matched Cook County transactions from 2018–2021, Cotality-reported sale prices differ from ground-truth sale prices by more than 5%. Output Quality | negative | Accuracy of broker-reported property sale prices |
Reading fidelity
high
Study strength
high
|
n=129606
1,919 transactions, approximately 1–2% of matched sales
|
| Brokered datasets exhibit coverage errors affecting roughly 12–15% of ground-truth transactions. Output Quality | negative | Coverage and representativeness of property transaction data |
Reading fidelity
high
Study strength
high
|
n=153044
12–15% of ground-truth transactions
|
| Nearly half of ATTOM match failures are attributable to broker misreporting of sale or property characteristics. Output Quality | negative | Transaction coverage resulting from reported sale and property characteristics |
Reading fidelity
high
Study strength
high
|
n=18617
49.8% of ATTOM match failures
|
| For ATTOM transactions with sale-price errors exceeding 5%, the conditional mean error was hundreds of thousands of dollars per transaction in each year from 2018 to 2021. Output Quality | negative | Magnitude of broker-reported sale-price error |
Reading fidelity
high
Study strength
high
|
n=1976
$398,934 (2018); $166,999 (2019); $200,411 (2020); $162,269 (2021)
|
| The predominant ATTOM sale-price discrepancies are exact ratios, especially two-thirds, one-third, and two times the ground-truth price. Output Quality | negative | Pattern and plausibility of sale-price misreporting |
Reading fidelity
high
Study strength
high
|
n=1976
998 cases (50.51%) had a 2/3 ratio
|
| The observed exact-ratio sale-price discrepancies are plausibly caused by brokers imputing prices from incorrect local transfer-tax rates or incorrect allocations of buyer and seller tax shares. Output Quality | negative | Validity of imputed property sale prices |
Reading fidelity
high
Study strength
medium
|
n=75
|
| Broker-derived estimates of property-tax assessment regressivity differ significantly depending on the data source because of broker data errors. Inequality | negative | Estimated regressivity of property-tax assessments |
Reading fidelity
high
Study strength
medium
|
n=153044
|
| County administrative sale-price data were highly accurate relative to underlying deed records in the authors’ validation sample. Output Quality | positive | Accuracy of county administrative sale-price records |
Reading fidelity
high
Study strength
medium
|
n=74
99.8% overall accuracy for Cotality discrepancies; 99.9% overall accuracy for ATTOM discrepancies
|
| Aggregate distributions of sale prices, assessed values, and assessment ratios appear comparable across ground-truth, ATTOM, and Cotality data despite substantial transaction-level discrepancies. Output Quality | null_result | Aggregate distributions of property prices, assessed values, and assessment ratios |
Reading fidelity
high
Study strength
high
|
n=153044
|