The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Linking population and family-registration databases lifted regional tax performance in three Indonesian districts — compliance rose by about a quarter and the tax base expanded by nearly a fifth while administrative costs fell by up to 42%. Machine-learning targeting (84% accuracy) and better forecasts (68% to 89%) amplified revenue gains, though legal, interoperability and capacity constraints may limit scale-up.

Integration of Individual Data and Family Cards in Optimizing Tax Revenue: a Public Accounting and Predictive Analytics Approach to Regional Development Planning
Sarsiti Sarsiti, Tamam Rosid, Juli Prastyorini, Putri Nilam Aisyah · February 04, 2026 · International Journal of Finance and Business Management
openalex descriptive medium evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Sarsiti Sarsiti provider ID
  2. Tamam Rosid provider ID
  3. Juli Prastyorini provider ID
  4. Putri Nilam Aisyah provider ID

Semantic Scholar

Latest observation:

  1. Sarsiti - Sarsiti provider ID
  2. Tamam Rosid provider ID
  3. Juli Prastyorini provider ID
  4. P. Aisyah provider ID
Integrating population and family registration data in three Indonesian districts improved tax compliance (23–31%), expanded the tax base (18–26%), cut administrative costs (35–42%), and—with ML targeting that was 84% accurate—boosted revenue-forecasting accuracy from 68% to 89%.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

This study examines the integration of individual population data with family registration records to optimize regional tax revenue in Indonesia. Using a mixed-methods approach across three districts (2021–2024), the findings show that integrated data systems increase tax compliance by 23–31% and expand the tax base by 18–26%. Predictive analytics using machine learning achieved 84% accuracy in identifying potential taxpayers and 78% precision in predicting payment behavior. Data integration also reduced administrative costs by 35–42% and improved revenue forecasting accuracy from 68% to 89%. Despite challenges related to legal frameworks, interoperability, and institutional capacity, integrated data systems enhance tax administration efficiency and support more equitable regional development planning.

Summary

Main Finding

Integrating individual population records with family-card (household) data and linking these to tax administration databases materially improves subnational revenue mobilization and planning. In three Indonesian districts (2021–2024) the integrated system (plus predictive analytics) raised tax compliance and the registered tax base, reduced administrative costs, and substantially improved revenue-forecast accuracy—effects attributed mainly to better taxpayer identification, risk-based enforcement, and data-informed planning.

Key Points

  • Scope and scale
    • Dataset: 2.23 million individual records, 587k family units, 1.47 million tax accounts across three districts (urban, peri-urban, rural).
    • Integration match rates: 94.3% individual→household, 87.6% household→tax account.
  • Quantitative outcomes (authors’ reported aggregates)
    • Tax compliance increased by ~23–31%; registered tax base expanded by ~18–26%.
    • Total local tax revenue rose by 41–48% across districts; difference-in-differences attributes ~31.7–37.4% of that growth to data integration.
    • Administrative cost ratios fell by ~35–42% (examples: 18.6%→10.8% in District A).
    • Revenue forecasting accuracy improved from ~68% to ~89% (MAPE fell ~32%→11%).
  • Predictive analytics / model performance
    • Models compared: Logistic Regression, Decision Tree, Random Forest, GBM, XGBoost.
    • Best performer: XGBoost — accuracy 85.1%, precision 83.4%, recall 81.5%, AUC 0.921 (test set).
    • Random Forest competitive (84.3% accuracy) with lower compute time.
    • Household-derived variables (from family-card) accounted for ~42.7% of predictive power.
  • Feature importance (top predictors of compliance)
    • Household income proxy, prior payment history, education, property ownership, occupation, number of economically active members, geographic accessibility, business registration, age.
  • Targeting and segmentation
    • K-means segmentation produced six taxpayer clusters. High-value non-compliant segment = 8.3% of taxpayers but ~23.7% of potential revenue → prioritized for enforcement.
  • Qualitative findings & barriers
    • Benefits confirmed by 47 semi-structured interviews (tax/IT/planning officials).
    • Key implementation challenges: legal/privacy frameworks, interoperability, data quality, and institutional capacity.

Data & Methods

  • Research design
    • Convergent mixed-methods: quantitative administrative analysis + predictive modeling + qualitative interviews.
    • Three Indonesian districts chosen to span urban–peri–urban–rural contexts.
  • Data integration pipeline
    • Steps: identifier standardization, deterministic/probabilistic matching, household enrichment, linkage to tax records, geospatial/socioeconomic enrichment.
    • Privacy: anonymization, role-based access, ethics approval, compliance with national data rules.
  • Quantitative analysis
    • Descriptive comparisons before/after integration and difference-in-differences with control areas to estimate causal effects.
    • Dependent variables: collection efficiency (actual/potential), compliance rate, tax base size, forecasting accuracy (MAPE), administrative cost ratio.
    • Controls: local GDP growth, inflation, unemployment, policy changes, administrative capacity.
  • Predictive modeling
    • Split: training 70% / validation 15% / test 15% (≈1.56M training records).
    • Algorithms: Logistic Regression, Decision Tree, Random Forest, GBM, XGBoost; hyperparameter tuning via grid search + cross-validation.
    • Evaluation metrics: accuracy, precision, recall, F1, AUC.
  • Forecasting & network methods
    • Time series: ARIMA and Prophet (with demographic-fiscal covariates).
    • Network analysis: graph methods and community detection to spot linked evasion schemes and high-value networks.
  • Qualitative
    • 47 interviews; thematic analysis to identify success factors and institutional constraints.

Implications for AI Economics

  • Fiscal capacity and subnational public finance
    • Integrated administrative data + ML materially expand the tax base, increase collections, and make subnational revenues more predictable — strengthening fiscal autonomy and enabling better multi-year development planning.
    • More predictable revenues reduce financing risk for public investments and can improve allocative efficiency (needs-based targeting, spatial equity).
  • Value of non-fiscal administrative data
    • Household-level registers (family cards) are high-value inputs for economic prediction tasks; adding demographic and asset proxies substantially improves model performance beyond standard tax records.
  • Targeting efficiency and enforcement design
    • ML-enabled segmentation supports risk-based enforcement that concentrates resources on small groups that hold outsized revenue potential (improves cost-effectiveness of audits/enforcement).
  • Operational and institutional considerations
    • Gains depend on data interoperability, identifier quality, IT capacity, and legal/privacy frameworks; investments in these areas are preconditions for scale-up.
    • Computational choices matter: tree-based ensemble methods (XGBoost/RF) offered best predictive trade-offs here; cost/latency considerations may favor RF in constrained settings.
  • Governance, fairness, and risks
    • Integrating administrative data raises privacy and surveillance risks, and can entrench bias if models mirror unequal reporting or registration patterns. Transparency, interpretability, and auditability of models are essential.
    • There are potential behavioral feedbacks: taxpayers may change behavior in response to targeting; adversarial adaptation (evasion tactics) can erode model effectiveness over time—requiring ongoing monitoring and model re-training.
  • Research directions for AI economics
    • Evaluate long-run behavioral responses and general equilibrium effects of data-driven enforcement.
    • Cost–benefit and distributional analyses: who bears compliance costs, and how do revenue gains affect public service provision and equity?
    • Robustness and fairness audits across socioeconomic groups and regions; causal identification of effects in varied institutional contexts.
    • Transferability studies: replicating results across countries with different registration completeness, informality, and legal frameworks.
  • Policy prescription (brief)
    • Prioritize: (1) secure, interoperable ID systems and metadata standards; (2) capacity building for tax authorities in data science and governance; (3) legal frameworks protecting privacy while allowing lawful data use; (4) operational plans for model governance (monitoring, fairness checks, update cycles).

Limitations noted by the authors: context-specific study (three districts in Indonesia), reliance on administrative records with remaining unmatched records and data-quality issues, and institutional heterogeneity that may affect external validity.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study reports substantial administrative outcomes (compliance, tax base, cost reductions, forecasting accuracy) and presents predictive-model performance, which lend credibility; however, it lacks a clear causal identification strategy (no randomization or strong quasi-experimental design reported), covers only three districts, and may be subject to selection and confounding, limiting causal claims. Methods Rigormedium — Uses mixed methods and machine-learning predictive analytics with concrete performance metrics, and assesses administrative outcomes, but the description lacks detail on model validation (out-of-sample testing, cross-validation, feature robustness), causal inference methods, counterfactuals, sample sizes, and robustness checks; qualitative components and how cost savings were measured are not fully described. SampleAdministrative integration pilot across three Indonesian districts (2021–2024) combining individual population data and family registration records, linked to regional tax records and payment histories; includes machine-learning models trained to identify potential taxpayers and predict payment behavior; exact sample sizes not reported. Themesgovernance org_design GeneralizabilityLimited to three districts in Indonesia — findings may not generalize to other regions or countries, Pilot districts may have had atypical institutional capacity, political support, or data quality, biasing results, Legal and regulatory context in Indonesia affects feasibility and may differ elsewhere, Short evaluation window (2021–2024) may not capture long-run effects or sustainability, ML model accuracy and precision depend on local data patterns and may not transfer to other datasets or jurisdictions

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Integrated data systems increase tax compliance by 23–31%. Fiscal And Macroeconomic positive tax compliance
Reading fidelity high
Study strength medium
n=3
23–31% increase
0.18
Integrated data systems expand the tax base by 18–26%. Fiscal And Macroeconomic positive tax base size
Reading fidelity high
Study strength medium
n=3
18–26% increase
0.18
Predictive analytics using machine learning achieved 84% accuracy in identifying potential taxpayers. Decision Quality positive classification accuracy for identifying potential taxpayers
Reading fidelity high
Study strength medium
84% accuracy
0.18
Predictive analytics achieved 78% precision in predicting payment behavior. Decision Quality positive precision of payment-behavior predictions
Reading fidelity high
Study strength medium
78% precision
0.18
Data integration reduced administrative costs by 35–42%. Organizational Efficiency positive administrative costs
Reading fidelity high
Study strength medium
n=3
35–42% reduction
0.18
Revenue forecasting accuracy improved from 68% to 89% after data integration. Fiscal And Macroeconomic positive revenue forecasting accuracy
Reading fidelity high
Study strength medium
n=3
forecasting accuracy improved from 68% to 89%
0.18
Despite benefits, the integration faced challenges related to legal frameworks, interoperability, and institutional capacity. Governance And Regulation negative barriers to data integration (legal frameworks, interoperability, institutional capacity)
Reading fidelity high
Study strength medium
n=3
0.18
Integrated data systems enhance tax administration efficiency and support more equitable regional development planning. Governance And Regulation positive tax administration efficiency and regional development planning equity
Reading fidelity high
Study strength medium
n=3
0.18
The study used a mixed-methods approach across three districts (2021–2024). Other null_result study design / methodology
Reading fidelity high
Study strength high
n=3
0.3

Notes