1 cumulative citations
View corpus contextLinking population and family-registration databases lifted regional tax performance in three Indonesian districts — compliance rose by about a quarter and the tax base expanded by nearly a fifth while administrative costs fell by up to 42%. Machine-learning targeting (84% accuracy) and better forecasts (68% to 89%) amplified revenue gains, though legal, interoperability and capacity constraints may limit scale-up.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
1 cumulative citations
View corpus contextThis study examines the integration of individual population data with family registration records to optimize regional tax revenue in Indonesia. Using a mixed-methods approach across three districts (2021–2024), the findings show that integrated data systems increase tax compliance by 23–31% and expand the tax base by 18–26%. Predictive analytics using machine learning achieved 84% accuracy in identifying potential taxpayers and 78% precision in predicting payment behavior. Data integration also reduced administrative costs by 35–42% and improved revenue forecasting accuracy from 68% to 89%. Despite challenges related to legal frameworks, interoperability, and institutional capacity, integrated data systems enhance tax administration efficiency and support more equitable regional development planning.
Summary
Main Finding
Integrating individual population records with family-card (household) data and linking these to tax administration databases materially improves subnational revenue mobilization and planning. In three Indonesian districts (2021–2024) the integrated system (plus predictive analytics) raised tax compliance and the registered tax base, reduced administrative costs, and substantially improved revenue-forecast accuracy—effects attributed mainly to better taxpayer identification, risk-based enforcement, and data-informed planning.
Key Points
- Scope and scale
- Dataset: 2.23 million individual records, 587k family units, 1.47 million tax accounts across three districts (urban, peri-urban, rural).
- Integration match rates: 94.3% individual→household, 87.6% household→tax account.
- Quantitative outcomes (authors’ reported aggregates)
- Tax compliance increased by ~23–31%; registered tax base expanded by ~18–26%.
- Total local tax revenue rose by 41–48% across districts; difference-in-differences attributes ~31.7–37.4% of that growth to data integration.
- Administrative cost ratios fell by ~35–42% (examples: 18.6%→10.8% in District A).
- Revenue forecasting accuracy improved from ~68% to ~89% (MAPE fell ~32%→11%).
- Predictive analytics / model performance
- Models compared: Logistic Regression, Decision Tree, Random Forest, GBM, XGBoost.
- Best performer: XGBoost — accuracy 85.1%, precision 83.4%, recall 81.5%, AUC 0.921 (test set).
- Random Forest competitive (84.3% accuracy) with lower compute time.
- Household-derived variables (from family-card) accounted for ~42.7% of predictive power.
- Feature importance (top predictors of compliance)
- Household income proxy, prior payment history, education, property ownership, occupation, number of economically active members, geographic accessibility, business registration, age.
- Targeting and segmentation
- K-means segmentation produced six taxpayer clusters. High-value non-compliant segment = 8.3% of taxpayers but ~23.7% of potential revenue → prioritized for enforcement.
- Qualitative findings & barriers
- Benefits confirmed by 47 semi-structured interviews (tax/IT/planning officials).
- Key implementation challenges: legal/privacy frameworks, interoperability, data quality, and institutional capacity.
Data & Methods
- Research design
- Convergent mixed-methods: quantitative administrative analysis + predictive modeling + qualitative interviews.
- Three Indonesian districts chosen to span urban–peri–urban–rural contexts.
- Data integration pipeline
- Steps: identifier standardization, deterministic/probabilistic matching, household enrichment, linkage to tax records, geospatial/socioeconomic enrichment.
- Privacy: anonymization, role-based access, ethics approval, compliance with national data rules.
- Quantitative analysis
- Descriptive comparisons before/after integration and difference-in-differences with control areas to estimate causal effects.
- Dependent variables: collection efficiency (actual/potential), compliance rate, tax base size, forecasting accuracy (MAPE), administrative cost ratio.
- Controls: local GDP growth, inflation, unemployment, policy changes, administrative capacity.
- Predictive modeling
- Split: training 70% / validation 15% / test 15% (≈1.56M training records).
- Algorithms: Logistic Regression, Decision Tree, Random Forest, GBM, XGBoost; hyperparameter tuning via grid search + cross-validation.
- Evaluation metrics: accuracy, precision, recall, F1, AUC.
- Forecasting & network methods
- Time series: ARIMA and Prophet (with demographic-fiscal covariates).
- Network analysis: graph methods and community detection to spot linked evasion schemes and high-value networks.
- Qualitative
- 47 interviews; thematic analysis to identify success factors and institutional constraints.
Implications for AI Economics
- Fiscal capacity and subnational public finance
- Integrated administrative data + ML materially expand the tax base, increase collections, and make subnational revenues more predictable — strengthening fiscal autonomy and enabling better multi-year development planning.
- More predictable revenues reduce financing risk for public investments and can improve allocative efficiency (needs-based targeting, spatial equity).
- Value of non-fiscal administrative data
- Household-level registers (family cards) are high-value inputs for economic prediction tasks; adding demographic and asset proxies substantially improves model performance beyond standard tax records.
- Targeting efficiency and enforcement design
- ML-enabled segmentation supports risk-based enforcement that concentrates resources on small groups that hold outsized revenue potential (improves cost-effectiveness of audits/enforcement).
- Operational and institutional considerations
- Gains depend on data interoperability, identifier quality, IT capacity, and legal/privacy frameworks; investments in these areas are preconditions for scale-up.
- Computational choices matter: tree-based ensemble methods (XGBoost/RF) offered best predictive trade-offs here; cost/latency considerations may favor RF in constrained settings.
- Governance, fairness, and risks
- Integrating administrative data raises privacy and surveillance risks, and can entrench bias if models mirror unequal reporting or registration patterns. Transparency, interpretability, and auditability of models are essential.
- There are potential behavioral feedbacks: taxpayers may change behavior in response to targeting; adversarial adaptation (evasion tactics) can erode model effectiveness over time—requiring ongoing monitoring and model re-training.
- Research directions for AI economics
- Evaluate long-run behavioral responses and general equilibrium effects of data-driven enforcement.
- Cost–benefit and distributional analyses: who bears compliance costs, and how do revenue gains affect public service provision and equity?
- Robustness and fairness audits across socioeconomic groups and regions; causal identification of effects in varied institutional contexts.
- Transferability studies: replicating results across countries with different registration completeness, informality, and legal frameworks.
- Policy prescription (brief)
- Prioritize: (1) secure, interoperable ID systems and metadata standards; (2) capacity building for tax authorities in data science and governance; (3) legal frameworks protecting privacy while allowing lawful data use; (4) operational plans for model governance (monitoring, fairness checks, update cycles).
Limitations noted by the authors: context-specific study (three districts in Indonesia), reliance on administrative records with remaining unmatched records and data-quality issues, and institutional heterogeneity that may affect external validity.
Assessment
Claims (9)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Integrated data systems increase tax compliance by 23–31%. Fiscal And Macroeconomic | positive | tax compliance |
Reading fidelity
high
Study strength
medium
|
n=3
23–31% increase
|
| Integrated data systems expand the tax base by 18–26%. Fiscal And Macroeconomic | positive | tax base size |
Reading fidelity
high
Study strength
medium
|
n=3
18–26% increase
|
| Predictive analytics using machine learning achieved 84% accuracy in identifying potential taxpayers. Decision Quality | positive | classification accuracy for identifying potential taxpayers |
Reading fidelity
high
Study strength
medium
|
84% accuracy
|
| Predictive analytics achieved 78% precision in predicting payment behavior. Decision Quality | positive | precision of payment-behavior predictions |
Reading fidelity
high
Study strength
medium
|
78% precision
|
| Data integration reduced administrative costs by 35–42%. Organizational Efficiency | positive | administrative costs |
Reading fidelity
high
Study strength
medium
|
n=3
35–42% reduction
|
| Revenue forecasting accuracy improved from 68% to 89% after data integration. Fiscal And Macroeconomic | positive | revenue forecasting accuracy |
Reading fidelity
high
Study strength
medium
|
n=3
forecasting accuracy improved from 68% to 89%
|
| Despite benefits, the integration faced challenges related to legal frameworks, interoperability, and institutional capacity. Governance And Regulation | negative | barriers to data integration (legal frameworks, interoperability, institutional capacity) |
Reading fidelity
high
Study strength
medium
|
n=3
|
| Integrated data systems enhance tax administration efficiency and support more equitable regional development planning. Governance And Regulation | positive | tax administration efficiency and regional development planning equity |
Reading fidelity
high
Study strength
medium
|
n=3
|
| The study used a mixed-methods approach across three districts (2021–2024). Other | null_result | study design / methodology |
Reading fidelity
high
Study strength
high
|
n=3
|