The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Configured LLMs speed routine VAT analyses and draft legally grounded justifications for tax professionals, but persistent hallucination and missing client context limit them to decision-support rather than unsupervised advice.

Using large language models for legal decision-making in Austrian value-added tax law: a comparative study
Marina Luketina, Andrea Benkel, Christoph G. Schuetz · August 31, 2026 · Journal of Business Analytics
openalex descriptive medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Marina Luketina provider ID
  2. Andrea Benkel provider ID
  3. Christoph G. Schuetz provider ID

Semantic Scholar

Latest observation:

  1. Marina Luketina provider ID
  2. Andrea Benkel provider ID
  3. Christoph G. Schuetz provider ID
Fine-tuned and retrieval-augmented LLMs can produce legally grounded initial analyses that meaningfully support Austrian/EU VAT advisory work, but hallucinations and gaps in client-specific context prevent safe full automation.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Background Tax consulting clients often describe cases in natural language, making large language models (LLMs) attractive for supporting legal decision-making in Austrian and EU value-added tax (VAT) law. However, the requirement for legally grounded, well-justified analyses is challenged by LLMs’ propensity to hallucinate.Methods This study experimentally evaluates two common approaches for enhancing LLM performance—fine-tuning and retrieval-augmented generation (RAG)—applied to both textbook VAT cases and real-world cases from a tax consulting firm. The aim is to identify optimal configurations of LLM-based systems and assess their legal-reasoning capabilities.Results Properly configured LLMs can effectively support tax professionals in VAT-related tasks, automating routine work, providing initial analyses, and generating legally grounded justifications for decisions. However, current prototypes are not yet ready for full automation given the sensitivity of the legal domain, and limitations persist in handling implicit client knowledge and context-specific documentation.Conclusion LLMs show strong potential to assist tax consultants by reducing workload and supporting VAT decision-making, though future work should focus on integrating structured background information to address remaining gaps in contextual understanding.

Summary

Main Finding

Properly configured large language models (LLMs)—using fine-tuning and/or retrieval-augmented generation (RAG)—can meaningfully support tax professionals on Austrian and EU VAT matters by automating routine analyses, producing initial case assessments, and generating legally grounded justifications. However, current prototypes are not yet reliable enough for full automation because they struggle with hallucination, implicit client knowledge, and context-specific documentation.

Key Points

  • Objective: Experimentally evaluate two common LLM-improvement strategies (fine-tuning and RAG) on textbook VAT problems and real-world tax-consulting cases to assess legal-reasoning performance and find effective system configurations.
  • Data: Two corpora were used — canonical textbook VAT cases (for known legal issues) and authentic cases collected from a tax consulting firm (to evaluate real-world applicability).
  • Approaches tested: fine-tuning LLMs on domain material; RAG that retrieves relevant statutes, case law, and firm documents at query time; and combinations/configurations thereof.
  • Outcomes:
    • Properly configured models can produce useful, legally grounded analyses suitable for preliminary advice and supporting professionals in routine tasks.
    • Hallucination remains a key risk; RAG and grounding mechanisms reduce but do not eliminate unsupported assertions.
    • Models have persistent limits in incorporating implicit client knowledge and interpreting case-specific documents not present in the retrieval corpus.
  • Practical status: Useful as decision-support tools and for automating low-risk routine work, but not yet safe for unsupervised, high-stakes legal decisions.

Data & Methods

  • Data sources:
    • Textbook VAT cases that capture canonical legal issues and standard reasoning paths.
    • Real-world cases from a tax consulting firm, including natural-language client descriptions and associated documentation.
  • Experimental interventions:
    • Fine-tuning LLMs on VAT-focused corpora to adapt language, reasoning patterns, and domain knowledge.
    • Retrieval-augmented generation (RAG): using an external, searchable knowledge base (statutes, guidance, firm precedents) to ground model outputs.
    • Testing mixed configurations (e.g., fine-tuned base model + RAG).
  • Evaluation criteria:
    • Legal grounding (citation accuracy and relevance).
    • Quality and coherence of legal justification.
    • Ability to handle implicit context and document-specific facts.
    • Practical utility for tax professionals (time saved, error rates, appropriateness of automation).
  • Findings summary:
    • Grounding via retrieval materially improves justifications and reduces hallucinations.
    • Fine-tuning improves stylistic and domain-aligned output but does not on its own eliminate unsupported claims.
    • Remaining failures often stem from missing structured background (client specifics, non-digitized documents, unstated assumptions).

Implications for AI Economics

  • Productivity and task shifting:
    • LLMs can raise productivity in tax advisory services by automating repetitive analyses and drafting, shifting human effort toward higher-complexity, judgement-heavy tasks.
    • Demand may shift toward roles focused on oversight, client elicitation, and handling complex exceptions.
  • Labor and wages:
    • Potential reduction in time spent on routine tasks could compress billable hours for junior staff but increase value capture for senior experts who supervise AI outputs.
    • Firms that invest in safe, well-integrated LLM systems may gain competitive advantage.
  • Investment and complementary capital:
    • Realizing benefits requires investment in structured data infrastructure (document ingestion, client profiles, knowledge bases) and retrieval pipelines to mitigate hallucinations.
    • Ongoing costs include model maintenance, fine-tuning, and compliance/legal review.
  • Risk, regulation, and liability:
    • Hallucination and context-missing errors create legal and reputational risk; human-in-the-loop and audit trails are economically necessary safeguards.
    • Regulators and professional bodies may need to clarify liability, disclosure, and quality standards for AI-assisted legal advice.
  • Adoption constraints and market structure:
    • Smaller firms may face barriers to adoption due to data and engineering costs; platform providers offering turnkey, legally grounded RAG stacks could consolidate market share.
    • Standardization of legal knowledge representations (structured statutes, precedents) would lower costs and speed safe deployment.
  • Research and policy priorities:
    • Focus on methods that integrate structured client data and provenance-aware retrieval.
    • Economic assessments should quantify trade-offs between automation gains and increased compliance/monitoring costs.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper reports controlled experimental comparisons across modeling strategies (fine-tuning, RAG, and combinations) on both canonical textbook VAT problems and authentic firm cases, showing consistent improvements from retrieval grounding and stylistic gains from fine-tuning. However, evidence is limited to prototype evaluations rather than large-scale field deployments or randomized trials, sample sizes and quantitative effect-metrics are not specified, and remaining failure modes (hallucination, missing client context) limit claims about safe automation. Methods Rigormedium — The study uses an appropriate experimental design for an ML/AI evaluation (multiple interventions, mixed datasets, and relevant evaluation criteria including citation accuracy and practical utility). Rigor is reduced by omission of key details in the supplied text (sample sizes, annotation procedures, inter-rater reliability, blind evaluation or independent replication), no randomized field validation of productivity gains, and likely reliance on subjective judgments for some outcomes. SampleTwo corpora: (1) canonical textbook VAT cases representing standard legal issues and reasoning paths; (2) authentic real-world tax-consulting cases collected from a single tax consulting firm, comprising natural-language client descriptions and associated documentation. The supplied summary does not report sample sizes, case mix, or representativeness. Themesproductivity human_ai_collab skills_training adoption org_design governance GeneralizabilitySingle-jurisdiction focus (Austrian and EU VAT) limits transferability to other legal systems and tax regimes, Real-world cases come from one consulting firm, so firm-level practices and documentation standards may bias results, Unclear sample size and case complexity distributions reduce confidence in applicability to high-stakes or rare cases, Performance depends on availability and quality of digitized, structured client data and retrieval corpora, which varies across firms, Prototype results may not generalize to other areas of law or to non-legal professional services without further validation

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Properly configured LLMs using fine-tuning and/or retrieval-augmented generation can support tax professionals on Austrian and EU VAT matters by automating routine analyses, producing initial case assessments, and generating legally grounded justifications. Organizational Efficiency positive Usefulness and quality of LLM-generated VAT analyses and legal justifications
Reading fidelity high
Study strength medium
not reported
0.18
Retrieval-based grounding materially improves the quality of legal justifications and reduces hallucinations, but does not eliminate unsupported assertions. Decision Quality positive Legal justification quality and frequency of unsupported assertions
Reading fidelity high
Study strength medium
not reported
0.18
Fine-tuning improves stylistic and domain-aligned output but does not by itself eliminate unsupported claims. Output Quality mixed Domain alignment, output style, and unsupported-claim rate
Reading fidelity high
Study strength medium
not reported
0.18
Current LLM prototypes are not reliable enough for fully automated, high-stakes VAT or legal decisions. Ai Safety And Ethics negative Reliability and suitability for unsupervised high-stakes legal decision-making
Reading fidelity high
Study strength medium
not reported
0.18
LLMs have persistent limitations in incorporating implicit client knowledge and interpreting case-specific documents that are not present in the retrieval corpus. Decision Quality negative Ability to use implicit context and case-specific documentation
Reading fidelity high
Study strength medium
not reported
0.18
LLMs can raise productivity in tax advisory services by automating repetitive analyses and drafting, while shifting human effort toward complex, judgment-heavy tasks. Organizational Efficiency positive Productivity and allocation of human effort in tax advisory work
Reading fidelity high
Study strength speculative
not reported
0.03
The use of LLMs may reduce time spent on routine tasks, potentially compressing billable hours for junior staff while increasing the value captured by senior experts who supervise AI outputs. Wages mixed Routine-task time, billable hours, and value capture across staff seniority levels
Reading fidelity high
Study strength speculative
not reported
0.03
Realizing productivity benefits from LLMs in tax advisory requires investment in structured data infrastructure, document ingestion, client profiles, knowledge bases, and retrieval pipelines. Organizational Efficiency positive Infrastructure requirements for effective and safe LLM deployment
Reading fidelity high
Study strength medium
not reported
0.18
Hallucination and context-missing errors create legal and reputational risks, making human oversight and audit trails necessary safeguards. Ai Safety And Ethics negative Legal, reputational, and governance risk from AI-assisted tax advice
Reading fidelity high
Study strength medium
not reported
0.18
Smaller tax-consulting firms may face greater barriers to adopting LLM systems because of data and engineering costs, while turnkey legally grounded RAG platforms could consolidate market share. Market Structure negative Adoption barriers and potential market-share concentration
Reading fidelity high
Study strength speculative
not reported
0.03

Notes