0 cumulative citations
View corpus contextConfigured LLMs speed routine VAT analyses and draft legally grounded justifications for tax professionals, but persistent hallucination and missing client context limit them to decision-support rather than unsupervised advice.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
0 cumulative citations
View corpus contextBackground Tax consulting clients often describe cases in natural language, making large language models (LLMs) attractive for supporting legal decision-making in Austrian and EU value-added tax (VAT) law. However, the requirement for legally grounded, well-justified analyses is challenged by LLMs’ propensity to hallucinate.Methods This study experimentally evaluates two common approaches for enhancing LLM performance—fine-tuning and retrieval-augmented generation (RAG)—applied to both textbook VAT cases and real-world cases from a tax consulting firm. The aim is to identify optimal configurations of LLM-based systems and assess their legal-reasoning capabilities.Results Properly configured LLMs can effectively support tax professionals in VAT-related tasks, automating routine work, providing initial analyses, and generating legally grounded justifications for decisions. However, current prototypes are not yet ready for full automation given the sensitivity of the legal domain, and limitations persist in handling implicit client knowledge and context-specific documentation.Conclusion LLMs show strong potential to assist tax consultants by reducing workload and supporting VAT decision-making, though future work should focus on integrating structured background information to address remaining gaps in contextual understanding.
Summary
Main Finding
Properly configured large language models (LLMs)—using fine-tuning and/or retrieval-augmented generation (RAG)—can meaningfully support tax professionals on Austrian and EU VAT matters by automating routine analyses, producing initial case assessments, and generating legally grounded justifications. However, current prototypes are not yet reliable enough for full automation because they struggle with hallucination, implicit client knowledge, and context-specific documentation.
Key Points
- Objective: Experimentally evaluate two common LLM-improvement strategies (fine-tuning and RAG) on textbook VAT problems and real-world tax-consulting cases to assess legal-reasoning performance and find effective system configurations.
- Data: Two corpora were used — canonical textbook VAT cases (for known legal issues) and authentic cases collected from a tax consulting firm (to evaluate real-world applicability).
- Approaches tested: fine-tuning LLMs on domain material; RAG that retrieves relevant statutes, case law, and firm documents at query time; and combinations/configurations thereof.
- Outcomes:
- Properly configured models can produce useful, legally grounded analyses suitable for preliminary advice and supporting professionals in routine tasks.
- Hallucination remains a key risk; RAG and grounding mechanisms reduce but do not eliminate unsupported assertions.
- Models have persistent limits in incorporating implicit client knowledge and interpreting case-specific documents not present in the retrieval corpus.
- Practical status: Useful as decision-support tools and for automating low-risk routine work, but not yet safe for unsupervised, high-stakes legal decisions.
Data & Methods
- Data sources:
- Textbook VAT cases that capture canonical legal issues and standard reasoning paths.
- Real-world cases from a tax consulting firm, including natural-language client descriptions and associated documentation.
- Experimental interventions:
- Fine-tuning LLMs on VAT-focused corpora to adapt language, reasoning patterns, and domain knowledge.
- Retrieval-augmented generation (RAG): using an external, searchable knowledge base (statutes, guidance, firm precedents) to ground model outputs.
- Testing mixed configurations (e.g., fine-tuned base model + RAG).
- Evaluation criteria:
- Legal grounding (citation accuracy and relevance).
- Quality and coherence of legal justification.
- Ability to handle implicit context and document-specific facts.
- Practical utility for tax professionals (time saved, error rates, appropriateness of automation).
- Findings summary:
- Grounding via retrieval materially improves justifications and reduces hallucinations.
- Fine-tuning improves stylistic and domain-aligned output but does not on its own eliminate unsupported claims.
- Remaining failures often stem from missing structured background (client specifics, non-digitized documents, unstated assumptions).
Implications for AI Economics
- Productivity and task shifting:
- LLMs can raise productivity in tax advisory services by automating repetitive analyses and drafting, shifting human effort toward higher-complexity, judgement-heavy tasks.
- Demand may shift toward roles focused on oversight, client elicitation, and handling complex exceptions.
- Labor and wages:
- Potential reduction in time spent on routine tasks could compress billable hours for junior staff but increase value capture for senior experts who supervise AI outputs.
- Firms that invest in safe, well-integrated LLM systems may gain competitive advantage.
- Investment and complementary capital:
- Realizing benefits requires investment in structured data infrastructure (document ingestion, client profiles, knowledge bases) and retrieval pipelines to mitigate hallucinations.
- Ongoing costs include model maintenance, fine-tuning, and compliance/legal review.
- Risk, regulation, and liability:
- Hallucination and context-missing errors create legal and reputational risk; human-in-the-loop and audit trails are economically necessary safeguards.
- Regulators and professional bodies may need to clarify liability, disclosure, and quality standards for AI-assisted legal advice.
- Adoption constraints and market structure:
- Smaller firms may face barriers to adoption due to data and engineering costs; platform providers offering turnkey, legally grounded RAG stacks could consolidate market share.
- Standardization of legal knowledge representations (structured statutes, precedents) would lower costs and speed safe deployment.
- Research and policy priorities:
- Focus on methods that integrate structured client data and provenance-aware retrieval.
- Economic assessments should quantify trade-offs between automation gains and increased compliance/monitoring costs.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Properly configured LLMs using fine-tuning and/or retrieval-augmented generation can support tax professionals on Austrian and EU VAT matters by automating routine analyses, producing initial case assessments, and generating legally grounded justifications. Organizational Efficiency | positive | Usefulness and quality of LLM-generated VAT analyses and legal justifications |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Retrieval-based grounding materially improves the quality of legal justifications and reduces hallucinations, but does not eliminate unsupported assertions. Decision Quality | positive | Legal justification quality and frequency of unsupported assertions |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Fine-tuning improves stylistic and domain-aligned output but does not by itself eliminate unsupported claims. Output Quality | mixed | Domain alignment, output style, and unsupported-claim rate |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Current LLM prototypes are not reliable enough for fully automated, high-stakes VAT or legal decisions. Ai Safety And Ethics | negative | Reliability and suitability for unsupervised high-stakes legal decision-making |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLMs have persistent limitations in incorporating implicit client knowledge and interpreting case-specific documents that are not present in the retrieval corpus. Decision Quality | negative | Ability to use implicit context and case-specific documentation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| LLMs can raise productivity in tax advisory services by automating repetitive analyses and drafting, while shifting human effort toward complex, judgment-heavy tasks. Organizational Efficiency | positive | Productivity and allocation of human effort in tax advisory work |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| The use of LLMs may reduce time spent on routine tasks, potentially compressing billable hours for junior staff while increasing the value captured by senior experts who supervise AI outputs. Wages | mixed | Routine-task time, billable hours, and value capture across staff seniority levels |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Realizing productivity benefits from LLMs in tax advisory requires investment in structured data infrastructure, document ingestion, client profiles, knowledge bases, and retrieval pipelines. Organizational Efficiency | positive | Infrastructure requirements for effective and safe LLM deployment |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Hallucination and context-missing errors create legal and reputational risks, making human oversight and audit trails necessary safeguards. Ai Safety And Ethics | negative | Legal, reputational, and governance risk from AI-assisted tax advice |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Smaller tax-consulting firms may face greater barriers to adopting LLM systems because of data and engineering costs, while turnkey legally grounded RAG platforms could consolidate market share. Market Structure | negative | Adoption barriers and potential market-share concentration |
Reading fidelity
high
Study strength
speculative
|
not reported
|