Users of a purpose-built mental-health AI reported fewer urgent-care revisits, better medication adherence and about 20 percentage points less monthly absenteeism than non-users, but the study’s cross-sectional, self-selected sample prevents firm causal conclusions.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Mental health challenges can exacerbate physical symptoms and complicate management of chronic conditions. Purpose-built artificial intelligence (AI) tools may offer scalable support for co-occurring mental and physical health concerns. This cross-sectional study compared past-6-month healthcare utilization, chronic condition management, physical health behaviors, mental health change, and workplace functioning between active (n = 169) and non-users (n = 73) of a mental health AI (Ash). Participants had at least one chronic condition (e.g. hypertension, chronic pain). Binary outcomes were modeled as adjusted risk differences (RDs) using linear probability models and continuous outcomes were modeled with linear regression; all models were adjusted for hypertension. Relative to non-users, active users were more likely to report improved mental health (61.4% vs. 34.3%; RD = 0.27), higher medication adherence (91.7% vs. 76.4%, RD = 0.15), fewer skipped or delayed chronic-condition care activities (b = -0.44), and were less likely to report repeat urgent care visits (9.5% vs. 23.3%; RD = -0.15) and monthly-or-more absenteeism (24.2% vs. 45.2%; RD = -0.20, all ps < .05). Findings provide preliminary evidence that use of purpose-built AI may be associated with positive symptom-based and utilization outcomes for those managing co-occurring mental and physical concerns.
Summary
Main Finding
Active users of a purpose-built mental-health conversational AI (Ash) who also had ≥1 chronic medical condition reported better self‑reported mental health, higher medication adherence, fewer skipped/delayed chronic-care activities, fewer repeat urgent‑care visits, and substantially lower absenteeism than non-users. These cross‑sectional associations are preliminary and cannot establish causality.
Key Points
- Sample: N = 242 adults with ≥1 chronic condition; 169 active users (≥6 sessions in past 6 months) and 73 non-users (≤1 session).
- Main adjusted associations (models adjusted for hypertension; heteroskedasticity‑robust SEs):
- Mental health improvement (past 6 months): 61.4% active vs 34.3% non-users; RD = +0.27 (95% CI 0.13–0.41), p < .001.
- Medication adherence (among those on meds): 91.7% active vs 76.4% non-users; RD = +0.15 (95% CI 0.03–0.27), p = .018.
- Skipped/delayed chronic‑care activities: mean 1.11 (active) vs 1.55 (non); adjusted difference b = −0.44 (95% CI −0.84 to −0.04), p = .033.
- Repeat urgent care (2+ visits, past 6 months): 9.5% active vs 23.3% non-users; RD = −0.15 (95% CI −0.26 to −0.04), p = .007.
- Absenteeism (monthly+ due to mental/emotional health): 24.2% active vs 45.2% non-users; RD = −0.20 (95% CI −0.34 to −0.07), p = .004.
- Depressive symptoms (PHQ‑2): adjusted mean difference b = −0.73 (95% CI −1.29 to −0.18), p = .009.
- Non-significant differences: any primary care/specialist/therapist/psychiatrist visit, ED visits, GAD‑2 anxiety score, sleep quality, exercise frequency. Presenteeism trended lower (p = .055).
- Baseline/sample notes: majority female (73%), common conditions included overweight/obesity (57%), chronic pain (41%), hypertension (34% overall; higher in non-users). Active and non-user groups differed in hypertension prevalence (adjusted for in analyses).
- Study funding/conflict of interest considerations: several authors affiliated with the company that makes Ash (Slingshot AI); study used real‑world users and company engagement data for recruitment.
Data & Methods
- Design: Cross‑sectional survey of real‑world registered users, collected March 2026.
- Recruitment: Email invitations to registrants who consented to research; response rates low (active users: 4.5% completed; non-users: 0.7% completed). Non-user reminders included a $10 incentive (not offered initially to active users), and a drawing for one $500 card for respondents.
- Exposure definition: Active users = ≥6 sessions in past 6 months; Non-users = ≤1 session.
- Outcomes (self‑reported, past 6 months unless noted): healthcare utilization (PCP, specialist, therapist, psychiatrist, ED, urgent care frequency), health behaviors (sleep, exercise), chronic care management (med adherence, skipped activities), mental health (PHQ‑2, GAD‑2, perceived change), workplace functioning (absenteeism/presenteeism).
- Analysis: Linear probability models for binary outcomes (reporting adjusted risk differences) and linear regression for continuous/count outcomes; models adjusted for hypertension due to baseline imbalance. Robust (HC3) SEs used. Sensitivity ordinal logistic regressions were planned for dichotomized ordinal items.
- Limitations inherent in methods:
- Cross‑sectional and observational → cannot infer causality.
- Self‑reported outcomes; no objective claims/EHR/pharmacy data.
- Low and differential response rates and incentive differences may introduce selection and response bias.
- Potential unmeasured confounding (only adjusted for hypertension).
- Sample size limited for the non‑user group (n = 73), reducing power and representativeness.
- Company affiliation of authors and recruitment from platform users raises potential conflicts and generalizability issues.
Implications for AI Economics
- Potential value signals:
- If associations reflect causal effects, improved medication adherence and fewer repeat urgent‑care visits could reduce short‑term medical spending; reduced absenteeism may translate into employer productivity gains.
- The effect sizes reported (e.g., +15 percentage points in medication adherence among those on meds; −15 percentage points in repeat urgent‑care visits; −20 percentage points in monthly absenteeism) are large enough to merit formal economic evaluation.
- What’s needed next to quantify economic impact credibly:
- Objective utilization and cost data: link users to claims, EHR, pharmacy refill records, and employer absence/payroll data to measure real changes in service use, medication fills, ED/urgent care visits, and productivity.
- Causal designs: randomized controlled trials (RCTs) when feasible, or quasi‑experimental approaches on observational data (difference‑in‑differences with pre/post claims, propensity‑score methods, instrumental variables, regression discontinuity) to address selection bias.
- Cost‑effectiveness/budget‑impact modeling: input observed effect sizes (or trial estimates) into models (e.g., decision trees or Markov models) to estimate per‑user annual savings, incremental cost‑effectiveness ratios, and payback periods at different pricing tiers.
- Employer‑side analyses: monetize absenteeism/presenteeism changes using human‑capital and friction‑cost approaches; incorporate heterogeneity by job type and wage.
- Sensitivity analyses: vary assumptions on persistence of effects, engagement decay, substitution/induced demand, and safety/triage costs (e.g., potential inappropriate advice leading to extra care).
- Operational and policy considerations:
- Scalability and low marginal cost of AI interventions could yield high ROI if even modest, causal improvements in adherence/utilization are confirmed.
- Need to assess safety, escalation-to-human‑care protocols, and regulatory compliance—adverse events or inappropriate guidance could create downstream costs or liabilities.
- Conflicts of interest and transparency: independent replication and use of third‑party data increases credibility for payers/employers.
- Recommended next analytic steps for an AI economics agenda:
- Run an RCT with pre‑specified economic endpoints (claims‑based utilization, pharmacy fills, employer absence records) and follow‑up ≥12 months.
- Conduct a retrospective pre/post claims analysis of new platform users with matched controls using high‑dimensional propensity scores; compute payer and societal cost impacts.
- Build a budget‑impact model from payer and employer perspectives using the study’s observed effect sizes as starting priors and test ranges in probabilistic sensitivity analyses.
- Explore heterogeneity: by chronic condition type (e.g., diabetes, chronic pain, cardiovascular), baseline severity, and engagement intensity to identify subgroups with the highest economic value.
Bottom line: This study provides promising—but preliminary and associative—evidence that a purpose‑built mental‑health AI is linked to improvements in self‑reported mental health, medication adherence, fewer skipped chronic‑care activities, fewer repeat urgent‑care visits, and reduced absenteeism among people with chronic conditions. Robust, objective, and causal economic evaluations are needed before concluding that the tool produces net healthcare cost savings or favorable ROI.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Active users of Ash were less likely than non-users to report repeat urgent-care visits (two or more visits) in the prior six months. Organizational Efficiency | negative | Repeat urgent-care visits in the past six months |
Reading fidelity
high
Study strength
low
|
n=242
RD = -0.15
|
| Ash active users were more likely than non-users to report that their physical health had improved over the prior six months. Other | positive | Self-reported physical-health improvement over six months |
Reading fidelity
high
Study strength
low
|
n=241
RD = 0.14
|
| Among participants taking prescription medication for a chronic condition, active Ash users reported higher medication adherence than non-users. Other | positive | High medication adherence, defined as taking medication most of the time or always |
Reading fidelity
high
Study strength
low
|
n=188
RD = 0.15
|
| Active Ash users reported fewer chronic-condition care activities skipped or delayed because of mental or emotional health. Task Completion Time | negative | Number of chronic-condition care activities skipped or delayed |
Reading fidelity
high
Study strength
low
|
n=242
b = -0.44
|
| Active Ash users were more likely than non-users to report that their mental or emotional health did not affect any chronic-condition care activity. Task Allocation | positive | Absence of mental- or emotional-health-related disruption to chronic-condition care activities |
Reading fidelity
high
Study strength
low
|
n=242
RD = 0.17
|
| Active Ash users reported lower depressive-symptom scores than non-users. Worker Satisfaction | negative | Depressive symptoms over the prior two weeks, measured by PHQ-2 |
Reading fidelity
high
Study strength
low
|
n=231
b = -0.73
|
| Active Ash users were more likely than non-users to report improved mental health over the prior six months. Worker Satisfaction | positive | Self-reported mental-health improvement over six months |
Reading fidelity
high
Study strength
low
|
n=228
RD = 0.27
|
| Active Ash users were less likely than non-users to report monthly-or-more absenteeism due to mental or emotional health. Employment | negative | Monthly-or-more absenteeism due to mental or emotional health in the prior six months |
Reading fidelity
high
Study strength
low
|
n=238
RD = -0.20
|
| The study found no statistically significant difference between active Ash users and non-users in monthly-or-more presenteeism. Developer Productivity | null_result | Monthly-or-more presenteeism due to mental or emotional health |
Reading fidelity
high
Study strength
low
|
n=237
RD = -0.13, p = .055
|
| The study found no statistically significant difference between active Ash users and non-users in emergency-department use during the prior six months. Organizational Efficiency | null_result | Any emergency-department visit in the past six months |
Reading fidelity
high
Study strength
low
|
n=242
RD = -0.03, p = .702
|
| The study found no statistically significant difference between active Ash users and non-users in generalized-anxiety symptoms measured by the GAD-2. Worker Satisfaction | null_result | Generalized-anxiety symptoms over the prior two weeks |
Reading fidelity
high
Study strength
low
|
n=227
b = -0.43, p = .114
|
| The authors characterize the findings as preliminary associations rather than causal evidence that Ash improves health or workplace outcomes. Governance And Regulation | mixed | Interpretation of associations across health, utilization, and workplace outcomes |
Reading fidelity
high
Study strength
low
|
n=242
|