The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Users of a purpose-built mental-health AI reported fewer urgent-care revisits, better medication adherence and about 20 percentage points less monthly absenteeism than non-users, but the study’s cross-sectional, self-selected sample prevents firm causal conclusions.

Healthcare Utilization, Chronic Condition Management, and Workplace Functioning Among Users of a Purpose-Built Mental Health AI (Ash): Cross-Sectional Study
Kristen M. Van Swearingen, Thomas D. Hull, Jeffrey Swigert, Caitlin A. Stamatis · September 08, 2026
arxiv correlational low evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Kristen M. Van Swearingen unresolved corpus identity
  2. Thomas D. Hull unresolved corpus identity
  3. Jeffrey Swigert unresolved corpus identity
  4. Caitlin A. Stamatis unresolved corpus identity
Compared with non-users, active users of a purpose-built mental health AI reported better self-rated mental and physical-health changes, higher medication adherence, fewer repeat urgent-care visits, and substantially lower absenteeism, but the cross-sectional, self-selected design precludes causal claims.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Mental health challenges can exacerbate physical symptoms and complicate management of chronic conditions. Purpose-built artificial intelligence (AI) tools may offer scalable support for co-occurring mental and physical health concerns. This cross-sectional study compared past-6-month healthcare utilization, chronic condition management, physical health behaviors, mental health change, and workplace functioning between active (n = 169) and non-users (n = 73) of a mental health AI (Ash). Participants had at least one chronic condition (e.g. hypertension, chronic pain). Binary outcomes were modeled as adjusted risk differences (RDs) using linear probability models and continuous outcomes were modeled with linear regression; all models were adjusted for hypertension. Relative to non-users, active users were more likely to report improved mental health (61.4% vs. 34.3%; RD = 0.27), higher medication adherence (91.7% vs. 76.4%, RD = 0.15), fewer skipped or delayed chronic-condition care activities (b = -0.44), and were less likely to report repeat urgent care visits (9.5% vs. 23.3%; RD = -0.15) and monthly-or-more absenteeism (24.2% vs. 45.2%; RD = -0.20, all ps < .05). Findings provide preliminary evidence that use of purpose-built AI may be associated with positive symptom-based and utilization outcomes for those managing co-occurring mental and physical concerns.

Summary

Main Finding

Active users of a purpose-built mental-health conversational AI (Ash) who also had ≥1 chronic medical condition reported better self‑reported mental health, higher medication adherence, fewer skipped/delayed chronic-care activities, fewer repeat urgent‑care visits, and substantially lower absenteeism than non-users. These cross‑sectional associations are preliminary and cannot establish causality.

Key Points

  • Sample: N = 242 adults with ≥1 chronic condition; 169 active users (≥6 sessions in past 6 months) and 73 non-users (≤1 session).
  • Main adjusted associations (models adjusted for hypertension; heteroskedasticity‑robust SEs):
    • Mental health improvement (past 6 months): 61.4% active vs 34.3% non-users; RD = +0.27 (95% CI 0.13–0.41), p < .001.
    • Medication adherence (among those on meds): 91.7% active vs 76.4% non-users; RD = +0.15 (95% CI 0.03–0.27), p = .018.
    • Skipped/delayed chronic‑care activities: mean 1.11 (active) vs 1.55 (non); adjusted difference b = −0.44 (95% CI −0.84 to −0.04), p = .033.
    • Repeat urgent care (2+ visits, past 6 months): 9.5% active vs 23.3% non-users; RD = −0.15 (95% CI −0.26 to −0.04), p = .007.
    • Absenteeism (monthly+ due to mental/emotional health): 24.2% active vs 45.2% non-users; RD = −0.20 (95% CI −0.34 to −0.07), p = .004.
    • Depressive symptoms (PHQ‑2): adjusted mean difference b = −0.73 (95% CI −1.29 to −0.18), p = .009.
  • Non-significant differences: any primary care/specialist/therapist/psychiatrist visit, ED visits, GAD‑2 anxiety score, sleep quality, exercise frequency. Presenteeism trended lower (p = .055).
  • Baseline/sample notes: majority female (73%), common conditions included overweight/obesity (57%), chronic pain (41%), hypertension (34% overall; higher in non-users). Active and non-user groups differed in hypertension prevalence (adjusted for in analyses).
  • Study funding/conflict of interest considerations: several authors affiliated with the company that makes Ash (Slingshot AI); study used real‑world users and company engagement data for recruitment.

Data & Methods

  • Design: Cross‑sectional survey of real‑world registered users, collected March 2026.
  • Recruitment: Email invitations to registrants who consented to research; response rates low (active users: 4.5% completed; non-users: 0.7% completed). Non-user reminders included a $10 incentive (not offered initially to active users), and a drawing for one $500 card for respondents.
  • Exposure definition: Active users = ≥6 sessions in past 6 months; Non-users = ≤1 session.
  • Outcomes (self‑reported, past 6 months unless noted): healthcare utilization (PCP, specialist, therapist, psychiatrist, ED, urgent care frequency), health behaviors (sleep, exercise), chronic care management (med adherence, skipped activities), mental health (PHQ‑2, GAD‑2, perceived change), workplace functioning (absenteeism/presenteeism).
  • Analysis: Linear probability models for binary outcomes (reporting adjusted risk differences) and linear regression for continuous/count outcomes; models adjusted for hypertension due to baseline imbalance. Robust (HC3) SEs used. Sensitivity ordinal logistic regressions were planned for dichotomized ordinal items.
  • Limitations inherent in methods:
    • Cross‑sectional and observational → cannot infer causality.
    • Self‑reported outcomes; no objective claims/EHR/pharmacy data.
    • Low and differential response rates and incentive differences may introduce selection and response bias.
    • Potential unmeasured confounding (only adjusted for hypertension).
    • Sample size limited for the non‑user group (n = 73), reducing power and representativeness.
    • Company affiliation of authors and recruitment from platform users raises potential conflicts and generalizability issues.

Implications for AI Economics

  • Potential value signals:
    • If associations reflect causal effects, improved medication adherence and fewer repeat urgent‑care visits could reduce short‑term medical spending; reduced absenteeism may translate into employer productivity gains.
    • The effect sizes reported (e.g., +15 percentage points in medication adherence among those on meds; −15 percentage points in repeat urgent‑care visits; −20 percentage points in monthly absenteeism) are large enough to merit formal economic evaluation.
  • What’s needed next to quantify economic impact credibly:
    • Objective utilization and cost data: link users to claims, EHR, pharmacy refill records, and employer absence/payroll data to measure real changes in service use, medication fills, ED/urgent care visits, and productivity.
    • Causal designs: randomized controlled trials (RCTs) when feasible, or quasi‑experimental approaches on observational data (difference‑in‑differences with pre/post claims, propensity‑score methods, instrumental variables, regression discontinuity) to address selection bias.
    • Cost‑effectiveness/budget‑impact modeling: input observed effect sizes (or trial estimates) into models (e.g., decision trees or Markov models) to estimate per‑user annual savings, incremental cost‑effectiveness ratios, and payback periods at different pricing tiers.
    • Employer‑side analyses: monetize absenteeism/presenteeism changes using human‑capital and friction‑cost approaches; incorporate heterogeneity by job type and wage.
    • Sensitivity analyses: vary assumptions on persistence of effects, engagement decay, substitution/induced demand, and safety/triage costs (e.g., potential inappropriate advice leading to extra care).
  • Operational and policy considerations:
    • Scalability and low marginal cost of AI interventions could yield high ROI if even modest, causal improvements in adherence/utilization are confirmed.
    • Need to assess safety, escalation-to-human‑care protocols, and regulatory compliance—adverse events or inappropriate guidance could create downstream costs or liabilities.
    • Conflicts of interest and transparency: independent replication and use of third‑party data increases credibility for payers/employers.
  • Recommended next analytic steps for an AI economics agenda:
    • Run an RCT with pre‑specified economic endpoints (claims‑based utilization, pharmacy fills, employer absence records) and follow‑up ≥12 months.
    • Conduct a retrospective pre/post claims analysis of new platform users with matched controls using high‑dimensional propensity scores; compute payer and societal cost impacts.
    • Build a budget‑impact model from payer and employer perspectives using the study’s observed effect sizes as starting priors and test ranges in probabilistic sensitivity analyses.
    • Explore heterogeneity: by chronic condition type (e.g., diabetes, chronic pain, cardiovascular), baseline severity, and engagement intensity to identify subgroups with the highest economic value.

Bottom line: This study provides promising—but preliminary and associative—evidence that a purpose‑built mental‑health AI is linked to improvements in self‑reported mental health, medication adherence, fewer skipped chronic‑care activities, fewer repeat urgent‑care visits, and reduced absenteeism among people with chronic conditions. Robust, objective, and causal economic evaluations are needed before concluding that the tool produces net healthcare cost savings or favorable ROI.

Assessment

Paper Typecorrelational Evidence Strengthlow — The study is cross-sectional with self-selected participants, low and differential response rates, limited covariate adjustment (only hypertension), and reliance on self-reported outcomes, so associations may reflect confounding, selection, or reporting bias rather than causal effects. Methods Rigormedium — Analytic methods are appropriate for descriptive comparative work (robust SEs, sensitivity ordinal models), but the design lacks pre-registration, randomization, longitudinal follow-up, or rich control for confounders; recruitment/incentive differences and low response rates further weaken internal validity. SampleAnalytic sample of 242 adults (169 active users, 73 non-users) who registered for the Ash mental health chatbot and reported ≥1 diagnosed chronic condition; active users had ≥6 sessions in prior 6 months, non-users had ≤1 session; recruited via email invitations (5,879 active users invited, 15,270 non-users invited) with low completion rates (267/5,879 and 104/15,270 completed overall) and final analytic sample restricted to those with chronic conditions; data collected March 2026; measures are self-reported (healthcare utilization, PHQ-2, GAD-2, medication adherence, workplace functioning); non-user reminders included a $10 gift incentive and all respondents entered a $500 gift card drawing. Themesproductivity human_ai_collab IdentificationCross-sectional observational comparison between active users (≥6 sessions in past 6 months) and non-users (≤1 session) of a purpose-built mental health chatbot; models adjust only for hypertension and use linear probability and linear regression with HC3 robust SEs; sensitivity checks with ordinal logistic regression; no randomization, longitudinal design, or comprehensive confounder adjustment to support causal inference. GeneralizabilitySelf-selected sample of registrants of a single mental health chatbot (Ash) — not representative of general population or all digital mental health users, Low and differential survey response rates (much lower among non-users) introduce nonresponse bias, Predominantly female sample and high share with Medicaid/Medicare/ACA insurance limits representativeness, Results rely on self-reported outcomes (recall and reporting biases), Findings reflect one purpose-built product and specific engagement thresholds; may not generalize to other AI tools or usage patterns, Cross-sectional snapshot; cannot generalize to longer-term outcomes or causal effects

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Active users of Ash were less likely than non-users to report repeat urgent-care visits (two or more visits) in the prior six months. Organizational Efficiency negative Repeat urgent-care visits in the past six months
Reading fidelity high
Study strength low
n=242
RD = -0.15
0.15
Ash active users were more likely than non-users to report that their physical health had improved over the prior six months. Other positive Self-reported physical-health improvement over six months
Reading fidelity high
Study strength low
n=241
RD = 0.14
0.15
Among participants taking prescription medication for a chronic condition, active Ash users reported higher medication adherence than non-users. Other positive High medication adherence, defined as taking medication most of the time or always
Reading fidelity high
Study strength low
n=188
RD = 0.15
0.15
Active Ash users reported fewer chronic-condition care activities skipped or delayed because of mental or emotional health. Task Completion Time negative Number of chronic-condition care activities skipped or delayed
Reading fidelity high
Study strength low
n=242
b = -0.44
0.15
Active Ash users were more likely than non-users to report that their mental or emotional health did not affect any chronic-condition care activity. Task Allocation positive Absence of mental- or emotional-health-related disruption to chronic-condition care activities
Reading fidelity high
Study strength low
n=242
RD = 0.17
0.15
Active Ash users reported lower depressive-symptom scores than non-users. Worker Satisfaction negative Depressive symptoms over the prior two weeks, measured by PHQ-2
Reading fidelity high
Study strength low
n=231
b = -0.73
0.15
Active Ash users were more likely than non-users to report improved mental health over the prior six months. Worker Satisfaction positive Self-reported mental-health improvement over six months
Reading fidelity high
Study strength low
n=228
RD = 0.27
0.15
Active Ash users were less likely than non-users to report monthly-or-more absenteeism due to mental or emotional health. Employment negative Monthly-or-more absenteeism due to mental or emotional health in the prior six months
Reading fidelity high
Study strength low
n=238
RD = -0.20
0.15
The study found no statistically significant difference between active Ash users and non-users in monthly-or-more presenteeism. Developer Productivity null_result Monthly-or-more presenteeism due to mental or emotional health
Reading fidelity high
Study strength low
n=237
RD = -0.13, p = .055
0.15
The study found no statistically significant difference between active Ash users and non-users in emergency-department use during the prior six months. Organizational Efficiency null_result Any emergency-department visit in the past six months
Reading fidelity high
Study strength low
n=242
RD = -0.03, p = .702
0.15
The study found no statistically significant difference between active Ash users and non-users in generalized-anxiety symptoms measured by the GAD-2. Worker Satisfaction null_result Generalized-anxiety symptoms over the prior two weeks
Reading fidelity high
Study strength low
n=227
b = -0.43, p = .114
0.15
The authors characterize the findings as preliminary associations rather than causal evidence that Ash improves health or workplace outcomes. Governance And Regulation mixed Interpretation of associations across health, utilization, and workplace outcomes
Reading fidelity high
Study strength low
n=242
0.15

Notes