The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Paid AI models outperform free counterparts at producing functional assistive-device software: Gemini Pro implemented 14 of 16 requested features with just 1.25 prompts on average, while free models produced at most four functions, suggesting subscription models materially improve nontechnical professionals' ability to generate usable code.

Using Natural Language Prompts With AI Models for Low-Cost Assistive Software Design: Exploratory Comparative Evaluation
Francesc Antoni Bañuls-Lapuerta, Vicent Marti-Miralles, Rómulo Jacobo Gónzalez-García, Gabriel Martínez-Rico · February 19, 2026 · JMIR Rehabilitation and Assistive Technologies
openalex descriptive medium evidence 7/10 relevance Summary only summary available; pdf_status=paywall DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Francesc Antoni Bañuls-Lapuerta provider ID
  2. Vicent Marti-Miralles provider ID
  3. Rómulo Jacobo Gónzalez-García provider ID
  4. Gabriel Martínez-Rico provider ID

Semantic Scholar

Latest observation:

  1. F. Bañuls-Lapuerta provider ID
  2. Vicent Marti-Miralles provider ID
  3. R. J. Gónzalez-García provider ID
  4. Gabriel Mártínez-Rico provider ID
Paid large-language models (Gemini Pro and GPT-5/ChatGPT Plus) produced substantially more functional assistive-device code with fewer prompts than free models, indicating stronger support for nontechnical professionals in creating low-cost personalized assistive tools.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Background: This study investigates the capacity of 7 artificial intelligence (AI) models, 5 free and 2 paid, to generate functional software for designing low-cost, personalized assistive products. Objective: The objective was to determine which models are most effective, accessible, and consistent in supporting nontechnical professionals in developing inclusive digital solutions and to assess the capabilities of commercially available and easy-to-access AI models to generate code from natural language interactions in the shape of a nontechnical assistive technology design process. Methods: Each AI model was prompted using natural language, without any technical input, to create a Python program that converts an arcade gamepad into an adapted mouse-like controller. Sixteen progressively complex functions were requested through standardized prompts, delivered without additional feedback or correction. Model performance was evaluated based on the number of successfully implemented functions and the average number of prompts required. Results: Paid models demonstrated markedly superior performance. Gemini Pro (Google) successfully implemented 14 of 16 requested functions with an average of 1.25 (SD 0.45) prompts, while ChatGPT Plus (GPT-5) achieved 11 functions with an average of 1.31 (SD 0.48) prompts. In contrast, free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking the highest within their category. The enhanced outcomes of paid models were linked to improved contextual understanding, greater tolerance for natural language, and reduced conversational drift. Conclusions: Paid AI models, particularly Gemini Pro and ChatGPT Plus, exhibit strong potential as tools for bridging the gap between health or education professionals and software development. They enable the creation of affordable, user-centered assistive technology without requiring advanced programming skills. Nevertheless, human oversight and foundational literacy in prompt design remain crucial to guarantee functionality, reliability, and ethical use.

Summary

Main Finding

Paid large AI models substantially outperform free models at converting natural-language, nontechnical prompts into working assistive-device software. In this study Gemini Pro (Google) and ChatGPT Plus (GPT-5) produced far more functional features with fewer prompts than free alternatives, suggesting paid models are more effective, robust, and usable by nontechnical professionals designing low-cost personalized assistive products.

Key Points

  • Study task: ask AI models (no technical input) to generate Python software that turns an arcade gamepad into an adapted, mouse-like controller through 16 progressively complex functions.
  • Population of models: 7 total — 2 paid models (Gemini Pro, ChatGPT Plus) and 5 free models (including Gemini Free, DeepSeek).
  • Performance summary:
    • Gemini Pro: 14 of 16 functions implemented; average 1.25 prompts per function (SD 0.45).
    • ChatGPT Plus (GPT-5): 11 of 16 functions implemented; average 1.31 prompts per function (SD 0.48).
    • Free models: between 0 and 4 functional outcomes; DeepSeek and Gemini Free were the best among free options.
  • Qualitative advantage of paid models: superior contextual understanding, greater tolerance for natural-language phrasing, and reduced conversational drift.
  • Human oversight and prompt-literacy remain necessary to ensure correctness, reliability, and ethical deployment.

Data & Methods

  • Experimental design:
    • 16 standardized, progressively complex natural-language prompts requesting specific functional capabilities.
    • No technical input from researchers; prompts framed as nontechnical design requests to simulate use by health/education professionals.
    • Models were prompted and allowed to produce code; researchers recorded the number of functions successfully implemented and the average number of prompts required (with standard deviations reported for top performers).
    • No technical corrections or coding-level instruction were provided by researchers (interactions aimed to remain within a natural-language design process).
  • Outcome measures:
    • Primary: count of implemented functions (out of 16).
    • Secondary: average number of prompts required to reach a working implementation per function.
  • Sample limitations:
    • Single application domain (assistive controller) and single implementation language (Python).
    • Seven models tested — small model pool and snapshot in time.
    • “Success” defined by implemented functionality; nuance of code quality, robustness, security, or maintainability not deeply quantified here.

Implications for AI Economics

  • Value of model quality and monetization:
    • Paid models show clear productivity returns for nontechnical users; this supports subscription/API pricing and justifies willingness to pay among professionals who gain time- and cost-savings.
    • Model quality translates into economic value beyond raw compute: better contextual understanding reduces iteration costs and human supervision.
  • Distributional effects and access inequality:
    • Superior paid models can widen gaps between well-resourced organizations (who can afford subscriptions) and resource-constrained practitioners, potentially worsening the digital divide in assistive-technology development.
    • Public-sector and low-income providers may need subsidized access or public alternatives to capture these productivity gains.
  • Labor and skill complementarities:
    • These models act as skill multipliers for nontechnical workers (health, education), reducing the need for specialist programmers for many custom assistive solutions — a form of task automation/augmentation that is skill-biased toward prompt-literacy rather than coding.
    • Investment in prompt-design training (human capital) is likely to have high returns for organizations adopting these tools.
  • Market structure and competition:
    • Freemium models that maintain only limited capability risk failing to deliver real utility for complex applied tasks; this may concentrate value in better-funded incumbents unless open or competitive alternatives improve.
    • Demand for reliable, domain-specific performance may drive firms toward premium tiers or specialized enterprise offerings.
  • Externalities and policy considerations:
    • Safety, reliability, and ethical deployment require oversight; poor outputs from free models could produce harms if used without checks.
    • Policymakers should consider subsidizing access for public-health/education providers, supporting open-model development, and funding evaluation frameworks for AI-generated assistive tech.
  • Research and economic evaluation needs:
    • Cost–benefit analyses comparing subscription costs to labor/time saved in real-world deployments.
    • Broader, longitudinal studies across domains and larger model samples to quantify general equilibrium effects (labor demand, market entry for small providers, welfare gains from cheaper assistive tech).

Limitations to keep in mind: narrow domain, limited model set, and focus on immediate functional outputs rather than long-term reliability, maintenance costs, or user testing. Further work should evaluate real-world deployments, user outcomes, and the tradeoffs between paid-access productivity and public-access equity.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study uses a clear, standardized benchmarking protocol (16 progressively complex functions, fixed natural-language prompts) and objective outcome metrics (functions implemented, prompts required), producing consistent differences between paid and free models; however, the evidence is limited to a single task domain, a small set of models and prompts, lacks statistical inference or replication, and is vulnerable to model-versioning and prompt-dependence. Methods Rigormedium — Strengths include standardized prompts, no additional feedback (reducing variability from human correction), and objective scoring of function success; weaknesses include a narrow task set (one assistive-device conversion), small model sample, likely single runs per request (limited replication), no user testing with target nontechnical professionals, potential nondeterminism in model outputs, and limited reporting of inter-rater coding or error analysis. SampleSeven commercial AI models (5 free-tier models and 2 paid-tier models) were prompted via natural language to produce Python code converting an arcade gamepad into an adapted mouse-like controller; 16 progressively complex functions were requested per model using standardized prompts, and outcomes recorded were number of functions successfully implemented and average number of prompts required (SD reported); paid models included Gemini Pro (Google) and ChatGPT Plus (GPT-5); free models included Gemini Free and DeepSeek among others. Themeshuman_ai_collab productivity adoption skills_training GeneralizabilitySingle task domain (one assistive-device conversion) limits applicability to other software projects or domains, Only Python code and one hardware interaction pattern assessed; results may not generalize to other languages, platforms, or hardware, Small set of models and time-bound versions — model performance can change with updates or different API settings, Nontechnical end-users were not directly studied, so claims about accessibility for nontechnical professionals are inferred rather than observed, Outcomes focus on functional implementation, not on usability, reliability, maintainability, or safety of generated code

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study tested 7 AI models (5 free and 2 paid) using natural-language prompts to create a Python program that converts an arcade gamepad into an adapted mouse-like controller. Other null_result none (methodological description of experimental setup)
Reading fidelity high
Study strength high
n=7
0.3
Sixteen progressively complex functions were requested from each model through standardized prompts, delivered without additional feedback or correction. Other null_result number and complexity of requested functions
Reading fidelity high
Study strength high
n=16
0.3
Paid models demonstrated markedly superior performance compared with free models. Developer Productivity positive number of successfully implemented functions
Reading fidelity high
Study strength medium
n=2
Gemini Pro 14/16; ChatGPT Plus 11/16; free models 0–4 functions
0.18
Gemini Pro (Google) successfully implemented 14 of 16 requested functions with an average of 1.25 (SD 0.45) prompts. Developer Productivity positive number of successfully implemented functions; average number of prompts required
Reading fidelity high
Study strength medium
n=16
14 of 16; average 1.25 (SD 0.45) prompts
0.18
ChatGPT Plus (GPT-5) achieved 11 of 16 requested functions with an average of 1.31 (SD 0.48) prompts. Developer Productivity positive number of successfully implemented functions; average number of prompts required
Reading fidelity high
Study strength medium
n=16
11 of 16; average 1.31 (SD 0.48) prompts
0.18
Free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking highest among free models. Developer Productivity negative number of successfully implemented functions
Reading fidelity high
Study strength medium
n=5
0 to 4 functions implemented (range across free models)
0.18
The superior outcomes of paid models were linked to improved contextual understanding, greater tolerance for natural language, and reduced conversational drift. Developer Productivity positive model contextual understanding / tolerance for natural language / conversational drift (qualitative)
Reading fidelity medium
Study strength speculative
not reported
0.02
Paid AI models, particularly Gemini Pro and ChatGPT Plus, exhibit strong potential as tools for bridging the gap between health or education professionals and software development, enabling creation of affordable, user-centered assistive technology without requiring advanced programming skills. Developer Productivity positive ability of nontechnical professionals to develop assistive technology (as proxied by model-generated code functionality and prompt efficiency)
Reading fidelity medium
Study strength low
not reported
0.05
Human oversight and foundational literacy in prompt design remain crucial to guarantee functionality, reliability, and ethical use of AI-generated assistive technologies. Governance And Regulation mixed requirement for oversight/prompt literacy to ensure safe/reliable outputs
Reading fidelity high
Study strength speculative
not reported
0.03
Model performance was evaluated based on the number of successfully implemented functions and the average number of prompts required. Developer Productivity null_result number of successfully implemented functions; average number of prompts required
Reading fidelity high
Study strength high
not reported
0.3

Notes