0 cumulative citations
View corpus contextPaid AI models outperform free counterparts at producing functional assistive-device software: Gemini Pro implemented 14 of 16 requested features with just 1.25 prompts on average, while free models produced at most four functions, suggesting subscription models materially improve nontechnical professionals' ability to generate usable code.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
1 cumulative citations
View corpus contextBackground: This study investigates the capacity of 7 artificial intelligence (AI) models, 5 free and 2 paid, to generate functional software for designing low-cost, personalized assistive products. Objective: The objective was to determine which models are most effective, accessible, and consistent in supporting nontechnical professionals in developing inclusive digital solutions and to assess the capabilities of commercially available and easy-to-access AI models to generate code from natural language interactions in the shape of a nontechnical assistive technology design process. Methods: Each AI model was prompted using natural language, without any technical input, to create a Python program that converts an arcade gamepad into an adapted mouse-like controller. Sixteen progressively complex functions were requested through standardized prompts, delivered without additional feedback or correction. Model performance was evaluated based on the number of successfully implemented functions and the average number of prompts required. Results: Paid models demonstrated markedly superior performance. Gemini Pro (Google) successfully implemented 14 of 16 requested functions with an average of 1.25 (SD 0.45) prompts, while ChatGPT Plus (GPT-5) achieved 11 functions with an average of 1.31 (SD 0.48) prompts. In contrast, free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking the highest within their category. The enhanced outcomes of paid models were linked to improved contextual understanding, greater tolerance for natural language, and reduced conversational drift. Conclusions: Paid AI models, particularly Gemini Pro and ChatGPT Plus, exhibit strong potential as tools for bridging the gap between health or education professionals and software development. They enable the creation of affordable, user-centered assistive technology without requiring advanced programming skills. Nevertheless, human oversight and foundational literacy in prompt design remain crucial to guarantee functionality, reliability, and ethical use.
Summary
Main Finding
Paid large AI models substantially outperform free models at converting natural-language, nontechnical prompts into working assistive-device software. In this study Gemini Pro (Google) and ChatGPT Plus (GPT-5) produced far more functional features with fewer prompts than free alternatives, suggesting paid models are more effective, robust, and usable by nontechnical professionals designing low-cost personalized assistive products.
Key Points
- Study task: ask AI models (no technical input) to generate Python software that turns an arcade gamepad into an adapted, mouse-like controller through 16 progressively complex functions.
- Population of models: 7 total — 2 paid models (Gemini Pro, ChatGPT Plus) and 5 free models (including Gemini Free, DeepSeek).
- Performance summary:
- Gemini Pro: 14 of 16 functions implemented; average 1.25 prompts per function (SD 0.45).
- ChatGPT Plus (GPT-5): 11 of 16 functions implemented; average 1.31 prompts per function (SD 0.48).
- Free models: between 0 and 4 functional outcomes; DeepSeek and Gemini Free were the best among free options.
- Qualitative advantage of paid models: superior contextual understanding, greater tolerance for natural-language phrasing, and reduced conversational drift.
- Human oversight and prompt-literacy remain necessary to ensure correctness, reliability, and ethical deployment.
Data & Methods
- Experimental design:
- 16 standardized, progressively complex natural-language prompts requesting specific functional capabilities.
- No technical input from researchers; prompts framed as nontechnical design requests to simulate use by health/education professionals.
- Models were prompted and allowed to produce code; researchers recorded the number of functions successfully implemented and the average number of prompts required (with standard deviations reported for top performers).
- No technical corrections or coding-level instruction were provided by researchers (interactions aimed to remain within a natural-language design process).
- Outcome measures:
- Primary: count of implemented functions (out of 16).
- Secondary: average number of prompts required to reach a working implementation per function.
- Sample limitations:
- Single application domain (assistive controller) and single implementation language (Python).
- Seven models tested — small model pool and snapshot in time.
- “Success” defined by implemented functionality; nuance of code quality, robustness, security, or maintainability not deeply quantified here.
Implications for AI Economics
- Value of model quality and monetization:
- Paid models show clear productivity returns for nontechnical users; this supports subscription/API pricing and justifies willingness to pay among professionals who gain time- and cost-savings.
- Model quality translates into economic value beyond raw compute: better contextual understanding reduces iteration costs and human supervision.
- Distributional effects and access inequality:
- Superior paid models can widen gaps between well-resourced organizations (who can afford subscriptions) and resource-constrained practitioners, potentially worsening the digital divide in assistive-technology development.
- Public-sector and low-income providers may need subsidized access or public alternatives to capture these productivity gains.
- Labor and skill complementarities:
- These models act as skill multipliers for nontechnical workers (health, education), reducing the need for specialist programmers for many custom assistive solutions — a form of task automation/augmentation that is skill-biased toward prompt-literacy rather than coding.
- Investment in prompt-design training (human capital) is likely to have high returns for organizations adopting these tools.
- Market structure and competition:
- Freemium models that maintain only limited capability risk failing to deliver real utility for complex applied tasks; this may concentrate value in better-funded incumbents unless open or competitive alternatives improve.
- Demand for reliable, domain-specific performance may drive firms toward premium tiers or specialized enterprise offerings.
- Externalities and policy considerations:
- Safety, reliability, and ethical deployment require oversight; poor outputs from free models could produce harms if used without checks.
- Policymakers should consider subsidizing access for public-health/education providers, supporting open-model development, and funding evaluation frameworks for AI-generated assistive tech.
- Research and economic evaluation needs:
- Cost–benefit analyses comparing subscription costs to labor/time saved in real-world deployments.
- Broader, longitudinal studies across domains and larger model samples to quantify general equilibrium effects (labor demand, market entry for small providers, welfare gains from cheaper assistive tech).
Limitations to keep in mind: narrow domain, limited model set, and focus on immediate functional outputs rather than long-term reliability, maintenance costs, or user testing. Further work should evaluate real-world deployments, user outcomes, and the tradeoffs between paid-access productivity and public-access equity.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study tested 7 AI models (5 free and 2 paid) using natural-language prompts to create a Python program that converts an arcade gamepad into an adapted mouse-like controller. Other | null_result | none (methodological description of experimental setup) |
Reading fidelity
high
Study strength
high
|
n=7
|
| Sixteen progressively complex functions were requested from each model through standardized prompts, delivered without additional feedback or correction. Other | null_result | number and complexity of requested functions |
Reading fidelity
high
Study strength
high
|
n=16
|
| Paid models demonstrated markedly superior performance compared with free models. Developer Productivity | positive | number of successfully implemented functions |
Reading fidelity
high
Study strength
medium
|
n=2
Gemini Pro 14/16; ChatGPT Plus 11/16; free models 0–4 functions
|
| Gemini Pro (Google) successfully implemented 14 of 16 requested functions with an average of 1.25 (SD 0.45) prompts. Developer Productivity | positive | number of successfully implemented functions; average number of prompts required |
Reading fidelity
high
Study strength
medium
|
n=16
14 of 16; average 1.25 (SD 0.45) prompts
|
| ChatGPT Plus (GPT-5) achieved 11 of 16 requested functions with an average of 1.31 (SD 0.48) prompts. Developer Productivity | positive | number of successfully implemented functions; average number of prompts required |
Reading fidelity
high
Study strength
medium
|
n=16
11 of 16; average 1.31 (SD 0.48) prompts
|
| Free models produced between 0 and 4 functional outcomes, with DeepSeek and Gemini Free ranking highest among free models. Developer Productivity | negative | number of successfully implemented functions |
Reading fidelity
high
Study strength
medium
|
n=5
0 to 4 functions implemented (range across free models)
|
| The superior outcomes of paid models were linked to improved contextual understanding, greater tolerance for natural language, and reduced conversational drift. Developer Productivity | positive | model contextual understanding / tolerance for natural language / conversational drift (qualitative) |
Reading fidelity
medium
Study strength
speculative
|
not reported
|
| Paid AI models, particularly Gemini Pro and ChatGPT Plus, exhibit strong potential as tools for bridging the gap between health or education professionals and software development, enabling creation of affordable, user-centered assistive technology without requiring advanced programming skills. Developer Productivity | positive | ability of nontechnical professionals to develop assistive technology (as proxied by model-generated code functionality and prompt efficiency) |
Reading fidelity
medium
Study strength
low
|
not reported
|
| Human oversight and foundational literacy in prompt design remain crucial to guarantee functionality, reliability, and ethical use of AI-generated assistive technologies. Governance And Regulation | mixed | requirement for oversight/prompt literacy to ensure safe/reliable outputs |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Model performance was evaluated based on the number of successfully implemented functions and the average number of prompts required. Developer Productivity | null_result | number of successfully implemented functions; average number of prompts required |
Reading fidelity
high
Study strength
high
|
not reported
|