Papers
Topics
Authors
Recent
Search
2000 character limit reached

Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a Large Language Model

Published 13 Apr 2026 in cs.AI and q-bio.OT | (2604.11287v1)

Abstract: Background: LLMs have been explored as tools for generating personalized exercise prescriptions, yet the consistency of outputs under identical conditions remains insufficiently examined. Objective: This study evaluated the intra-model consistency of LLM-generated exercise prescriptions using a repeated generation design. Methods: Six clinical scenarios were used to generate exercise prescriptions using Gemini 2.5 Flash (20 outputs per scenario; total n = 120). Consistency was assessed across three dimensions: (1) semantic consistency using SBERT-based cosine similarity, (2) structural consistency based on the FITT principle using an AI-as-a-judge approach, and (3) safety expression consistency, including inclusion rates and sentence-level quantification. Results: Semantic similarity was high across scenarios (mean cosine similarity: 0.879-0.939), with greater consistency in clinically constrained cases. Frequency showed consistent patterns, whereas variability was observed in quantitative components, particularly exercise intensity. Unclassifiable intensity expressions were observed in 10-25% of resistance training outputs. Safety-related expressions were included in 100% of outputs; however, safety sentence counts varied significantly across scenarios (H=86.18, p less than 0.001), with clinical cases generating more safety expressions than healthy adult cases. Conclusions: LLM-generated exercise prescriptions demonstrated high semantic consistency but showed variability in key quantitative components. Reliability depends substantially on prompt structure, and additional structural constraints and expert validation are needed before clinical deployment.

Authors (1)

Summary

  • The paper quantifies intra-model consistency in AI-generated exercise prescriptions across clinically constrained and healthy scenarios.
  • The study employs repeated prompt testing with SBERT cosine similarity and FITT classification to highlight semantic reliability and numerical instability, particularly in exercise intensity.
  • The findings imply that while AI models reliably include safety guidelines, they require structured prompts and expert review for accurate quantitative outputs in clinical settings.

Consistency in AI-Generated Exercise Prescriptions: Analysis and Implications

Study Objective and Methodological Framework

The study "Consistency of AI-Generated Exercise Prescriptions: A Repeated Generation Study Using a LLM" (2604.11287) conducts a rigorous examination of intra-model consistency in AI-generated exercise prescriptions. Utilizing Gemini 2.5 Flash via Vertex AI API, the research addresses a critical gap: quantifying how stable and reproducible LLM-generated prescriptions are under repeated, identical prompt conditions. Prescriptions were generated for six clinically defined and healthy scenarios, each repeated 20 times, and analyzed for consistency in semantic structure, FITT-component prescriptions, and inclusion of safety content.

Notably, the evaluation eschews mere surface-level assessment, instead decomposing consistency into three orthogonal dimensions:

  • Semantic Consistency: Measured by SBERT-based cosine similarity across paired outputs.
  • Structural Consistency: Assessed via AI-as-a-judge FITT (Frequency, Intensity, Time, Type) classification, with explicit mapping of quantitative variables.
  • Safety Expression Consistency: Both binary inclusion and sentence-level quantification across key safety domains.

Quantitative Findings on Consistency

Semantic Consistency

Semantic analysis revealed high intra-scenario cosine similarity (means ranging 0.879–0.939), confirming that the LLM outputs are textually stable. However, this stability is modulated by the clinical specificity of the scenario; cases with more tightly constrained clinical guidelines (e.g., post-colon cancer surgery, knee OA with fall risk) displayed the least variability and highest semantic coherence. Scenarios with broader latitude (e.g., muscle hypertrophy in healthy adults) demonstrated lower mean similarity and greater variance, indicating increased model stochasticity when fewer clinical constraints are present.

Structural Consistency and FITT Classification

Despite stable semantic output, key FITT prescription components exhibited notable heterogeneity:

  • Frequency showed the most structural alignment with guidelines, across cases.
  • Intensity was the most unstable aspect. For resistance training in clinical scenarios, 10–25% of outputs could not be classified by usual criteria (e.g., explicit %1RM omitted or ambiguous). This reflects inherent limitations in LLMs' capability for reliable, fine-grained quantitative recommendations even with explicit prompt engineering.
  • Duration and Exercise Type were generally scenario-appropriate but occasionally estimated or omitted precise values, further highlighting numerical instability.
  • Structural consistency thus decays whenever prompts permit, or require, more nuanced reasoning, or the guidelines offer latitude.

Safety Inclusion Consistency

Prompts explicitly required safety components. Accordingly, every output (100%) included all target safety classes—contraindications, precautions, symptom monitoring, risk warnings. However, sentence-level quantification revealed substantial scenario-dependent variability in the magnitude of safety content. Complex clinical scenarios yielded more extensive safety guidance (e.g., multimorbidity case: 61.4 sentences/output), while healthy cases included minimal safety commentary. This underscores that while explicit prompt engineering ensures inclusion, it does not standardize depth, which is regulated by scenario complexity and possibly LLM interpretation of clinical risk.

Implications for Digital Health and Clinical Deployment

Reliability, Prompt Engineering, and the Limits of LLMs

These results critically inform both practical deployment and theoretical understanding:

  • Prompt Structure as Primary Modulator: The LLM’s reliability in exercise prescription is less a function of the foundational LM architecture and more contingent on the structure and specificity of the input prompt. Explicit constraints and detailed instructions reduce intra-model variability; unconstrained or open-ended prompts lead to non-equivalent outputs, even for identical scenarios.
  • Quantitative Output Instability: The inability of LLMs to reliably generate and maintain stable numerical values (e.g., %1RM, duration, intensity) is a significant limitation, especially for domains where dosage precision is clinically non-negotiable.
  • Safety Content Is Achievable with Structured Prompts: Binary inclusion of safety elements appears enforceable with prompt constraints, but controlling expression depth remains challenging. This finding corroborates the value of integrating domain-specific prompt templates for safety-critical AI outputs, but also signals the need for post-generation expert review to ensure clinical adequacy.

Theoretical Significance and Directions for Future Research

This study operationalizes reproducibility as a distinct performance axis in LLM-generated clinical/health content, complementing metrics such as accuracy and safety. The introduction of hybrid evaluation pipelines—semantic similarity metrics alongside domain-specific structural and safety checks—establishes a framework for multi-dimensional assessment of generative models in medical application.

Future research should systematically compare multiple LLMs under identical prompt and scenario structures to delineate model-dependent versus prompt-dependent effects. Additionally, formal expert adjudication of content validity is necessary to benchmark AI-as-a-judge pipelines. Expansion beyond six scenarios will be needed to ensure robustness across broader clinical heterogeneity.

Conclusion

The stability and reliability of LLM-generated exercise prescriptions are substantially mediated by prompt structure and scenario constraints. While high semantic and categorical safety consistency are feasible, the persistence of quantitative instability in key FITT parameters—particularly exercise intensity—signals that LLMs alone are insufficient for autonomous clinical prescription. Structured prompt engineering, expert validation, and possibly hybrid symbolic-numeric constraint systems are prerequisites for responsible AI deployment in exercise and rehabilitation medicine.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.