- The paper demonstrates that naturalness (human-likeness) and appropriateness are distinct metrics, as high naturalness does not guarantee domain suitability.
- It employs a controlled perceptual analysis with 150 native speakers and a Latin Square design to evaluate varied TTS systems across multiple use-cases.
- The findings advocate for domain-specific, context-aware evaluation protocols to better align TTS performance with listener expectations.
Evaluating TTS: Disentangling Naturalness and Appropriateness Across Domains
Introduction and Motivation
This paper interrogates the long-standing reliance on "naturalness"—often equated with human-likeness—as the primary criterion for evaluating Text-to-Speech (TTS) systems. Despite significant advances in fidelity and human-like synthesis, the field lacks robust methods for assessing domain suitability, more properly framed as "appropriateness." The authors empirically examine how perceptions of naturalness and appropriateness diverge across domains, specifically focusing on AI assistant, reader, actor, animated character, and spontaneous speaker use-cases. The analysis highlights the inadequacy of one-size-fits-all metrics, especially for expressive TTS, and emphasizes the necessity for context-aware evaluation protocols.
Experimental Methodology
The study employs a controlled perceptual analysis leveraging a bespoke Gradio interface, with 150 native English speakers rating TTS and ground truth (GT) samples generated across five domains. Utilizing carefully selected SOTA TTS systems covering a range of timbral, prosodic, and expressive capabilities (e.g., Kokoro/StyleTTS2, Gemini TTS/Flash 2.5, Kyutai-TTS, GPT-4o-mini-tts, ElevenLabs multilingual_v2), the study systematically measures perceived appropriateness and human-likeness on a sentence-by-sentence basis. Sentences span narration (LibriQuote), spontaneous dialogue (MSP-Podcast), affective conversation (MELD, AnimeVox), and informational statements. The evaluation design uses a Latin Square experimental setup to counterbalance anchoring and minimize voice preference bias.
Results: Appropriateness is Domain-Dependent
The analysis reveals that perceived appropriateness is highly domain-sensitive and often decoupled from traditional naturalness metrics. For example, Kokoro excels in reading and assistant roles but underperforms in conversational contexts, while Kyutai-TTS is rated highly for spontaneous speech yet fares poorly as an assistant or character. Commercial systems like Eleven Labs and Gemini perform strongly in acting domains but not in spontaneous conversation. Notably, the inter-rater agreement on appropriateness judgments is modest, affirming the subjective and context-dependent nature of these ratings.
Figure 1: Appropriateness scores across five personas and all speech tasks, showcasing variability by system and context.
Figure 2: Appropriateness and human-likeness scores for each TTS and persona, including statistical significance of observed differences.
Decoupling Human-Likeness and Appropriateness
The data explicitly demonstrates that human-likeness (as a proxy for naturalness) and appropriateness are not universally correlated. For instance, while positive correlations exist in reading, acting, and spontaneous speech domains, the correlation vanishes or even reverses for AI assistant or animated character roles. Notably, human-likeness scores penalize stylized or unconventional prosody, rewarding spontaneous or neutral renditions. In several domains, high naturalness models are deemed inappropriate, indicating a need for task-specific evaluation.
Impact of Task and Acoustic Features
Human-likeness ratings for ground truth samples also vary by domain, with conversational and acted dialogue achieving higher mean scores than more stylized or regionally accented narration datasets. TTS systems show similar task-related fluctuations, implying that system performance reflects both model capability and congruence with domain expectations.
Acoustic-prosodic analyses uncover domain-specific patterns: animated character synthesis correlates strongly with pacing variability and nPVI, suggesting a premium on rhythmic dynamism, whereas narration rewards steady delivery. Assistant roles disfavor wide pitch or high valence, preferring neutral timbre. For spontaneous domains, "human" imperfections such as jitter and creak are positively weighted by listeners, which traditional MOS metrics would penalize.
Figure 3: Radar profiles highlight domain-specific appropriateness patterns by TTS system.
Figure 4: Human-likeness scores for ground-truth datasets, evidencing context sensitivity in perception.
Figure 5: Task-dependent variation in human-likeness scores across TTS systems.
Figure 6: Heatmap of correlations between acoustic features and appropriateness, illustrating domain-dependent signal attributes.
Limitations of Automatic Metrics and Implications
When correlating appropriateness with established automatic metrics (e.g., UTMOS, DNSMOS, embedding-based similarity), strong negative correlations emerge for more expressive domains, underscoring that such metrics penalize features that are positively correlated with human perception in those settings. Conversely, metrics that perform well in reading or assistant tasks do not generalize, reinforcing the inadequacy of global-quality predictors for expressive TTS.
Figure 7: Correlations of appropriateness and automatic metrics, revealing metric-task misalignment.
Practical and Theoretical Implications
This work underscores the necessity for domain-profiling both TTS models and evaluation metrics. Current SOTA systems remain specialized, excelling only in certain expressive or neutral roles. The authors highlight a systematic perceptual acceptance bias towards less expressive, more "AI-like" delivery for assistant scenarios. Importantly, the results caution against relying on conventional MOS or objective metrics for expressive tasks as they fail to reliably distinguish appropriateness.
Theoretically, the findings advocate for a multidimensional, context-sensitive approach to TTS assessment and development. Future TTS models should expose domain-adaptive profiles, with evaluation protocols that explicitly incorporate listener expectations, acoustic-prosodic variability, and situational appropriateness rather than collapsing system quality into a single axis.
Conclusion
The separation of "naturalness" and "appropriateness" as distinct evaluative axes for TTS systems is empirically validated. Domain context determines appropriateness, often in ways that are orthogonal to, or even in tension with, human-likeness as measured via standard MOS or automatic metrics. Progress in TTS synthesis should thus be measured by contextually nuanced, multi-metric protocols. Future research should address multi-turn, socially grounded scenarios and further elucidate the role of speaker identity and listener bias in perceived appropriateness, guiding the development of more generalizable, controllable, and contextually adaptive TTS systems.