Papers
Topics
Authors
Recent
Search
2000 character limit reached

SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses

Published 2 Jul 2026 in cs.CL, cs.AI, and cs.CY | (2607.02049v1)

Abstract: LLMs are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-lingual abilities in these circumstances remain underexplored. Existing benchmarks emphasize multilingual performance but rarely examine crisis-related empathy and cultural grounding in low-to-mid-resource languages. We introduce SPLIT, a 500-prompt benchmark designed to evaluate LLM consistency in generating emotionally grounded responses across five categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. We evaluate three technically diverse LLMs across three dimensions: Empathetic Accuracy, Linguistic Naturalness, and Contextual & Cultural Grounding. The framework aims to assess and compare the quality of LLM responses in both English and Ukrainian languages, as well as to explore the reliability of the LLM-as-a-jury paradigm. Our findings reveal that Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct degrade when transitioning to Ukrainian, while DeepSeek-V3 remains comparatively stable within our benchmark. We additionally find that human and AI evaluators agree weakly on empathy and naturalness but diverge on cultural grounding. We further argue that producing Ukrainian text is not equivalent to producing Ukrainian emotional support. Our findings may assist in the future development of more culturally tailored benchmark designs, as well as encourage a stronger emphasis on human-centered evaluation.

Authors (1)

Summary

  • The paper demonstrates that LLMs experience marked drops in empathetic accuracy and cultural grounding when shifting from English to Ukrainian.
  • It employs a 500-prompt benchmark across diverse crisis scenarios to compare models such as DeepSeek-V3, Gemini-2.5-Flash, and LLaMA-3.3-70B-Instruct.
  • Results reveal that specialized MoE models sustain better cross-lingual performance, while automated evaluations often overlook nuanced cultural misalignments.

SPLIT: Evaluating Cross-Lingual Empathy and Cultural Grounding in LLMs

Introduction

The paper "SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses" (2607.02049) presents a rigorous evaluation of LLMs in emotionally intensive, crisis-related contexts across English and Ukrainian. The study distinguishes itself by shifting the focus away from mere multilingual fluency and instead evaluates nuanced empathetic and culturally grounded response generation. Leveraging a purpose-built 500-prompt SPLIT benchmark, the work highlights significant cross-lingual degradations in empathy and grounding, particularly for mid- and low-resource languages. Human and automated evaluators yield divergent judgments, emphasizing the limitations of relying solely on LLM-as-a-judge evaluation in such settings.

Benchmark and Methodology

The SPLIT benchmark comprises 500 emotion-centric prompts, equally balanced across five psychosocial crisis categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. This prompt set is constructed to specifically probe LLMs' ability to deliver support within the cultural and psychological realities of crisis-affected Ukrainians.

Three technically diverse state-of-the-art models were evaluated:

Responses were assessed on three axes:

  1. Empathetic Accuracy: Appropriateness and depth of emotional support.
  2. Linguistic Naturalness: Authenticity, fluency, idiomatic usage.
  3. Contextual and Cultural Grounding: Localization, reference to norms, and adaptation to the Ukrainian context.

Both human evaluation and a multi-agent LLM jury (GPT-4o, Mistral Large, Claude 4.5 Sonnet) were used for assessment, and the agreement between them was statistically analyzed.

Cross-Lingual Trajectories and Major Findings

LLMs exhibit pronounced cross-lingual performance degradation for Ukrainian in empathy-centric and culturally specific response tasks. Relative trajectories for macro-average human evaluation show clear model-specific divergence: Figure 1

Figure 1: Cross-lingual performance trajectories showing macro-average human evaluation scores from EN to UA.

Notably, Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct both undergo substantial performance drops when generating Ukrainian responses. In contrast, DeepSeek-V3 maintains stable performance, and even shows marginal gains for certain categories.

Manual Human Evaluation

Detailed manual assessments affirm that:

  • DeepSeek-V3 is robust across both languages, with Empathetic Accuracy and Contextual/Cultural Grounding slightly improving for Ukrainian.
  • Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct exhibit sharp drops in all dimensions when moving from English to Ukrainian, with LLaMA’s Linguistic Naturalness undergoing the most severe degradation.
  • Human annotators penalize overly formal, literal, or clichéd translations that fail to resonate contextually. Figure 2

    Figure 2: Human Evaluation Baseline scores across the three evaluated dimensions in English (EN) and Ukrainian (UA).

Automated LLM Jury Evaluations vs. Human Judgments

Algorithmic evaluation by LLM judges exhibits limited alignment with human assessments, particularly for cultural grounding: Figure 3

Figure 3: Automated Baseline scores across the three evaluated dimensions in English (EN) and Ukrainian (UA).

Key observations:

  • The LLM-as-a-judge paradigm assigns higher, more stable scores across both languages and most models, neglecting nuanced declines that are strongly detected by human raters.
  • Weak positive Pearson correlations are found only for Empathetic Accuracy and Linguistic Naturalness.
  • For Contextual and Cultural Grounding, the correlation is negative and statistically insignificant, indicating a clear limitations in recognizing culturally grounded responses. Figure 4

    Figure 4: Overall agreement between Automated and Human Evaluation Baselines

Quantitatively, the systemically lenient MAE and ME values further substantiate AI jury’s overscoring tendency and superficiality regarding authentic cultural alignment.

Implications: Multilingualism vs. Multiculturalism

The data reinforce a critical distinction: “producing Ukrainian text is not equivalent to producing Ukrainian emotional support.” High lexical fluency does not guarantee culturally attuned empathetic responses. The SPLIT findings provide compelling evidence that:

  • LLMs trained predominantly on English or large-scale web corpora lack sufficient exposure to culturally embedded emotional expression in Ukrainian.
  • MoE architectures with specialized expertise routing (DeepSeek-V3) offer advantages for sustaining cross-lingual empathetic continuity.
  • Even with extensive pretraining on 200+ languages, closed commercial models like Gemini may sacrifice local nuance for generalized politeness and formulaic content, failing to achieve deep cultural grounding.

Limitations

Notably, the study relies on a single highly proficient human annotator, potentially introducing subjectivity, and focuses only on Ukrainian as a mid-resource target. Additionally, the SPLIT benchmark is thematically narrow and not yet validated across domains such as healthcare or education.

Future Research Directions

  • Expanding the SPLIT paradigm to additional low- and mid-resource languages and new sociocultural domains.
  • Developing more sophisticated, culturally tailored alignment and RLHF pipelines.
  • Engineering evaluation frameworks that incorporate a wider range of human raters and inter-annotator agreement measures.
  • Enriching model pretraining with high-quality, culturally annotated datasets targeting emotional support expressions in underrepresented languages.

Conclusion

The SPLIT benchmark (2607.02049) demonstrates that current LLMs exhibit stark discrepancies between surface-level multilingual capabilities and the ability to generate grounded empathetic support in low-/mid-resource languages. Whereas dense-transformer and web-scale MoE models experience severe drops in emotional alignment for Ukrainian, specialized MoE architectures sustain cross-lingual efficacy. Importantly, the LLM-as-a-judge paradigm fails to reliably proxy human judgments in the domain of cultural grounding, underscoring the necessity for direct human-in-the-loop evaluation.

Theoretical and practical implications highlight that multilingualism does not entail multicultural competence. As LLMs see increasing deployment in culturally sensitive applications, systematic integration of cultural embedding, idiomatic adaptation, and multi-annotator human evaluation is imperative. The SPLIT findings mandate a paradigm shift toward multicultural benchmarks and training regimes as prerequisites for responsible and effective emotional-support AI in global contexts.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.