- The paper demonstrates that LLMs experience marked drops in empathetic accuracy and cultural grounding when shifting from English to Ukrainian.
- It employs a 500-prompt benchmark across diverse crisis scenarios to compare models such as DeepSeek-V3, Gemini-2.5-Flash, and LLaMA-3.3-70B-Instruct.
- Results reveal that specialized MoE models sustain better cross-lingual performance, while automated evaluations often overlook nuanced cultural misalignments.
SPLIT: Evaluating Cross-Lingual Empathy and Cultural Grounding in LLMs
Introduction
The paper "SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses" (2607.02049) presents a rigorous evaluation of LLMs in emotionally intensive, crisis-related contexts across English and Ukrainian. The study distinguishes itself by shifting the focus away from mere multilingual fluency and instead evaluates nuanced empathetic and culturally grounded response generation. Leveraging a purpose-built 500-prompt SPLIT benchmark, the work highlights significant cross-lingual degradations in empathy and grounding, particularly for mid- and low-resource languages. Human and automated evaluators yield divergent judgments, emphasizing the limitations of relying solely on LLM-as-a-judge evaluation in such settings.
Benchmark and Methodology
The SPLIT benchmark comprises 500 emotion-centric prompts, equally balanced across five psychosocial crisis categories: Stress, Panic, Loneliness, Internal Displacement, and Tension. This prompt set is constructed to specifically probe LLMs' ability to deliver support within the cultural and psychological realities of crisis-affected Ukrainians.
Three technically diverse state-of-the-art models were evaluated:
Responses were assessed on three axes:
- Empathetic Accuracy: Appropriateness and depth of emotional support.
- Linguistic Naturalness: Authenticity, fluency, idiomatic usage.
- Contextual and Cultural Grounding: Localization, reference to norms, and adaptation to the Ukrainian context.
Both human evaluation and a multi-agent LLM jury (GPT-4o, Mistral Large, Claude 4.5 Sonnet) were used for assessment, and the agreement between them was statistically analyzed.
Cross-Lingual Trajectories and Major Findings
LLMs exhibit pronounced cross-lingual performance degradation for Ukrainian in empathy-centric and culturally specific response tasks. Relative trajectories for macro-average human evaluation show clear model-specific divergence:
Figure 1: Cross-lingual performance trajectories showing macro-average human evaluation scores from EN to UA.
Notably, Gemini-2.5-Flash and LLaMA-3.3-70B-Instruct both undergo substantial performance drops when generating Ukrainian responses. In contrast, DeepSeek-V3 maintains stable performance, and even shows marginal gains for certain categories.
Manual Human Evaluation
Detailed manual assessments affirm that:
Automated LLM Jury Evaluations vs. Human Judgments
Algorithmic evaluation by LLM judges exhibits limited alignment with human assessments, particularly for cultural grounding:
Figure 3: Automated Baseline scores across the three evaluated dimensions in English (EN) and Ukrainian (UA).
Key observations:
Quantitatively, the systemically lenient MAE and ME values further substantiate AI jury’s overscoring tendency and superficiality regarding authentic cultural alignment.
Implications: Multilingualism vs. Multiculturalism
The data reinforce a critical distinction: “producing Ukrainian text is not equivalent to producing Ukrainian emotional support.” High lexical fluency does not guarantee culturally attuned empathetic responses. The SPLIT findings provide compelling evidence that:
- LLMs trained predominantly on English or large-scale web corpora lack sufficient exposure to culturally embedded emotional expression in Ukrainian.
- MoE architectures with specialized expertise routing (DeepSeek-V3) offer advantages for sustaining cross-lingual empathetic continuity.
- Even with extensive pretraining on 200+ languages, closed commercial models like Gemini may sacrifice local nuance for generalized politeness and formulaic content, failing to achieve deep cultural grounding.
Limitations
Notably, the study relies on a single highly proficient human annotator, potentially introducing subjectivity, and focuses only on Ukrainian as a mid-resource target. Additionally, the SPLIT benchmark is thematically narrow and not yet validated across domains such as healthcare or education.
Future Research Directions
- Expanding the SPLIT paradigm to additional low- and mid-resource languages and new sociocultural domains.
- Developing more sophisticated, culturally tailored alignment and RLHF pipelines.
- Engineering evaluation frameworks that incorporate a wider range of human raters and inter-annotator agreement measures.
- Enriching model pretraining with high-quality, culturally annotated datasets targeting emotional support expressions in underrepresented languages.
Conclusion
The SPLIT benchmark (2607.02049) demonstrates that current LLMs exhibit stark discrepancies between surface-level multilingual capabilities and the ability to generate grounded empathetic support in low-/mid-resource languages. Whereas dense-transformer and web-scale MoE models experience severe drops in emotional alignment for Ukrainian, specialized MoE architectures sustain cross-lingual efficacy. Importantly, the LLM-as-a-judge paradigm fails to reliably proxy human judgments in the domain of cultural grounding, underscoring the necessity for direct human-in-the-loop evaluation.
Theoretical and practical implications highlight that multilingualism does not entail multicultural competence. As LLMs see increasing deployment in culturally sensitive applications, systematic integration of cultural embedding, idiomatic adaptation, and multi-annotator human evaluation is imperative. The SPLIT findings mandate a paradigm shift toward multicultural benchmarks and training regimes as prerequisites for responsible and effective emotional-support AI in global contexts.