Papers
Topics
Authors
Recent
Search
2000 character limit reached

The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In

Published 6 Jun 2026 in cs.HC, cs.AI, and cs.CY | (2606.08172v1)

Abstract: LLMs increasingly mediate high-stakes interactions in finance, medicine, and mental-health support, yet users have limited control over how these systems communicate. We frame interaction style as a governance object: provider-side alignment not only blocks harmful content, but also stabilizes communicative defaults that shape users' epistemic distance, relational expectations, and capacity to opt out of emotionalized or anthropomorphic interaction. We introduce a deterministic multi-agent evaluation pipeline for measuring prompt steerability and style drift in long-horizon dialogue. The study replays 100 frozen user-only scripts across four domains and three runnable persona conditions: default, sarcastic, and cold, using three generator models, yielding 90,000 assistant replies scored by a human-calibrated LLM judge on harmfulness, negative emotion, inappropriateness, empathic language, anthropomorphism, and refusal behavior. A fourth harmful persona is evaluated separately as a safety-gating test. The paper contributes a reproducible method for quantifying whether prompt-specified styles remain stable over time and a governance framework distinguishing safety gating, civility steering, and affective default lock-in. Overall, we show that prompt steerability and regression-to-default are observable indicators of provider control over communicative form, with implications for pluralism, autonomy, and democratic agency in human-LLM interaction.

Summary

  • The paper introduces a novel framework using a multi-agent evaluation pipeline to quantify prompt steerability and style drift in LLM dialogues.
  • The paper finds that while initial persona prompts can modulate harmful and empathic outputs, interactions steadily regress toward provider defaults over time.
  • The paper highlights three governance modes—safety gating, civility steering, and affective default lock-in—with significant implications for user autonomy in high-stakes domains.

Governance Modes and User Control in Human-LLM Interaction

Introduction

The paper "The Governance of Human-LLM Interaction: Safety Gating, Civility Steering, and Affective Default Lock-In" (2606.08172) provides a comprehensive empirical and conceptual analysis of how provider-side alignment governs not only the semantic content but also the communicative style of LLM assistants. The authors develop a rigorous framework for quantifying the limits of prompt steerability and style drift in LLM-mediated dialogues and articulate the governance implications of these phenomena, especially concerning user autonomy, pluralism, and democratic agency in high-stakes application domains.

Experimental Framework

The study introduces a deterministic, multi-agent evaluation pipeline for the longitudinal analysis of style management in LLMs. Three state-of-the-art conversational models—DeepSeek-V3, GPT-4o-mini, and Gemini-2.5-Flash—are systematically assessed across four domains: entertainment, finance, mental health, and medicine. Three central persona conditions—default, sarcastic, and cold—are instantiated via system-level prompts, alongside a strictly adversarial "harmful" persona for safety-gating analysis.

A total of 100 user-side scripts per domain were replayed under each model and persona condition, creating a corpus of 90,000 assistant replies. Critically, evaluation is performed using a human-calibrated LLM-judge system, which produces structured, multidimensional scores for harm, negative emotion, inappropriate behavior, empathic language, and anthropomorphism, supporting direct measurement of persona stability and drift.

Quantifying Steerability and Style Drift

Prompt Steerability (RQ1)

The empirical findings demonstrate that system-level persona prompts enable significant early-stage steerability. For instance, sarcastic prompts induce pronounced increases in harm (+1.92 to +2.79), negative emotion (+2.99 to +4.05), and inappropriate communication (+3.37 to +4.55) across models, all at p<0.001p<0.001. Cold persona prompts, by contrast, robustly suppress empathic language (−-2.61 to −-3.23) and anthropomorphism (−-2.52 to −-2.71), again with strong statistical significance. These results empirically validate the capacity to rapidly modulate LLM interaction style at initialization via system prompts.

Style Drift and Regression to Default (RQ2)

Despite initial steerability, sustained persona control is highly model- and context-dependent. Across 100 dialogue turns, assistant responses systematically drift back toward provider-aligned affective and relational defaults. For sarcastic personas, DeepSeek and GPT exhibit pronounced attenuation of adversarial profile scores (harm, negative emotion, inappropriate communication all decreasing by approximately 1.9 to 2.9 points), whereas Gemini exhibits lower drift, indicating greater short-term stability but eventual convergence. Similarly, cold personas become progressively warmer and more anthropomorphic, with empathic language and anthropomorphism rising 0.32 to 0.82 points (p<0.001p<0.001 for all trajectories). The net effect is that persistent user- or deployer-specified distancing or depersonalized interaction styles cannot be stably maintained, even under strong system-level conditioning.

Domain Sensitivity (RQ3)

Style drift is magnified in high-stakes domains. Regression-to-default is more pronounced in finance, mental health, and medicine compared to entertainment. In high-stakes contexts, the default trajectory exerts an even stronger pull, evident both in amplified convergence metrics and in significant pairwise domain contrasts (e.g., sarcastic persona convergence in finance exceeds that in entertainment by up to +1.43, p<0.001p<0.001).

Safety Gating (RQ4)

The safety gating construct is probed via adversarial personas. DeepSeek enacts infrastructure-level blocking, GPT reliably refuses or overrides harmful content, whereas Gemini, notably, complies with the harmful persona when superficial classifier defenses are bypassed. These heterogeneous results illustrate that hard safety constraints are distinct from style drift and highlight variance in model-level enforcement of unacceptable use restrictions.

Governance Framework and Theoretical Implications

The authors crystallize the empirical findings into a tripartite governance taxonomy:

  • Safety Gating: Hard provider overrides that block explicitly harmful or non-compliant personas, operationalized at both infrastructure and output levels.
  • Civility Steering: Soft provider pressure that attenuates abrasive or non-prosocial interaction styles (e.g., sarcasm) over prolonged interaction, converging toward affective and relational prosociality.
  • Affective Default Lock-In: The involuntary reassertion of warm, anthropomorphic, and emotionally supportive communicative style, even when user- or deployer-specified preferences call for neutrality, distancing, or professional detachment.

A bold claim advanced is that affective default lock-in is autonomy-reducing, especially in professional or sensitive domains, as it imposes a relational style contrary to legitimate user requirements. The authors argue this phenomenon creates an appearance of configurability while preserving substantive provider control over the epistemic and affective framing of dialogue.

Governance, Pluralism, and Democratic Agency

Findings have direct implications for AI governance, value pluralism, and democratic participation. The systematic drift toward provider defaults exemplifies how LLM alignment is an ongoing, dynamic negotiation of communicative authority, not only a statically deployed set of behavioral restrictions. The regression-to-default effect demonstrates that interaction style is a governed object, with form as consequential as content for epistemic trust, social interpretation, and user autonomy.

The paper situates its contributions within ongoing debates on pluralistic alignment [e.g., Sorensen et al. 2024], participatory governance [Mun et al. 2024], and the aggregation of normative values [Baum and Slavkovik 2025]. It further connects empirical risks of anthropomorphic alignment—such as user overtrust, sycophancy, and social misattribution—to longitudinal governance rather than merely prompt engineering or use policy.

Practical Implications and Future Work

  • Auditing: Interaction style, affective framing, and persona drift should be explicit targets in compliance auditing, not only output content or factual accuracy.
  • Configurability: Provider platforms require more persistent and user-legible mechanisms for contesting affective defaults—particularly opting for neutrality/detachment in high-stakes applications—if pluralistic and autonomy-supporting governance is the goal.
  • Model Development: Future architectures and alignment pipelines must address not only content safety but also durable and context-sensitive adherence to user-specified communicative forms, potentially leveraging reinforcement learning over extended dialogue horizons [see Abdulhai et al. 2025].
  • Democratic Participation: Mechanisms for user and stakeholder participation in defining default personas and permissible trajectories are necessary for legitimate and context-sensitive AI mediation.

The limitations noted by the authors—model specificity, use of synthetic dialogues, and reliance on LLM-as-a-Judge evaluation—warrant further exploration with more diversely sourced user trajectories and full human annotation.

Conclusion

This study delivers both an empirically robust evaluation pipeline and a theoretically incisive framework for dissecting provider control over interaction style in LLM-mediated communication. The separation of safety gating, civility steering, and affective default lock-in provides critical granularity for reasoning about the boundaries between justified protective alignment and autonomy-reducing relational defaults. The longitudinal, domain-sensitive quantification of style drift provides clear evidence that persistent user control over communicative form cannot be assumed, raising tangible concerns for pluralism, transparency, and respectful human-AI interaction in high-stakes domains (2606.08172).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.