Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games
Abstract: People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of LLMs use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding LLMs to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Explain it Like I'm 14
What is this paper about?
This paper looks at how to make AI LLMs act more like real people when they’re trying to persuade or make decisions. Instead of just giving the AI a “persona” (like “pretend you’re a judge”), the authors use simple math rules from psychology and economics—called cognitive models—to tell the AI exactly how to update its beliefs when it sees new evidence. They test this in a legal setting using historical court cases and show that this “equation-to-behavior” approach makes AI simulations more realistic and more controllable.
What questions did the researchers ask?
The researchers focused on questions you can understand as “How do people change their minds—and can AI copy that?”
- Can AI follow different belief‑updating styles people use (like careful, biased, or overly confident) using clear, math-like instructions?
- Do bigger AI models follow these rules better than smaller ones?
- Can we train smaller models to follow these rules more reliably?
- If we train persuasive AIs against a variety of human decision styles, do they learn better strategies than if they only practice against “perfectly rational” opponents?
- Does this method match real human decisions (like judges’ verdicts) better than typical “pretend to be a judge” prompts?
How did they study it?
They built a realistic practice world using old court transcripts from London’s Old Bailey (1674–1913). Think of each trial as a “persuasion game”:
- The Sender (like a prosecutor) reveals pieces of evidence over several rounds.
- The Receiver (like a judge) updates their belief about “guilty or not” and then decides to convict or acquit.
To make this work at scale, they:
- Broke each trial into small pieces of evidence.
- Used an AI (checked by human reviewers) to label how strong each piece of evidence is and how pieces relate to each other.
- Changed only the order of the same evidence (e.g., strongest-first vs. weakest-first) to see how order affects decisions.
They then controlled the Receiver’s behavior using two methods:
- Equation-to-Behavior Prompting: They wrote prompts that literally tell the AI to follow specific belief-updating equations from cognitive science, such as:
- Bayesian updating (carefully adjusting beliefs with new evidence),
- Affine distortion (mixing your old opinion with new evidence, staying a bit “stuck” to your prior),
- Motivated updating (leaning toward the conclusion you prefer),
- Grether’s α–β model (turning up or down how much you trust prior beliefs or new evidence).
- Equation-to-Behavior RL (Reinforcement Learning): They “taught” smaller AI models by rewarding them when their beliefs matched what the equations predict and gently penalizing them for drifting too far from their original style. This helps the models learn the rules more consistently.
What did they find?
Here are the main findings:
- Larger models can follow the math rules from cognitive science just by prompting. Smaller models often can’t do this reliably without training.
- Even by default (with no special rules), many AIs act somewhat like Bayesians—adjusting beliefs sensibly—but they also show human‑like biases (for example, being overly swayed by early evidence, a “primacy” effect).
- Training smaller models with Equation-to-Behavior RL made them much better at following the rules, even in new situations they hadn’t seen before—reducing belief errors by about 26.5% on tricky test cases.
- When they trained persuasive AIs against a mix of different “types” of decision-makers (not just perfectly rational ones), the AIs improved their average impact by about 2.5%–12%, even when trying to persuade a strong model (like a “GPT-5-mini”).
- Using equation-based instructions matched real historical verdicts better than just prompting a “judge persona,” improving accuracy by about 4.7%–9.1% for stronger models.
Why is this important?
If we want AI to interact safely and effectively with people—whether in training lawyers, advising policymakers, or designing fair systems—we need it to understand the many ways humans actually update their beliefs, not just the “textbook rational” way. This paper shows a practical, testable way to control an AI’s decision style using simple math rules from human psychology and economics. That can:
- Make simulations more realistic for training and evaluation,
- Help identify and reduce risks from overly persuasive AI,
- Improve AI’s strategic thinking by preparing it for diverse human behaviors,
- Open the door to studying more complex human decision patterns in a controlled, measurable way.
In short, the paper provides a toolkit for turning equations about human thinking into reliable AI behavior—making AI better at understanding, anticipating, and interacting with real people.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
Below is a single, consolidated list of concrete gaps the paper leaves unresolved that future work could address:
- External validity: No direct human-subject evaluation shows that Equation-to-Behavior (E2B) Prompting or E2B-RL reproduces actual human belief updating or actions in persuasion games beyond small-scale annotation checks.
- Parameter calibration to humans: Cognitive model parameters (e.g., ) are not empirically estimated from human updating trajectories in the presented tasks; fits to historical judges (Appendix) are limited and do not validate trial-level belief dynamics.
- Identifiability: The paper does not address that distinct cognitive models can yield indistinguishable behavior in some persuasion problems, nor propose identification strategies or diagnostics to disambiguate models from observed actions/beliefs.
- Generalization across domains: All experiments are in historical legal transcripts (Old Bailey); there is no test of transfer to contemporary legal contexts or other strategic domains (e.g., medical, financial, political persuasion).
- Dataset representativeness: The Old Bailey corpus (1674–1913) differs culturally, legally, and linguistically from modern settings; the impact of these shifts on model behavior and conclusions is unexamined.
- Annotation reliability and bias: LLM-generated evidential strength and dependency labels are only lightly validated; the extent, direction, and impact of systematic annotation errors or biases on downstream evaluations are not quantified.
- Dependency modeling: Exact posterior evaluation is limited to “approximately independent” subsets; methods to compute or approximate posteriors when signals are dependent (using the provided dependency structure) are not developed or validated.
- Explanation-based updating formalization: The main text lacks a formal, executable definition of the “explanation-based” model (only a parameter is mentioned in training); it is unclear how explanatory coherence is computed or validated.
- Restricted state/action space: The framework assumes binary states (guilty/innocent) and binary actions (convict/acquit); extensions to multi-state/multi-action settings or continuous outcomes are not explored.
- Utility specification: Receiver/Sender utilities are simplified and symmetric; the effect of asymmetric error costs (e.g., false conviction vs. acquittal), legal standards like “beyond a reasonable doubt,” or risk aversion is not modeled or calibrated.
- Evidence sequencing design: The sender in core evaluations uses deterministic orderings (prosecution vs. defense), not fully strategic or interactive disclosure; the effect of adaptive sender strategies and receiver questioning/cross-examination remains open.
- Multi-round dynamics: Interactions are capped at three rounds with equal-sized evidence batches; stopping rules, early decision-making, and dynamic evidence selection by either agent are not studied.
- Mixed populations and heterogeneity: Although “mixed-receiver” training is used, the paper does not quantify how well models infer or adapt online to an unknown distribution over receiver types, nor how quickly they personalize to an individual.
- Robustness to prompting and inference settings: Sensitivity to instruction wording, temperature/decoding parameters, and prompt order is not systematically assessed; reproducibility across runs and providers is uncertain (some large-model results use subsets).
- Cross-model divergence: Large models implement qualitative features of the same rule but produce materially different quantitative trajectories; the causes (e.g., latent priors, implicit heuristics) are not analyzed.
- Compression/distillation: There is no method to distill high-fidelity equation-following behavior from large models into smaller ones without RL, nor a study of knowledge distillation from E2B-prompted large models.
- RL reward design: E2B-RL optimizes absolute belief error ; alternatives (e.g., KL divergence, Brier scores, action-level regret) and their impact on behavior fidelity and calibration are not compared.
- Task-general performance: The effect of E2B-RL on general capabilities (reasoning, safety benchmarks, hallucination) is unknown; potential degradation or reward hacking is not probed.
- Sample efficiency and scaling laws: The training uses ~1102 instances and 500 steps; there is no ablation on data size, rollout count, KL strength, or steps to understand sample efficiency and stability.
- Safety and ethics: While risks are noted, there is no design or evaluation of guardrails (e.g., persuasion constraints, intent detection, transparency mechanisms) for E2B-trained agents, especially when optimizing persuasive impact.
- Mechanistic interpretability: Adherence to equations is assessed behaviorally; no analysis probes whether intermediate representations or computations inside the LLM align with the specified cognitive models.
- Realistic sender–receiver equilibria: Joint training toward equilibrium (co-adaptation), non-stationarity, and convergence properties are not analyzed; how optimal sender strategies change under heterogeneous non-Bayesian receivers remains largely unexplored.
- Model misspecification: How E2B-RL behaves when the cognitive model is wrong or partially specified (e.g., misestimated priors, unmodeled dependencies) is not examined; robustness to model error is unknown.
- Persona + equation composition: Although noted as complementary, the paper does not systematically test how persona prompts interact with equation prompts, or how to resolve conflicts between persona traits and equation adherence.
- Group decision-making: Extensions to panels/juries (aggregation of heterogeneous updaters, social influence) and the impact on strategic disclosure are not addressed.
- Online parameter inference: Methods to infer a receiver’s cognitive parameters on-the-fly during interaction (and adapt sender strategy accordingly) are not developed or evaluated.
- Counterfactuals and auditability: The framework does not provide tools to generate counterfactual explanations (e.g., which evidence or parameter change flipped the decision) to audit simulated decision processes.
- Modern legal realism: Many real-world determinants of verdicts (procedural rulings, legal standards, judge/jury instructions, attorney performance) are not modeled, leaving a gap between simulated and actual courtroom dynamics.
Practical Applications
Below are actionable, real-world applications that build on the paper’s findings and methods. Each item links to sectors, suggests concrete tools/products/workflows, and notes assumptions or dependencies affecting feasibility.
Immediate Applications
These can be deployed now with current models and tooling (especially using Equation-to-Behavior Prompting with large LLMs and the released dataset/protocols).
- LegalTech — training and evaluation
- Use case: AI-assisted moot courts and trial preparation that expose attorneys to diverse, non-Bayesian “judges.”
- Tool/product: Receiver Simulator parameterized by Bayesian, affine distortion, motivated updating, and Grether α–β models; Evidence-ordering sandbox using Old Bailey–style evidence units.
- Dependencies/assumptions: Access to strong LLMs; domain shift from historical to modern cases; human-in-the-loop oversight; ethical constraints on strategic persuasion.
- AI safety and model governance
- Use case: Red-teaming and auditing of LLMs for covert or overly persuasive behavior; quantifying order effects, primacy biases, and over-inference.
- Tool/product: Persuasion Risk Sandbox with Equation-to-Behavior Prompting; “Persuasion Safety Scorecard” that reports invariance violations and belief error under multiple receiver models.
- Dependencies/assumptions: Agreement on metrics; access to standardized prompts and evaluation harness; organizational adoption in model release processes.
- Academia (cognitive science, economics, HCI)
- Use case: Rapid, controlled testing of cognitive models at scale; hypothesis screening before costly human studies.
- Tool/product: Executable Cognitive-Model Prompts; Old Bailey–based benchmark and evaluation scripts; estimation routines fitting α–β parameters to human or archival decisions.
- Dependencies/assumptions: LLMs as imperfect but useful proxies; need to triangulate with human data; careful interpretation to avoid overgeneralization.
- Public policy and regulation
- Use case: Regulatory sandbox to test the efficacy of disclosures, warnings, and transparency cues in conversational systems.
- Tool/product: Compliance pre-check harness that measures whether receivers (Bayesian vs biased) still get misled under different labeling and ordering regimes.
- Dependencies/assumptions: Regulator buy-in; alignment with evolving AI transparency rules; careful scope to avoid facilitating manipulative deployment.
- Healthcare communications
- Use case: Prototype message sequencing for patient adherence or informed consent that is robust to conservative or motivated updating.
- Tool/product: Message Ordering Advisor that previews performance across receiver profiles (e.g., conservative Bayesian vs over-inference).
- Dependencies/assumptions: Domain adaptation beyond legal text; IRB/ethics review; not for direct clinical deployment without trials; guardrails against undue influence.
- Education (statistics, critical thinking, law)
- Use case: Teaching modules demonstrating Bayesian vs non-Bayesian updating and order effects in narratives.
- Tool/product: Interactive classroom labs using Equation-to-Behavior prompts and historical cases; instructor dashboards visualizing belief trajectories.
- Dependencies/assumptions: Curricular alignment; appropriate scaffolding for students; content accessibility.
- Software/product UX
- Use case: Information architecture tests to detect harmful order effects in onboarding, disclosures, and decision aids.
- Tool/product: Order-Effect A/B Emulator that stress-tests content sequences across receiver models.
- Dependencies/assumptions: Simulated receivers approximate real users; tethered to user research and A/B testing.
- Finance (advice, disclosures, compliance)
- Use case: Pre-deployment checks of risk-disclosure language to ensure fairness and reduce undue persuasion.
- Tool/product: Compliance Pre-Check for Persuasive Risk that evaluates how different receivers interpret disclosures and likelihoods.
- Dependencies/assumptions: Legal sign-off; alignment with investor-protection rules; not a substitute for legal compliance review.
- Customer support and CRM training
- Use case: Training curricula for agents to handle diverse decision-making styles (e.g., order-sensitive, motivated updating) ethically.
- Tool/product: Mixed-Receiver Practice Scenarios with feedback on clarity, sequencing, and non-coercive framing.
- Dependencies/assumptions: No personalization without explicit consent; adhere to company trust-and-safety policies.
- Trust & safety operations
- Use case: Guardrails that detect and limit over-inference or manipulative sequencing in assistant responses.
- Tool/product: Cognitive-constraint policy tuning (e.g., cap β-like behavior); monitoring for primacy/recency exploit patterns.
- Dependencies/assumptions: Policy hooks in serving stack; continuous QA to avoid suppressing legitimate information.
- Research engineering and tooling
- Use case: Rapid prototyping of Equation-to-Behavior RL for small/medium models to reduce belief error on structured tasks.
- Tool/product: veRL/GRPO recipe and prompts for belief-accuracy rewards; lightweight evaluation on independent-evidence subsets.
- Dependencies/assumptions: Compute budget; careful reward shaping; verifiable posteriors require independence or tractable models.
Long-Term Applications
These require further research, scaling, domain adaptation, or standards development (notably Equation-to-Behavior RL generalization, domain-specific datasets, and governance frameworks).
- Ethical, profile-aware assistants (opt-in, transparent)
- Use case: Assistants that adapt explanations to a user’s consented cognitive preferences (e.g., preferring causal narratives vs numbers), while minimizing manipulation.
- Tool/product: On-device profile estimator with explicit consent; “explanation mode” selector; bias-aware content planner.
- Dependencies/assumptions: Privacy and consent; robust, validated parameter estimation; independent oversight; clear harm-minimization policies.
- Small/on-device models trained via Equation-to-Behavior RL
- Use case: Edge deployments that faithfully implement specified cognitive behaviors for evaluation or pedagogy without relying on massive APIs.
- Tool/product: Compact E2B-RL model zoo and verification suite; reproducible training pipelines.
- Dependencies/assumptions: Broader datasets to prevent overfitting; formal verification of belief updates; energy/compute constraints.
- Policy impact simulators for public communication
- Use case: Simulating heterogeneous populations’ responses to health advisories, safety recalls, or emergency guidance.
- Tool/product: Population Persuasion Simulator with demographic weighting and cognitive-model mixtures; scenario planning UI.
- Dependencies/assumptions: Representative modeling of populations; fairness and equity audits; human decision-maker input; transparent limitations.
- Legal decision-support for evidence presentation
- Use case: Decision-support that proposes ethically defensible evidence sequences and narratives robust to order effects.
- Tool/product: Evidence Sequencer with judge-profile scenarios; counterargument forecaster; “robust-to-bias” presentation scorer.
- Dependencies/assumptions: Admissibility and professional responsibility rules; judicial acceptance; rigorous validation against real cases.
- Industry standards for persuasion risk benchmarks
- Use case: Model release gates that include Equation-to-Behavior evaluations for order sensitivity, over-inference, and motivated updating.
- Tool/product: Persuasion Safety Benchmark and certification process interoperable across vendors.
- Dependencies/assumptions: Multi-stakeholder standards bodies; alignment with law/regulation; periodic refresh to cover new cognitive models.
- Negotiation and procurement agents robust to diverse counterpart styles
- Use case: Agents trained against mixed receiver distributions to avoid brittle strategies in real negotiations.
- Tool/product: Mixed-Receiver Curriculum for negotiation simulators; adaptive strategy library (e.g., when to emphasize causal links).
- Dependencies/assumptions: Strong guardrails to prevent manipulative use; human approval loops; domain-specific constraints.
- Cross-domain datasets with evidential strength and dependency annotations
- Use case: Extending beyond legal: medical information, financial news, consumer product safety, civic information.
- Tool/product: Semi-automated annotation pipelines with expert verification; dependency modeling tools.
- Dependencies/assumptions: Expert time and cost; licensing; domain generalization; careful handling of sensitive data.
- Estimating cognitive parameters from real logs for evaluation (not targeting)
- Use case: Retrospective audits of systems’ persuasive impact and user decision patterns (with privacy protections).
- Tool/product: Parameter Estimation Toolkit (e.g., fit α–β or affine χ) with aggregation and privacy guarantees.
- Dependencies/assumptions: Access to appropriately consented, anonymized logs; measurement governance; no deployment for covert targeting.
- Human–robot teaming and safety communication
- Use case: Robots that calibrate explanations to improve operator understanding under time pressure without over-influencing.
- Tool/product: Cognitive-model–aware explanation policies; real-time confidence and evidence-order selectors.
- Dependencies/assumptions: Real-time inference reliability; safety certification; human-factors validation.
- Public health and crisis communications
- Use case: Design of announcement sequences that retain accuracy while minimizing harmful primacy/recency effects in stressed audiences.
- Tool/product: Scenario rehearsal suite measuring comprehension and action intent across receiver mixes.
- Dependencies/assumptions: Interdisciplinary oversight; field trials; explicit non-manipulation commitments.
- Academic method development: “Executable” cognitive model zoo
- Use case: Rapid integration of richer models (e.g., correlation neglect, explanation-based updating, selective exposure) as prompts and RL objectives.
- Tool/product: Open repository of model specs with unit tests, synthetic generators, and validation protocols.
- Dependencies/assumptions: Mathematical tractability for reward computation; community curation; benchmarking norms.
- Compliance and consumer protection auditing
- Use case: Third-party audits of conversational systems for covert commercial steering and resilience to motivated updating.
- Tool/product: Independent audit kits derived from Equation-to-Behavior tests; reports for regulators and the public.
- Dependencies/assumptions: Access for auditors; legal frameworks; clear pass/fail criteria and remediation processes.
- Debiasing advisors (finance, healthcare, education)
- Use case: Advisory systems that explicitly counteract known biases (e.g., under-inference, base-rate neglect) with consent and transparency.
- Tool/product: Bias-Aware Explainer that adaptively presents priors, likelihoods, and counterexamples; user-adjustable “rationality aids.”
- Dependencies/assumptions: Randomized controlled trials demonstrating benefit; careful UX to avoid reactance; ongoing monitoring for unintended effects.
Glossary
- Affine Distortion: A non-Bayesian belief-updating model that linearly mixes a reference belief with the Bayesian posterior, attenuating updates toward the reference. "A simple departure from Bayesian updating is affine distortion."
- base-rate neglect: A cognitive bias where prior probabilities are underweighted relative to new evidence. "base-rate neglect for "
- Bayesian persuasion: A game-theoretic framework where a sender designs information signals to influence a receiver’s action under Bayesian updating. "as in Bayesian persuasion~\citep{kamenica2011bayesian}"
- Bayesian updating: The normative rule for revising beliefs using Bayes’ rule after observing evidence. "Under Bayesian updating, the posterior belief after observing the realization is given by Bayes' rule:"
- best-response: A decision rule selecting the action that maximizes expected utility given current beliefs. "The Receiver selects a best-response action at round ."
- cheap talk: Costless, non-binding communication between strategic agents that may convey information depending on incentives. "Previous work such as signaling games~\citep{spence1973job}, cheap talk~\citep{crawford1982strategic}, verifiable disclosure~\citep{grossman1981informational, milgrom1981good}, and Bayesian persuasion~\citep{kamenica2011bayesian} examines strategic information transmission among rational agents, usually predicting partial unraveling of private information in equilibrium."
- conservative Bayesianism: A tendency to under-update from priors, pulling posteriors closer to the prior than Bayes would. "and captures the idea of ``conservative Bayesianism''~\citep{edwards1968conservatism}"
- correlation neglect: A bias where decision-makers ignore dependence between signals and treat them as independent. "and correlation neglect, which are also empirically observed in human decision-making and theoretically studied in previous literature."
- Equation-to-Behavior Prompting: A method that translates formal cognitive model equations into prompts to control LLM behavior. "We propose an approach that we call Equation-to-Behavior Prompting for guiding LLMs to match cognitive models,"
- Equation-to-Behavior RL: A reinforcement learning approach that trains LLMs to follow mathematical cognitive-model specifications. "However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5\% in out-of-distribution parameterizations."
- executable specification: A formal, computable description used as a ground-truth rule to evaluate and reward model behavior. "We treat the cognitive model as an executable specification that provides verifiable rewards."
- explanation-based updating: A belief-updating tendency that favors coherent causal or narrative explanations over purely probabilistic sufficiency. "including affine distortion, motivated updating, Grether's - model, and explanation-based updating."
- framing effects: Changes in decisions caused by how information is presented rather than its substantive content. "\citet{duetting2025information} study framing effects in information design, using LLMs as proxies for framing-induced beliefs, and characterize the tractability of framing-only versus joint framing-and-signaling optimization."
- Grether's α-β model: A parametric generalization of Bayes’ rule that raises priors and likelihoods to exponents, capturing biases like base-rate neglect or over-/under-inference. "Grether's - model generalizes Bayesian updating by introducing exponents on the prior and likelihood~\citep{grether1980bayes, grether1992testing}:"
- Group Relative Policy Optimization (GRPO): A policy-gradient RL algorithm that estimates advantages relative to grouped rollouts. "We fine-tune Llama-3.1-8B-Instruct, Qwen-2.5-7B-Instruct, and Mistral-7B-Instruct-v0.2 using the veRL~\citep{sheng2025verl} framework with Group Relative Policy Optimization (GRPO)~\citep{shao2024grpo}."
- information design: The study of crafting information structures or signals to influence decisions in strategic settings. "In parallel, theoretical models of information design have also been applied to analyze human-AI interactions~\citep{xu2024persuasion, fudenberg2025delegation, collina2025emergent}."
- invertible distortion functions: Monotone belief-transformation functions that can be inverted, used to formalize systematic deviations from Bayes. "for systematically distorted updated beliefs with invertible distortion functions, no two distinct updating rules can be unambiguously ranked across all persuasion problems."
- Kullback–Leibler divergence: A measure of discrepancy between probability distributions used as a regularizer in RL training. "$\mathrm{KL}(\pi_\theta \| \pi_{\text{ref})$ denotes the Kullback--Leibler divergence between $\pi_{\text{ref}$ and , and controls the strength of the regularization."
- motivated reasoning: A cognitive process where desires or goals bias how evidence is interpreted or integrated. "Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning."
- motivated updating: A model where agents trade off accuracy (Bayesian proximity) against closeness to a preferred reference belief. "Motivated updating models belief updating as the outcome of a trade-off between the Bayesian posterior and a reference belief."
- over-inference: Excessively strong updating from evidence relative to Bayes, often modeled with β > 1 in Grether’s model. "and over-inference for ."
- partial unraveling: An equilibrium phenomenon where not all private information is revealed despite incentives to signal. "usually predicting partial unraveling of private information in equilibrium."
- persona prompts: Natural-language instructions that specify an agent’s role or traits to shape LLM behavior. "Previous methods largely rely on persona prompts~\citep{maiya2025open, abdulhai2025consistently}, or preference learning~\citep{poddar2024personalizing, li2024personalized}."
- persuasion games: Interactive settings where a sender strategically reveals information to sway a receiver’s beliefs or actions. "evaluate this approach on persuasion games based on legal decision-making."
- primacy bias: A tendency to over-weight early evidence relative to later evidence. "This asymmetry suggests models exhibit a primacy bias analogous to the motivated updating common among humans:"
- Proceedings of the Old Bailey: A historical corpus of London criminal trial transcripts used as the paper’s legal decision-making dataset. "Our dataset is built from The Proceedings of the Old Bailey~\citep{hitchcock2023oldbailey}, a collection of criminal trial transcripts from Londonâs central criminal court, spanning 1674 to 1913."
- signal scheme: A mapping from states to distributions over signals that defines how evidence is generated. "according to the signal scheme "
- signaling games: Games where one party with private information sends signals to influence the beliefs or actions of another. "Previous work such as signaling games~\citep{spence1973job}, cheap talk~\citep{crawford1982strategic}, verifiable disclosure~\citep{grossman1981informational, milgrom1981good}, and Bayesian persuasion~\citep{kamenica2011bayesian} examines strategic information transmission among rational agents,"
- strategic information transmission: Communication where messages are chosen to influence a receiver’s decision under strategic incentives. "examines strategic information transmission among rational agents, usually predicting partial unraveling of private information in equilibrium."
- under-inference: Insufficiently strong updating from evidence relative to Bayes, often modeled with 0 < β < 1 in Grether’s model. "under-inference for "
- verifiable disclosure: Settings where senders can credibly reveal information (or the lack thereof) and receivers update accordingly. "verifiable disclosure~\citep{grossman1981informational, milgrom1981good}"
- verifiable rewards: Objective signals derived from executable specifications used to train agents via reinforcement learning. "We treat the cognitive model as an executable specification that provides verifiable rewards."
Collections
Sign up for free to add this paper to one or more collections.

