Papers
Topics
Authors
Recent
Search
2000 character limit reached

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

Published 12 Jul 2026 in cs.CL | (2607.10511v1)

Abstract: When an LLM judge calls a peer review analytical and a human committee calls another review high quality, are they tracking the same thing? We argue they are not, and that the difference matters philosophically. We operationalise Kahneman's dual-process theory into a structured rubric for peer review and release Kahneman4Review, a benchmark of 3,563 rated reviews scored along nine theoretically motivated textual dimensions, eight bias diagnostics, and a continuous reasoning-quality score. Three findings bear on trustworthiness: decision tier is not detectably aligned with the rubric's text-grounded epistemic-quality proxy; public-showcase agentic reviews receive higher raw scores than pooled human reviews, but length and venue explain most of the gap and the samples are not paper-paired; and ICLR review-text diagnostics shift at the 2022--2023 transition, temporally coincident with widespread LLM availability but without identifying its cause. A matched function-probe pilot further shows that the rubric distinguishes textual probes designed to contrast genuine fault-finding with surface fluency. We argue that a trustworthy reliability benchmark for LLM judges must separate analytical form from epistemic function, and propose concrete design choices toward that goal. An interactive demo is available at https://huggingface.co/spaces/nuojohnchen/Kahneman4Review.

Summary

  • The paper reveals that LLM-based reviewers mainly capture superficial analytic fluency rather than genuine fault-finding, challenging epistemic reliability.
  • It introduces Kahneman4Review, a benchmark with 3,563 annotated reviews using dual-process theory to evaluate analytical trace quality across top ML conferences.
  • Empirical results show that while LLM-generated reviews score higher due to verbosity, these scores do not reflect underlying epistemic function.

Articulate Intuition or Genuine Analysis? Benchmarking Epistemic Reliability in LLM-as-a-Judge Peer Reviews

Introduction and Problem Statement

This paper critically addresses the utility and limitations of LLMs as judges in the peer review process, dissecting whether LLMs can reliably distinguish between surface-level analytical form and authentic epistemic analysis in review texts. The investigation is situated at the intersection of empirical benchmarking, philosophical epistemology, and scientific evaluation, fundamentally questioning if textual analytic traces captured by LLM metrics faithfully correspond to genuine scientific reasoning or merely articulate fluency.

Methodological Framework: Theory-Driven Rubric and Dataset

The authors develop Kahneman4Review, a benchmark comprising 3,563 annotated reviews, stratified across leading ML venues (ICLR, ICML, NeurIPS) and a set of agentic LLM-generated reviews. The annotation rubric is grounded in dual-process theory, operationalizing nine textual dimensions (e.g., impression reliance, reasoning chain, evidence specificity, falsifiability, heuristic usage, uncertainty handling, issue prioritization, reasoning density, core criticality) and eight recognizable bias diagnostics. Reviews are also mapped into System 1, System 2, Mixed, or Non-evaluative categories, reflecting the dominance of associative or deliberative reasoning traces.

LLM judges (primarily claude-sonnet-4-6) rate reviews without access to source papers, furnishing multi-dimensional scores, epistemic labels, and bias tags. The rubric explicitly separates "analytical form" (the textual architecture of reasoning) from "epistemic function" (evidence of genuine, falsifiable, and response-sensitive critique).

Empirical Findings

Venue-Level and Temporal Analysis

Analysis of the human-review corpus reveals a field-wide analytical-trace ceiling: Reasoning Quality Score (RQS) distributions are tightly concentrated (mean RQS∈[2.80,2.94]\mathrm{mean~RQS} \in [2.80, 2.94]) across top ML conferences, with the ICLR venue skewed toward System-1-like traces compared to ICML/NeurIPS. Figure 1

Figure 1: Label distribution across three human venues and the Stanford agentic reviewer, indicating significant inter-venue differences.

Temporal analysis shows a significant post-2022 shift toward System-1-like reviews, temporally coincident with widespread LLM deployment. Specifically, System-1-like diagnostic prevalence in ICLR reviews nearly triples from 2021/2022 to 2023/2025, while System-2-like traces correspondingly diminish. Mean RQS drops monotonically, with a length-standardized gap of −0.21-0.21 RQS units (p ≪ 0.001) from 2021 to 2025. Figure 2

Figure 2: ICLR review-text diagnostics across years; left panel shows label distribution shift, right panel shows monotonic RQS decline.

Relationship Between Review Outcomes and Analytical-Trace Quality

Outcome signals (acceptance tier, area chair judgments) are not predictive of RQS or other rubric-grounded proxies of reasoning quality; within-venue and across-venue comparisons show null effects (p=0.260p=0.260 for tier, p=0.578p=0.578 for AC commentary group vs peers). Figure 3

Figure 3: Mean RQS by acceptance tier per venue; pooled ANOVA reveals no significant differences.

Figure 4

Figure 4: RQS distribution by AC annotation group; rubric-based signals and AC judgments are weakly aligned at best.

Evaluating Agentic LLM Reviewers

LLM-generated agentic reviews (e.g., Stanford agentic reviewer) earn substantially higher analytic-trace scores than human reviews: 74% System-2-classified vs 21% for humans, mean RQS 3.50 vs 2.86 (d=1.26d=1.26, 95% CI [1.19,1.33][1.19, 1.33]). All nine rubric dimensions show large, significant effect sizes post-multiple-testing correction.

However, length is a strong confound: length and venue-specific controls shrink the raw AI-human gap by 74%, yielding an adjusted estimate (β^1=+0.17\hat\beta_1 = +0.17, p = 0.014), indicating that most of the LLM advantage is attributable to extended verbosity and venue effects rather than underlying epistemic quality. Figure 5

Figure 5: Bias-diagnostic frequency per group; Checklist Inflation dominates, with agentic LLM reviews exhibiting less Question Substitution and Authority Substitution compared to humans.

Figure 6

Figure 6: PCA of rubric dimensions; humans cluster around the origin, while the agentic reviewer shifts +1.7 on the "analytical-trace" axis, reflecting unidimensional score saturation by LLMs.

Bias Profiles and Co-occurrence in Human and LLM Reviews

Bias diagnostics reveal frequent Checklist Inflation in both human and LLM reviews, but with differential profiles: agentic reviewers exhibit very little Question or Authority Substitution. Human reviews demonstrate weak but notable co-occurrence patterns for System-1-like heuristics (e.g., Representativeness and Question Substitution). Figure 7

Figure 7: Bias co-occurrence in human reviews; primary clusters support weak System-1-like trace patterning.

Validation and Probing

Test-retest reliability for the rubric under a fixed LLM prompt is robust (κ=0.77,α=0.94\kappa=0.77, \alpha=0.94), but model-version sensitivity is high (κ=0.14\kappa=0.14 across LLM versions), underscoring model-conditional interpretation of scores.

A targeted probe experiment contrasting "genuine-fault-finding" and "surface fluency" reviews yields strong rubric discrimination: in 33/35 pairs, the genuine-fault review is scored higher (dz=2.52,p<0.001d_z=2.52, p<0.001), all labeled System 2. Figure 8

Figure 8: Latent (−0.21-0.210, −0.21-0.211) space; agentic AI reviewer is distinctly displaced from the human diagonal, reflecting high System-2 and low System-1 trace incidence.

Theoretical and Practical Implications

The principal finding is that current LLM judges capture the surface form of analytical reasoning but cannot guarantee underlying epistemic function; PCA reveals that the nine-dimensional rubric is empirically collapsed into a single analytical-trace axis, reflecting LLM-induced near-unidimensionality. Thus, LLMs reward text exhibiting superficial markers of analysis without verifying whether the review manifests genuine fault-finding, falsifiability, or domain-sensitive critique.

On a practical level, this challenges the use of LLM-based review quality assessments for gatekeeping unless corroborated by span-level audit and paper-grounded validation. Theoretical implications are aligned with analytic philosophy: epistemic reliability cannot be imputed from text-form alone, as analytic fluency and genuine understanding are not isomorphic.

Concrete recommendations for benchmarking design include: (1) Span-level auditable judgments rather than global scores to enforce falsifiability; (2) Enforced counterfactual justification for labels to stimulate contrastive reasoning; (3) Calibration against curated, public-errata probes to disentangle analytical form from epistemic function.

Conclusion

The research exposes critical limitations in the epistemic reliability of LLM judges in peer review contexts. While length-controlled, rubric-driven, and systematically validated, the Kahneman4Review framework demonstrates that textual analytic-trace measurements are dominated by surface form. Current LLM-based evaluation pipelines are susceptible to over-crediting analytical fluency at the expense of genuine critique. Future work must focus on span-level audits, public-errata-grounded probes, and explicit model-conditional reporting to advance trustworthy automation in scientific evaluation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.