Papers
Topics
Authors
Recent
Search
2000 character limit reached

EduArt: An educational-level benchmark for evaluating art history knowledge in large language models

Published 2 Jul 2026 in cs.CL and cs.CV | (2607.02007v1)

Abstract: LLMs now score near ceiling on general benchmarks, but these aggregate measures reveal little about how models behave within single disciplines. Existing art-focused evaluations rely on synthetic questions and rarely report item-level properties. This paper introduces EduArt, an educational-level benchmark for art-historical knowledge and visual reasoning in multimodal LLMs. EduArt comprises 871 human-authored questions from Italian secondary-school exercises and US Advanced Placement Art History exams, spanning two languages and seven formats from multiple choice to in-text word placement and error identification. Twelve models from six provider families were evaluated under a default answer-only condition and a motivation condition requiring written justification, and characterized using Classical Test Theory and a logistic regression isolating the effects of format, language, image presence, and model. The benchmark showed strong psychometric properties (mean discrimination 0.514, 82.3 percent good discriminators), while multiple-choice accuracy saturated near ceiling for six models, showing recognition formats alone cannot distinguish frontier models. Format was a strong independent predictor of accuracy: models exceeding 94 percent on multiple choice fell to 23.9 percent on open completion (Claude Opus 4.6) and 6.2 percent on error identification (Claude Sonnet 4.6). The motivation condition changed accuracy in a predominantly negative, family-dependent direction. These dissociations indicate that art-historical knowledge and the ability to deploy it are distinct capabilities, and that single-format benchmarks overestimate what models can reliably do. Mapping this capability profile is a precondition for responsible use of multimodal LLMs in art-historical scholarship, where tasks demand producing and manipulating content rather than selecting from fixed options.

Summary

  • The paper demonstrates that art history evaluation using educational assessments reveals differential capabilities across diverse question formats.
  • It employs rigorous psychometric analysis and logistic regression to measure task-specific challenges in multimodal and open-format contexts.
  • The benchmark shows that reliance on MCQs inflates performance scores while masking true deficits in visual and reasoning abilities.

EduArt: Educational-Scale Benchmarking of Art History Competence in Multimodal LLMs

Motivation and Benchmark Design

The saturation of benchmark performance in general academic domains by current LLMs and MLLMs motivates discipline-specific, diagnostic evaluation frameworks. General-purpose benchmarks such as MMLU, BIG-Bench, and Humanity’s Last Exam provide broad coverage but lack depth within individual scholarly domains, particularly those with distinct multimodal and interpretative requirements. Existing art-focused benchmarks have relied predominantly on synthetically generated items and have failed to comprehensively report item-level psychometrics, limiting diagnostic capacity (Spinaci et al., 23 Sep 2025), [11375590], [3590773].

EduArt establishes an educational-level, art history-focused evaluation benchmark for multimodal LLMs, drawing exclusively from human-authored educational assessments: Italian secondary-school exercises centered on Renaissance paintings (MyZanichelli) and the US Advanced Placement (AP) Art History exam. The dataset consists of 871 items, spanning seven linguistic and structurally diverse formats (multiple choice, open/completion, spatial placement, error detection, etc.), with a substantial subset demanding actual visual reasoning on associated low-to-medium resolution artworks. The extraction pipeline is rigorous, involving both LLM-assisted structuring and manual validation to ensure high-quality item curation. This dual-source, dual-language, and multimodal corpus provides a non-trivial, content-rich evaluation platform for probing both knowledge and application skills.

Methodology: Psychometrics and Experimental Setup

EduArt adopts Classical Test Theory (CTT) for rigorous psychometric analysis, reporting item difficulty (pp) and discrimination (rpbr_{pb}), and employs logistic regression to decouple the effects of format, language, and image presence on accuracy. Twelve current-generation models, spanning six provider families (OpenAI, Google, Anthropic, Qwen, Mistral, Meta), are evaluated under two conditions: i) answer-only (default) and ii) answer-with-justification (“motivation”). The evaluation targets heterogeneity in question structure, scoring via format-controlled rules, and the presence of images in the input.

Core Findings: Differential Capabilities and Diagnostic Power

Psychometrics: The benchmark demonstrates strong psychometric quality, with a mean discrimination of $0.514$ and 82.3% of items classified as “good” discriminators, ensuring that item performance meaningfully separates model capabilities (Figure 1). Figure 1

Figure 1: Item difficulty (pp) vs. discrimination (rpbr_{pb}) by question type, confirming a broad spread and high diagnosticity for non-MCQ formats.

Format-Sensitive Saturation: Standard closed-set MCQ formats are saturated by six models (MCQ exact-match >90%), mirroring trends on MMLU and other general benchmarks. However, performance collapses on open-format items: for instance, models scoring >94% MCQ (e.g., Claude Opus 4.6) drop to 23.9% on open completion and as low as 6.2% on error identification (Claude Sonnet 4.6). Logistic regression isolates format as a robust independent predictor, and the spread between a model’s best and worst format performance exceeds 50 percentage points for most models (Figure 2). Figure 2

Figure 2: Accuracy of each model across the seven question formats, revealing stark format-dependent variability and highlighting the inadequacy of MCQ-only evaluations.

Image Dependency and Visual Reasoning: Contrary to naive expectations, raw accuracy is higher on image-present items for almost all models, but multivariate analysis reveals that, controlling for format and language, image presence actually reduces the odds of correct response (OR = 0.765, p<p<0.001). This aligns with findings that current multimodal LLMs often exploit language-driven priors rather than genuine visual reasoning, a pattern confirmed in other multimodal and art-focused VQA evaluations [MathVerse, CVQA, 11375590].

Effect of Justification Requirement: Requiring models to provide structured justification (“motivation” condition) generally decreases performance, but in a strongly model-family-dependent manner (Figure 3). All Claude models experience modest gains, while GPT and Gemini models universally degrade (max: −12.8 pp on Gemini 3.1 Flash Lite Preview), suggesting that justification impedes precise structured output except where family-specific training improves joint reasoning and response production. Figure 3

Figure 3: Change in macro-averaged accuracy from default to motivation condition for each model/family, evidencing family-dependent degradation or small improvement when justification is required.

Theoretical Implications

EduArt provides compelling evidence that recognition-based assessment (MCQ) grossly overestimates the range and depth of art-historical problem-solving capabilities in current multimodal LLMs. The marked drop in performance for open, completion, or error-identification tasks—even among frontier models—demonstrates a dissociation between knowledge “possession” and knowledge “deployment.” The item-level psychometric reporting facilitates a shift in benchmarking from aggregate accuracy towards capability profiling, crucial for understanding and managing failure modes as models approach human-level general performance on standard metrics.

Furthermore, logistic modeling demonstrates that genuine multimodal and cross-linguistic competence is still lacking. Adverse format and image coefficients indicate that benchmarking visual reasoning in cultural heritage domains requires explicit design to preclude language-only shortcutting, reinforcing conclusions in MathVerse [MathVerse, 2025] and VisuLogic.

Practical Impact and Prospects for Future Research

EduArt establishes a diagnostic foundation for responsible model deployment in art-historical research, curation, and education. Constraining evaluation to MCQs is demonstrably insufficient; adequate benchmarking must prioritize open, generative, and manipulative content formats, as these are directly aligned with scholarly practice needs (catalogation, attribution, narrative description).

Future research directions should include:

  • Extending benchmarks with higher-resolution art images for tasks requiring fine-grained visual discrimination.
  • Developing automated or LLM-as-a-judge evaluation protocols for open-format items to scale non-MCQ benchmarking [2025.emnlp-main.138].
  • Investigating continuous or IRT-based scoring to increase sensitivity in the top-performing regime (Zhou et al., 21 May 2025).
  • Incorporating adversarial and contamination-resistant design principles to ensure the continued validity of benchmarks as training data evolves (White et al., 2024, Golchin et al., 2023).

Conclusion

EduArt demonstrates that high-fidelity, discipline-specific, and psychometrically validated benchmarks are essential as frontier LLMs close the gap on general academic tasks. Art history, as a multimodal, interpretative, and culturally laden domain, exposes pronounced gaps between surface-level recognition and robust content production and manipulation. Fine-grained item-level reporting and diverse task formats are mandatory prerequisites for meaningful measurement and responsible downstream use, especially in cultural heritage contexts where scholarly reliability is non-negotiable.


References:

  • (2607.02007) EduArt: An educational-level benchmark for evaluating art history knowledge in LLMs
  • [11375590] VQArt-Bench: A Semantically Rich VQA Benchmark for Art and Cultural Heritage
  • (Spinaci et al., 23 Sep 2025) Benchmarking Vision-Language and Multimodal LLMs in Zero-shot and Few-shot Scenarios
  • (Zhou et al., 21 May 2025) Lost in Benchmarks? Rethinking LLM Benchmarking with Item Response Theory
  • [MathVerse, 2025], [CVQA, 2024] and further literature as cited in the paper

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.