2000 character limit reached
Scientific Reasoning: Assessment of Multimodal Generative LLMs
Published 3 Mar 2025 in cs.CL, cs.AI, and cs.CV | (2503.01064v1)
Abstract: LLMs can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data.
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.