Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Authors
Patrick Phuoc Do (University of Notre Dame), Chau Ta (University of Notre Dame), Chaoli Wang (University of Notre Dame)
Presentation
- Session
- Me, Myself, and AI
- Time
- Wednesday, Nov 11, 13:45 – 13:54 (US/Eastern) · session 13:00 – 14:30
- Location
- Hall America center
Links
Sign in to access the preprint PDF.
Sign in- Download Supplemental Material
Keywords
Visualization Literacy; Scientific Visualization; Multimodal Large Language Model
Abstract
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (SciVis). We benchmark six MLLMs on the scientific visualization literacy assessment test, a standardized SciVis literacy assessment comprising 49 items based on 18 scientific visualizations and illustrations, spanning 8 techniques and 11 task types. We evaluate three closed-source and three open-source models under a closed-world protocol and compare their performance using data from 485 human participants. Results show that current MLLMs do not exhibit uniform SciVis literacy. Gemini is the strongest model overall, exceeding the human mean across the evaluated subsets, whereas the open-source models remain below the human baseline. Performance is highly uneven across techniques and tasks: models perform best on scientific illustration, search, and spatial understanding, but struggle on texture-based and integration-based visualizations and on quantitative estimation. Error analysis reveals recurring failures in fine-grained quantitative estimation, flow-direction interpretation, and grounded encoding interpretation. These findings position SciVis literacy as a necessary benchmark dimension for evaluating multimodal AI systems. Our code and model outputs are publicly available at https://github.com/patdmp/mllm-scivis-lit-benchmark.
For Practitioners
Visualization researchers, AI/ML practitioners integrating vision-language models into data analysis pipelines, and science communicators or educators who use scientific visualizations to convey findings would find this paper relevant. AI/ML practitioners can use our results to understand where current MLLMs fail on specific visualization types and avoid over-relying on them for automated interpretation of scientific visualizations. They also gain a ready-made benchmark (SVLAT) for evaluating new models on scientific visualizations. Educators and science communicators can learn which scientific visualization techniques and tasks are hardest for AI to interpret, informing their material design when AI-assisted reading is part of the workflow.