A Multimodal Reasoning Typology for Grounding Chart-Image Coherence in Science Communication
Authors
Avina Nakarmi (New Jersey Institute of Technology), Sohom Sen (New Jersey Institute of Technology), Xun Song (New Jersey Institute of Technology), Sreyashi Samaddar (Brooklyn College), Aritra Dasgupta (New Jersey Institute of Technology)
Presentation
- Session
- Me, Myself, and AI
- Time
- Wednesday, Nov 11, 13:36 – 13:45 (US/Eastern) · session 13:00 – 14:30
- Location
- Hall America center
Links
Sign in to access the preprint PDF.
Sign in- Download Supplemental Material
Keywords
Multimodal reasoning, science communication, grounding, neuroscience, chart comprehension
Abstract
Charts and images appear together throughout scientific publications, yet most computational work does not characterize their coherence. We argue that a chart, its accompanying image, and the caption that links them form a multimodal unit, and that the inferential work required to read it varies systematically. To capture this variation, we develop a typology of reasoning gaps, R1 through R5, that characterizes how chart, image, and text jointly convey a scientific claim, and the interpretive work this demands of the reader. Some pairs restate the same data, while in other pairs, charts are used to quantify a structure the image localizes, project image content onto an external variable, audit an image-based claim, or jointly construct a frame that neither panel can establish alone. The typology is anchored in the grounding theory of communication and was derived bottom-up, with a neuroscience expert, from a corpus of 79 traumatic brain injury papers and 32 chart-image pairs. Crucially, the levels provide a systematic mechanism for identifying where grounding succeeds or breaks down, rather than leaving it to subjective inference. We show this in a study in which a domain expert and three non-experts judge vision-language model (VLM) descriptions of 25 pairs: the level predicts where their judgments align and where they diverge, isolating the points at which contextual knowledge, not the figure, carries coherence. This typology thus offers figure designers a systematic way to balance text against chart-image pairs, bridging the expert-to-non-expert divide in reading a scientific takeaway.
For Practitioners
Visualization designers, domain experts, and science communicators/journalists can critically evaluate how the reasoning gap framework fits their domain/problem and adopt it accordingly to ensure clear semantic interpretation of multimodal units (e.g., chart-image pairs complemented by text descriptions).