Verify-First: Visual Triage of Structurally Consequential Instability in LLM-Generated Analytical Explanations

Authors

Suhyun Park (Hanwha Systems)

Presentation

Session
Me, Myself, and AI
Time
Wednesday, Nov 11, 14:03 – 14:12 (US/Eastern) · session 13:00 – 14:30
Location
Hall America center

Keywords

Visual analytics, human-centered AI, large language models, uncertainty, verification, triage

Abstract

Large language models (LLMs) can produce different analytical claims across repeated runs, but analysts rarely have enough time to verify every unstable claim. I frame this challenge as a visual triage problem: deciding where limited verification effort should be spent first. I present Verify-First, an interactive visual analytics prototype that extracts and aligns claims across multiple LLM-generated explanations, externalizes their schema-induced relations in a constrained explanation graph, and ranks claims using Verification Leverage. This interpretable heuristic combines cross-run instability, graph-based structural influence, and estimated verification cost; it is intended as a prioritization aid rather than a ground-truth measure of analytical importance. In an initial probe-based evaluation on one OSMI mental-health task, Verify-First improved AvgS@3 from 0.605 to 0.808 over instability-only ranking, while performing similarly to a no-downstream ablation on early-stage metrics. Because claim generation, LLM-provided claim fields, and counterfactual probing are model-dependent, these results constitute an initial structural stress test rather than external validation. The work shifts verification support from merely displaying disagreement toward helping analysts allocate scarce verification attention.

For Practitioners

Data scientists, BI analysts, visualization practitioners, and teams deploying LLM-assisted analytical systems may use this work to prioritize which unstable model-generated claims should be verified first. The proposed triage framing helps practitioners allocate limited verification effort by considering cross-run instability, structural influence within an explanation, and estimated verification cost, while keeping the rationale inspectable and open to human override.