NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

Authors

Chuhan Zhang (Zhejiang University), Ye Zhang (Zhejiang University), Bowen Shi (Zhejiang University), Yuyou Gan (Zhejiang University), Tianyu Du (Zhejiang University), Shouling Ji (Zhejiang University), Dazhen Deng (Zhejiang University), Yingcai Wu (Zhejiang University)

Presentation

Session
Let's figure out how things work behind the scenes
Time
Wednesday, Nov 11, 13:48 – 14:00 (US/Eastern) · session 13:00 – 14:30
Location
Hall America north

Keywords

Large Language Model Security, Machine Learning Explainability, Visual Analytics

Abstract

Jailbreak attacks bypass the safety alignment of large language models (LLMs) to elicit harmful outputs, yet the vast parameter space makes diagnosing the underlying failure mechanisms extremely challenging. We present NeuroBreak, a visual analytics system that helps experts progressively unpack jailbreak mechanisms from layer-level semantics down to neuron-level behaviors. A layer-wise probing pipeline traces how harmful representations evolve across layers, while a dual-dimensional character--behavior categorization reveals each safety-related neuron's inherent tendency and contextual contribution. These analyses are made interpretable through tailored visualization designs: a task-driven probing projection that reveals safety decision boundaries, a dual-stream semantic evolution flow that traces cross-layer semantic shifts, and a character--behavior chord graph that unifies neuron roles, attribution scores, and collaborative relations in a single view with in-situ causal verification. Quantitative evaluations and case studies show that NeuroBreak uncovers safety failure causes and provides actionable insights for strengthening LLM defenses.

For Practitioners

Practitioners in LLM safety, model alignment, and mechanistic interpretability may find this work useful. NeuroBreak provides a workflow for diagnosing jailbreak failures by tracing harmful semantic changes across layers and examining the roles and interactions of safety-related neurons. Practitioners can use these analyses to identify vulnerable internal mechanisms, compare failure patterns across attack contexts, and guide more targeted model interventions and safety fine-tuning.