Confirmation Bias Shapes Outlier Identification and Removal in Data Cleaning with Scatterplots
Authors
Shiyao Li (Emory University), Akul Rana (Emory University), Cindy Xiong Bearfield (Georgia Tech), Emily Wall (Emory University)
Presentation
- Session
- Great, now you scattered the data everywhere!
- Time
- Tuesday, Nov 10, 15:24 – 15:36 (US/Eastern) · session 15:00 – 16:30
- Location
- Hall Essex center
Keywords
Confirmation bias, data visualization, outlier removal
Abstract
Confirmation bias can shape how prior beliefs influence both the perception of scatterplot correlations and the identification of outliers during data cleaning. We characterize this bias in two stages: perception bias, in which people interpret a scatterplot’s correlation in ways that align with their beliefs, and outlier bias, in which they selectively remove points to make the resulting data better match those beliefs. In a pre-registered crowdsourced study with 507 participants, we designed an interactive ‘Find Outliers’ task in which participants first expressed their prior beliefs about the relationship shown in a scatterplot before viewing it. Then, if they believed outliers were present, they identified and removed those outliers. Overall, 59% of participants removed data points. Among them, stronger prior beliefs predicted greater perception and outlier bias. They also tended to remove data points in ways that made the remaining data more consistent with their beliefs, especially when the outliers were less visually salient and the data were therefore more open to interpretation. Our findings show that data cleaning is not purely objective. Prior beliefs can guide how people visually identify and remove outliers. We therefore discuss how visualization systems can better support data cleaning workflows to mitigate confirmation bias.
For Practitioners
Data analysts and data scientists who use visualizations for data cleaning may be particularly interested in this work. Practitioners can apply our findings by pre-specifying outlier-handling criteria, comparing data patterns before and after removing points, and documenting the rationale for removal decisions.