Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human–AI Collaboration

Authors

Yiran Xiao (University of California, Santa Barbara), George Legrady (University of California, Santa Barbara)

Presentation

Session
Is the Model Even Thinking?
Time
Thursday, Nov 12, 08:09 – 08:18 (US/Eastern) · session 08:00 – 09:30
Location
Hall America south

Keywords

Diffusion Models, Visual Analytics, Explainable AI, Human-AI Collaboration, Interactive Systems

Abstract

Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret. Many existing workflows rely on aggregated attention or scalar summaries that separate temporal change from image-space evidence. To address this gap, we present a visual analytics framework for exploring attention dynamics in diffusion models: the step-indexed evolution of token-level cross-attention maps, their temporal concentration, and their spatial relationships. Our approach enables structured analysis of attention behavior across generation steps by integrating quantitative measures with data-driven stage identification in an interactive workflow. Case studies on a structured 60-prompt Stable-Diffusion-class benchmark illustrate recurring, interpretable patterns within this setting and show how linked temporal and spatial views facilitate the observation and discussion of generative processes, supporting more effective human–AI collaboration.

For Practitioners

This paper will be of interest to visualization and HCI practitioners, generative-AI tool builders, ML interpretability and model-evaluation teams, creative-technology developers, and digital artists who work with text-to-image systems. Practitioners can apply the workflow by capturing step-resolved cross-attention, organizing prompts into controlled semantic families, and using linked temporal and spatial views to inspect when prompt tokens stabilize, overlap, or compete during generation. These practices can inform more transparent image-generation interfaces, prompt-analysis tools, and human-AI collaboration workflows.