PROVA-GT: PROvenance-centered Visual Analytics for Grounded Theory
Authors
Qifan Zhou (Zhejiang University), Qingqing Bi (Nanyang Technological University), Weihua Zhou (Zhejiang University), Duo Chen (Zhejiang University), Siwei Fu (Zhejiang University)
Presentation
- Session
- Let's dig into the data (from France)
- Time
- Wednesday, Nov 11, 08:00 – 08:12 (US/Eastern) · session 08:00 – 09:30
- Location
- Hall America north
Keywords
Grounded Theory, Agentic Workflow, Design Study
Abstract
In social science research, grounded theory (GT) is widely adopted to analyze and code textual data. While large language models (LLMs) open new possibilities for coding at scale, they complicate the GT workflow by generating excessive codes and introducing hallucinations, posing challenges for experts to audit and revise the coding results. To address these challenges, we conducted a two-year design study and derived a series of domain needs and design tasks via interviews with scientists in entrepreneurial management. Subsequently, we presented PROVA-GT, a visual analytics system facilitating the GT workflow by supporting both the forward process of code generation and the backward process of validation and refinement, which are interconnected and iterative. Specifically, our system features a provenance workflow that leverages the agentic LLM and novel visual designs to assist experts in the backward process. To evaluate the usefulness of PROVA-GT and its generalizability in different domains, we reported two in-depth expert studies on entrepreneurial inspirations and trust erosion, respectively. We also conducted a user study to assess the effectiveness of the agentic workflow.
For Practitioners
Who would be interested in reading this paper? Qualitative researchers and social scientists (e.g., sociologists, psychologists, management and marketing scholars, public health researchers) who use grounded theory or other qualitative coding methods to analyze interview, focus group, or open-ended survey data. Data scientists and NLP practitioners who build or evaluate AI-assisted text‑analysis pipelines, especially those integrating large language models (LLMs) with human oversight. UX/HCI researchers and visual analytics designers studying human‑AI collaboration, explainable AI, and provenance‑centered interfaces for high‑stakes analytical tasks. Developers of qualitative data analysis (QDA) software (commercial or open source, such as ATLAS.ti, NVivo, MAXQDA) looking for new interaction paradigms that combine LLM assistance with transparent, verifiable refinement. Journalists and analysts who systematically code textual sources (e.g., investigative journalists, policy analysts) and want to speed up thematic coding while maintaining evidential traceability. How could practitioners apply what they learn from this paper? Improve rigor and transparency in AI‑assisted coding. The paper’s provenance workflow shows how to trace every generated code and relationship back to the original text fragments, enabling systematic auditing, reducing hallucination risks, and building trust in LLM‑generated codes. Adopt an agentic refinement loop for coding. The ProvAgent approach of tool‑based planning, reasoning, acting, and observing can be replicated in other qualitative tools—allowing researchers to efficiently split, merge, create, or delete codes and relationships with LLM support, while keeping the human in control through explicit confirmation steps. Design bi‑directional visual analytics systems. Practitioners can use the system’s coordinated views (Content, Hierarchy, Linkage, ProvAgent) as a blueprint for building tools that tightly couple code generation and validation, reducing context‑switching and cognitive load. Scale grounded theory analyses. By automating parts of open coding and categorization with LLMs, and by using sequential pattern mining to surface frequent relationships across large text corpora, teams can handle datasets that would be too large for purely manual coding. Train students and new researchers. The system can be used as a teaching platform that makes the iterative nature of grounded theory concrete, demonstrating how codes emerge, evolve, and are validated against evidence—an approach applicable in qualitative methods classrooms. Extend the techniques to other qualitative methods. Although the system is tailored for grounded theory, the provenance‑centered, agent‑assisted workflow can be adapted to thematic analysis, content analysis, or any text‑coding tradition where auditability and iterative refinement are required.