KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
Authors
Yan Zhu (Tulane University), Yongbo Chen (Tulane University), Rebecca Faust (Tulane University)
Presentation
- Session
- Please don't just stare at the picture
- Time
- Tuesday, Nov 10, 10:24 – 10:36 (US/Eastern) · session 10:00 – 11:30
- Location
- Hall America center
Links
Sign in to access the preprint PDF.
Sign in- Download Supplemental Material
Keywords
Text Data, Dimensionality Reduction, Semantic Interaction
Abstract
In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.
For Practitioners
Practitioners: Researchers and practitioners working in human-AI interaction, interactive AI, visual analytics, and human-in-the-loop machine learning, as well as data scientists and domain experts who analyze large text collections. Application: Practitioners can use KeySI to interactively refine text embeddings, improve the semantic organization of document collections, and better understand relationships among documents during exploratory analysis. The approach can support document exploration, corpus analysis, topic discovery, and other human-in-the-loop visual analytics workflows.