CATS QDA
A local-first workspace where qualitative researchers analyze sensitive interview data with AI assistance they can inspect, override, and audit.

The problem
Rigorous qualitative coding, memoing, and theme development take substantial researcher time, so valuable interview and document collections are often analyzed only at limited scale — and AI shortcuts usually hide exactly the judgments researchers need to see.
The approach
An end-to-end analysis workspace that runs entirely on the researcher’s own computer: transcripts, codebooks, themes, memos, search, and data-grounded chat, with agentic workflows that propose codes and organize themes while recording an execution trace the researcher can review stage by stage.
Research basis
The workbench builds the GATOS codebook-generation method — published in Humanities and Social Sciences Communications and applied at scale in the Journal of Engineering Education — into an interactive system. The pipeline a researcher watches run (corpus understanding, candidate generation, novelty analysis, theme organization, quality review) is the published method, made inspectable.
What it does today
- Organizes projects around research questions, with transcript ingestion, speaker attribution, dialogue acts, filtering, and version controls.
- Answers researcher questions from the corpus with grounded chat that cites source turns and reports which retrieval tools ran.
- Generates candidate codebooks through a multi-stage agent pipeline and keeps a persistent execution trace of every stage.
- Structures codebooks with definitions, abstraction levels, theme grouping, and explicit acceptance state — the researcher accepts or rejects, not the model.
- Links analytic memos to files, codes, and themes with provenance labels.
- Runs local-first — database, vector search, and language models on the researcher’s own machine — so sensitive corpora never leave it.
What it does not do
- Not hosted and not installable without a developer-oriented local stack; explicitly single-user and loopback-only, with no authentication or multi-user isolation.
- No completed external pilot or user study.
- The quality indicators it reports (saturation, coverage, coherence, balance) have not been validated against benchmarks.
Inside the prototype



Evidence
Working prototype, demonstrated and documented
Nine interface captures from one running production-mode session against a fully synthetic fixture, each documented with the capability it shows and the limitation it must not obscure.
As of 2026-08-28
Engineering substance
Implemented Next.js frontend and FastAPI backend with PostgreSQL migrations, vector search, durable execution records, automated tests, and architecture decision records covering the local trust boundary.
As of 2026-08-28
Published method underneath
The codebook pipeline implements the GATOS method: introduced in Humanities and Social Sciences Communications (2026), applied to 10,000+ posts in the Journal of Engineering Education (2025).
As of 2026
What we are not claiming
- A working demonstration against synthetic data is evidence that the system exists and runs — not that it improves research quality, productivity, or coding validity. Those claims wait on the planned pilot and evaluation.
- Agent-proposed codes and themes are candidates for researcher judgment, not findings. The design thesis is inspectability, and the honest corollary is that the human review it enables is still required.
Responsible use, privacy, and rights
- Use with real research data stays inside the documented local trust boundary and applicable IRB and data-use agreements; the platform makes analysis local, not consent unnecessary.
- All interface imagery shows a fully synthetic demonstration fixture, labeled as such in-product. Invention-disclosure material and novelty claims remain private.
Where it came from
- Using Large Language Models and Generative AI to Scale Qualitative Data Analysis
Virginia Tech Academy of Data Science Discovery Fund · 2024-2025
Published evidence
- Thematic analysis with open-source generative AI and machine learning: A new method for inductive qualitative codebook development
Humanities and Social Sciences Communications, 2026
- Using generative AI for large-scale qualitative analysis of social media posts to understand why people leave computer science
Journal of Engineering Education, 2025
More from the Studio
All entriesUse it
Course Sphynx
A faculty-facing platform for helping instructors redesign courses through guided, transformation-oriented support.
Explore it
GATOS qualitative analysis workflow
A research workflow for using open-source generative AI and machine learning to support inductive qualitative codebook development.
Explore it
Design Team Meeting Intelligence
A local-first research instrument that turns engineering design-team meetings into traceable evidence about participation, deliberation, and responsible design.