Studio

CATS QDA

A local-first workspace where qualitative researchers analyze sensitive interview data with AI assistance they can inspect, override, and audit.

Explore itResearch prototype · Internal · Pilot and evaluation ahead
CATS QDA project workspace showing a synthetic study with research questions, data file, code, theme, and participant counts.
ScreenshotA synthetic qualitative-analysis workspace with research questions and live project counts. The counts come from a deliberately small documentation fixture, not a study result.

The problem

Rigorous qualitative coding, memoing, and theme development take substantial researcher time, so valuable interview and document collections are often analyzed only at limited scale — and AI shortcuts usually hide exactly the judgments researchers need to see.

The approach

An end-to-end analysis workspace that runs entirely on the researcher’s own computer: transcripts, codebooks, themes, memos, search, and data-grounded chat, with agentic workflows that propose codes and organize themes while recording an execution trace the researcher can review stage by stage.

Research basis

The workbench builds the GATOS codebook-generation method — published in Humanities and Social Sciences Communications and applied at scale in the Journal of Engineering Education — into an interactive system. The pipeline a researcher watches run (corpus understanding, candidate generation, novelty analysis, theme organization, quality review) is the published method, made inspectable.

What it does today

  • Organizes projects around research questions, with transcript ingestion, speaker attribution, dialogue acts, filtering, and version controls.
  • Answers researcher questions from the corpus with grounded chat that cites source turns and reports which retrieval tools ran.
  • Generates candidate codebooks through a multi-stage agent pipeline and keeps a persistent execution trace of every stage.
  • Structures codebooks with definitions, abstraction levels, theme grouping, and explicit acceptance state — the researcher accepts or rejects, not the model.
  • Links analytic memos to files, codes, and themes with provenance labels.
  • Runs local-first — database, vector search, and language models on the researcher’s own machine — so sensitive corpora never leave it.

What it does not do

  • Not hosted and not installable without a developer-oriented local stack; explicitly single-user and loopback-only, with no authentication or multi-user isolation.
  • No completed external pilot or user study.
  • The quality indicators it reports (saturation, coverage, coherence, balance) have not been validated against benchmarks.

Inside the prototype

Turn-level transcript review with speaker roles, dialogue acts, filtering, and version controls.
ScreenshotTurn-level transcript review with confirmed speaker roles, dialogue acts, filtering, and version controls. The transcript is fictional text; this does not demonstrate live transcription quality.
A data-grounded chat response citing source turns from the synthetic interview corpus, with tool usage reported.
ScreenshotA data-grounded synthesis citing source turns from the synthetic corpus. The displayed response is a seeded successful run, not a claim of general model reliability.
A five-stage codebook generation execution trace from corpus understanding through quality review.
ScreenshotAn execution trace from corpus understanding through candidate generation, novelty analysis, and theme organization. The run was seeded as a completed example; it does not establish autonomous validity.

Evidence

Working prototype, demonstrated and documented

Nine interface captures from one running production-mode session against a fully synthetic fixture, each documented with the capability it shows and the limitation it must not obscure.

As of 2026-08-28

Engineering substance

Implemented Next.js frontend and FastAPI backend with PostgreSQL migrations, vector search, durable execution records, automated tests, and architecture decision records covering the local trust boundary.

As of 2026-08-28

Published method underneath

The codebook pipeline implements the GATOS method: introduced in Humanities and Social Sciences Communications (2026), applied to 10,000+ posts in the Journal of Engineering Education (2025).

As of 2026

What we are not claiming

  • A working demonstration against synthetic data is evidence that the system exists and runs — not that it improves research quality, productivity, or coding validity. Those claims wait on the planned pilot and evaluation.
  • Agent-proposed codes and themes are candidates for researcher judgment, not findings. The design thesis is inspectability, and the honest corollary is that the human review it enables is still required.

Responsible use, privacy, and rights

  • Use with real research data stays inside the documented local trust boundary and applicable IRB and data-use agreements; the platform makes analysis local, not consent unnecessary.
  • All interface imagery shows a fully synthetic demonstration fixture, labeled as such in-product. Invention-disclosure material and novelty claims remain private.

Where it came from

More from the Studio

All entries