Studio

GATOS qualitative analysis workflow

A research workflow for using open-source generative AI and machine learning to support inductive qualitative codebook development.

Explore itResearch prototype · In development · Packaging the published method for release

The problem

Large qualitative corpora are difficult to analyze rigorously without hiding methodological decisions or overwhelming human coding capacity.

The approach

The workflow connects model-assisted theme generation, researcher validation, and documented codebook decisions so AI support remains inspectable rather than replacing qualitative judgment.

Research basis

GATOS stands for Generative AI-enabled Theme Organization and Structuring. It was introduced and validated in Humanities and Social Sciences Communications, then applied at scale in the Journal of Engineering Education. The workflow is a published method first and a piece of software second, which is why the method is citable today while the packaged tooling is not yet public.

What it does today

  • Summarizes each unit of raw text against the researcher’s stated research question, producing atomic summary points analogous to observation memos.
  • Embeds those summary points, reduces their dimensionality, and clusters them so semantically related observations group together.
  • Generates candidate codes from each cluster and organizes those codes into higher-level themes.
  • Keeps every intermediate artifact — summaries, clusters, candidate codes — inspectable, so the path from raw text to theme can be audited rather than taken on trust.
  • Runs on open-source models, so a corpus never has to leave the researcher’s own environment.

What it does not do

  • There is no installable package or hosted service yet. Using the method today means implementing it from the papers.
  • It does not ingest data, manage projects, or produce analysis reports for you.
  • It has no interface. The published work describes a workflow, not an application.

Evidence

Method validated against known themes

Across three synthetic datasets built so their underlying themes were known in advance, the workflow generated themes closely matching most of the original sub-themes, and produced progressively fewer new codes as it processed more clusters rather than inventing one per cluster.

As of 2026

Applied at scale to a real corpus

Used to analyze more than 10,000 Reddit posts about why people leave computer science, published in the Journal of Engineering Education.

As of 2025

Funded work

Developed under a $10,000 Virginia Tech Academy of Data Science Discovery Fund award, 2024–2025.

As of 2025

What we are not claiming

  • Validation to date rests on three synthetic datasets whose themes were known in advance, plus one applied study. That is real evidence for recovering known structure, not a general validity claim across arbitrary corpora or domains.
  • It is decision support for qualitative researchers, not a replacement for interpretive judgment. The published method assumes human review at the points where it matters.
  • Agreement between the workflow and a human coder has not been established as a reliability statistic that would satisfy every methodological tradition.

Responsible use, privacy, and rights

  • Researchers remain responsible for the ethics and permissions covering their own corpus; the workflow makes analysis faster, not consent broader.
  • Model temperature is set to zero for determinism, but generative models remain probabilistic and outputs should be treated as candidates for review.

Where it came from

More from the Studio

All entries