Last updated: 2026-10-01
Grounded Theory at AI Speed: Borrowing Agile Patterns for a Domain That Won't Hold Still
Grounded theory was built for domains that sit still long enough to be understood: gather data, code it, gather more to test the emerging categories against, and keep cycling until new data stops producing new properties — the point Glaser and Strauss named theoretical saturation [1]. Built into that cycle is an assumption that the phenomenon under study changes more slowly than the research chasing it. Study how experienced nurses triage patients, or how small businesses adopt a new accounting practice, and the underlying phenomenon will still resemble itself by the time the write-up is finished.
Study how people use large language models, and it won't. Overtaken by test-time compute, reasoning models, and agentic tool-use pipelines that didn't exist when data collection started, a codebook built around "prompt engineering" can be half-obsolete by the time a paper clears review. Classic saturation assumes the target holds still while the researcher converges on it; here, the target moves too.like chasing a moving target
What Grounded Theory Assumes, and Where the Assumption Breaks FoundationalKnowledge that endures for decades — core principles
The method's core discipline is well established, and runs in three stages. Open coding breaks the data into concepts, properties, and dimensions with no prior category scheme imposed; axial coding then relates those concepts to each other — conditions, actions, interactions, consequences — building them into a structure; and selective coding integrates that structure around a core category, producing the theory itself [2]. Running throughout is theoretical sampling: deliberately seeking out data likely to challenge or extend the current categories, and stopping only when it stops doing so [1].
Grounded theory was designed to build theory from data precisely because the phenomenon wasn't already understood, so nothing about the discipline itself requires a stable target. What it does assume is that the phenomenon changes more slowly than the coding cycle converges on it. In AI and LLM research, among other fast-moving socio-technical domains, that assumption can simply be wrong. A category built from careful axial coding — how users work around a model's context limit, say — can be overtaken by a context-window increase before the analysis is even finished. The ground the researcher is standing on keeps shifting underfoot — not because the domain is unusually hard to study, but because the method's own convergence criterion was never built to outrun a target moving this fast.the theory must be faster than the change
the categories?"} D -->|yes| A D -->|no| E["Theoretical
saturation"] F["Domain shifts:
new capability,
new paradigm"] -.->|invalidates categories
before E is reached| B style E fill:#8FBF6A style F fill:#F2B8B5
Reconceiving Saturation as a Bounded Definition of Done Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Long ago, software teams facing a moving target gave up waiting for stable, complete requirements before shipping anything. Scrum's answer is the sprint: a fixed timebox, at the end of which a team delivers something meeting an agreed Definition of Done, rather than working indefinitely toward a moving notion of "finished" [4]. Applying that same discipline to grounded theory in a fast-moving domain means giving up on waiting for the categories to stop moving altogether, and instead asking a narrower, answerable question: has this sub-theory, bounded to this timebox and this slice of the phenomenon, stopped producing new properties from fresh data drawn from within that same window? That is local saturation — saturation of a bounded sub-theory or architectural behaviour, checked on a sprint cadence — as distinct from the global saturation classic grounded theory pursues across the whole phenomenon, which a genuinely fast-moving domain may not hold still long enough to ever reach.
Grounded theory's own practice already has a precedent for this move, even if it sounds radical at first. Guest, Bunce and Johnson's landmark study of interview-based saturation operationalised the concept empirically, tracking how many new codes each successive interview produced; the great majority of themes in their data set had emerged within the first twelve interviews, with new-code discovery falling off sharply after that [3]. In other words, a diminishing-returns curve stands in for the judgement call "saturation has been reached." Treating that same curve as a live signal — the rate at which a coding pass turns up genuinely new properties, checked against a threshold at the end of each timeboxed increment — extends the precedent from a post-hoc justification into an operational stopping rule for a single sprint's local scope. Connecting Scrum's Definition of Done to grounded theory's saturation criterion this way is this page's own move, not a technique either Glaser and Strauss or Schwaber and Sutherland set out to build.diminishing returns is a statistical signal
Team Consensus as a Second Dynamic Threshold Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
Classic grounded theory builds another discipline on the same assumption of a stable target. Inter-coder consensus — independent coders converging on the same labels for the same data, reconciled until agreement is high and stable across individual codes, properties, and edge cases [2] — carries an identical vulnerability. Holding out for that kind of fine-grained agreement produces the same theoretical lag as holding out for fine-grained saturation: by the time two coders have reconciled every edge case, the phenomenon they're reconciling it about has moved on.
The fix is the same abstraction-ladder move, applied to a different discipline. Rather than treating the required bar for agreement as fixed, scale it against how fast the domain is actually moving:
| Domain velocity | Where consensus is required | What counts as disagreement |
|---|---|---|
| Low (the classic case) | Micro-level: individual codes, properties, sub-categories, edge-case behaviour | Any unreconciled label difference blocks closure. |
| High (LLM and agentic-AI research) | Macro-level: the core relational structure and its primary interactions only | Micro-level variation is logged as temporal variance — a property of the category, not a reason to delay it. |
That reframes the question a team is actually answering at a sprint boundary. It stops being "do we agree on every label" and becomes "do we agree this high-level construct explains the phenomenon as it currently stands, and is it stable enough to publish as a baseline." Scaling the required consensus bar against domain velocity this way is this page's own extension of the sprint-bounded saturation idea above, not a documented technique in either the grounded-theory or Agile literature.
Beyond simply settling disagreements faster, a scheduled team reconciliation point does double duty. Left to work alone, an individual researcher is prone to getting caught up in whatever the newest model release or capability happens to be — a high-frequency signal that may or may not mean anything structurally. Routing observations through the whole team at a fixed sprint boundary, rather than letting consensus emerge whenever it happens to, filters that noise: a pattern only one researcher has noticed stays a candidate, and a pattern the team independently converges on becomes a category. When two coders do split on how to divide a category under time pressure, the tie-breaking rule that keeps a fast-moving research programme usable favours consolidation over proliferation — merging the split, not forking it — since runaway proliferation of ever-finer categories is exactly the endless-refactoring failure mode a velocity-aware framework exists to avoid.
observation"] --> G{"Sprint-boundary
consensus gate"} R2["Researcher B's
observation"] --> G G -->|independently
replicated| C["Stable macro
category"] G -->|single-researcher
only| N["Logged as candidate,
not yet a category"] style C fill:#8FBF6A style N fill:#FFF3CD
The practical stopping test that falls out of this is predictive rather than definitional: the team has reached working consensus once its members, categorising new incoming data independently, agree on most of it — not once every taxonomy boundary has been argued to a close. Take a domain whose working assumptions turn over roughly every three months, a plausible rough estimate for how fast agentic-AI tooling conventions have been shifting recently. Spend six months chasing 95%-plus micro-level agreement on a fine-grained taxonomy in that domain, and the resulting theory describes a phenomenon that has already moved on. Settle for roughly 80% macro-level agreement inside a three-week sprint instead, and the field gets something it can still use while it's current — with the unresolved 20% logged as the seed of the next sprint's refactoring, not a defect to clear before anything ships.predictive consensus: can we classify the next case?
Mapping Agile Constructs onto Grounded-Theory Practice Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice
The mapping holds at more than one point, which is what makes it worth taking seriously rather than treating as a one-off analogy:
| Agile construct | Grounded-theory equivalent | What it does in a fast-moving domain |
|---|---|---|
| Backlog | The pool of uncoded data — transcripts, logs, benchmark runs, literature not yet reviewed | Makes explicit that not all evidence gets analysed at once, and lets the queue be reprioritised as new theoretical gaps open up. |
| Sprint / timebox | One bounded coding-and-analysis increment | Forces a decision point before data collection and refinement can run on indefinitely chasing a moving target. |
| Minimum viable theory | A published mid-range theory, scoped explicitly to the current paradigm | Gets findings into circulation while they're still accurate, rather than holding out for a completeness the domain won't sit still for. |
| Refactoring | Re-running axial or selective coding on existing categories | Updates higher-level categories when a new capability or paradigm invalidates the lower-level constructs they were built on. |
| Pivot | A formal, documented narrowing or shift of theoretical scope | Names the moment a foundational assumption breaks (a shift from prompt engineering to agentic workflows, say) as a deliberate scope decision, not a quiet abandonment of earlier work. |
uncoded data"] --> SP["Timeboxed
coding sprint"] SP --> V{"New-code rate
below threshold?"} V -->|no| SP V -->|yes| MVT["Publish minimum
viable theory"] MVT --> BL SP -.->|foundational assumption
breaks| PIV["Pivot:
redefine theory scope"] PIV --> BL style MVT fill:#8FBF6A style PIV fill:#FFC857
Separating What Decays Fast from What Doesn't FoundationalKnowledge that endures for decades — core principles
Not every category a fast-moving domain produces decays at the same rate. One built around a specific prompt phrasing, a specific model's quirks, or last quarter's state-of-the-art benchmark score goes stale quickly, because the artefact it describes is itself ephemeral. Another — built around how people recover from a tool's failure, how trust shifts after a visible error, or how responsibility gets negotiated between a human and an automated system — tends to persist across model generations, because it describes a pattern in the interaction rather than a property of one implementation. Sorting categories deliberately into these two layers while coding, and aiming the theory itself at the slower-moving one, keeps a mid-range theory useful past the lifespan of whichever model or technique happened to be current when the data was gathered.
Two further habits protect that separation in practice. Treating a code's timestamp, and the model generation it was collected under, as a property recorded within the coding matrix — rather than as an embarrassment to be smoothed over — turns an apparent contradiction between old and new data into a dimension of the category itself: the theory doesn't just say what happens, it says what happens under GPT-4-era tool use versus what happens once agentic pipelines are common. And publishing in increments — working papers and scoped mid-range theories released as each timebox closes — spreads the risk that any single release goes stale, rather than concentrating years of unpublished synthesis into one monolithic theory that a single paradigm shift can retire overnight.timestamp as a variable, not noise
Where the Analogy Breaks Down FoundationalKnowledge that endures for decades — core principles
This framing has real limits, and they matter more here than they would for a looser analogy.
- There is no product owner for a scientific theory. A sprint review has someone in the room with the authority to accept the increment as done. For a grounded theory, the real acceptance test is the wider research community instead — peer review, replication, citation by others working the same seam — and that feedback loop runs on a timescale of years, not two weeks. A category that survives a sprint-bounded local saturation check has cleared a much lower bar than one the field has actually stress-tested, and reporting the two with the same confidence would be a mistake this framing doesn't excuse.
- A velocity metric can be gamed, including by accident. Once "new-code rate below a threshold" becomes the criterion for calling a sprint done, a researcher under time pressure has an incentive to code more coarsely, or to stop looking as hard for disconfirming data — the same target-distorts-behaviour risk this site's discussion of Goodhart's law in rubric design raises for a different measurement. A velocity threshold is a useful trigger for a human's saturation judgement, not a replacement for it.
- Data quality doesn't reduce to queue depth. A software backlog's items are broadly comparable units of work; an interview transcript, a benchmark log, and a single line from a forum post are not — and treating the backlog metaphor too literally risks flattening those genuinely different evidentiary weights into one undifferentiated queue.
Where This Is Already Being Tried Ephemeral / ToolingKnowledge that evolves in months to a year — check for updates
Independent of any Agile framing, the broader idea of operationalising saturation at all is an active, unsettled area rather than a solved one. A 2025 preprint proposes a machine-learning decision-support system that predicts an appropriate sample size in advance from ten input parameters — research scope, information power, and researcher competence among them — rather than monitoring code emergence during analysis itself [5]. That's a different mechanism from the sprint-bounded, in-progress velocity check described above — it operates before data collection rather than during it — but it's evidence that treating saturation as something other than pure researcher intuition is a live methodological question well beyond this page, not a niche concern invented for AI research specifically.
Related Topics
- Project Marking as BDD — the same move in a different direction: borrowing a software-engineering discipline (Given/When/Then) to sharpen judgement in a domain that isn't software.
- Evaluating Detailed Rubrics — the Goodhart's-law risk of a measurement becoming a target, raised there for marking criteria and here for a saturation velocity metric.
- How LLMs Reason, Self-Correct and Get Checked — background on the reasoning-model and test-time-compute shifts used above as examples of categories a codebook can be overtaken by.
- Synthesis by Humans and Machines — the wider question of how synthesis work divides between a human researcher and an automated system.
References
- Glaser, B. G., & Strauss, A. L. (1967). The Discovery of Grounded Theory: Strategies for Qualitative Research. Aldine.
- Strauss, A., & Corbin, J. (1990). Basics of Qualitative Research: Grounded Theory Procedures and Techniques. Sage Publications.
- Guest, G., Bunce, A., & Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), 59–82. https://doi.org/10.1177/1525822X05279903
- Schwaber, K., & Sutherland, J. (2020). The Scrum Guide. https://scrumguides.org/scrum-guide.html
- Tutar, H., Erden, C., & Şentürk, Ü. (2025). Q-Sat AI: Machine learning-based decision support for data saturation in qualitative studies. arXiv:2511.01935 (preprint).