Data Saturation Qualitative Research: How Many Interviews Are Enough?

Data saturation qualitative research is one of those phrases scholars only search for once panic has already set in. A supervisor asks how you decided on your interview count, and you realise you never really decided. You just stopped when you got tired. I have guided dissertation scholars through this exact moment for over twelve years. The honest answer is that no fixed number works for everyone, but there is a defensible way to arrive at one, and that is what this article gives you.

For most qualitative dissertations, 9 to 17 interviews is enough to reach code saturation, and 16 to 24 if you need the deeper meaning saturation that grounded theory or phenomenology studies often demand. That is the short way of answering how many interviews for qualitative research is enough, though the full picture below explains why that range moves depending on your topic and your participants.

What Is Data Saturation in Qualitative Research, Really?

Data saturation is the point in a qualitative study where new interviews stop producing new themes, codes, or insights, and the researcher begins hearing the same patterns repeated in different words. This is the definition I give every client, because it is the one examiners actually look for in a methods chapter. The idea traces back to Glaser and Strauss’s original work on grounded theory, long before software or structured checklists existed to test it.

People often use “data saturation,” “code saturation,” and “thematic saturation” as if they mean the same thing. They do not, and mixing them up is one of the most common errors I see in draft methodology chapters.

  • Code saturation is reached when no new codes emerge from fresh transcripts.
  • Thematic saturation goes one level up, when the broader themes built from those codes stop changing.
  • Meaning saturation is the deepest level, where your understanding of why a theme exists stops evolving, not just whether it appears.

If your dissertation only claims code saturation but your research question needs a rich, interpretive understanding, an examiner will notice the mismatch. This distinction alone has saved several of my clients from a difficult viva question.

What the Research Actually Says (And Where I Disagree With It)

Three studies get quoted endlessly on this topic, and I think each one is used more casually than it deserves.

Guest, Bunce and Johnson (2006): The Study Everyone Misquotes

Guest, Bunce and Johnson found that 92 percent of themes in their study of West African sex workers appeared within the first 12 interviews. Students now treat “12 interviews” as a universal rule. It is not. That study used a fairly homogeneous population and a narrowly defined question, so applying it to a diverse, cross-cultural sample will likely leave you short.

Hennink, Kaiser and Marconi (2017): A More Honest Split

Hennink, Kaiser and Marconi refined this further, separating code saturation, 9 to 17 interviews, from meaning saturation, 16 to 24 interviews. I find this the more honest framework of the two, because it forces you to ask which kind of saturation your research question actually demands, rather than chasing one convenient number.

Mason (2010): What 560 Theses Actually Reveal

Mason analysed 560 PhD theses using qualitative interviews and found a mean sample size of 31, with a median of 28. Here is the part I find most telling and rarely discussed: a statistically significant share of these studies reported sample sizes that were exact multiples of ten. That is not saturation, that is rounding. In my experience advising scholars, committees reward tidy numbers more than they reward genuine methodological honesty, and this data point backs up what I have seen for years.

FrameworkSuggested RangeBest Fit
Guest, Bunce & Johnson (2006)12 interviewsHomogeneous, narrow topic
Hennink, Kaiser & Marconi (2017)9-17 (code), 16-24 (meaning)Most standard dissertations
Mason (2010, empirical average)28-31General PhD benchmark

Sample Size by Research Design

The number also depends heavily on which qualitative approach you are running, a point that gets lost when people search for “sample size for qualitative study” as if one answer fits every design.

  • Phenomenology: 3 to 10 participants, since the goal is depth of lived experience, not breadth.
  • Grounded theory: 20 to 30 participants, because the method is built around developing a theory from patterns.
  • Case study: typically 1 to 5 cases, with saturation applying within each case rather than across a large sample.
  • Generic thematic analysis: 5 to 25, the widest and least precise category, which is exactly why so many dissertation scholars struggle here.

These figures broadly follow Creswell’s widely used methodology guidelines, though I always tell clients to treat them as a starting range, not a rulebook. If your dissertation blends approaches, for example a design that mixes case analysis with survey data, your justification needs to address saturation separately for each strand, not as one combined figure.

What Actually Drives Your Number Up or Down

The frameworks above are starting points, not fixed rules. In my consulting practice, four factors consistently push a scholar’s real number away from the textbook range.

  • Population homogeneity. A tightly defined group, for example only first-year nursing students at one hospital, saturates faster than a mixed group of professionals across industries.
  • Scope and complexity of the question. A narrow question about one specific workplace policy saturates faster than a broad question about career identity.
  • Interviewer depth, not just interview count. A skilled, semi-structured interview that probes follow-up answers reaches saturation with fewer sessions than a rushed, surface-level one.
  • Purposive sampling choices. How deliberately you select participants for maximum variation, rather than convenience, changes how quickly genuine saturation appears versus a false plateau.

Malterud’s “Information Power” Model, and Why I Only Half Recommend It

Malterud, Siersma and Guassora (2016) argued that sample size should depend on information power rather than a fixed count. In short, the more relevant information each participant carries, the fewer participants you need. Their model rests on four factors:

  • How broad or narrow your study aim is
  • How specific your sample is to that aim
  • How strong the dialogue quality is in each interview
  • What analysis strategy you plan to use

It is a more sophisticated way to think about qualitative rigour than a fixed number. My honest opinion, after using this with clients, is that examiners still want a number at the proposal stage, even if information power is the better underlying logic. I tell scholars to plan a starting range using Hennink’s figures, then justify any deviation using Malterud’s reasoning once data collection is underway. This hybrid approach has held up well in vivas I have supported.

Fugard and Potts later pushed this further by turning saturation into an actual calculation, estimating the odds of spotting a theme based on how common that theme is in your wider population. It is a useful cross-check once you have a draft codebook, though I have rarely seen a dissertation committee ask for it by name.

How Do You Know You Have Actually Reached Saturation?

Guest, Namey and Chen (2020) proposed a practical three-part method that I now use as a working checklist with clients:

  1. Set a base size. Code your first batch of interviews (commonly 6 to 10) to build an initial codebook.
  2. Track run length. Keep coding in small batches and note how many transcripts pass with zero new codes in a row.
  3. Apply a new information threshold. Decide in advance what counts as “new enough” to matter, for example a code appearing in fewer than 5 percent of transcripts.

Watch out for false saturation. If your sample is narrow, for instance all participants from one department of one company, you may hit apparent saturation quickly simply because you never asked a different kind of person a question. Real saturation should survive being tested against a more diverse sub-sample. A second coder reviewing a sample of your transcripts, often called inter-rater checking, adds real weight to a saturation claim if your committee asks how you know your own coding was consistent.

Tracking Saturation in NVivo

Most of my clients use NVivo for coding, and saturation tracking is where the software earns its cost. Build one running codebook from the start, and after every two or three interviews, run a code frequency query. When that query stops surfacing new codes and existing ones simply gain more references, you are approaching saturation. ATLAS.ti offers similar frequency tracking if that is what your university licenses. The underlying principle stays the same regardless of software.

I worked with an MBA scholar last year studying employee trust in hybrid work policies. She was convinced 8 interviews were enough because her transcripts “felt repetitive.” Our coaching session on NVivo coding showed two entirely new sub-themes appearing in interview 11, tied to a department she had under-sampled. We stopped at 15, and her examiners specifically praised the saturation table in her defence.

Justifying Your Sample Size in the Methods Chapter

Examiners are not impressed by a number alone, they want the reasoning behind it. A line I often help scholars write looks something like this: “Data collection continued until no new codes emerged across three consecutive interviews, consistent with Hennink et al.’s (2017) code saturation threshold, resulting in a final sample of 14 participants.”

If your supervisor insists on a number before you have collected any data, be upfront that it is a planning estimate, not a saturation claim, and revisit it once coding begins. This single clarification has resolved more supervisor disagreements than anything else I suggest.

My Take, From the Consulting Chair

I once had a client researching consumer trust in fintech apps who was told by a well-meaning peer that 30 interviews was the “safe” academic number. She had a very specific, homogeneous participant group and reached genuine thematic saturation at interview 13. Padding her sample to 30 would have wasted months and diluted her analysis with repetitive data. Every interview beyond genuine saturation also means more transcription hours and more coding time, which adds up fast when you are funding your own fieldwork. Saturation is a quality signal, not a quantity target, and treating it as a checkbox is the single biggest mistake I see in dissertation planning.

If you are deciding between qualitative and mixed approaches at the stage where you choose your data sources, or weighing a cross-sectional versus longitudinal structure, saturation planning should happen at that same stage, not as an afterthought once fieldwork has started.

FAQ

How many interviews are enough for qualitative research? Most studies settle between 9 and 17 interviews for code saturation, rising to 16 to 24 for deeper meaning saturation. Grounded theory studies often run higher, and narrow phenomenological studies can finish sooner.

What is the difference between data saturation and sample size? Sample size is simply the number of participants you interview. Data saturation is the analytical judgment that no new information is being added, which is what justifies that number after the fact.

Is 10 interviews enough for a qualitative dissertation? It depends on your design. For a narrow phenomenological study with a homogeneous group, yes. For grounded theory or a diverse population, likely not.

How do you report saturation in a dissertation methodology chapter? State the point at which no new codes or themes emerged, reference the framework you used (such as Hennink et al., 2017), and give the final participant count with a brief justification.

Can NVivo tell you when you have reached saturation? Not automatically. NVivo helps you track code frequency and spot the plateau, but the judgment call is still yours to make and defend.

What if my supervisor demands an exact number before I start collecting data? Give a planning estimate based on your methodology type, and be clear that it is provisional. Confirm the final figure once your coding shows new codes have actually levelled off, which is where a structured approach to analysing your data makes the real difference.

Does saturation apply to focus groups the same way it applies to interviews? The principle is the same, but the numbers shift. Most studies reach saturation within 4 to 8 focus groups rather than individual interview counts, since each session already brings multiple voices into the room.


References

Glaser and Strauss (grounded theory foundations);

Guest, Bunce and Johnson (2006);

Hennink, Kaiser and Marconi (2017);

Mason (2010); Creswell (2013);

Malterud, Siersma and Guassora (2016);

Fugard and Potts;

Guest, Namey and Chen (2020).

Perfect for students, researchers, and professionals looking to build real statistical skills.