Skip to content

Tensions in Practice: Population estimation in tension with model development

Studying how a support service is used

One research team wants to estimate how often people encounter a service problem. Another wants to understand the kinds of situations in which it occurs. A population study selects cases through a declared sampling design; a model-building study can choose the next case because earlier analysis exposed a gap. Those case selections can be useful for different goals, but a set deliberately rich in unusual cases cannot simply be counted as if it mirrored the population.

Estimate population frequency

Use a sampling procedure that supports a stated inference about how common the problem is.

Develop an explanation

Seek cases that can challenge or extend an emerging account of the problem.

Why these aims pull against each other

The next case most useful for filling a model’s gap need not be the next case prescribed by a population sampling design. Spending effort for one goal does not automatically produce evidence for the other.

Compare the arrangements

Follow a population design

Define the target population and use a declared probability sampling design to obtain observations for a frequency estimate.

What it protects
The selection mechanism can support population inference when its coverage, response, measurement and analysis requirements are met.
What it costs
It can spend observations on cases that add little to a particular model’s weak point; maintaining the sampling design also takes resources.
When it fits
Fits a population-frequency question with a suitable frame and analysis. This diagram chooses a fixed design, not a claim that every probability design is non-adaptive.

Illustration note: This is the representative-sampling counter-arrangement. The source’s conceptual contrast is narrowed here to a specific fixed design, without claiming that representative samples contain only typical cases.

Choose by the model’s gaps

Analyze a case, identify an unresolved part of the account, and choose another case that can probe that gap. Stop when further cases cease to add structural insight.

What it protects
Selection effort can concentrate on cases likely to change the developing account.
What it costs
Each selection requires analysis. The resulting case mix is purposefully distorted for population counting and may miss an unrecognized gap.
When it fits
Fits concept or explanation development. The selector must target informative challenges, not merely examples that agree with the current model.

Illustration note: The return arrow is the defining analysis-to-selection feedback. The stopping condition concerns insight, not a fixed interview quota or proven completeness.

What this illustration does—and does not—establish

Theoretical Sampling: Model Development versus Population Estimation (Conflicting Objectives) distinguishes estimation from model development; Theoretical Sampling: Closed Loop versus Batch (the Interleaving Requirement) requires analysis between selections. The service context is editorial, and the fixed probability-design counterexample does not rule out more sophisticated mixed research programs.

  • The two arrangements answer different questions. They are not competing accuracy levels for the same estimand.
  • A deliberately selected case can reveal a useful mechanism without establishing how common it is.
  • Neither a probability label nor an insight-based stopping rule alone guarantees good evidence; coverage, response, measurement, analysis and the honesty of case selection remain relevant.

Source entries

Theoretical Sampling

Prime · Source of the tension

This source passage supplies the contextual tension. The concrete arrangements and schematic examples are editorial illustrations, not measured findings.

Model Development versus Population Estimation (Conflicting Objectives)

T1 — Model Development versus Population Estimation (Conflicting Objectives). Theoretical sampling optimizes for developing a model; representative sampling optimizes for estimating a population.

Read the source section

The source operation

Theoretical sampling is the structural pattern in which the next case to study is selected by what it would teach about the emerging theory, not by what it would say about the population. The selection criterion is informativeness for concept development — pick the cases that would maximally challenge, refine, or extend the current model — and the procedure is interleaved with analysis: each new case is chosen in light of what previous cases revealed, what gaps remain, and what boundary conditions are still untested. The catalogue of cases is not fixed in advance; it grows by what the developing theory needs.

Read the source section

Analysis must steer the next selection

T4 — Closed Loop versus Batch (the Interleaving Requirement). The mechanism depends on interleaving analysis with selection, so each case is chosen in light of what prior cases revealed. Collect everything first and analyze second, and the loop is severed — the selection was never steered by the developing model. The failure mode is doing field research or bulk labeling and calling it theoretical sampling: real work is done, but the feedback that makes each selection responsive to the model's state is missing. The diagnostic is to ask whether analysis actually happens *between* selections: if the sample was fixed in advance and never revised mid-collection, the loop is open, and the procedure has forfeited the entire advantage of state-driven selection regardless of how the cases were chosen.

Read the source section