Skip to content

Tensions in Practice: Representing familiar people in tension with testing unfamiliar people

A small historical panel used to plan an evaluation split

People P and Q each have two recorded visits. Reserving each second visit leaves both people represented in the fitting evidence. Reserving both of Q’s visits leaves an entire person unseen during fitting. The first split targets additional visits for familiar people; the second targets a new person. Disjoint rows alone do not make those evaluation questions equivalent.

Assess a familiar-person use

Retain earlier information about each evaluated person.

Assess an unfamiliar-person use

Keep the evaluated identity absent from fitting.

Why these aims pull against each other

Holding out whole people improves identity-level separation but leaves fewer identities for fitting. Splitting within each person preserves that coverage but cannot support an unseen-person claim.

Compare the arrangements

Hold out later visits

Fit using P1 and Q1; evaluate on P2 and Q2 without using their outcomes or derived features during development.

Reserve each second visit
Evidence use
P / visit 1Fit
P / visit 2Evaluate
Q / visit 1Fit
Q / visit 2Evaluate
What it protects
Both evaluated people have earlier records available, matching the familiar-person target.
What it costs
The score cannot establish performance on people wholly absent from fitting.
When it fits
Fits later-visit use with no future-to-past feature leakage and the declared earlier records legitimately available.

Illustration note: The labels designate temporal order within each person. No accuracy values or independence between visits are assumed.

Hold out an entire person

Fit using P1 and P2; keep Q1 and Q2 out of every fitting and model-selection step.

Reserve all of Q
Evidence use
P / visit 1Fit
P / visit 2Fit
Q / visit 1Evaluate
Q / visit 2Evaluate
What it protects
Q is a genuinely unseen identity in this historical evaluation.
What it costs
Only P contributes fitting data, and Q alone offers an extremely narrow assessment.
When it fits
Fits an unfamiliar-person target when more groups will be needed for a credible estimate.

Illustration note: This small table illustrates segregation only; it is not a sufficient evaluation sample or a deployment guarantee.

What this illustration does—and does not—establish

The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.

  • All records are historical. The second split concerns identity generalization, not a simulation of calendar-time deployment.
  • No result is calculated. Real inference must handle within-person dependence and a representative set of people.
  • Both arrangements require whole-cycle segregation; changing a split cannot rescue prior exposure to the reserved outcomes.

Source entries

Holdout Set

Prime · Source of the tension

The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.

Representativeness versus Disjointness (coupling)

Two requirements pull against each other: the holdout must be *disjoint* from fitting yet *representative* of the cases that will matter. Carving out a disjoint slice can leave it unrepresentative.

Read the source section

The source operation

A holdout set is a portion of available evidence deliberately withheld from the process that produces a candidate — a model, a theory, a plan, a policy, a design — so that the withheld portion can be used to score the candidate *without* the candidate having been shaped by it. The structural commitment is the segregation of two evidence streams: one stream is spent building, the other is reserved for evaluating, and the two are kept disjoint by a procedural guarantee that must hold across the candidate's whole development cycle, not merely at the first split.

Read the source section