Tensions in Practice: Representing familiar people in tension with testing unfamiliar people¶
A small historical panel used to plan an evaluation split
People P and Q each have two recorded visits. Reserving each second visit leaves both people represented in the fitting evidence. Reserving both of Q’s visits leaves an entire person unseen during fitting. The first split targets additional visits for familiar people; the second targets a new person. Disjoint rows alone do not make those evaluation questions equivalent.
Assess a familiar-person use
Retain earlier information about each evaluated person.
Assess an unfamiliar-person use
Keep the evaluated identity absent from fitting.
Why these aims pull against each other
Holding out whole people improves identity-level separation but leaves fewer identities for fitting. Splitting within each person preserves that coverage but cannot support an unseen-person claim.
Choose an arrangement to see what changes and what remains difficult.
Each record stays in the same row. The changed evidence-use labels reveal whether the evaluated person appeared in fitting.
What this choice protects
What it costs
When it fits
Compare the arrangements
Hold out later visits
Fit using P1 and Q1; evaluate on P2 and Q2 without using their outcomes or derived features during development.
| Evidence use | |
|---|---|
| P / visit 1 | Fit |
| P / visit 2 | Evaluate |
| Q / visit 1 | Fit |
| Q / visit 2 | Evaluate |
- What it protects
- Both evaluated people have earlier records available, matching the familiar-person target.
- What it costs
- The score cannot establish performance on people wholly absent from fitting.
- When it fits
- Fits later-visit use with no future-to-past feature leakage and the declared earlier records legitimately available.
Illustration note: The labels designate temporal order within each person. No accuracy values or independence between visits are assumed.
Hold out an entire person
Fit using P1 and P2; keep Q1 and Q2 out of every fitting and model-selection step.
| Evidence use | |
|---|---|
| P / visit 1 | Fit |
| P / visit 2 | Fit |
| Q / visit 1 | Evaluate |
| Q / visit 2 | Evaluate |
- What it protects
- Q is a genuinely unseen identity in this historical evaluation.
- What it costs
- Only P contributes fitting data, and Q alone offers an extremely narrow assessment.
- When it fits
- Fits an unfamiliar-person target when more groups will be needed for a credible estimate.
Illustration note: This small table illustrates segregation only; it is not a sufficient evaluation sample or a deployment guarantee.
What this illustration does—and does not—establish
The source supplies the structural tension; the invented example makes one relation inspectable. Costs and conditions are part of each arrangement, not exceptions to a universal recommendation.
- All records are historical. The second split concerns identity generalization, not a simulation of calendar-time deployment.
- No result is calculated. Real inference must handle within-person dependence and a representative set of people.
- Both arrangements require whole-cycle segregation; changing a split cannot rescue prior exposure to the reserved outcomes.
Source entries
Holdout Set
The canonical tension motivates this comparison. The setting, finite values and arrangements are declared editorial illustrations, not measured findings.
Representativeness versus Disjointness (coupling)
Two requirements pull against each other: the holdout must be *disjoint* from fitting yet *representative* of the cases that will matter. Carving out a disjoint slice can leave it unrepresentative.
The source operation
A holdout set is a portion of available evidence deliberately withheld from the process that produces a candidate — a model, a theory, a plan, a policy, a design — so that the withheld portion can be used to score the candidate *without* the candidate having been shaped by it. The structural commitment is the segregation of two evidence streams: one stream is spent building, the other is reserved for evaluating, and the two are kept disjoint by a procedural guarantee that must hold across the candidate's whole development cycle, not merely at the first split.