Unit testing¶
Checking a bounded software unit against expected behavior with repeatable test cases.
Core Idea¶
Unit testing checks a bounded software behavior by setting up inputs, invoking the unit, and comparing the observed result with an explicit expectation. A unit can be a function, method, or component behavior. The important boundary is the locally stated claim: the test runner's pass or failure concerns this behavior under these conditions, not the entire application.
Official pytest and JUnit examples show the same relation in Python and Java. One deliberately failing pytest assertion makes the expected-versus-actual comparison visible; a JUnit calculator assertion illustrates a passing local check. Fixtures and controlled dependencies can improve repeatability, but isolation is a design choice and integration behavior still needs its own tests.
Structural Signature¶
Sig role-phrases:
- Bounded software unit — Selects a function, method, component behavior, or other local implementation boundary. It is constitutive. Counterfactual: A full end-to-end workflow has a different test scope even if it contains this unit.
- Test input and setup — Supplies controlled state and selected inputs or collaborator behavior. It is constitutive. Counterfactual: An outcome with unknown setup is not a repeatable local test case.
- Expected-behavior oracle — States a value, exception, invariant, or tolerance against which actual behavior is checked. It is constitutive. Counterfactual: Executing a function without an expectation is a demonstration, not a unit test verdict.
- Unit invocation — Runs the selected code under the setup and observes its output or effect. It is constitutive. Counterfactual: Static reading of source alone is not a test execution.
- Localized verdict — Reports pass or failure for the case while preserving failure evidence. It is constitutive. Counterfactual: A passing case must not be promoted into a proof of all behavior.
- Dependency boundary — Specifies which collaborators are real, controlled, or outside the local scope. It is central. Counterfactual: An unbounded integration chain weakens the local causal interpretation.
What It Is Not¶
- Not integration testing. A chain of connected components has a broader failure boundary.
- Not a proof of correctness. Passing examples do not cover all possible inputs or deployments.
- Not code review. Static inspection does not invoke the unit against an oracle.
- Not test-driven development itself. Writing tests first is one workflow, not a defining condition for a unit test.
- Closest near-miss. A unit may call real collaborators; isolation is a matter of boundary design rather than an absolute claim of no dependencies.
Scope of Application¶
- Developer feedback. Catch local behavior regressions quickly.
- Refactoring. Check preserved unit contracts after implementation changes.
- API behavior design. Make value, exception, and tolerance expectations explicit.
- Defect localization. Narrow a failure to one bounded behavior before wider integration diagnosis.
Clarity¶
Choose a small software behavior, provide an input, say what should happen, run it, and compare actual with expected. A failed assertion is evidence for that case. A passing assertion is not a claim that the surrounding system works.
Manages Complexity¶
Software behavior has vast possible input and dependency combinations. Unit tests reduce the immediate search space to explicit cases and controlled setup, making failures legible. That compression is useful only if the boundary and its omissions are recorded, rather than mistaken for whole-system coverage.
Abstract Reasoning¶
- Identify a bounded behavior and its public expectation.
- Choose representative and edge inputs with controlled setup.
- State a value, exception, invariant, or tolerance oracle.
- Invoke the unit and collect actual behavior.
- Compare actual and expected with useful failure evidence.
- Revise the case or implementation, then test integrations separately.
Knowledge Transfer¶
The test-case pattern transfers among languages and frameworks because local invocation and criterion comparison are stable roles. An experiment on a physical component is structurally similar but is not software unit testing without a software unit and executable test harness.
Examples¶
Canonical¶
pytest's own getting-started example supplies a defining minimal test: func(x)=x+1 is called with 3, but assert func(3)==5 fails because the actual value is 4. The deliberately failing tutorial maps a bounded unit, input, oracle, invocation, and local verdict; it is a worked construction, not a production defect report.
Mapped back: Bounded software unit → func(x)=x+1; Test input and setup → input 3 in a simple test module; Expected-behavior oracle → assert returned value equals 5; Unit invocation → pytest calls func(3); Localized verdict → failed assertion, actual 4 versus expected 5; Dependency boundary → pure function without external collaborators in this demonstration.
Applied / In Practice¶
The pytest project's own test_mark_closest checks its marker-lookup behavior in a generated class: it creates class and method markers, collects the methods, asserts the method's closest marker has location 'function', and asserts a missing marker returns None. This is a real repository regression test of bounded project behavior, not a claim that all marker interactions are correct.
Mapped back: Bounded software unit → pytest get_closest_marker behavior; Test input and setup → generated class with class-level and method-level markers; Expected-behavior oracle → method marker location is function; missing marker returns None; Unit invocation → project test collects methods then invokes get_closest_marker; Localized verdict → assertions for returned marker and None case; Dependency boundary → local generated test module and pytester fixture, not a deployed application.
Structural Tensions¶
T1 — Unit Isolation versus Realistic Collaboration. A tightly controlled unit makes failures local but can hide interface mistakes; real collaborators add fidelity but blur the unit boundary.
Diagnostic: Which dependencies must be real for this claim?
T2 — Fast Narrow Checks versus Behavioral Coverage. Many cheap local assertions aid feedback, yet none alone covers all branches or system interactions.
Diagnostic: Which behavior is actually sampled?
T3 — Strict Equality versus Tolerance Oracle. Exact expected values reveal mismatches but can incorrectly reject legitimate floating-point results; tolerances admit variation at the cost of weaker discrimination.
Diagnostic: What output uncertainty is legitimate?
Structural–Framed Character¶
Unit testing is mixed-framed: local comparison with an expected result is structurally clear, while the unit boundary and oracle are chosen in software practice. Evaluative weight: a pass reports agreement with specified expectations in that setup; it does not certify the whole application, correctness of the oracle, or absence of defects. Human-practice-bound: without developers selecting a software unit, invocation, and assertion, there is no unit test, even if the program executes. Institutional origin: frameworks and team conventions affect isolation and naming, but no one tool is constitutive of all unit testing. Vocabulary travels: evaluation against a criterion occurs in many domains, whereas executable code, a local harness, and a repeatable software result make this literal. Import versus recognize: another language's test suite checking bounded functions is still unit testing; a laboratory trial of one physical component is an analogy.
The portable skeleton is the live parent prime Evaluation: apply a criterion-bearing frame to a bounded object and report a verdict. Unit testing narrows the object to software behavior and the criterion to an executable expected-behavior oracle. Its character: a local, repeatable software evaluation whose evidence remains limited by its unit boundary and test design.
Structural Core vs. Domain Accent¶
Skeletal core. A bounded object is exercised under controlled conditions and evaluated against a criterion. Domain-bound accent. The object is executable software and the criterion is an assertable behavior oracle in a repeatable test harness. Transfer boundary. A general evaluation or full-system trial lacks this local software test boundary.
Instantiates / Related Primes¶
This entry is a kind of Evaluation.
-
Strict parent: prime Evaluation. Every unit test applies an explicit criterion to observed behavior and produces a pass/fail judgment; software-local execution narrows, rather than replaces, the evaluation genus.
-
Neighbor. Integration testing checks interactions among units and can discover failures a local test misses.
Relationships to Other Abstractions¶
Current abstraction Unit testing Domain-specific
Parents (1) — more general patterns this builds on
-
Unit testing is a kind of Evaluation Prime
Unit testing is a strict kind of Evaluation: Checking a bounded software unit against expected behavior with repeatable test cases.The current Evaluation prime applies a criterion-bearing frame to a bounded object and yields a verdict or score. Each unit test applies an expected-behavior oracle to a bounded software unit and reports a local pass/fail verdict, satisfying child→parent subsumption while adding execution and software-specific restrictions.
Hierarchy path (1) — routes to 1 parentless root
- Unit testing → Evaluation → Comparison → Self Checking
Neighborhood in Abstraction Space¶
Unit testing sits in a crowded region of the domain-specific corpus (38th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Generic System & Interface Definitions (27 abstractions)
Nearest neighbors
- Service-Oriented Programming — 0.90
- Interface transparency (computing) — 0.87
- Time-sharing — 0.87
- Content Analysis — 0.87
- Engineered System — 0.87
Computed from structural-signature embeddings · 2026-10-08
Not to Be Confused With¶
- Integration testing. Tell: Exercises collaboration across unit boundaries.
- Static analysis. Tell: Inspects code without necessarily invoking the behavior.
- Test-driven development. Tell: A sequencing practice that may use unit tests.
- Formal verification. Tell: Uses proof obligations stronger than a finite set of test executions.
References¶
- pytest, Get Started — deliberately failing
func(3)test, exception oracle, approximate comparison, and reported verdict. - pytest, How to use fixtures — controlled setup and dependency behavior.
- JUnit 5 User Guide, Writing Tests — independent calculator-addition demonstration.
- Frozen Wikipedia discovery revision — candidate provenance, not authority for the repaired scope.
The official framework examples show the method, not measured production effectiveness. The local-test verdict remains bounded to the cases and setup actually exercised.
- pytest project, testing/test_mark.py, test_mark_closest — real project regression test of local marker lookup; not a measure of all production defects.