Feature Factory¶
The product-organisation anti-pattern of measuring success by features shipped per unit time while never asking whether any feature changed an outcome — a proxy output displacing the target it was meant to track, Goodhart's Law in the product operating loop.
Core Idea¶
A feature factory is the product-organisation anti-pattern in which the organisation's success is measured by the number of features shipped per unit time, the operating rhythm is calibrated to keep that throughput high, and the question of whether any shipped feature actually changed user behaviour, revenue, retention, or whatever outcome the feature was nominally for goes unasked, unmeasured, or is asked only ceremonially after the fact. The team's output metric — features delivered, story points completed, sprints closed — substitutes for the outcome metric, and the organisation optimises on the substitute: the product accretes shipped features, the roadmap stays full, the team looks productive on the charts the executives see, and the underlying business or user metrics drift independently of what is being built. The structural mechanism is a four-stage substitution: a target outcome (retention, revenue, mission impact) motivates the product organisation's existence; a proxy output (features shipped, velocity) is measurable, controllable, and locally generated; the operating loop begins rewarding and reporting on the proxy as though it were the target; and an output-outcome decoupling emerges because nothing in that loop punishes shipping features that move no outcome metric — so gaming follows (smaller features, more sprints, more launches) and the proxy and target drift apart. The term was codified by John Cutler in the product-discovery and outcomes-versus-outputs literature; the broader structural pattern — a proxy measure displacing the target it was meant to track — is Goodhart's Law applied to the product operating loop.
Structural Signature¶
Sig role-phrases:
- the target outcome — the user, business, or mission metric (retention, revenue, impact) that motivates the product organisation's existence
- the proxy output — features shipped, story points, sprints closed: measurable, controllable, locally generated, and structurally an upstream input to the outcome rather than the outcome itself
- the management substitution — the operating loop (planning cadence, executive review, performance evaluation) closing on and rewarding the proxy as though it were the target
- the output-outcome decoupling — because nothing in the loop punishes shipping features that move no outcome metric, the proxy and target drift progressively apart
- the gaming response — local optimisation on the proxy: features split smaller, sprints multiplied, launches counted, raising the output number without moving the target
- the answer-rate diagnostic — the health signal: the organisation's rate of supplying, per shipped feature, a pre-launch outcome hypothesis and a post-launch measurement (high in a healthy loop, near-zero in a factory)
- the late, widening surfacing — a rising output line beside a slowly deteriorating outcome line over quarters; the gap discovered only once the outcome has drifted far enough to alarm leadership
- the forcing-function rewire — pair every committed output with an outcome hypothesis and kill criterion before commit, and re-point the review, roadmap, and evaluation onto outcome metrics; throughput is predicted to fall when the fix works
What It Is Not¶
- Not a talent or execution problem. The team may be correctly optimising the metric placed in front of it; the defect is in which metric the operating loop is closed on, not in the people closing it. Reading flat business numbers as a need to "ship better or faster" — or to replace the team — misdiagnoses a metric-substitution failure as an effort failure, and pushing harder on output only feeds the factory.
- Not high throughput being evidence of health. Features shipped, story points, and sprints closed are outputs — structurally an upstream input to retention or revenue, not the outcome itself — so a rising output line carries no information about whether the target moved. The health signal is the organisation's answer-rate to "for each shipped feature, what outcome was hypothesised, and was it measured?", which a healthy loop supplies for most work and a factory for almost none.
- Not every high-output organisation. It is a feature factory only when an outcome metric exists, matters, and is being abdicated in favour of the output proxy. Genuinely exploratory work that defers the outcome link deliberately (research, platform investment), or execution where outcomes are tracked and merely modest, is not the pattern however high the throughput. The test is whether anything in the loop punishes shipping a feature that moves no outcome.
- Not fixed by shipping more or better features. Because nothing in the loop punishes outcome-less shipping, adding output cannot repair it; the corrective is a forcing function on the proxy-target link — pair every committed output with an outcome hypothesis and a kill criterion before commit, and re-point the executive review, roadmap, and evaluation onto outcome metrics. Improving the artifact leaves a loop closed on the wrong number still closed on it.
- Not a pattern where falling output is a regression. When the fix works, feature throughput is predicted to drop — the factory was manufacturing motion — while the outcome metrics finally begin to move. Reading that decline as lost productivity inverts the diagnosis; the drop is the signature of the loop being rewired onto outcomes, not a cost of the cure.
Scope of Application¶
Feature factory lives within product organisations; its reach is within that domain, across every surface form the proxy-for-outcome substitution takes in software-product and digital-delivery work. It is the product operating loop's instance of Goodhart's Law — the genuinely distant cousins (test-score gaming, hospital waiting-time gaming, citation counting, command-economy output tonnage) belong to that target-measure-substitution parent, not to "feature factory" by name.
- Roadmap-by-feature-list planning — a planning artefact that is a quarter-grouped list of features carrying no outcome hypotheses, success metrics, or kill criteria.
- Velocity / story-points-as-headline-metric engineering cultures — team health measured by story throughput, with outcomes occasionally measured but never gating the next sprint.
- Launch-then-move-on practice — features shipped, announced, and left behind as attention rotates to the next, with no instrumented validation of impact.
- Sales-led roadmaps — every feature promised to a deal, with no one revisiting whether the feature is used or whether the deal would have closed without it.
- Innovation theatre in regulated digital-transformation programmes — feature counts substituting for movement in patient, learning, or service-quality outcomes.
- After-the-fact success-bar inflation — when outcomes are measured at all, the bar reset to "it launched" or "no rollback was needed" rather than any prior impact hypothesis.
Clarity¶
Naming the feature factory reframes a familiar leadership complaint — "we are shipping a lot but the business numbers aren't moving" — from an execution story into a metric-substitution diagnosis. Without the label, the reflex is to read flat outcomes as a talent or effort problem and to push the team to ship better or faster; the concept makes legible that the team may be correctly optimising the metric placed in front of it, and the defect is in which metric the operating loop is closed on, not in the people closing it. That separates two questions an executive review routinely fuses: is the team productive? and is the work moving the outcome it exists to move? — and points the correction away from replacing the team and toward rewiring the metric the planning cadence, the executive review, and performance evaluation actually report against.
It also sharpens the output / outcome distinction that throughput charts silently erase. Features shipped, story points, sprints closed are outputs — measurable, controllable, locally generated — and the label makes legible that they are structurally an upstream input to retention or revenue, not the outcome itself, so a rising output line carries no information about whether the target moved. The sharper question the concept licenses is diagnostic and answerable: for each shipped feature, what was the hypothesised outcome, and was it measured? An organisation's answer-rate to that question is itself the diagnostic — a healthy product loop can answer it for most of its work, a feature factory for almost none. Locating the breakdown at the proxy-target link (rather than in shipping skill) is what tells the leader the remedy is a forcing function on that link — pair every committed output with an outcome hypothesis and a kill criterion before commit — since shipping more features cannot repair a loop whose problem is that nothing in it punishes shipping features that move no outcome.
Manages Complexity¶
Product organisations present a scatter of seemingly distinct dysfunctions — a sales-led roadmap where every feature is promised to a deal, a velocity-obsessed engineering culture chasing story points, a digital-transformation programme counting features instead of patient or learning outcomes, success bars reset after the fact to "it launched" — each of which a leader might try to fix on its own terms. Feature factory collapses that scatter into one mechanism: the operating loop is closed on a proxy output rather than the target outcome it exists to move. With that compression the diagnosis reduces to a single measurable quantity — the organisation's answer-rate to "for each shipped feature, what outcome was hypothesised, and was it measured?" — which a healthy loop can supply for most of its work and a feature factory for almost none, regardless of which surface pathology is showing. And because all the variants share the proxy-target structure, one intervention family fixes them: pair every committed output with an outcome hypothesis and kill criterion before commit. The leader tracks where output metrics gate decisions that are nominally about outcomes and reads the health of the loop off that one ratio, instead of diagnosing sales pressure, engineering rhythm, and transformation theatre as separate problems.
Abstract Reasoning¶
Feature factory licenses inference moves that all turn on the proxy-target decoupling the compression isolates: an output metric closed on in place of the outcome it was meant to move.
Diagnostic — infer a closed-on-proxy loop from the answer-rate, not from the throughput. The characteristic move is to ignore the rising output line (features shipped, velocity) entirely as evidence of health — because by construction it carries no information about the outcome — and instead measure the organisation's answer-rate to a single question asked of each shipped feature: what outcome was hypothesised, and was it measured? From "we shipped sixty features and the business metrics are flat," the move is to infer not that the team ships badly but that the operating loop is closed on the proxy, and to confirm it by finding that the team can supply a pre-launch outcome hypothesis for only a handful of the sixty and a post-launch measurement for fewer still. The diagnostic signature that distinguishes a feature factory from ordinary under-performance is precisely this gap between output legibility and outcome legibility: outputs are exhaustively tracked, outcomes barely at all. A second diagnostic reads the gaming tells — features split smaller, sprints multiplied, launches counted — as evidence that local optimisation is running on the proxy, since those moves raise the output number without necessarily moving any target.
Interventionist — rewire which metric the loop closes on; do not exhort or replace the team. Because the team is correctly optimising the metric placed in front of it, the licensed intervention acts on the operating-metric graph, not on effort or talent. The master move is a forcing function on the proxy-target link: pair every committed output with an outcome hypothesis and a kill criterion before commit, so nothing enters the build queue without a named outcome it is accountable to. Around that, the move rewires the three places the proxy currently displaces the target — change the headline number in the executive review from features-shipped to movement in named outcome metrics, recast the roadmap from a feature list into a set of outcome bets, and tie performance evaluation to outcome movement rather than throughput. Each is a prediction: close the loop on outcomes and the loop will begin punishing features that move nothing, so observed feature throughput should fall (the factory was manufacturing motion) while the outcome metrics finally begin to move. The predicted drop in output is not a cost but the signature of the fix working — a counterintuitive inference the concept makes available.
Boundary-drawing — is this a feature factory, or legitimately output-driven work? The concept draws the line at whether the outcome is knowable and being abdicated. It is a feature factory only when an outcome metric exists, matters, and the loop has nonetheless closed on the output proxy instead — distinguishing it from genuinely exploratory work where the outcome link is deliberately deferred (research, platform investment) and from execution where outcomes are tracked and merely modest. The move is to ask does nothing in the loop punish shipping a feature that moves no outcome? — if outcomes are measured and do gate the next commitment, it is not a feature factory however high the throughput; if they are not, it is, however sound the individual features. It also bounds against mis-attribution: a flat outcome with a loop that does close on outcomes is a hard-problem or market story, not metric substitution, and reaching for the feature-factory diagnosis there is itself an error.
Order-of-events / predictive — the decoupling widens over time and surfaces late. The concept predicts a temporal drift: once the loop closes on the proxy, the proxy and target diverge progressively, because nothing in the loop corrects features that fail to move the outcome, so each cycle adds shipped output while the target metric wanders independently. The move reads the shape of the divergence as confirmation — a steadily rising output line beside a slowly deteriorating retention or revenue line over many quarters is the order-of-events fingerprint of a loop closed on the wrong metric — and reads it forward as a forecast: absent a forcing function on the proxy-target link, the next cycle will reproduce the same accretion of motion without outcome, and the gap will be discovered only when the outcome metric has drifted far enough to alarm leadership, long after the loop went wrong.
Knowledge Transfer¶
Within product organisations the diagnosis transfers as mechanism across every surface form the pathology takes, because the underlying structure — an operating loop closed on a proxy output in place of the target outcome — is the same regardless of which dysfunction is showing. The diagnostics (ignore the throughput line; measure the answer-rate to "for each shipped feature, what outcome was hypothesised, and was it measured?"; read the gaming tells of features split smaller and sprints multiplied), the single intervention family (pair every committed output with an outcome hypothesis and kill criterion before commit; rewire the executive review, the roadmap, and performance evaluation onto outcome metrics), and the counterintuitive prediction (feature throughput should fall when the fix works) carry intact across roadmap-by-feature-list planning, velocity/story-points-as-headline-metric engineering cultures, launch-then-move-on practice, sales-led roadmaps where every feature is promised to a deal, innovation theatre in regulated digital-transformation programmes, and after-the-fact success-bar inflation ("it launched" as the bar). These are not separate problems; they are one mechanism with the surface pathology swapped, so the answer-rate diagnostic proven against a velocity-obsessed team applies unchanged to a sales-led roadmap. The supporting apparatus travels with it: the outcomes-versus-outputs writing (Cutler, Reinertsen, Cagan, Torres, Perri), the Lean/TPS output-versus-outcome distinction, and opportunity-tree experimentation practice all operate across these contexts.
The reach of the named concept stops at the edge of product organisations, and honesty requires marking why its apparent breadth is one substrate replayed. "Feature factory" is Cutler's product-discovery idiom; it is not used as a structural primitive outside software product organisations, and the candidate's four-domain transfer sketch (public innovation, product development, education reform, health-tech) sits inside the product-organisation substrate at the level of analogy — each is still an organisation shipping artefacts under a roadmap with output metrics substituting for outcome ones. The transfer between those happens because they share that substrate, not because "feature factory" reaches into a foreign domain.
What genuinely travels to distinct substrates is the pattern the feature factory explicitly instantiates, and that — not the named concept — is what should carry any cross-domain lesson (case B). The feature factory is Goodhart's Law applied to the product operating loop: once a proxy measure becomes the target, it ceases to be a good measure of the outcome it was meant to track, and local optimisation games the proxy as it drifts from the target. That target_measure_substitution / Goodhart pattern recurs as genuine co-instances across domains that share no product-management machinery — school-test-score gaming, NHS and hospital waiting-time gaming, academic citation-counting, police arrest quotas, factory output-tonnage in command economies (the Soviet nail-factory archetype), social-media engagement optimisation, ML training-loss diverging from real capability, financial capital-ratio gaming — and across all of them the same four-stage structure (a target, a measurable proxy, a management substitution, an emergent decoupling with gaming) holds. The cross-domain insight therefore belongs to the Goodhart / proxy-target-drift parent, not to "feature factory," whose home-bound cargo is the product-management vocabulary (features, story points, sprints, roadmaps, the outcomes-vs-outputs codification) and the Cutler name. The honest report is: across product organisations' variants the diagnosis transfers as mechanism with only the surface form changed; for genuinely distant systems, carry the general Goodhart's-Law / target-measure-substitution pattern, while the feature-factory framing and name stay home as the domain accent. (See Structural Core vs. Domain Accent.)
Examples¶
Canonical¶
The defining articulation is John Cutler's 2016 essay "12 Signs You're Working in a Feature Factory," which crystallised the anti-pattern from patterns recurring across product teams. Picture the archetype it describes: a product team whose roadmap is a quarter-by-quarter list of features, whose stand-ups and reviews report velocity and story points, and whose definition of "done" is shipped-and-announced. Features go out continuously — a redesigned onboarding flow, a new settings panel, a dozen more each quarter — the throughput charts the executives see stay full and green, and everyone looks busy. Yet no one can say, for any given feature, what user or business outcome it was supposed to move, and almost nothing is instrumented after launch to check. Retention and revenue drift on their own, uncorrelated with the shipping pace.
Mapped back: Retention and revenue are the target outcome; the features, story points, and sprints are the proxy output. Reporting velocity as the headline in reviews is the management substitution, and the inability to state a per-feature hypothesis is a near-zero answer-rate diagnostic — the fingerprint of the output-outcome decoupling.
Applied / In Practice¶
The pattern is pervasive in large enterprise and public-sector "digital transformation" programmes. A representative case: an agency modernising a citizen-facing service commits to a multi-year roadmap and reports progress to oversight boards as the count of modules and features delivered against schedule. Delivery dashboards show steady feature completion — new portals, forms, integrations — and the programme is declared on track. But the outcomes that justified the spend (application processing time, error rates, citizen satisfaction, cost-to-serve) are measured late, rarely, or only after the bar has quietly been reset to "it launched." Reviews reward the shipped-features number; nothing in the loop kills a module that fails to improve service.
Mapped back: Processing time and citizen satisfaction are the target outcome; modules and features delivered are the proxy output. Reporting delivery counts to oversight is the management substitution, and the drift between a full delivery dashboard and untracked service outcomes over years is the late, widening surfacing — corrected only by a forcing-function rewire pairing each module with a named outcome and kill criterion.
Structural Tensions¶
T1: Proxy's controllability versus its emptiness (the very properties that make features-shipped a good management metric are what make it a bad outcome measure). The output proxy is chosen for real virtues: features shipped, story points, and sprints closed are measurable, controllable, and locally generated — a team can move them by its own effort, on a predictable cadence, and report them cleanly. Those are exactly the properties a manager wants in a metric. But they are also precisely why the proxy carries no information about the outcome: retention and revenue are none of measurable-by-the-team, controllable, or locally generated, so the metric that is easy to steer is structurally upstream of and decoupled from the target it stands in for. The feature factory is not a lapse of rigor; it is the predictable result of the operating loop preferring the metric it can actually close on. The convenience and the emptiness are the same set of properties viewed from two ends. Diagnostic: Is the headline metric being used because it is controllable and legible, or because it actually carries information about whether the target outcome moved?
T2: Falling throughput as cure versus as regression (the fix's success signature is indistinguishable from failure on the old dashboard). The concept makes a counterintuitive prediction: when the loop is rewired onto outcomes, feature throughput should fall, because the factory was manufacturing motion and the loop now kills features that move nothing. But this is precisely the signal that every incumbent instrument reads as a problem — a declining velocity line, fewer launches, a thinner roadmap look identical to a team losing productivity. The remedy therefore produces, on the metrics leadership has been trained to watch, the exact picture of things getting worse, and only the outcome line (slower, laggier, less legible) shows the cure working. An organisation must tolerate an apparent regression on its most visible number to realise the fix, which is exactly the tolerance a feature factory lacks by construction. Diagnostic: When output drops after the rewire, is the outcome line being watched to confirm the loop is now killing motion-without-outcome, or is the falling throughput being read as lost productivity and reversed?
T3: Outcome accountability as forcing function versus as a brake on genuine exploration (the same demand that fixes the factory can strangle work whose outcome is properly deferred). The master intervention is to pair every committed output with an outcome hypothesis and kill criterion before commit, so nothing enters the queue without a named outcome it answers to. This is the correct cure where an outcome is knowable and being abdicated. But the concept itself draws a boundary: genuinely exploratory work — research, platform investment — deliberately defers the outcome link, and demanding a measured outcome hypothesis for every commitment there would penalise exactly the work whose value cannot yet be pinned to a metric. The forcing function that disciplines a feature factory can, applied indiscriminately, become its own pathology: outcome-theater that suppresses long-horizon bets to satisfy a pre-commit ritual. The remedy's power and its overreach share one mechanism — mandatory outcome accountability. Diagnostic: Is the missing outcome hypothesis a knowable target being abdicated, or work whose outcome is legitimately deferred — and would forcing a hypothesis here discipline the loop or strangle exploration?
T4: Answer-rate as clean diagnostic versus its gameability (the moment the answer-rate becomes the metric, it is itself subject to the substitution it detects). The concept's sharpest diagnostic is the organisation's answer-rate to "for each shipped feature, what outcome was hypothesised, and was it measured?" — near-zero in a factory, high in a healthy loop. But this diagnostic is a metric, and the concept's own parent (Goodhart) predicts that any metric made a target gets gamed. Once leadership rewards a high answer-rate, teams can manufacture perfunctory outcome hypotheses and after-the-fact measurements — the ceremonial "asked only after the fact" behavior the definition already names — restoring the appearance of a closed outcome loop while the substitution persists one level up. The tool that detects proxy-target drift is not immune to becoming a proxy itself. The diagnostic works only while it is used to understand rather than to score, which is exactly the discipline the pathology erodes. Diagnostic: Are the outcome hypotheses and measurements genuine bets that gate the next commitment, or ceremonial artifacts produced to raise a now-rewarded answer-rate?
T5: Autonomy versus reduction (its own product-org anti-pattern or the product-loop instance of Goodhart's Law / target-measure substitution). "Feature factory" is Cutler's named product-discovery idiom, with proprietary cargo — features, story points, sprints, roadmaps, the outcomes-versus-outputs codification — that transfers as mechanism across every surface form the pathology takes inside product organisations. Yet the entry is explicit that the concept is Goodhart's Law applied to the product operating loop, and that the pattern genuinely recurs in test-score gaming, hospital waiting-time gaming, citation counting, and command-economy output tonnage — none of which share any product-management machinery. The tension is between a named anti-pattern that anchors its own product literature and the recognition that its portable content belongs to the target_measure_substitution / Goodhart parent. Diagnostic: Resolve toward Goodhart / target-measure-substitution when carrying the lesson to non-product systems; toward "feature factory" and its forcing-function rewire when diagnosing an actual product organisation's operating loop.
Structural–Framed Character¶
Feature factory sits at mixed on the structural–framed spectrum, leaning framed — the same profile as its batch-mate feature creep: a construct constituted by a human management practice, but built on a mechanism that recurs across domains as genuine co-instances rather than analogy. The criteria pull both ways. On human-practice-bound it is decisively framed: the anti-pattern presupposes a product organisation with an operating loop — a planning cadence, an executive review, performance evaluation — a designed human management institution, and it dissolves the instant that practice is removed; there is no feature factory without a loop closing on a metric, only artifacts being produced. On institutional origin it is likewise framed: the pattern is codified by Cutler within product-management culture and is an artifact of how organizations report, reward, and evaluate work, not a fact of nature obtaining observer-free. On evaluative weight it is mixed on the same seam feature creep is: "anti-pattern," "factory," and "gaming" carry a pathology charge and the outcome is a failure, yet the diagnosis is pointedly not a verdict on the people — "not a talent or execution problem," the team is correctly optimizing the metric placed in front of it — so the evaluative content is a diagnosis of a metric-substitution failure mode rather than a conviction of a reasoner. On vocab-travels it is framed: the distinctive vocabulary (features, story points, sprints, roadmaps, the answer-rate diagnostic, the forcing-function rewire) is product-management furniture that must be renamed off-substrate.
What pulls it firmly off the framed pole, into mixed, is import-vs-recognize, where it is an especially strong case-(B) instance. The entry states outright that the feature factory is Goodhart's Law applied to the product operating loop, and the underlying mechanism recurs as full co-instances across domains that share no product machinery whatever: school test-score gaming, hospital waiting-time gaming, academic citation-counting, police arrest quotas, command-economy output tonnage, ML training-loss diverging from capability, financial capital-ratio gaming. Each carries the identical four-stage structure — target, measurable proxy, management substitution, emergent decoupling with gaming — so the reuse is recognition of the same mechanism, not metaphor borrowing product management's shape.
The portable structural skeleton is that parent: Goodhart's Law / target_measure_substitution — once a proxy measure becomes the target it ceases to track the outcome it was meant to measure, and local optimization games the proxy as it drifts from the target. That skeleton is genuinely and widely substrate-portable, and it is exactly what the feature factory instantiates from its umbrella, not what makes "feature factory" itself travel: the cross-domain reach belongs to Goodhart, while the outcomes-versus-outputs codification, the story-point/sprint vocabulary, and the forcing-function rewire stay home; even the concept's apparent four-domain breadth (public innovation, education reform, health-tech) is one product-organisation substrate replayed, not the name reaching into foreign ground. Its character: a management-practice-constituted, mildly pathology-charged product anti-pattern whose distinctive vocabulary is home-bound, but which is pulled well into mixed territory because the proxy-displaces-target mechanism it instantiates is a robust, widely recognized cross-domain structure (Goodhart) — structural in that borrowed skeleton, framed in its product-operating-loop specifics.
Structural Core vs. Domain Accent¶
This section decides why the feature factory is a domain-specific abstraction and not a prime — and it is an especially clear case, because the entry states outright that it is Goodhart's Law applied to the product operating loop.
What is skeletal (could lift toward a cross-domain prime). Strip the product management and a robust relational structure survives: a measurable proxy is substituted for the target outcome it was meant to track; once the operating loop closes on and rewards the proxy, the proxy ceases to be a good measure of the target, and local optimization games the proxy as it drifts progressively away from the outcome. The portable pieces are abstract — a target outcome, a measurable-and-controllable proxy that is structurally upstream of it, a management substitution that rewards the proxy as though it were the target, and an emergent decoupling with gaming that widens over time. That skeleton is genuinely and widely substrate-portable, recurring as full co-instances across domains that share no product machinery whatever: school test-score gaming, hospital waiting-time gaming, academic citation-counting, police arrest quotas, command-economy output tonnage, ML training-loss diverging from capability, financial capital-ratio gaming. Precisely because it recurs as mechanism, it is carried by the parent the feature factory instantiates — target_measure_substitution / Goodhart's Law. That parent is the core the feature factory shares, not what makes it distinctive.
What is domain-bound. What makes this specifically a feature factory is product-organization furniture and none of it survives extraction. It presupposes a product operating loop — a planning cadence, an executive review, performance evaluation — a designed human management institution. Its worked vocabulary is discipline-internal: features shipped, story points, sprints closed, roadmaps, velocity, the outcomes-versus-outputs codification (Cutler, Reinertsen, Cagan, Torres, Perri), the answer-rate diagnostic ("for each shipped feature, what outcome was hypothesised, and was it measured?"), and the forcing-function rewire (pair every committed output with an outcome hypothesis and kill criterion before commit; re-point the review, roadmap, and evaluation onto outcomes). The empirical cases (Cutler's "12 Signs," the digital-transformation delivery dashboard) are drawn from it. The decisive test: the concept's own apparent breadth (public innovation, education reform, health-tech) is one product-organization substrate replayed, not the name reaching foreign ground — and for a genuinely distant system like a command economy's output tonnage, calling it "a feature factory" imports story-point-and-roadmap cargo that has no referent there, even though the Goodhart mechanism is fully present. The product-loop vocabulary is the accent that stays home.
Why this does not clear the prime bar. A prime is a relational structure whose vocabulary travels and whose cross-domain transfer is recognition of the same mechanism, not analogy. The feature factory's transfer is bimodal, and the second mode is strong. Within product organizations it moves intact as mechanism — the answer-rate diagnostic, the single intervention family, and the counterintuitive prediction (throughput should fall when the fix works) carry without translation across roadmap-by-feature-list planning, velocity-obsessed engineering cultures, launch-then-move-on practice, sales-led roadmaps, transformation innovation theatre, and after-the-fact success-bar inflation, because each is the same operating loop with the surface pathology swapped. Beyond product organizations the mechanism still recurs — but as co-instances of the parent Goodhart pattern, which each domain exhibits in its own terms (test scores, waiting times, citations), not by importing "feature factory." So the cross-domain reach belongs to target_measure_substitution / Goodhart's Law, whose baggage-free form is what actually generalizes. When the bare structural lesson is needed off the product substrate, it is already carried, in more general form, by that parent; "feature factory," as named, carries product-management baggage that should stay home.
Relationships to Other Abstractions¶
Current abstraction Feature Factory Domain-specific
Parents (2) — more general patterns this builds on
-
Feature Factory is a kind of, typical Build Trap Domain-specific
A Feature Factory is the characteristic request-and-throughput realization of the broader Build Trap, though the full three-dial regime is not guaranteed.Both diagnose high delivery output with no owned or decision-gating outcome. Feature Factory adds the concrete feature-list intake, velocity dashboard, and launch-counting behavior; Build Trap states the broader measurement, calendar, and incentive regime that usually sustains it.
-
Feature Factory is a decomposition of Proxy-Target Divergence Prime
Stripped of product vocabulary, the factory is an operating apparatus closed on a controllable proxy after that proxy has decoupled from its target.Feature count or velocity begins as an upstream indicator of delivery, becomes the rewarded operating target, and then rises independently of retention, revenue, or mission impact. The source explicitly identifies this as its portable Goodhart-family structure.
Hierarchy paths (2) — routes to 2 parentless roots
- Feature Factory → Build Trap
- Feature Factory → Proxy-Target Divergence → Proxy–Target Fidelity → Representation → Abstraction
Not to Be Confused With¶
-
Feature creep. Its batch-mate and closest confusable — also a software-product pathology in which features accrete over time. But the two name different failures: feature creep is cost-blindness (a per-feature approval gate that never sees the compounding maintenance and complexity cost of all additions together), while the feature factory is outcome-blindness (an operating loop that rewards features shipped and never checks whether any moved a user or business metric). A product can suffer either without the other. Tell: is the un-tracked quantity the aggregate cost of the accumulated features (feature creep), or the outcome impact of each shipped feature (feature factory)?
-
Vanity metrics. Flattering numbers that look like progress but carry no information about real outcomes — raw sign-ups, page views, features shipped. The feature factory's proxy output is a vanity metric, but the anti-pattern is the whole operating loop closed on such a proxy — planning cadence, executive review, and performance evaluation all reporting against it — not the number in isolation. Tell: is the concern a single misleading indicator (vanity metric), or an entire product operating loop that substitutes such a proxy for the outcome it exists to move (feature factory)?
-
The cobra effect / perverse incentive. An incentive that induces behaviour producing the opposite of what was intended — breeding cobras to collect the bounty. This is a Goodhart cousin but differs in emphasis: the feature factory's harm is not a reversed outcome but a decoupled one — motion manufactured against the proxy while the target metric drifts independently and uncorrected. Tell: did the metric induce actively counterproductive behaviour that worsens the target (cobra effect), or merely output that is uninformative about and unmoored from the target (feature factory)?
-
The build trap. Melissa Perri's near-synonymous framing (her writing is among the outcomes-versus-outputs sources the entry cites) for the same product pathology — an organisation "stuck in the build trap" measures success by output rather than outcome. This is not a rival concept to sort against but the same mechanism under a different author's name; the difference is terminological lineage (Cutler's "feature factory" vs Perri's "build trap"), not structure. Tell: there is no structural discriminator — treat them as co-labels for the output-over-outcome product loop, not as two distinct patterns.
-
Goodhart's Law /
target_measure_substitution(the parent it instances). The substrate-neutral pattern — once a proxy measure becomes the target it ceases to track the outcome it was meant to measure, and local optimisation games the proxy as it drifts — that the feature factory explicitly is, applied to the product operating loop. It recurs as full co-instances sharing no product machinery (test-score gaming, hospital waiting-time gaming, citation counting, command-economy output tonnage), and Goodhart's Law carries that cross-domain structure. Campbell's Law is the narrower domain-specific species for consequential quantitative social indicators that distort the institutions they govern. Tell: off the product substrate — a hospital, a school, a command economy — the operative structure is Goodhart / target-measure-substitution; "feature factory," with its story-point and roadmap accent, applies only to a product organisation's operating loop. (Treated fully in the sections above.)
Neighborhood in Abstraction Space¶
Feature Factory sits in a crowded region of the domain-specific corpus (35th percentile for distinctiveness): several abstractions share nearly its structure, so a description that fits it tends to fit its neighbors too.
Family — Proxy Metrics & Venture Adaptation (13 abstractions)
Nearest neighbors
- Build Trap — 0.88
- Progress Illusion — 0.87
- Vanity-Metric Addiction — 0.86
- McNamara fallacy — 0.84
- Pivot Thrashing — 0.83
Computed from structural-signature embeddings · 2026-07-12