Ethics of Technology & AI Governance¶
← Back to Mechanisms by Origin Domain
The applied-ethical and policy discipline concerned with the responsible development, deployment, and governance of technology — particularly AI, algorithmic systems, data, and bio/cybertech. Canonical traditions: computer ethics lineage (Moor, Floridi), fairness-accountability-transparency in ML, AI safety and alignment, bioethics, privacy theory.
Reviewed origins (43)¶
These attributions have been reviewed as historical or practice origins and promoted to mechanism frontmatter.
- Algorithmic Ranking Audit — Tests an automated ranking or recommendation gate for the hidden demotion, bias, drift, and objective-mismatch that its published outputs alone never reveal.
- Autonomous Agent Safety Constraints — Bounds the permissions, rates, and objectives of autonomous agents inside a defined interaction boundary, re-tuning the limits as the agents adapt, so their local actions cannot aggregate into unsafe system behavior.
- Auxiliary-Prior Review Workshop — Convenes domain experts and adversarial reviewers to enumerate what an outside observer already knows, so a release is judged against real background knowledge rather than in isolation.
- Blanket Variable Quality Audit — Audits an established blanket for governance quality — that it collects no more than the minimal sufficient interface, and that the same interface holds across subgroups.
- Classification Fairness Review — Measures how a classification's errors — false inclusions and false exclusions — distribute across affected groups, exposing the burden and invisibility a formally neutral boundary can still produce.
- Collingridge Curve Workshop — Facilitates mapping of consequence-information growth against intervention-cost growth before a major commitment.
- Content Moderation Action Threshold — Locates a platform's enforcement line — remove versus leave up — by weighing wrongful restriction of a user's speech against the harm of content left to spread, and pairs it with an appeal path for the calls it gets wrong.
- Dataset Datasheet or Data Card — A standardized document shipped with a dataset that answers a fixed question set — provenance, composition, collection process, recommended and discouraged uses, and known limitations — tailored to its different audiences.
- Deployment Impact Dashboard — Tracks outcome, harm, dependency, and lock-in signals during staged deployment.
- Derived Eligibility or Status Answer — Answers the consumer's actual question with a computed predicate or status — 'meets the income threshold: yes' — returned live in place of the underlying record, so the source releases a conclusion instead of the data behind it.
- Distortion Model Card — Documents the assumed distortion pattern, supporting evidence, scope, counterevidence, and expiry conditions.
- Distributional Impact and Tail Audit — Compares benefit, burden, access, and error across parties and distribution tails.
- Ethical Impact Assessment — Captures foreseeable ethical consequences of applying or adapting a decision across contexts, including who benefits and who bears risk.
- Exception Review Queue — Routes rare legitimate-but-blocked cases through a governed human review so exceptions are granted without reopening the misuse path for everyone.
- Exit and Interoperability Rule — Requires portability, migration paths, or alternative access so stakeholders are not trapped before consequences are known.
- Explainability Review — Asks not whether a system is correct but whether its reasons are legible — whether the explanation it offers is understandable, and faithful, enough for the people who must rely on or contest the decision.
- Fairness Audit by Stratum — Checks whether differential treatment is producing intended fit without unacceptable disparate harm, exclusion, or hidden under-service.
- Fairness or Bias Audit — Checks whether optimization disproportionately burdens groups, hides inequity, or shifts harm to less visible stakeholders.
- Fairness-Metric and Exception Stress Test — Probes proxies, gaming, baseline shifts, temporal drift, and exception capture.
- Filter Transparency Dashboard — Shows recommendation, moderation, citation, agenda, or selection patterns that shape ordinary intake.
- Generated Content Disclosure Gate — Holds internally- or model-generated content at the point of release until it carries a label saying it was generated and is phrased so a downstream reader can weight it as such.
- Hallucination Check — A review pass over generated or inferred content that flags every unsupported detail and verifies each nontrivial claim against a real source before it is trusted.
- Human Review Escalation Cutoff — Sets the confidence line at which an automated decision system stops deciding and hands a case to a human — a line bounded above all by how many cases the reviewers can actually handle.
- Human Review Trigger — Requires accountable review when marginal gains are small but protected values, human impacts, or legitimacy concerns are at stake.
- Manual Boundary Review Queue — Routes the cases an automated cut can't safely resolve to human reviewers, whose accumulated rulings become a deliberative record of how the boundary is actually applied.
- Method Card or Model Card — A published, standardized card that states a method's intended and out-of-scope uses, its performance broken out by condition, and the tradeoffs each stakeholder inherits — so downstream users receive the method's limits, not just its headline number.
- Misuse Monitoring Dashboard — Instruments the deployed design for residual misuse, bypass attempts, false blocks, and workaround traces, and feeds them back into revision.
- Model Applicability Card — A short published document that states what a model is validated for — its intended use, input populations, excluded uses, and the assumptions that must hold — so it isn't trusted outside the conditions it was built and tested under.
- Model Card or Datasheet Linkage — Attaches model, dataset, or artifact metadata to the abstraction so downstream users can inspect provenance, intended use, excluded use, evaluation, and limitations.
- Model Card Value Section — Adds explicit value assumptions, intended uses, excluded uses, affected groups, and evaluation priorities to technical documentation for AI or analytic systems.
- Model Limitations Card — A short document that travels with a model, dataset, or calculation and states where it is valid, where it is uncertain, and where it is unsafe to use — so an authoritative-looking output cannot be trusted beyond the conditions it was built for.
- Model or Rule Card — Documents intended use, constraints, known failure modes, data assumptions, explainability, and review owner for the selected method.
- Model Scope Review — Audits what a model, dataset, or metric leaves out of its frame — and how those exclusions inflate the claims made from it.
- Out-of-Domain Prompt — Fires an interrupt the moment a claim crosses out of validated scope, forcing a pause and a downgrade before the recommendation is accepted.
- Participatory Fairness Deliberation — Structures affected-party evidence and reasons around candidate standards, consequences, and conflicts.
- Platform Access Rule — Defines the published conditions under which each side can join, list, sell, or build on a platform — and be removed — so access turns on stated criteria rather than the operator's discretion.
- Platform Moderation Strike System — Running software that detects rule violations, records strikes against a specific account, escalates restrictions as strikes accumulate, and gives the user notice and a route to appeal.
- Platform Rule — Constrains participant behavior in a platform, marketplace, forum, or shared infrastructure by defining allowed actions and consequences.
- Purpose-Based Access Request — Makes a consumer declare, before any data flows, the specific purpose and the task-justified fields it needs — so access is granted against a stated need rather than a standing entitlement.
- Residual Leakage Review Board — A standing cross-functional body that reviews the leakage remaining after controls, sets the tolerated distinguishability budget, and records — with named accountability — what residual risk is formally accepted.
- Reward Signal Red Team — A standing adversarial team that tries to break a reward signal before it trains anyone — hunting for ways to score high while defeating the intent, and for who gets hurt in the process.
- Stakeholder Exclusion Audit — Checks which affected parties sit outside the boundary of evidence, participation, or remedy — and gives them a channel to put that exclusion on the record.
- Subgroup Outcome-Validity Dashboard — Combines protected subgroup patterns with validity, scoring, persistence, review, and experience evidence.
Also Draws from This Domain (223)¶
These mechanisms have another primary origin but were reviewed as also drawing materially from this domain.
Because this set contains more than 100 mechanisms, it is divided by solution family—the governing move the mechanism makes. This is a browsing subdivision only; it does not change the origin attribution. Click a family below to jump to its fully visible section, or click a column header to sort.
| Solution family | Mechanisms | Description |
|---|---|---|
| Access, Admission & Permissions | 1 | Solutions that decide who or what may enter, act, consume capacity, or cross a protected boundary, including eligibility rules, quotas, credentials, and scoped authority. |
| Adaptation & Reconfiguration | 11 | Solutions that alter structure, parameters, roles, or behavior in response to changing conditions while preserving the system's purpose. |
| Aggregation & Synthesis | 7 | Solutions that combine many observations, judgments, signals, or parts into a useful whole while managing weighting, dependence, and loss of detail. |
| Alignment & Incentives | 6 | Solutions that make individual choices, rewards, responsibilities, or local objectives support a larger goal instead of working against it. |
| Allocation & Prioritization | 5 | Solutions that distribute scarce attention, effort, money, capacity, or opportunity among competing claims and make the order of service explicit. |
| Anticipation & Forecasting | 6 | Solutions that look ahead, surface plausible futures, identify leading indicators, or prepare options before a consequential state arrives. |
| Attention, Salience & Focus | 1 | Solutions that direct limited attention toward what matters, protect focus from interference, or deliberately change what becomes noticeable. |
| Boundary & Scope Control | 3 | Solutions that define, move, or police what is inside a problem, system, role, claim, or responsibility and what remains outside it. |
| Buffering & Reserves | 1 | Solutions that absorb variability, delay, shocks, or temporary imbalance through slack, queues, inventories, reserves, or intermediate storage. |
| Calibration & Tuning | 3 | Solutions that compare behavior with a reference and adjust parameters, thresholds, mappings, or tolerances until performance falls within an acceptable range. |
| Causal Diagnosis | 1 | Solutions that distinguish symptoms from causes, compare explanations, localize a fault, or identify the intervention point responsible for an outcome. |
| Classification & Taxonomy | 1 | Solutions that sort cases into meaningful classes, establish membership criteria, or organize concepts so distinctions can guide action. |
| Communication & Signaling | 3 | Solutions that convey meaning, intent, state, or credibility across people or systems while accounting for interpretation, noise, and strategic response. |
| Comparison & Evaluation | 2 | Solutions that place alternatives, cases, or outcomes against shared criteria so differences become visible and judgments become defensible. |
| Compression & Simplification | 4 | Solutions that reduce complexity, detail, or dimensionality while retaining the structure needed for the current decision or task. |
| Constraints & Guardrails | 14 | Solutions that prevent unacceptable states or actions by encoding limits, invariants, preconditions, safe envelopes, or error-proofing rules. |
| Containment & Isolation | 3 | Solutions that keep faults, hazards, conflicts, contamination, or overload from spreading by separating regions, flows, or responsibilities. |
| Coordination & Synchronization | 10 | Solutions that align interdependent actors, tasks, clocks, states, or handoffs so joint work progresses without collision or drift. |
| Decoupling & Interfaces | 3 | Solutions that reduce harmful dependency by inserting contracts, adapters, abstractions, or replaceable boundaries between interacting parts. |
| Emergence & Self-Organization | 4 | Solutions that shape local rules, interactions, or environmental cues so useful global order can arise without direct central specification. |
| Evidence, Inference & Validation | 8 | Solutions that gather, test, triangulate, or qualify evidence so claims and decisions match what the observations can actually support. |
| Feedback & Regulation | 6 | Solutions that sense the effects of action and use the result to stabilize, steer, damp, amplify, or otherwise regulate subsequent behavior. |
| Flow & Routing | 1 | Solutions that direct material, information, demand, work, or traffic through paths and stages to improve movement and avoid congestion. |
| Governance & Accountability | 18 | Solutions that allocate decision rights, oversight, responsibility, transparency, and consequences so power remains answerable and action-owned. |
| Identity, Reference & Matching | 2 | Solutions that establish what an entity is, bind records to the right referent, resolve names, or match cases without confusing near-equivalents. |
| Knowledge, Memory & Provenance | 3 | Solutions that capture, retain, retrieve, transfer, and trace knowledge or records so later users can recover both content and origin. |
| Learning & Scaffolding | 2 | Solutions that sequence practice, feedback, examples, and support so capability grows and transfers beyond the original learning setting. |
| Mapping & Transformation | 2 | Solutions that translate between representations, coordinate systems, scales, formats, or states while preserving the relationships that matter. |
| Measurement & Observability | 6 | Solutions that make hidden state inferable through instruments, indicators, probes, sampling, or diagnostic views with known limits. |
| Negotiation & Strategic Interaction | 6 | Solutions that account for other agents' incentives, reactions, commitments, bargaining power, and counter-moves when outcomes are interdependent. |
| Optimization & Search | 9 | Solutions that explore alternatives under objectives and constraints, prune infeasible regions, and improve a candidate toward a chosen criterion. |
| Participation, Norms & Culture | 10 | Solutions that shape belonging, legitimacy, shared expectations, collective practice, and the willingness of people to contribute or comply. |
| Planning & Staging | 1 | Solutions that turn an intended outcome into phases, milestones, option points, and coordinated preparations before execution. |
| Recovery & Restoration | 2 | Solutions that return a damaged, degraded, or interrupted system to service through repair, rollback, reentry, regeneration, or reconstruction. |
| Reframing & Sensemaking | 11 | Solutions that change the interpretive frame, surface hidden assumptions, or organize ambiguous experience into a more useful account. |
| Representation & Modeling | 10 | Solutions that construct schemas, models, diagrams, abstractions, or formal descriptions that make structure available for reasoning. |
| Risk, Robustness & Uncertainty | 1 | Solutions that make uncertainty explicit, limit downside, preserve acceptable behavior across variation, or prepare contingencies for adverse outcomes. |
| Scaling & Capacity | 1 | Solutions that match capability to load, grow or shrink safely, and manage how structure and performance change with size. |
| Selection & Filtering | 5 | Solutions that admit, retain, rank, or reject candidates according to fitness, relevance, quality, or another discriminating rule. |
| Stress Testing & Rehearsal | 2 | Solutions that expose a system or organization to controlled difficulty, adversarial conditions, or practice scenarios before real failure stakes apply. |
| Substitution & Fallback | 2 | Solutions that replace unavailable or unsuitable means with alternatives while preserving the essential function, contract, or outcome. |
| Thresholds & Phase Change | 4 | Solutions that detect, create, avoid, or govern nonlinear transitions when accumulating conditions cross a consequential boundary. |
| Tradeoffs & Decision Support | 5 | Solutions that expose competing objectives, preference structure, stopping rules, and consequences so a choice can be made under constraint. |
| Transmission, Propagation & Networks | 16 | Solutions that shape how signals, behaviors, effects, or resources spread through channels and network topology over space or time. |
| Variation & Experimentation | 1 | Solutions that deliberately vary conditions, compare trials, preserve controls, and learn from differential outcomes without overclaiming. |
Access, Admission & Permissions¶
Solutions that decide who or what may enter, act, consume capacity, or cross a protected boundary, including eligibility rules, quotas, credentials, and scoped authority.
1 mechanism · View full solution family
- Appeals or Reconsideration Workflow — Gives a contested decision a defined second path — timed review, new-evidence handling, and a route to escalation — so a denial can be corrected or confirmed by someone other than the gate that made it.
Adaptation & Reconfiguration¶
Solutions that alter structure, parameters, roles, or behavior in response to changing conditions while preserving the system's purpose.
11 mechanisms · View full solution family
- Abuse-Case Replay Harness — Replays sanitized abuse scenarios to check whether a proposed mitigation catches the pattern without unacceptable collateral damage.
- Equitable Access and Consent Review — An independent oversight review that checks a time-critical intervention reaches everyone fairly and consensually — and that the claimed window is real, not urgency manufactured from shaky group evidence.
- Identity-Question Timing Protocol — Governs when identity data are requested, who can see them, and how collection is separated from performance when appropriate.
- Identity-Safe Review Channel — Provides confidential intake, independent review, correction authority, pattern escalation, and retaliation protection.
- IP and Provenance Checklist — Runs the source through a fixed list of rights, confidentiality, attribution, and safety questions before and during extraction — gating whether this source may be learned from at all, and how.
- Platform-Ecosystem Rule Change — Changes the rules of a live digital ecosystem and governs the fast, often adversarial way participants re-adapt to them.
- Rate-Limited Friction Escalation — Adds proportional friction to suspicious repeated behavior while preserving legitimate access and appeal paths.
- Responsible Disclosure Absorption Pipeline — Turns good-faith external reports into verified fixes and learning updates without amplifying exploit detail.
- Staged Rule Rollout with Rollback — Limits the blast radius of rapid updates by releasing them gradually and reverting if harm indicators rise.
- Stakeholder Boundary Review — Decides who counts as inside the system being shaped — constructors, beneficiaries, and the affected outsiders who bear the spillovers — before the boundary is drawn implicitly by whoever holds the pen.
- Weakened Adversarial Example Set — A curated corpus of real attack patterns deliberately weakened to below the harm threshold and chosen to span the threat family, so a learner or model can train against safe specimens of the whole attack space.
Aggregation & Synthesis¶
Solutions that combine many observations, judgments, signals, or parts into a useful whole while managing weighting, dependence, and loss of detail.
7 mechanisms · View full solution family
- Appeal and Exception Review — A standing channel through which a filtered-out or softened output can be contested, and exceptions granted on stated public reasons — a release valve on the intersection.
- Diverse Recommendation Exposure — Rebalances a feed or search ranking so raw popularity is offset by source diversity, minority evidence, uncertainty, and independent quality signals.
- Filter Rationale Register — A standing record of why each filter exists — its stated purpose, owner, and the legitimacy standard it claims — so every screening rule can be traced, justified, or challenged.
- Progressive Friction — Makes additional accumulation increasingly costly or review-heavy as the compounding variable approaches a dangerous range.
- Shadow Review Board — A standing independent panel that re-adjudicates a running stream of filtered-out outputs under its own declared criteria, revealing what the live filter set would have passed had the judgment been someone else's.
- Structural Filter Postmortem — After an output that should have surfaced didn't, reconstructs the trace of every filter it hit and how they combined — blamelessly, because each filter was locally reasonable and each producer sincere.
- Viewpoint Presence Dashboard — Tracks, period over period, which viewpoints and sources are present in the surviving output and which stay absent — turning slow filter drift toward homogeneity into something you can watch.
Alignment & Incentives¶
Solutions that make individual choices, rewards, responsibilities, or local objectives support a larger goal instead of working against it.
6 mechanisms · View full solution family
- Anti-Gaming Scoring Rule — Scores behavior so the top score is earned by producing the real outcome, not by manipulating the measured proxy, and re-tunes as gaming emerges.
- Over-Suppression Red Team — Deliberately attacks the suppression rule to surface the valid weak signals it has been quietly erasing — the minority views, faint evidence, and rare safety-critical cases hidden among the losers.
- Product Purpose Review — Audits a product's features, roadmap, and growth targets against the user outcome it exists to serve, flagging the ones that grow engagement while failing the beneficiary.
- Spurious Association Probe Set — A standing battery of targeted test cases that deliberately try to trip a learned link into revealing that it rides on a shortcut, a stereotype, or a leaked cue rather than the real signal.
- Structural Harm Scan — Surfaces the systemic conditions a simple blame story hides — incentives, resource gaps, design defects, hidden beneficiaries — and gives omitted stakeholders a voice.
- Subgroup Disaggregation Audit — Breaks down aggregate impacts by subgroup, geography, role, income, exposure, access, or vulnerability to reveal hidden losses and uneven gains.
Allocation & Prioritization¶
Solutions that distribute scarce attention, effort, money, capacity, or opportunity among competing claims and make the order of service explicit.
5 mechanisms · View full solution family
- API Response Projection — Shapes the outgoing response at the producer, composing it from an allow-list of only the fields a given consumer's role and purpose justify, so surplus data is never serialized and never leaves the source.
- Attribute-Based Access Policy — Computes at request time what a consumer may receive by evaluating attributes of the actor, resource, purpose, and context against per-field necessity rules — so the disclosed view narrows or widens with the situation instead of being a fixed grant.
- Claim Certificate or Verifiable Credential — Packages a single attested fact — 'over 21', 'currently licensed', 'in good standing' — as a portable, cryptographically-verifiable credential the holder presents in place of the underlying record, and that can expire or be revoked.
- Field-Level Redaction — Removes or blacks out the specific fields flagged sensitive or surplus from an outgoing record, at the producer, so what leaves carries only what the recipient may see.
- Privacy Impact Review — A pre-release assessment that maps what a source record actually contains and what a recipient could infer or re-identify from a proposed disclosure, before the disclosure is designed.
Anticipation & Forecasting¶
Solutions that look ahead, surface plausible futures, identify leading indicators, or prepare options before a consequential state arrives.
6 mechanisms · View full solution family
- Forecast Impact Audit — Examines, after release, how a forecast actually moved behavior — comparing the reaction that occurred against the reaction that was modeled, and testing whether anyone gamed it — to tell a self-defeating forecast apart from a merely wrong one.
- Forecast Release Decision Log — A dated, append-only record of each forecast released — the exact claim, who could see it, and the disclosure boundary applied — so the decision to publish a reactive forecast can be reviewed against what was known at the time, not what happened after.
- Forecast-as-Intervention Label — A standing tag attached to a forecast that declares it can change the outcome it predicts, telling readers to treat it as guidance to act on — and stating why it is being disclosed at all.
- Leakage Scan — Systematically interrogates a candidate feature set for information that would not be available in real use — future outcomes, post-decision fields, or forbidden proxies — before any of it ships.
- Shortcut Probe Holdout Set — A curated held-out test set where the suspected shortcut cue is deliberately broken, exposing whether the system learned the real signal or a convenient proxy that merely correlated with reward.
- Transition Harm Review — Surfaces the users, workers, dependencies, and public goods a disruption could harm as it displaces the incumbent, and designs the mitigations and guardrails before scale locks the harm in.
Attention, Salience & Focus¶
Solutions that direct limited attention toward what matters, protect focus from interference, or deliberately change what becomes noticeable.
1 mechanism · View full solution family
- Salience Red Team — A standing adversarial group chartered to ask what the loudest items are crowding out and who engineered their prominence.
Boundary & Scope Control¶
Solutions that define, move, or police what is inside a problem, system, role, claim, or responsibility and what remains outside it.
3 mechanisms · View full solution family
- Content Moderation Gate — A platform boundary that reviews user-generated content and allows, removes, labels, or downranks it by safety, legality, and community rules — with a path to appeal.
- Risk Capital Requirement — Forces an actor to hold capital, reserves, or insurance against the low-probability, high-consequence harms it could impose on others — so the risk it creates sits on its own balance sheet.
- Stakeholder-Inclusive Redesign — Redraws the problem frame around the lived constraints of the people it affects, so their experience — not just the technical unit — decides what belongs inside and what success means.
Buffering & Reserves¶
Solutions that absorb variability, delay, shocks, or temporary imbalance through slack, queues, inventories, reserves, or intermediate storage.
1 mechanism · View full solution family
- Substrate Risk Release Gate — A pass/block control at the release point that refuses to ship substrate whose inherited risk is unaccounted-for or exceeds a blast-radius-scaled bar.
Calibration & Tuning¶
Solutions that compare behavior with a reference and adjust parameters, thresholds, mappings, or tolerances until performance falls within an acceptable range.
3 mechanisms · View full solution family
- Label Proxy Screen — Scans every candidate feature for the tell-tale signature of a target proxy — a column that is suspiciously predictive because it is really a downstream trace of the outcome — and files the suspects for confirmation.
- Low-Confidence Escalation Trigger — Diverts any case whose heuristic confidence falls below a set threshold out of the fast path and into human review, logging each hand-off as an exception.
- Subgroup Coverage Calibration Table — A table that reports nominal versus realized coverage broken out by subgroup, site, period, or risk stratum, so local undercoverage cannot hide inside a healthy overall average.
Causal Diagnosis¶
Solutions that distinguish symptoms from causes, compare explanations, localize a fault, or identify the intervention point responsible for an outcome.
1 mechanism · View full solution family
- Subgroup Residual Heatmap — Tiles average residual across two crossed segmentations so a subgroup the overall fit hides lights up as a hot cell.
Classification & Taxonomy¶
Solutions that sort cases into meaningful classes, establish membership criteria, or organize concepts so distinctions can guide action.
1 mechanism · View full solution family
- Classification Disagreement Audit — Measures where classifications diverge — reviewer vs reviewer, human vs model — to expose systematic bias and human-model misalignment.
Communication & Signaling¶
Solutions that convey meaning, intent, state, or credibility across people or systems while accounting for interpretation, noise, and strategic response.
3 mechanisms · View full solution family
- Action-Link Landing Path — Routes post-story motivation into a proportionate choice, resource, support, commitment, or deliberation step.
- Metaphor Drift, Harm, and Downstream-Decision Audit — Samples reuse reinterpretation translation policy incidents identity effects and outcomes to trigger revision or retirement.
- Narrative Ethics Review — Checks truthfulness, agency, disclosure, vulnerability, stigma, emotional coercion, and overgeneralization risk against explicit legitimacy boundaries.
Comparison & Evaluation¶
Solutions that place alternatives, cases, or outcomes against shared criteria so differences become visible and judgments become defensible.
2 mechanisms · View full solution family
- Anti-Discrimination Check — Holds a case fixed and flips only a protected characteristic — race, sex, religion, disability, age — to see whether treatment moves; a targeted symmetry test for the markers the law and ethics forbid from counting.
- Fairness-Standard Comparison Table — Applies candidate standards to the same decision and displays reasons, winners, burdens, conflicts, and uncertainty.
Compression & Simplification¶
Solutions that reduce complexity, detail, or dimensionality while retaining the structure needed for the current decision or task.
4 mechanisms · View full solution family
- High-Risk Targeting List — Ranks cases, sites, or suppliers by predicted contribution to harm or cost so scarce scrutiny lands on the riskiest few — and holds the risk scores themselves to account.
- Manual Review Route — Diverts cases that automated rules cannot safely decide to human judgment, so ambiguous tail cases get context instead of a confident wrong answer.
- Omission Checklist — A standardized prompt sheet that forces reviewers to name what a simplification removed — the variables, cases, stakeholders, and steps dropped for simplicity — before anyone judges whether the loss matters.
- Stakeholder Review — Asks the people who actually use, operate, or are affected by a simplified artifact which omitted cases and constraints they consider important — surfacing losses invisible to its designers.
Constraints & Guardrails¶
Solutions that prevent unacceptable states or actions by encoding limits, invariants, preconditions, safe envelopes, or error-proofing rules.
14 mechanisms · View full solution family
- Appeal and Rapid Restoration Workflow — Gives a wrongly-engaged legitimate party a fast, independent path to contest the action and have the harm reversed before it hardens into permanent loss.
- Cue-Hijack Red-Team Review — Tasks an adversarial reviewer with actively finding ways a design could capture attention, appetite, fear, or reward-seeking beyond user intent, focusing on vulnerable slices.
- Default-Off High-Stimulation Setting — Keeps especially vivid, autoplaying, variable-reward, or urgent cue modes switched off out of the box, so they turn on only when a user or governing actor intentionally enables them.
- False-Positive Harm Budget Dashboard — Meters the running cost of wrongful self-engagements against a pre-set allowance, weighting each by harm intensity, so the defense's autoimmune damage is priced and capped.
- Frequency Cap and Cooldown — Limits how many times a cue may be shown in a window and inserts a quiet period before it can be repeated, refreshed, or re-escalated.
- Graduated Response Matrix — A lookup table that maps classifier confidence and self-status ambiguity against response harm, so uncertain judgments are routed to weaker, more reversible actions.
- Granular Opt-In Flow — Splits the single 'I agree' into separate, independently refusable consents, so accepting core access never silently carries optional data uses, communications, or add-ons.
- High-Arousal Content Throttle — Reduces the distribution, ranking boost, or autoplay escalation of live content whose cue profile provokes disproportionate arousal relative to its substantive value.
- Purpose-Bound Security Condition Record — Records each security or safety condition that stays attached to access — its purpose, scope, evidence, expiry, and the less-restrictive alternatives weighed — and binds it to that purpose so a legitimate guardrail cannot quietly become an all-purpose leash.
- Rebundling Drift Audit — Periodically re-checks a decoupled bundle to catch conditions that have crept back through defaults, pricing, renewal friction, degraded alternatives, or quiet policy edits.
- Review Queue — Routes proposed archetype matches into review by stakes, holding the higher-consequence ones for scrutiny and escalation instead of letting every confident match act on itself.
- Sandbox or Pilot Pathway — Runs a candidate hidden path inside a bounded, reversible enclosure with a limited blast radius, so its feasibility can be measured before the full system is exposed to it.
- Supernormal Cue Audit — Systematically reviews a product, environment, or interface against a checklist of known cue channels, flagging any that exceed a natural or validated calibration range.
- Variable-Reward Schedule Limit — Restricts intermittent, surprise, or loot-like reward loops when their unpredictability is what drives compulsive checking, redirecting toward predictable or earned reinforcement.
Containment & Isolation¶
Solutions that keep faults, hazards, conflicts, contamination, or overload from spreading by separating regions, flows, or responsibilities.
3 mechanisms · View full solution family
- Adaptive Circumvention Red Team — Plays the motivated adversary against a control to find how it will be evaded and which under-defended destination the blocked pressure will be pushed toward.
- Minimal Interface Dashboard — A standing operational view that surfaces only the validated blanket variables and wires each to the decision it informs — turning the minimal sufficient interface into the one screen people actually watch and act on.
- Synthetic Data Testbed — Swaps sensitive live data for a generated stand-in so pipelines and models can be exercised without exposing real records.
Coordination & Synchronization¶
Solutions that align interdependent actors, tasks, clocks, states, or handoffs so joint work progresses without collision or drift.
10 mechanisms · View full solution family
- Agency Health Dashboard — Turns the live health of an agency loop — is feedback timely, is the actor actually acting, is discretion being used — into a small set of continuously-watched signals.
- Challenge Window and Correction Protocol — Gives a party classified or scored on a private record a bounded, defined window to contest it and force a re-check — turning a one-sided datum into something its subject can see and correct before it hardens into a decision.
- Harm Reduction Dashboard — A live instrument that tracks, after a remedy ships, whether harm and burden actually fell over time — including whether the burden was merely shifted somewhere less visible.
- Human-in-the-Loop Operating Model — Defines what humans review, decide, override, escalate, maintain, or learn from when automation participates in the work.
- Platform Ecosystem Change Council — The standing, representative body that holds decision authority over ecosystem-wide changes — breaking contracts, participation terms, ranking and fees, deprecation — so the rules that decide who captures value are made with the builders who live by them.
- Platform Extension Review and Certification — A risk-tiered gate that combines automated evidence with human judgment to certify an extension safe and compatible enough to admit — with the depth of scrutiny scaled to the potential harm.
- Safety Case Review — Examines whether technical controls, human practices, incentives, and governance together make the system acceptably safe.
- Sociotechnical Design Workshop — Brings technical owners, operators, affected users, managers, and governance owners together to map coupled social-technical changes.
- Structural Audit — A systematic inspection of one organization's own rules, workflows, incentives, and resource flows to surface the arrangements that quietly generate harm.
- Systems Harm Analysis — Follows a person or group across the several institutions whose separate rules compound into harm no single agency owns, and asks what the interaction — not any one part — produces.
Decoupling & Interfaces¶
Solutions that reduce harmful dependency by inserting contracts, adapters, abstractions, or replaceable boundaries between interacting parts.
3 mechanisms · View full solution family
- Managed Discourse Retreat Plan — A plan for withdrawing a position from active normalization when further exposure causes harm or when the legitimacy rationale fails.
- Platform Policy Harmonization — Aligns a platform's rules across its regions, categories, and seller tiers so users can't escape a safeguard by relabeling the same activity.
- Quarantine and Manual Review Queue — Holds inbound artifacts that are neither cleanly acceptable nor safely rejectable in a queue, so a human resolves the ambiguous middle instead of the parser silently guessing.
Emergence & Self-Organization¶
Solutions that shape local rules, interactions, or environmental cues so useful global order can arise without direct central specification.
4 mechanisms · View full solution family
- Anti-Spam Rules — Places local posting, account, and message constraints — with allow-listed exceptions — on the channels where many small sends aggregate into systemic spam or abuse.
- Emergent-Risk Moderation — Moderates behavior by its contribution to a forming harmful macro-pattern rather than by isolated rule violations, adjusting thresholds as the pattern shifts.
- Platform Abuse Controls — Runs distributed abuse through an end-to-end pipeline — detect the pattern, throttle or restrict, adjudicate appeals, and watch for displacement — to contain coordinated misuse.
- Recommendation Diversity Constraint — Adds relevance-bounded diversity, serendipity, or source-independence rules to ranking and feed systems.
Evidence, Inference & Validation¶
Solutions that gather, test, triangulate, or qualify evidence so claims and decisions match what the observations can actually support.
8 mechanisms · View full solution family
- Adversarial Example Generation — Constructs hard inputs deliberately engineered to make a rule fail, then keeps only the ones that stay realistic enough to matter in the real operating scope.
- Anonymous Membership Proof — Proves that the prover belongs to an authorized set without identifying which member they are.
- Benchmark Suite Coverage Matrix — Maps every benchmark case against the tasks, subgroups, operating conditions, and failure modes it exercises, so the blank cells — the parts of the domain nothing tests — become visible before a headline score is mistaken for a passing grade.
- Distributional-Assumption Card — A one-page record that pins a distributional commitment — modeled quantity, family, support, rationale, evidence, decision use, owner, and expiry — into an inspectable, hand-off-safe contract.
- Hallucination Intrusion Triage — Takes items already flagged as possible fabrications or memory intrusions and sorts them by how much rides on them, quarantining, escalating, or releasing each before it is trusted.
- Multiverse Analysis Report — Runs the analysis across every defensible analytic choice at once and shows the whole spread of results, exposing whether the headline depends on one lucky path.
- Privacy-Preserving Compliance Oracle — Checks a private record against public rules and emits only a bounded compliance verdict.
- Subgroup Dashboard with Warning Flags — Shows aggregate and subgroup figures side by side with rules that flag masked harm, unstable small cells, and equity-relevant gaps as they arise.
Feedback & Regulation¶
Solutions that sense the effects of action and use the result to stabilize, steer, damp, amplify, or otherwise regulate subsequent behavior.
6 mechanisms · View full solution family
- Capability-Scoped Tool Gateway — Checks policy and capability scope before interpreted content can call tools or affect protected state.
- LLM Instruction/Data Boundary — Separates system, developer, tool, user, and retrieved-context roles so untrusted text cannot become tool-authoritative instruction.
- Model Registry — The system of record for every regulating model — its lineage, assumptions, owner, approvals, and deployment status — so any model in production can be traced, re-approved, or rolled back.
- Model-Failure Red Team — An independent team whose mandate is to make the model fail — hunting the conditions under which it gives wrong answers, mapping that failure frontier, and checking the system degrades safely past it.
- Prediction Impact Audit — Checks whether a risk score, forecast, or warning label is changing how a system treats its subjects in ways that help bring about the very outcome it predicts.
- Training-Data Exclusion List — A standing denylist that stops marked synthetic or planted artifacts from being ingested into models, dashboards, search indexes, and decision-support datasets.
Flow & Routing¶
Solutions that direct material, information, demand, work, or traffic through paths and stages to improve movement and avoid congestion.
1 mechanism · View full solution family
- Vulnerability-Based Support Workflow — Directs additional protection, outreach, simplification, or case management toward strata with lower capacity or higher exposure to harm.
Governance & Accountability¶
Solutions that allocate decision rights, oversight, responsibility, transparency, and consequences so power remains answerable and action-owned.
18 mechanisms · View full solution family
- API Governance Policy — Governs how third parties build on an operator's API — access tiers, stability guarantees, deprecation notice, and security rules — so an ecosystem can depend on an interface that won't shift without warning.
- Appeal and Dispute Process — Gives a participant hit by a suspension, delisting, or access denial a real channel to contest it before a reviewer who didn't make the original call — with a path back if it was wrong.
- Consent Capture and Revocation Workflow — Records what has been agreed to, under what scope, and how consent can be renewed, limited, or withdrawn where appropriate.
- Consent Renewal Prompt — A change-triggered workflow that re-discloses what shifted and asks the person to affirmatively re-agree to the updated scope before the new use begins.
- Data Portability Rule — Requires that a participant can export their own data — and, where it applies, their reputation and connections — in a usable, machine-readable format, so leaving costs no more than the network's value honestly justifies.
- Emergency Override Protocol — Allows temporary external intervention only when predefined harm, rights, safety, legal, or systemic thresholds are met.
- Ethics Checklist — Prompts reviewers to ask value, harm, consent, fairness, transparency, and responsibility questions. It implements the archetype only when answers are recorded and contestable.
- Granular Permission Dashboard — A single review surface that lays out every permission a person has granted — split by purpose, actor, and duration — so the whole consent landscape can be inspected and spotted for staleness at a glance.
- Metric Value Review — Examines what a metric rewards, ignores, normalizes, or sacrifices, especially when the metric is treated as objective evidence of success.
- Moderation and Abuse Response — Detects harms that network scale amplifies — fraud, harassment, spam, manipulation — and responds with proportionate, reviewable action fast enough to matter.
- Moderation Appeal — A platform-specific workflow for challenging content, account, or community enforcement decisions.
- Moderation Appeal Process — Reviews contested platform enforcement decisions using evidence, rules, reasons, and remedy options.
- Moderation Strike System — Implements scaled platform responses for repeated or severe violations, provided it preserves context, appeal, and de-escalation.
- Opt-In Flow — A pre-entry interface gate that presents one permission request with its material terms at the threshold of a feature, and requires an affirmative choice against a genuine, equally-weighted decline.
- Permission–Entitlement Crosswalk — Maps technical permissions against substantive entitlements to catch capabilities with no claim behind them and claims the system cannot actually execute.
- Transparency Impact Review — Periodically asks whether all the disclosure is actually producing accountability, and at what burden, rather than just accumulating published volume.
- Value Audit — Reviews a policy, model, metric, or process to identify hidden value priorities, displaced alternatives, affected parties, and unsupported legitimacy claims.
- Withdrawal Procedure — The end-to-end workflow that executes a person's decision to take back or narrow an existing permission — receiving it, propagating the stop downstream, and stating what can and cannot be undone.
Identity, Reference & Matching¶
Solutions that establish what an entity is, bind records to the right referent, resolve names, or match cases without confusing near-equivalents.
2 mechanisms · View full solution family
- Exemplar Feedback Registry — Logs what happened every time an exemplar was reused and uses those outcomes to broaden, narrow, or retire each stored case's authority — so the case memory sharpens instead of fossilizing.
- Stereotype / Tokenism Red Team — Reviews whether the identity appeal relies on caricature, mascot examples, demographic essentialism, or borrowed legitimacy.
Knowledge, Memory & Provenance¶
Solutions that capture, retain, retrieve, transfer, and trace knowledge or records so later users can recover both content and origin.
3 mechanisms · View full solution family
- API Reuse Boundary Header — Rides boundary facts — version, deprecation date, required scope, rate limits, privacy constraints — on the API call itself, so a developer meets the constraints at the moment they invoke the endpoint.
- Source-to-Score Lineage Graph — Visualizes lineage from substrate records through transformations to the final score, label, dashboard value, or decision artifact.
- Stakeholder Knowledge Forum — Creates a structured setting where different knowledge holders compare evidence, explain assumptions, surface interpretive gaps, and record changes to the knowledge base.
Learning & Scaffolding¶
Solutions that sequence practice, feedback, examples, and support so capability grows and transfers beyond the original learning setting.
2 mechanisms · View full solution family
- Consequence Design Review — Reviews proposed rewards, penalties, feedback, recognition, and natural consequences for alignment, proportionality, timing, fairness, and side effects.
- Uncertainty Tagging — Attaches a travel-with-the-claim status label — observed, inferred, assumed, estimated, unverified, verified — to each part of a completion, and logs when that status changes.
Mapping & Transformation¶
Solutions that translate between representations, coordinate systems, scales, formats, or states while preserving the relationships that matter.
2 mechanisms · View full solution family
- Model Specification — States the inputs a model accepts, the outputs and ranges it produces, and the assumptions and scope of validity under which those outputs can be trusted — so downstream users know where the model applies and where it must not be used.
- Viewpoint-Omission Audit — Reviews what the chosen station, crop, and projection hide, shrink, or flatter — and requires alternate evidence wherever an omission carries real consequence.
Measurement & Observability¶
Solutions that make hidden state inferable through instruments, indicators, probes, sampling, or diagnostic views with known limits.
6 mechanisms · View full solution family
- Certification Regime — A standing institution that codifies the evidence a system must present before it is approved — keyed to its risk class — and defines the events that force the credential to be re-earned.
- Claim-Scope Watermark — Stamps every output with the vantage it came from and the slice of the world it can honestly speak for, so the scope travels with the claim instead of being stripped the moment it's quoted.
- Deviation Review Queue — Routes flagged departures to human or automated review, annotation, escalation, or follow-up, with a fairness check on who gets scrutinized.
- Risk Score Proxy Metric — Uses a composite score as an indirect estimate of risk, quality, eligibility, or likely behavior, requiring strong fairness and validity safeguards.
- Sentinel Outcome Dashboard — A standing, owner-facing display that lines up the proxy against downstream outcome and harm signals so silent decoupling becomes visible at a glance.
- Validity Limitation Memo — A short written statement travelling with the measure that fixes what its scores may and may not be used to claim, for whom, and what harms to watch when it's used.
Negotiation & Strategic Interaction¶
Solutions that account for other agents' incentives, reactions, commitments, bargaining power, and counter-moves when outcomes are interdependent.
6 mechanisms · View full solution family
- Abuse Complaint and Appeals Process — Gives affected parties a reviewable channel to challenge denial, degradation, or retaliation — surfacing abuse, adjudicating it, and escalating through a ladder of remedies.
- Adaptive Stage-Gate Protocol — Allows progression only when evidence, reversibility, stakeholder legitimacy, and option-preservation criteria are satisfied.
- Audit or Attestation Record — Has an independent examiner test a commitment against a defined standard and issue a relied-upon record, turning 'trust us' into a checkable attestation.
- Interoperability and Portability Mandate — Requires the controller to expose standardized interfaces and let users take their data and connections elsewhere, so rivals can plug in and dependence on the bottleneck falls over time.
- Safe-Stop and De-escalation Trigger — A pre-set rule that halts and unwinds the sequence when ethical, legal, or collateral harm rises past the justification for continuing.
- Stakeholder Harm Reporting Channel — Lets affected parties report consequences that operators may not observe through internal telemetry.
Optimization & Search¶
Solutions that explore alternatives under objectives and constraints, prune infeasible regions, and improve a candidate toward a chosen criterion.
9 mechanisms · View full solution family
- Causal Feature Review Panel — Convenes domain experts to judge which of a model's influential features are causally or semantically meaningful and which are artifacts, proxies, or coincidences — and to name the intended structure it should be using instead.
- Group-Stratified Validation — Reports performance broken out by subgroup, source, instrument, and annotator, so a healthy-looking aggregate can't hide the slice where the shortcut has quietly failed.
- Guardrail Dashboard — Displays constraint, safety, fairness, quality, or side-effect indicators alongside the main objective score.
- Optimization Target Review — Periodically reviews whether the current objective, metric, or reward target still produces the intended outcomes under observed behavior.
- Override and Exception Log — Records when users depart from the default method, why, and whether exceptions reveal a boundary failure.
- Safety or Compliance Exclusion — Removes any candidate that crosses a safety, legal, or ethical red line — a hard, non-negotiable cut deliberately biased toward over-exclusion, with a controlled waiver as the only way back.
- Sample Audit of Exclusions — Re-examines a representative sample of what was pruned — not what was kept — to catch false negatives, bias, and drift before a filter quietly discards the answers that mattered.
- Shortcut-Risk Model Card Section — A standing section of the model's documentation that records the suspected shortcuts, what was tested, what residual risk remains, and the conditions that force revalidation.
- Stratified Benchmark Suite — Builds the test set as explicit per-regime strata — noise levels, subgroups, scales, scenario types — and reports each separately, so a method cannot win by acing the common cases while quietly failing the ones that matter.
Participation, Norms & Culture¶
Solutions that shape belonging, legitimacy, shared expectations, collective practice, and the willingness of people to contribute or comply.
10 mechanisms · View full solution family
- Appeal and Correction Workflow — Gives a subject a governed path to contest and fix reputational information that is false, irrelevant, malicious, or stale.
- Fact-Checking with Harm Awareness — Verifies a circulating claim to a defensible standard while preserving the legitimate concern underneath, so debunking does not slide into denial.
- Localization with Safeguards — Localizes language, practice, or implementation while preserving explicit minimum protections and escalation paths.
- Longitudinal Fit, Equity, and Burden Audit — Samples outcome, effort, abandonment, error, disclosure, stigma, support load, failure, repair, and disparities over time.
- Mediation-Layer Transparency Review — Exposes the stack of intermediaries — algorithms, gatekeepers, layers of process — sitting between a participant and the system, judges which add value versus only distance, and plans to make the necessary ones legible and cut the rest.
- Moderation Record with Reentry — Logs rule violations and their repair while defining the conditions under which standing is restored.
- Multimodal Equivalence and Assistive-Compatibility Test — Tests outcome, agency, timing, interoperability, safety, privacy, remedy, and effort across alternate modes and assistive tools.
- Taboo-Topic Research Safeguard — Structures research, data gathering, or documentation around taboo matters so inquiry does not sensationalize, coerce disclosure, expose participants, or flatten local meaning.
- Taxonomy or Form Revision — Executes an approved change to the actual category infrastructure — labels, form fields, database codes, rubrics, and intake questions — so a revised construction takes hold in the systems people transact through.
- Visual or Spatial Cue Redesign — Changes signs, seating, layouts, imagery, maps, interface states, or spatial defaults that silently mark some people or practices as central and others as peripheral.
Planning & Staging¶
Solutions that turn an intended outcome into phases, milestones, option points, and coordinated preparations before execution.
1 mechanism · View full solution family
- Survivorship Bias Audit — Tests whether a funnel that looks healthy among the people it measures is quietly ignoring those excluded, abandoned, refused, or dropped before they were ever counted.
Recovery & Restoration¶
Solutions that return a damaged, degraded, or interrupted system to service through repair, rollback, reentry, regeneration, or reconstruction.
2 mechanisms · View full solution family
- Right-to-Repair Interface — Grants owners and independent shops governed access to the parts, tools, and diagnostics needed to repair a product the maker doesn't service directly.
- Stakeholder Meaning Check — Tests a reading against the people it is about or for — do the affected communities recognize the meaning as theirs? — before the interpretation is acted on.
Reframing & Sensemaking¶
Solutions that change the interpretive frame, surface hidden assumptions, or organize ambiguous experience into a more useful account.
11 mechanisms · View full solution family
- Anonymous Aggregate Response — Collects responses under a rule that severs each answer from the person who gave it and reports only aggregates suppressed below a safe cell size, so a distribution can be seen without anyone being singled out.
- Expert Adjudication Panel — A convened, cross-disciplinary review body that rules on the small set of high-impact violations whose interpretation needs human judgment no filter can encode.
- Graduated Visibility Takedown Workflow — A staged response path that moves from quiet repair or limited access control to public notice or formal escalation only when thresholds are met.
- Low-Detail Policy Notice Template — A short notice that states category, authority, and appeal or context path without repeating the restricted content in promotional or searchable form.
- Model-Use Impact and Performativity Review — Tests how putting a model to use — publishing, scoring, ranking, forecasting — changes the very system it measures, and whether that feedback has quietly invalidated the model.
- Randomized Response or Privacy-Preserving Survey — Injects known random noise into each individual answer so that no single response reveals the person's true state, yet the population prevalence can still be recovered by removing the noise statistically.
- Retaliation and Re-identification Audit — Attacks its own protection like an adversary would — modeling who could infer, disclose, or punish a protected participant — and traces whether adverse actions actually followed, then orders repair.
- Ritual Drift And Harm Audit — A periodic, deliberately skeptical review that asks whether a repeated practice has drifted into rote habit, hypocrisy, coercion, exclusion, or harm — and whether its stated claim still matches lived conduct.
- Staged Identity Disclosure — Holds identity with a custodian and releases only the minimum needed fields, to only the audience that needs them, only when a defined trigger fires — so exposure tracks necessity instead of defaulting to public.
- Threshold Release Rule — Releases aggregate or representative signals only when safety, sample size, and anti-identification conditions are met.
- Trust and Legitimacy Checklist — Verifies the credibility, consent, fairness, accountability, and repair conditions an adoption needs to be seen as legitimate — as a gate before launch, not an apology after.
Representation & Modeling¶
Solutions that construct schemas, models, diagrams, abstractions, or formal descriptions that make structure available for reasoning.
10 mechanisms · View full solution family
- Bias Review Checklist — A fixed, reusable set of prompts that lets any maintainer screen a knowledge structure for hidden bias in a single pass and decide whether a deeper review is warranted.
- Cluster Label Review Workshop — Convenes domain experts to inspect candidate clusters, name them cautiously, adjudicate boundary and outlier cases, and set the terms under which the labels may be used downstream.
- Consent and Privacy Boundary Checklist — Gates whether it is legitimate to build, keep, share, and act on a model of another agent's private state — before the model is used, not after.
- Red-Team Schema Review — Assigns reviewers to attack a schema from the perspectives it is most likely to exclude, surfacing where it breaks and recording the dissent it provokes.
- Stakeholder Category Review — Convenes maintainers and affected users around concrete categories to adjudicate changes together, producing either revisions or a recorded reason they were declined.
- Stakeholder Inclusion Review — Examines whether the affected stakeholder set includes overlooked groups, boundary populations, indirect beneficiaries, or negatively affected parties.
- Stereotype Audit — Scans labels, personas, narratives, rubrics, and model variables for generalized trait claims about a group — naming the claim, pinning exactly whom it targets, examining the wording that carries it, and tracing the decision it feeds.
- Taxonomy Bias Audit — A deep, evidence-driven investigation of a single taxonomy — its labels, residual buckets, and missing distinctions — that traces how its category choices shape downstream outcomes.
- Taxonomy Redesign Workshop — A facilitated workshop that reworks a classification's categories, parent-child relations, and inclusion rules so the tree fits real and edge cases — then maps the old tree onto the new.
- Variability Analysis — Measures within-category spread, between-category overlap, subgroup differences, and interaction effects — and checks whether an apparent group difference is a measurement artifact — to show a fixed-essence claim does not fit the data.
Risk, Robustness & Uncertainty¶
Solutions that make uncertainty explicit, limit downside, preserve acceptable behavior across variation, or prepare contingencies for adverse outcomes.
1 mechanism · View full solution family
- Manual Supervision Mode — Routes actions that are normally automated through a human reviewer, so a person approves each consequential step while the system's autonomy can't be trusted.
Scaling & Capacity¶
Solutions that match capability to load, grow or shrink safely, and manage how structure and performance change with size.
1 mechanism · View full solution family
- Expert-Governed Modality Change — In safety- or human-affecting settings, requires qualified review before switching intervention modality after a plateau — so a change of approach protects the people it affects rather than merely relabeling failure.
Selection & Filtering¶
Solutions that admit, retain, rank, or reject candidates according to fitness, relevance, quality, or another discriminating rule.
5 mechanisms · View full solution family
- Adverse Adaptation Red Team — A chartered, safety-bounded exercise in which defenders imagine how an adaptive adversary would evolve to slip past the current barrier set — and whether the nominally independent layers would fall to the same move.
- False-Capture Audit — An arm's-length review that samples what the selector actually caught, sorts true target from non-target, and reports a false-capture rate the operator can't self-certify away.
- Fitness Proxy Audit — Audits what your barrier and its metrics actually reward for surviving — exposing proxies that let an escape variant look 'handled' precisely because it has become harder to see.
- Non-Target Impact Pre-Mortem — Before deployment, imagines the intervention has already caused off-target harm and works backward to name who gets caught and how — turning bycatch into a design input rather than a post-mortem finding.
- Success Metric Reweighting — Rewrites the scorecard so a bycatch term counts against success, making off-target harm subtract from the headline number instead of sitting outside it, and names who owns that term.
Stress Testing & Rehearsal¶
Solutions that expose a system or organization to controlled difficulty, adversarial conditions, or practice scenarios before real failure stakes apply.
2 mechanisms · View full solution family
- Progressive Disclosure of Risky Options — Keeps risky options out of the ordinary path and reveals them only to users who deliberately seek them out.
- Rate Limit or Cooling Hold — Caps how often or how fast an action can be repeated, and imposes a cooling delay, so impulsive or bulk misuse is blunted.
Substitution & Fallback¶
Solutions that replace unavailable or unsuitable means with alternatives while preserving the essential function, contract, or outcome.
2 mechanisms · View full solution family
- Quality Guardrail Gate — Blocks or escalates an optimization change when protected quality floors or customer, learner, patient, worker, or user outcomes degrade.
- Staged Commitment and Irreversibility-Acceptance Gate — Before exposure grows, reviews readiness, residue, burden, alternatives, horizon, consent, remedy, and fallback — then decides proceed, pause, reverse, narrow, or explicitly accept the next stage's new irreversibility.
Thresholds & Phase Change¶
Solutions that detect, create, avoid, or govern nonlinear transitions when accumulating conditions cross a consequential boundary.
4 mechanisms · View full solution family
- Associative Cue Preloading — Pre-loads the interpretive frame — language, examples, and signals — so an audience reads the eventual trigger the intended way, within an explicit consent boundary.
- Risk Score Threshold Recalibration — Moves the score boundary that routes cases to auto-approve, review, or deny when a deployed model's population or performance has drifted, keeping a human channel for contested cases.
- Stakeholder Interpretation Review — Tests an audit's inferred status reading against real affected and expert readers, so a marking is judged stigmatizing or justified by the people who live with it, not by the auditor's guess.
- Symmetry Labeling Matrix — Lays parallel cases out as rows and their describable dimensions as columns, so the cells left empty for the default case make the hidden norm visible at a glance.
Tradeoffs & Decision Support¶
Solutions that expose competing objectives, preference structure, stopping rules, and consequences so a choice can be made under constraint.
5 mechanisms · View full solution family
- Ethical Guardrail Review — A convened human review, triggered when a decision may sacrifice a protected ethical value — fairness, dignity, privacy, accountability — that no numeric threshold can adequately capture.
- Ethical Preference Inference Review — Governs whether inferring and acting on someone's revealed preferences is permissible — checking consent, the evidence's limits, and whether the use exploits rather than serves the chooser.
- Exception, Appeal, and Manual Review — Allows legitimate users to correct false denials, accessibility conflicts, credential gaps, or unusual but valid use cases.
- Open Standard or Portability Rule — Guarantees open interfaces, data portability, and exit rights so a compounding platform's participants keep the freedom to leave — bounding lock-in before the loop becomes too entrenched to govern.
- Stakeholder Weight Review Panel — Reviews proposed weights for legitimacy, impact, and acceptability.
Transmission, Propagation & Networks¶
Solutions that shape how signals, behaviors, effects, or resources spread through channels and network topology over space or time.
16 mechanisms · View full solution family
- Coarsening and Generalization Policy — Lowers the resolution of a release — coarser geography, time, categories, or numbers — until any individual hides inside a group large enough that no member stands out.
- Linkage Attack Test — Tests whether released records can be joined to outside datasets on shared quasi-identifiers to re-identify individuals or infer their protected attributes.
- Membership Inference Probe — Estimates whether a release or model reveals that a specific individual's record was in the underlying dataset — where mere presence is itself the secret.
- Metadata Minimization Filter — Strips or coarsens the incidental metadata riding along with an output — timestamps, identifiers, headers, geotags — so what's attached to the payload can't reveal the protected fact.
- Model Inversion Red Team — Has an adversarial team try to reconstruct hidden training data or attributes from a model's outputs — confidence scores, embeddings, explanations, generated text — under controlled conditions before release.
- Noise or Randomization Release — Adds calibrated random noise to outputs so they stay accurate in aggregate while no single protected input can be confidently recovered from them.
- Post-Release Reconstruction Monitor — Watches, after a release is already out, for signs that recipients or downstream tools are recombining it toward the protected originals — so protection can be revised before the risk is realized.
- Privacy Budget Accounting — Keeps a running ledger of how much reconstruction risk every query, view, and version has already spent against an explicit budget, and refuses releases once the budget would be overdrawn.
- Privacy Relay or Anonymizing Proxy — Relays a source's requests while stripping the identifying signals that would link them back, so a counterparty or observer sees the traffic but not who sent it.
- Privacy-Preserving Telemetry View — A sanitized view over internal logs, metrics, and traces that lets operators watch system health without the observability data itself becoming a channel that leaks protected state.
- Query Rate and Composition Limit — Caps how many queries an observer may make and which combinations they may compose, so a protected fact can't be reconstructed by differencing many individually-permitted answers.
- Query Rate and Overlap Limit — Caps the volume, overlap, and adaptivity of queries a recipient can make, so that no sequence of individually-safe requests can be composed into a reconstruction.
- Reaction Metric Throttling — Damps the amplification channel itself — hiding or rate-limiting the reaction counts and virality signals that let a feeling snowball faster than anyone can check it.
- Response Padding or Coarsening — Pads response size and coarsens response precision to fixed buckets, so that size and granularity — not just content — reveal nothing that distinguishes one protected state from another.
- Synthetic or Perturbed Data Validation — Tests a synthetic or perturbed release to confirm it still carries the utility it was made for and does not regenerate or memorize any real protected record.
- Threshold Suppression — Withholds any output that rests on too few underlying records — suppressing small cells so a released aggregate can't be narrowed down to expose an individual protected state.
Variation & Experimentation¶
Solutions that deliberately vary conditions, compare trials, preserve controls, and learn from differential outcomes without overclaiming.
1 mechanism · View full solution family
- Backlash Stop Rule — A pre-committed threshold that halts, dials down, or reroutes the exposure sequence to repair the moment signals show it is making the aversion worse.