{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp06_four_proposal_generalization60_20260803","cell_id":"bounded_rivalry_governance__accounting_auditing","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_id":"brg-aa-03-holdout-cost-driver-tournament","proposal_index":3,"version":0,"title":"Holdout Cost-Driver Tournament for Shared-Service Allocations","problem":"Business units receiving shared services may sponsor competing cost-allocation models because only one enterprise allocation basis will determine the next fiscal year's reported unit costs. Since lowering one unit's allocated share generally raises another's, sponsors can improve their position by choosing favorable historical periods, redefining usage boundaries, excluding inconvenient transactions, proposing opaque transformations, shifting activity to unrepresented units, or lobbying for weights that favor their operating model. Competition among models can expose alternative cost drivers and measurement weaknesses, but an unmanaged selection can reward the model that reallocates burden most effectively rather than the model that most defensibly represents resource use.","actors":["Business-unit finance teams sponsoring allocation models","Shared-service provider whose costs form the allocation pool","Corporate controller as accounting-policy owner","Managerial-accounting or FP&A staff administering the comparison","Data-governance staff maintaining source definitions and lineage","Internal audit or an independent model-validation team","Business units affected by allocations but not sponsoring a model","Executives using unit margins and costs for budgeting, pricing, and make-or-buy decisions"],"observable_state":"For a defined shared-service pool, submitted models can be compared using ledger reconciliation, driver definitions, source-data lineage, sponsor identity, allocated shares by unit, model-development hours, historical windows selected, excluded records, sensitivity results, and communications with judges. Strategic warning patterns include each sponsor's model lowering its own share relative to the baseline, unexplained exclusions concentrated in sponsor activity, synchronized assumptions that shift cost to nonparticipants, favorable windows chosen after outcomes are visible, escalating model complexity, or an incumbent controlling the only usable driver dataset.","consequence":"A selected allocation basis can distort reported unit economics, transfer costs to weakly represented units, obscure the shared service's actual consumption pattern, and influence budgets, pricing, sourcing, or performance evaluations through a measurement rule chosen partly for distributive advantage. Persistent control of the driver data or annual selection process can also make later challenges nominal rather than effective.","affected_objective":"An auditable and decision-useful managerial allocation of a reconciled shared-service cost pool, using a driver whose selection is not determined by the sponsor's preferred distribution of costs.","intervention":"Run a bounded, identity-blinded tournament among proposed cost-driver models for one shared-service pool. The stated purpose is to select an auditable representation of resource use; the scarce prize is designation as the official managerial allocation basis for one fiscal year. Any affected unit or neutral accounting team may sponsor a model if it discloses beneficiaries, supplies reproducible logic, uses governed data, and reconciles the entire pool. A frozen rulebook permits driver construction, documented assumptions, preregistered transformations, and challenges to data quality while prohibiting post-result window selection, selective transaction exclusion, hidden manual overrides, reciprocal burden shifting, off-channel judge contact, and proprietary logic that validators cannot inspect. Models first pass eligibility and ledger-reconciliation gates, then receive scores from blinded testing on withheld periods and stress scenarios using usage linkage, out-of-period stability, reconciliation integrity, auditability, data-collection burden, and sensitivity to discretionary assumptions. The model's favorable or unfavorable distribution to its sponsor is disclosed to validators but is not a scoring input. Leaders undergo independent re-performance before selection, and affected units may appeal factual or computational errors. Entry labor and model complexity are capped. Cross-submission patterns that suggest coordinated exclusions or rotating beneficiaries trigger investigation rather than an automatic verdict. A portion of implementation capacity is retained centrally to correct material allocation defects discovered during the trial year. The designation expires after one year, the driver data remains common infrastructure rather than winner-controlled property, and an ex-post review can reopen, revise, or retire the contest.","structural_mapping":[{"archetype_element":"Explicit rivalry purpose","domain_realization":"Use rivalry to compare alternative representations of shared-service consumption, not to negotiate which unit bears less cost."},{"archetype_element":"Scarce prize or selection constraint","domain_realization":"Designation as the single official managerial allocation basis for one defined cost pool and one fiscal year."},{"archetype_element":"Competitor eligibility boundary","domain_realization":"Affected units and neutral accounting teams may sponsor reproducible models after beneficiary disclosure, data-lineage documentation, and full-pool reconciliation; judges and data custodians may not compete."},{"archetype_element":"Contest arena boundary","domain_realization":"Documented driver design and data-quality challenge are allowed; selective exclusions, post-result window changes, hidden overrides, reciprocal transfers, opaque logic, and off-channel influence are prohibited."},{"archetype_element":"Performance metric and scoring basis","domain_realization":"Blinded holdout testing scores usage linkage, temporal stability, ledger reconciliation, auditability, collection burden, and sensitivity to discretionary assumptions rather than sponsor savings."},{"archetype_element":"Fair process and due process","domain_realization":"Eligibility, gates, weights, holdout-selection procedure, tie-breaks, conflict rules, evidence disclosure, and a time-boxed computational appeal are frozen before submissions are opened."},{"archetype_element":"Anti-sabotage and anti-collusion guardrail","domain_realization":"Published fouls cover data interference, selective omissions, reciprocal burden shifting, shared hidden assumptions, and judge lobbying; cross-model screens refer suspicious patterns for separate inquiry."},{"archetype_element":"Externality and spillover boundary","domain_realization":"Every model is stress-tested for unexplained transfers to non-sponsoring units, implementation burden, and downstream correction exposure, with centrally retained capacity available for remediation."},{"archetype_element":"Escalation and arms-race damper","domain_realization":"Development hours, permitted data sources, parameter count, and implementation complexity are capped or penalized so victory cannot be purchased through an opaque modeling arms race."},{"archetype_element":"Winner power and lock-in review","domain_realization":"The designation expires annually, governed driver data remains equally accessible to challengers, and the winner cannot write later rules or control validation infrastructure."},{"archetype_element":"Learning and recalibration loop","domain_realization":"Realized usage linkage, allocation volatility, corrections, collection cost, appeals, behavior changes, and control of infrastructure are reviewed before the next designation."}],"mechanism_mapping":[{"mechanism_slug":"contest_rulebook","role":"Freezes eligibility, legal model-building choices, prohibited distributive tactics, scoring, tie-breaks, conflicts, disclosure, and appeals before sponsor identities and outputs are evaluated.","counterfactual_removal":"Without a binding rulebook, administrators could adjust weights or evidence requirements after seeing which units gain, and sponsors could not distinguish legitimate modeling choices from prohibited burden shifting."},{"mechanism_slug":"ranked_leaderboard_with_audit","role":"Ranks eligible models on preregistered holdout measures and requires independent reconstruction of the leaders before the official basis is selected.","counterfactual_removal":"Without leader re-performance, a high score could depend on unreproducible code, hidden overrides, data leakage, or a selectively prepared submission."},{"mechanism_slug":"anti_collusion_monitoring","role":"Compares submissions and later rounds for common unexplained exclusions, reciprocal favorable assumptions, synchronized parameter changes, and rotating beneficiaries, referring anomalies for investigation.","counterfactual_removal":"Without pattern monitoring, sponsors could coordinate models that appear independent while jointly transferring burden to nonparticipants or alternating the favored unit."},{"mechanism_slug":"sabotage_or_foul_penalty_schedule","role":"Precommits graduated responses—correction, score adjustment, round disqualification, or temporary ineligibility—for data interference, concealed overrides, selective exclusion, reciprocal shifting, and judge contact.","counterfactual_removal":"Without stated consequences, distributive manipulation could remain a rational path to selection and enforcement could be improvised against disfavored sponsors."},{"mechanism_slug":"spending_cap_or_resource_cap","role":"Caps development labor, parameter count, specialized data acquisition, and implementation burden within declared categories.","counterfactual_removal":"Without a bounded resource envelope, sponsors could escalate consulting spend and model complexity, making rank track budget or opacity rather than measurement quality."},{"mechanism_slug":"externality_bond_or_liability_rule","role":"Retains part of the implementation allocation under controller control until trial-year validation and uses it to investigate and correct material defects or unexplained transfers.","counterfactual_removal":"Without a holdback, a sponsor could receive the full benefit of designation while affected units and central accounting absorb subsequent correction costs."},{"mechanism_slug":"challenger_access_window","role":"Expires the designation after one year, schedules reopening, and guarantees qualified challengers access to the same governed driver data and documentation.","counterfactual_removal":"Without recurring access and shared infrastructure, the initial winner could convert implementation knowledge or data control into permanent standard ownership."},{"mechanism_slug":"antitrust_or_competition_review","role":"Reviews whether the winner has acquired control over driver definitions, data collection, interfaces, or rulemaking needed by future challengers and mandates common access when necessary.","counterfactual_removal":"Without a winner-power review, a defensible initial victory could still foreclose later competition through ownership of the measurement infrastructure rather than superior model performance."},{"mechanism_slug":"post_contest_impact_review","role":"Compares the selected model's realized stability, corrections, data burden, behavior responses, allocation shifts, and infrastructure control with the contest's stated purpose.","counterfactual_removal":"Without ex-post review, a model that passed historical tests but induced new gaming or proved unstable could remain entrenched as the official basis."}],"causal_chain":["Only one allocation basis can govern a shared-service pool, and its output redistributes reported costs among units.","Affected units therefore have both an informational contribution and a distributive incentive when sponsoring competing models.","Visible historical data, discretionary definitions, and incumbent data control allow a sponsor to make its preferred cost distribution appear technically superior.","A frozen arena separates the purpose of representing resource use from the contestants' preferred allocations and defines legitimate entrants and model-building actions.","Full-pool reconciliation and governed data close routes based on omitted or privately transformed transactions.","Blinded holdout periods, stress scenarios, and independent re-performance make reproducible measurement performance—not known sponsor benefit—the declared route to winning.","Foul rules, coordination screens, complexity caps, and a remediation holdback constrain manipulation, collusion, modeling escalation, and exported correction costs.","A one-year designation preserves the needed single standard while shared infrastructure and scheduled challenge prevent the winner from owning future selection.","Ex-post comparison of realized behavior and allocations determines whether the model, scoring rule, or rivalry itself should continue."],"baseline":"The controller or finance committee selects or renews an allocation key through professional judgment and negotiation, often using visible historical examples and incumbent driver data. Affected units advocate alternatives, but there is no single frozen arena governing sponsorship conflicts, test windows, permissible transformations, comparative scoring, appeals, resource escalation, or future access.","nearest_rivals":["Controller selection of an allocation basis without competition","Negotiated cost-sharing percentages among affected business units","Direct metering or transaction tracing that removes the need to select a proxy driver","An independent consultant's model recommendation","A simple fixed driver such as headcount, revenue, or transaction volume","Separate allocation methods for each business unit or service component instead of one contested enterprise basis"],"remaining_contrastive_claim":"The proposal applies only where a single proxy allocation basis remains necessary and affected sponsors possess both useful model knowledge and distributive incentives. Its contrastive claim is that governing rivalry among measurement models requires conflict-bounded entry, permissible transformations, hidden comparative evidence, anti-coordination controls, resource limits, correction responsibility, common infrastructure, temporary designation, and reopening; central judgment, negotiation, or a technical validation check addresses only parts of that strategic selection problem.","authority_safety":{"decision_authority":"The corporate controller may authorize a retrospective shadow comparison for managerial accounting, with independent validation by internal audit or a model-risk function and data access governed by existing custodians. Any later live adoption requires the organization's normal budgeting and accounting-policy approvals.","authorized_first_step":"Select one already-closed fiscal year and one shared-service pool that does not affect statutory reporting in the pilot. Freeze the rubric, eligibility gates, allowable sources, and a withheld quarter before inspecting candidate outputs. Reconstruct the incumbent model and up to three previously proposed alternatives using existing governed data, blind sponsor identity where feasible, and compare reconciliation, stability, sensitivity, burden, and sponsor-benefit patterns. Make no ledger, budget, compensation, transfer-pricing, or performance changes.","excluded_actions":["Changing statutory, tax, regulatory, transfer-pricing, inventory, or external-reporting accounting treatments through the shadow exercise","Posting shadow allocations to any ledger or using them in budgets, pricing, compensation, or performance evaluation","Creating new charges or collecting funds from business units during the pilot","Changing holdout periods, weights, exclusions, or stress scenarios after candidate outputs are visible","Allowing a sponsor, beneficiary, incumbent data owner, or prospective winner to judge its own model","Treating a coordination screen as proof of intent or misconduct","Accessing personal, customer, or restricted operational data outside existing authorization","Penalizing sponsors or employees based on retrospective model differences","Replacing direct tracing with a proxy contest where reliable direct measurement is already feasible"],"halt_rollback":"Stop the shadow comparison if models cannot reconcile to the same pool, source definitions cannot be applied consistently, protected data would be required, sponsor blinding or validator independence materially fails, or participants begin using shadow outputs in live decisions. Quarantine the rankings, retain only authorized validation records, notify affected data and accounting owners, and continue the approved allocation basis."},"negative_tests":{"strongest_counterevidence":"The apparent rivalry may be ordinary technical disagreement: sponsors may not influence selection, their proposed models may not systematically favor themselves, and differences may arise from legitimate service heterogeneity. Direct metering or decomposition of the pool may also make a competitive proxy-selection arena unnecessary.","problem_falsifier":"The inferred problem is falsified if there is no scarce common standard, model sponsors lack distributive exposure, sponsorship does not affect selection, and favorable-window, selective-exclusion, lobbying, coordinated-transfer, complexity-escalation, or infrastructure-control patterns are absent after accounting for documented service differences.","intervention_falsifier":"The intervention is not supported for live testing if rankings are unstable across reasonable holdouts or stress scenarios, blinded validators cannot reproduce model logic, no candidate clears reconciliation and auditability gates, sponsor benefit remains the dominant explanation of rank, or direct tracing is feasible at acceptable authorized burden.","risks":["Holdout performance may not represent future service consumption.","A composite metric may conceal value judgments about stability, causality, and administrative burden.","Blinding may fail because model structure identifies its sponsoring unit.","Complexity caps may exclude a justified model for heterogeneous services.","Governed historical data may embed prior allocation choices or missing usage.","Sponsors may shift operational behavior after selection to improve future allocations.","A winning proxy may create false precision despite weak causal linkage.","Distributional stress tests may be mistaken for proof that a large but justified allocation change is unfair.","The remediation holdback may create incentives to underreport defects.","Annual reopening may cause policy churn or discourage investment in measurement infrastructure.","Incumbent data custodians may retain practical control despite formal challenger access.","Using a managerial model beyond its authorized purpose could affect statutory or personnel decisions." ]},"next_evidence_step":"Pre-register and run only the authorized closed-year shadow tournament. Record gate failures, ranking sensitivity, independent reproducibility, sponsor-benefit correlations, unexplained transfers, data and labor burden, and whether direct tracing is feasible. The bounded decision is whether one prospective parallel-ledger simulation is warranted; it is not authorization to change an allocation basis.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Proposal 1 governs business-unit rivalry for scarce consolidation-review slots, where the strategic object is apparent close readiness and the intervention is a recurring verified queue. Proposal 2 governs rivalry among internal auditors for scarce engagement-lead mandates, where finding production and evidence behavior are tested through a replicable assurance challenge. This proposal instead addresses zero-sum sponsorship of managerial cost-allocation models for one official measurement standard. Its intervention is a blinded holdout tournament among accounting models with full-pool reconciliation, distributive-conflict disclosure, common data infrastructure, and a time-limited policy designation. It concerns neither close workflow priority nor auditor career opportunity; it has different competitors, evidence, prize, externalities, causal path, decision authority, and adoption boundary from both earlier proposals and can be adopted independently of them.","revision_record":{"parent_version":null,"progress_targets_addressed":["Produced a complete proposal-index-3 candidate grounded only in the supplied archetype, authored mechanisms, domain card, and earlier sealed proposals.","Addressed a materially different accounting problem centered on rival sponsorship of a single cost-allocation standard.","Specified purpose, scarcity, eligibility, allowed actions, holdout scoring, due process, anti-abuse controls, spillover responsibility, escalation limits, winner access, recalibration, authority, safeguards, falsifiers, and bounded evidence.","Explained diversity separately from proposals 1 and 2 without asserting novelty, prevalence, demand, or effect size."],"conceptual_changes":["Initial version; introduced a measurement-standard selection problem rather than a review-queue or audit-lead allocation problem."],"operational_changes":["Initial version; defined a closed-year blinded model tournament, reconciliation gates, holdout testing, conflict disclosure, resource caps, shared driver infrastructure, annual reopening, and rollback."],"evidence_changes":["Initial version; limited initial evidence to governed historical data and a nonposting shadow comparison; prior art remains unsearched."],"claim_changes":["Initial version; retained only the conditional claim that bounded rivalry is relevant when sponsor knowledge and distributive incentives coexist around a necessary single allocation basis."]}}