{"schema_version":1,"experiment_id":"eoa_inverse_innovation_exp09_archetype_breadth150_20260804","cell_id":"prediction_error_learning_calibration__accounting_auditing","arm":"BREADTH_PROBE_ONE_SHOT","candidate_id":"prediction_error_learning_calibration__accounting_auditing__P1","proposal_index":1,"version":0,"title":"Expectation-Adjusted Learning from Accounts-Payable Control Exceptions","problem":"An internal audit function revises risk ratings and future testing for an accounts-payable three-way-match override control from raw exception counts. High-volume business units repeatedly attract attention because they produce more exceptions, while a small but unexpectedly deteriorating unit may receive no additional scrutiny because its absolute count remains low. The workpapers do not preserve what exception rate or severity the auditors expected before each test window, so routine outcomes, genuine deterioration, transaction-mix changes, and sampling noise cannot be distinguished when the audit plan is updated.","actors":["Internal audit engagement lead","Internal auditors performing recurring control tests","Accounts-payable process owner","Business-unit control owners","Chief audit executive"],"observable_state":"Across recurring test windows, workpapers contain sampled transactions, observed override exceptions, and revised unit risk ratings, but no prospectively timestamped unit-level prediction of exception probability or severity. Changes in testing effort correlate with raw exception totals or threshold crossings even when transaction volume, transaction mix, and sample uncertainty differ across units.","consequence":"Audit effort can be redirected toward expected high-volume exceptions while an informative increase above a low-risk unit's expectation is missed; alternatively, a noisy cluster can trigger unnecessary expansion. The resulting risk ratings and test allocations may reflect volume and salience more than evidence of changed control performance.","affected_objective":"Calibrated allocation of internal-audit testing that remains sensitive to material control deterioration without treating every observed exception as equally informative.","intervention":"Add a shadow prediction-error calibration layer to the recurring control test. Before each test window, the engagement team records each unit's predicted exception probability and predicted severity distribution, the assumptions supporting those predictions, and the intended update target. After testing, it records realized results, computes signed deviations from the predictions, checks sampling uncertainty, transaction mix, measurement reliability, and possible control or process changes, and assigns the deviation to the most plausible cause. Only reliable residual deviations influence a bounded recommendation for the next window's risk rating or sample emphasis. Severe exceptions continue through existing escalation rules regardless of whether they were expected.","structural_mapping":[{"archetype_element":"Prior Prediction Record","domain_realization":"A timestamped pre-test estimate of override-exception probability and severity for each business unit and test window."},{"archetype_element":"Value Reference Frame","domain_realization":"Control-failure exposure, represented separately by exception incidence, financial magnitude, policy significance, and indicators of fraud or regulatory risk."},{"archetype_element":"Received Outcome Record","domain_realization":"The observed exceptions and non-exceptions, with sample size, transaction population, unit, date, amount, override reason, and evidence quality."},{"archetype_element":"Signed Error Signal","domain_realization":"Observed exception incidence or severity minus its pre-recorded expectation, retaining whether control performance was better or worse than expected."},{"archetype_element":"Credit Assignment Window","domain_realization":"A documented assessment of whether a deviation is attributable to the control, process execution, system configuration, transaction mix, sampling, measurement error, or a changed operating condition."},{"archetype_element":"Noise and Volatility Filter","domain_realization":"Confidence bounds, minimum evidence requirements, cross-window persistence checks, and review of population or sampling changes before an error affects planning."},{"archetype_element":"Learning Gain Rule","domain_realization":"A bounded rule that gives greater planning weight to reliable, repeated, or safety-relevant deviations and less weight to sparse or unstable observations."},{"archetype_element":"Update Target","domain_realization":"The unit-level control-risk estimate and proposed emphasis of the next scheduled internal-audit test, not the control owner's performance evaluation."},{"archetype_element":"Shortcut Learning Guard","domain_realization":"A check that recommendations are not driven merely by business-unit transaction volume, exception dollar size, auditor familiarity, or ease of obtaining evidence."},{"archetype_element":"Ethical Reward Safety Review","domain_realization":"A review that prohibits using prediction errors as employee incentives or blame scores and preserves confidentiality, due process, and existing escalation duties."}],"mechanism_mapping":[{"mechanism_slug":"prediction_outcome_delta_log","role":"Captures each prospective exception prediction beside the realized result, signed gap, context, credit assessment, and proposed update.","counterfactual_removal":"Without the log, expectations can be rewritten after outcomes are known, and the team reverts to interpreting raw exception counts."},{"mechanism_slug":"calibration_curve_review","role":"Periodically compares predicted exception probabilities with observed frequencies across comparable units and windows to identify systematic over- or underprediction.","counterfactual_removal":"Without calibration review, signed errors may accumulate without revealing that the forecasting process itself is biased or poorly scaled."},{"mechanism_slug":"negative_prediction_error_review","role":"Requires diagnostic review when observed control exposure is reliably worse than expected, while preserving immediate escalation for severe findings.","counterfactual_removal":"Without this review, small absolute counts that materially exceed expectation can remain hidden below raw-count thresholds."},{"mechanism_slug":"shortcut_probe_holdout_set","role":"Recalculates recommendations on a held-out, volume-balanced subset to test whether the update is learning control deterioration rather than transaction volume or sample composition.","counterfactual_removal":"Without the probe, apparent learning can be produced by the incidental fact that high-volume units generate more visible exceptions."}],"causal_chain":["Auditors prospectively record unit-level exception and severity expectations before seeing the test results.","Testing produces received outcomes with sample, population, timing, and transaction-mix context.","The delta log preserves the signed gap between predicted and observed control exposure.","Credit and noise checks separate plausible control deterioration from sampling variation, mix shifts, measurement problems, and one-off events.","A bounded learning-gain rule converts only the reliable residual gap into a proposed change to the risk estimate or next-window test emphasis.","Calibration and shortcut checks determine whether later predictions improve without merely tracking volume or other incidental cues.","Existing materiality, fraud, regulatory, and severe-exception escalation rules remain independent safety backstops."],"baseline":"The engagement team records raw exception counts and severity, compares them with fixed thresholds or prior-period totals, and uses professional judgment to revise risk ratings and sample plans without a timestamped pre-outcome expectation or an explicit signed-error, credit-assignment, and noise-filtering step.","nearest_rivals":["Risk-based audit planning: prioritizes areas using assessed risk but does not necessarily make each recurring test's prior expectation explicit or govern learning by the signed prediction-outcome gap.","Continuous-auditing exception dashboards: surface raw anomalies or threshold breaches but can preserve volume bias unless exceptions are interpreted relative to prospectively recorded expectations.","Statistical audit sampling and stratification: controls inference from a sampled population but does not by itself specify how surprising results should update future audit practice.","After-action review: can explain completed work but permits hindsight reconstruction unless the prediction was recorded before the outcome."],"remaining_contrastive_claim":"The candidate's distinguishing claim is limited to the update rule: future control-risk estimates and test emphasis are informed by reliable signed deviations from prospectively recorded expectations after credit and noise checks, rather than by raw exception magnitude, fixed threshold crossing, or retrospective explanation alone.","authority_safety":{"decision_authority":"The chief audit executive may authorize the shadow methodology; the engagement lead may record predictions and recommendations but may not unilaterally reduce required testing or suppress findings. Existing governance determines any later audit-plan change.","authorized_first_step":"Run a non-decisional shadow log alongside one already scheduled recurring test of the accounts-payable override control; retain the approved scope, sampling plan, reporting thresholds, and escalation process.","excluded_actions":["Reducing or canceling required audit procedures based on pilot scores","Changing management controls or accounting records","Using prediction errors to rank, reward, discipline, or publicly identify employees","Delaying or suppressing material, fraud-related, regulatory, or safety-relevant findings","Representing the shadow output as external-audit evidence or an assurance opinion","Allowing process owners to revise predictions after test outcomes are visible"],"halt_rollback":"Stop the shadow method and revert to the approved audit procedure if predictions are backfilled, protected information is exposed, the log influences findings outside its authority, severe exceptions are attenuated, or participants use it for blame. Preserve the original workpapers and decision trail so all pilot-derived recommendations can be disregarded."},"negative_tests":{"strongest_counterevidence":"Across comparable test windows, raw exception severity and statistically adjusted exception rates support later control assessments at least as consistently as the signed-error method, while the added prediction layer supplies no stable information and mainly introduces forecast gaming or anchoring.","problem_falsifier":"Prospective review shows that the team already records usable pre-test expectations and bases updates on expectation-adjusted, uncertainty-aware deviations, or that future test allocation is not influenced by raw outcomes in the alleged way.","intervention_falsifier":"Pre-recorded predictions are too unstable or manipulable to form a usable reference, credit assignment cannot distinguish control changes from mix and sampling effects, or shadow recommendations are less reproducible and no better calibrated against subsequent windows than the baseline.","risks":["Auditors or process owners may sandbag predictions to avoid negative deviations.","Numeric predictions may create false precision or anchor professional judgment.","A narrow value frame may underweight qualitative, fraud, compliance, or reputational significance.","Noise filters may delay response to a rare but consequential exception.","Small samples may produce unstable signed errors.","Credit assignment may become blame directed at the nearest visible employee.","The method may learn transaction volume, evidence availability, or auditor familiarity as shortcuts.","Added logging may consume audit capacity without improving decisions."]},"next_evidence_step":"During one already scheduled six-week internal-audit test of the accounts-payable override control, prospectively log predictions for no more than four business units before samples are opened. Keep all official procedures unchanged. After outcomes are known, have two auditors independently compute signed deviations, classify credit and noise, and make shadow next-window recommendations. Compare inter-rater agreement, hindsight-edit attempts, volume sensitivity, and whether the recommendations correctly order the units' prospectively observed results in the immediately following scheduled test window. Do not use the pilot to change scope or issue assurance.","prior_art_status":"UNSEARCHED","diversity_from_prior_proposals":"Not evaluated against other proposals under runtime isolation; the candidate is derived solely from the supplied archetype and domain card and is instantiated as prospective, unit-level learning from recurring accounts-payable control tests.","revision_record":{"parent_version":null,"progress_targets_addressed":[],"conceptual_changes":[],"operational_changes":[],"evidence_changes":[],"claim_changes":[]}}