{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"invariant_mode_decomposition_design__information_theory:P2:v0","cell_id":"invariant_mode_decomposition_design__information_theory","search_queries":["site:nist.gov SP 800-226 differential privacy final 2025 composition correlated queries","site:census.gov reconstruction re-identification attack 2010 census data official","matrix mechanism differential privacy correlated queries SVD workload primary paper","low rank mechanism differential privacy workload SVD VLDB 2012 paper","privacy funnel data disclosure channel sensitive attributes utility optimization primary paper","linear privacy filter SVD sensitive attribute data release utility primary research","joint disclosure risk multiple aggregate statistics linear combinations reconstruction attack paper","site:opendp.org workload-aware correlated noise matrix mechanism documentation","Dinur Nissim Revealing information while preserving privacy PODS 2003 PDF","The matrix mechanism optimizing linear counting queries differential privacy official journal PDF 2015","From information bottleneck to privacy funnel 2014 paper PDF official","MaSS multi-attribute selective suppression ICML 2024 paper"],"sources":[{"source_id":"S1","title":"A Simulated Reconstruction and Reidentification Attack on the 2010 U.S. Census: Full Technical Report","publisher":"U.S. Census Bureau","url":"https://www.census.gov/library/working-papers/2023/adrm/CES-WP-23-63.html","source_class":"PRIMARY_RESEARCH","publication_date":"2023-12","accessed_at":"2026-08-03","claims_supported":["Bundles of aggregate tables can enable reconstruction that is not apparent from individual statistics.","Using 34 published table sets, the study reconstructed five-variable microdata and reported perfect reconstruction for all records in 70% of census blocks.","The study reported correct race-and-ethnicity inference for 3.4 million vulnerable population uniques at 95% accuracy.","The 2020 disclosure framework defended against the tested reconstruction attacks, demonstrating the value of stronger release-wide controls."]},{"source_id":"S2","title":"NIST SP 800-226: Guidelines for Evaluating Differential Privacy Guarantees","publisher":"National Institute of Standards and Technology","url":"https://csrc.nist.gov/pubs/sp/800/226/final","source_class":"OFFICIAL_GUIDANCE","publication_date":"2025-03-06","accessed_at":"2026-08-03","claims_supported":["Differential privacy is a mathematical framework for quantifying privacy loss when an entity's data appears in a dataset.","Practitioners need implementation-level evaluation because privacy hazards can arise when mathematical guarantees are realized in software.","A modal audit cannot substitute for evaluation of the complete formal differential-privacy implementation."]},{"source_id":"S3","title":"Disclosure Avoidance: Latest Frequently Asked Questions","publisher":"U.S. Census Bureau","url":"https://www.census.gov/programs-surveys/decennial-census/decade/2020/planning-management/process/disclosure-avoidance/2020-das-updates/2020-das-faqs.html","source_class":"OFFICIAL_GUIDANCE","publication_date":"2024-01-24","accessed_at":"2026-08-03","claims_supported":["The Census Bureau adopted differential privacy to guard against reconstruction and reidentification of released census data.","The Bureau expressly identifies a dual mandate to produce useful statistics while protecting respondent confidentiality.","The Bureau is an identifiable high-authority adopter of release-wide privacy controls, although it has not expressed demand for this particular spectral method.","Operational release systems preserve mandated invariants and use carefully calibrated noise, illustrating governance constraints on any additional filter."]},{"source_id":"S4","title":"Revealing Information while Preserving Privacy","publisher":"Association for Computing Machinery","url":"https://systems.cs.columbia.edu/private-systems-class/papers/DinurNissim2003Revealing.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2003-06-09","accessed_at":"2026-08-03","claims_supported":["Multiple noisy subset-sum answers can collectively encode enough information for database reconstruction.","Theoretical reconstruction results establish that apparently innocuous statistical answers cannot be assessed safely only in isolation.","Output perturbation, query restriction, and auditing are longstanding disclosure-control approaches rather than novel categories."]},{"source_id":"S5","title":"The Matrix Mechanism: Optimizing Linear Counting Queries under Differential Privacy","publisher":"The VLDB Journal / Springer","url":"https://people.cs.umass.edu/~miklau/assets/pubs/dp/Li15matrix.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2015-08-18","accessed_at":"2026-08-03","claims_supported":["The matrix mechanism optimizes a workload of linear counting queries jointly rather than treating each query independently.","It derives workload answers from a strategy workload and can induce complex correlated noise while preserving differential privacy.","Workload-specific joint optimization and correlated-noise allocation substantially overlap the candidate's proposed control layer."]},{"source_id":"S6","title":"Low-Rank Mechanism: Optimizing Batch Queries under Differential Privacy","publisher":"Proceedings of the VLDB Endowment","url":"https://vldb.org/pvldb/vol5/p1352_ganzhaoyuan_vldb2012.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"2012-08-27","accessed_at":"2026-08-03","claims_supported":["Correlated query batches can be processed jointly for better accuracy than independent-query mechanisms.","The Low-Rank Mechanism factors a workload matrix and uses its low-rank structure to allocate privacy-preserving noise.","Experiments reported large accuracy improvements over then-current batch-query methods, making low-rank spectral workload treatment close prior art."]},{"source_id":"S7","title":"From the Information Bottleneck to the Privacy Funnel","publisher":"IEEE Information Theory Workshop / arXiv","url":"https://arxiv.org/abs/1402.1774","source_class":"PRIMARY_RESEARCH","publication_date":"2014-09-30","accessed_at":"2026-08-03","claims_supported":["A data-release mapping can be optimized to reduce inference about correlated private data subject to a utility constraint.","Mutual information provides an inference-oriented privacy metric, with a bound relating it to gains under bounded inference costs.","Direct privacy-utility optimization over a disclosure channel predates the candidate and is a stronger general rival than singular-value ranking alone."]},{"source_id":"S8","title":"MaSS: Multi-attribute Selective Suppression for Utility-preserving Data Transformation from an Information-theoretic Perspective","publisher":"Proceedings of Machine Learning Research","url":"https://proceedings.mlr.press/v235/chen24f.html","source_class":"PRIMARY_RESEARCH","publication_date":"2024-07-21","accessed_at":"2026-08-03","claims_supported":["Learned transformations can selectively suppress multiple sensitive attributes while preserving useful attributes.","The method supplies recent primary evidence that targeted, utility-preserving sensitive-attribute suppression is technically implementable.","Selective suppression is adjacent prior art, although MaSS addresses learned data representations rather than a fixed bundle of aggregate releases under unchanged formal privacy accounting."]}],"problem_evidence":{"support":"STRONG","rationale":"The problem is visible both theoretically and operationally. Dinur and Nissim show that collections of noisy statistical answers can permit reconstruction, while the Census Bureau demonstrated reconstruction and protected-attribute inference from published aggregate tables at national scale. NIST and Census materials confirm that complete-release privacy hazards and privacy-utility tradeoffs matter to real release authorities. The evidence does not establish prevalence across ordinary analytics services or show that per-output controls are the specific cause in most deployments.","source_ids":["S1","S2","S3","S4"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"The U.S. Census Bureau is an identifiable adopter and authorizer with an expressly stated need to protect confidentiality while retaining statistical quality, and NIST addresses practitioners evaluating differential-privacy software. No source shows that the Bureau, another named service owner, or a funder wants the proposed SVD-based audit specifically; candidate-specific pull and access to an eligible workload remain unverified.","source_ids":["S2","S3"]},"prior_art":{"proximity":"SUBSTANTIAL_COLLISION","closest_analogues":[{"name":"Matrix Mechanism","similarity":"Jointly models a linear query workload and chooses a workload-adapted strategy that yields correlated noise under a formal differential-privacy guarantee.","remaining_difference":"It optimizes query-answer error from the workload matrix rather than fitting a protected-source-to-output perturbation operator, ranking modes by held-out attribute-inference consequence, and monitoring modal drift.","source_ids":["S5"]},{"name":"Low-Rank Mechanism","similarity":"Uses low-rank matrix factorization to exploit correlations across a batch of releases and improve utility under fixed differential privacy.","remaining_difference":"Its factorization targets workload accuracy and privacy sensitivity, not a locally estimated protected-contrast disclosure Jacobian with residual inference tests and spectral-gap withdrawal rules.","source_ids":["S6"]},{"name":"Privacy Funnel","similarity":"Treats disclosure as a channel and directly optimizes inference privacy against utility for correlated private and disclosed variables.","remaining_difference":"It is a general probabilistic privacy mapping optimized by mutual information, not an interpretable SVD-based audit-and-control layer constrained to preserve an existing release mechanism's formal settings.","source_ids":["S7"]},{"name":"MaSS selective suppression","similarity":"Selectively removes information about multiple sensitive attributes while attempting to preserve useful information.","remaining_difference":"It learns transformations for high-dimensional representations and does not specifically address aggregate-query bundles, mandatory disclosure rules, local source-to-release modes, or formal privacy-accounting preservation.","source_ids":["S8"]},{"name":"Census Disclosure Avoidance System","similarity":"An operational release-wide system designed to resist reconstruction while balancing confidentiality and statistical utility.","remaining_difference":"The official description supports calibrated noise and invariants but does not document the candidate's protected-source SVD, consequence-ranked modal controls, or spectral-gap monitoring.","source_ids":["S1","S3"]}],"distinctive_claim_remaining":"On a frozen multi-output release workload and unchanged formal privacy settings, a filter selected from an empirically estimated protected-source-to-release Jacobian, ranked by held-out disclosure sensitivity rather than singular value or workload error alone, will reduce a predeclared joint protected-attribute inference metric more than independent noise, query suppression, the matrix/low-rank mechanism, and a direct privacy-utility optimizer at the same utility budget; the advantage must persist across admissible data splits and perturbations and disappear when residual disclosure or modal-instability falsifiers fire.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The mathematical and software primitives are established: joint workload matrices, low-rank factorizations, correlated noise, privacy-utility optimization, and targeted sensitive-attribute suppression all have primary research support. A bounded offline implementation is technically credible. Evidence is absent for stable estimation of the candidate's local operator on a real aggregate-release service, safe generation of protected-source perturbations, reliable interpretation under near-degenerate singular values, compatibility with a specific formal privacy accountant, secure handling of mode artifacts, and superiority to the strongest optimizers. Legal authority is deployment-specific; only synthetic or properly authorized de-identified sandbox work is presently supported.","source_ids":["S2","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":4,"rationale":"Joint-release failures can expose protected attributes at large scale, as the Census reconstruction study demonstrates, although impact for the proposed target service is not quantified.","source_ids":["S1","S4"]},"stakeholder_pull":{"score":3,"rationale":"Census and NIST establish strong institutional need for usable privacy protection, but no named organization requests this spectral layer.","source_ids":["S2","S3"]},"incremental_advantage":{"score":2,"rationale":"Protected-contrast sensitivity ranking, residual inference testing, and drift withdrawal could add interpretability, but matrix/low-rank mechanisms and privacy-funnel optimizers already address most of the privacy-utility allocation problem.","source_ids":["S5","S6","S7","S8"]},"distinctiveness_plausibility":{"score":2,"rationale":"The exact governed combination was not found, but its core operations substantially collide with established spectral workload optimization and selective privacy mappings; world novelty remains unmeasured.","source_ids":["S5","S6","S7","S8"]},"technical_implementability":{"score":4,"rationale":"SVD, low-rank optimization, correlated noise, and selective transformations are well demonstrated. Reliability of a fitted local disclosure operator is the principal unresolved technical issue.","source_ids":["S5","S6","S8"]},"adoption_authority_feasibility":{"score":3,"rationale":"A data steward or privacy officer can authorize a sandbox audit, but live filtering must preserve mandated rules, invariants, formal accounting, and organization-specific legal authority.","source_ids":["S2","S3"]},"evidence_readiness":{"score":2,"rationale":"The problem and adjacent methods are well supported, but the incremental claim has no head-to-head evidence and requires access to a governed workload and held-out evaluation data.","source_ids":["S1","S5","S6","S7"]},"safety_net_benefit":{"score":3,"rationale":"Keeping formal privacy and suppression rules as hard constraints plus rollback limits downside, but a local audit can be misread as a guarantee and its mode artifacts may themselves be sensitive.","source_ids":["S2","S3"]},"scalability":{"score":3,"rationale":"Matrix and low-rank workload methods scale beyond coordinate-wise review, but repeated operator estimation, adversarial evaluation, and drift monitoring may be costly for changing or very large release bundles.","source_ids":["S5","S6"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"One frozen workload; threat-model and metric specification; secure synthetic or authorized de-identified evaluation data; operator estimation; four comparators; held-out inference, utility, residual, conditioning, and stability analysis.","confidence":"LOW","assumptions":["Approximately 8–20 person-weeks across privacy engineering, statistics or ML, data stewardship, and security review.","An existing sandbox, formal privacy implementation, and query workload are available.","No live publication, procurement, or new protected-data collection is included.","The eight sources establish technical scope but provide no 2026 labor or infrastructure prices."] ,"source_ids":["S2","S5","S6","S7"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Production-quality filter implementation, privacy-accountant compatibility testing, secure artifact storage, audit logging, monitoring integration, analyst documentation, governance and legal review, and shadow deployment for one service.","confidence":"LOW","assumptions":["One established analytics service with accessible release code and test infrastructure.","No weakening of existing privacy parameters, suppression rules, or access controls.","Includes engineering hardening and independent privacy/security validation but not organization-wide rollout.","Source literature supports feasibility but not commercial deployment costs."] ,"source_ids":["S2","S3","S5","S6"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Controlled launch for one governed release family, including parallel baseline operation, incident and rollback procedures, analyst acceptance testing, documentation, and authorization evidence.","confidence":"LOW","assumptions":["Launch follows a successful shadow evaluation and written approval by the release authority.","Existing secure compute and privacy-accounting systems are reused.","Public release is prohibited until all formal guarantees, utility floors, and residual-disclosure tests pass.","No liability reserve or large-scale data remediation is included."] ,"source_ids":["S2","S3"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Scheduled mode and gain re-estimation, residual and drift monitoring, periodic adversarial evaluation, governance review after workload changes, restricted artifact management, and incident response readiness for one release family.","confidence":"LOW","assumptions":["Quarterly review plus event-triggered refitting after schema, model, query, distribution, or policy change.","Compute demand is moderate relative to the host analytics pipeline.","One part-time privacy engineer or equivalent cross-functional allocation maintains the system.","No evidence source directly benchmarks recurring cost."] ,"source_ids":["S2","S3","S5","S6"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Primary theory and the Census reconstruction study directly establish that collections of individually released statistics can jointly reveal sensitive records or attributes.","source_ids":["S1","S4"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"The Census Bureau is an identifiable release authority that expressly balances confidentiality with statistical quality and adopted release-wide formal privacy protection. This verifies the adopter class and need, not interest in the candidate.","source_ids":["S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The remaining claim specifies a protected-source Jacobian, consequence-ranked modal filter, matched privacy and utility settings, four comparators, held-out outcomes, stability requirements, and explicit falsifiers.","source_ids":["S5","S6","S7","S8"]},"bounded_next_evidence_step":{"status":"YES","reason":"A single frozen workload can be evaluated offline with fixed data partitions, perturbation bounds, metrics, privacy settings, comparators, and stop rules without modifying a live release.","source_ids":["S2","S5","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The next step is limited to a restricted sandbox using synthetic or specifically authorized de-identified records, retains all mandatory controls, and cannot authorize publication. Live deployment remains outside this gate and requires written governance and legal approval.","source_ids":["S2","S3"]},"credible_cost_scope_and_range":{"status":"UNCERTAIN","reason":"The scopes and resource assumptions are bounded, but none of the eight direct sources provides 2026 labor, infrastructure, assurance, or integration cost benchmarks; all bands therefore have low confidence.","source_ids":[]}},"next_evidence_step":"With a willing data steward, pre-register one frozen release workload: source and output schemas, admissible perturbations, threat model, joint attribute-inference metric, utility floors, unchanged formal privacy parameters, partitions, and halt rules. In a restricted sandbox, estimate the local operator on training data and compare (1) the unchanged release, (2) matched independent-noise allocation, (3) query-level suppression, (4) a matrix or low-rank workload mechanism, (5) a direct privacy-funnel or adversarial privacy-utility optimizer, and (6) the proposed disclosure-sensitivity-ranked modal filter. On untouched records and perturbations, report confidence intervals for disclosure and utility, residual protected predictability, singular-subspace stability rather than individual vectors when values are near-degenerate, conditioning, and failure after simulated workload drift. Falsify the candidate if it does not materially improve the preregistered disclosure criterion at matched utility and formal settings, if a strongest rival is noninferior, if the selected subspace is unstable across admissible samples, or if residual disclosure, mandatory-control, or utility checks fail. A positive result permits only a separately authorized shadow run.","blocking_evidence":["No named release owner has expressed candidate-specific demand, offered a workload, or committed authorization and staff.","No held-out head-to-head result compares the candidate with matrix/low-rank workload mechanisms and direct privacy-utility optimization under matched formal privacy and utility settings.","No evidence establishes that the protected-source-to-release operator or selected singular subspace is stable across admissible samples, perturbations, nonlinear thresholds, and workload changes.","Compatibility with a specific production privacy accountant, suppression policy, invariant set, legal regime, and artifact-handling policy is untested.","The prevalence of the proposed failure mode outside the Census example is unknown.","Cost bands lack external 2026 price benchmarks.","World novelty, patentability, freedom to operate, market size, and realized impact are unmeasured."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The search found substantial prior art for reconstruction from joint statistics, workload-wide correlated noise, low-rank or spectral query optimization, privacy-channel optimization, and selective suppression of sensitive attributes. It did not establish whether the precise combination of an empirically estimated protected-source-to-release Jacobian, disclosure-consequence modal ranking, residual inference testing, and spectral-drift withdrawal under unchanged formal privacy controls has appeared previously. No conclusion is made about world novelty, patentability, freedom to operate, market size, or realized impact.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":false,"progress_targets":["Obtain written participation and sandbox authority from one named data steward or release-governance officer.","Pre-register the protected-source basis, threat model, disclosure metric, utility floors, formal privacy settings, perturbation range, statistical analysis, and halt rules.","Implement matched independent-noise, query-suppression, matrix/low-rank, and direct privacy-utility comparators rather than comparing only with the weak coordinate-wise baseline.","Demonstrate a held-out disclosure reduction at matched utility and formal privacy settings, with uncertainty intervals and a prespecified materiality threshold.","Show singular-subspace stability across admissible splits and perturbations, and report failure under near-degeneracy, nonlinearity, and simulated drift.","Verify that residuals do not retain consequential protected inference and that audit artifacts can be handled under the applicable governance and legal controls.","Replace low-confidence cost assumptions with measured engineering time, secure-compute usage, assurance effort, and monitoring load from the sandbox study."],"reason":"Bounded web research verifies the problem and reveals substantial collision with matrix mechanisms, low-rank workload optimization, privacy funnels, and selective suppression. The only defensible remaining advantage is empirical and workload-specific: whether consequence-ranked protected-source modes outperform those strong rivals while remaining stable and compatible with formal privacy controls. Resolving that claim requires governed data, pipeline access, and live or shadow computation rather than more public-web research; under the controller rule this requires an empirical-research stop, which is non-repairable."},"proposal_index":2}