{"schema_version":1,"research_id":"eoa_inverse_innovation_exp05_external_evaluation_20260803","source_assessment_id":"computability_boundary_mapping__film_media_production:P2:v0","cell_id":"computability_boundary_mapping__film_media_production","search_queries":["site:aswf.io OpenColorIO VFX color management official documentation deterministic","site:netflixtechblog.com media encoding nondeterministic output reproducibility","render pipeline equivalence verification proof carrying rewrite VFX","translation validation compiler equivalence Alive2 CompCert proof carrying optimization","site:opencolorio.readthedocs.io exact reproducible color processing cache ID official","site:openimageio.readthedocs.io image comparison idiff official documentation","site:docs.foundry.com nuke compare frames difference quality control hash cache","site:netflixtechblog.com VFX pipeline quality control pixels color metadata","official IMF media package hash integrity validation Netflix Photon QC","site:netflixtechblog.com Photon IMF validation media official","site:theiabm.org media quality control file based workflow errors cost","site:aswf.io VFX reference platform compatibility software versions production","MaterialX specification standard graph equivalence optimization transformations validation official","site:materialx.org specification node graph implementation deterministic validation","site:academysoftwarefoundation.github.io MaterialX graph optimization equivalence","site:openfx.readthedocs.io plugin deterministic render specification","site:bls.gov OEWS May 2025 software developers annual mean wage 2025","site:bls.gov occupational outlook software developers median pay 2025","site:bls.gov software developers quality assurance analysts testers median annual wage May 2024","Turing 1936 On Computable Numbers undecidable machine halting original paper PDF","Alive2 Bounded Translation Validation LLVM PLDI 2021 DOI official paper","Netflix asset QC official post delivery quality control IMF source management"],"sources":[{"source_id":"S1","title":"Welcome to Asset QC: A Guide to the Ecosystem","publisher":"Netflix Partner Help Center","url":"https://partnerhelp.netflixstudios.com/hc/en-us/articles/360057627193-Welcome-to-Asset-QC-A-Guide-to-the-Ecosystem","source_class":"COMMERCIAL_FIRST_PARTY","publication_date":"2020-12-03","accessed_at":"2026-08-03","claims_supported":["Netflix requires comprehensive technical-quality and consistency review before launch.","Asset QC preserves issue history and mediates reports, review, fixes, and communication.","Final IMF and audio deliverables receive QC, with 24- or 48-hour reporting service levels and identified production and QC roles."]},{"source_id":"S2","title":"Rendering — OpenFX 1.5.1 documentation","publisher":"OpenFX Project","url":"https://openfx.readthedocs.io/en/main/Reference/ofxRendering.html","source_class":"STANDARD","publication_date":"undated; version 1.5.1","accessed_at":"2026-08-03","claims_supported":["Real image-effect plug-ins expose thread-safety, sequential-rendering, render-window, draft-mode, GPU, and render-farm behavior that an equivalence contract must model.","Some effects depend on previous frames and require ordered execution on one instance.","Host and OpenCL-environment changes require recompilation or hash-based invalidation."]},{"source_id":"S3","title":"Comparing Images With idiff — OpenImageIO 3.1.16 documentation","publisher":"OpenImageIO Project","url":"https://openimageio.readthedocs.io/en/stable/idiff.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"undated; version 3.1.16","accessed_at":"2026-08-03","claims_supported":["Existing media tooling compares rendered images pixel by pixel and supports exact, thresholded, and perceptual criteria.","The tool can emit a difference image and distinguishes pass, warning, failure, dimension mismatch, and file error.","A comparison concerns supplied image instances; it is not a universal proof over unrendered sources, frames, parameters, or environments."]},{"source_id":"S4","title":"MaterialX Specification v1.38","publisher":"Lucasfilm Ltd. / MaterialX Project","url":"https://materialx.org/assets/MaterialX.v1.38.Spec.pdf","source_class":"STANDARD","publication_date":"2020","accessed_at":"2026-08-03","claims_supported":["A film-oriented standard already represents typed computational node graphs, compositing operators, inputs, outputs, custom nodes, and target environments.","MaterialX distinguishes universal standard nodes from application-specific customization and allows host-environment and state substitutions.","It is a plausible substrate for a restricted render-expression fragment, but the specification does not supply proof-carrying equivalence approval."]},{"source_id":"S5","title":"On Computable Numbers, with an Application to the Entscheidungsproblem","publisher":"Proceedings of the London Mathematical Society","url":"https://www.cs.miami.edu/home/burt/learning/csc427.202/Turing1936.pdf","source_class":"PRIMARY_RESEARCH","publication_date":"1936","accessed_at":"2026-08-03","claims_supported":["Turing formalized effective computation with machines and proved that no general process decides whether an arbitrary machine has specified unbounded behavior.","The source establishes the undecidable source problem needed by the proposal, but does not validate the proposal's render-graph reduction or its model assumptions."]},{"source_id":"S6","title":"Alive2: Bounded Translation Validation for LLVM","publisher":"ACM SIGPLAN PLDI / paper authors","url":"https://web.ist.utl.pt/nuno.lopes/pubs.php?id=alive2-pldi21","source_class":"PRIMARY_RESEARCH","publication_date":"2021-06","accessed_at":"2026-08-03","claims_supported":["Alive2 is an implemented bounded semantic-equivalence validator for individual compiler transformations.","Bounding loop unrolling controls resources but can miss bugs, demonstrating why bounded checking must not inherit an unrestricted guarantee.","Deployment found 47 previously unreported LLVM bugs, while semantic ambiguity required changes to the language reference."]},{"source_id":"S7","title":"CompCert C: a trustworthy compiler","publisher":"CompCert Project","url":"https://compcert.org/man/manual001.html","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2026-02-13","accessed_at":"2026-08-03","claims_supported":["CompCert demonstrates machine-checked semantic preservation for transformations over formally specified source and target languages.","Its contract permits explicit refusal instead of requiring an answer or output for every case.","Formal preservation is established relative to defined observable behavior and excludes unverified portions of the surrounding toolchain."]},{"source_id":"S8","title":"National Employment and Wage Data by Occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm?mod=article_inline","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["The May 2025 national mean annual wage was $148,100 for software developers and $111,490 for software quality-assurance analysts and testers.","These wages provide a public labor-cost anchor for 2026 resource-equivalent estimates, before benefits, overhead, specialist premiums, compute, and contingency."]}],"problem_evidence":{"support":"MODERATE","rationale":"The broader problem visibly matters: Netflix subjects final audiovisual assets to mandatory technical-consistency QC, while OpenFX documents stateful, order-sensitive, hardware-dependent, and preview-specific plug-in behavior that makes an underspecified equivalence claim unsafe. OpenImageIO confirms that instance-level pixel comparison is an established practice. However, no searched source directly documents a post-production team claiming universal graph equivalence from golden frames or treating verifier timeout as equality or inequality; prevalence and realized consequences of that exact misuse remain unmeasured.","source_ids":["S1","S2","S3"]},"stakeholder_evidence":{"support":"WEAK","rationale":"Netflix identifies production teams, QC operators, vendors, and managers who are accountable for technical consistency and fixes, making a post-production or release custodian credible in principle. No source expresses demand for proof-carrying render substitutions, identifies a facility willing to adopt them, commits funding, or confirms that equivalence-gate decisions fall within an identified supervisor's current authority.","source_ids":["S1"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"Alive2 bounded translation validation","similarity":"Very close computational analogue: it checks semantic preservation of individual optimizations, makes the semantics explicit, bounds expensive reasoning, and acknowledges missed bugs outside the bound.","remaining_difference":"It targets LLVM IR rather than audiovisual render graphs and does not provide the proposed EQUIVALENT, BOUNDED-AGREEMENT, DIFFERENT-WITNESS, and UNKNOWN production-routing record.","source_ids":["S6"]},{"name":"CompCert verified semantic preservation","similarity":"Close proof-based analogue: transformations are accepted under formal source and target semantics, with machine-checked preservation and explicit refusal possible.","remaining_difference":"It verifies a compiler and language rather than certificates attached to VFX, color, compositing, or transcoding graph substitutions.","source_ids":["S7"]},{"name":"MaterialX typed node-graph standard","similarity":"Direct domain substrate for expressing typed shading and compositing graphs, standard operators, custom nodes, targets, and environment-dependent inputs.","remaining_difference":"It validates representation and interchange, not semantic equivalence of two graphs or correctness certificates for substitutions.","source_ids":["S4"]},{"name":"OpenImageIO idiff","similarity":"Direct domain baseline for exact or thresholded comparison and concrete difference evidence with multiple non-success outcomes.","remaining_difference":"It compares supplied rendered files and cannot establish reusable universal equivalence over all sources, frames, parameters, states, and environments.","source_ids":["S3"]}],"distinctive_claim_remaining":"Conditional on a model-matched reduction and a sound restricted calculus, a film-production gate can improve claim honesty by auto-approving only certificate-backed substitutions in an enforceable render fragment while mechanically preventing finite agreement, timeout, and failed search from being reported as universal equivalence. The falsifiable contrast is fewer false universal approvals and correctly preserved UNKNOWN or bounded labels than a golden-frame/idiff workflow on the same seeded cases—not world novelty, patentability, market size, or realized impact.","confidence":"HIGH"},"implementation_evidence":{"support":"MODERATE","rationale":"The main technical pieces are credible: MaterialX supplies a domain-relevant graph representation; CompCert demonstrates machine-checked semantic preservation; Alive2 demonstrates practical transformation validation with bounded limitations; and OpenImageIO supplies differential comparators. OpenFX simultaneously shows why the unrestricted deployed model is difficult: state, ordering, host behavior, render farms, preview shortcuts, GPUs, and environment changes all affect semantics. No source establishes a sound pure media fragment, a reviewed render reduction, a checker implementation, production integration, proprietary plug-in coverage, certificate privacy controls, or acceptable artist and supervisor workload.","source_ids":["S2","S3","S4","S6","S7"]},"scores":{"meaningful_impact":{"score":3,"rationale":"Preventing an altered frame, channel, or sample and avoiding futile universal-verifier work could matter, and final-deliverable consistency is operationally important. Frequency and loss magnitude for the specific substitution error are unknown.","source_ids":["S1","S2"]},"stakeholder_pull":{"score":2,"rationale":"Externally visible QC responsibility exists, but no adopter, authorizer, or funder has requested proof-carrying render equivalence.","source_ids":["S1"]},"incremental_advantage":{"score":3,"rationale":"Certificates and explicit UNKNOWN labels offer a logically stronger guarantee than rendered-file comparison within the certified fragment. Whether useful production substitutions fall inside that fragment, and whether the gate reduces total review cost, require testing.","source_ids":["S3","S4","S6","S7"]},"distinctiveness_plausibility":{"score":3,"rationale":"No direct film-production proof-carrying equivalence gate was found, but the core architecture closely transfers established compiler verification and translation-validation practice onto an existing graph standard.","source_ids":["S4","S6","S7"]},"technical_implementability":{"score":3,"rationale":"A small pure fragment and checker are plausible; arbitrary OpenFX behavior cannot be covered honestly without severe restrictions and explicit environment contracts.","source_ids":["S2","S4","S6","S7"]},"adoption_authority_feasibility":{"score":3,"rationale":"Existing QC workflows have identifiable production, vendor, and management roles, and a non-production advisory gate need not displace creative authority. Authority over actual graph substitution and proof-rule stewardship is not externally verified.","source_ids":["S1"]},"evidence_readiness":{"score":2,"rationale":"Strong theoretical and adjacent implementation precedents exist, but there is no reviewed domain reduction, prototype result, real substitution corpus, prevalence audit, or adopter commitment.","source_ids":["S5","S6","S7"]},"safety_net_benefit":{"score":4,"rationale":"Separating certified equivalence, bounded agreement, witnessed difference, out-of-fragment input, and UNKNOWN directly reduces deceptive automation risk and supports reversible refusal.","source_ids":["S3","S6","S7"]},"scalability":{"score":2,"rationale":"Certificate checking can scale once certificates exist, but semantic formalization, rule maintenance, proprietary plug-ins, stateful effects, hardware variation, and fragment coverage are likely to require substantial per-tool and per-version work.","source_ids":["S2","S4","S7"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"50K_TO_250K","scope":"Six- to eight-week partnered, non-production study: independent reduction review, minimal semantics, small checker, 20 seeded synthetic pairs, 10 de-identified historical decisions, baseline comparator, security handling, and a written result.","confidence":"MODERATE","assumptions":["Approximately 0.5 formal-methods FTE, 1 pipeline-engineering FTE, 0.25 QA FTE, and intermittent supervisor/reviewer time for six to eight weeks.","Loaded labor is estimated above the BLS mean wages to cover benefits, overhead, specialist premiums, and contingency.","Existing non-production render infrastructure is available; no proprietary plug-in license procurement or production outage is included."],"source_ids":["S8"]},"initial_deployment_startup":{"band_2026_usd":"50K_TO_250K","scope":"Turn a successful prototype into a controlled internal service for one narrow graph fragment, including versioned semantics, certificate storage, CI, access controls, and operator documentation.","confidence":"LOW","assumptions":["One to two engineers for roughly three to six months, with part-time formal review and pipeline integration.","Scope excludes arbitrary executable plug-ins and covers one facility and one fragment.","Existing identity, artifact storage, and CI systems can be reused."],"source_ids":["S2","S4","S8"]},"operational_launch":{"band_2026_usd":"250K_TO_1M","scope":"Production hardening at one organization: checker and router integration, audit trail, monitoring, independent soundness review, incident rollback, training, security review, and validation against representative substitutions.","confidence":"LOW","assumptions":["Two to five engineering and QA resource-years including specialist review and organizational overhead.","No attempt is made to certify every commercial VFX or codec plug-in.","Compute demand remains dominated by bounded comparison and regression rendering already available to the facility."],"source_ids":["S1","S2","S8"]},"annual_recurring":{"band_2026_usd":"250K_TO_1M","scope":"Maintain semantics, approved rewrite rules, checker versions, regression fixtures, certificate invalidation, security controls, operator support, and revalidation after toolchain changes.","confidence":"LOW","assumptions":["Approximately 1.5 to 4 loaded FTE equivalents plus compute and independent review.","Every supported renderer, codec, hardware target, or plug-in version expands maintenance scope.","The gate remains limited to high-value substitutions rather than all graph edits."],"source_ids":["S2","S4","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"UNCERTAIN","reason":"Technical consistency risk, complex plug-in semantics, and bounded image comparison are externally supported, but the defining behavioral claim—teams overstating golden-frame matches or timeouts as universal equivalence—was not directly observed in the eight sources.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"UNCERTAIN","reason":"Netflix demonstrates credible production and QC actors with responsibility for final-asset consistency, but no organization or named supervisor expresses need for or authority to adopt this equivalence gate.","source_ids":["S1"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"Against the same seeded graph pairs and resource bound, test whether the proposed gate produces zero false EQUIVALENT labels, preserves UNKNOWN or BOUNDED-AGREEMENT for unresolved cases, invalidates stale certificates, and still approves useful certified rewrites more reliably than golden-frame/idiff comparison.","source_ids":["S3","S6","S7"]},"bounded_next_evidence_step":{"status":"YES","reason":"A time-bounded, non-production partner study with a fixed semantics, 20 seeded pairs, 10 historical decisions, explicit comparator, recorded metrics, and predetermined falsifiers is feasible.","source_ids":["S3","S4","S6","S7"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"A synthetic, read-only, non-production study can preserve existing release authority and prohibit graph replacement. Production authorization, proprietary-data handling, and certificate-security questions remain later gates rather than stops for this first step.","source_ids":["S1"]},"credible_cost_scope_and_range":{"status":"YES","reason":"The phase scopes and staffing assumptions are explicit, and 2025 BLS wages provide a 2026 resource-equivalent labor anchor. Confidence remains low beyond first evidence because no facility quote, compute profile, or integration estimate exists.","source_ids":["S8"]}},"next_evidence_step":"With one post-production facility, pre-register a six- to eight-week non-production study. First audit 10 de-identified historical substitution decisions to determine whether any reusable exact-equivalence claim, timeout-as-verdict, or sampled-frame overclaim actually occurred. Independently review the proposed program-to-frame-function reduction under a frozen OpenFX/MaterialX-like model, including totality, source-to-target direction, the complement step, and dependence on unbounded frame indices. Then implement a minimal checker and evaluate the stated 20 seeded pairs. Run the facility's ordinary golden-frame/OpenImageIO comparison as the comparator with the same render budget. Measure false-EQUIVALENT count, correct status-label rate, stale-certificate invalidation, counterexample yield, useful certified-approval count, runtime, and staff minutes. Falsify the problem if none of the audited decisions asserts reusable unrestricted equality or collapses timeout/incompleteness; falsify the theoretical boundary if the reviewed reduction is not valid for the frozen model; falsify safety if any seeded inequivalent pair is labeled EQUIVALENT or any semantic-version mismatch preserves approval; and falsify adoptability if no representative useful substitution fits the fragment or total staff cost exceeds the baseline without a stronger accepted guarantee.","blocking_evidence":["No direct prevalence evidence that film or post-production teams claim unrestricted exact graph equivalence from sampled frames or timeouts.","No named facility, post-production supervisor, release custodian, or funder has expressed demand or committed to a pilot.","The proposed reduction has not been independently checked against an actual render-graph representation and execution model.","No prototype evidence establishes checker soundness, deterministic behavior, late-difference handling, or certificate invalidation.","No real substitution corpus establishes that a useful share of optimizations fits the proposed pure total fragment.","Proprietary-content, plug-in intellectual-property, diagnostic-trace privacy, and certificate-access controls have not been reviewed.","Deployment and recurring costs are labor-anchored estimates without facility quotes, compute measurements, or integration estimates."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"This evaluation compared the candidate only with the eight listed direct sources. It found close cross-domain precedents in compiler semantic preservation and bounded translation validation, plus domain substrates and comparison tools, but no direct source describing the complete film-production gate. That absence is not evidence of world novelty. World novelty, patentability, freedom to operate, market size, and realized impact remain unmeasured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Obtain a named facility partner and written confirmation of the decision owner, data boundary, and non-production pilot authority.","Audit at least 10 historical substitution decisions and document whether the proposal's exact overclaim behavior is present.","Complete independent, model-specific review of every reduction obligation, including unbounded-index and complement assumptions.","Run the 20-pair seeded study against an equal-budget golden-frame/OpenImageIO comparator with zero false EQUIVALENT as a hard safety threshold.","Demonstrate at least one representative, operationally useful substitution that fits the certified fragment and quantify staff time and runtime.","Document proprietary-data, certificate-access, checker-security, rollback, and version-invalidation controls before any production inquiry."],"reason":"Bounded web research established the theoretical boundary, close prior art, plausible building blocks, a credible comparator, and labor-anchored cost ranges. The decisive remaining questions—actual problem prevalence, adopter pull, model-matched reduction validity, checker behavior, fragment usefulness, workflow burden, and production authority—require proprietary records, independent review, and live non-production testing rather than more web search."},"proposal_index":2}