{"schema_version":1,"research_id":"eoa_inverse_innovation_exp06_external_evaluation_20260803","source_assessment_id":"predictive_residual_processing__religious_studies_theology:P4:v0","cell_id":"predictive_residual_processing__religious_studies_theology","search_queries":["site:eric.ed.gov university instructors grading feedback workload formative assessment study","religious studies assessment feedback rubric interpretive essays workload instructors","AI assisted grading group similar answers rubric Gradescope official","Bayesian knowledge tracing student model predicts next response formative assessment primary paper","higher education assessment feedback workload marking time primary study instructors","site:advance-he.ac.uk feedback workload assessment higher education marking","site:educause.edu faculty grading workload feedback survey","religious studies assessment valid alternative interpretations rubric higher education","primary study automated essay scoring open ended higher education rubric bias validity human graders 2024","large language models grading essays rubric agreement bias primary study 2025","learning analytics student model humanities writing prediction longitudinal primary research","automated formative feedback essays human oversight student challenge research","site:studentprivacy.ed.gov FERPA education technology student data official school officials legitimate educational interest","site:ed.gov AI toolkit education human oversight privacy assessment 2024","site:unesco.org generative AI education human agency data privacy assessment guidance","site:nist.gov AI RMF education bias human oversight official","site:bls.gov Occupational Employment and Wage Statistics software developers May 2025 postsecondary teachers wages","site:bls.gov Occupational Outlook Handbook software developers median pay 2025 postsecondary teachers","2026 university instructional designer salary software developer hourly cost official","site:aqa.org.uk religious studies assessment objective analyse evaluate arguments evidence official specification","site:ibo.org religious studies subject brief assessment interpretation official","site:cambridgeinternational.org religious studies syllabus assessment objectives interpretation evidence official","\"Interpretable Knowledge Tracing\" \"Simple and Efficient Student Modeling\" PDF","\"Can AI grade your essays?\" comparative analysis large language models teacher ratings PDF","site:dl.acm.org 10.1145/3706468.3706527","site:aclanthology.org automated essay scoring fairness 2024 primary study"],"sources":[{"source_id":"S1","title":"An Empirical Analysis Exploring the Impact of Traditional Exams and Multi-Stage Assignments on Academic Workload in a Final Year Engineering Context","publisher":"Practitioner Research in Higher Education / ERIC","url":"https://eric.ed.gov/?id=EJ1334835","source_class":"PRIMARY_RESEARCH","publication_date":"2021","accessed_at":"2026-08-03","claims_supported":["A one-semester quantitative study recorded assessment time in two higher-education courses.","The course using a multi-stage assignment with instructor feedback imposed 23% more instructor assessment workload than the comparison course.","The study supports the existence of a feedback-workload problem but is limited to two engineering courses and does not establish prevalence in religious studies."]},{"source_id":"S2","title":"EduMark AI: rethinking assessment and feedback with ethical AI","publisher":"Advance HE","url":"https://advance-he.ac.uk/news-and-views/edumark-ai-rethinking-assessment-and-feedback-ethical-ai/","source_class":"OFFICIAL_ORGANIZATION_DATA","publication_date":"2025-10-01","accessed_at":"2026-08-03","claims_supported":["Queen Mary University of London staff estimated more than 60 hours of marking and feedback per academic for one module.","QMUL pursued timely, consistent, personalized feedback while retaining educator oversight, and the Drapers’ Fund for Innovation in Teaching and Learning funded EduMark AI.","The initiative tested rubric-aligned feedback on more than 200 anonymized submissions and reported a 50–60% marking-time reduction, although this is a first-party project report rather than an independent evaluation.","EduMark requires educator review because AI can miss subtle reasoning or originality."]},{"source_id":"S3","title":"A-level Religious Studies 7062: Scheme of assessment","publisher":"AQA","url":"https://www.aqa.org.uk/subjects/religious-studies/a-level/religious-studies-7062/specification/scheme-of-assessment","source_class":"OFFICIAL_GUIDANCE","publication_date":"2026-01-29","accessed_at":"2026-08-03","claims_supported":["An official religious-studies assessment framework requires evidence-substantiated reasoning, critical interpretation of texts and sources, and recognition that others may hold different views.","The framework distinguishes knowledge and understanding from analysis and evaluation and covers several religious traditions.","These objectives make rubric-level reasoning features plausible targets, but the source concerns A-level examinations rather than recurring university formative assignments."]},{"source_id":"S4","title":"AI-assisted grading and answer groups","publisher":"Gradescope","url":"https://guides.gradescope.com/hc/en-us/articles/24838908062093-AI-assisted-grading-and-answer-groups","source_class":"OFFICIAL_PRODUCT_DOCUMENTATION","publication_date":"2025-12-16","accessed_at":"2026-08-03","claims_supported":["Gradescope already groups similar answers for rubric-based batch grading and lets instructors confirm groups or grade submissions individually.","Its documented AI grouping is limited to fixed-template multiple-choice, math fill-in, and one-line text fill-in responses.","The product also supports manually grouped open responses and regrade requests, but it does not document student-specific next-response prediction or model-plus-residual reconstruction."]},{"source_id":"S5","title":"Interpretable Knowledge Tracing: Simple and Efficient Student Modeling with Causal Relations","publisher":"arXiv; presented at EAAI/AAAI","url":"https://arxiv.org/abs/2112.11209","source_class":"PRIMARY_RESEARCH","publication_date":"2021-12-15","accessed_at":"2026-08-03","claims_supported":["Knowledge tracing infers student skill mastery from prior interactions and predicts future performance.","The IKT model uses interpretable features for individual mastery, transfer-related ability, and problem difficulty.","This supports the technical plausibility of a versioned learner-state predictor, but the demonstrated task uses sparse structured practice data rather than plural, open-ended religious-studies interpretation."]},{"source_id":"S6","title":"Can AI grade your essays? A comparative analysis of large language models and teacher ratings in multidimensional essay scoring","publisher":"arXiv; accepted at the Learning Analytics and Knowledge Conference 2025","url":"https://arxiv.org/abs/2411.16337","source_class":"PRIMARY_RESEARCH","publication_date":"2024-11-25","accessed_at":"2026-08-03","claims_supported":["Five language models were compared with 37 teachers across ten criteria on 20 authentic German student essays.","The best reported model reached Spearman r=0.74 with human overall scores and ICC=0.80 across repeated model ratings.","Models tended toward higher scores and were weaker at content quality than language-related criteria, supporting human validation and full-response fallback.","The small secondary-school narrative corpus does not establish performance on theology or religious-studies reasoning."]},{"source_id":"S7","title":"Family Educational Rights and Privacy Act regulations","publisher":"U.S. Department of Education, Student Privacy Policy Office","url":"https://studentprivacy.ed.gov/ferpa?exp=8","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-08-03 (current regulation page)","accessed_at":"2026-08-03","claims_supported":["Student submissions and maintained learner profiles can be education records, and students have inspection and correction-related rights.","Access without consent is limited to officials with legitimate educational interests; contractors must be under institutional control and restricted in use and redisclosure.","Education research may use reasonably de-identified records, while identifiable studies require purpose, scope, duration, security, and destruction controls under an applicable exception or consent.","FERPA is only a U.S. baseline; local privacy, accessibility, labor, assessment, and research-governance rules remain institution-specific."]},{"source_id":"S8","title":"National employment and wage data by occupation, May 2025","publisher":"U.S. Bureau of Labor Statistics","url":"https://www.bls.gov/news.release/ocwage.t01.htm","source_class":"GOVERNMENT_OR_REGULATOR","publication_date":"2026-05-15","accessed_at":"2026-08-03","claims_supported":["May 2025 mean annual wages were $148,100 for software developers, $126,800 for data scientists, $92,040 for postsecondary philosophy and religion teachers, and $47,670 for postsecondary teaching assistants.","These wage anchors support resource-equivalent labor estimates, but they exclude institutional benefits, overhead, procurement, hosting, and local wage variation."]}],"problem_evidence":{"support":"MODERATE","rationale":"Higher-education evidence shows that iterative instructor feedback can materially increase assessment workload, and an identifiable university reports more than 60 marking hours for one module. Official religious-studies objectives confirm that interpretive work requires evidence, critical analysis, and respect for differing views—the kinds of features the proposal would repeatedly inspect. No source measures how much religious-studies marking is predictable repetition, whether subtle changes are routinely missed, or whether staffing and assignment design rather than review representation cause the bottleneck.","source_ids":["S1","S2","S3","S6"]},"stakeholder_evidence":{"support":"MODERATE","rationale":"QMUL is an identifiable analogous adopter, the Drapers’ Fund an identifiable funder, and their expressed need is faster, consistent, personalized feedback with educator oversight. AQA is an identifiable assessment authority defining related religious-studies competencies, while Gradescope documents institutional demand for grading-efficiency tools. No religious-studies department, instructor, privacy office, or accessibility office has expressed demand for this specific predictive-residual workflow or committed data, authority, or funds.","source_ids":["S2","S3","S4"]},"prior_art":{"proximity":"ADJACENT_PRIOR_ART","closest_analogues":[{"name":"EduMark AI","similarity":"Uses anonymized assignments, rubric-aligned AI feedback, ethical approval, educator oversight, and a stated marking-time objective.","remaining_difference":"It evaluates each whole submission; the source does not describe a frozen pre-response learner prediction, signed residual representation, compatible reconstruction, sequential residual update, or independent raw-response audit.","source_ids":["S2"]},{"name":"Gradescope answer groups and AI-assisted grading","similarity":"Reduces review repetition by clustering similar answers, applying rubrics by group, supporting individual review, and accepting regrade requests.","remaining_difference":"It groups contemporaneous answers across students and is AI-limited to constrained response formats; it does not predict an individual student's next open interpretive response or reconstruct a rubric profile from a learner model plus residuals.","source_ids":["S4"]},{"name":"Interpretable Knowledge Tracing","similarity":"Maintains an interpretable student model and predicts later performance from prior skill evidence.","remaining_difference":"It predicts structured practice performance to adapt instruction, not model-relative residuals for human review of open interpretive writing, and it lacks raw-submission audits and reviewer-time compression.","source_ids":["S5"]},{"name":"Multidimensional LLM essay scoring","similarity":"Applies multiple predefined criteria to authentic essays and compares automated judgments with teachers.","remaining_difference":"It scores full essays independently rather than predicting the next rubric profile from the student's sequence; it also exhibits content-related divergence that the proposal must detect rather than assume away.","source_ids":["S6"]}],"distinctive_claim_remaining":"For recurring, low-stakes assignments within one frozen religious-studies task family, a learner-specific rubric profile predicted before opening each held-out response, followed by human-coded signed residuals and one-click full-text access, will reduce median total review time by at least 20% versus complete manual rubric review while keeping complete-profile reconstruction disagreement at or below 5%, producing no additional missed academically defensible alternative or mandatory-bypass case, and costing less in total after prediction, validation, audit, challenge, and fallback work are counted. Failure on any threshold, or an error pattern associated with an ethically reviewable language or interpretive-position grouping, falsifies the bounded claim.","confidence":"MODERATE"},"implementation_evidence":{"support":"MODERATE","rationale":"Interpretable learner-state prediction, multidimensional rubric coding, AI-assisted answer grouping, human oversight, regrade handling, and privacy-controlled research are all separately feasible. The proposal wisely preserves full submissions and excludes automatic consequential grading. However, no source demonstrates accurate prediction of open religious-studies reasoning from two prior assignments, reliable detection of defensible alternative interpretations, model-plus-residual reconstruction, unbiased residual routing, or net time savings after human coding and audits. A secure retrospective shadow study is feasible; operational deployment remains unverified.","source_ids":["S2","S4","S5","S6","S7","S8"]},"scores":{"meaningful_impact":{"score":3,"rationale":"The workload and feedback-timeliness problem can matter, but its magnitude in the target discipline and the educational value of redirecting saved attention are unmeasured.","source_ids":["S1","S2","S3"]},"stakeholder_pull":{"score":3,"rationale":"A funded university analogue and an established institutional grading product show real pull for workload-efficient, human-supervised feedback, but not for this exact workflow or domain.","source_ids":["S2","S4"]},"incremental_advantage":{"score":2,"rationale":"The claimed advantage over full review, whole-essay AI, and answer grouping is clear but entirely untested; coding, validation, audit, and fallback may consume any first-pass savings.","source_ids":["S2","S4","S6"]},"distinctiveness_plausibility":{"score":3,"rationale":"The searched analogues cover learner prediction, essay scoring, and repeated-answer compression separately; none documents their proposed reconstructive residual-and-audit combination. This is search-bounded distinctiveness, not world novelty.","source_ids":["S2","S4","S5","S6"]},"technical_implementability":{"score":3,"rationale":"Component technologies exist, and a transparent small-sample prototype is feasible, but open interpretive prediction and reliable valid-alternative detection are substantially harder than structured knowledge tracing or constrained answer grouping.","source_ids":["S4","S5","S6"]},"adoption_authority_feasibility":{"score":3,"rationale":"Faculty can authorize learning targets and low-stakes feedback, while an institution can authorize a de-identified study. Any identifiable or vendor-hosted learner profile additionally requires privacy, accessibility, research, procurement, and assessment-policy authority not yet secured.","source_ids":["S2","S3","S7"]},"evidence_readiness":{"score":2,"rationale":"There is adjacent evidence and a testable design, but there are no target-course prevalence data, held-out predictions, reconstruction results, reviewer-time measurements, equity audits, or committed data partner.","source_ids":["S1","S2","S5","S6"]},"safety_net_benefit":{"score":4,"rationale":"Full-response access, human validation, student challenge, independent audits, and mandatory fallback directly address known content-quality divergence and education-record contestability. Their effectiveness still requires testing.","source_ids":["S6","S7"]},"scalability":{"score":2,"rationale":"Existing tools show that AI-assisted assessment can scale, but this proposal needs per-student longitudinal data, domain-literate coding, independent marking, challenge handling, and frequent fallback; cross-course transfer is explicitly excluded and net operating cost is unknown.","source_ids":["S2","S4","S5","S8"]}},"score_confidence":"MODERATE","costs":{"first_evidence":{"band_2026_usd":"10K_TO_50K","scope":"Retrospective shadow evaluation of 48 already-authorized, de-identified responses from 12 students, including frozen rubric/model setup, counterbalanced residual-first and full-response review, blinded adjudication, analysis, privacy review, and a short report.","confidence":"MODERATE","assumptions":["Approximately 120–300 combined faculty, teaching-assistant, research, data-analysis, and project-management hours.","Existing secure institutional storage and rubric tooling are reused; no production LMS integration is built.","May 2025 BLS wages are treated as labor anchors and rounded upward for 2026 benefits and overhead.","Course records are already authorized and can be de-identified without recruitment or licensing expense."],"source_ids":["S7","S8"]},"initial_deployment_startup":{"band_2026_usd":"250K_TO_1M","scope":"Build and validate a secure, accessible, institution-controlled prototype for one institution, including learner-state versioning, residual and full-response interfaces, audit logging, LMS integration, threat modeling, privacy review, accessibility testing, and faculty configuration.","confidence":"LOW","assumptions":["Roughly 1–3 staff-years across software engineering, data science, product/accessibility design, security, and faculty subject-matter work.","The institution already has an LMS, identity management, approved hosting, and assessment records infrastructure.","No foundation-model training or multi-institution deployment is included.","The wide band reflects unknown integration, procurement, accessibility-remediation, and local governance costs."],"source_ids":["S2","S7","S8"]},"operational_launch":{"band_2026_usd":"50K_TO_250K","scope":"One prospective semester in a single low-stakes course, with notices or consent as locally required, staff training, shadow-only operation, independent audits, challenge handling, technical support, monitoring, and final evaluation.","confidence":"LOW","assumptions":["One course and one task family; no summative grades or automated feedback.","Includes 0.2–0.5 technical FTE-equivalent during launch plus compensated faculty, graders, accessibility, privacy, and evaluation work.","Full marking remains available and is performed for audit and every bypass case.","Model/API and hosting charges are assumed small relative to labor."],"source_ids":["S2","S7","S8"]},"annual_recurring":{"band_2026_usd":"50K_TO_250K","scope":"Steady-state operation for one department covering approximately 5–10 low-stakes course sequences, including hosting, model evaluation, 0.25–0.5 technical FTE-equivalent, faculty governance, independent audit marking, support, accessibility maintenance, and incident response.","confidence":"LOW","assumptions":["No cross-institution learner profile, high-stakes grading, or custom foundation-model training.","Audit and fallback rates remain bounded; a high fallback rate would eliminate savings and could move cost upward.","Institutional security, identity, and LMS services are already available.","Labor dominates cost, anchored to BLS software, data, faculty, and teaching-assistant wages."],"source_ids":["S7","S8"]}},"verified_pipeline_gates":{"externally_supported_problem":{"status":"YES","reason":"Independent higher-education research records a workload increase from iterative feedback, and an identifiable university reports substantial module marking time; official religious-studies criteria confirm the interpretive assessment burden. Target-discipline prevalence remains unknown but does not negate visible existence.","source_ids":["S1","S2","S3"]},"externally_credible_adopter_or_authorizer":{"status":"YES","reason":"QMUL is a credible analogous adopter, the Drapers’ Fund an identified funder, and AQA an identified authority defining relevant religious-studies assessment objectives. No party has committed to this exact proposal, which remains an adoption gap rather than absence of a credible adopter class.","source_ids":["S2","S3"]},"distinct_testable_incremental_claim":{"status":"YES","reason":"The pre-response, learner-specific, reconstructive residual workflow is contrastive against complete manual review, whole-essay AI, answer grouping, and knowledge tracing, with explicit time, disagreement, valid-alternative, bypass, equity, and total-effort falsifiers.","source_ids":["S2","S4","S5","S6"]},"bounded_next_evidence_step":{"status":"YES","reason":"A 48-response retrospective design can freeze the model before held-out work, compare residual-first with complete review, use blinded adjudication, and halt without changing grades or live learner profiles.","source_ids":["S6","S7","S8"]},"no_unresolved_safety_or_authority_stop":{"status":"YES","reason":"The authorized first step can be limited to institution-approved, reasonably de-identified records and read-only shadow analysis, with no live feedback or grading. It must not proceed until the institution confirms data authorization and de-identification; those conditions are operational prerequisites, while the bounded design itself presents no unavoidable safety or authority stop.","source_ids":["S7"]},"credible_cost_scope_and_range":{"status":"YES","reason":"All four estimates specify deployment scale and labor assumptions and use current official wage anchors. Confidence is low beyond the retrospective study because integration, governance, audit, and fallback rates are unknown.","source_ids":["S8"]}},"next_evidence_step":"With one institutionally authorized course partner, assemble 48 de-identified responses from 12 students completing four sequential low-stakes assignments in one task family. Use only assignments 1–2 to create transparent learner-state predictions; freeze the rubric, model, thresholds, bypasses, and analysis before opening assignments 3–4. Counterbalance qualified reviewers between complete-response and residual-first conditions, require every residual-first reviewer to have one-click full-text access, and obtain separate blinded double-full-mark adjudication for all 24 held-out responses. Measure total reviewer minutes including coding, validation, expansion, audit, challenge, and maintenance; complete-profile disagreement; missed material criteria and defensible alternatives; false residuals; feedback-action differences; confidence calibration; expansions; version failures; and fallback. Preregister success as at least 20% lower median total review time, at most 5% reconstruction disagreement, no additional missed defensible alternative or mandatory-bypass case versus adjudication, no evident language or interpretive-position error pattern, and lower total effort than full review. Any failure falsifies or narrows the claim; the study must remain shadow-only and cannot alter grades, feedback, or learner profiles.","blocking_evidence":["No measured prevalence of predictable repeated rubric confirmation in religious-studies or theology courses.","No held-out evidence that two prior open responses predict the declared features in later responses.","No empirical comparison of residual-first review with complete manual review on total time, reconstruction fidelity, or feedback quality.","No evidence that valid alternative interpretations and subtle contextual qualifications are preserved across traditions, languages, accessibility contexts, and argumentative styles.","No identified religious-studies course partner has committed authorized records, reviewers, privacy/accessibility approval, or pilot funding.","No measured audit, challenge, fallback, maintenance, or model-version failure rate from which net recurring cost or scalability can be estimated.","The proposed 12-student study is too small to establish subgroup equity or realized learning impact; it can only detect gross failure and justify or reject a larger study."],"research_disposition":"PARTNERED_RESEARCH_PROGRAM","world_novelty_boundary":"The eight-source search found adjacent funded assessment practice, grading products, learner-model research, and essay-scoring research, but no exact documented combination of pre-response student-specific prediction, reconstructive reasoning residuals, human-validated sequential updating, independent full-response audits, and mandatory interpretive fallback. This is only a bounded prior-art contrast. World novelty, patentability, freedom to operate, market size, and realized impact were not measured.","arm":"COMPLETE_PROPOSAL_PORTFOLIO","candidate_version":0,"controller_recommendation":{"action":"STOP_EMPIRICAL_RESEARCH_NEEDED","repairable":false,"material_progress_observed":true,"progress_targets":["Secure an identified course partner, institutional data authorization, de-identification confirmation, and qualified independent reviewers.","Preregister and execute the 48-response temporal holdout comparison with complete-review and residual-first comparators.","Demonstrate at least 20% net review-time reduction while meeting the 5% reconstruction limit and zero-extra-miss requirements for defensible alternatives and mandatory bypasses.","Report total effort including coding, validation, audits, challenges, maintenance, and fallback; reject the intervention if it does not beat complete review.","Use a subsequent adequately powered study, only if the bounded test passes, to test language, accessibility, tradition, and interpretive-position equity and later learning outcomes.","Obtain explicit conditional willingness from faculty and institutional privacy, accessibility, and assessment authorities before any prospective shadow launch."],"reason":"Web evidence verifies a meaningful general workload problem, credible adopter classes, adjacent prior art, component feasibility, and a bounded falsifiable claim. The decisive questions—predictability of open religious-studies reasoning, preservation of defensible alternatives, net reviewer-time advantage, equity, fallback rate, and workflow acceptability—require proprietary student work, qualified human adjudication, and observed reviewer behavior. They cannot be resolved by further bounded web research, so empirical partnered research is required."},"proposal_index":4}