Material-Divergence Red Team¶
Adversarial red team — instantiates Summary-Substance Alignment Audit
Puts an adversarial team on the summary alone to manufacture the most damaging defensible misreading — and log it before a hostile outsider finds it.
A Material-Divergence Red Team points a small adversarial group at the summary with one brief: given only this surface, what is the most damaging reading a person could take and defend as reasonable — that the substance would not support? Where the review and check mechanisms ask "does the summary match the body," the red team asks the generative, hostile question — "what could someone do with this summary that the body wouldn't back" — and treats the gap as an attack surface. It works from the summary outward, as an outsider would, invents the worst defensible misreadings, confirms against the body that they are genuinely unsupported, and logs the material ones as risks before a journalist, regulator, or competitor reaches them first.
Example¶
Before a company publishes its annual sustainability report, a red team takes only the summary and its headline figures. The summary says "50% of our energy is now renewable." The team probes the numbers an outsider will: renewable share of what — electricity only, or total energy including the vehicle fleet and process heat? Over what boundary — owned sites, or the far larger supply chain? They find the headline is defensible only under the narrowest reading the body permits, while a reasonable reader will infer the broadest. That gap goes into the summary-surface risk register as a material-divergence finding, rated by how damaging the misreading is and how easily an outsider reaches it — and flagged for a fix before publication, not after an exposé forces a correction.
How it works¶
- Take only the summary, as a hostile outsider would see it, with no privileged reading of the body.
- Generate the strongest damaging-but-defensible misreadings the surface invites.
- Confirm each against the substance — keep only the misreadings the body genuinely does not support.
- Rate by damage × reachability against the divergence threshold, and log the survivors to the risk register with a recommended fix.
Its distinguishing trait is that it is generative and adversarial: it invents misreadings rather than checking stated claims, and prioritizes by attack value rather than by claim order.
Tuning parameters¶
- Adversary model — who is reading hostilely: a journalist, a regulator, a competitor, a litigator; the model shapes which misreadings the team hunts.
- Interpretive latitude — how much benefit of the doubt the "reasonable reader" is denied; more latitude finds more gaps but risks far-fetched ones.
- Threshold — how damaging and how reachable a misreading must be to log, set against the material-divergence cutoff.
- Coverage — headline figures only, or the whole summary surface.
- Independence — an outside red team versus an internal rotation, trading candor for context.
When it helps, and when it misleads¶
Its strength is finding the failure the compliance-style checks miss — the summary that is technically true and still misleads — by simulating the adversary before the adversary arrives; this is classic red-teaming, attacking your own artifact to surface its failures first.[1] Its limit is that it produces hypotheses, not verdicts: an over-aggressive team manufactures far-fetched misreadings and floods the register with noise. The classic misuse is running it as theater — logging findings that are never fixed, so the register documents known risks without closing them. The discipline is to rate by reachability as well as damage, and to tie every logged finding to an owner and a fix.
How it implements the components¶
This team fills the adversarial-detection side of the audit:
summary_surface_risk_register— the logged, rated inventory of damaging misreadings the summary invites; the red team's deliverable.material_divergence_threshold— the cutoff a misreading's damage and reachability must clear to count as a material finding worth logging.
It generates and rates misreadings; it does not own the approval gate that assigns accountability (Dual-Surface Sign-Off), nor empirically test what ordinary readers actually infer (Summary-Only Reader Test), nor inventory the specific qualifiers a summary dropped (Qualifier-Drop Scan). It shares the divergence threshold with Body-Change Summary Invalidation, which enforces the same cutoff mechanically on edits rather than probing it adversarially.
Related¶
- Instantiates: Summary-Substance Alignment Audit — the adversarial pass that surfaces the technically-true-but-misleading summary before an outsider does.
- Sibling mechanisms: Dual-Surface Sign-Off · Qualifier-Drop Scan · Summary-Only Reader Test · Body-Change Summary Invalidation · Certainty & Causality Inflation Check
Notes¶
The red team and Summary-Only Reader Test approach the same worry from opposite ends: the reader test observes what real, non-adversarial readers conclude; the red team constructs the worst reading a motivated adversary could defend. The first bounds the typical misunderstanding, the second the weaponizable one — and a summary can be safe against one while exposed to the other.
References¶
[1] Red-teaming — the practice, originally from military and security work, of assigning a team to attack one's own artifact or plan in order to surface its failures before a real adversary exploits them. ↩