Skip to content

Bechdel test

The Bechdel–Wallace test is a minimal three-part screen for women's independent conversational presence in a fictional work.

Version
v2 · 2026-10-03 · History
Domain-specific #
13007
Domain group
Social Sciences
Origin domain
Communication & Media Studies
Subdomain
Film Criticism → Communication & Media Studies
Aliases
Bechdel Wallace Test

Core Idea

The Bechdel–Wallace test asks a deliberately small question about a fictional work: does it contain two women who talk to each other about something other than a man? Its three conditions are presence of the pair, a mutual exchange, and a topic independent of men. Alison Bechdel put the rule in her 1985 Dykes to Watch Out For strip “The Rule”; her own site credits her friend Liz Wallace for the idea. Calling it only Bechdel's invention would misstate that provenance.[1]

Passing this low threshold does not establish well-written women, gender equity or a feminist work. Failing can be a revealing absence, but a particular work may have a legitimate narrow setting or narrative purpose. Across a declared corpus, the rule can draw attention to a repeated representational pattern. The interpretation of any pass rate depends on how the films were selected and how borderline conversations were coded.[1][2]

Structural Signature

Sig role-phrases:

  • Work: a bounded film or other fictional narrative under examination.
  • Two women: more than one focal-group character must appear.
  • Mutual exchange: those women speak to each other, not merely in separate scenes.
  • Independent topic: at least one such exchange concerns something besides a man.
  • Coding convention: a stated rule for ambiguities, including what counts as a character, conversation, or topic.[1][2]

The three conditions are conjunctive. If the pair does not exist, later tests cannot be met. If both women appear but never converse, mere cast presence is insufficient. If they converse only about men, the third condition is unmet. That staged form is why researchers can code intermediate results rather than only a final pass/fail.[2]

What It Is Not

It is not a test of a film's overall politics, artistic quality, screen time, agency or the complexity of its characters. A short exchange can satisfy it while the rest of the film marginalizes women; a film centered on one woman can fail for lacking a second. Neither result, in isolation, resolves a critical judgment. The point is the conspicuously minimal conversational threshold.[1][2]

Nor is the popular “two named women” requirement part of the original three-part wording. It is an operational variant used by databases and by Agarwal and colleagues' computational study. Naming can make coding more consistent and avoid counting a fleeting unidentified role, but changes the threshold. Studies must say which version they use before comparing rates.[1][2]

Scope of Application

The source context was filmgoing in a comic strip, and film remains the main application. The underlying question can be asked of other narrative media, but transferring it requires decisions about what counts as “talk” in text, interactive stories or other forms. Games and literature are possible extensions, but the cited primary sources establish film and a potentially reusable fiction screen, not a specific measured result in every medium.[1][2]

Agarwal and colleagues' 2015 original study made the rule computational. They acquired 964 screenplays, found 457 with labels from an existing ratings site, then used 367 for training/development and 90 for testing. In that bounded labeled set, 191 passed and 266 failed the full test. Those are sample counts, not a claim that a given percentage of all films fail. The authors' stages operationalize named women, conversational links and aboutness of dialogue. The two script-level applications below are separately checked against original screenplays; they do not rely on a corpus total as a substitute for showing the three conditions.[2][3][4]

Clarity

The first condition is about two women in one work, not two women listed in a franchise. The second demands some exchange between them, not two separate conversations with a man. The third asks what their exchange is about; merely mentioning “he” does not automatically mean a conversation is about a man. The 2015 authors explicitly discuss the difference between a word mention and semantic aboutness when evaluating automated features.[2]

One of their error analyses is especially concrete. For 2001: A Space Odyssey, they report only one named woman, Elena, under their coding. An automated method treated “Stewardess” as another named woman and therefore falsely cleared the first condition. That is a first-stage coding error, not a reported full three-condition film verdict; it demonstrates why the named-character variant and cast metadata need audit.[2]

Manages Complexity

The rule compresses a broad concern—women can be present yet narratively oriented only around men—into three inspectable questions. A viewer can ask them without evaluating every line's significance. In a research project, stage-wise coding shows where a film fails, rather than reducing all failures to one opaque score. Agarwal and colleagues could therefore analyze separate errors in identifying women, linking speakers, and deciding conversational topic.[2]

Compression has a cost. Dialogue aboutness is not reducible to a word ban: the researchers show that a sentence mentioning a man may concern a medical situation, while a sentence without an explicit male name may still be about a romantic relationship. Their machine-learning work treats this as a difficult semantic problem. Human coders can disagree too; a binary result is not automatically reproducible just because the rule is short.[2]

Abstract Reasoning

The rule is a conjunction: pair ∧ exchange ∧ independent topic. The order matters for audit but not for final logic. A film with two women and non-man dialogue elsewhere, but never between those women, does not satisfy it. A film with one rich female protagonist also fails the first condition. These counterfactuals show the identity is conversational independence, not general presence or quality.[1]

Aggregating outcomes asks a different question from judging an individual work. If a carefully defined sample of films frequently fails such a low screen, that may prompt investigation of production or narrative conventions. It cannot by itself establish causes, intent, budget effects or audience impact. Agarwal's 457 labeled screenplays are a selected subset of the 964 they acquired; treating 191/457 as universal industry prevalence would cross that boundary.[2]

Knowledge Transfer

The three-part form can inspire other representation screens, but replacing “women” or “man” with a different group changes the substantive question and should be named as an adaptation, not a result of the original test. Even within film, requiring named women rather than any two women changes the admission threshold. A good transfer carries explicit coding rules and limits, not merely a familiar label.[1][2]

The live representation prime is broader: it concerns how an entity is depicted or stood for in a medium. The Bechdel–Wallace test is one intentionally thin conversational screen. The overlap is obvious, but a strict parent edge is not asserted without catalog review.

Examples

  1. Frozen (2013), Anna and Elsa at the ice palace. Jennifer Lee's final shooting script, pp. 68–70, labels both sisters by name and shows alternating speech between them. Anna tells Elsa about the winter affecting Arendelle; Elsa answers that she cannot reverse it, and the two discuss the storm and danger. A prior portion of the encounter touches Elsa's solitude, but the identified exchange about ending the winter is not about a man. Mapped back: work = this film script; two women = Anna and Elsa; mutual exchange = their back-and-forth in the ice-palace scene; independent topic = the uncontrolled winter and how to reverse it; coding convention = original three conditions, also satisfying the later named-character variant. Thus this source-located scene clears all three conditions. The claim is about the shooting-script scene, not an audit of every cut in the released film.[3]

  2. The Devil Wears Prada (2006), the belt selection. Aline Brosh McKenna's shooting-script revisions, scene 45, pp. 27–28A, show Miranda evaluating two belts and Andy saying she cannot tell them apart; Miranda responds by explaining the significance of fashion choices and Andy's sweater. Andy and Miranda are named women addressing one another, and this exchange concerns clothing and work rather than a man. Mapped back: work = this film script; two women = Andy and Miranda; mutual exchange = Andy's response and Miranda's reply in the same scene; independent topic = belt selection, fashion and the industry; coding convention = the original rule and, independently, the named-character variant. The scene clears all three conditions without making the film a certificate of good representation. The cited script is a revision-stage primary artifact, so this is not a claim about every released-film edit.[4]

Structural Tensions

Simple minimal screen versus rich portrayal assessment. Three low-cost questions make an absence easy to notice and let researchers compare a declared sample. The same low bar lets shallow token dialogue pass and excludes information about agency, plot centrality and quality. Adding those judgments can improve interpretation but sacrifices the original test's simplicity and comparability. Diagnostic: is a pass reported only as a minimal conversational-presence finding, with richer claims supported by separate evidence?[1][2]

Structural–Framed Character

The conjunction of pair, exchange and topic has a clear structural shape, but its evaluative weight comes from a feminist critique of how fiction positions women. It depends on authored characters, dialogue and human judgments of “aboutness,” not on an observer-independent natural threshold. The rule originated in a comic and Wallace's conversation, then traveled into databases and computational research with institutional coding choices such as “named.” Importing it into other media or groups is warranted only when the new unit of dialogue and focal group are specified, not merely because three boxes can be ticked. Its character: a normatively motivated, deliberately weak representational screen whose simple form is portable but whose interpretation is medium- and coding-dependent.[1][2]

Structural Core vs. Domain Accent

The skeleton is a conjunctive minimum: a pair from the focal group interacts on a topic not defined by the excluded group. The live Verification prime supplies the conformance-check genus, with a declared criterion, procedure and limited pass/fail verdict. The domain mechanism is conversational presence in a fictional narrative; the named-character requirement and screenplay parsing are later accents. The named test fails the prime bar because its normative point and measurement problem depend on gendered fictional dialogue. A generic prime about threshold tests would need independent unlike-domain cases and would not inherit the Bechdel–Wallace result's meaning.[1][2]

This entry is a kind of Verification.

The live Verification prime is the strict parent: the three stated dialogue conditions are checked against a bounded fictional work and yield a limited conformance verdict. A pass does not verify artistic quality or overall gender equity, and named-character requirements are later coding variants. The live Representation prime is a comparison, not the checked parent.

Relationships to Other Abstractions

Local relationship map for Bechdel testParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Bechdel testDOMAINPrime abstraction: Verification — is a kind ofVerificationPRIME

Current abstraction Bechdel test Domain-specific

Parents (1) — more general patterns this builds on

  • Bechdel test is a kind of Verification Prime

    The Bechdel–Wallace test is a bounded verification procedure.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Bechdel test sits in a sparse region of the domain-specific corpus (69th percentile for distinctiveness): few abstractions share its structure, so a faithful description tends to retrieve it precisely.

Family — Narrative Structure & Storytelling Devices (24 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

  • A certificate that a film is feminist, fair, or artistically good.
  • The later named-character variant presented as the original 1985 wording.
  • Counting two women who never talk to each other.
  • Treating a convenience set of labeled scripts as the whole film industry.[2]

References

[1] Alison Bechdel's official Dykes to Watch Out For site, “The Rule”, reproducing the circa-1985 strip and crediting Liz Wallace for the rule; the 2005 post is by site writer Cathy, as it states. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k

[2] Apoorv Agarwal, Jiehan Zheng, Shruti Kamath, Sriramkumar Balasubramanian and Shirin Ann Dey, “Key Female Characters in Film Have More to Talk About Besides Men: Automating the Bechdel Test,” NAACL (2015), 830–840, original study, §§1, 4–7, Table 1 and error analysis. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p ↩q

[3] Jennifer Lee, Frozen, final shooting script, 23 September 2013, original screenplay copy, pp. 68–70, ice-palace exchange between Anna and Elsa; third-party host, not a claim that the PDF is currently studio-hosted. registry ↩a ↩b

[4] Aline Brosh McKenna, The Devil Wears Prada, shooting draft and revisions, 2005, original screenplay copy hosted on screenwriter John August's site, scene 45, pp. 27–28A. This entry paraphrases the exchange rather than reproducing screenplay dialogue. registry ↩a ↩b