Skip to content

Visual Capture

When vision conflicts with another sense about a shared property, the combined percept can be pulled toward the visual cue under conditions favoring vision.

Version
v1 · 2026-10-03 · History
Domain-specific #
13693
Domain group
Social Sciences
Origin domain
Psychology & Behavioral Sciences
Subdomain
Perceptual Psychology → Psychology & Behavioral Sciences
Aliases
Visual dominance in cross-modal conflict

Core Idea

Visual capture names a cross-modal perceptual result: when seeing and another sense give discordant information about a common property, the reported combined percept can shift toward what vision indicates. In Rock and Victor's 1964 experiment, observers grasped an object while optical distortion made its seen shape differ from its felt shape; drawing and matching reports strongly favored vision, often without reported awareness of the mismatch. In Alais and Burr's audiovisual localization experiment, a well-localized visual target pulled the perceived location of a spatially discrepant sound toward itself.[1][2]

The effect is conditional, not a standing rule that vision defeats all other senses. Ernst and Banks measured visual and haptic precision in a height-judgment task and found behavior close to a maximum-likelihood cue-combination model: vision dominates when its estimate is less variable. Alais and Burr progressively blurred the visual stimulus; moderate blur yielded a compromise, and severe blur reversed the direction so auditory location captured vision. Their result supports a reliability-based explanation for their tested task, not a proof that every visual-capture episode has one neural mechanism.[3][2]

Thus the defining observation is a visually directed shift of a measured percept under cross-modal conflict. A puppet, a prism-distorted object and a blurred test flash do not simply share a word “visual”; they test whether estimates about one judged property are combined and which cue pulls the report.

Structural Signature

Sig role-phrases: jointly judged property; visual estimate; nonvisual estimate; imposed discrepancy; relative cue precision; combined perceptual report; visually directed bias.

  1. Shared property: shape, height or spatial location is judged through more than one sense. Without a common target property, “capture” is an ungrounded metaphor.
  2. Visual cue: supplies one estimate. It need not be physically correct; Rock and Victor deliberately distorted its shape information.[1]
  3. Other modality: touch or hearing supplies another estimate, allowing a conflict and unimodal comparison.
  4. Discrepancy and co-presentation: the experiment makes the signals disagree while participants judge them in a combined setting. A universal synchrony or maximum-distance threshold is not asserted from these three sources.
  5. Relative precision: determines weight in the Ernst–Banks and Alais–Burr models; it is explanatory in those experiments, not a mandatory measurement in every historical report.[3][2]
  6. Report: drawing/matching or location responses show a shift toward vision. “Vision is dominant” is a claim about measured judgments, not direct observation of one modality being deleted.

Condensed: shared property + discordant visual/nonvisual cues + combined judgment pulled toward vision = visual capture.

What It Is Not

  • Not universal vision supremacy. Severe blur in Alais and Burr produced auditory capture of vision, the opposite direction.[2]
  • Not synonymous with Bayesian Cue Integration. That live prime describes a broader estimation rule that can favor either cue; visual capture is one directional outcome under suitable conditions.
  • Not identical to the ventriloquism effect. Ventriloquism is the audio–visual location instance. Visual capture also occurs in visual–touch shape judgments such as Rock and Victor's.[1][2]
  • Not ordinary vision-only perception. A comparison requires another sensory cue about the same property.
  • Not proof that the nonvisual signal is absent. A biased integrated report, even without conscious conflict awareness, does not establish loss of touch or hearing input.
  • Not all multisensory conflict. When cues are kept separate, compromise evenly, or the nonvisual cue captures vision, the visually directed result is absent.
  • Not necessarily an error by the nervous system. Precision-weighted integration may improve estimates when cues share a source; deliberately displaced stimuli make the same rule yield a shifted report.[3][2]

Scope of Application

Rock and Victor created a visual–tactual conflict over a viewed and grasped object's shape by optical distortion. The observers then drew what they perceived or matched another object. The original abstract reports strong visual dominance and frequent lack of conflict awareness. It does not report an exact weight for each sense, nor justify extending the result to every property or viewing condition.[1]

Ernst and Banks studied height through combined visual and haptic cues. They first measured each modality's variance, then compared human bimodal judgments with a maximum-likelihood integrator. The original study reports close model agreement and makes dominance depend on relative variance. It qualifies Rock and Victor's categorical reading: haptics can affect a visual–haptic estimate and, depending on precision, vision need not always receive the larger weight.[3]

Alais and Burr studied spatial location of paired auditory and visual targets. Good visual localization yielded the familiar ventriloquist/visual-capture direction. Blur manipulated visual precision; with intermediate blur the combined location lay between cues, and with extensive blur hearing could pull visual location. The paper's measured reversal is not a separate “visual-capture” instance, but a decisive boundary case against an immutable vision hierarchy.[2]

Clarity

“Capture” describes the direction of a perceptual report, not a literal force between senses. In a combined location task, compare the bimodal report with auditory-only and visual-only localization; if it is drawn toward the visual coordinate relative to hearing's estimate, the visual-capture description is apt. In a shape task, compare drawing or object matching under combined viewing/feeling with the discrepant source cues. A physical stimulus can be wrong in one modality because the experiment deliberately distorts it; a coherent final percept need not match the actual tactual or sound-source location.[1][2]

Precision and correctness also differ. A visual cue may be tightly localized but deliberately displaced. The brain can weight a precise cue strongly and still reach a biased answer about the externally defined “true” source. Conversely, extreme blur makes vision uninformative and can reverse which cue pulls the combined judgment.[3][2]

Manages Complexity

A single percept avoids issuing two incompatible shape or location reports for what the task treats as one object or event. Precision weighting offers a principled way to combine noisy evidence and, in Alais and Burr's experiment, bimodal localization often improved over unimodal precision. But that compression can obscure the discrepancy: Rock and Victor reported strong visual favoring even when the viewed shape was optically distorted. The system's estimate is adaptive to estimated cue quality, not guaranteed veridicality under experimentally manipulated conflicts.[1][2]

Abstract Reasoning

For a claimed visual-capture case, identify the shared judged variable, the visual and nonvisual values, the combined presentation, unimodal reliability if measured, and the actual response measure. Ask whether the combined report moves toward vision and whether the visual cue was more precise or simply salient. Test a reversal by degrading vision or improving the other modality rather than assuming a fixed sensory hierarchy. Keep the inference to integration separate from claims about a particular brain pathway unless that pathway was independently measured.[3][2]

Diagnostic: Under the tested cue reliabilities, which modality actually draws the combined report toward itself, and what is the comparison baseline?

Knowledge Transfer

The directional pattern transfers from visual–touch shape to visual–auditory location because both have a shared estimated property, discordant cues and a measurable visual pull. The exact stimulus, response and weighting model do not transfer unchanged. Ventriloquism is the live domain-specific auditory-location instance; Bayesian Cue Integration is a live broad prime that can explain both visual and reversed nonvisual influence. Calling every persuasive image “visual capture” imports a multisensory experiment that may not exist.

Examples

Rock–Victor distorted shape (1964)

Observers looked at an object through optical distortion while simultaneously grasping it, so seen and felt shape differed. They indicated the resulting shape by drawing it or matching another object. The original report says the judgments strongly followed visual shape, often without noticed conflict. This is a direct visual–tactual case, not an inference from a ventriloquist stage act.[1]

Mapped back: shared property = object shape; visual cue = distorted seen shape; nonvisual cue = grasped tactual shape; discrepancy = experimental optics; relative precision = not quantified by this paper; report = drawing/matching favors vision; boundary = lack of awareness is reported often, not asserted universally.

Alais–Burr good-vision audiovisual location (2004)

With a well-localized visual stimulus paired with a sound at a different location, observers' combined location judgment was drawn toward vision. The same experiment changed visual blur and found less visual dominance and ultimately auditory capture when the visual target was severely blurred. The positive visual-capture condition therefore includes a tested boundary in the same original source.[2]

Mapped back: shared property = spatial location; visual cue = well-localized target; nonvisual cue = sound; discrepancy = experimental spatial offset; relative precision = good visual localization; report = perceived bimodal position toward vision; boundary = severe blur reverses direction.

Severe-blur reversal

In Alais and Burr, sufficiently blurred visual targets were localized poorly enough that sound captured the visual position. This is an observed negative visual-capture case, not a failed replication of all cross-modal interaction.[2]

Mapped back: common location, two modalities and a combined report remain; visual-directed dominance does not.

Structural Tensions

Precision gain versus conflict-induced bias. Combining cues can reduce uncertainty when they concern one source; Ernst–Banks's model and Alais–Burr's improved bimodal localization illustrate the benefit. Yet if an experimental lens or offset places a precise visual cue at a misleading value, giving it greater weight can pull the percept away from the nonvisual signal. Ignoring vision avoids that particular distortion but forfeits its precision when it is informative. Diagnostic: Is the visual cue merely precise, or also aligned with the physical source the task aims to recover?[1][3][2]

This tension belongs to cue-combination under conflict, not an assertion that every act of seeing involves a tradeoff. Rock and Victor's unawareness result is a reported outcome, not an independent universal “coherence versus awareness” cost.

Structural–Framed Character

The entry lies between structural and framed: the roles of shared property, discrepant cues and visually pulled report can be operationalized across experiments, but which cue dominates depends on stimulus quality, task instructions and response measure. The evaluative language “dominance” can mislead if read as vision's permanent superiority; the original reversal and variance model make it a condition-specific description. Human perceptual practice supplies judgments, while experimental institutions create lens distortion, blur, spatial offsets and unimodal baselines that make the effect visible. Vocabulary travels faithfully from shape conflict to location conflict if the visual-direction criterion remains; it travels too far if “visual capture” is used for ordinary attention to images or if Bayesian reliability weighting is treated as proof of a universal neural route. The sources let us recognize a repeatable bias without importing an unmeasured thalamic hierarchy from a secondary description. Its character: an experimentally framed, direction-specific perceptual effect with a reusable cross-modal role structure and reliability-dependent limits.[1][3][2]

Structural Core vs. Domain Accent

The portable skeleton is combining discrepant evidence about a shared property, possibly weighted by reliability; verified live Bayesian Cue Integration can own that broad model. The domain-bound mechanism is a human visual cue pulling a measured auditory or haptic percept under particular cross-modal stimulus conditions. This named entry fails the prime bar because cue interaction can favor touch or sound, and many evidence-combination settings have no visual sense or perceptual report. Multisensory Integration is a necessary cross-modal prerequisite, not a subsumption parent of the visual-capture outcome. Ventriloquism Effect is a narrower auditory-location neighbor; it does not exhaust visual–touch capture. No single neural or Bayesian mechanism is asserted for every report.

This entry presupposes Multisensory integration.

Strict compositional prerequisite: Multisensory integration (presupposes, strict). Visual capture needs interacting discrepant senses, but multisensory interaction need not produce visual dominance. Bayesian Cue Integration is a model neighbor rather than a universal mechanism claim. Ventriloquism Effect is a narrower auditory-location phenomenon.

Relationships to Other Abstractions

Local relationship map for Visual CaptureParents appear above the current abstraction, mutual partners to the right, and children below. Node labels state whether each abstraction is prime or domain-specific; colors identify relation types.Visual CaptureDOMAINDomain-specific abstraction: Multisensory integration — presupposesMultisensoryintegrationDOMAIN

Current abstraction Visual Capture Domain-specific

Parents (1) — more general patterns this builds on

  • Visual Capture presupposes Multisensory integration Domain-specific

    Visual capture presupposes cross-modal cue interaction.

Hierarchy path (1) — routes to 1 parentless root

Neighborhood in Abstraction Space

Visual Capture sits in a moderately populated region (53rd percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.

Family — Multisensory Perception & Binding (13 abstractions)

Nearest neighbors

Computed from structural-signature embeddings · 2026-10-08

Not to Be Confused With

Ventriloquism effect: the audiovisual sound-location case. Haptic or auditory capture: the reverse direction when touch or sound dominates a poorer visual cue. Colavita effect: a response/attention phenomenon, not automatically the same location-or-shape estimate. Visual illusion without another modality: lacks cross-modal conflict. Maximum-likelihood cue integration: a broader explanatory model, not the measured visual-bias result itself.[3][2]

References

[1] Irvin Rock and Jack Victor, “Vision and Touch: An Experimentally Created Conflict Between the Two Senses,” Science 143 (1964), 594–596, original indexed abstract. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i

[2] David Alais and David Burr, “The ventriloquist effect results from near-optimal bimodal integration,” Current Biology 14 (2004), 257–262, original paper, visual-blur manipulation and reversal. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i ↩j ↩k ↩l ↩m ↩n ↩o ↩p

[3] Marc O. Ernst and Martin S. Banks, “Humans integrate visual and haptic information in a statistically optimal fashion,” Nature 415 (2002), 429–433, original abstract and figures. registry ↩a ↩b ↩c ↩d ↩e ↩f ↩g ↩h ↩i