Visual Capture¶
When vision conflicts with another sense about a shared property, the combined percept can be pulled toward the visual cue under conditions favoring vision.
Core Idea¶
Visual capture occurs when a combined judgment about one property is pulled toward a visual cue despite conflicting touch or auditory information. Rock and Victor optically distorted the visible shape of a grasped object; drawing and matching judgments strongly favored what was seen, often without noticed conflict. The effect describes a measured direction of bias, not literal deletion of the other sense.[^ref-9db774f1195a]
Scope of Application¶
Rock–Victor's shape case spans vision and touch. Alais and Burr's location task spans vision and hearing: good visual localization pulled apparent sound position toward sight. Severe visual blur reversed the effect so sound captured vision. Ernst and Banks found visual–haptic height judgments close to a precision-weighted combination model, qualifying any permanent vision-dominance hierarchy.[ref-9db774f1195a][ref-e67ee1182cb3][^ref-0ff1bfaf4cb3]
Clarity¶
The cues must concern a shared judged property, be compared under a discrepancy and produce a combined report. A visual-only response is not cross-modal capture; a sound-driven shift of blurred vision is the reverse case. Relative cue precision explains tested shifts but is not proof of one universal neural pathway.[ref-0ff1bfaf4cb3][ref-e67ee1182cb3]
Manages Complexity¶
Integration yields one estimate rather than incompatible sensory answers. It may improve precision when cues share a source, yet a precise, experimentally displaced visual cue can bias that estimate away from the other signal. The physical truth of a cue and its measurement precision are different.[ref-9db774f1195a][ref-e67ee1182cb3]
Abstract Reasoning¶
Identify the property, each unimodal cue and the combined response. Compare the response with unimodal baselines; ask which direction it shifted and how blur or noise changed cue reliability. Do not infer vision always dominates from one positive case.[^ref-e67ee1182cb3]
Knowledge Transfer¶
The visual-directed role pattern transfers from shape to sound location, while the task and stimulus do not. Visual capture presupposes Multisensory Integration, but is an outcome rather than a subtype of that process. The live Ventriloquism Effect is a narrower audiovisual instance; Bayesian Cue Integration is a broader model whose use here does not prove one mechanism for every report.
[^ref-9db774f1195a]: Rock and Victor, original Science visual–touch experiment (1964), indexed abstract. [^ref-0ff1bfaf4cb3]: Ernst and Banks, original visual–haptic cue-integration study (2002). [^ref-e67ee1182cb3]: Alais and Burr, original audiovisual location and blur experiment (2004).
Relationships to Other Abstractions¶
Current abstraction Visual Capture Domain-specific
Parents (1) — more general patterns this builds on
-
Visual Capture presupposes Multisensory integration Domain-specific
Visual capture presupposes cross-modal cue interaction.
Hierarchy path (1) — routes to 1 parentless root
- Visual Capture → Multisensory integration → Perceptual Process
Neighborhood in Abstraction Space¶
Visual Capture sits in a moderately populated region (53rd percentile for distinctiveness): it has near-neighbors but no dense thicket of look-alikes.
Family — Multisensory Perception & Binding (13 abstractions)
Nearest neighbors
- Ventriloquism Effect — 0.86
- Cue Validity — 0.86
- Leitmotif — 0.86
- Navon Figure — 0.86
- Eye Tracking — 0.85
Computed from structural-signature embeddings · 2026-10-08