PubMed Health⌕ Search

Biomedical subjects

Heinrich H Bülthoff

Publications and source records attributed to Heinrich H Bülthoff.

At least 19 recordsLinked to original sources

A Bayesian model of the disambiguation of gravitoinertial force by visual cues.

The otoliths are stimulated in the same fashion by gravitational and inertial forces, so otolith signals are ambiguous indicators of self-orientation. The ambiguity can be resolved with added visual information indicating orientation and acceleration with respect to the earth. Here we present a Bayesian model of the statistically optimal combination of noisy vestibular and visual signals. Likelihoods associated with sensory measurements are represented in an orientation/acceleration space. The likelihood function associated with the otolith signal illustrates the ambiguity; there is no unique solution for self-orientation or acceleration. Likelihood functions associated with other sensory signals can resolve this ambiguity. In addition, we propose two priors, each acting on a dimension in the orientation/acceleration space: the idiotropic prior and the no-acceleration prior. We conducted experiments using a motion platform and attached visual display to examine the influence of visual signals on the interpretation of the otolith signal. Subjects made pitch and acceleration judgments as the vestibular and visual signals were manipulated independently. Predictions of the model were confirmed: (1) visual signals affected the interpretation of the otolith signal, (2) less variable signals had more influence on perceived orientation and acceleration than more variable ones, and (3) combined estimates were more precise than single-cue estimates. We also show that the model can explain some well-known phenomena including the perception of upright in zero gravity, the Aubert effect, and the somatogravic illusion.

Acceleration↗

Spatial updating in virtual reality: the sufficiency of visual information.

Robust and effortless spatial orientation critically relies on "automatic and obligatory spatial updating", a largely automatized and reflex-like process that transforms our mental egocentric representation of the immediate surroundings during ego-motions. A rapid pointing paradigm was used to assess automatic/obligatory spatial updating after visually displayed upright rotations with or without concomitant physical rotations using a motion platform. Visual stimuli displaying a natural, subject-known scene proved sufficient for enabling automatic and obligatory spatial updating, irrespective of concurrent physical motions. This challenges the prevailing notion that visual cues alone are insufficient for enabling such spatial updating of rotations, and that vestibular/proprioceptive cues are both required and sufficient. Displaying optic flow devoid of landmarks during the motion and pointing phase was insufficient for enabling automatic spatial updating, but could not be entirely ignored either. Interestingly, additional physical motion cues hardly improved performance, and were insufficient for affording automatic spatial updating. The results are discussed in the context of the mental transformation hypothesis and the sensorimotor interference hypothesis, which associates difficulties in imagined perspective switches to interference between the sensorimotor and cognitive (to-be-imagined) perspective.

Adolescent↗

Multimodal similarity and categorization of novel, three-dimensional objects.

Similarity has been proposed as a fundamental principle underlying mental object representations and capable of supporting cognitive-level tasks such as categorization. However, much of the research has considered connections between similarity and categorization for tasks performed using a single perceptual modality. Considering similarity and categorization within a multimodal context opens up a number of important questions: Are the similarities between objects the same when they are perceived using different modalities or using more than one modality at a time? Is similarity still able to explain categorization performance when objects are experienced multimodally? In this study, we addressed these questions by having subjects explore novel, 3D objects which varied parametrically in shape and texture using vision alone, touch alone, or touch and vision together. Subjects then performed a pair-wise similarity rating task and a free sorting categorization task. Multidimensional scaling (MDS) analysis of similarity data revealed that a single underlying perceptual map whose dimensions corresponded to shape and texture could explain visual, haptic, and bimodal similarity ratings. However, the relative dimension weights varied according to modality: shape dominated texture when objects were seen, whereas shape and texture were roughly equally important in the haptic and bimodal conditions. Some evidence was found for a multimodal connection between similarity and categorization: the probability of category membership increased with similarity while the probability of a category boundary being placed between two stimuli decreased with similarity. In addition, dimension weights varied according to modality in the same way for both tasks. The study also demonstrates the usefulness of 3D printing technology and MDS techniques in the study of visuohaptic object processing.

Concept Formation↗

An advantage for detecting dynamic targets in natural scenes.

In the present study, we tested the extent to which observers use dynamic information to detect targets in natural scenes. For this purpose, we used composite stimuli in which target sequences were superimposed onto distractor sequences. We varied target visibility in the composite sequence, and the presence or absence of motion. Across four experiments, we found a dynamic advantage for target detection: Observers performed more accurately with dynamic than static target scenes. This advantage depended on the availability of target motion, irrespective of whether the target was upright or inverted in the image plane (Experiments 1-4). The magnitude of this advantage also depended on the availability of segmentation cues (Experiments 1 and 2) and on the distractors used (Experiments 2 and 4). Overall, the dynamic advantage reported extends previous work using isolated dynamic objects to more complex scenes.

Attention↗

Classification of faces in man and machine.

We attempt to shed light on the algorithms humans use to classify images of human faces according to their gender. For this, a novel methodology combining human psychophysics and machine learning is introduced. We proceed as follows. First, we apply principal component analysis (PCA) on the pixel information of the face stimuli. We then obtain a data set composed of these PCA eigenvectors combined with the subjects' gender estimates of the corresponding stimuli. Second, we model the gender classification process on this data set using a separating hyperplane (SH) between both classes. This SH is computed using algorithms from machine learning: the support vector machine (SVM), the relevance vector machine, the prototype classifier, and the K-means classifier. The classification behavior of humans and machines is then analyzed in three steps. First, the classification errors of humans and machines are compared for the various classifiers, and we also assess how well machines can recreate the subjects' internal decision boundary by studying the training errors of the machines. Second, we study the correlations between the rank-order of the subjects' responses to each stimulus-the gender estimate with its reaction time and confidence rating-and the rank-order of the distance of these stimuli to the SH. Finally, we attempt to compare the metric of the representations used by humans and machines for classification by relating the subjects' gender estimate of each stimulus and the distance of this stimulus to the SH. While we show that the classification error alone is not a sufficient selection criterion between the different algorithms humans might use to classify face stimuli, the distance of these stimuli to the SH is shown to capture essentials of the internal decision space of humans. Furthermore, algorithms such as the prototype classifier using stimuli in the center of the classes are shown to be less adapted to model human classification behavior than algorithms such as the SVM based on stimuli close to the boundary between the classes.

Algorithms↗

Comparing view sensitivity in shape discrimination with shape sensitivity in view discrimination.

In three picture-picture matching experiments, the effects of a view change on our ability to detect a shape change (Experiments 1 and 2) were contrasted with the effects of a shape change on our ability to detect a view change (Experiment 3). In each experiment, both view changes and shape changes influenced performance. However, shape changes had more influence than did view changes in the shape change detection task Conversely, view changes were more influential when the task was to detect view changes. Thus, the participants could often discriminate between the effects of shape changes and the effects of view changes. The disruptive effect of task-irrelevant changes (view changes in the first two experiments; shape changes in the final experiment) does not support Stankiewicz's (2002) claim that information about viewpoint and about shape can be estimated independently by human observers. However, the greater effect of variation in the task-relevant than in the task-irrelevant dimension indicates that the observers were moderately successful at disregarding irrelevant changes.

Attention↗

A search advantage for faces learned in motion.

Recently there has been growing interest in the role that motion might play in the perception and representation of facial identity. Most studies have considered old/new recognition as a task. However, especially for non-rigid motion, these studies have often produced contradictory results. Here, we used a delayed visual search paradigm to explore how learning is affected by non-rigid facial motion. In the current studies we trained observers on two frontal view faces, one moving non-rigidly, the other a static picture. After a delay, observers were asked to identify the targets in static search arrays containing 2, 4 or 6 faces. On a given trial target and distractor faces could be shown in one of five viewpoints, frontal, 22 degrees or 45 degrees to the left or right. We found that familiarizing observers with dynamic faces led to a constant reaction time advantage across all setsizes and viewpoints compared to static familiarization. This suggests that non-rigid motion affects identity decisions even across extended periods of time and changes in viewpoint. Furthermore, it seems as if such effects may be difficult to observe using more traditional old/new recognition tasks.

Adolescent↗

Is prior knowledge of object geometry used in visually guided reaching?

We investigated whether humans use prior knowledge of the geometry of faces in visually guided reaching. When viewing the inside of a mask of a face, the mask is often perceived as being a normal (convex) face, instead of the veridical, hollow (concave) shape. In this "hollow-face illusion," prior knowledge of the shape of faces dominates perception, even when in conflict with information from binocular disparity. Computer images of normal and hollow faces were presented, such that depth information from binocular disparity was consistent or in conflict with prior knowledge of the geometry. Participants reached to touch either the nose or cheek of the faces or gave verbal estimates of the corresponding distances. We found that reaching to touch was dominated by prior knowledge of face geometry. However, hollow faces were estimated to be flatter than normal faces. This suggests that the visual system combines binocular disparity and prior assumptions, rather than completely discounting one or the other. When comparing the magnitude of the hollow-face illusion in reaching and verbal tasks, we found that the flattening effect of the illusion was similar for verbal and reaching tasks.

Cognition↗

3D shape perception from combined depth cues in human visual cortex.

Our perception of the world's three-dimensional (3D) structure is critical for object recognition, navigation and planning actions. To accomplish this, the brain combines different types of visual information about depth structure, but at present, the neural architecture mediating this combination remains largely unknown. Here, we report neuroimaging correlates of human 3D shape perception from the combination of two depth cues. We measured fMRI responses while observers judged the 3D structure of two sequentially presented images of slanted planes defined by binocular disparity and perspective. We compared the behavioral and fMRI responses evoked by changes in one or both of the depth cues. fMRI responses in extrastriate areas (hMT+/V5 and lateral occipital complex), rather than responses in early retinotopic areas, reflected differences in perceived 3D shape, suggesting 'combined-cue' representations in higher visual areas. These findings provide insight into the neural circuits engaged when the human brain combines different information sources for unified 3D visual perception.

Brain Mapping↗

Similar cortical correlates underlie visual object identification and orientation judgment.

Visual object perception has been suggested to follow two different routes in the human brain: a ventral, view-invariant occipital-temporal route processes object identity, whereas a dorsal, view-dependent occipital-parietal route processes spatial properties of an object. Using fMRI, we addressed the question whether these routes are exclusively involved in either object recognition or spatial representation. We presented subjects with images of natural objects and involved them either in object identification or object orientation judgment task. For both tasks, we observed activation in ventro-temporal as well as parietal areas bilaterally, with significantly stronger responses for the orientation judgment in both ventro-temporal as well as parietal areas. Our findings suggest that object identification and orientation judgment do not follow strictly separable cortical pathways, but rather involve both the dorsal and the ventral stream.

Adult↗

The dynamics of visual pattern masking in natural scene processing: a magnetoencephalography study.

We investigated the dynamics of natural scene processing and mechanisms of pattern masking in a scene-recognition task. Psychophysical recognition performance and the magnetoencephalogram (MEG) were recorded simultaneously. Photographs of natural scenes were briefly displayed and in the masked condition immediately followed by a pattern mask. Viewing the scenes without masking elicited a transient occipital activation that started approximately 70 ms after the pattern onset, peaked at 110 ms, and ended after 170 ms. When a mask followed the target an additional transient could be reliably identified in the MEG traces. We assessed psychophysical performance levels at different latencies of this transient. Recognition rates were reduced only when the additional activation produced by the pattern mask overlapped with the initial 170 ms of occipital activation from the target. Our results are commensurate with an early cortical locus of pattern masking and indicate that 90 ms of undistorted cortical processing is necessary to reliably recognize a scene. Our data also indicate that as little as 20 ms of undistorted processing is sufficient for above-chance discrimination of a scene from a distracter.

Brain Mapping↗

Effects of rearranged vision on event-related lateralizations of the EEG during pointing.

We used event-related lateralizations of the EEG (ERLs) and reversed vision to study visuomotor processing with conflicting proprioceptive and visual information during pointing. Reversed vision decreased arm-related lateralization, probably reflecting the simultaneous activity of left and right arm specific neurons: neurons in the hemisphere contralateral to the observed action were probably activated by visual feedback, neurons in the hemisphere contralateral to the response side by the somatomotor feedback. Lateralization related to the target in parietal cortex increased, indicating that visual to motor transformation in parietal cortex required additional time and resources with reversed vision. A short period of adaptation to an additional lateral displacement of the visual field increased arm-contralateral activity in parietal cortex during the movement. This is in agreement with the, which showed that adaptation to a lateral displacement of the visual field is reflected in increased parietal involvement during pointing.

Adult↗

Learning from humans: computational modeling of face recognition.

In this paper, we propose a computational architecture of face recognition based on evidence from cognitive research. Several recent psychophysical experiments have shown that humans process faces by a combination of configural and component information. Using an appearance-based implementation of this architecture based on low-level features and their spatial relations, we were able to model aspects of human performance found in psychophysical studies. Furthermore, results from additional computational recognition experiments show that our framework is able to achieve excellent recognition performance even under large view rotations. Our interdisciplinary study is an example of how results from cognitive research can be used to construct recognition systems with increased performance. Finally, our modeling results also make new experimental predictions that will be tested in further psychophysical studies, thus effectively closing the loop between psychophysical experimentation and computational modeling.

Area Under Curve↗

Visual, haptic and crossmodal recognition of scenes.

Real-world scene perception can often involve more than one sensory modality. Here we investigated the visual, haptic and crossmodal recognition of scenes of familiar objects. In three experiments participants first learned a scene of objects arranged in random positions on a platform. After learning, the experimenter swapped the position of two objects in the scene and the task for the participant was to identify the two swapped objects. In experiment 1, we found a cost in scene recognition performance when there was a change in sensory modality and scene orientation between learning and test. The cost in crossmodal performance was not due to the participants verbally encoding the objects (experiment 2) or by differences between serial and parallel encoding of the objects during haptic and visual learning, respectively (experiment 3). Instead, our findings suggest that differences between visual and haptic representations of space may affect the recognition of scenes of objects across these modalities.

Adult↗

Merging the senses into a robust percept.

To perceive the external environment our brain uses multiple sources of sensory information derived from several different modalities, including vision, touch and audition. All these different sources of information have to be efficiently merged to form a coherent and robust percept. Here we highlight some of the mechanisms that underlie this merging of the senses in the brain. We show that, depending on the type of information, different combination and integration strategies are used and that prior knowledge is often required for interpreting the sensory signals.

Brain↗

Perceptual organization of local elements into global shapes in the human visual cortex.

The question of how local image features on the retina are integrated into perceived global shapes is central to our understanding of human visual perception. Psychophysical investigations have suggested that the emergence of a coherent visual percept, or a "good-Gestalt", is mediated by the perceptual organization of local features based on their similarity. However, the neural mechanisms that mediate unified shape perception in the human brain remain largely unknown. Using human fMRI, we demonstrate that not only higher occipitotemporal but also early retinotopic areas are involved in the perceptual organization and detection of global shapes. Specifically, these areas showed stronger fMRI responses to global contours consisting of collinear elements than to patterns of randomly oriented local elements. More importantly, decreased detection performance and fMRI activations were observed when misalignment of the contour elements disturbed the perceptual coherence of the contours. However, grouping of the misaligned contour elements by disparity resulted in increased performance and fMRI activations, suggesting that similar neural mechanisms may underlie grouping of local elements to global shapes by different visual features (orientation or disparity). Thus, these findings provide novel evidence for the role of both early feature integration processes and higher stages of visual analysis in coherent visual perception.

Humans↗

The use of facial motion and facial form during the processing of identity.

Previous research has shown that facial motion can carry information about age, gender, emotion and, at least to some extent, identity. By combining recent computer animation techniques with psychophysical methods, we show that during the computation of identity the human face recognition system integrates both types of information: individual non-rigid facial motion and individual facial form. This has important implications for cognitive and neural models of face perception, which currently emphasize a separation between the processing of invariant aspects (facial form) and changeable aspects (facial motion) of faces.

Adolescent↗

A chimeric point-light walker.

Ambiguity has long been used as a probe into visual processing. Here, we describe a new dynamic ambiguous figure-the chimeric point-light walker--which we hope will prove to be a useful tool for exploring biological motion. We begin by describing the construction of the stimulus and discussing the compelling finding that, when presented in a mask, observers consistently fail to notice anything odd about the walker, reporting instead that they are watching an unambiguous figure moving either to the left or right. Some observers report that the initial percept fluctuates, moving first to the left, then to the right, or vice versa; others always perceive a constant direction. All observers, when briefly shown the unmasked ambiguous figure, have no difficulty in perceiving the novel motion pattern once the mask is returned. These two findings--the initial report of unambiguous motion and the subsequent 'primed' perception of the ambiguity--are both consistent with an important role for top-down processing in biological motion. We conclude by suggesting several domains within the realm of biological-motion processing where this simple stimulus may prove to be useful.

Humans↗