PubMed Health⌕ Search

Biomedical subjects

Maximilian Riesenhuber

Publications and source records attributed to Maximilian Riesenhuber.

8 recordsLinked to original sources

Robust object recognition with cortex-like mechanisms.

We introduce a new general framework for the recognition of complex visual scenes, which is motivated by biology: We describe a hierarchical system that closely follows the organization of visual cortex and builds an increasingly complex and invariant feature representation by alternating between a template matching and a maximum pooling operation. We demonstrate the strength of the approach on a range of recognition tasks: From invariant single object recognition in clutter to multiclass categorization problems and complex scene understanding tasks that rely on the recognition of both shape-based as well as texture-based objects. Given the biological constraints that the system had to satisfy, the approach performs surprisingly well: It has the capability of learning from only a few training examples and competes with state-of-the-art systems. We also discuss the existence of a universal, redundant dictionary of features that could handle the recognition of most object categories. In addition to its relevance for computer vision, the success of this approach suggests a plausibility proof for a class of feedforward models of object recognition in cortex.

Algorithms↗

Evaluation of a shape-based model of human face discrimination using FMRI and behavioral techniques.

Understanding the neural mechanisms underlying object recognition is one of the fundamental challenges of visual neuroscience. While neurophysiology experiments have provided evidence for a "simple-to-complex" processing model based on a hierarchy of increasingly complex image features, behavioral and fMRI studies of face processing have been interpreted as incompatible with this account. We present a neurophysiologically plausible, feature-based model that quantitatively accounts for face discrimination characteristics, including face inversion and "configural" effects. The model predicts that face discrimination is based on a sparse representation of units selective for face shapes, without the need to postulate additional, "face-specific" mechanisms. We derive and test predictions that quantitatively link model FFA face neuron tuning, neural adaptation measured in an fMRI rapid adaptation paradigm, and face discrimination performance. The experimental data are in excellent agreement with the model prediction that discrimination performance should asymptote as faces become dissimilar enough to activate different neuronal populations.

Adolescent↗

Experience-dependent sharpening of visual shape selectivity in inferior temporal cortex.

Whereas much is known about the visual shape selectivity of neurons in the inferior temporal cortex (ITC), less is known about the role of visual learning in the development and refinement of ITC shape selectivity. To address this, we trained monkeys to perform a visual categorization task with a parametric set of highly familiar stimuli. During training, the stimuli were always presented at the same orientation. In this experiment, we recorded from ITC neurons while monkeys viewed the trained stimuli in addition to image-plane rotated versions of those stimuli. We found that, concomitant with the monkeys' behavioral performance, neuronal stimulus selectivity was stronger for stimuli presented at the trained orientation than for rotated versions of the same stimuli. We also recorded from ITC neurons while monkeys viewed sets of novel and familiar (but not explicitly trained) randomly chosen complex stimuli. We again found that ITC stimulus selectivity was sharper for familiar than novel stimuli, suggesting that enhanced shape tuning in ITC can arise for both passively experienced and explicitly trained stimuli.

Algorithms↗

Face processing in humans is compatible with a simple shape-based model of vision.

Understanding how the human visual system recognizes objects is one of the key challenges in neuroscience. Inspired by a large body of physiological evidence, a general class of recognition models has emerged, which is based on a hierarchical organization of visual processing, with succeeding stages being sensitive to image features of increasing complexity. However, these models appear to be incompatible with some well-known psychophysical results. Prominent among these are experiments investigating recognition impairments caused by vertical inversion of images, especially those of faces. It has been reported that faces that differ 'featurally' are much easier to distinguish when inverted than those that differ 'configurally'; a finding that is difficult to reconcile with the physiological models. Here, we show that after controlling for subjects' expectations, there is no difference between 'featurally' and 'configurally' transformed faces in terms of inversion effect. This result reinforces the plausibility of simple hierarchical models of object representation and recognition in the cortex.

Face↗

Intracellular measurements of spatial integration and the MAX operation in complex cells of the cat primary visual cortex.

We have examined the spatial integration properties of complex cells to determine whether some of their responses can be described by a maximum operation (MAX)-like computation, as suggested by Riesenhuber and Poggio's model of object recognition. Membrane potential was recorded from anesthetized cats while optimally oriented bars were presented, either alone or in pairs, in different parts of the cells' receptive field. In most cells, the membrane potential response to two bars presented simultaneously could not be predicted by the sum of the responses to individual bars. In many cells, however, the responses closely approximated a MAX-like model. That is, the response of the cell to two bars was similar to the larger of the two individual responses ("soft-MAX"). The degree of nonlinear summation varied from cell to cell and varied within single cells from one stimulus configuration to another but on average fit most closely to the MAX model. The firing response of the cells was also well predicted by the MAX-like model. The MAX-like behavior was independent of the distance between the bars (orthogonal to the preferred orientation), independent of the relative amplitude of the responses, and slightly less pronounced at low levels of contrast. This MAX-like behavior of a subset of complex cells may play an important role in invariant object recognition in clutter.

Animals↗

A comparison of primate prefrontal and inferior temporal cortices during visual categorization.

Previous studies have suggested that both the prefrontal cortex (PFC) and inferior temporal cortex (ITC) are involved in high-level visual processing and categorization, but their respective roles are not known. To address this, we trained monkeys to categorize a continuous set of visual stimuli into two categories, "cats" and "dogs." The stimuli were parametrically generated using a computer graphics morphing system (Sheltonelton, 2000) that allowed precise control over stimulus shape. After training, we recorded neural activity from the PFC and the ITC of monkeys while they performed a category-matching task. We found that the PFC and the ITC play distinct roles in category-based behaviors: the ITC seems more involved in the analysis of currently viewed shapes, whereas the PFC showed stronger category signals, memory effects, and a greater tendency to encode information in terms of its behavioral meaning.

Analysis of Variance↗

Neural mechanisms of object recognition.

Single-unit recordings from behaving monkeys and human functional magnetic resonance imaging studies have continued to provide a host of experimental data on the properties and mechanisms of object recognition in cortex. Recent advances in object recognition, spanning issues regarding invariance, selectivity, representation and levels of recognition have allowed us to propose a putative model of object recognition in cortex.

Animals↗

Visual categorization and the primate prefrontal cortex: neurophysiology and behavior.

The ability to group stimuli into meaningful categories is a fundamental cognitive process. To explore its neuronal basis, we trained monkeys to categorize computer-generated stimuli as "cats" and "dogs." A morphing system was used to systematically vary stimulus shape and precisely define a category boundary. Psychophysical testing and analysis of eye movements suggest that the monkeys categorized the stimuli by attending to multiple stimulus features. Neuronal activity in the lateral prefrontal cortex reflected the category of visual stimuli and changed with learning when a monkey was retrained with the same stimuli assigned to new categories. Further, many neurons showed activity that appeared to reflect the monkey's decision about whether two stimuli were from the same category or not. These results suggest that the lateral prefrontal cortex is an important part of the neuronal circuitry underlying category learning and category-based behaviors.

Action Potentials↗