Neuroscience: making faces in the brain.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to James J DiCarlo.
Explore the source record for details and available documents.
The capability of the adult primate visual system for rapid and accurate recognition of targets in cluttered, natural scenes far surpasses the abilities of state-of-the-art artificial vision systems. Understanding this capability remains a fundamental challenge in visual neuroscience. Recent experimental evidence suggests that adaptive coding strategies facilitated by underlying neural plasticity enable the adult brain to learn from visual experience and shape its ability to integrate and recognize coherent visual objects.
Local field potentials (LFPs) arise largely from dendritic activity over large brain regions and thus provide a measure of the input to and local processing within an area. We characterized LFPs and their relationship to spikes (multi and single unit) in monkey inferior temporal cortex (IT). LFP responses in IT to complex objects showed strong selectivity at 44% of the sites and tolerance to retinal position and size. The LFP preferences were poorly predicted by the spike preferences at the same site but were better explained by averaging spikes within approximately 3 mm. A comparison of separate sites suggests that selectivity is similar on a scale of approximately 800 microm for spikes and approximately 5 mm for LFPs. These observations imply that inputs to IT neurons convey selectivity for complex shapes and that such input may have an underlying organization spanning several millimeters.
Understanding the brain computations leading to object recognition requires quantitative characterization of the information represented in inferior temporal (IT) cortex. We used a biologically plausible, classifier-based readout technique to investigate the neural coding of selectivity and invariance at the IT population level. The activity of small neuronal populations (approximately 100 randomly selected cells) over very short time intervals (as small as 12.5 milliseconds) contained unexpectedly accurate and robust information about both object "identity" and "category." This information generalized over a range of object positions and scales, even for novel objects. Coarse information about position and scale could also be read out from the same population.
The highest stages of the visual ventral pathway are commonly assumed to provide robust representation of object identity by disregarding confounding factors such as object position, size, illumination, and the presence of other objects (clutter). However, whereas neuronal responses in monkey inferotemporal cortex (IT) can show robust tolerance to position and size changes, previous work shows that responses to preferred objects are usually reduced by the presence of nonpreferred objects. More broadly, we do not yet understand multiple object representation in IT. In this study, we systematically examined IT responses to pairs and triplets of objects in three passively viewing monkeys across a broad range of object effectiveness. We found that, at least under these limited clutter conditions, a large fraction of the response of each IT neuron to multiple objects is reliably predicted as the average of its responses to the constituent objects in isolation. That is, multiple object responses depend primarily on the relative effectiveness of the constituent objects, regardless of object identity. This average effect becomes virtually perfect when populations of IT neurons are pooled. Furthermore, the average effect cannot simply be explained by attentional shifts but behaves as a primarily feedforward response property. Together, our observations are most consistent with mechanistic models in which IT neuronal outputs are normalized by summed synaptic drive into IT or spiking activity within IT and suggest that normalization mechanisms previously revealed at earlier visual areas are operating throughout the ventral visual stream.
While it is often assumed that objects can be recognized irrespective of where they fall on the retina, little is known about the mechanisms underlying this ability. By exposing human subjects to an altered world where some objects systematically changed identity during the transient blindness that accompanies eye movements, we induced predictable object confusions across retinal positions, effectively 'breaking' position invariance. Thus, position invariance is not a rigid property of vision but is constantly adapting to the statistics of the environment.
We describe a new technique that uses the timing of neuronal and behavioral responses to explore the contributions of individual neurons to specific behaviors. The approach uses both the mean neuronal latency and the trial-by-trial covariance between neuronal latency and behavioral response. Reliable measurements of these values were obtained from single-unit recordings made from anterior inferotemporal (AIT) cortex and the frontal eye fields (FEF) in monkeys while they performed a choice reaction time task. These neurophysiological data show that the responses of AIT neurons and some FEF neurons have little covariance with behavioral response, consistent with a largely "sensory" response. The responses of another group of FEF neurons with longer mean latency covary tightly with behavioral response, consistent with a largely "motor" response. A very small fraction of FEF neurons had responses consistent with an intermediate position in the sensory-motor pathway. These results suggest that this technique is a valuable tool for exploring the functional organization of neuronal circuits that underlie specific behaviors.
Visual object recognition is computationally difficult because changes in an object's position, distance, pose, or setting may cause it to produce a different retinal image on each encounter. To robustly recognize objects, the primate brain must have mechanisms to compensate for these variations. Although these mechanisms are poorly understood, it is thought that they elaborate neuronal representations in the inferotemporal cortex that are sensitive to object form but substantially invariant to other image variations. This study examines this hypothesis for image variation resulting from changes in object position. We studied the effect of small differences (+/-1.5 degrees ) in the retinal position of small (0.6 degrees wide) visual forms on both the behavior of monkeys trained to identify those forms and the responses of 146 anterior IT (AIT) neurons collected during that behavior. Behavioral accuracy and speed were largely unaffected by these small changes in position. Consistent with previous studies, many AIT responses were highly selective for the forms. However, AIT responses showed far greater sensitivity to retinal position than predicted from their reported receptive field (RF) sizes. The median AIT neuron showed a approximately 60% response decrease between positions within +/-1.5 degrees of the center of gaze, and 52% of neurons were unresponsive to one or more of these positions. Consistent with previous studies, each neuron's rank order of target preferences was largely unaffected across position changes. Although we have not yet determined the conditions necessary to observe this marked position sensitivity in AIT responses, we rule out effects of spatial-frequency content, eye movements, and failures to include the RF center. To reconcile this observation with previous studies, we hypothesize that either AIT position sensitivity strongly depends on object size or that position sensitivity is sharpened by extensive visual experience at fixed retinal positions or by the presence of flanking distractors.
More than 350 neurons with fingerpad receptive fields (RFs) were studied in cortical area 3b of three alert monkeys. Random dot patterns, which contain all stimulus patterns with equal probability, were scanned across these RFs at three velocities and eight directions to reveal the RFs' spatial and temporal structure. Area 3b RFs are characterized by three components: (1) a single, central excitatory region of short duration, (2) one or more inhibitory regions, also of short duration, that are adjacent to and nearly synchronous with the excitation, and (3) a region of inhibition that overlaps the excitation partially or totally and is temporally delayed with respect to the first two components. As a result of these properties, RF spatial structure depends on scanning direction but is virtually unaffected by changes in scanning velocity. This RF characterization, which is derived solely from responses to scanned random-dot patterns, predicts a neuron's responses to random patterns accurately, as expected, but it also predicts orientation sensitivity and preferred orientation measured with a scanned bar. Both orientation sensitivity and the ratio of coincident inhibition (number 2 above) to excitation are stronger in the supra- and infragranular layers than in layer IV.