PubMed Health⌕ Search

Biomedical subjects

Denis G Pelli

Publications and source records attributed to Denis G Pelli.

9 recordsLinked to original sources

Feature detection and letter identification.

Seeking to understand how people recognize objects, we have examined how they identify letters. We expected this 26-way classification of familiar forms to challenge the popular notion of independent feature detection ("probability summation"), but find instead that this theory parsimoniously accounts for our results. We measured the contrast required for identification of a letter briefly presented in visual noise. We tested a wide range of alphabets and scripts (English, Arabic, Armenian, Chinese, Devanagari, Hebrew, and several artificial ones), three- and five-letter words, and various type styles, sizes, contrasts, durations, and eccentricities, with observers ranging widely in age (3 to 68) and experience (none to fluent). Foreign alphabets are learned quickly. In just three thousand trials, new observers attain the same proficiency in letter identification as fluent readers. Surprisingly, despite this training, the observers-like clinical letter-by-letter readers-have the same meager memory span for random strings of these characters as observers seeing them for the first time. We compare performance across tasks and stimuli that vary in difficulty by pitting the human against the ideal observer, and expressing the results as efficiency. We find that efficiency for letter identification is independent of duration, overall contrast, and eccentricity, and only weakly dependent on size, suggesting that letters are identified by a similar computation across this wide range of viewing conditions. Efficiency is also independent of age and years of reading. However, efficiency does vary across alphabets and type styles, with more complex forms yielding lower efficiencies, as one might expect from Gestalt theories of perception. In fact, we find that efficiency is inversely proportional to perimetric complexity (perimeter squared over "ink" area) and nearly independent of everything else. This, and the surprisingly fixed ratio of detection and identification thresholds, indicate that identifying a letter is mediated by detection of about 7 visual features.

Adolescent↗

Noise masking reveals channels for second-order letters.

We investigate the channels underlying identification of second-order letters using a critical-band masking paradigm. We find that observers use a single 1-1.5 octave-wide channel for this task. This channel's best spatial frequency (c/letter) did not change across different noise conditions (indicating the inability of observers to switch channels to improve signal-to-noise ratio) or across different letter sizes (indicating scale invariance), for a fixed carrier frequency (c/letter). However, the channel's best spatial frequency does change with stimulus carrier frequency (both in c/letter); one is proportional to the other. Following Majaj et al. (Majaj, N. J., Pelli, D. G., Kurshan, P., & Palomares, M. (2002). The role of spatial frequency channels in letter identification. Vision Research, 42, 1165-1184), we define "stroke frequency" as the line frequency (strokes/deg) in the luminance image. That is, for luminance-defined letters, stroke frequency is the number of lines (strokes) across each letter divided by letter width. For second-order letters, letter texture stroke frequency is the number of carrier cycles (luminance lines) within the letter ink area divided by the letter width. Unlike the nonlinear dependence found for first-order letters (implying scale-dependent processing), for second-order letters the channel frequency is half the letter texture stroke frequency (suggesting scale-invariant processing).

Contrast Sensitivity↗

Are faces processed like words? A diagnostic test for recognition by parts.

Do we identify an object as a whole or by its parts? This simple question has been surprisingly hard to answer. It has been suggested that faces are recognized as wholes and words are recognized by parts. Here we answer the question by applying a test for crowding. In crowding, a target is harder to identify in the presence of nearby flankers. Previous work has described crowding between objects. We show that crowding also occurs between the parts of an object. Such internal crowding severely impairs perception, identification, and fMRI face-area activation. We apply a diagnostic test for crowding to a word and a face, and we find that the critical spacing of the parts required for recognition is proportional to distance from fixation and independent of size and kind. The critical spacing defines an isolation field around the target. Some objects can be recognized only when each part is isolated from the rest of the object by the critical spacing. In that case, recognition is by parts. Recognition is holistic if the observer can recognize the object even when the whole object fits within a critical spacing. Such an object has only one part. Multiple parts within an isolation field will crowd each other and spoil recognition. To assess the robustness of the crowding test, we manipulated familiarity through inversion and the face- and word-superiority effects. We find that threshold contrast for word and face identification is the product of two factors: familiarity and crowding. Familiarity increases sensitivity by a factor of x1.5, independent of eccentricity, while crowding attenuates sensitivity more and more as eccentricity increases. Our findings show that observers process words and faces in much the same way: The effects of familiarity and crowding do not distinguish between them. Words and faces are both recognized by parts, and their parts -- letters and facial features -- are recognized holistically. We propose that internal crowding be taken as the signature of recognition by parts.

Contrast Sensitivity↗

Crowding is unlike ordinary masking: distinguishing feature integration from detection.

A letter in the peripheral visual field is much harder to identify in the presence of nearby letters. This is "crowding." Both crowding and ordinary masking are special cases of "masking," which, in general, refers to any effect of a "mask" pattern on the discriminability of a signal. Here we characterize crowding, and propose a diagnostic test to distinguish it from ordinary masking. In ordinary masking, the signal disappears. In crowding, it remains visible, but is ambiguous, jumbled with its neighbors. Masks are usually effective only if they overlap the signal, but the crowding effect extends over a large region. The width of that region is proportional to signal eccentricity from the fovea and independent of signal size, mask size, mask contrast, signal and mask font, and number of masks. At 4 deg eccentricity, the threshold contrast for identification of a 0.32 deg signal letter is elevated (up to six-fold) by mask letters anywhere in a 2.3 deg region, 7 times wider than the signal. In ordinary masking, threshold contrast rises as a power function of mask contrast, with a shallow log-log slope of 0.5 to 1, whereas, in crowding, threshold is a sigmoidal function of mask contrast, with a steep log-log slope of 2 at close spacing. Most remarkably, although the threshold elevation decreases exponentially with spacing, the threshold and saturation contrasts of crowding are independent of spacing. Finally, ordinary masking is similar for detection and identification, but crowding occurs only for identification, not detection. More precisely, crowding occurs only in tasks that cannot be done based on a single detection by coarsely coded feature detectors. These results (and observers' introspections) suggest that ordinary masking blocks feature detection, so the signal disappears, while crowding (like "illusory conjunction") is excessive feature integration - detected features are integrated over an inappropriately large area because there are no smaller integration fields - so the integrated signal is ambiguous, jumbled with the mask. In illusory conjunction, observers see an object that is not there made up of features that are. A survey of the illusory conjunction literature finds that most of the illusory conjunction results are consistent with the spatial crowding described here, which depends on spatial proximity, independent of time pressure. The rest seem to arise through a distinct phenomenon that one might call "temporal crowding," which depends on time pressure ("overloading attention"), independent of spatial proximity.

Contrast Sensitivity↗

Using visual noise to characterize amblyopic letter identification.

Amblyopia is a much-studied but poorly understood developmental visual disorder that reduces acuity, profoundly reducing contrast sensitivity for small targets. Here we use visual noise to probe the letter identification process and characterize its impairment by amblyopia. We apply five levels of analysis - threshold, threshold in noise, equivalent noise, optical MTF, and noise modeling - to obtain a two-factor model of the amblyopic deficit: substantially reduced efficiency for small letters and negligibly increased cortical noise. Cortical noise, expressed as an equivalent input noise, varies among amblyopes but is roughly 1.4x normal, as though only 1/1.4 the normal number of cortical spikes are devoted to the amblyopic eye. This raises threshold contrast for large letters by a factor of radical1.4 = 1.2x, a negligible effect. All 16 amblyopic observers showed near-normal efficiency for large letters (> 4x acuity) and greatly reduced efficiency for small letters: 1/4 normal at 2x acuity and approaching 1/16 normal at acuity. Finding that the acuity loss represents a loss of efficiency rules out all models of amblyopia except those that predict the same sensitivity loss on blank and noisy backgrounds. One such model is the last-channel hypothesis, which supposes that the highest-spatial-frequency channels are missing, leaving the remaining highest-frequency channel struggling to identify the smallest letters. However, this hypothesis is rejected by critical band masking of letter identification, which shows that the channels used by the amblyopic eye have normal tuning for even the smallest letters. Finally, based on these results, we introduce a new "Dual Acuity" chart that promises to be a quick diagnostic test for amblyopia.

Adult↗

Covert attention enhances letter identification without affecting channel tuning.

Directing covert attention to the target location enhances sensitivity, but it is not clear how this enhancement comes about. Knowing that a single spatial frequency channel mediates letter identification, we use the critical-band-masking paradigm to investigate whether covert attention affects the spatial frequency tuning of that channel. We find that directing attention to the target location halves threshold energy without affecting the channel's spatial frequency tuning.

Attention↗

The remarkable inefficiency of word recognition.

Do we recognize common objects by parts, or as wholes? Holistic recognition would be efficient, yet people detect a grating of light and dark stripes by parts. Thus efficiency falls as the number of stripes increases, in inverse proportion, as explained by probability summation among independent feature detectors. It is inefficient to detect correlated components independently. But gratings are uncommon artificial stimuli that may fail to tap the full power of visual object recognition. Familiar objects become special as people become expert at judging them, possibly because the processing becomes more holistic. Letters and words were designed to be easily recognized, and, through a lifetime of reading, our visual system presumably has adapted to do this as well as it possibly can. Here we show that in identifying familiar English words, even the five most common three-letter words, observers have the handicap predicted by recognition by parts: a word is unreadable unless its letters are separately identifiable. Efficiency is inversely proportional to word length, independent of how many possible words (5, 26 or thousands) the test word is drawn from. Human performance never exceeds that attainable by strictly letter- or feature-based models. Thus, everything seen is a pattern of features. Despite our virtuosity at recognizing patterns and our expertise from reading a billion letters, we never learn to see a word as a feature; our efficiency is limited by the bottleneck of having to rigorously and independently detect simple features.

Contrast Sensitivity↗

Flicker flutter: is an illusory event as good as the real thing?

Verghese and Stone (1995) showed that reducing the perceived number of objects by grouping also reduces objective performance. Shams, Kamitani, and Shimojo (2000) showed that a single flash accompanied by multiple beeps appears to flash more than once. We show that objective orientation-discrimination performance depends solely on the perceived number of flashes, independent of the actual number of beeps and flashes. Thus the unit of perceptual analysis seems to be a perceived event, independent of how it is induced.

Auditory Perception↗

The role of spatial frequency channels in letter identification.

How we see is today explained by physical optics and retinal transduction, followed by feature detection, in the cortex, by a bank of parallel independent spatial-frequency-selective channels. It is assumed that the observer uses whichever channels are best for the task at hand. Our current results demand a revision of this framework: Observers are not free to choose which channels they use. We used critical-band masking to characterize the channels mediating identification of broadband signals: letters in a wide range of fonts (Sloan, Bookman, Künstler, Yung), alphabets (Roman and Chinese), and sizes (0.1-55 degrees ). We also tested sinewave and squarewave gratings. Masking always revealed a single channel, 1.6+/-0.7 octaves wide, with a center frequency that depends on letter size and alphabet. We define an alphabet's stroke frequency as the average number of lines crossed by a slice through a letter, divided by the letter width. For sharp-edged (i.e. broadband) signals, we find that stroke frequency completely determines channel frequency, independent of alphabet, font, and size. Moreover, even though observers have multiple channels, they always use the same channel for the same signals, even after hundreds of trials, regardless of whether the noise is low-pass, high-pass, or all-pass. This shows that observers identify letters through a single channel that is selected bottom-up, by the signal, not top-down by the observer. We thought shape would be processed similarly at all sizes. Bandlimited signals conform more to this expectation than do broadband signals. Here, we characterize processing by channel frequency. For sinewave gratings, as expected, channel frequency equals sinewave frequency f(channel)=f. For bandpass-filtered letters, channel frequency is proportional to center frequency f(channel) proportional, variantf(center) (log-log slope 1) when size is varied and the band (c/letter) is fixed, but channel frequency is less than proportional to center frequency f(channel) proportional, variantf(center)(2/3) (log-log slope 2/3) when the band is varied and size is fixed. Finally, our main result, for sharp-edged (i.e. broadband) letters and squarewaves, channel frequency depends solely on stroke frequency, f(channel)/10c/deg=(2/3), with a log-log slope of 2/3. Thus, large letters (and coarse squarewaves) are identified by their edges; small letters (and fine squarewaves) are identified by their gross strokes.

Contrast Sensitivity↗