PubMed HealthSearch

PubMed · 8533342

Human efficiency for recognizing 3-D objects in luminance noise.

Abstract

The purpose of this study was to establish how efficiently humans use visual information to recognize simple 3-D objects. The stimuli were computer-rendered images of four simple 3-D objects--wedge, cone, cylinder, and pyramid--each rendered from 8 randomly chosen viewing positions as shaded objects, line drawings, or silhouettes. The objects were presented in static, 2-D Gaussian luminance noise. The observer's task was to indicate which of the four objects had been presented. We obtained human contrast thresholds for recognition, and compared these to an ideal observer's thresholds to obtain efficiencies. In two auxiliary experiments, we measured efficiencies for object detection and letter recognition. Our results showed that human object-recognition efficiency is low (3-8%) when compared to efficiencies reported for some other visual-information processing tasks. The low efficiency means that human recognition performance is limited primarily by factors intrinsic to the observer rather than the information content of the stimuli. We found three factors that play a large role in accounting for low object-recognition efficiency: stimulus size, spatial uncertainty, and detection efficiency. Four other factors play a smaller role in limiting object-recognition efficiency: observers' internal noise, stimulus rendering condition, stimulus familiarity, and categorization across views.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

B S Tjan, W L Braje, G E Legge, D Kersten. 1995. Human efficiency for recognizing 3-D objects in luminance noise.. https://doi.org/10.1016/0042-6989(95)00070-g

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Pitfalls in estimating motion detector receptive field geometry.

A number of psychophysical investigations have used spatial-summation methods to estimate the receptive field (RF) geometry of motion detectors by exploring how psychophysical thresholds change with stimulus height and/or width. This approach is based on the idea that an observer's ability to detect motion direction is strongly determined by the relationship between the stimulus geometry (height and width) and the RF of the activated motion detectors. Our results show that previous estimates of RF geometry can depend significantly on stimulus position in the visual field as well as on the stimulus height-to-width ratio. The data further show that RF estimates depend on the stimulus in a manner that is inconsistent with basic predictions derived from current motion detector models. Hence previous estimates of height, width, and height-to-width ratios of motion detector RFs are inaccurate and unreliable. This inaccuracy/unreliability is attributed to a number of sources. These include incorrect fixed-parameter values in model fits, as well as the confounding of physiological spatial summation area through combined use of contrast thresholds and Gaussian-windowed stimuli. A third source of error is an asymmetric variation of spatiotemporal correlation in the stimulus as either its height or width is varied (and the other dimension held constant). Most importantly, a fourth source of unreliability is attributed to the existence of a nonlinear, nonmonotonic distribution of motion detectors in the visual field that has been previously described and is a natural result of visual anatomy.

Contrast Sensitivity

Perceived location of bars and edges in one-dimensional images: computational models and human vision.

Observers used a cursor to mark the location and polarity of all the bar and edge features seen in compound (f + 3f) gratings of moderate frequency and contrast. They almost always reported six bars and six edges per cycle of the fundamental frequency (f = 0.4 c/deg, contrast 32%), for all phases of the third harmonic (3f = 1.2 c/deg, contrast 10.7%). This general pattern of features was predicted by the positions of peaks and troughs in the outputs of even and odd filters applied to the stimulus waveform, but not by peaks of "local energy" since there were only two energy peaks per cycle. We considered a family of filters whose amplitude spectrum has slope p on a log-log plot. The best-fitting filter slope was determined for bars (even filter) and edges (odd filter) in conjunction with a classification rule in which all peaks and troughs in the response profile are counted as features. If bars were seen at luminance peaks, and edges seen at gradient peaks (zero-crossings in the second derivative) we should have found p = 0 for bars and p = 1 for edges. In fact, for both bars and edges the best-fitting slope was about p = 0.5. For edges, this is consistent with the use of a smoothed (Gaussian) derivative operator. The filters form a quadrature pair, as in the energy model, but features are not constrained to lie at energy peaks. A compressive transducer preceding the filters improved the goodness-of-fit for predicted edge locations, but did not affect the estimate of filter slopes, nor the goodness-of-fit for bar locations. In an experiment with single blurred edges we confirmed that the perceived location of edges is shifted towards the darker side of the edge in direct proportion to the contrast of the edge. This was well predicted by adding a compressive transducer to the filter model.

Contrast Sensitivity