PubMed HealthSearch

Biomedical subjects

G Sperling

Publications and source records attributed to G Sperling.

At least 19 recordsLinked to original sources

Object spatial frequencies, retinal spatial frequencies, noise, and the efficiency of letter discrimination.

To determine which spatial frequencies are most effective for letter identification, and whether this is because letters are objectively more discriminable in these frequency bands or because can utilize the information more efficiently, we studied the 26 upper-case letters of English. Six two-octave wide filters were used to produce spatially filtered letters with 2D-mean frequencies ranging from 0.4 to 20 cycles per letter height. Subjects attempted to identify filtered letters in the presence of identically filtered, added Gaussian noise. The percent of correct letter identifications vs s/n (the root-mean-square ratio of signal to noise power) was determined for each band at four viewing distances ranging over 32:1. Object spatial frequency band and s/n determine presence of information in the stimulus; viewing distance determines retinal spatial frequency, and affects only ability to utilize. Viewing distance had no effect upon letter discriminability: object spatial frequency, not retinal spatial frequency, determined discriminability. To determine discrimination efficiency, we compared human discrimination to an ideal discriminator. For our two-octave wide bands, s/n performance of humans and of the ideal detector improved with frequency mainly because linear bandwidth increased as a function of frequency. Relative to the ideal detector, human efficiency was 0 in the lowest frequency bands, reached a maximum of 0.42 at 1.5 cycles per object and dropped to about 0.104 in the highest band. Thus, our subjects best extract upper-case letter information from spatial frequencies of 1.5 cycles per object height, and they can extract it with equal efficiency over a 32:1 range of retinal frequencies, from 0.074 to more than 2.3 cycles per degree of visual angle.

Adult

The kinetic depth effect and optic flow--II. First- and second-order motion.

We use a difficult shape identification task to analyze how humans extract 3D surface structure from dynamic 2D stimuli--the kinetic depth effect (KDE). Stimuli composed of luminous tokens moving on a less luminous background yield accurate 3D shape identification regardless of the particular token used (either dots, lines, or disks). These displays stimulate both the 1st-order (Fourier-energy) motion detectors and 2nd-order (nonFourier) motion detectors. To determine which system supports KDE, we employ stimulus manipulations that weaken or distort 1st-order motion energy (e.g. frame-to-frame alternation of the contrast polarity of tokens) and manipulations that create microbalanced stimuli which have no useful 1st-order motion energy. All manipulations that impair 1st-order motion energy correspondingly impair 3D shape identification. In certain cases, 2nd-order motion could support limited KDE, but it was not robust and was of low spatial resolution. We conclude that 1st-order motion detectors are the primary input to the kinetic depth system. To determine minimal conditions for KDE, we use a two frame display. Under optimal conditions, KDE supports shape identification performance at 63-94% of full-rotation displays (where baseline is 5%). Increasing the amount of 3D rotation portrayed or introducing a blank inter-stimulus interval impairs performance. Together, our results confirm that the human KDE computation of surface shape uses a global optic flow computed primarily by 1st-order motion detectors with minor 2nd-order inputs. Accurate 3D shape identification requires only two views and therefore does not require knowledge of acceleration.

Depth Perception

The visible persistence of stimuli in stroboscopic motion.

This paper reports an improved paradigm to measure visible persistence. The stimulus is a pair of lines stroboscopically displayed in successive positions moving in opposite directions. The subjects' judgement of simultaneous appearance of all the presented lines is used to estimate visible persistence. This paradigm permitted independent manipulation of spatial and temporal stimulus separations in linear motion. The resulting estimates of visible persistence increase with spatial separation up to 0.24 deg of visual angle and approaches a maximum value at larger spatial separations. The results are consistent with the existence of a hypothetical visual gain mechanism that operates over small retinal distances to effectively decrease persistence duration with decreasing spatial separation.

Afterimage

Intelligent temporal subsampling of American Sign Language using event boundaries.

How well can a sequence of frames be represented by a subset of the frames? Video sequences of American Sign Language (ASL) were investigated in two modes: dynamic (ordinary video) and static (frames printed side by side on the display). An activity index was used to choose critical frames at event boundaries, times when the difference between successive frames is at a local minimum. Sign intelligibility was measured for 32 experienced ASL signers who viewed individual signs. For full gray-scale dynamic signs activity-index subsampling yielded sequences that were significantly more intelligible than when every mth frame was chosen. This result was even more pronounced for static images. For binary images, the relative advantage of activity subsampling was smaller. We conclude that event boundaries can be defined computationally and that subsampling from event boundaries is better than choosing at regular intervals.

Adolescent

[Stenosing ureteritis and factor XIII deficiency in anaphylactoid purpura].

A case report of a 6-year-old boy is presented. The patient suffered from a severe Schönlein-Henoch purpura. It was demonstrated that ureteral stenosis develops during the clinical course of this disease. This complication has to be considered in anaphylactoid purpura, since it is usually self-limiting and does not require surgical intervention. This confirms once again the necessity to look for a decrease of factor XIII activity in these patients and of the value of substituting this compound if there are severe abdominal complaints. The theoretical background of this therapeutical intervention is discussed.

Child

How to study the kinetic depth effect experimentally.

Sperling, Landy, Dosher, and Perkins (1989) proposed an objective 3D shape identification task with 2D artifactual cues removed and with full feedback (FB) to the subjects to measure KDE and to circumvent algorithmically equivalent KDE-alternative computations and artifactual non-KDE processing. (1) The 2D velocity flow-field was necessary and sufficient for true KDE. (2) Only the first-order (Fourier-based) perceptual motion system could solve our task because the second-order (rectifying) system could not simultaneously process more than two locations. (3) To ensure first-order motion processing, KDE tasks must require simultaneous processing at more than two locations. (4) Practice with FB is essential to measure ultimate capacity (aptitude) and, thereby, to enable comparisons with ideal observers. Experiments without FB measure ecological achievement--the ability of subjects to extrapolate their past experience to the current stimuli.

Algorithms

Kinetic depth effect and optic flow--I. 3D shape from Fourier motion.

Fifty-three different 3D shapes were defined by sequences of 2D views (frames) of dots on a rotating 3D surface. (1) Subjects' accuracy of shape identifications dropped from over 90% to less than 10% when either the polarity of the stimulus dots was alternated from light-on-gray to dark-on-gray on successive frames or when neutral gray interframe intervals were interposed. Both manipulations interfere with motion extraction by spatio-temporal (Fourier) and gradient first-order detectors. Second-order (non-Fourier) detectors that use full-wave rectification are unaffected by alternating-polarity but disrupted by interposed gray frames. (2) To equate the accuracy of two-alternative forced-choice (2AFC) planar direction-of-motion discrimination in standard and polarity-alternated stimuli, standard contrast was reduced. 3D shape discrimination survived contrast reduction in standard stimuli whereas it failed completely with polarity-alternation even at full contrast. (3) When individual dots were permitted to remain in the image sequence for only two frames, performance showed little loss compared to standard displays where individual dots had an expected lifetime of 20 frames, showing that 3D shape identification does not require continuity of stimulus tokens. (4) Performance in all discrimination tasks is predicted (up to a monotone transformation) by considering the quality of first-order information (as given by a simple computation on Fourier power) and the number of locations at which motion information is required. Perceptual first-order analysis of optic flow is the primary substrate for structure-from-motion computations in random dot displays because only it offers sufficient quality of perceptual motion at a sufficient number of locations.

Depth Perception

Ratings of kinetic depth in multidot displays.

Subjects saw kinetic depth displays whose shape (sphere or cylinder) was defined by luminous dots distributed randomly on the surface or in the volume of the object. Subjects rated perceived 3-D depth, rigidity, and coherence. Despite individual differences, all 3 ratings increased with the number of dots. Dots in the volume yielded ratings equal to or greater than surface dots. Each rating varied with 3 of 4 factors (shape, distribution, numerosity, and perspective), but the ratings either between trials or between conditions were often uncorrelated. Object shape affected rigidity but not depth ratings. Veridically perceived polar displays had slightly lower rigidity but higher depth ratings than parallel projection displays. (Reversed polar displays were always grossly nonrigid.) The interaction of ratings and stimulus parameters requires theories and experiments in which different KDE ratings are not treated interchangeably.

Attention

Kinetic depth effect and identification of shape.

We introduce an objective shape-identification task for measuring the kinetic depth effect (KDE). A rigidly rotating surface consisting of hills and valleys on an otherwise flat ground was defined by 300 randomly positioned dots. On each trial, 1 of 53 shapes was presented; the observer's task was to identify the shape and its overall direction of rotation. Identification accuracy was an objective measure, with a low guessing base rate of the observer's perceptual ability to extract 3D structure from 2D motion via KDE. (1) Objective accuracy data were consistent with previously obtained subjective rating judgments of depth and coherence. (2) Along with motion cues, rotating real 3D dot-defined shapes inevitably produced a cue of changing dot density. By shortening dot lifetimes to control dot density, we showed that changing density was neither necessary nor sufficient to account for accuracy; motion alone sufficed. (3) Our shape task was solvable with motion cues from the 6 most relevant locations. We extracted the dots from these locations and used them in a simplified 2D direction-labeling motion task with 6 perceptually flat flow fields. Subjects' performance in the 2D and 3D tasks was equivalent, indicating that the information processing capacity of KDE is not unique. (4) Our proposed structure-from-motion algorithm for the shape task first finds relative minima and maxima of local velocity and then assigns 3D depths proportional to velocity.

Adult

Texture interactions determine perceived contrast.

For a patch of random visual texture embedded in a surrounding background of similar texture, we demonstrate that the perceived contrast of the texture patch depends substantially on the contrast of the background. When the texture patch is surrounded by high-contrast texture, the bright points of the texture patch appear dimmer, and simultaneously, its dark points appear less dark than when it is surrounded by a uniform background. The induced reduction of apparent contrast is greatly diminished when (i) the texture patch and background are filtered into nonoverlapping spatial frequency bands or (ii) the texture patch and background are presented to different eyes. Our results are unanticipated by all current theories of lightness perception and point to a perceptual mechanism for contrast gain control occurring at an early cortical or precortical neural locus.

Contrast Sensitivity

Three stages and two systems of visual processing.

Three stages of visual processing determine how internal noise appears to an external observer: light adaptation, contrast gain control and a postsensory/decision stage. Dark noise occurs prior to adaptation, determines dark-adapted absolute thresholds and mimics stationary external noise. Sensory noise occurs after dark adaptation, determines contrast thresholds for sine gratings and similar stimuli, and mimics external noise that increases with mean luminance. Postsensory noise incorporates perceptual, decision and mnemonic processes. It occurs after contrast-gain control and mimics external noise that increases with stimulus contrast (i.e., multiplicative noise). Dark noise and sensory noise are frequency specific and primarily affect weak signals. Only postsensory noise significantly affects the discriminability of strong signals masked by stimulus noise; postsensory noise has constant power over a wide spatial frequency range in which sensory noise varies enormously. Two parallel perceptual regimes jointly serve human object recognition and motion perception: a first-order linear (Fourier) regime that computes relations directly from stimulus luminance, and a second-order nonlinear (nonFourier) rectifying regime that uses the absolute value (or power) of stimulus contrast. When objects or movements are defined by high spatial frequencies (i.e., texture carrier frequencies whose wavelengths are small compared to the object size), the responses of high-frequency receptors are demodulated by rectification to facilitate discrimination at the higher processing levels. Rectification sacrifices the statistical efficiency (noise resistance) of the first-order regime for efficiency of neural connectivity and computation.

Adaptation, Ocular

Spatial-frequency bands in complex visual stimuli: American Sign Language.

Dynamic images of individual signs of American Sign Language (ASL) with a resolution of 96 X 64 pixels were bandpass filtered in adjacent frequency bands. Intelligibility was determined by testing deaf subjects fluent in ASL. The following results were obtained. (1) By iteratively varying the center frequencies and bandwidths of the spatial bandpass filters, it was possible to divide the original signal into four different component bands of high intelligibility. (2) The measured temporal-frequency spectrum was approximately the same in all bands. (3) The masking of signals in band i by noise in band j was found to be inversely proportional to log [f signal/f noise]. At constant performance, the ratio of root-mean-square signal amplitude to noise amplitude, s/n, was the same for bands 2,3, and 4 and higher for band 1. (4) When weak signals i and j were added linearly, there was a slight intelligibility advantage for signals in the same band (i = j) compared with signals in adjacent bands and for signals in adjacent bands compared with signals in distant bands.

Humans

Drift-balanced random stimuli: a general basis for studying non-Fourier motion perception.

To some degree, all current models of visual motion-perception mechanisms depend on the power of the visual signal in various spatiotemporal-frequency bands. Here we show how to construct counterexamples: visual stimuli that are consistently perceived as obviously moving in a fixed direction yet for which Fourier-domain power analysis yields no systematic motion components in any given direction. We provide a general theoretical framework for investigating non-Fourier motion-perception mechanisms; central are the concepts of drift-balanced and microbalanced random stimuli. A random stimulus S is drift balanced if its expected power in the frequency domain is symmetric with respect to temporal frequency, that is, if the expected power in S of every drifting sinusoidal component is equal to the expected power of the sinusoid of the same spatial frequency, drifting at the same rate in the opposite direction. Additionally, S is microbalanced if the result WS of windowing S by any space-time-separable function W is drift balanced. We prove that (i) any space-time-separable random (or nonrandom) stimulus is microbalanced; (ii) any linear combination of pairwise independent microbalanced (respectively, drift-balanced) random stimuli is microbalanced and drift balanced if the expectation of each component is uniformly zero; (iii) the convolution of independent microbalanced and drift-balanced random stimuli is microbalanced and drift balanced; (iv) the product of independent microbalanced random stimuli is microbalanced; and (v) the expected response of any Reichardt detector to any microbalanced random stimulus is zero at every instant in time. Examples are provided of classes of microbalanced random stimuli that display consistent and compelling motion in one direction. All the results and examples from the domain of motion perception are transposable to the space-domain problem of detecting orientation in a texture pattern.

Fourier Analysis

Dynamics of automatic and controlled visual attention.

The time course of attention was experimentally observed using two kinds of stimuli: a cue to begin attending or to shift attention, and a stimulus to be attended. Precise measurements of the time course of attention show that it consists of two partially concurrent processes: a fast, effortless, automatic process that records the cue and its neighboring events; and a slower, effortful, controlled process that records the stimulus to be attended and its neighboring events.

Acoustic Stimulation

Limits of visual communication: the effect of signal-to-noise ratio on the intelligibility of American Sign Language.

To determine the limits of human observers' ability to identify visually presented American Sign Language (ASL), the contrast s and the amount of additive noise n in dynamic ASL images were varied independently. Contrast was tested over a 4:1 range; the rms signal-to-noise ratios (s/n) investigated were s/n = 1/4, 1/2, 1, and infinity (which is used to designate the original, uncontaminated images). Fourteen deaf subjects were tested with an intelligibility test composed of 85 isolated ASL signs, each 2-3 sec in length. For these ASL signs (64 x 96 pixels, 30 frames/sec), subjects' performance asymptotes between s/n = 0.5 and 1.0; further increases in s/n do not improve intelligibility. Intelligibility was found to depend only on s/n and not on contrast. A formulation in terms of logistic functions was proposed to derive intelligibility of ASL signs from s/n, sign familiarity, and sign difficulty. Familiarity (ignorance) is represented by additive signal-correlated noise; it represents the likelihood of a subject's knowing a particular ASL sign, and it adds to s/n. Difficulty is represented by a multiplicative difficulty coefficient; it represents the perceptual vulnerability of an ASL sign to noise and it adds to log(s/n).

Adult

Tradeoffs between stereopsis and proximity luminance covariance as determinants of perceived 3D structure.

A 2D polar projection of a 3D wire cube (Necker cube) in clockwise rotation can be perceived either veridically as a clockwise-rotating cube (rigid percept) or as a counterclockwise-rotating rubbery, truncated pyramid (nonrigid percept). The 3D percept is influenced by various cues: linear perspective, stereo disparity, and proximity-luminance covariance (PLC, the intensification of edges in proportion to their proximity to the observer). Perspective, by itself or in combination, is a very weak cue whereas PLC is a powerful cue [Schwartz and Sperling (1983) Bull. Psychon. Soc. 21, 456-458]. Here we determined psychometric functions for perceptual resolution in static displays and dynamic rotating displays (with and without a static preview) as determined by stereopsis and PLC in isolation and with both cues jointly, possibly in conflict. Stereopsis was the dominant cue in static displays and in most dynamic displays. When a static display preceded a dynamic display, it strongly influenced the subsequent dynamic percept. Perceptual resolution in all conditions was accurately described by a winner-take-all model in which the strength of evidence for each percept from different cues is simply algebraically added.

Depth Perception

A signal-to-noise theory of the effects of luminance on picture memory: comment on Loftus.

In studies of picture memory, subjects typically view a sequence of pictures. Their memory is tested either after each picture is presented (short-term recall) or at the end of the sequence (long-term recall). The increase in performance as a function of picture viewing time defines "the rate of information acquisition." Loftus (1985) found that reducing the luminance of a picture reduces the rate at which information is acquired (for both short-term and long-term tests) and, for long viewing times, reduces the total amount of recall. The theory proposed here assumes that both of these effects are consequences of intrinsic noise in the visual system that becomes relatively more prominent as signal (picture luminance or contrast) is reduced. Noise shares a limited capacity channel with signal, and thus noise reduces the rate of information acquisition; noise, as well as signal, occupies space in memory, and thus noise reduces recall performance.

Attention