PubMed Health⌕ Search

Biomedical subjects

DeLiang Wang

Publications and source records attributed to DeLiang Wang.

6 recordsLinked to original sources

Speech segregation based on sound localization.

At a cocktail party, one can selectively attend to a single voice and filter out all the other acoustical interferences. How to simulate this perceptual ability remains a great challenge. This paper describes a novel, supervised learning approach to speech segregation, in which a target speech signal is separated from interfering sounds using spatial localization cues: interaural time differences (ITD) and interaural intensity differences (IID). Motivated by the auditory masking effect, the notion of an "ideal" time-frequency binary mask is suggested, which selects the target if it is stronger than the interference in a local time-frequency (T-F) unit. It is observed that within a narrow frequency band, modifications to the relative strength of the target source with respect to the interference trigger systematic changes for estimated ITD and IID. For a given spatial configuration, this interaction produces characteristic clustering in the binaural feature space. Consequently, pattern classification is performed in order to estimate ideal binary masks. A systematic evaluation in terms of signal-to-noise ratio as well as automatic speech recognition performance shows that the resulting system produces masks very close to ideal binary ones. A quantitative comparison shows that the model yields significant improvement in performance over an existing approach. Furthermore, under certain conditions the model produces large speech intelligibility improvements with normal listeners.

Adult↗

The role of priming in conjunctive visual search.

To assess the role of priming in conjunctive visual search tasks, we systematically varied the consistency of the target and distractor identity between different conditions. Search was fastest in the standard conjunctive search paradigm where identities remained constant. Search was slowest when potential target identity varied predictably for each successive trial (the 'switch' condition). The role of priming was also demonstrated on a trial-by-trial basis in a 'streak' condition where target and distractor identity was unpredictable yet was consistent within streaks. When the target to be found was the same for a few trials in a row, search performance became similar to that when the potential target was the same on all trials. A similar pattern was found for the target absent trials, suggesting that priming is based on the whole search array rather than just the target in each case. Further analysis indicated that the effects of priming are sufficiently strong to account for the advantage seen for the conjunctive search task. We conclude that the role of priming in visual search is underestimated in current theories of visual search and that differences in search times often attributed to top-down guidance may instead reflect the benefits of priming.

Attention↗

A spectral histogram model for texton modeling and texture discrimination.

We suggest a spectral histogram, defined as the marginal distribution of filter responses, as a quantitative definition for a texton pattern. By matching spectral histograms, an arbitrary image can be transformed to an image with similar textons to the observed. We use the chi(2)-statistic to measure the difference between two spectral histograms, which leads to a texture discrimination model. The performance of the model well matches psychophysical results on a systematic set of texture discrimination data and it exhibits the nonlinearity and asymmetry phenomena in human texture discrimination. A quantitative comparison with the Malik-Perona model is given, and a number of issues regarding the model are discussed.

Discrimination, Psychological↗

A dynamically coupled neural oscillator network for image segmentation.

We propose a dynamically coupled neural oscillator network for image segmentation. Instead of pair-wise coupling, an ensemble of oscillators coupled in a local region is used for grouping. We introduce a set of neighborhoods to generate dynamical coupling structures associated with a specific oscillator. Based on the proximity and similarity principles, two grouping rules are proposed to explicitly consider the distinct cases of whether an oscillator is inside a homogeneous image region or near a boundary between different regions. The use of dynamical coupling makes our segmentation network robust to noise on an image, and unlike image processing algorithms no iterative operation is needed for noise removal. For fast computation, a segmentation algorithm is abstracted from the underlying oscillatory dynamics, and has been applied to synthetic and real images. Simulation results demonstrate the effectiveness of our oscillator network in image segmentation.

Algorithms↗

An oscillatory correlation model of visual motion analysis.

We describe and evaluate a model of motion perception based on the integration of information from two parallel pathways: a motion pathway and a luminance pathway. The motion pathway has two stages. The first stage measures and pools local motion across the input animation sequence and assigns reliability indices to these pooled measurements. The second stage groups locations on the basis of these measurements. In the luminance pathway, the input scene is segmented into regions on the basis of similarities in luminance. In a subsequent integration stage, motion and luminance segments are combined to obtain the final estimates of object motion. The neural network architecture we employ is based on LEGION (locally excitatory globally inhibitory oscillator networks), a scheme for feature binding and region labeling based on oscillatory correlation. Many aspects of the model are implemented at the neural network level, whereas others are implemented at a more abstract level. We apply this model to the computation of moving, uniformly illuminated, two-dimensional surfaces that are either opaque or transparent. Model performance replicates a number of distinctive features of human motion perception.

Humans↗

Intrinsic generalization analysis of low dimensional representations.

Low dimensional representations of images impose equivalence relations in the image space; the induced equivalence class of an image is named as its intrinsic generalization. The intrinsic generalization of a representation provides a novel way to measure its generalization and leads to more fundamental insights than the commonly used recognition performance, which is heavily influenced by the choice of training and test data. We demonstrate the limitations of linear subspace representations by sampling their intrinsic generalization, and propose a nonlinear representation that overcomes these limitations. The proposed representation projects images nonlinearly into the marginal densities of their filter responses, followed by linear projections of the marginals. We use experiments on large datasets to show that the representations that have better intrinsic generalization also lead to better recognition performance.

Generalization, Psychological↗