PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “computer vision”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Vision as Bayesian inference: analysis by synthesis?

We argue that the study of human vision should be aimed at determining how humans perform natural tasks with natural images. Attempts to understand the phenomenology of vision from artificial stimuli, although worthwhile as a starting point, can lead to faulty generalizations about visual systems, because of the enormous complexity of natural images. Dealing with this complexity is daunting, but Bayesian inference on structured probability distributions offers the ability to design theories of vision that can deal with the complexity of natural images, and that use 'analysis by synthesis' strategies with intriguing similarities to the brain. We examine these strategies using recent examples from computer vision, and outline some important implications for cognitive science.

Algorithms↗

Generalizing Swendsen-Wang to sampling arbitrary posterior probabilities.

Many vision tasks can be formulated as graph partition problems that minimize energy functions. For such problems, the Gibbs sampler provides a general solution but is very slow, while other methods, such as Ncut and graph cuts are computationally effective but only work for specific energy forms and are not generally applicable. In this paper, we present a new inference algorithm that generalizes the Swendsen-Wang method to arbitrary probabilities defined on graph partitions. We begin by computing graph edge weights, based on local image features. Then, the algorithm iterates two steps. 1) Graph clustering: It forms connected components by cutting the edges probabilistically based on their weights. 2) Graph relabeling: It selects one connected component and flips probabilistically, the coloring of all vertices in the component simultaneously. Thus, it realizes the split, merge, and regrouping of a "chunk" of the graph, in contrast to Gibbs sampler that flips a single vertex. We prove that this algorithm simulates ergodic and reversible Markov chain jumps in the space of graph partitions and is applicable to arbitrary posterior probabilities or energy functions defined on graphs. We demonstrate the algorithm on two typical problems in computer vision--image segmentation and stereo vision. Experimentally, we show that it is 100-400 times faster in CPU time than the classical Gibbs sampler and 20-40 times faster then the DDMCMC segmentation algorithm. For stereo, we compare performance with graph cuts and belief propagation. We also show that our algorithm can automatically infer generative models and obtain satisfactory results (better than the graphic cuts or belief propagation) in the same amount of time.

Algorithms↗

A 3-D reconstruction system for the human jaw using a sequence of optical images.

This paper presents a model-based vision system for dentistry that will assist in diagnosis, treatment planning, and surgical simulation. Dentistry requires an accurate three-dimensional (3-D) representation of the teeth and jaws for diagnostic and treatment purposes. The proposed integrated computer vision system constructs a 3-D model of the patient's dental occlusion using an intraoral video camera. A modified shape from shading (SFS) technique, using perspective projection and camera calibration, extracts the 3-D information from a sequence of two-dimensional (2-D) images of the jaw. Data fusion of range data and 3-D registration techniques develop the complete jaw model. Triangulation is then performed, and a solid 3-D model is reconstructed. The system performance is investigated using ground truth data, and the results show acceptable reconstruction accuracy.

Algorithms↗

Optical computer recognition of facial expressions associated with stress induced by performance demands.

Application of computer vision to track changes in human facial expressions during long-duration spaceflight may be a useful way to unobtrusively detect the presence of stress during critical operations. To develop such an approach, we applied optical computer recognition (OCR) algorithms for detecting facial changes during performance while people experienced both low- and high-stressor performance demands. Workload and social feedback were used to vary performance stress in 60 healthy adults (29 men, 31 women; mean age 30 yr). High-stressor scenarios involved more difficult performance tasks, negative social feedback, and greater time pressure relative to low workload scenarios. Stress reactions were tracked using self-report ratings, salivary cortisol, and heart rate. Subjects also completed personality, mood, and alexithymia questionnaires. To bootstrap development of the OCR algorithm, we had a human observer, blind to stressor condition, identify the expressive elements of the face of people undergoing high- vs. low-stressor performance. Different sets of videos of subjects' faces during performance conditions were used for OCR algorithm training. Subjective ratings of stress, task difficulty, effort required, frustration, and negative mood were significantly increased during high-stressor performance bouts relative to low-stressor bouts (all p < 0.01). The OCR algorithm was refined to provide robust 3-d tracking of facial expressions during head movement. Movements of eyebrows and asymmetries in the mouth were extracted. These parameters are being used in a Hidden Markov model to identify high- and low-stressor conditions. Preliminary results suggest that an OCR algorithm using mouth and eyebrow regions has the potential to discriminate high- from low-stressor performance bouts in 75-88% of subjects. The validity of the workload paradigm to induce differential levels of stress in facial expressions was established. The paradigm also provided the basic stress-related facial expressions required to establish a prototypical OCR algorithm to detect such changes. Efforts are underway to further improve the OCR algorithm by adding facial touching and automating application of the deformable masks and OCR algorithms to video footage of the moving faces as a prelude to blind validation of the automated approach.

Adult↗

Computer-assisted reading of mammograms.

Techniques developed in computer vision and automated pattern recognition can be applied to assist radiologists in reading mammograms. With the introduction of direct digital mammography this will become a feasible approach. A radiologist in breast cancer screening can use findings of the computer as a second opinion, or as a pointer to suspicious regions. This may increase the sensitivity and specificity of screening programs, and it may avoid the need for double reading. In this paper methods which have been developed for automated detection of mammographic abnormalities are reviewed. Programs for detecting microcalcification clusters and stellate lesions have reached a level of performance which makes application in practice viable. Current programs for recognition of masses and asymmetry perform less well. Large-scale studies still have to demonstrate if radiologists in a screening situation can deal with the relatively large number of false positives which are marked by computer programs, where the number of normal cases is much higher than in observer experiments conducted thus far.

Breast Neoplasms↗

Generalized system for plant growth analysis using infrared LED.

A computer vision system was developed to analyze plant growth. The developed system has a great advantage that various shapes of plant can be applied without any modification of software. The system consisted mainly of CCD camera, image capture board, computer, infrared LED and mirror system, and was easily set up. The infrared LED was hooked onto the plant part which was the location to be measured. Images, containing both the plant and the LED, with the different angle were simultaneously obtained using the mirror system, digitized and stored into a magneto optical diskette at a fixed interval. The centroid value of the LED was computed from the stored images and determined as the LED location. Experiments were conducted to evaluate the newly developed system by analyzing the elongation rate of Verbena bonariensis L. The results were compared to that obtained by the former non-contact analyzing system and it was found that the new system was applicable to plant growth analysis. This approach can not be substituted for the non-contact analyzing system formerly developed because it is contact type. However, this system can be applied to various plants without any modification in software, and has a potential for a wide use.

Algorithms↗

Orbital hemorrhage with loss of vision.

A 58-year-old woman had the sudden onset of unilateral painful proptosis, ophthalmoplegia, vomiting, and loss of vision. Computed axial tomography showed a mass that was greatly attenuated in the orbit. The initial reading of the internal carotid angiogram was normal, but a subtraction study showed a hypervascular lesion within the orbit with features indicating a hemangioma. Orbital decompression failed to restore the vision as intraorbital hemorrhage had irreparably damaged the optic nerve.

Carotid Arteries↗

A comparison of algorithms for inference and learning in probabilistic graphical models.

Research into methods for reasoning under uncertainty is currently one of the most exciting areas of artificial intelligence, largely because it has recently become possible to record, store, and process large amounts of data. While impressive achievements have been made in pattern classification problems such as handwritten character recognition, face detection, speaker identification, and prediction of gene function, it is even more exciting that researchers are on the verge of introducing systems that can perform large-scale combinatorial analyses of data, decomposing the data into interacting components. For example, computational methods for automatic scene analysis are now emerging in the computer vision community. These methods decompose an input image into its constituent objects, lighting conditions, motion patterns, etc. Two of the main challenges are finding effective representations and models in specific applications and finding efficient algorithms for inference and learning in these models. In this paper, we advocate the use of graph-based probability models and their associated inference and learning algorithms. We review exact techniques and various approximate, computationally efficient techniques, including iterated conditional modes, the expectation maximization (EM) algorithm, Gibbs sampling, the mean field method, variational techniques, structured variational techniques and the sum-product algorithm ("loopy" belief propagation). We describe how each technique can be applied in a vision model of multiple, occluding objects and contrast the behaviors and performances of the techniques using a unifying cost function, free energy.

Algorithms↗

3D laser scanning for image guided stereotactic neurosurgery.

While this work is in its very early stages, the 3D laser scanner shows significant promise as a surgical localization device with advantages over other sensing methods. Accurate 3D surface extraction and matching, a central problem in computer vision, is the key to frameless stereotaxic neurosurgery using this technique.

Brain↗

On the sample complexity of learning for networks of spiking neurons with nonlinear synaptic interactions.

We study networks of spiking neurons that use the timing of pulses to encode information. Nonlinear interactions model the spatial groupings of synapses on the neural dendrites and describe the computations performed at local branches. Within a theoretical framework of learning we analyze the question of how many training examples these networks must receive to be able to generalize well. Bounds for this sample complexity of learning can be obtained in terms of a combinatorial parameter known as the pseudodimension. This dimension characterizes the computational richness of a neural network and is given in terms of the number of network parameters. Two types of feedforward architectures are considered: constant-depth networks and networks of unconstrained depth. We derive asymptotically tight bounds for each of these network types. Constant depth networks are shown to have an almost linear pseudodimension, whereas the pseudodimension of general networks is quadratic. Networks of spiking neurons that use temporal coding are becoming increasingly more important in practical tasks such as computer vision, speech recognition, and motor control. The question of how well these networks generalize from a given set of training examples is a central issue for their successful application as adaptive systems. The results show that, although coding and computation in these networks is quite different and in many cases more powerful, their generalization capabilities are at least as good as those of traditional neural network models.

Action Potentials↗

Clustered blockwise PCA for representing visual data.

Principal Component Analysis (PCA) is extensively used in computer vision and image processing. Since it provides the optimal linear subspace in a least-square sense, it has been used for dimensionality reduction and subspace analysis in various domains. However, its scalability is very limited because of its inherent computational complexity. We introduce a new framework for applying PCA to visual data which takes advantage of the spatio-temporal correlation and localized frequency variations that are typically found in such data. Instead of applying PCA to the whole volume of data (complete set of images), we partition the volume into a set of blocks and apply PCA to each block. Then, we group the subspaces corresponding to the blocks and merge them together. As a result, we not only achieve greater efficiency in the resulting representation of the visual data, but also successfully scale PCA to handle large data sets. We present a thorough analysis of the computational complexity and storage benefits of our approach. We apply our algorithm to several types of videos. We show that, in addition to its storage and speed benefits, the algorithm results in a useful representation of the visual data.

Algorithms↗

Bilateral symmetry analysis of breast MRI.

Mammographic interpretation often uses symmetry between left and right breasts to indicate the site of potential tumour masses. This approach has not been applied to breast images obtained from MRI. We present an automatic technique for breast symmetry detection based on feature extraction techniques which does not require any efforts to co-register breast MRI data. The approach applies computer-vision techniques to detect natural biological symmetries in breast MR scans based on three objective measures of similarity: multiresolution non-orthogonal wavelet representation, three-dimensional intensity distributions and co-occurrence matrices. Statistical distributions that are invariant to feature localization are computed for each of the extracted image features. These distributions are later compared against each other to account for perceptual similarity. Studies based on 51 normal MRI scans of randomly selected patients showed that the sensitivity of symmetry detection rate approached 94%. The symmetry analysis procedure presented in this paper can be applied as an aid in detecting breast tissue changes arising from disease.

Breast↗

Edge and corner detection by photometric quasi-invariants.

Feature detection is used in many computer vision applications such as image segmentation, object recognition, and image retrieval. For these applications, robustness with respect to shadows, shading, and specularities is desired. Features based on derivatives of photometric invariants, which we will call full invariants, provide the desired robustness. However, because computation of photometric invariants involves nonlinear transformations, these features are Instable and, therefore, impractical for many applications. We propose a new class of derivatives which we refer to as quasi-invariants. These quasi-invariants are derivatives which share with full photometric invariants the property that they are insensitive for certain photometric edges, such as shadows or specular edges, but without the inherent instabilities of full photometric invariants. Experiments show that the quasi-invariant derivatives are less sensitive to noise and introduce less edge displacement than full invariant derivatives. Moreover, quasi-invariants significantly outperform the full invariant derivatives in terms of discriminative power.

Algorithms↗

Deficits in stereoscopic depth perception by mildly mentally retarded adults.

The ability of mildly mentally retarded adults to perceive specific perceptual phenomena attendant to global stereopsis produced by random element stereograms was investigated. From the standpoint of computational vision, these phenomena are difficult to process, yet nonretarded persons perceive them effortlessly and without error. Retarded subjects in this study, however, exhibited large qualitative deficits not attributable to an absence of stereopsis or a failure to comprehend. These results suggest that the computational requirements of the stimuli exceeded resources and imply the presence of a substantial structural deficit in an automatic preattentive perceptual stage quite distant from the domain of cognition.

Adult↗

An advanced system for the simulation and planning of orthodontic treatment.

This paper presents a new system for three-dimensional (3-D) orthodontic treatment planning and movement of teeth. We describe a computer vision technique for the acquisition and processing of 3-D images of the profile of hydrocolloid dental imprints. Profile measurement is based on the triangulation method which detects deformation of the projection of a laser line on the dental imprints. The system is computer-controlled and designed to achieve depth and lateral resolutions of 0.1 and 0.2 mm, respectively, within a depth range of 40 mm. The 3-D image of the imprint is segmented in order to identify different teeth. Two operators are presented: one for the detection of molars and premolars based on a directional gradient, and one for incisors and canines based on 3-D registration with dental models contained in a database. We apply these 3-D dental models to simulate the 3-D movement of teeth, including rotations, during orthodontic treatment. With this objective, we have developed an original simplified model of arch-wire behaviour and a viscoplastic behaviour law for the alveolar bone in order to simulate teeth displacements during orthodontic treatment. The contribution of the paper is part of a diagnosis system (called MAGALLANES) that is designed to replace manual measurement methods, which use costly plaster models, with computer measurement methods and teeth movement simulation using cheap hydrocolloid dental wafers. This procedure will reduce the cost and acquisition time of orthodontic data and facilitate the conduct of epidemiological studies.

Biomechanical Phenomena↗

Generalized principal component analysis (GPCA).

This paper presents an algebro-geometric solution to the problem of segmenting an unknown number of subspaces of unknown and varying dimensions from sample data points. We represent the subspaces with a set of homogeneous polynomials whose degree is the number of subspaces and whose derivatives at a data point give normal vectors to the subspace passing through the point. When the number of subspaces is known, we show that these polynomials can be estimated linearly from data; hence, subspace segmentation is reduced to classifying one point per subspace. We select these points optimally from the data set by minimizing certain distance function, thus dealing automatically with moderate noise in the data. A basis for the complement of each subspace is then recovered by applying standard PCA to the collection of derivatives (normal vectors). Extensions of GPCA that deal with data in a high-dimensional space and with an unknown number of subspaces are also presented. Our experiments on low-dimensional data show that GPCA outperforms existing algebraic algorithms based on polynomial factorization and provides a good initialization to iterative techniques such as K-subspaces and Expectation Maximization. We also present applications of GPCA to computer vision problems such as face clustering, temporal video segmentation, and 3D motion segmentation from point correspondences in multiple affine views.

Algorithms↗

The humanID gait challenge problem: data sets, performance, and analysis.

Identification of people by analysis of gait patterns extracted from video has recently become a popular research problem. However, the conditions under which the problem is "solvable" are not understood or characterized. To provide a means for measuring progress and characterizing the properties of gait recognition, we introduce the HumanID Gait Challenge Problem. The challenge problem consists of a baseline algorithm, a set of 12 experiments, and a large data set. The baseline algorithm estimates silhouettes by background subtraction and performs recognition by temporal correlation of silhouettes. The 12 experiments are of increasing difficulty, as measured by the baseline algorithm, and examine the effects of five covariates on performance. The covariates are: change in viewing angle, change in shoe type, change in walking surface, carrying or not carrying a briefcase, and elapsed time between sequences being compared. Identification rates for the 12 experiments range from 78 percent on the easiest experiment to 3 percent on the hardest. All five covariates had statistically significant effects on performance, with walking surface and time difference having the greatest impact. The data set consists of 1,870 sequences from 122 subjects spanning five covariates (1.2 Gigabytes of data). The gait data, the source code of the baseline algorithm, and scripts to run, score, and analyze the challenge experiments are available at http://www.GaitChallenge.org. This infrastructure supports further development of gait recognition algorithms and additional experiments to understand the strengths and weaknesses of new algorithms. The more detailed the experimental results presented, the more detailed is the possible meta-analysis and greater is the understanding. It is this potential from the adoption of this challenge problem that represents a radical departure from traditional computer vision research methodology.

Adult↗

Supervised range-constrained thresholding.

A novel thresholding approach to confine the intensity frequency range of the object based on supervision is introduced. It consists of three steps. First, the region of interest (ROI) is determined in the image. Then, from the histogram of the ROI, the frequency range in which the proportion of the background to the ROI varies is estimated through supervision. Finally, the threshold is determined by minimizing the classification error within the constrained variable background range. The performance of the approach has been validated against 54 brain MR images, including images with severe intensity inhomogeneity and/or noise, CT chest images, and the Cameraman image. Compared with nonsupervised thresholding methods, the proposed approach is substantially more robust and more reliable. It is also computationally efficient and can be applied to a wide class of computer vision problems, such as character recognition, fingerprint identification, and segmentation of a wide variety of medical images.

Algorithms↗