PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “computer vision”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Computer-aided mammographic screening for spiculated lesions.

PURPOSE: To study the use of a computer vision method as a second reader for the detection of spiculated lesions on screening mammograms. MATERIALS AND METHODS: An algorithmic computer process for the detection of spiculated lesions on digitized screen-film mammograms was applied to 85 four-view clinical cases: 36 cases with cancer proved by means of biopsy and 49 cases with negative findings at examination and follow-up. The computer detections were printed as film with added outlines that indicated the suspected cancers. Four radiologists screened the 85 cases twice, once without and once with the computer reports as ancillary films. RESULTS: The algorithm alone achieved 100% sensitivity, with a specificity of 82%. The computer reports increased the average radiologist sensitivity by 9.7% (P = .005), moving from 80.6% to 90.3%, with no decrease in average specificity. CONCLUSION: The study demonstrated that computer analysis of mammograms can provide a substantial and statistically significant increase in radiologist screening efficacy.

Algorithms↗

Multistage integration model for human egomotion perception.

Human computational vision models that attempt to account for the dynamic perception of egomotion and relative depth typically assume a common three-stage process: first, compute the optical flow field based on the dynamically changing image; second, estimate the egomotion states based on the flow; and third, estimate the relative depth/shape based on the egomotion states and possibly on a model of the viewed surface. We propose a model more in line with recent work in human vision, employing multistage integration. Here the dynamic image is first processed to generate spatial and temporal image gradients that drive a mutually interconnected state estimator and depth/shape estimator. The state estimator uses the image gradient information in combination with a depth/shape estimate of the viewed surface and an assumed model of the viewer's dynamics to generate current state estimates; in tandem, the depth/shape estimator uses the image gradient information in combination with the viewer's state estimate and assumed shape model to generate current depth/shape estimates. In this paper, we describe the model and compare model predictions with empirical data.

Aircraft↗

Architecture and design of a computerised stereogram generator for vision test.

A new microcomputer-based stereogram generator was designed and implemented to generate various visual stimuli that are used for testing the binocular vision system. The system is capable of generating static and dynamic stereoscopic stereograms that can be varied in size, shape, speed and disparity. It can also be used to generate a luminous stimulus on a dark background which, except for the depth parameters, can be varied in a similar way to the stereoscopic stimulus. A 16/32-bit microprocessor has been employed for the overall control of the stereogram parameters, which provides flexibility, versatility, compactness and speed at reduced cost. We have applied this system to the measurement of eye movement and computer vision.

Diagnosis, Computer-Assisted↗

Aspect graphs for visual recognition of three-dimensional objects.

Visual representation of three-dimensional (3-D) objects in our environment is a crucial question, for human as well as for machine vision. Some basics are reviewed of a viewer-centred model of 3-D objects, aspect graphs, which represents a 3-D object by all its topologically stable visible image contours (its aspects) and by the transitions between stable image contours (the visual events). This representation takes only geometrical information about discontinuities in depth and in surface orientation into account, and other clues, such as shadows, markings, texture, etc, are disregarded. Mathematical results give some insight into the relationships between the geometry of a 3-D object and the aspect of its image contours, the techniques used to compute an aspect graph effectively, and the state of the art of this type of model in computer vision. Current research is reviewed on viewer-centred representation in cognitive science that seems to indicate that aspect graphs could also have some relevance for human vision.

Attention↗

Statistical characterization of real-world illumination.

Although studies of vision and graphics often assume simple illumination models, real-world illumination is highly complex, with reflected light incident on a surface from almost every direction. One can capture the illumination from every direction at one point photographically using a spherical illumination map. This work illustrates, through analysis of photographically acquired, high dynamic range illumination maps, that real-world illumination possesses a high degree of statistical regularity. The marginal and joint wavelet coefficient distributions and harmonic spectra of illumination maps resemble those documented in the natural image statistics literature. However, illumination maps differ from typical photographs in that illumination maps are statistically nonstationary and may contain localized light sources that dominate their power spectra. Our work provides a foundation for statistical models of real-world illumination, thereby facilitating the understanding of human material perception, the design of robust computer vision systems, and the rendering of realistic computer graphics imagery.

Contrast Sensitivity↗

A model of visual recognition and categorization.

To recognize a previously seen object, the visual system must overcome the variability in the object's appearance caused by factors such as illumination and pose. Developments in computer vision suggest that it may be possible to counter the influence of these factors, by learning to interpolate between stored views of the target object, taken under representative combinations of viewing conditions. Daily life situations, however, typically require categorization, rather than recognition, of objects. Due to the open-ended character of both natural and artificial categories, categorization cannot rely on interpolation between stored examples. Nonetheless, knowledge of several representative members, or prototypes, of each of the categories of interest can still provide the necessary computational substrate for the categorization of new instances. The resulting representational scheme based on similarities to prototypes appears to be computationally viable, and is readily mapped onto the mechanisms of biological vision revealed by recent psychophysical and physiological studies.

Animals↗

Modelling the world in real time: how robots engineer information.

Programming robots and other autonomous systems to interact with the world in real time is bringing into sharp focus general questions about representation, inference and understanding. These artificial agents use digital computation to interpret the data gleaned from sensors and produce decisions and actions to guide their future behaviour. In a physical system, however, finite computational resources unavoidably impose the need to approximate and make selective use of the information available to reach prompt deductions. Recent research has led to widespread adoption of the methodology of Bayesian inference, which provides the absolute framework to understand this process fully via modelling as informed, fully acknowledged approximation. The performance of modern systems has improved greatly on the heuristic methods of the early days of artificial intelligence. We discuss the general problem of real-time inference and computation, and draw on examples from recent research in computer vision and robotics: specifically visual tracking and simultaneous localization and mapping.

Artificial Intelligence↗

A Modified ART 1 Algorithm more Suitable for VLSI Implementations.

This paper presents a modification to the original ART 1 algorithm ([Carpenter and Grossberg, 1987a], A massively parallel architecture for a self-organizing neural pattern recognition machine, Computer Vision, Graphics, and Image Processing, 37, 54-115) that is conceptually similar, can be implemented in hardware with less sophisticated building blocks, and maintains the computational capabilities of the originally proposed algorithm. This modified ART 1 algorithm (which we will call here ART 1(m)) is the result of hardware motivated simplifications investigated during the design of an actual ART 1 chip [Serrano-Gotarredona et al., 1994, Proc. 1994 IEEE Int. Conf. Neural Networks (Vol. 3, pp. 1912-1916); [Serrano-Gotarredona and Linares-Barranco, 1996], IEEE Trans. VLSI Systems, (in press)]. The purpose of this paper is simply to justify theoretically that the modified algorithm preserves the computational properties of the original one and to study the difference in behavior between the two approaches. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

An efficient motion estimator with application to medical image registration.

Image registration is a very important problem in computer vision and medical image processing. Numerous algorithms for registering single and multi-modal image data have been reported in these areas. Robustness as well as computational efficiency are prime factors of importance in image data registration. In this paper, a robust/reliable and efficient algorithm for estimating the transformation between two image data sets of a patient taken from the same modality over time is presented. Estimating the registration between two image data sets is formulated as a motion-estimation problem. We use a hierarchical optical flow motion model which allows for both global as well as local motion between the data sets. In this hierarchical motion model, we represent the flow field with a B-spline basis which implicitly incorporates smoothness constraints on the field. In computing the motion, we minimize the expectation of the squared differences energy function numerically via a modified Newton iteration scheme. The main idea in the modified Newton method is that we precompute the Hessian of the energy function at the optimum without explicitly knowing the optimum. This idea is used for both global and local motion estimation in the hierarchical motion model. We present examples of motion estimation on synthetic and real data (from a patient acquired during pre- and post-operative stages) and compare the performance of our algorithm with that of competing ones.

Algorithms↗

The role of chromatin state in intron retention: A case study in leveraging large scale deep learning models.

Complex deep learning models trained on very large datasets have become key enabling tools for current research in natural language processing and computer vision. By providing pre-trained models that can be fine-tuned for specific applications, they enable researchers to create accurate models with minimal effort and computational resources. Large scale genomics deep learning models come in two flavors: the first are large language models of DNA sequences trained in a self-supervised fashion, similar to the corresponding natural language models; the second are supervised learning models that leverage large scale genomics datasets from ENCODE and other sources. We argue that these models are the equivalent of foundation models in natural language processing in their utility, as they encode within them chromatin state in its different aspects, providing useful representations that allow quick deployment of accurate models of gene regulation. We demonstrate this premise by leveraging the recently created Sei model to develop simple, interpretable models of intron retention, and demonstrate their advantage over models based on the DNA language model DNABERT-2. Our work also demonstrates the impact of chromatin state on the regulation of intron retention. Using representations learned by Sei, our model is able to discover the involvement of transcription factors and chromatin marks in regulating intron retention, providing better accuracy than a recently published custom model developed for this purpose.

Deep Learning↗

Automatic 3-D grayscale volume matching and shape analysis.

Recently, shape matching in three dimensions (3-D) has been gaining importance in a wide variety of fields such as computer graphics, computer vision, medicine, and biology, with applications such as object recognition, medical diagnosis, and quantitative morphological analysis of biological operations. Automatic shape matching techniques developed in the field of computer graphics handle object surfaces, but ignore intensities of inner voxels. In biology and medical imaging, voxel intensities obtained by computed tomography (CT), magnetic resonance imagery (MRI), and confocal microscopes are important to determine point correspondences. Nevertheless, most biomedical volume matching techniques require human interactions, and automatic methods assume matched objects to have very similar shapes so as to avoid combinatorial explosions of point. This article is aimed at decreasing the gap between the two fields. The proposed method automatically finds dense point correspondences between two grayscale volumes; i.e., finds a correspondent in the second volume for every voxel in the first volume, based on the voxel intensities. Mutiresolutional pyramids are introduced to reduce computational load and handle highly plastic objects. We calculate the average shape of a set of similar objects and give a measure of plasticity to compare them. Matching results can also be used to generate intermediate volumes for morphing. We use various data to validate the effectiveness of our method: we calculate the average shape and plasticity of a set of fly brain cells, and we also match a human skull and an orangutan skull.

Algorithms↗

Camera-based calibration techniques for seamless multiprojector displays.

Multiprojector, large-scale displays are used in scientific visualization, virtual reality, and other visually intensive applications. In recent years, a number of camera-based computer vision techniques have been proposed to register the geometry and color of tiled projection-based display. These automated techniques use cameras to "calibrate" display geometry and photometry, computing per-projector corrective warps and intensity corrections that are necessary to produce seamless imagery across projector mosaics. These techniques replace the traditional labor-intensive manual alignment and maintenance steps, making such displays cost-effective, flexible, and accessible. In this paper, we present a survey of different camera-based geometric and photometric registration techniques reported in the literature to date. We discuss several techniques that have been proposed and demonstrated, each addressing particular display configurations and modes of operation. We overview each of these approaches and discuss their advantages and disadvantages. We examine techniques that address registration on both planar (video walls) and arbitrary display surfaces and photometric correction for different kinds of display surfaces. We conclude with a discussion of the remaining challenges and research opportunities for multiprojector displays.

Algorithms↗

A computational framework for cortical learning.

Recent physiological findings have revealed that long-term adaptation of the synaptic strengths between cortical pyramidal neurons depends on the temporal order of presynaptic and postsynaptic spikes, which is called spike-timing-dependent plasticity (STDP) or temporally asymmetric Hebbian (TAH) learning. Here I prove by analytical means that a physiologically plausible variant of STDP adapts synaptic strengths such that the presynaptic spikes predict the postsynaptic spikes with minimal error. This prediction error model of STDP implies a mechanism for cortical memory: cortical tissue learns temporal spike patterns if these spike patterns are repeatedly elicited in a set of pyramidal neurons. The trained network finishes these patterns if their beginnings are presented, thereby recalling the memory. Implementations of the proposed algorithms may be useful for applications in voice recognition and computer vision.

Action Potentials↗

Variational denoising of partly textured images by spatially varying constraints.

Denoising algorithms based on gradient dependent regularizers, such as nonlinear diffusion processes and total variation denoising, modify images towards piecewise constant functions. Although edge sharpness and location is well preserved, important information, encoded in image features like textures or certain details, is often compromised in the process of denoising. We propose a mechanism that better preserves fine scale features in such denoising processes. A basic pyramidal structure-texture decomposition of images is presented and analyzed. A first level of this pyramid is used to isolate the noise and the relevant texture components in order to compute spatially varying constraints based on local variance measures. A variational formulation with a spatially varying fidelity term controls the extent of denoising over image regions. Our results show visual improvement as well as an increase in the signal-to-noise ratio over scalar fidelity term processes. This type of processing can be used for a variety of tasks in partial differential equation-based image processing and computer vision, and is stable and meaningful from a mathematical viewpoint.

Algorithms↗

The use of morphological characteristics and texture analysis in the identification of tissue composition in prostatic neoplasia.

Quantitative examination of prostate histology offers clues in the diagnostic classification of lesions and in the prediction of response to treatment and prognosis. To facilitate the collection of quantitative data, the development of machine vision systems is necessary. This study explored the use of imaging for identifying tissue abnormalities in prostate histology. Medium-power histological scenes were recorded from whole-mount radical prostatectomy sections at x 40 objective magnification and assessed by a pathologist as exhibiting stroma, normal tissue (nonneoplastic epithelial component), or prostatic carcinoma (PCa). A machine vision system was developed that divided the scenes into subregions of 100 x 100 pixels and subjected each to image-processing techniques. Analysis of morphological characteristics allowed the identification of normal tissue. Analysis of image texture demonstrated that Haralick feature 4 was the most suitable for discriminating stroma from PCa. Using these morphological and texture measurements, it was possible to define a classification scheme for each subregion. The machine vision system is designed to integrate these classification rules and generate digital maps of tissue composition from the classification of subregions; 79.3% of subregions were correctly classified. Established classification rates have demonstrated the validity of the methodology on small scenes; a logical extension was to apply the methodology to whole slide images via scanning technology. The machine vision system is capable of classifying these images. The machine vision system developed in this project facilitates the exploration of morphological and texture characteristics in quantifying tissue composition. It also illustrates the potential of quantitative methods to provide highly discriminatory information in the automated identification of prostatic lesions using computer vision.

Humans↗

An experimental comparison of min-cut/max-flow algorithms for energy minimization in vision.

After [15], [31], [19], [8], [25], [5], minimum cut/maximum flow algorithms on graphs emerged as an increasingly useful tool for exact or approximate energy minimization in low-level vision. The combinatorial optimization literature provides many min-cut/max-flow algorithms with different polynomial time complexity. Their practical efficiency, however, has to date been studied mainly outside the scope of computer vision. The goal of this paper is to provide an experimental comparison of the efficiency of min-cut/max flow algorithms for applications in vision. We compare the running times of several standard algorithms, as well as a new algorithm that we have recently developed. The algorithms we study include both Goldberg-Tarjan style "push-relabel" methods and algorithms based on Ford-Fulkerson style "augmenting paths." We benchmark these algorithms on a number of typical graphs in the contexts of image restoration, stereo, and segmentation. In many cases, our new algorithm works several times faster than any of the other methods, making near real-time performance possible. An implementation of our max-flow/min-cut algorithm is available upon request for research purposes.

Algorithms↗

Robust fusion of uncertain information.

A technique is presented to combine n data points, each available with point-dependent uncertainty, when only a subset of these points come from N < n sources, where N is unknown. We detect the significant modes of the underlying multivariate probability distribution using a generalization of the nonparametric mean shift procedure. The number of detected modes automatically defines N, while the belonging of a point to the basin of attraction of a mode provides the fusion rule. The robust data fusion algorithm was successfully applied to two computer vision problems: estimating the multiple affine transformations, and range image segmentation.

Algorithms↗