PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “computer vision”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

A non-contact mouse for surgeon-computer interaction.

We have developed a system that uses computer vision to replace standard computer mouse functions with hand gestures. The system is designed to enable non-contact human-computer interaction (HCI), so that surgeons will be able to make more effective use of computers during surgery. In this paper, we begin by discussing the need for non-contact computer interfaces in the operating room. We then describe the design of our non-contact mouse system, focusing on the techniques used for hand detection, tracking, and gesture recognition. Finally, we present preliminary results from testing and planned future work.

Computer Peripherals↗

Is vision continuous with cognition? The case for cognitive impenetrability of visual perception.

Although the study of visual perception has made more progress in the past 40 years than any other area of cognitive science, there remain major disagreements as to how closely vision is tied to cognition. This target article sets out some of the arguments for both sides (arguments from computer vision, neuroscience, psychophysics, perceptual learning, and other areas of vision science) and defends the position that an important part of visual perception, corresponding to what some people have called early vision, is prohibited from accessing relevant expectations, knowledge, and utilities in determining the function it computes--in other words, it is cognitively impenetrable. That part of vision is complex and involves top-down interactions that are internal to the early vision system. Its function is to provide a structured representation of the 3-D surfaces of objects sufficient to serve as an index into memory, with somewhat different outputs being made available to other systems such as those dealing with motor control. The paper also addresses certain conceptual and methodological issues raised by this claim, such as whether signal detection theory and event-related potentials can be used to assess cognitive penetration of vision. A distinction is made among several stages in visual processing, including, in addition to the inflexible early-vision stage, a pre-perceptual attention-allocation stage and a post-perceptual evaluation, selection, and inference stage, which accesses long-term memory. These two stages provide the primary ways in which cognition can affect the outcome of visual perception. The paper discusses arguments from computer vision and psychology showing that vision is "intelligent" and involves elements of "problem solving." The cases of apparently intelligent interpretation sometimes cited in support of this claim do not show cognitive penetration; rather, they show that certain natural constraints on interpretation, concerned primarily with optical and geometrical properties of the world, have been compiled into the visual system. The paper also examines a number of examples where instructions and "hints" are alleged to affect what is seen. In each case it is concluded that the evidence is more readily assimilated to the view that when cognitive effects are found, they have a locus outside early vision, in such processes as the allocation of focal attention and the identification of the stimulus.

Agnosia↗

Visual space distortion.

We are surrounded by surfaces that we perceive by visual means. Understanding the basic principles behind this perceptual process is a central theme in visual psychology, psychophysics, and computational vision. In many of the computational models employed in the past, it has been assumed that a metric representation of physical space can be derived by visual means. Psychophysical experiments, as well as computational considerations, can convince us that the perception of space and shape has a much more complicated nature, and that only a distored version of actual, physical space can be computed. This paper develops a computational geometric model that explains why such distortion might take place. The basic idea is that, both in stereo and motion, we perceive the world from multiple views. Given the rigid transformation between the views and the properties of the image correspondence, the depth of the scene can be obtained. Even a slight error in the rigid transformation parameters causes distortion of the computed depth of the scene. The unified framework introduced here describes this distortion in computational terms. We characterize the space of distortions by its level sets, that is, we characterize the systematic distortion via a family of iso-distortion surfaces which describes the locus over which depths are distorted by some multiplicative factor. Given that humans' estimation of egomotion or estimation of the extrinsic parameters of the stereo apparatus is likely to be imprecise, the framework is used to explain a number of psychophysical experiments on the perception of depth from motion or stereo.

Cybernetics↗

Informatics in radiology (infoRAD): NeatVision: visual programming for computer-aided diagnostic applications.

A free visual programming-based image analysis development environment for medical imaging applications called NeatVision was developed to provide high-level access to a wide range of image processing algorithms through a well-defined, easy-to-use graphical interface. The system contains over 300 image manipulation, processing, and analysis algorithms. For more advanced users, an upgrade path is provided to extend the core library with use of the developer's interface, giving users access to additional plug-in features, automatic source code generation, compilation with full error feedback, and dynamic algorithm updates. NeatVision was designed to allow users at all levels of expertise to focus on the computer vision design task for computer-aided diagnostic (CAD) applications rather than the subtleties of a particular programming language. The environment allows the designers of image analysis-based CAD techniques to implement their ideas in a dynamic and straightforward manner. Both NeatVision standard and developer's versions can be downloaded free of charge from the Internet and can run on a variety of computer platforms.

Diagnosis, Computer-Assisted↗

Image recognition: visual grouping, recognition, and learning.

Vision extracts useful information from images. Reconstructing the three-dimensional structure of our environment and recognizing the objects that populate it are among the most important functions of our visual system. Computer vision researchers study the computational principles of vision and aim at designing algorithms that reproduce these functions. Vision is difficult: the same scene may give rise to very different images depending on illumination and viewpoint. Typically, an astronomical number of hypotheses exist that in principle have to be analyzed to infer a correct scene description. Moreover, image information might be extracted at different levels of spatial and logical resolution dependent on the image processing task. Knowledge of the world allows the visual system to limit the amount of ambiguity and to greatly simplify visual computations. We discuss how simple properties of the world are captured by the Gestalt rules of grouping, how the visual system may learn and organize models of objects for recognition, and how one may control the complexity of the description that the visual system computes.

Humans↗

The forms of knowledge mobilized in some machine vision systems.

This paper describes a number of computer vision systems that we have constructed, and which are firmly based on knowledge of diverse sorts. However, that knowledge is often represented in a way that is only accessible to a limited set of processes, that make limited use of it, and though the knowledge is amenable to change, in practice it can only be changed in rather simple ways. The rest of the paper addresses the questions: (i) what knowledge is mobilized in the furtherance of a perceptual task?; (ii) how is that knowledge represented?; and (iii) how is that knowledge mobilized? First we review some cases of early visual processing where the mobilization of knowledge seems to be a key contributor to success yet where the knowledge is deliberately represented in a quite inflexible way. After considering the knowledge that is involved in overcoming the projective nature of images, we move the discussion to the knowledge that was required in programs to match, register, and recognize shapes in a range of applications. Finally, we discuss the current state of process architectures for knowledge mobilization.

Humans↗

Using the low-resolution properties of correlated images to improve the computational efficiency of eigenspace decomposition.

Eigendecomposition is a common technique that is performed on sets of correlated images in a number of computer vision and robotics applications. Unfortunately, the computation of an eigendecomposition can become prohibitively expensive when dealing with very high-resolution images. While reducing the resolution of the images will reduce the computational expense, it is not known a priori how this will affect the quality of the resulting eigendecomposition. The work presented here provides an analysis of how different resolution reduction techniques affect the eigendecomposition. A computationally efficient algorithm for calculating the eigendecomposition based on this analysis is proposed. Examples show that this algorithm performs well on arbitrary video sequences.

Algorithms↗

Neural gradient models for the measurement of image velocity.

Although gradient schemes for detecting the motion of images and measuring their velocities are commonly used in computer vision, and although there is increasing evidence to support the existence of such schemes in biological vision, little attention has been directed to suggesting how such computations might be realized by neural hardware. This paper proposes two simple models, consisting of physiologically realistic networks of neurons, that approximate the gradient scheme. Computer simulations demonstrate that the models measure the speed of an object or pattern independently of its structural properties.

Computer Simulation↗

Parallel integration of vision modules.

Computer algorithms have been developed for several early vision processes, such as edge detection, stereopsis, motion, texture, and color, that give separate cues to the distance from the viewer of three-dimensional surfaces, their shape, and their material properties. Not surprisingly, biological vision systems still greatly outperform computer vision programs. One of the keys to the reliability, flexibility, and robustness of biological vision systems is their ability to integrate several visual cues. A computational technique for integrating different visual cues has now been developed and implemented with encouraging results on a parallel supercomputer.

Algorithms↗

Recursive implementations of temporal filters for image motion computation.

Efficient algorithms for image motion computation are important for computer vision applications and the modelling of biological vision systems. Intensity-based image motion computation proceeds in two stages: the convolution of linear spatiotemporal filter kernels with the image sequence, followed by the non-linear combination of the filter outputs. If the spatiotemporal extent of the filter kernels is large, then the convolution stage can be very intensive computationally. One effective means of reducing the storage required and computation involved in implementing the temporal convolutions is the introduction of recursive filtering. Non-recursive methods require the number of frames of the image sequence stored at any given time to be equal to the temporal extent of the slowest temporal filter. In contrast, recursive methods encode recent stimulus history implicitly in the values of a small number of variables updated through a series of feedback equations. Recursive filtering reduces the number of values stored in memory during convolution and the number of mathematical operations involved in computing the filters' outputs. This paper extends previous recursive implementations of gradient- and correlation-based motion analysis algorithms [Fleet DJ, Langley K (1995) IEEE PAMI 17: 61-67; Clifford CWG, Ibbotson MR, Langley K (1997) Vis Neurosci 14: 741-749], describing a recursive implementation of causal band-pass temporal filters suitable for use in energy- and phase-based algorithms for image motion computation. It is shown that the filters' temporal frequency tuning curves fit psychophysical estimates of the temporal properties of human visual filters.

Algorithms↗

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform↗

AI echo INSIGHT study: A prospective blinded randomized trial of artificial intelligence echocardiogram interpretation.

BACKGROUND: Transthoracic echocardiography (TTE) is the most commonly performed cardiac imaging modality with over 30 million studies annually. Demand for timely expert interpretation continues to outpace capacity, creating diagnostic delays and inter-observer variability that impact patient care. Recent research has suggested computer vision artificial intelligence (AI) models can generate accurate preliminary comprehensive TTE reports, however, prospective evaluation is needed to determine whether AI-assisted TTE interpretation can improve clinician efficiency while preserving diagnostic accuracy. METHODS: AI ECHO INSIGHT is a prospective randomized blinded clinical trial conducted at Kaiser Permanente Northern California that will evaluate 1200 historical TTE studies (1000 consecutive unselected studies plus 200 with moderate or greater valvular disease) interpreted using three workflows: (1) AI-generated preliminary report finalized by a blinded cardiologist (AI-assisted); (2) cardiologist-generated preliminary report finalized by a blinded cardiologist (cardiologist-assisted); and (3) sonographer-generated preliminary report finalized by a blinded cardiologist (sonographer-assisted). The primary outcome is the rate of substantial change between preliminary and final reports, comparing the AI-assisted workflow to the pooled cardiologist-assisted and sonographer-assisted workflows. Secondary outcomes include cardiologist interpretation time for report finalization, superiority testing for diagnostic accuracy, and reporting consistency. CONCLUSION: AI ECHO INSIGHT is a prospective randomized blinded clinical trial evaluating the clinical impact of AI-assisted TTE interpretation on diagnostic accuracy, cardiologist efficiency, and reporting consistency in real-world echocardiography workflows. TRIAL REGISTRATION: ClinicalTrials.gov registration number NCT07229300.

Humans↗

Implementation of digital stereo imaging for analysis of metaphyses and joints in skeletal collections.

The surface structure of the growing portion of bones, called the metaphysis, contains clues about the locomotor characteristics of various species. Present methods of capturing this anthropologically interesting surface are time-consuming and subject to human error. The research implements a digital stereo imaging technique for bone metaphyses and joints in skeletal collections. The corresponding points in two images collected from different angles are determined using an area-based correlation matching method. The depths of matched points are computed from the difference in location of the points in the two images. The paper presents a practical implementation of computer vision for anthropology using an 80286-based personal computer, a camera and a video digitiser. The stereo matching algorithm, a practical implementation of classical stereo imaging methods, takes less than 1 min and produces reasonable representations of mammal bones. The accuracy of the depth measurements ranged from 0.7 to 12 per cent for 45-150 cm object-camera distances. False matches occurred in approximately 6 per cent of the total matched points.

Animals↗

Efficient molecular surface generation using level-set methods.

Molecules interact through their surface residues. Calculation of the molecular surface of a protein structure is thus an important step for a detailed functional analysis. One of the main considerations in comparing existing methods for molecular surface computations is their speed. Most of the methods that produce satisfying results for small molecules fail to do so for large complexes. In this article, we present a level-set-based approach to compute and visualize a molecular surface at a desired resolution. The emerging level-set methods have been used for computing evolving boundaries in several application areas from fluid mechanics to computer vision. Our method provides a uniform framework for computing solvent-accessible, solvent-excluded surfaces and interior cavities. The computation is carried out very efficiently even for very large molecular complexes with tens of thousands of atoms. We compared our method to some of the most widely used molecular visualization tools (Swiss-PDBViewer, PyMol, and Chimera) and our results show that we can calculate and display a molecular surface 1.5-3.14 times faster on average than all three of the compared programs. Furthermore, we demonstrate that our method is able to detect all of the interior inaccessible cavities that can accommodate one or more water molecules.

Algorithms↗

Edge detection revisited.

The present manuscript aims at solving four problems of edge detection: the simultaneous detection of all step edges from a fine to a coarse scale; the detection of thin bars with a width of very few pixels; the detection of trihedral junctions; the development of an algorithm with image-independent parameters. The proposed solution of these problems combines an extensive spatial filtering with classical methods of computer vision and newly developed algorithms. Step edges are computed by extracting local maxima from the energy summed over a large bank of directional odd filters with a different scale. Thin roof edges are computed by considering maxima of the energy summed over narrow odd and even filters along the direction providing maximal response. Junctions are precisely detected and recovered using the output of directional filters. The proposed algorithm has a threshold for the minimum contrast of detected edges: for the large number of tested images this threshold was fixed equal to three times the standard deviation of the noise present in usual acquisition system (estimated to be between 1 and 1.3 gray levels out of 256), therefore, the proposed scheme is in fact parameter free. This scheme for edge detection performs better than the classical Canny edge detector in two quantitative comparisons: the recovery of the original image from the edge map and the structure from motion task. As the Canny detector in previous comparisons was shown to be the best or among the best detectors, the proposed scheme represents a significant improvement over previous approaches.

Algorithms↗

Finding perceptually dominant orientations in natural textures.

An algorithm for detecting orientation in texture is developed and compared with results of humans detecting orientation in the same textures. The algorithm is based on the steerable filters of Freeman and Adelson (IEEE Trans. PAMI 13, 891-906, 1991), orientation-selective filters derived from derivatives of Gaussians. The filters are applied over multiple scales and their outputs non-linearly contrast-normalized. The data for humans were collected from forty subjects who were asked to identify 'the minimum number of dominant orientations' they perceived, and the 'strength' with which they perceived each orientation. Test data consisted of 111 grey-level images of natural textures taken from the Brodatz album, a standard collection used in computer vision and image processing. Results show that the computer and humans chose at least one of the same dominant orientations on 95 of the natural textures. Of these textures, 74 were also in 100% agreement on the location of all the dominant orientations chosen by both humans and computer. Disagreements are analyzed and possible causes are discussed. Some apparent limitations in the current filter shapes and sizes are illustrated, as well as some (surprisingly small) effects believed to be caused by semantic recognition and gestalt grouping.

Algorithms↗

Clifford Fourier transform on vector fields.

Image processing and computer vision have robust methods for feature extraction and the computation of derivatives of scalar fields. Furthermore, interpolation and the effects of applying a filter can be analyzed in detail and can be advantages when applying these methods to vector fields to obtain a solid theoretical basis for feature extraction. We recently introduced the Clifford convolution, which is an extension of the classical convolution on scalar fields and provides a unified notation for the convolution of scalar and vector fields. It has attractive geometric properties that allow pattern matching on vector fields. In image processing, the convolution and the Fourier transform operators are closely related by the convolution theorem and, in this paper, we extend the Fourier transform to include general elements of Clifford Algebra, called multivectors, including scalars and vectors. The resulting convolution and derivative theorems are extensions of those for convolution and the Fourier transform on scalar fields. The Clifford Fourier transform allows a frequency analysis of vector fields and the behavior of vector-valued filters. In frequency space, vectors are transformed into general multivectors of the Clifford Algebra. Many basic vector-valued patterns, such as source, sink, saddle points, and potential vortices, can be described by a few multivectors in frequency space.

Algorithms↗

Optic flow and autonomous navigation.

Many animals, especially insects, compute and use optic flow to control their motion direction and to avoid obstacles. Recent advances in computer vision have shown that an adequate optic flow can be computed from image sequences. Therefore studying whether artificial systems, such as robots, can use optic flow for similar purposes is of particular interest. Experiments are reviewed that suggest the possible use of optic flow for the navigation of a robot moving in indoor and outdoor environments. The optic flow is used to detect and localise obstacles in indoor scenes, such as corridors, offices, and laboratories. These routines are based on the computation of a reduced optic flow. The robot is usually able to avoid large obstacles such as a chair or a person. The avoidance performances of the proposed algorithm critically depend on the optomotor reaction of the robot. The optic flow can be used to understand the ego-motion in outdoor scenes, that is, to obtain information on the absolute velocity of the moving vehicle and to detect the presence of other moving objects. A critical step is the correction of the optic flow for shocks and vibrations present during image acquisition. The results obtained suggest that optic flow can be successfully used by biological and artificial systems to control their navigation. Moreover, both systems require fast and accurate optomotor reactions and need to compensate for the instability of the viewed world.

Computer Simulation↗