PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Computer vision”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A network that learns to recognize three-dimensional objects.

The visual recognition of three-dimensional (3-D) objects on the basis of their shape poses at least two difficult problems. First, there is the problem of variable illumination, which can be addressed by working with relatively stable features such as intensity edges rather than the raw intensity images. Second, there is the problem of the initially unknown pose of the object relative to the viewer. In one approach to this problem, a hypothesis is first made about the viewpoint, then the appearance of a model object from such a viewpoint is computed and compared with the actual image. Such recognition schemes generally employ 3-D models of objects, but the automatic learning of 3-D models is itself a difficult problem. To address this problem in computational vision, we have developed a scheme, based on the theory of approximation of multivariate functions, that learns from a small set of perspective views a function mapping any viewpoint to a standard view. A network equivalent to this scheme will thus 'recognize' the object on which it was trained from any viewpoint.

Artificial Intelligence↗

A possible neuronal basis for representation of acoustic scenes in auditory cortex of the big brown bat.

Behavioural studies and field observations demonstrate that echolocating bats simultaneously perceive range, direction and shape of multiple objects in the environment as acoustic images derived from echoes. Cortical echo delay-tuned neurons contribute to the perception of object range, because focal inactivation of these neurons disrupts behavioural discrimination of range. We report here that response properties of delay-tuned neurons in the cortical tonotopic area of the bat, Eptesicus, transform the sequential arrival times of echoes with different delays into a concurrent, accumulating neural representation of multiple objects at different ranges. The sharpness of delay tuning systematically increases at each best delay in a subpopulation of these neurons while responses to echoes at different delays are accumulated. The resulting concurrent, multiresolution representation of echo delay corresponds to neural implementation of a common representation of images used in computational vision and may provide the basis for representing acoustic images of multiple objects as acoustic 'scenes'.

Acoustic Stimulation↗

The generic viewpoint assumption in a framework for visual perception.

A visual system makes assumptions in order to interpret visual data. The assumption of 'generic view' states that the observer is not in a special position relative to the scene. Researchers commonly use a binary decision of generic or accidental view to disqualify scene interpretations that assume accidental viewpoints. Here we show how to use the generic view assumption, and others like it, to quantify the likelihood of a view, adding a new term to the probability of a given image interpretation. The resulting framework better models the visual world and reduces the reliance on other prior assumptions. It may lead to computer vision algorithms of greater power and accuracy, or to better models of human vision. We show applications to the problems of inferring shape, surface reflectance properties, and motion from images.

Humans↗

Models of object recognition.

Understanding how biological visual systems recognize objects is one of the ultimate goals in computational neuroscience. From the computational viewpoint of learning, different recognition tasks, such as categorization and identification, are similar, representing different trade-offs between specificity and invariance. Thus, the different tasks do not require different classes of models. We briefly review some recent trends in computational vision and then focus on feedforward, view-based models that are supported by psychophysical and physiological data.

Animals↗

Acquisition of 3-dimensional shapes from images.

Advances in computer vision have started to infiltrate the specialty of orthodontics. During the past few years, a number of new products have appeared that are capable of extracting the 3-dimensional (3-D) structure of an object just by "looking." Examples include laser scanners for creating 3-D models of the face, and hand-held scanners for creating virtual models of the teeth. Such noninvasive methods will surely evolve rapidly and be applied to a multitude of diagnostic and therapeutic modalities, changing the way we think and practice. This article introduces the basic principles behind such technology so that we can better appreciate its advantages, limitations and possibilities. From the large number of methods for acquiring 3-D shapes from images, 4 were selected and are described below. For more comprehensive coverage, see the book by Klette et al (1).

Humans↗

Action categories and the perception of biological motion.

Johansson filmed walkers and runners in a dark room with lights attached to their main joints and demonstrated that such moving light spots were perceived as human movements. To extend this finding the detection and recognition of Johansson displays of different kinds of movements under three light-spot conditions were studied to determine how human actions are perceived on the basis of biological-motion information. Locomotory, instrumental, and social actions were presented in each condition, namely in normal Johansson (light attached to joints), inter-joint (light attached between joints), and upside-down Johansson. Subjects' verbal responses and recognition times were measured. Locomotory actions were recognised better and faster than social and instrumental actions. Furthermore, biological motions were recognised much better and faster when the light-spot displays were presented in the normal orientation rather than upside down. Recognition rate was only slightly impaired under the inter-joint condition. It is argued that the perceptual analysis of actions and movements starts primarily on an intermediate level of action coding and comprises more than just the similarity of movement patterns or simple structures. Additionally, coding of dynamic phase relations and semantic coding take place at very early stages of the processing of biological motion. Implications of these results for computer vision, perceptual models, and mental representations are discussed.

Female↗

Shape from shading. I: Surface curvature and orientation.

The human visual system makes effective use of shading alone in recovering the shape of objects. Pictures of sculptures are readily interpreted--a situation where shading provides virtually the sole cue to shape. However, shading has been considered a poor cue to depth in comparison with retinal disparity and kinetic cues. Curvature discrimination thresholds were measured with the use of a surface-alignment task for a range of surface curvatures from 0.16 cm-1 to 1.06 cm-1. Weber fractions were around 0.1, demonstrating considerable precision in this task. Weber fractions did not vary substantially as a function of surface curvature. Rotation of the light source around the line of sight had no effect on curvature discrimination but rotation towards the viewer increased discrimination thresholds. In contrast, slant discrimination declined with rotation of the light-source vector towards the viewpoint. When a band-limited random grey-level texture was mapped onto the sphere, curvature discrimination thresholds increased gradually as a function of texture contrast, even though texture and shading provided consistent cues to depth. Adding texture also increased slant discrimination thresholds, demonstrating that texture can act as a source of noise in shape-from-shading tasks. The psychophysical findings have been used to evaluate whether current algorithms for shape from shading in computer vision could serve as models of human three-dimensional shape analysis and to highlight low-level intramodular interactions between depth cues. It is demonstrated that, in the case of surfaces defined by shading, curvature descriptions are primary and do not depend upon the prior encoding of surface orientation, and Koenderink's local-shape index is suggested as an alternative intermediate representation of surface shape in the human visual system.

Attention↗

Colour in a larger perspective: the rebirth of Gestalt psychology.

This overview takes the reader from the classical contrast and assimilation studies of the past to today's colour research, in a broad sense, with its renewed emphasis on the phenomenological qualities of visual perception. It shows how the shift in paradigm from local to global effects in single-unit recordings prompted a reappraisal of appearance in visual experiments, not just in colour, but in the perception of motion, texture, and depth as well. Gestalt ideas placed in the context of modern concepts are shown to inspire psychophysicists, neurophysiologists, and computational vision scientists alike. Feedforward, horizontal interactions, and feedback are discussed as potential neuronal mechanisms to account for phenomena such as uniform surfaces, filling-in, and grouping arising from processes beyond the classical receptive field. A look forward towards future developments in the field of figure-ground segregation (Gestalt formation) concludes the article.

Color Perception↗

Subjective surfaces: a method for completing missing boundaries.

We present a model and algorithm for segmentation of images with missing boundaries. In many situations, the human visual system fills in missing gaps in edges and boundaries, building and completing information that is not present. This presents a considerable challenge in computer vision, since most algorithms attempt to exploit existing data. Completion models, which postulate how to construct missing data, are popular but are often trained and specific to particular images. In this paper, we take the following perspective: We consider a reference point within an image as given and then develop an algorithm that tries to build missing information on the basis of the given point of view and the available information as boundary data to the algorithm. We test the algorithm on some standard images, including the classical triangle of Kanizsa and low signal/noise ratio medical images.

Journal Article↗

Repeated low-level red-light therapy for improving asthenopic symptoms and accommodation in presbyopia.

BACKGROUND: To assess the short-term effectiveness of repeated low-level red light (RLRL) therapy in relieving asthenopia and enhancing accommodation in presbyopia. METHODS: This randomized, parallel-group, double-masked clinical trial enrolled adults with presbyopia and self-reported asthenopia. Participants were allocated using computer-generated randomization and randomly assigned at a 1:1 ratio to RLRL or sham groups. Blinding included participants, examiners, assessors, and statisticians. The primary outcome was the change from baseline in the Computer Vision Syndrome Questionnaire (CVS-Q) score at day 31. Secondary outcomes were the change in accommodative amplitude (AA), Near Activity Visual Questionnaire (NAVQ) score, habitual near visual acuity, near-addition power, accommodative facility, positive and negative relative accommodation, binocular cross-cylinder response, and accommodative convergence-to-accommodation ratio. Continuous outcomes were analyzed using linear mixed-effects models. RESULTS: Sixty-four of 66 randomized participants (aged 41-62 years) completed the 1-month trial. At day 31, RLRL showed greater improvement than sham in CVS-Q score (adjusted mean difference, -1.75 points; 95% CI, -3.10 to -0.39), binocular AA (1.09 D; 95% CI, 0.37 to 1.82), and NAVQ score (-8.07 points; 95% CI, -14.17 to -1.97). The effect on AA was most pronounced in a subgroup of eyes with baseline amplitude >2.0 D (adjusted mean difference 1.33 D; 95% CI 0.32-2.34). Other measures did not differ between groups at each visit. No treatment-related adverse events were reported. Adherence was similar between groups (mean compliance: 98.2% vs 97.5%). CONCLUSIONS: Short-term treatment with RLRL significantly reduced asthenopic symptoms and improved accommodative amplitude in individuals with presbyopia.Trial registration: NCT06745661 (registered December 8, 2024).

Humans↗

A simple radiographic measurement method for polyethylene wear in total knee arthroplasty.

This study describes a new method for evaluating polyethylene wear in total knee arthroplasty. Since the amount of wear is dependent on a number of variables such as the weight and activity of the patient, it should be estimated based on in vivo measurements. We used a computer vision technique called three-dimensional/two-dimensional (3-D/2-D) matching to perform in vivo assessment using a single-plane radiograph. Using the 3-D/2-D matching algorithm we estimated the 3-D position and orientation of each knee implant and then measured the femorotibial distance, which is defined as the shortest perpendicular distance from the tibial tray to the femoral component. The accuracy of the proposed 3-D/2-D matching method was determined by in vitro investigations. The worst errors in in-plane/out-of-plane translations and rotations were 0.20/1.95 mm and 0.17/0.29 degrees, respectively. The root-mean-square error in femorotibial distance measurements using real polyethylene inserts was 0.04 mm. Results of in vivo femorotibial distance measurements are also described.

Humans↗

Quantitative light microscopy of combined perfusion and freezing processes.

The rational design of cryopreservation protocols for living tissues demands an understanding of the mechanisms of mass transport between cells and their environment throughout the entire process. We have developed a new microscope stage to enable a specimen to be viewed continuously during a preservation protocol, including the addition and removal of cryoprotective additives and freezing and thawing. The specimen is contained in a sealed chamber having inlet and outlet ports for admitting and collecting perfusate solution, the entire volume of which may be exchanged with a time constant of 1-5 s, depending on the solution viscosity. The temperature of the active area of the stage is regulated by the standard techniques of convection cryomicroscopy over a range in excess of 50 to -100 degrees C. A series of experiments has been performed on this system to measure the osmotic behaviour of rat pancreas islets during the addition and removal of dimethyl sulphoxide at temperatures between 25 and -10 degrees C. The technique involves mounting a single islet onto the low-temperature stage so that it is constrained from lateral movement by a specially sized mesh. Both the system temperature and chemical composition are monitored and controlled simultaneously and independently; as a consequence, virtually any defined cryopreservation protocol may be imposed on the specimen. For making permeability measurements, the bathing medium of the specimen may be changed very rapidly to produce a defined osmotic stress. Alternatively, the specimen may be subcooled to a specific and fixed subzero temperature, at which point ice is nucleated in the extracellular medium, creating a near instantaneous change in composition. The temporal alteration in specimen size is monitored by video microscopy and quantified by computer vision analysis methods. One of several mass transfer models is fitted to the data to estimate the membrane permeability based on the assumption of either transport dominated by the movement of water or simultaneous coupled flows of water and cryoprotective agent.

Animals↗

Three-dimensional reconstruction of vascular trees: experimental evaluation.

This paper is the second of two that together present a novel approach to the problem of reconstructing vascular trees from a small number of projections. Previously, we described the reconstruction algorithm and how it effectively circumvents the matching or "correspondence problem" found in most photogrammetric or computer-vision-based approaches. The algorithm is fully automatic and assumes that the imaging geometry is known, the vascular tree is a connected structure, and that its center-lines have been identified in three or more images. It employs consistency and connectivity constraints and comprises three steps: The first generates a connected structure representing the multiplicity of solutions that are consistent with the first two views; the second assigns a measure of agreement to each branch in this structure based on one or more additional projections; and the third step employs this measure to distinguish between those branches comprising the vasculature and the accompanying artifacts. This paper addresses the issue of validation via simulations and experiments. In addition to a clinical case, we examine the performance of the algorithm when applied to simulated projections of two 3-D vascular models, both representative of the complexity faced in coronary and cerebral angiography. The results in each instance are impressive and demonstrate that adequate reconstructions may be obtained with as few as three distinct views.

Algorithms↗

Diffuse capillary telangiectasia of the brain manifested as a slowly progressive course.

Brain capillary telangiectasia (BCT) are usually small, solitary, benign in clinical manifestation, or found incidentally at autopsy. However, diffuse BCT are rarely reported. A 39-year-old woman had the first generalized seizure 10 years previously. Thereafter, she had sustained progressive spastic paraparesis and blurred vision. Computed tomography (CT) showed diffuse brain atrophy with numerous calcified spots. Contrast T1-weighted magnetic resonance images showed diffuse faintly enhancing lesions with stippled appearance in the whole brain. Cerebral angiography revealed multiple small nets of dilated capillaries and Xenon CT showed diffuse low cerebral blood flow in the resting and good cerebral reserve capacity in the acetazolamide challenge test. Therefore, diffuse BCT was diagnosed clinically. Although BCT are not rare, diffuse BCT with slowly progressive neurological symptoms had never been reported before. It causes global cerebral ischemia that lead to brain atrophy and a degenerative course.

Adult↗

Pulsatile influxes of H+, K+ and Ca2+ lag growth pulses of Lilium longiflorum pollen tubes.

Fluxes of H+, K+ and Ca2+ were measured with self-referencing ion-selective probes, near the plasma membrane of growing Lilium longiflorum pollen tubes. Measurements from three regions around short, steady-growing tubes showed small, steady influx of H+ over the distal 40 microm and a region of the tube within 50-100 microm of the grain with larger magnitude efflux from the grain. K+ fluxes were immeasurable in short tubes. Measurements of longer tubes that were growing in a pulsatile manner revealed a pulsatile influx of both H+ and K+ at the growing tip. The average fluxes at the cell surface during the peaks of the H+ and K+ pulses were 489+/-81 and 688+/-144 pmol cm-2 second-1, respectively. Growth was measured by tracking the pollen tips with a computer vision system that achieved a spatial resolution of approximately 1/10 pixel. The high spatial resolution enabled the detection of growth, and thus the changes in growth rates, with a temporal sampling rate of 1 frame/second. These data show that the H+ and K+ pulses have a phase lag of 103+/-9 and 100+/-11 degrees, respectively, with respect to the growth pulses. Calcium fluxes were also measured in growing tubes. During steady growth, the calcium influx was relatively steady. When pulsatile growth began, the basal Ca2+ influx decreased and a pulsatile component appeared, superimposed on the reduced basal Ca2+ flux. The peaks of the Ca2+ pulses at the cell surface averaged 38.4+/-2.5 pmol cm-2 second-1. Longer tubes had large pulsatile Ca2+ fluxes with smaller baseline fluxes. The Ca2+ influx pulses had a phase lag of 123+/-9 degrees with respect to the growth pulses.

Calcium↗

Spectral sharpening with positivity

Spectral sharpening is a method for developing camera or other optical-device sensor functions that are more narrowband than those in hardware, by means of a linear transform of sensor functions. The utility of such a transform is that many computer vision and color-correction algorithms perform better in a sharpened space, and thus such a space can be used as an intermediate representation for carrying out calculations. In this paper we consider how one may sharpen sensor functions such that the transformed sensors are all positive. We show that constrained optimization can be used to produce positive sensors in two fundamentally different ways: by constraining the coefficients in the transform or by constraining the functions directly. In the former method, we prove that convexity can be used to constrain the solution exactly. In a sense, we are continuing the work of MacAdam and of Pearson and Yule, who formed positive combinations of the color-matching functions. However, the advantage of the spectral sharpening approach is that not only can we produce positive curves, but the process is "steerable" in that we can produce positive curves with as good or better properties for sharpening within a given set of sharpening intervals. At base, however, it is positive colors in the transformed space that are the prime objective. Therefore we also carry out sharpening of sensor curves governed not by positivity of the curves themselves but of colors resulting from them. Curves that result have negative lobes but generate positive colors. We find that this type of constrained sharpening generates the best results, which are almost as good as for unconstrained sharpening but without the penalty of negative colors. All methods discussed may be used with any number of sensors.

Journal Article↗

Color constancy at a pixel.

In computational terms we can solve the color constancy problem if device red, green, and blue sensor responses, or RGB's, for surfaces seen under an unknown illuminant can be mapped to corresponding RGB's under a known reference light. In recent years almost all authors have argued that this three-dimensional problem is too hard. It is argued that because a bright light striking a dark surface results in the same physical spectra as those of a dim light incident on a light surface, the magnitude of RGB's cannot be recovered. Consequently, modern color constancy algorithms attempt only to recover image chromaticities under the reference light: They solve a two-dimensional problem. While significant progress has been made toward achieving chromaticity constancy, recent work has shown that the most advanced algorithms are unable to render chromaticity stable enough so that it can be used as a cue for object recognition [B. V. Funt, K. Bernard, and L. Martin, in Proceedings of the Fifth European Conference on Computer Vision (European Vision Society, Springer-Verlag, Berlin, 1998), Vol. II, p. 445.] We take this reductionist approach a little further and look at the one-dimensional color constancy problem. We ask, Is there a single color coordinate, a function of image chromaticities, for which the color constancy problem can be solved? Our answer is an emphatic yes. We show that there exists a single invariant color coordinate, a function of R, G, and B, that depends only on surface reflectance. Two corollaries follow. First, given an RGB image of a scene viewed under any illuminant, we can trivially synthesize the same gray-scale image (we simply code the invariant coordinate as a gray scale). Second, this result implies that we can solve the one-dimensional color constancy problem at a pixel (in scenes with no color diversity whatsoever). We present experiments that show that invariant gray-scale histograms are a stable feature for object recognition. Indexing on invariant distributions supports almost perfect recognition for a dataset of 11 objects viewed under five colored lights. In contrast, object recognition based on chromaticity histograms (post-color constancy preprocessing) delivers much poorer recognition.

Artificial Intelligence↗

On the relationship between radiance and irradiance: determining the illumination from images of a convex Lambertian object.

We present a theoretical analysis of the relationship between incoming radiance and irradiance. Specifically, we address the question of whether it is possible to compute the incident radiance from knowledge of the irradiance at all surface orientations. This is a fundamental question in computer vision and inverse radiative transfer. We show that the irradiance can be viewed as a simple convolution of the incident illumination, i.e., radiance and a clamped cosine transfer function. Estimating the radiance can then be seen as a deconvolution operation. We derive a simple closed-form formula for the irradiance in terms of spherical harmonic coefficients of the incident illumination and demonstrate that the odd-order modes of the lighting with order greater than 1 are completely annihilated. Therefore these components cannot be estimated from the irradiance, contradicting a theorem that is due to Preisendorfer. A practical realization of the radiance-from-irradiance problem is the estimation of the lighting from images of a homogeneous convex curved Lambertian surface of known geometry under distant illumination, since a Lambertian object reflects light equally in all directions proportional to the irradiance. We briefly discuss practical and physical considerations and describe a simple experimental test to verify our theoretical results.

Journal Article↗