PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Pattern Recognition, Automated”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Automated derivation and refinement of sequence length patterns for protein sequences using evolutionary computation.

Several stratagems are used in protein bioinformatics for the classification of proteins based on sequence, structure or function. We explore the concept of a minimal signature embedded in a sequence that defines the likely position of a protein in a classification. Specifically, we address the derivation of sparse profiles for the G-protein coupled receptor (GPCR) clan of integral membrane proteins. We present an evolutionary algorithm (EA) for the derivation of sparse profiles (signatures) without the need to supply a multiple alignment. We also apply an evolution strategy (ES) to the problem of pattern and profile refinement. Patterns were derived for the GPCR 'superfamily' and GPCR families 1-3 individually from starting populations of randomly generated signatures, using a database of integral membrane protein sequences and an objective function using a modified receiver operator characteristic (ROC) statistic. The signature derived for the family 1 GPCR sequences was shown to perform very well in a stringent cross-validation test, detecting 76% of unseen GPCR sequences at 5% error. Application of the ES refinement method to a signature developed by a previously described method [Sadowski, M.I., Parish, J.H., 2003. Automated generation and refinement of protein signatures: case study with G-protein coupled receptors. Bioinformatics 19, 727-734] resulted in a 6% increase of coverage for 5% error as measured in the validation test. We note that there might be a limit to this or any classification of proteins based on patterns or schemata.

Algorithms↗

Noise reduction in spine videofluoroscopic images using the undecimated wavelet transform.

Videofluoroscopy permits using sequences of low quality images to study the spine movement. In this work the problem of enhancing the quality of these images is considered in order to facilitate the extraction of kinematic parameters. The method is based on the undecimated wavelet transform and on a preliminary training of a sub-set of images. The anatomical features are preserved using a mask. Key element of the method is its fast and automated implementation. The concept of improving the extraction of kinematic parameters by improving the image representation instead of the technique to extract these is also innovative. The technique has been tested on two sequences of images and the results demonstrates that the method enhances images not related with training sub-set.

Algorithms↗

A word-oriented approach to alignment validation.

MOTIVATION: Multiple sequence alignment at the level of whole proteomes requires a high degree of automation, precluding the use of traditional validation methods such as manual curation. Since evolutionary models are too general to describe the history of each residue in a protein family, there is no single algorithm/model combination that can yield a biologically or evolutionarily optimal alignment. We propose a 'shotgun' strategy where many different algorithms are used to align the same family, and the best of these alignments is then chosen with a reliable objective function. We present WOOF, a novel 'word-oriented' objective function that relies on the identification and scoring of conserved amino acid patterns (words) between pairs of sequences. RESULTS: Tests on a subset of reference protein alignments from BAliBASE showed that WOOF tended to rank the (manually curated) reference alignment highest among 1060 alternative (automatically generated) alignments for a majority of protein families. Among the automated alignments, there was a strong positive relationship between the WOOF score and similarity to the reference alignment. The speed of WOOF and its independence from explicit considerations of three-dimensional structure make it an excellent tool for analyzing large numbers of protein families. AVAILABILITY: On request from the authors.

Algorithms↗

Clustering of diverse genomic data using information fusion.

MOTIVATION: Genome sequencing projects and high-through-put technologies like DNA and Protein arrays have resulted in a very large amount of information-rich data. Microarray experimental data are a valuable, but limited source for inferring gene regulation mechanisms on a genomic scale. Additional information such as promoter sequences of genes/DNA binding motifs, gene ontologies, and location data, when combined with gene expression analysis can increase the statistical significance of the finding. This paper introduces a machine learning approach to information fusion for combining heterogeneous genomic data. The algorithm uses an unsupervised joint learning mechanism that identifies clusters of genes using the combined data. RESULTS: The correlation between gene expression time-series patterns obtained from different experimental conditions and the presence of several distinct and repeated motifs in their upstream sequences is examined here using publicly available yeast cell-cycle data. The results show that the combined learning approach taken here identifies correlated genes effectively. The algorithm provides an automated clustering method, but allows the user to specify apriori the influence of each data type on the final clustering using probabilities. AVAILABILITY: Software code is available by request from the first author. CONTACT: jkasturi@cse.psu.edu.

Algorithms↗

Automated matching of temporally sequential CT sections.

In the evaluation of patient response to therapy through measurements on thoracic computed tomography (CT) scans, the selection of anatomically equivalent sections in temporally sequential scans is required. We developed an automated method based on normalized mutual information (NMI) to expedite the selection of anatomically equivalent sections. The method requires as input two temporally sequential CT scans from the same patient. A specified section from the baseline scan is then compared with the sections of a follow-up scan. Each section in the follow-up scan is successively translated and rotated relative to the baseline section, and NMI is calculated. The section in the follow-up scan that yields the highest NMI with respect to the baseline section is selected as the matching section. The method was applied to a database of 22 pairs of temporally sequential CT scans from mesothelioma patients. Five observers manually selected their choice of the best anatomically matched section for each of three predetermined sections in the 22 baseline scans, and the range of selected sections was recorded. The automated method was applied to the same baseline sections to determine the computer-based anatomically matched sections in the corresponding follow-up scan. The automated process was performed using both original CT sections and sections automatically segmented so that only intrathoracic pixels contributed to NMI calculations. The accuracy of the automated method was quantified on a section-by-section basis by comparison with the range of sections selected by the observers. The automated method without segmentation selected equivalent sections within the observers' range for 54 of the 66 matching tasks (81.8%). An 11% improvement was achieved when thoracic segmentation was performed as a pre-processing step.

Adult↗

Automatic navigation path generation based on two-phase adaptive region-growing algorithm for virtual angioscopy.

In this paper, we propose a fast and automated navigation path generation algorithm to visualize inside of carotid artery using MR angiography images. The carotid artery is one of the body regions not accessible by real optical probe but can be visualized with virtual endoscopy. By applying two-phase adaptive region-growing algorithm, the carotid artery segmentation is started at the initial seed, which is located on the initially thresholded binary image. This segmentation algorithm automatically detects the branch position with stack feature. Combining with a priori knowledge of anatomic structure of carotid artery, the detected branch position is used to separate the carotid artery into internal carotid artery and external carotid artery. A fly-through path is determined to automatically move the virtual camera based on the intersecting coordinates of two bisectors on the circumscribed quadrangle of segmented carotid artery. In consideration of the interactive rendering speed and the usability of standard graphic hardware, endoscopic view of carotid artery is generated by using surface rendering algorithm with perspective projection method. In addition, the endoscopic view is provided with ray casting algorithm for off-line navigation of carotid artery. Experiments have been conducted on both mathematical phantom and clinical data sets. This algorithm is more effective than key-framing and topological thinning method in terms of automated features and computing time. This algorithm is also applicable to generate the centerline of renal artery, coronary artery, and airway tree which has tree-like cylinder shape of organ structures in the medical imagery.

Algorithms↗

Cell detection in phase-contrast images used for alpha-particle track-etch dosimetry: a semi-automated approach.

A novel alpha-particle irradiator has recently been developed that provides the ability to characterize cell response. The irradiator is comprised of a collimated, planar alpha-particle source which, from below, irradiates cells cultured on a track-etch material. Cells are imaged using phase-contrast microscopy before and following irradiation to obtain geometric information and survival rates; these can be used with data from alpha-particle track images to assess cell response. A key step in this process is determining cell location within the pre-irradiation images. Although this can be done completely by a human observer, the number of images requiring analysis makes the process time-consuming and tedious. To reduce the potential human error and decrease user interaction time, a semi-automated, computer-aided method of cell detection has been developed. The method employs a two-level adaptive thresholding technique to obtain size and position information about potential cell cytoplasms and nuclei. Proximity and geometry-based thresholds are then used to mark structures as cells. False-positive detections from the automated algorithm are due mostly to imperfections in the track-etch background, camera effects and cellular residue. To correct for these, a human observer reviews all detected structures, discarding false positives. When analysing two randomly selected cell dish image databases, the semi-automated method detected 92-94% of all cells and 94-97% of cells with a well-defined cytoplasm and nucleus while reducing human workload by 32-83%.

Algorithms↗

Matching and anatomical labeling of human airway tree.

Matching of corresponding branchpoints between two human airway trees, as well as assigning anatomical names to the segments and branchpoints of the human airway tree, are of significant interest for clinical applications and physiological studies. In the past, these tasks were often performed manually due to the lack of automated algorithms that can tolerate false branches and anatomical variability typical for in vivo trees. In this paper, we present algorithms that perform both matching of branchpoints and anatomical labeling of in vivo trees without any human intervention and within a short computing time. No hand-pruning of false branches is required. The results from the automated methods show a high degree of accuracy when validated against reference data provided by human experts. 92.9% of the verifiable branchpoint matches found by the computer agree with experts' results. For anatomical labeling, 97.1% of the automatically assigned segment labels were found to be correct.

Algorithms↗

Classification of breast masses in ultrasonic B scans using Nakagami and K distributions.

Classification of breast masses in greyscale ultrasound images is undertaken using a multiparameter approach. Five parameters reflecting the non-Rayleigh nature of the backscattered echo were used. These parameters, based mostly on the Nakagami and K distributions, were extracted from the envelope of the echoes at the site, boundary, spiculated region and shadow of the mass. They were combined to create a linear discriminant. The performance of this discriminant for the classification of breast masses was studied using a data set consisting of 70 benign and 29 malignant cases. The Az value for the discriminant was 0.96 +/- 0.02, showing great promise in the classification of masses into benign and malignant ones. The discriminant was combined with the level of suspicion values of the radiologist leading to an Az value of 0.97 +/- 0.014. The parameters used here can be calculated with minimal clinical intervention, so the method proposed here may therefore be easily implemented in an automated fashion. These results also support the recent reports suggesting that ultrasound may help as an adjunct to mammography in breast cancer diagnostics to enhance the classification of breast masses.

Adult↗

Accuracy of short-axis cardiac MRI automatically derived from scout acquisitions in free-breathing and breath-holding modes.

To qualitatively assess the accuracy of automated cardiovascular magnetic resonance planning procedures devised from scout acquisitions in free-breathing and breath-holding modes, to quantitatively evaluate the accuracy of the derived left ventricular volumes, mass and function and compare these parameters with the ones obtained from the manually planned acquisitions. Ten healthy volunteers underwent cardiovascular MR (CMR) acquisitions for ventricular function assessment. Short-axis data sets of the left ventricle (LV) were manually planned and generated twice in an automatic fashion. Automated planning parameters were derived from gated scout acquisitions in free-breathing and breath-holding modes. End-diastolic volume (EDV), end-systolic volume (ESV), ejection fraction (EF), and left ventricular mass (LVM) were measured. The agreement between the manual and automatic planning methods, as well as the variability of the aforementioned measurements were assessed. The differences between two automated planning methods were also compared. The mean differences between the manual and automated CMR planning derived from gated scouts in free-breathing mode were 8.05 ml (EDV), 1.84 ml (ESV), 0.69% (EF), and 4.72 g (LVM). The comparison between manual and automated CMR planning derived from gated scouts in breath-holding mode yielded the following differences: 4.22 ml (EDV), 0.34 ml (ESV), 0.3% (EF), and -0.72 mg (LVM). The variability coefficients were 3.72 and 3.66 (EDV), 5.6 and 8.19 (ESV), 3.46 and 4.31 (EF), 6.49 and 5.20 (LVM) for the automated CMR planning methods derived from scouts in free-breathing and breath-holding modes, respectively. Automated CMR planning methods can provide accurate measurements of LV dimensions in normal subjects, and therefore may be utilized in the clinical environment to provide a cost-effective solution for functional assessment of the human cardiovascular system.

Adult↗

Automatic extraction of acronym-meaning pairs from MEDLINE databases.

Acronyms are widely used in biomedical and other technical texts. Understanding their meaning constitutes an important problem in the automatic extraction and mining of information from text. Here we present a system called ACROMED that is part of a set of Information Extraction tools designed for processing and extracting information from abstracts in the Medline database. In this paper, we present the results of two strategies for finding the long forms for acronyms in biomedical texts. These strategies differ from previous automated acronym extraction methods by being tuned to the complex phrase structures of the biomedical lexicon and by incorporating shallow parsing of the text into the acronym recognition algorithm. The performance of our system was tested with several data sets obtaining a performance of 72 % recall with 97 % precision. These results are found to be better for biomedical texts than the performance of other acronym extraction systems designed for unrestricted text.

Abbreviations as Topic↗

PRINTS prepares for the new millennium.

PRINTS is a diagnostic collection of protein fingerprints. Fingerprints exploit groups of motifs to build characteristic family signatures, offering improved diagnostic reliability over single-motif approaches by virtue of the mutual context provided by motif neighbours. Around 1000 fingerprints have now been created and stored in PRINTS. The September 1998 release (version 20.0), encodes approximately 5700 motifs, covering a range of globular and membrane proteins, modular polypeptides and so on. The database is accessible via the DbBrowser Web Server at http://www.biochem.ucl.ac.uk/bsm/dbbrowser /. In addition to supporting its continued growth, recent enhancements to the resource include a BLAST server, and more efficient fingerprint search software, with improved statistics for estimating the reliability of retrieved matches. Current efforts are focused on the design of more automated methods for database maintenance; implementation of an object-relational schema for efficient data management; and integration with PROSITE, profiles, Pfam and ProDom, as part of the international InterPro project, which aims to unify protein pattern databases and offer improved tools for genome analysis.

Amino Acid Sequence↗

Rapid analysis of hematology image data: the ADC-500 preprocessor.

A sequential, pipeline processor (that we have named the ADC-500 preprocessor) has been developed which scene segments the three color image data from the ADC-500 optics one image element at a time, groups together image elements from each object in the scene and extracts features from each object. The processing occurs at television frame rates, requiring 16.7 msec to process the entire image. This speed was instrumental in allowing the ADC-500 automated differential analyzer to perform routine 500-cell differentials. The preprocessor also contains hardware which simplifies compilation of the three color histograms. The segmentation algorithms implemented in the preprocessor are multicolor extensions of the classical monochrome density histogram threshold method. For most cell image analysis tasks, a sequential pipeline processor of this type should be more economical and as fast or faster than a parallel processor.

Blood Cells↗

An energy-based three-dimensional segmentation approach for the quantitative interpretation of electron tomograms.

Electron tomography allows for the determination of the three-dimensional structures of cells and tissues at resolutions significantly higher than that which is possible with optical microscopy. Electron tomograms contain, in principle, vast amounts of information on the locations and architectures of large numbers of subcellular assemblies and organelles. The development of reliable quantitative approaches for the analysis of features in tomograms is an important problem, and a challenging prospect due to the low signal-to-noise ratios that are inherent to biological electron microscopic images. This is, in part, a consequence of the tremendous complexity of biological specimens. We report on a new method for the automated segmentation of HIV particles and selected cellular compartments in electron tomograms recorded from fixed, plastic-embedded sections derived from HIV-infected human macrophages. Individual features in the tomogram are segmented using a novel robust algorithm that finds their boundaries as global minimal surfaces in a metric space defined by image features. The optimization is carried out in a transformed spherical domain with the center an interior point of the particle of interest, providing a proper setting for the fast and accurate minimization of the segmentation energy. This method provides tools for the semi-automated detection and statistical evaluation of HIV particles at different stages of assembly in the cells and presents opportunities for correlation with biochemical markers of HIV infection. The segmentation algorithm developed here forms the basis of the automated analysis of electron tomograms and will be especially useful given the rapid increases in the rate of data acquisition. It could also enable studies of much larger data sets, such as those which might be obtained from the tomographic analysis of HIV-infected cells from studies of large populations.

Algorithms↗

A novel approach for high-quality microarray processing using third-dye array visualization technology.

Historically, microarray image processing has been technically challenging in obtaining quality gene expression data. After hybridization of Cy3- and Cy5-labeled samples, images are collected and processed to obtain gene expression ratio measurements for each of the elements on the array. The hybridization process often brings in contaminating noise, which can make correct identification of the signal difficult. In addition, spot intensity levels are highly variable due to the expression differences of different genes, and weak spots are often difficult to detect. These conditions are further complicated by inherent irregularities in spot position, shape, and size commonly found on high-density microarrays, making image processing an often labor-intensive task that is difficult to reliably automate. We previously reported a novel third-dye array visualization (TDAV) technology that allows prehybridization visualization and quality control of printed arrays. Here, we present a new microarray image processing approach utilizing TDAV. By incorporating the third-dye image, we show that overall quality of the microarray data is significantly improved, and automation of processing is feasible and reliable. Furthermore, we demonstrate use of the third-dye image to better quality control microarray image analysis. Both the principle and implementation of the approach are presented in detail, with experimental results.

Algorithms↗

Gene clustering by latent semantic indexing of MEDLINE abstracts.

MOTIVATION: A major challenge in the interpretation of high-throughput genomic data is understanding the functional associations between genes. Previously, several approaches have been described to extract gene relationships from various biological databases using term-matching methods. However, more flexible automated methods are needed to identify functional relationships (both explicit and implicit) between genes from the biomedical literature. In this study, we explored the utility of Latent Semantic Indexing (LSI), a vector space model for information retrieval, to automatically identify conceptual gene relationships from titles and abstracts in MEDLINE citations. RESULTS: We found that LSI identified gene-to-gene and keyword-to-gene relationships with high average precision. In addition, LSI identified implicit gene relationships based on word usage patterns in the gene abstract documents. Finally, we demonstrate here that pairwise distances derived from the vector angles of gene abstract documents can be effectively used to functionally group genes by hierarchical clustering. Our results provide proof-of-principle that LSI is a robust automated method to elucidate both known (explicit) and unknown (implicit) gene relationships from the biomedical literature. These features make LSI particularly useful for the analysis of novel associations discovered in genomic experiments. AVAILABILITY: The 50-gene document collection used in this study can be interactively queried at http://shad.cs.utk.edu/sgo/sgo.html.

Abstracting and Indexing↗

Computerized characterization of breast masses on three-dimensional ultrasound volumes.

We are developing computer vision techniques for the characterization of breast masses as malignant or benign on radiologic examinations. In this study, we investigated the computerized characterization of breast masses on three-dimensional (3-D) ultrasound (US) volumetric images. We developed 2-D and 3-D active contour models for automated segmentation of the mass volumes. The effect of the initialization method of the active contour on the robustness of the iterative segmentation method was studied by varying the contour used for its initialization. For a given segmentation, texture and morphological features were automatically extracted from the segmented masses and their margins. Stepwise discriminant analysis with the leave-one-out method was used to select effective features for the classification task and to combine these features into a malignancy score. The classification accuracy was evaluated using the area Az under the receiver operating characteristic (ROC) curve, as well as the partial area index Az(0.9), defined as the relative area under the ROC curve above a sensitivity threshold of 0.9. For the purpose of comparison with the computer classifier, four experienced breast radiologists provided malignancy ratings for the 3-D US masses. Our dataset consisted of 3-D US volumes of 102 biopsied masses (46 benign, 56 malignant). The classifiers based on 2-D and 3-D segmentation methods achieved test Az values of 0.87+/-0.03 and 0.92+/-0.03, respectively. The difference in the Az values of the two computer classifiers did not achieve statistical significance. The Az values of the four radiologists ranged between 0.84 and 0.92. The difference between the computer's Az value and that of any of the four radiologists did not achieve statistical significance either. However, the computer's Az(0.9) value was significantly higher than that of three of the four radiologists. Our results indicate that an automated and effective computer classifier can be designed for differentiating malignant and benign breast masses on 3-D US volumes. The accuracy of the classifier designed in this study was similar to that of experienced breast radiologists.

Algorithms↗

Computerized analysis of abnormal asymmetry in digital chest radiographs: evaluation of potential utility.

The purpose of this study was to develop and test a computerized method for the fully automated analysis of abnormal asymmetry in digital posteroanterior (PA) chest radiographs. An automated lung segmentation method was used to identify the aerated lung regions in 600 chest radiographs. Minimal a priori lung morphology information was required for this gray-level thresholding-based segmentation. Consequently, segmentation was applicable to grossly abnormal cases. The relative areas of segmented right and left lung regions in each image were compared with the corresponding area distributions of normal images to determine the presence of abnormal asymmetry. Computerized diagnoses were compared with image ratings assigned by a radiologist. The ability of the automated method to distinguish normal from asymmetrically abnormal cases was evaluated by using receiver operating characteristic (ROC) analysis, which yielded an area under the ROC curve of 0.84. This automated method demonstrated promising performance in its ability to detect abnormal asymmetry in PA chest images. We believe this method could play a role in a picture archiving and communications (PACS) environment to immediately identify abnormal cases and to function as one component of a multifaceted computer-aided diagnostic scheme.

Databases as Topic↗