PubMed Health⌕ Search

Biomedical subjects

Walter L Ruzzo

Publications and source records attributed to Walter L Ruzzo.

11 recordsLinked to original sources

A regression-based K nearest neighbor algorithm for gene function prediction from heterogeneous data.

BACKGROUND: As a variety of functional genomic and proteomic techniques become available, there is an increasing need for functional analysis methodologies that integrate heterogeneous data sources. METHODS: In this paper, we address this issue by proposing a general framework for gene function prediction based on the k-nearest-neighbor (KNN) algorithm. The choice of KNN is motivated by its simplicity, flexibility to incorporate different data types and adaptability to irregular feature spaces. A weakness of traditional KNN methods, especially when handling heterogeneous data, is that performance is subject to the often ad hoc choice of similarity metric. To address this weakness, we apply regression methods to infer a similarity metric as a weighted combination of a set of base similarity measures, which helps to locate the neighbors that are most likely to be in the same class as the target gene. We also suggest a novel voting scheme to generate confidence scores that estimate the accuracy of predictions. The method gracefully extends to multi-way classification problems. RESULTS: We apply this technique to gene function prediction according to three well-known Escherichia coli classification schemes suggested by biologists, using information derived from microarray and genome sequencing data. We demonstrate that our algorithm dramatically outperforms the naive KNN methods and is competitive with support vector machine (SVM) algorithms for integrating heterogenous data. We also show that by combining different data sources, prediction accuracy can improve significantly CONCLUSION: Our extension of KNN with automatic feature weighting, multi-class prediction, and probabilistic inference, enhance prediction accuracy significantly while remaining efficient, intuitive and flexible. This general framework can also be applied to similar classification problems involving heterogeneous datasets.

Algorithms↗

Bone morphogenetic protein 4: potential regulator of shear stress-induced graft neointimal atrophy.

OBJECTIVE: Placement in baboons of a distal femoral arteriovenous fistula increases shear stress through aortoiliac polytetrafluoroethylene (PTFE) grafts and induces regression of a preformed neointima. Atrophy of the neointima might be controlled by shear stress-induced genes, including the bone morphogenetic proteins (BMPs). We have investigated the expression and function of BMPs 2, 4, and 5 in the graft neointima and in cultured baboon smooth muscle cells (SMCs). METHODS: Baboons received bilateral aortoiliac PTFE grafts and 8 weeks later, a unilateral femoral arteriovenous fistula. RESULTS: Quantitative polymerase chain reaction showed that high shear stress increased BMP2, 4, and 5 messenger RNA (mRNA) in graft intima between 1 and 7 days, while noggin (a BMP inhibitor) mRNA was decreased. BMP4 most potently (60% inhibition) inhibited platelet-derived growth factor-stimulated SMC proliferation compared with BMP2 and BMP5 (31% and 26%, respectively). BMP4 also increased SMC death by 190% +/- 10%. Noggin reversed the antiproliferative and proapoptotic effects of BMP4. Finally, Western blotting confirmed BMP4 protein upregulation by high shear stress at 4 days. BMP4 expression demonstrated by in situ hybridization was confined to endothelial cells. CONCLUSIONS: Increased BMPs (particularly BMP4) coupled with decreased noggin may promote high shear stress-mediated graft neointimal atrophy by inhibiting SMC proliferation and increasing SMC death.

Animals↗

Macronuclear genome sequence of the ciliate Tetrahymena thermophila, a model eukaryote.

The ciliate Tetrahymena thermophila is a model organism for molecular and cellular biology. Like other ciliates, this species has separate germline and soma functions that are embodied by distinct nuclei within a single cell. The germline-like micronucleus (MIC) has its genome held in reserve for sexual reproduction. The soma-like macronucleus (MAC), which possesses a genome processed from that of the MIC, is the center of gene expression and does not directly contribute DNA to sexual progeny. We report here the shotgun sequencing, assembly, and analysis of the MAC genome of T. thermophila, which is approximately 104 Mb in length and composed of approximately 225 chromosomes. Overall, the gene set is robust, with more than 27,000 predicted protein-coding genes, 15,000 of which have strong matches to genes in other organisms. The functional diversity encoded by these genes is substantial and reflects the complexity of processes required for a free-living, predatory, single-celled organism. This is highlighted by the abundance of lineage-specific duplications of genes with predicted roles in sensing and responding to environmental conditions (e.g., kinases), using diverse resources (e.g., proteases and transporters), and generating structural complexity (e.g., kinesins and dyneins). In contrast to the other lineages of alveolates (apicomplexans and dinoflagellates), no compelling evidence could be found for plastid-derived genes in the genome. UGA, the only T. thermophila stop codon, is used in some genes to encode selenocysteine, thus making this organism the first known with the potential to translate all 64 codons in nuclear genes into amino acids. We present genomic evidence supporting the hypothesis that the excision of DNA from the MIC to generate the MAC specifically targets foreign DNA as a form of genome self-defense. The combination of the genome sequence, the functional diversity encoded therein, and the presence of some pathways missing from other model organisms makes T. thermophila an ideal model for functional genomic studies to address biological, biomedical, and biotechnological questions of fundamental importance.

Animals↗

CMfinder--a covariance model based RNA motif finding algorithm.

MOTIVATION: The recent discoveries of large numbers of non-coding RNAs and computational advances in genome-scale RNA search create a need for tools for automatic, high quality identification and characterization of conserved RNA motifs that can be readily used for database search. Previous tools fall short of this goal. RESULTS: CMfinder is a new tool to predict RNA motifs in unaligned sequences. It is an expectation maximization algorithm using covariance models for motif description, featuring novel integration of multiple techniques for effective search of motif space, and a Bayesian framework that blends mutual information-based and folding energy-based approaches to predict structure in a principled way. Extensive tests show that our method works well on datasets with either low or high sequence similarity, is robust to inclusion of lengthy extraneous flanking sequence and/or completely unrelated sequences, and is reasonably fast and scalable. In testing on 19 known ncRNA families, including some difficult cases with poor sequence conservation and large indels, our method demonstrates excellent average per-base-pair accuracy--79% compared with at most 60% for alternative methods. More importantly, the resulting probabilistic model can be directly used for homology search, allowing iterative refinement of structural models based on additional homologs. We have used this approach to obtain highly accurate covariance models of known RNA motifs based on small numbers of related sequences, which identified homologs in deeply-diverged species.

Algorithms↗

Sequence-based heuristics for faster annotation of non-coding RNA families.

MOTIVATION: Non-coding RNAs (ncRNAs) are functional RNA molecules that do not code for proteins. Covariance Models (CMs) are a useful statistical tool to find new members of an ncRNA gene family in a large genome database, using both sequence and, importantly, RNA secondary structure information. Unfortunately, CM searches are extremely slow. Previously, we created rigorous filters, which provably sacrifice none of a CM's accuracy, while making searches significantly faster for virtually all ncRNA families. However, these rigorous filters make searches slower than heuristics could be. RESULTS: In this paper we introduce profile HMM-based heuristic filters. We show that their accuracy is usually superior to heuristics based on BLAST. Moreover, we compared our heuristics with those used in tRNAscan-SE, whose heuristics incorporate a significant amount of work specific to tRNAs, where our heuristics are generic to any ncRNA. Performance was roughly comparable, so we expect that our heuristics provide a high-quality solution that--unlike family-specific solutions--can scale to hundreds of ncRNA families. AVAILABILITY: The source code is available under GNU Public License at the supplementary web site.

Algorithms↗

6S RNA is a widespread regulator of eubacterial RNA polymerase that resembles an open promoter.

6S RNA is an abundant noncoding RNA in Escherichia coli that binds to sigma70 RNA polymerase holoenzyme to globally regulate gene expression in response to the shift from exponential growth to stationary phase. We have computationally identified >100 new 6S RNA homologs in diverse eubacterial lineages. Two abundant Bacillus subtilis RNAs of unknown function (BsrA and BsrB) and cyanobacterial 6Sa RNAs are now recognized as 6S homologs. Structural probing of E. coli 6S RNA and a B. subtilis homolog supports a common secondary structure derived from comparative sequence analysis. The conserved features of 6S RNA suggest that it binds RNA polymerase by mimicking the structure of DNA template in an open promoter complex. Interestingly, the two B. subtilis 6S RNAs are discoordinately expressed during growth, and many proteobacterial 6S RNAs could be cotranscribed with downstream homologs of the E. coli ygfA gene encoding a putative methenyltetrahydrofolate synthetase. The prevalence and robust expression of 6S RNAs emphasize their critical role in bacterial adaptation.

Bacillus subtilis↗

A glycine-dependent riboswitch that uses cooperative binding to control gene expression.

We identified a previously unknown riboswitch class in bacteria that is selectively triggered by glycine. A representative of these glycine-sensing RNAs from Bacillus subtilis operates as a rare genetic on switch for the gcvT operon, which codes for proteins that form the glycine cleavage system. Most glycine riboswitches integrate two ligand-binding domains that function cooperatively to more closely approximate a two-state genetic switch. This advanced form of riboswitch may have evolved to ensure that excess glycine is efficiently used to provide carbon flux through the citric acid cycle and maintain adequate amounts of the amino acid for protein synthesis. Thus, riboswitches perform key regulatory roles and exhibit complex performance characteristics that previously had been observed only with protein factors.

5' Untranslated Regions↗

Exploiting conserved structure for faster annotation of non-coding RNAs without loss of accuracy.

MOTIVATION: Non-coding RNAs (ncRNAs)-functional RNA molecules not coding for proteins-are grouped into hundreds of families of homologs. To find new members of an ncRNA gene family in a large genome database, covariance models (CMs) are a useful statistical tool, as they use both sequence and RNA secondary structure information. Unfortunately, CM searches are slow. Previously, we introduced 'rigorous filters', which provably sacrifice none of CMs' accuracy, although often scanning much faster. A rigorous filter, using a profile hidden Markov model (HMM), is built based on the CM, and filters the genome database, eliminating sequences that provably could not be annotated as homologs. The CM is run only on the remainder. Some biologically important ncRNA families could not be scanned efficiently with this technique, largely due to the significance of conserved secondary structure relative to primary sequence in identifying these families. Current heuristic filters are also expected to perform poorly on such families. RESULTS: By augmenting profile HMMs with limited secondary structure information, we obtain rigorous filters that accelerate CM searches for virtually all known ncRNA families from the Rfam Database and tRNA models in tRNAscan-SE. These filters scan an 8 gigabase database in weeks instead of years, and uncover homologs missed by heuristic techniques to speed CM searches. AVAILABILITY: Software in development; contact the authors.

Algorithms↗

Atherosclerotic plaque smooth muscle cells have a distinct phenotype.

OBJECTIVE: The present study addresses the question, "Are plaque smooth muscles cells (SMCs) genetically distinct from medial SMCs as reflected by the ability to maintain a distinctive expression phenotype in vitro?" METHODS AND RESULTS: Multiple cell strains were developed from carotid endarcterectomy specimens, and quadruplicate array hybridizations were completed for each sample. A new normalization protocol was developed and used to analyze the data. Permutation analysis suggests that most of the significant differences in expression could not have occurred by chance. A broad pattern of significant expression differences, consisting of almost 5% of the genes probed, was detected. Quantitative polymerase chain reaction (QPCR) confirmation was found in 70% of a subset of genes selected for validation. CONCLUSIONS: The SMC cultures were nearly indistinguishable by morphological features, population doubling time, and sensitivity to cell death induced by Fas cross-linking. Surprisingly, array expression analysis identified differences so extensive that we conclude that plaque and medial SMCs are distinctly different SMC cell types.

Apoptosis↗

Pre-mRNA secondary structure prediction aids splice site prediction.

Accurate splice site prediction is a critical component of any computational approach to gene prediction in higher organisms. Existing approaches generally use sequence-based models that capture local dependencies among nucleotides in a small window around the splice site. We present evidence that computationally predicted secondary structure of moderate length pre-mRNA subsequencies contains information that can be exploited to improve acceptor splice site prediction beyond that possible with conventional sequence-based approaches. Both decision tree and support vector machine classifiers, using folding energy and structure metrics characterizing helix formation near the splice site, achieve a 5-10% reduction in error rate with a human data set. Based on our data, we hypothesize that acceptors preferentially exhibit short helices at the splice site.

Computer Simulation↗

Transcriptional analyses of Barrett's metaplasia and normal upper GI mucosae.

Over the last two decades, the incidence of esophageal adenocarcinoma (EA) has increased dramatically in the US and Western Europe. It has been shown that EAs evolve from premalignant Barrett's esophagus (BE) tissue by a process of clonal expansion and evolution. However, the molecular phenotype of the premalignant metaplasia, and its relationship to those of the normal upper gastrointestinal (GI) mucosae, including gastric, duodenal, and squamous epithelium of the esophagus, has not been systematically characterized. Therefore, we used oligonucleotide-based microarrays to characterize gene expression profiles in each of these tissues. The similarity of BE to each of the normal tissues was compared using a series of computational approaches. Our analyses included esophageal squamous epithelium, which is present at the same anatomic site and exposed to similar conditions as Barrett's epithelium, duodenum that shares morphologic similarity to Barrett's epithelium, and adjacent gastric epithelium. There was a clear distinction among the expression profiles of gastric, duodenal, and squamous epithelium whereas the BE profiles showed considerable overlap with normal tissues. Furthermore, we identified clusters of genes that are specific to each of the tissues, to the Barrett's metaplastic epithelia, and a cluster of genes that was distinct between squamous and non-squamous epithelia.

Barrett Esophagus↗