PubMed Health⌕ Search

Biomedical subjects

Giri Narasimhan

Publications and source records attributed to Giri Narasimhan.

10 recordsLinked to original sources

Serial NetEvolve: a flexible utility for generating serially-sampled sequences along a tree or recombinant network.

UNLABELLED: Serial NetEvolve is a flexible simulation program that generates DNA sequences evolved along a tree or recombinant network. It offers a user-friendly Windows graphical interface and a Windows or Linux simulator with a diverse selection of parameters to control the evolutionary model. Serial NetEvolve is a modification of the Treevolve program with the following additional features: simulation of serially-sampled data, the choice of either a clock-like or a variable rate model of sequence evolution, sampling from the internal nodes and the output of the randomly generated tree or network in our newly proposed NeTwick format. AVAILABILITY: From website http://biorg.cis.fiu.edu/SNE Contacts: giri@cis.fiu.edu SUPPLEMENTARY INFORMATION: Manual and examples available from http://biorg.cis.fiu.edu/SNE.

Algorithms↗

An eco-informatics tool for microbial community studies: supervised classification of Amplicon Length Heterogeneity (ALH) profiles of 16S rRNA.

Support vector machines (SVM) and K-nearest neighbors (KNN) are two computational machine learning tools that perform supervised classification. This paper presents a novel application of such supervised analytical tools for microbial community profiling and to distinguish patterning among ecosystems. Amplicon length heterogeneity (ALH) profiles from several hypervariable regions of 16S rRNA gene of eubacterial communities from Idaho agricultural soil samples and from Chesapeake Bay marsh sediments were separately analyzed. The profiles from all available hypervariable regions were concatenated to obtain a combined profile, which was then provided to the SVM and KNN classifiers. Each profile was labeled with information about the location or time of its sampling. We hypothesized that after a learning phase using feature vectors from labeled ALH profiles, both these classifiers would have the capacity to predict the labels of previously unseen samples. The resulting classifiers were able to predict the labels of the Idaho soil samples with high accuracy. The classifiers were less accurate for the classification of the Chesapeake Bay sediments suggesting greater similarity within the Bay's microbial community patterns in the sampled sites. The profiles obtained from the V1+V2 region were more informative than that obtained from any other single region. However, combining them with profiles from the V1 region (with or without the profiles from the V3 region) resulted in the most accurate classification of the samples. The addition of profiles from the V 9 region appeared to confound the classifiers. Our results show that SVM and KNN classifiers can be effectively applied to distinguish between eubacterial community patterns from different ecosystems based only on their ALH profiles.

Artificial Intelligence↗

Clustering genes using gene expression and text literature data.

Clustering of gene expression data is a standard technique used to identify closely related genes. In this paper, we develop a new clustering algorithm, MSC (Multi-Source Clustering), to perform exploratory analysis using two or more diverse sources of data. In particular, we investigate the problem of improving the clustering by integrating information obtained from gene expression data with knowledge extracted from biomedical text literature. In each iteration of algorithm MSC, an EM-type procedure is employed to bootstrap the model obtained from one data source by starting with the cluster assignments obtained in the previous iteration using the other data sources. Upon convergence, the two individual models are used to construct the final cluster assignment. We compare the results of algorithm MSC for two data sources with the results obtained when the clustering is applied on the two sources of data separately. We also compare it with that obtained using the feature level integration method that performs the clustering after simply concatenating the features obtained from the two data sources. We show that the z-scores of the clustering results from MSC are better than that from the other methods. To evaluate our clusters better, function enrichment results are presented using terms from the Gene Ontology database. Finally, by investigating the success of motif detection programs that use the clusters, we show that our approach integrating gene expression data and text data reveals clusters that are biologically more meaningful than those identified using gene expression data alone.

Artificial Intelligence↗

Distinct transcriptional profiles characterize oral epithelium-microbiota interactions.

Transcriptional profiling, bioinformatics, statistical and ontology tools were used to uncover and dissect genes and pathways of human gingival epithelial cells that are modulated upon interaction with the periodontal pathogens Actinobacillus actinomycetemcomitans and Porphyromonas gingivalis. Consistent with their biological and clinical differences, the common core transcriptional response of epithelial cells to both organisms was very limited, and organism-specific responses predominated. A large number of differentially regulated genes linked to the P53 apoptotic network were found with both organisms, which was consistent with the pro-apoptotic phenotype observed with A. actinomycetemcomitans and anti-apoptotic phenotype of P. gingivalis. Furthermore, with A. actinomycetemcomitans, the induction of apoptosis did not appear to be Fas- or TNF(alpha)-mediated. Linkage of specific bacterial components to host pathways and networks provided additional insight into the pathogenic process. Comparison of the transcriptional responses of epithelial cells challenged with parental P. gingivalis or with a mutant of P. gingivalis deficient in production of major fimbriae, which are required for optimal invasion, showed major expression differences that reverberated throughout the host cell transcriptome. In contrast, gene ORF859 in A. actinomycetemcomitans, which may play a role in intracellular homeostasis, had a more subtle effect on the transcriptome. These studies help unravel the complex and dynamic interactions between host epithelial cells and endogenous bacteria that can cause opportunistic infections.

Aggregatibacter actinomycetemcomitans↗

Alginate production affects Pseudomonas aeruginosa biofilm development and architecture, but is not essential for biofilm formation.

Extracellular polymers can facilitate the non-specific attachment of bacteria to surfaces and hold together developing biofilms. This study was undertaken to qualitatively and quantitatively compare the architecture of biofilms produced by Pseudomonas aeruginosa strain PAO1 and its alginate-overproducing (mucA22) and alginate-defective (algD) variants in order to discern the role of alginate in biofilm formation. These strains, PAO1, Alg+ PAOmucA22 and Alg- PAOalgD, tagged with green fluorescent protein, were grown in a continuous flow cell system to characterize the developmental cycles of their biofilm formation using confocal laser scanning microscopy. Biofilm Image Processing (BIP) and Community Statistics (COMSTAT) software programs were used to provide quantitative measurements of the two-dimensional biofilm images. All three strains formed distinguishable biofilm architectures, indicating that the production of alginate is not critical for biofilm formation. Observation over a period of 5 days indicated a three-stage development pattern consisting of initiation, establishment and maturation. Furthermore, this study showed that phenotypically distinguishable biofilms can be quantitatively differentiated.

Alginates↗

MinPD: distance-based phylogenetic analysis and recombination detection of serially-sampled HIV quasispecies.

A new computational method to study within-host viral evolution is explored to better understand the evolution and pathogenesis of viruses. Traditional phylogenetic tree methods are better suited to study relationships between contemporaneous species, which appear as leaves of a phylogenetic tree. However, viral sequences are often sampled serially from a single host. Consequently, data may be available at the leaves as well as the internal nodes of a phylogenetic tree. Recombination may further complicate the analysis. Such relationships are not easily expressed by traditional phylogenetic methods. We propose a new algorithm, called MinPD, based on minimum pairwise distances. Our algorithm uses multiple distance matrices and correlation rules to output a MinPD tree or network. We test our algorithm using extensive simmulations and apply it to a set of HIV sequence data isolated from one patient over a period of ten years. The proposed visualization of the phylogenetic tree\network further enhances the benefits of our methods.

Algorithms↗

Degenerate primer design via clustering.

This paper describes a new strategy for designing degenerate primers for a given multiple alignment of amino acid sequences. Degenerate primers are useful for amplifying homologous genes. However, when a large collection of sequences is considered, no consensus region may exist in the multiple alignment, making it impossible to design a single pair of primers for the collection. In such cases, manual methods are used to find smaller groups from the input collection so that primers can be designed for individual groups. Our strategy proposes an automatic grouping of the input sequences by using clustering techniques. Conserved regions are then detected for each individual group. Conserved regions are scored using a BlockSimilarity score, a novel alignment scoring scheme that is appropriate for this application. Degenerate primers are then designed by reverse translating the conserved amino acid sequences to the corresponding nucleotide sequences. Our program, DePiCt, was written in BioPerl and was tested on the Toll-Interleukin Receptor (TIR)and the non-TIR family of plant resistance genes. Existing programs for degenerate primer design were unable to find primers for these data sets.

Algorithms↗

Mining protein sequences for motifs.

We use methods from Data Mining and Knowledge Discovery to design an algorithm for detecting motifs in protein sequences. The algorithm assumes that a motif is constituted by the presence of a "good" combination of residues in appropriate locations of the motif. The algorithm attempts to compile such good combinations into a "pattern dictionary" by processing an aligned training set of protein sequences. The dictionary is subsequently used to detect motifs in new protein sequences. Statistical significance of the detection results are ensured by statistically determining the various parameters of the algorithm. Based on this approach, we have implemented a program called GYM. The Helix-Turn-Helix motif was used as a model system on which to test our program. The program was also extended to detect Homeodomain motifs. The detection results for the two motifs compare favorably with existing programs. In addition, the GYM program provides a lot of useful information about a given protein sequence.

Algorithms↗

Multiple comparisons model-based clustering and ternary pattern tree numerical display of gene response to treatment: procedure and application to the preclinical evaluation of chemopreventive agents.

Microarray technology has greatly aided the identification of genes that are expressed differentially. Statistical analysis of such data by multiple comparisons procedures has been slow to develop, in part, because methods to cluster the results of such comparisons in biologically meaningful ways have not been available. We isolated and analyzed, by Northern blot and GeneChip, replicate liver RNA samples (n = 4/group) from rats fed with control diet or diet containing one of three chemopreventive compounds, selected because their pharmacological activities, including RNA expression response, are relatively well understood. We report on a classification tree, based on the results of nonparametric multiple comparisons, which results in the bipolar hierarchical clustering of genes in relation to their response to treatment. In addition to identifying treatment-responsive genes, application of this procedure to our test study identified the known pharmacological relationships among the treatment groups without supervision. Also, small treatment-specific subsets of genes were identified that may be indicative of additional pharmacophores present in the test compounds.

Animals↗