PubMed Health⌕ Search

Biomedical subjects

Suraj Peri

Publications and source records attributed to Suraj Peri.

10 recordsLinked to original sources

Genome annotation of Anopheles gambiae using mass spectrometry-derived data.

BACKGROUND: A large number of animal and plant genomes have been completely sequenced over the last decade and are now publicly available. Although genomes can be rapidly sequenced, identifying protein-coding genes still remains a problematic task. Availability of protein sequence data allows direct confirmation of protein-coding genes. Mass spectrometry has recently emerged as a powerful tool for proteomic studies. Protein identification using mass spectrometry is usually carried out by searching against databases of known proteins or transcripts. This approach generally does not allow identification of proteins that have not yet been predicted or whose transcripts have not been identified. RESULTS: We searched 3,967 mass spectra from 16 LC-MS/MS runs of Anopheles gambiae salivary gland homogenates against the Anopheles gambiae genome database. This allowed us to validate 23 known transcripts and 50 novel transcripts. In addition, a novel gene was identified on the basis of peptides that matched a genomic region where no gene was known and no transcript had been predicted. The amino termini of proteins encoded by two predicted transcripts were confirmed based on N-terminally acetylated peptides sequenced by tandem mass spectrometry. Finally, six sequence polymorphisms could be annotated based on experimentally obtained peptide sequences. CONCLUSION: The peptide sequences from this study were mapped onto the genomic sequence using the distributed annotation system available at Ensembl and can be visualized in the context of all other existing annotations. The strategy described in this paper can be used to correct and confirm genome annotations and permit discovery of novel proteins in a high-throughput manner by mass spectrometry.

Animals↗

A functional annotation of subproteomes in human plasma.

The data collected by Human Proteome Organization's Plasma Proteome Pilot project phase was analyzed by members of our working group. Accordingly, a functional annotation of the human plasma proteome was carried out. Here, we report the findings of our analyses. First, bioinformatic analyses were undertaken to determine the likely sources of plasma proteins and to develop a protein interaction network of proteins identified in this project. Second, annotation of these proteins was performed in the context of functional subproteomes involved in the coagulation pathway, the mononuclear phagocytic system, the inflammation pathway, the cardiovascular system, and the liver; as well as the subset of proteins associated with DNA binding activities. Our analyses contributed to the Plasma Proteome Database (http://www.plasmaproteomedatabase.org), an annotated database of plasma proteins identified by HPPP as well as from other published studies. In addition, we address several methodological considerations including the selective enrichment of post-translationally modified proteins by the use of multi-lectin chromatography as well as the use of peptidomic techniques to characterize the low molecular weight proteins in plasma. Furthermore, we have performed additional analyses of peptide identification data to annotate cleavage of signal peptides, sites of intra-membrane proteolysis and post-translational modifications. The HPPP-organized, multi-laboratory effort, as described herein, resulted in much synergy and was essential to the success of this project.

Blood Coagulation↗

A bioinformatics analysis of protein tyrosine phosphatases in humans.

Protein tyrosine phosphatases (PTPs) cooperate with protein tyrosine kinases to regulate signal transduction pathways. Genome-wide surveys cataloging protein tyrosine phosphatases in humans have recently been carried out. Here, we present a bioinformatics analysis of protein tyrosine phosphatases in the human genome to examine their domain architecture, alternative splicing and pseudogenes. We present evidence that alternative transcripts exist for 25 out of 35 PTPs analyzed. These alternative transcripts include novel exons; skipped exons as well as cryptic donor/acceptor splice sites. We discovered a novel isoform of PTPN18 based on analysis of expressed sequence tags (ESTs). The deletion of 4 exons in the catalytic domain of the novel isoform may alter the enzymatic activity toward its substrates. We were able to experimentally validate 2 of our novel isoform predictions through RT-PCR. Finally, a user-friendly web-based resource that consolidates the gene and protein annotations for all human protein tyrosine phosphatases has been developed and is freely available at http://ptpr.ibioinformatics.org.

Alternative Splicing↗

BioBuilder as a database development and functional annotation platform for proteins.

BACKGROUND: The explosion in biological information creates the need for databases that are easy to develop, easy to maintain and can be easily manipulated by annotators who are most likely to be biologists. However, deployment of scalable and extensible databases is not an easy task and generally requires substantial expertise in database development. RESULTS: BioBuilder is a Zope-based software tool that was developed to facilitate intuitive creation of protein databases. Protein data can be entered and annotated through web forms along with the flexibility to add customized annotation features to protein entries. A built-in review system permits a global team of scientists to coordinate their annotation efforts. We have already used BioBuilder to develop Human Protein Reference Database http://www.hprd.org, a comprehensive annotated repository of the human proteome. The data can be exported in the extensible markup language (XML) format, which is rapidly becoming as the standard format for data exchange. CONCLUSIONS: As the proteomic data for several organisms begins to accumulate, BioBuilder will prove to be an invaluable platform for functional annotation and development of customizable protein centric databases. BioBuilder is open source and is available under the terms of LGPL.

Computational Biology↗

Human protein reference database as a discovery resource for proteomics.

The rapid pace at which genomic and proteomic data is being generated necessitates the development of tools and resources for managing data that allow integration of information from disparate sources. The Human Protein Reference Database (http://www.hprd.org) is a web-based resource based on open source technologies for protein information about several aspects of human proteins including protein-protein interactions, post-translational modifications, enzyme-substrate relationships and disease associations. This information was derived manually by a critical reading of the published literature by expert biologists and through bioinformatics analyses of the protein sequence. This database will assist in biomedical discoveries by serving as a resource of genomic and proteomic information and providing an integrated view of sequence, structure, function and protein networks in health and disease.

Computational Biology↗

Computational and experimental analysis reveals a novel Src family kinase in the C. elegans genome.

MOTIVATION: The complete genomes of a number of organisms have already been sequenced. However, the vast majority of annotated genes are derived by gene prediction methods. It is important to not only validate the predicted coding regions but also to identify genes that may have been missed by these programs. METHODS: We searched the entire C.elegans genomic sequence database maintained by the Sanger Center using human c-Src sequence in a TBLASN search. We have confirmed one of the predicted regions by isolation of a cDNA and carried out a phylogenetic analysis of Src kinase family members in the worm, fly and several vertebrate species. RESULTS: Our analysis identified a novel tyrosine kinase in the C.elegans genome that contains functional features typical of the Src family kinases that we have designated as Src-1. The open reading frame contains a conserved N-terminal myristoylation site and a tyrosine residue within the C-terminus that is crucial for regulating the activity of Src kinases. Our phylogenetic analysis of Src family members from C. elegans, Drosophila and other higher organisms revealed a relationship among Src kinases from C. elegans and Drosophila.

Amino Acid Sequence↗

From biological databases to platforms for biomedical discovery.

The use of high-throughput DNA sequencing and proteomic methods has led to an unprecedented increase in the amount of genomic and proteomic data. Application of computing technologies and development of computational tools to analyze and present these data has not kept pace with the accumulation of information. Here, we discuss the use of different database systems to store biological information and mention some of the key emerging computing technologies that are likely to have a key role in the future of bioinformatics.

Algorithms↗

Development of human protein reference database as an initial platform for approaching systems biology in humans.

Human Protein Reference Database (HPRD) is an object database that integrates a wealth of information relevant to the function of human proteins in health and disease. Data pertaining to thousands of protein-protein interactions, posttranslational modifications, enzyme/substrate relationships, disease associations, tissue expression, and subcellular localization were extracted from the literature for a nonredundant set of 2750 human proteins. Almost all the information was obtained manually by biologists who read and interpreted >300,000 published articles during the annotation process. This database, which has an intuitive query interface allowing easy access to all the features of proteins, was built by using open source technologies and will be freely available at http://www.hprd.org to the academic community. This unified bioinformatics platform will be useful in cataloging and mining the large number of proteomic interactions and alterations that will be discovered in the postgenomic era.

BRCA1 Protein↗

Toll and interleukin-1 receptor (TIR) domain-containing proteins in plants: a genomic perspective.

Toll and interleukin-1 receptor (TIR) domains were originally described from comparisons of proteins found in mammals and Drosophila. They are now known to occur in several organisms, with the most TIR proteins being found in Arabidopsis: our analysis of the sequenced Arabidopsis genome has revealed the presence of at least 135 proteins containing TIR domains. Several novel types of TIR-domain-containing proteins are found in Arabidopsis that are not found in other genomes. Here, we discuss the roles of TIR-domain-containing proteins in pathogen resistance and as candidate signaling modules.

Animals↗