PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Molecular Sequence Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

An atlas of differential gene expression during early Xenopus embryogenesis.

We have carried out a large-scale, semi-automated whole-mount in situ hybridization screen of 8369 cDNA clones in Xenopus laevis embryos. We confirm that differential gene expression is prevalent during embryogenesis since 24% of the clones are expressed non-ubiquitously and 8% are organ or cell type specific marker genes. Sequence analysis and clustering yielded 723 unique genes displaying a differential expression pattern. Of these, 18% were already described in Xenopus, 47% have homologs and 35% are lacking significant sequence similarity in databases. Many of them encode known developmental regulators. We classified 363 of the 723 genes for which a Gene Ontology annotation for molecular function could be attributed and found 'DNA binding' and 'enzyme' the most represented terms. The most common protein domains encoded in these embryonic, differentially expressed genes are the homeobox and RNA Recognition Motif (RRM). Fifty-nine putative orthologs of human disease genes, and 254 organ or cell specific marker genes were identified. Markers were found for nasal placode and archenteron roof, organs for which a specific marker was previously unavailable. Markers were also found for novel subdomains of various other organs. The tissues for which most markers were found are muscle and epidermis. Expression of cell cycle regulators fell in two classes, containing proliferation-promoting and anti-proliferative genes, respectively. We identified 66 new members of the BMP4, chromatin, endoplasmic reticulum, and karyopherin synexpression groups, thus providing a first glimpse of their probable cellular roles. Cluster analysis of tissues to measure tissue relatedness yielded some unorthodox affinities besides expectable lineage relationships. In conclusion, this study represents an atlas of gene expression patterns, which reveals embryonic regionalization, provides novel marker genes, and makes predictions about the functional role of unknown genes.

Animals↗

The first archaeal ATP-dependent glucokinase, from the hyperthermophilic crenarchaeon Aeropyrum pernix, represents a monomeric, extremely thermophilic ROK glucokinase with broad hexose specificity.

An ATP-dependent glucokinase of the hyperthermophilic aerobic crenarchaeon Aeropyrum pernix was purified 230-fold to homogeneity. The enzyme is a monomeric protein with an apparent molecular mass of about 36 kDa. The apparent K(m) values for ATP and glucose (at 90 degrees C and pH 6.2) were 0.42 and 0.044 mM, respectively; the apparent V(max) was about 35 U/mg. The enzyme was specific for ATP as a phosphoryl donor, but showed a broad spectrum for phosphoryl acceptors: in addition to glucose, which showed the highest catalytic efficiency (k(cat)/K(m)), the enzyme also phosphorylates glucosamin, fructose, mannose, and 2-deoxyglucose. Divalent cations were required for maximal activity: Mg(2+), which was most effective, could partially be replaced with Co(2+), Mn(2+), and Ni(2+). The enzyme had a temperature optimum of at least 100 degrees C and showed significant thermostability up to 100 degrees C. The coding function of open reading frame (ORF) APE2091 (Y. Kawarabayasi, Y. Hino, H. Horikawa, S. Yamazaki, Y. Haikawa, K. Jin-no, M. Takahashi, M. Sekine, S. Baba, A. Ankai, H. Kosugi, A. Hosoyama, S. Fukui, Y. Nagai, K. Nishijima, H. Nakazawa, M. Takamiya, S. Masuda, T. Funahashi, T. Tanaka, Y. Kudoh, J. Yamazaki, N. Kushida, A. Oguchi, and H. Kikuchi, DNA Res. 6:83-101, 145-152, 1999), previously annotated as gene glk, coding for ATP-glucokinase of A. pernix, was proved by functional expression in Escherichia coli. The purified recombinant ATP-dependent glucokinase showed a 5-kDa higher molecular mass on sodium dodecyl sulfate-polyacrylamide gel electrophoresis, but almost identical kinetic and thermostability properties in comparison to the native enzyme purified from A. pernix. N-terminal amino acid sequence of the native enzyme revealed that the translation start codon is a GTG 171 bp downstream of the annotated start codon of ORF APE2091. The amino acid sequence deduced from the truncated ORF APE2091 revealed sequence similarity to members of the ROK family, which comprise bacterial sugar kinases and transcriptional repressors. This is the first report of the characterization of an ATP-dependent glucokinase from the domain of Archaea, which differs from its bacterial counterparts by its monomeric structure and its broad specificity for hexoses.

Adenosine Triphosphate↗

Molecular Biocomputing Suite: a word processor add-in for the analysis and manipulation of nucleic acid and protein sequence data.

In all fields of molecular biology, researchers are increasingly challenged by experiments planned and evaluated on the basis of nucleic acid and protein sequence data generally retrieved from public databases. Despite the wide spectrum of available Web-based software tools for sequence analysis, the routine use of these tools has disadvantages, particularly because of the elaborate and heterogeneous ways of data input, output, and storage. Here we present a Visual Basic-encoded Microsoft Word Add-In, the Molecular BioComputing Suite (MBCS), available at the BioTechniques Software Library (www.BioTechniques.com). The MBCS software aims to manage and expedite a wide range of sequence analyses and manipulations using an integrated text editor environment including menu-guided commands. Its independence of sequence formats enables MBCS to be used as a pivotal application between other software tools for sequence analysis, manipulation, annotation, and editing.

Amino Acid Sequence↗

A study of the middle-scale nucleotide clustering in DNA sequences of various origin and functionality, by means of a method based on a modified standard deviation.

The deviation from randomness in the distribution of nucleotides in genomic sequences is quantified and studied, using a modified standard deviation (MSD). This method implies a "per block" computation of the standard deviation of the nucleotide frequencies of occurrence, using local means (means taken in a neighborhood of each block). This quantity may serve as a scale-dependent measure of the nucleotide clustering. In the present work, the meso-scale of tenths of nucleotides is principally explored, by means of suitably adjusted filter parameters. This length scale is of an order of magnitude not directly affected by the grammar and syntax rules of the protein-coding procedure, remaining shorter than the scale of appearance of large-scale characteristics of the genome. MSD has been found to distinguish systematically between the sequences of different origin and functionality. The most near-random are found to be coding sequences of prokaryotes, while in intronic and intergenic regions of eukaryotic genomes, extended clustering of similar nucleotides is observed. The distributions of MSD values of large collections of sequences are found to be in most cases characteristic of their biological role and origin. Protein- and non-coding, prokaryotic and eukaryotic DNA as well as promoter, rRNA, viral and organelle sequences have been examined. The presented results corroborate a recently proposed model for genome evolution. The method is also applied for an assessment of the annotation of ORFs taken from the complete genome of Saccharomyces cerevisiae.

Animals↗

JAFA: a protein function annotation meta-server.

With the high number of sequences and structures streaming in from genomic projects, there is a need for more powerful and sophisticated annotation tools. Most problematic of the annotation efforts is predicting gene and protein function. Over the past few years there has been considerable progress in automated protein function prediction, using a diverse set of methods. Nevertheless, no single method reports all the information possible, and molecular biologists resort to 'shopping around' using different methods: a cumbersome and time-consuming practice. Here we present the Joined Assembly of Function Annotations, or JAFA server. JAFA queries several function prediction servers with a protein sequence and assembles the returned predictions in a legible, non-redundant format. In this manner, JAFA combines the predictions of several servers to provide a comprehensive view of what are the predicted functions of the proteins. JAFA also offers its own output, and the individual programs' predictions for further processing. JAFA is available for use from http://jafa.burnham.org.

Internet↗

Differentially expressed genes in the Trypanosoma brucei life cycle identified by RNA fingerprinting.

RNA fingerprinting by arbitrarily primed polymerase chain reaction (RAP-PCR) was used to identify genes that were differentially expressed during the life cycle of Trypanosoma brucei, as well as in response to heat shock. The standard RAP-PCR protocol was varied in two ways. First, the PCR reactions sometimes included a primer derived from the 5' mini-exon sequence, to ensure that most of the products contained the 5' end of mRNAs. Second, differentially amplified products were reamplified, isolated on single strand conformation polymorphism (SSCP) gels, cloned, and sequenced. Clones representing 32 different expressed sequence tags (ESTs) were obtained. Twenty-four ESTs were confirmed as differentially expressed by RT-PCR between different stages of the parasite cycle, or in response to temperature elevation. Nine clones had significant similarities to sequences already in the database. These transcripts included genes encoding cell surface proteins, metabolic enzymes, and heat shock proteins, either from trypanosomes or other organisms. Of particular interest, ESAG1 was shown to be heat-inducible in the procyclic stage. Most of the transcripts were unrelated to any other sequences in the database, and were deposited as new ESTs. The identification of stage-specific and heat shock-regulated transcripts will complement the growing T. brucei database. In addition, this experimental approach allows previous entries in the sequence database to be annotated with regulatory information.

Animals↗

Enzyme function less conserved than anticipated.

The level of sequence similarity that implies similarity in protein structure is well established. Recently, many groups proposed thresholds for similarity in sequence implying similarity in enzymatic function. All previous results suggest the strong conservation of enzymatic function above levels of 50% pairwise sequence identity. Here, I argue that all groups substantially overestimated the conservation of enzyme function because their data sets were either too biased, or too small. An unbiased analysis suggested that less than 30% of the pair fragments above 50% sequence identity have entirely identical EC numbers. Another surprising finding was that even BLAST E-values below 10(-50) did not suffice to automatically transfer enzyme function without errors. As expected, most misclassifications originated from similarities in relatively short regions and/or from transferring annotations for different domains. Both problems cannot be corrected easily by adjusting the thresholds for automatic transfer of genome annotations. A score relating sequence identity to alignment length (distance from HSSP-threshold) outperformed statistical BLAST scores for high sequence similarity. In particular, the distance score allowed error-free transfer of enzyme function for the 10% most similar enzyme pairs. The results illustrated how difficult it is to assess the conservation of protein function and to guarantee error-free genome annotations, in general: sets with millions of pair comparisons might not suffice to arrive at statistically significant conclusions. In practice, the revised detailed estimates for the sequence conservation of enzyme function may provide important benchmarks for everyday sequence analysis and for more cautious automatic genome annotations.

Amino Acid Sequence↗

Protein sequence analysis in silico: application of structure-based bioinformatics to genomic initiatives.

The current pace of high-throughput genome sequencing programs coupled with high-throughput functional genomic screens has provided researchers with a bewildering array of sequence and biological data to contend with. Identification of proteins of interest from a particular biological study requires the application of bioinformatic tools to process and prioritise the data. From a protein function standpoint, transfer of annotation from known proteins to a novel target is currently the only practical way to convert vast quantities of raw sequence data into meaningful information. New bioinformatics tools now provide more sophisticated methods to transfer functional annotation, integrating sequence, family profile and structural search methodology. The importance of these approaches to medical research is increasing as we move to annotate the proteome through functional and structural genomic efforts.

Animals↗

Microarray-based analysis of early development in Xenopus laevis.

In order to examine transcriptional regulation globally, during early vertebrate embryonic development, we have prepared Xenopus laevis cDNA microarrays. These prototype embryonic arrays contain 864 sequenced gastrula cDNA. In order to analyze and store array data, a microarray analysis pipeline was developed and integrated with sequence analysis and annotation tools. In three independent experimental settings, we demonstrate the power of these global approaches and provide optimized protocols for their application to molecular embryology. In the first set, by comparing maternal versus zygotic transcription, we document groups of genes that are temporally regulated. This analytical approach resulted in the discovery of novel temporally regulated genes. In the second, we examine changes in gene expression spatially during development by comparing dorsal and ventral mesoderm dissected from early gastrula embryos. We have discovered novel genes with spatial enrichment from these experiments. Finally, we use the prototype microarray to examine transcriptional responses from embryonic explants treated with activin. We selected genes (two of which are novel) regulated by activin for further characterization. All results obtained by the arrays were independently tested by RT-PCR or by in situ hybridization to provide a direct assessment of the accuracy and reproducibility of these approaches in the context of molecular embryology.

Animals↗

Planetary biology--paleontological, geological, and molecular histories of life.

The history of life on Earth is chronicled in the geological strata, the fossil record, and the genomes of contemporary organisms. When examined together, these records help identify metabolic and regulatory pathways, annotate protein sequences, and identify animal models to develop new drugs, among other features of scientific and biomedical interest. Together, planetary analysis of genome and proteome databases is providing an enhanced understanding of how life interacts with the biosphere and adapts to global change.

Amino Acid Sequence↗

AVID: an integrative framework for discovering functional relationships among proteins.

BACKGROUND: Determining the functions of uncharacterized proteins is one of the most pressing problems in the post-genomic era. Large scale protein-protein interaction assays, global mRNA expression analyses and systematic protein localization studies provide experimental information that can be used for this purpose. The data from such experiments contain many false positives and false negatives, but can be processed using computational methods to provide reliable information about protein-protein relationships and protein function. An outstanding and important goal is to predict detailed functional annotation for all uncharacterized proteins that is reliable enough to effectively guide experiments. RESULTS: We present AVID, a computational method that uses a multi-stage learning framework to integrate experimental results with sequence information, generating networks reflecting functional similarities among proteins. We illustrate use of the networks by making predictions of detailed Gene Ontology (GO) annotations in three categories: molecular function, biological process, and cellular component. Applied to the yeast Saccharomyces cerevisiae, AVID provides 37,451 pair-wise functional linkages between 4,191 proteins. These relationships are approximately 65-78% accurate, as assessed by cross-validation testing. Assignments of highly detailed functional descriptors to proteins, based on the networks, are estimated to be approximately 67% accurate for GO categories describing molecular function and cellular component and approximately 52% accurate for terms describing biological process. The predictions cover 1,490 proteins with no previous annotation in GO and also assign more detailed functions to many proteins annotated only with less descriptive terms. Predictions made by AVID are largely distinct from those made by other methods. Out of 37,451 predicted pair-wise relationships, the greatest number shared in common with another method is 3,413. CONCLUSION: AVID provides three networks reflecting functional associations among proteins. We use these networks to generate new, highly detailed functional predictions for roughly half of the yeast proteome that are reliable enough to drive targeted experimental investigations. The predictions suggest many specific, testable hypotheses. All of the data are available as downloadable files as well as through an interactive website at http://web.mit.edu/biology/keating/AVID. Thus, AVID will be a valuable resource for experimental biologists.

Algorithms↗

A novel beta-glucanase gene from Bacillus halodurans C-125.

A novel endo-beta-1,3(4)-D-glucanase gene was found in the complete genome sequence of Bacillus halodurans C-125. The gene was previously annotated as an "unknown" protein and assigned an incorrect open reading frame (ORF). However, determining the biochemical characteristics has elucidated the function and correct ORF of the gene. The gene encodes 231 amino acids, and its calculated molecular mass was estimated to be 26743.16 Da. The amino acid sequence alignment showed that the highest sequence identity was only 28% with that of the beta-1,3-1,4-glucanase from Bacillus subtilis. Moreover, the nucleotide sequence did not match any other known Bacillus beta-glucanase gene. The member of the gene cluster that includes this novel gene was apparently different from that of the gene cluster including the putative beta-glucanase genes (bh3231 and bh3232) from B. halodurans C-125. Therefore, the novel gene is not a copy of either of these genes, and in B. halodurans cells, the putative role of the encoded protein may differ from that of bh3231 and bh3232. To examine the activity of the gene product, the gene was cloned as a His-tagged protein and expressed in Escherichia coli. The purified enzyme showed activity against lichenan, barley beta-glucan, laminarin, and carboxymethyl curdlan. Thin-layer chromatography showed that the enzyme hydrolyzes substrates in an endo-type manner. When beta-glucan was used as a substrate, the pH optimum was between 6 and 8, and the temperature optimum was 60 degrees C. After 2 h incubation at 50 and 60 degrees C, the residual activity remained 100% and 50%, respectively. The enzymatic activity was abolished after 30 min incubation at 70 degrees C. Based on the results, the gene encodes an endo-type beta-1,3(4)-D-glucanase (E.C. 3.2.1.6).

Amino Acid Sequence↗

Analysis of 10,000 ESTs from lymphocytes of the cynomolgus monkey to improve our understanding of its immune system.

BACKGROUND: The cynomolgus monkey (Macaca fascicularis) is one of the most widely used surrogate animal models for an increasing number of human diseases and vaccines, especially immune-system-related ones. Towards a better understanding of the gene expression background upon its immunogenetics, we constructed a cDNA library from Epstein-Barr virus (EBV)-transformed B lymphocytes of a cynomolgus monkey and sequenced 10,000 randomly picked clones. RESULTS: After processing, 8,312 high-quality expressed sequence tags (ESTs) were generated and assembled into 3,728 unigenes. Annotations of these uniquely expressed transcripts demonstrated that out of the 2,524 open reading frame (ORF) positive unigenes (mitochondrial and ribosomal sequences were not included), 98.8% shared significant similarities (E-value less than 1e-10) with the NCBI nucleotide (nt) database, while only 67.7% (E-value less than 1e-5) did so with the NCBI non-redundant protein (nr) database. Further analysis revealed that 90.0% of the unigenes that shared no similarities to the nr database could be assigned to human chromosomes, in which 75 did not match significantly to any cynomolgus monkey and human ESTs. The mapping regions to known human genes on the human genome were described in detail. The protein family and domain analysis revealed that the first, second and fourth of the most abundantly expressed protein families were all assigned to immunoglobulin and major histocompatibility complex (MHC)-related proteins. The expression profiles of these genes were compared with that of homologous genes in human blood, lymph nodes and a RAMOS cell line, which demonstrated expression changes after transformation with EBV. The degree of sequence similarity of the MHC class I and II genes to the human reference sequences was evaluated. The results indicated that class I molecules showed weak amino acid identities (<90%), while class II showed slightly higher ones. CONCLUSION: These results indicated that the genes expressed in the cynomolgus monkey could be used to identify novel protein-coding genes and revise those incomplete or incorrect annotations in the human genome by comparative methods, since the old world monkeys and humans share high similarities at the molecular level, especially within coding regions. The identification of multiple genes involved in the immune response, their sequence variations to the human homologues, and their responses to EBV infection could provide useful information to improve our understanding of the cynomolgus monkey immune system.

5' Untranslated Regions↗

BAliBASE 3.0: latest developments of the multiple sequence alignment benchmark.

Multiple sequence alignment is one of the cornerstones of modern molecular biology. It is used to identify conserved motifs, to determine protein domains, in 2D/3D structure prediction by homology and in evolutionary studies. Recently, high-throughput technologies such as genome sequencing and structural proteomics have lead to an explosion in the amount of sequence and structure information available. In response, several new multiple alignment methods have been developed that improve both the efficiency and the quality of protein alignments. Consequently, the benchmarks used to evaluate and compare these methods must also evolve. We present here the latest release of the most widely used multiple alignment benchmark, BAliBASE, which provides high quality, manually refined, reference alignments based on 3D structural superpositions. Version 3.0 of BAliBASE includes new, more challenging test cases, representing the real problems encountered when aligning large sets of complex sequences. Using a novel, semiautomatic update protocol, the number of protein families in the benchmark has been increased and representative test cases are now available that cover most of the protein fold space. The total number of proteins in BAliBASE has also been significantly increased from 1444 to 6255 sequences. In addition, full-length sequences are now provided for all test cases, which represent difficult cases for both global and local alignment programs. Finally, the BAliBASE Web site (http://www-bio3d-igbmc.u-strasbg.fr/balibase) has been completely redesigned to provide a more user-friendly, interactive interface for the visualization of the BAliBASE reference alignments and the associated annotations.

Amino Acid Sequence↗

Identification and characterization of two dipeptidyl-peptidase III isoforms in Drosophila melanogaster.

Dipeptidyl-peptidase III (DPP III) hydrolyses small peptides with a broad substrate specificity. It is thought to be involved in a major degradation pathway of the insect neuropeptide proctolin. We report the purification and characterization of a soluble DPP III from 40 g Drosophila melanogaster. Western blot analysis with anti-(DPP III) serum revealed the purification of two proteins of molecular mass 89 and 82 kDa. MS/MS analysis of these proteins resulted in the sequencing of 45 and 41 peptide fragments, respectively, confirming approximately 60% of both annotated D. melanogaster DPP III isoforms (CG7415-PC and CG7415-PB) predicted at 89 and 82 kDa. Sequencing also revealed the specific catalytic domain HELLGH in both isoforms, indicating that they are both effective in degrading small peptides. In addition, with a probe specific for D. melanogaster DPP III, northern blot analysis of fruit fly total RNA showed two transcripts at approximately 2.6 and 2.3 kb, consistent with the translation of 89-kDa and 82-kDa DPP III proteins. Moreover, the purified enzyme hydrolyzed the insect neuropeptide proctolin (Km approximately 4 microm) at the second N-terminal peptide bound, and was inhibited by the specific DPP III inhibitor tynorphin. Finally, anti-(DPP III) immunoreactivity was observed in the central nervous system of D. melanogaster larva, supporting a functional role for DPP III in proctolin degradation. This study shows that DPP III is in actuality synthesized in D. melanogaster as 89-kDa and 82-kDa isoforms, representing two native proteins translated from two alternative mRNA transcripts.

Animals↗

The maize viviparous15 locus encodes the molybdopterin synthase small subunit.

A new Zea mays viviparous seed mutant, viviparous15 (vp15), was isolated from the UniformMu transposon-tagging population. In addition to precocious germination, vp15 has an early seedling lethal phenotype. Biochemical analysis showed reduced activities of several enzymes that require molybdenum cofactor (MoCo) in vp15 mutant seedlings. Because MoCo is required for abscisic acid (ABA) biosynthesis, the viviparous phenotype is probably caused by ABA deficiency. We cloned the vp15 mutant using a novel high-throughput strategy for analysis of high-copy Mu lines: We used MuTAIL PCR to extract genomic sequences flanking the Mu transposons in the vp15 line. The Mu insertions specific to the vp15 line were identified by in silico subtraction using a database of MuTAIL sequences from 90 UniformMu lines. Annotation of the vp15-specific sequences revealed a Mu insertion in a gene homologous to human MOCS2A, the small subunit of molybdopterin (MPT) synthase. Molecular analysis of two allelic mutations confirmed that Vp15 encodes a plant MPT synthase small subunit (ZmCNX7). Our results, and a related paper reporting the cloning of maize viviparous10, demonstrate robust cloning strategies based on MuTAIL-PCR. The Vp15/CNX7, together with other CNX genes, is expressed in both embryo and endosperm during seed maturation. Expression of Vp15 appears to be regulated independently of MoCo biosynthesis. Comparisons of Vp15 loci in genomes of three cereals and Arabidopsis thaliana identified a conserved sequence element in the 5' untranslated region as well as a micro-synteny among the cereals.

Alleles↗

Functional analysis of a novel nonsense PPP1R12A variant in a Chinese family with infantile epilepsy.

BACKGROUND: Defects in PPP1R12A can lead to genitourinary and/or brain malformation syndrome (GUBS). GUBS is primarily characterized by neurological or genitourinary system abnormalities, but a few reported cases are associated with neonatal seizures. Here, we report a case of a female newborn with neonatal seizures caused by a novel variant in PPP1R12A, aiming to enhance the clinical and variant data of genetic factors related to epilepsy in early life. METHODS: Whole-exome and Sanger sequencing were used for familial variant assessment, and bioinformatics was employed to annotate the variant. A structural model of the mutant protein was simulated using molecular dynamics (MD), and the free binding energy between PPP1R12A and PPP1CB was analyzed. A mutant plasmid was constructed, and mutant protein expression was analyzed using western blotting (WB), and the interaction between the mutant and PPP1CB proteins using co-immunoprecipitation (Co-IP) experiments. RESULTS: The patient experienced tonic-clonic seizures on the second day after birth. Genetic testing revealed a heterozygous variant in PPP1R12A, NM_002480.3:c.2533&#xa0;C&#x2009;>&#x2009;T (p.Arg845Ter). Both parents had the wild-type gene. MD suggested that loss of the C-terminal structure in the mutant protein altered its structural stability and increased the binding energy with PPP1CB, indicating unstable protein-protein interactions. On WB, a low-molecular-weight band was observed, indicating that the protein was truncated. Co-IP indicated that the mutant protein no longer interacted with PPP1CB, indicating an effect on the structural stability of the myosin phase complex. CONCLUSION: The PPP1R12A c.2533&#xa0;C&#x2009;>&#x2009;T variant may explain the neonatal seizures in the present case. The findings of this study expand the spectrum of PPP1R12A variants and highlight the potential significance of truncated proteins in the pathogenesis of GUBS.

Female↗

ELISA: a unified, multidimensional view of the protein domain universe.

ELISA (http://romi.bu.edu/elisa/) is a database that was designed for flexibility in defining interesting queries about protein domain evolution. We have defined and included both the inherent characteristics of the domains such as structure and function and comparisons of these characteristics between domains. Thus, the database is useful in defining structural and functional links between related protein domains and by extension sequences that encode them. In this database we introduce and employ a novel method of functional annotation and comparison. For each protein domain we create a probabilistic functional annotation tree using GO. We have designed an algorithm that accurately compares these trees and thus provides a measure of "functional distance" between two protein domains. Along with functional annotation, we have also included structural comparison between protein domains and best sequence comparisons to all known genomes. The latter enables researchers to dynamically do searches for domains sharing similar phylogenetic profiles. This combination of data and tools enables the researcher to design complex queries to carry out research in the areas of protein domain evolution, structure prediction and functional annotation of novel sequences.

Amino Acid Sequence↗