PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 505 records · Page 28Linked to original sources

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗

Clustering of domains of functionally related enzymes in the interaction database PRECISE by the generation of primary sequence patterns.

The PRECISE database was developed by our laboratory to allow for the systematic study of the ligand interactions common to a set of functionally related enzymes, where an interaction site is defined broadly as any residue(s) that interact with a ligand. During the construction of PRECISE, enzyme chains are extracted from the protein data bank (PDB) and clustered according to functional homology as defined by the enzyme commission (EC) nomenclature system. A sequence representative is chosen from each cluster based on the criterion set forth by the non-redundant PDB set, and pair-wise alignments of each cluster member to the representative are performed. Atom-based residue-ligand interactions are calculated for each cluster member, and the summation of ligand interactions for all cluster members at each aligned position is determined. Although we were able to successfully align most clusters using a simple dynamic programming algorithm, several cluster created exhibited poor pair-wise alignments of each cluster member to its sequence representative. We hypothesized that the observed alignment problems were, in most cases, due to the incorrect separation and alignment of different domains in multi-domain proteins, a mistake that frequently causes error proliferation in functional annotation. Here we present the results of generating primary sequence patterns for each poorly aligned cluster in PRECISE to assess the extent to which multi-domain proteins that are incorrectly aligned contributes to poor pair-wise alignments of each cluster member to its representative. This requires the use of an iterative locally optimal pair-wise alignment algorithm to build a hierarchical similarity-based sequence pattern for a set of functionally related enzymes. Our results show that poor alignments in PRECISE are caused most frequently by the misalignment of multi-domain proteins, and that the generation of primary sequence patterns for the assignment of sequence family membership yields better alignments for the functionally related enzyme clusters in PRECISE than our original alignment algorithm.

Amino Acid Sequence↗

Beyond synexpression relationships: local clustering of time-shifted and inverted gene expression profiles identifies new, biologically relevant interactions.

The complexity of biological systems provides for a great diversity of relationships between genes. The current analysis of whole-genome expression data focuses on relationships based on global correlation over a whole time-course, identifying clusters of genes whose expression levels simultaneously rise and fall. There are, of course, other potential relationships between genes, which are missed by such global clustering. These include activation, where one expects a time-delay between related expression profiles, and inhibition, where one expects an inverted relationship. Here, we propose a new method, which we call local clustering, for identifying these time-delayed and inverted relationships. It is related to conventional gene-expression clustering in a fashion analogous to the way local sequence alignment (the Smith-Waterman algorithm) is derived from global alignment (Needleman-Wunsch). An integral part of our method is the use of random score distributions to assess the statistical significance of each cluster. We applied our method to the yeast cell-cycle expression dataset and were able to detect a considerable number of additional biological relationships between genes, beyond those resulting from conventional correlation. We related these new relationships between genes to their similarity in function (as determined from the MIPS scheme) or their having known protein-protein interactions (as determined from the large-scale two-hybrid experiment); we found that genes strongly related by local clustering were considerably more likely than random to have a known interaction or a similar cellular role. This suggests that local clustering may be useful in functional annotation of uncharacterized genes. We examined many of the new relationships in detail. Some of them were already well-documented examples of inhibition or activation, which provide corroboration for our results. For instance, we found an inverted expression profile relationship between genes YME1 and YNT20, where the latter has been experimentally documented as a bypass suppressor of the former. We also found new relationships involving uncharacterized yeast genes and were able to suggest functions for many of them. In particular, we found a time-delayed expression relationship between J0544 (which has not yet been functionally characterized) and four genes associated with the mitochondria. This suggests that J0544 may be involved in the control or activation of mitochondrial genes. We have also looked at other, less extensive datasets than the yeast cell-cycle and found further interesting relationships. Our clustering program and a detailed website of clustering results is available at http://www.bioinfo.mbb.yale.edu/expression/cluster (or http://www.genecensus.org/expression/cluster).

Algorithms↗

Ab initio protein structure prediction.

Steady progress has been made in the field of ab initio protein folding. A variety of methods now allow the prediction of low-resolution structures of small proteins or protein fragments up to approximately 100 amino acid residues in length. Such low-resolution structures may be sufficient for the functional annotation of protein sequences on a genome-wide scale. Although no consistently reliable algorithm is currently available, the essential challenges to developing a general theory or approach to protein structure prediction are better understood. The energy landscapes resulting from the structure prediction algorithms are only partially funneled to the native state of the protein. This review focuses on two areas of recent advances in ab initio structure prediction-improvements in the energy functions and strategies to search the caldera region of the energy landscapes.

Chemistry, Physical↗

ECgene: an alternative splicing database update.

ECgene (http://genome.ewha.ac.kr/ECgene) was developed to provide functional annotation for alternatively spliced genes. The applications encompass the genome-based transcript modeling for alternative splicing (AS), domain analysis with Gene Ontology (GO) annotation and expression analysis based on the EST and SAGE data. We have expanded the ECgene's AS modeling and EST clustering to nine organisms for which sufficient EST data are available in the GenBank. As for the human genome, we have also introduced several new applications to analyze differential expression. ECprofiler is an ontology-based candidate gene search system that allows users to select an arbitrary combination of gene expression pattern and GO functional categories. DEGEST is a database of differentially expressed genes and isoforms based on the EST information. Importantly, gene expression is analyzed at three distinctive levels-gene, isoform and exon levels. The user interfaces for functional and expression analyses have been substantially improved. ASviewer is a dedicated java application that visualizes the transcript structure and functional features of alternatively spliced variants. The SAGE part of the expression module provides many additional features including SNP, differential expression and alternative tag positions.

Alternative Splicing↗

ASPIC: a web resource for alternative splicing prediction and transcript isoforms characterization.

Alternative splicing (AS) is now emerging as a major mechanism contributing to the expansion of the transcriptome and proteome complexity of multicellular organisms. The fact that a single gene locus may give rise to multiple mRNAs and protein isoforms, showing both major and subtle structural variations, is an exceptionally versatile tool in the optimization of the coding capacity of the eukaryotic genome. The huge and continuously increasing number of genome and transcript sequences provides an essential information source for the computational detection of genes AS pattern. However, much of this information is not optimally or comprehensively used in gene annotation by current genome annotation pipelines. We present here a web resource implementing the ASPIC algorithm which we developed previously for the investigation of AS of user submitted genes, based on comparative analysis of available transcript and genome data from a variety of species. The ASPIC web resource provides graphical and tabular views of the splicing patterns of all full-length mRNA isoforms compatible with the detected splice sites of genes under investigation as well as relevant structural and functional annotation. The ASPIC web resource-available at http://www.caspur.it/ASPIC/--is dynamically interconnected with the Ensembl and Unigene databases and also implements an upload facility.

Algorithms↗

Yeast genomic expression studies using DNA microarrays.

The exploration and characterization of yeast genomic expression programs is providing a wealth of information about yeast biology, as well as other organisms. The intriguing biology of yeast species invites characterization of genomic expression patterns to illuminate the details of cellular physiology. In addition to its value as an interesting organism, yeast maintains its role as an excellent model in which to characterize genomic expression programs. Microarray studies are quickly spreading to plant, animal, and microbial organisms that remain in the early stages of characterization. The extensive knowledge of yeast biology, as well as the relative ease with which yeast studies can be performed and controlled, facilitates interpretation of the genomic expression data. Importantly, existing information about yeast biology, including functional annotations for each gene, is captured and efficiently presented in databases such as the Saccharomyces Genome Database (SGD), the Munich Information Center Yeast Genome Database (MIPS), the Yeast and Pombe Protein Databases (YPD and PPD, respectively), and others. A number of databases also allow the exploration of published genomic expression studies, including the "Expression Connection" at SGD and the Microarray Global Viewer (yMGV) organized by Marc et al. Consulting these databases to retrieve known details about gene function and regulation vastly facilitates interpretation of the genomic expression data, allowing biological hypotheses to be formulated and tested. These hypotheses can be applied to other organisms that may execute genomic expression programs similar to those seen in yeast. Furthermore, as more genomic expression studies in multiple organisms emerge, large-scale data comparisons can be conducted, within and across organisms. Incorporating the results of yeast studies into such comparisons is certain to increase our understanding about the function, regulation, and evolution of genomic expression programs.

Carbocyanines↗

Identification and validation of novel ERBB2 (HER2, NEU) targets including genes involved in angiogenesis.

V-erb-b2 erythroblastic leukemia viral oncogene homolog 2 (ERBB2; synonyms HER2, NEU) encodes a transmembrane glycoprotein with tyrosine kinase-specific activity that acts as a major switch in different signal-transduction processes. ERBB2 amplification and overexpression have been found in a number of human cancers, including breast, ovary and kidney carcinoma. Our aim was to detect ERBB2-regulated target genes that contribute to its tumorigenic effect on a genomewide scale. The differential gene expression profile of ERBB2-transfected and wild-type mouse fibroblasts was monitored employing DNA microarrays. Regulated expression of selected genes was verified by RT-PCR and validated by Western blot analysis. Genome wide gene expression profiling identified (i) known targets of ERBB2 signaling, (ii) genes implicated in tumorigenesis but so far not associated with ERBB2 signaling as well as (iii) genes not yet associated with oncogenic transformation, including novel genes without functional annotation. We also found that at least a fraction of coexpressed genes are closely linked on the genome. ERBB2 overexpression suppresses the transcription of antiangiogenic factors (e.g., Sparc, Timp3, Serpinf1) but induces expression of angiogenic factors (e.g., Klf5, Tnfaip2, Sema3c). Profiling of ERBB2-dependent gene regulation revealed a compendium of potential diagnostic markers and putative therapeutic targets. Identification of coexpressed genes that colocalize in the genome may indicate gene regulatory mechanisms that require further study to evaluate functional coregulation. (Supplementary material for this article can be found on the International Journal of Cancer website at http://www.interscience.wiley.com/jpages/0020-7136/suppmat/index.html.)

Animals↗

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins.

Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.

Computational Biology↗

Protein complex compositions predicted by structural similarity.

Proteins function through interactions with other molecules. Thus, the network of physical interactions among proteins is of great interest to both experimental and computational biologists. Here we present structure-based predictions of 3387 binary and 1234 higher order protein complexes in Saccharomyces cerevisiae involving 924 and 195 proteins, respectively. To generate candidate complexes, comparative models of individual proteins were built and combined together using complexes of known structure as templates. These candidate complexes were then assessed using a statistical potential, derived from binary domain interfaces in PIBASE (http://salilab.org/pibase). The statistical potential discriminated a benchmark set of 100 interface structures from a set of sequence-randomized negative examples with a false positive rate of 3% and a true positive rate of 97%. Moreover, the predicted complexes were also filtered using functional annotation and sub-cellular localization data. The ability of the method to select the correct binding mode among alternates is demonstrated for three camelid VHH domain-porcine alpha-amylase interactions. We also highlight the prediction of co-complexed domain superfamilies that are not present in template complexes. Through integration with MODBASE, the application of the method to proteomes that are less well characterized than that of S.cerevisiae will contribute to expansion of the structural and functional coverage of protein interaction space. The predicted complexes are deposited in MODBASE (http://salilab.org/modbase).

Algorithms↗

Revisiting the prediction of protein function at CASP6.

The ability to predict the function of a protein, given its sequence and/or 3D structure, is an essential requirement for exploiting the wealth of data made available by genomics and structural genomics projects and is therefore raising increasing interest in the computational biology community. To foster developments in the area as well as to establish the state of the art of present methods, a function prediction category was tentatively introduced in the 6th edition of the Critical Assessment of Techniques for Protein Structure Prediction (CASP) worldwide experiment. The assessment of the performance of the methods was made difficult by at least two factors: (a) the experimentally determined function of the targets was not available at the time of assessment; (b) the experiment is run blindly, preventing verification of whether the convergence of different predictions towards the same functional annotation was due to the similarity of the methods or to a genuine signal detectable by different methodologies. In this work, we collected information about the methods used by the various predictors and revisited the results of the experiment by verifying how often and in which cases a convergent prediction was obtained by methods based on different rationale. We propose a method for classifying the type and redundancy of the methods. We also analyzed the cases in which a function for the target protein has become available. Our results show that predictions derived from a consensus of different methods can reach an accuracy as high as 80%. It follows that some of the predictions submitted to CASP6, once reanalyzed taking into account the type of converging methods, can provide very useful information to researchers interested in the function of the target proteins.

Caspase 6↗

Genomes of the ex-type strains of Elsinoë mangiferae and E. perseae, the causal agents of scab on mango and avocado.

Elsinoë species are slow-growing, hemibiotrophic to necrotrophic fungi that cause scab diseases on economically important fruit crops. Genome resources for many host-specific species remain limited. We report high-quality draft genome assemblies for the ex-type strains of Elsinoë mangiferae (CBS 226.50) and E. perseae (CBS 406.34), causal agents of mango and avocado scab, respectively. Among 5 approaches tested, a Nanopore-only NextDenovo assembly produced the most contiguous genomes, yielding 24.5 Mb (E. mangiferae) and 25.1 Mb (E. perseae) assemblies with 13 and 18 contigs, respectively, BUSCO completeness scores of ∼94%, and multiple putative telomere-to-telomere chromosomes. Gene prediction identified 9,134 and 9,243 genes, respectively. Functional annotation revealed enrichment of metabolic and regulatory pathways, including those involved in posttranslational modification, protein transport, and secondary metabolism. Carbohydrate-active enzyme repertoires were small but conserved, consistent with stealth pathogenicity strategies and low plant cell wall degradation. Both genomes encoded large secretomes (>850 proteins), diverse protease repertoires (>300 proteins), Ecp2-like effector proteins, and multiple biosynthetic gene clusters, including clusters with similarity to those associated with elsinochrome and ACT-toxin II biosynthesis, some of which may contribute to host-pathogen interactions and disease development. A large fraction of genes lacked functional characterization, suggesting incomplete databases and/or the presence of lineage-specific genes potentially involved in virulence or host adaptation. These genome resources fill critical gaps for underrepresented Elsinoë species and provide taxonomically anchored references essential for diagnostics, comparative genomics, and research into the molecular basis of host specificity and pathogenicity in scab-causing fungi.

Persea↗

Comparative analysis of Saccharomyces cerevisiae WW domains and their interacting proteins.

BACKGROUND: The WW domain is found in a large number of eukaryotic proteins implicated in a variety of cellular processes. WW domains bind proline-rich protein and peptide ligands, but the protein interaction partners of many WW domain-containing proteins in Saccharomyces cerevisiae are largely unknown. RESULTS: We used protein microarray technology to generate a protein interaction map for 12 of the 13 WW domains present in proteins of the yeast S. cerevisiae. We observed 587 interactions between these 12 domains and 207 proteins, most of which have not previously been described. We analyzed the representation of functional annotations within the network, identifying enrichments for proteins with peroxisomal localization, as well as for proteins involved in protein turnover and cofactor biosynthesis. We compared orthologs of the interacting proteins to identify conserved motifs known to mediate WW domain interactions, and found substantial evidence for the structural conservation of such binding motifs throughout the yeast lineages. The comparative approach also revealed that several of the WW domain-containing proteins themselves have evolutionarily conserved WW domain binding sites, suggesting a functional role for inter- or intramolecular association between proteins that harbor WW domains. On the basis of these results, we propose a model for the tuning of interactions between WW domains and their protein interaction partners. CONCLUSION: Protein microarrays provide an appealing alternative to existing techniques for the construction of protein interaction networks. Here we built a network composed of WW domain-protein interactions that illuminates novel features of WW domain-containing proteins and their protein interaction partners.

Amino Acid Motifs↗

Intermediary metabolism in sea urchin: the first inferences from the genome sequence.

The genome sequence of the purple sea urchin Strongylocentrotus purpuratus recently became available. We report the results of functional annotation and initial analysis of more than 2300 proteins predicted to be involved in metabolite transport and enzymatic conversion in sea urchin. The comparison of various reconstructed biosynthetic and catabolic pathways in sea urchin to those known in other genomes suggests the overall similarity of the sea urchin metabolism to that of the vertebrates, with relatively small but non-trivial differences from both vertebrates and protostomes. There are several examples of two parallel, non-orthologous solutions for the same molecular function in sea urchin, in contrast with the other completely sequenced metazoans that tend to contain just one version of the same function. There are also genes that appear to be close phylogenetic neighbors of plant or bacterial homologs, as opposed to homologs in other Metazoa. The evolutionary and functional significance of these variations is discussed.

Amino Acids↗

The predicted secretome of Lactobacillus plantarum WCFS1 sheds light on interactions with its environment.

The predicted extracellular proteins of the bacterium Lactobacillus plantarum were analysed to gain insight into the mechanisms underlying interactions of this bacterium with its environment. Extracellular proteins play important roles in processes ranging from probiotic effects in the gastrointestinal tract to degradation of complex extracellular carbon sources such as those found in plant materials, and they have a primary role in the adaptation of a bacterium to changing environmental conditions. The functional annotation of extracellular proteins was improved using a wide variety of bioinformatics methods, including domain analysis and phylogenetic profiling. At least 12 proteins are predicted to be directly involved in adherence to host components such as collagen and mucin, and about 30 extracellular enzymes, mainly hydrolases and transglycosylases, might play a role in the degradation of substrates by L. plantarum to sustain its growth in different environmental niches. A comprehensive overview of all predicted extracellular proteins, their domains composition and their predicted function is provided through a database at http://www.cmbi.ru.nl/secretome which could serve as a basis for targeted experimental studies into the function of extracellular proteins.

Bacterial Adhesion↗

B cell pathways implicate shared genetic architecture between schizophrenia and immune-mediated diseases.

BACKGROUND: Schizophrenia and immune-mediated diseases are globally prevalent and highly heritable conditions that frequently co-occur, posing major public health burdens. However, their shared genetic architecture remains poorly understood. METHODS: We applied the bivariate causal mixture model (MiXeR) to investigate the polygenic overlap between schizophrenia and eight common immune-mediated diseases, using genome-wide association study summary statistics comprising 2,489 to 67,323 cases and 9,066 to 497,622 controls. Shared loci were identified through conditional/conjunctional false discovery rate (cond/conjFDR), local genetic correlation (LAVA), and colocalization analyses. Subsequently, gene mapping, functional annotation, expression-trait association, and drug-gene interaction analyses were performed to explore shared genes and enriched pathways, and genetic risk scores (GRS) from the UK Biobank were used to validate the findings. RESULTS: MiXeR estimated substantial polygenic overlap between schizophrenia and immune-mediated diseases, and conjFDR identified 133 shared loci, with eight prioritized through local genetic correlation and colocalization signals. These eight loci were mapped to 85 protein-coding genes enriched in pathways essential for B cell function. Among them, S-PrediXcan analyses identified 14 genes whose expression in brain tissues or blood was associated with both diseases. These genes also interact with immunomodulatory or antihypertensive drugs. Additionally, 11 of the 14 genes were linked to innate immunity and/or cognitive traits. Using UK Biobank data, we further confirmed that overall, shared gene, and B cell activation and receptor signaling pathway–specific genetic risk for schizophrenia is associated with immune-mediated disease susceptibility. CONCLUSIONS: These findings underscore the shared genetic architecture of schizophrenia and immune-mediated diseases, advancing insights at the interface of psychiatric genetics and immunology.

Schizophrenia↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from ≈32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

Gene discovery in Plasmodium vivax through sequencing of ESTs from mixed blood stages.

Despite the significance of Plasmodium vivax as the most widespread human malaria parasite and a major public health problem, gene expression in this parasite is poorly understood. To accelerate gene discovery and facilitate the annotation phase of the P. vivax genome project, we have undertaken a transcriptome approach to study gene expression in the mixed blood stages of a P. vivax field isolate. Using a cDNA library constructed from purified blood stages, we have obtained single-pass sequences for approximately 21,500 expressed sequence tags (ESTs), the largest number of transcript tags obtained so far for this species. Cluster analysis revealed that the library is highly redundant, resulting in 5407 clusters. Clustered ESTs were searched against public protein databases for functional annotation, and more than one-third showed a significant match, the majority of these to Plasmodium falciparum proteins. The most abundant clusters were to genes encoding ribosomal proteins and proteins involved in metabolism, consistent with the predominance of trophozoites in the field isolate sample. In spite of the scarcity of other parasite stages in the field isolate, we could identify genes that are expressed in rings, schizonts and gametocytes. This study should facilitate our understanding of the gene expression in P. vivax asexual stages and provide valuable data for gene prediction and annotation of the P. vivax genome sequence.

Animals↗