PubMed Health⌕ Search

Biomedical subjects

Kerr Wall

Publications and source records attributed to Kerr Wall.

5 recordsLinked to original sources

ChloroplastDB: the Chloroplast Genome Database.

The Chloroplast Genome Database (ChloroplastDB) is an interactive, web-based database for fully sequenced plastid genomes, containing genomic, protein, DNA and RNA sequences, gene locations, RNA-editing sites, putative protein families and alignments (http://chloroplast.cbio.psu.edu/). With recent technical advances, the rate of generating new organelle genomes has increased dramatically. However, the established ontology for chloroplast genes and gene features has not been uniformly applied to all chloroplast genomes available in the sequence databases. For example, annotations for some published genome sequences have not evolved with gene naming conventions. ChloroplastDB provides unified annotations, gene name search, BLAST and download functions for chloroplast encoded genes and genomic sequences. A user can retrieve all orthologous sequences with one search regardless of gene names in GenBank. This feature alone greatly facilitates comparative research on sequence evolution including changes in gene content, codon usage, gene structure and post-transcriptional modifications such as RNA editing. Orthologous protein sets are classified by TribeMCL and each set is assigned a standard gene name. Over the next few years, as the number of sequenced chloroplast genomes increases rapidly, the tools available in ChloroplastDB will allow researchers to easily identify and compile target data for comparative analysis of chloroplast genes and genomes.

Chloroplasts↗

Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA).

In the eight years since phylogenomics was introduced as the intersection of genomics and phylogenetics, the field has provided fundamental insights into gene function, genome history and organismal relationships. The utility of phylogenomics is growing with the increase in the number and diversity of taxa for which whole genome and large transcriptome sequence sets are being generated. We assert that the synergy between genomic and phylogenetic perspectives in comparative biology would be enhanced by the development and refinement of minimal reporting standards for phylogenetic analyses. Encouraged by the development of the Minimum Information About a Microarray Experiment (MIAME) standard, we propose a similar roadmap for the development of a Minimal Information About a Phylogenetic Analysis (MIAPA) standard. Key in the successful development and implementation of such a standard will be broad participation by developers of phylogenetic analysis software, phylogenetic database developers, practitioners of phylogenomics, and journal editors.

Genomics↗

Phylogeny and domain evolution in the APETALA2-like gene family.

The combined processes of gene duplication, nucleotide substitution, domain duplication, and intron/exon shuffling can generate a complex set of related genes that may differ substantially in their expression patterns and functions. The APETALA2-like (AP2-like) gene family exhibits patterns of both gene and domain duplication, coupled with changes in sequence, exon arrangement, and expression. In angiosperms, these genes perform an array of functions including the establishment of the floral meristem, the specification of floral organ identity, the regulation of floral homeotic gene expression, the regulation of ovule development, and the growth of floral organs. To determine patterns of gene diversification, we conducted a series of broad phylogenetic analyses of AP2-like sequences from green plants. These studies indicate that the AP2 domain was duplicated prior to the divergence of the two major lineages of AP2-like genes, euAP2 and AINTEGUMENTA (ANT). Structural features of the AP2-like genes as well as phylogenetic analyses of nucleotide and amino acid (aa) sequences of the AP2-like gene family support the presence of the two major lineages. The ANT lineage is supported by a 10-aa insertion in the AP2-R1 domain and a 1-aa insertion in the AP2-R2 domain, relative to all other members of the AP2-like family. MicroRNA172-binding sequences, the function of which has been studied in some of the AP2-like genes in Arabidopsis, are restricted to the euAP2 lineage. Within the ANT lineage, the euANT lineage is characterized by four conserved motifs: one in the 10-aa insertion in the AP2-R1 domain (euANT1) and three in the predomain region (euANT2, euANT3, and euANT4). Our expression studies show that the euAP2 homologue from Amborella trichopoda, the putative sister to all other angiosperms, is expressed in all floral organs as well as leaves.

Amino Acid Sequence↗

EST clustering error evaluation and correction.

MOTIVATION: The gene expression intensity information conveyed by (EST) Expressed Sequence Tag data can be used to infer important cDNA library properties, such as gene number and expression patterns. However, EST clustering errors, which often lead to greatly inflated estimates of obtained unique genes, have become a major obstacle in the analyses. The EST clustering error structure, the relationship between clustering error and clustering criteria, and possible error correction methods need to be systematically investigated. RESULTS: We identify and quantify two types of EST clustering error, namely, Type I and II in EST clustering using CAP3 assembling program. A Type I error occurs when ESTs from the same gene do not form a cluster whereas a Type II error occurs when ESTs from distinct genes are falsely clustered together. While the Type II error rate is <1.5% for both 5' and 3' EST clustering, the Type I error in the 5' EST case is approximately 10 times higher than the 3' EST case (30% versus 3%). An over-stringent identity rule, e.g., P >/= 95%, may even inflate the Type I error in both cases. We demonstrate that approximately 80% of the Type I error is due to insufficient overlap among sibling ESTs (ISO error) in 5' EST clustering. A novel statistical approach is proposed to correct ISO error to provide more accurate estimates of the true gene cluster profile.

Algorithms↗