PubMed Health⌕ Search

Biomedical subjects

Liangjiang Wang

Publications and source records attributed to Liangjiang Wang.

7 recordsLinked to original sources

BeetleBase: the model organism database for Tribolium castaneum.

BeetleBase (http://www.bioinformatics.ksu.edu/BeetleBase/) is an integrated resource for the Tribolium research community. The red flour beetle (Tribolium castaneum) is an important model organism for genetics, developmental biology, toxicology and comparative genomics, the genome of which has recently been sequenced. BeetleBase is constructed to integrate the genomic sequence data with information about genes, mutants, genetic markers, expressed sequence tags and publications. BeetleBase uses the Chado data model and software components developed by the Generic Model Organism Database (GMOD) project. This strategy not only reduces the time required to develop the database query tools but also makes the data structure of BeetleBase compatible with that of other model organism databases. BeetleBase will be useful to the Tribolium research community for genome annotation as well as comparative genomics.

Animals↗

BindN: a web-based tool for efficient prediction of DNA and RNA binding sites in amino acid sequences.

BindN (http://bioinformatics.ksu.edu/bindn/) takes an amino acid sequence as input and predicts potential DNA or RNA-binding residues with support vector machines (SVMs). Protein datasets with known DNA or RNA-binding residues were selected from the Protein Data Bank (PDB), and SVM models were constructed using data instances encoded with three sequence features, including the side chain pK(a) value, hydrophobicity index and molecular mass of an amino acid. The results suggest that DNA-binding residues can be predicted at 69.40% sensitivity and 70.47% specificity, while prediction of RNA-binding residues achieves 66.28% sensitivity and 69.84% specificity. When compared with previous studies, the SVM models appear to be more accurate and more efficient for online predictions. BindN provides a useful tool for understanding the function of DNA and RNA-binding proteins based on primary sequence data.

Amino Acids↗

Comparative analysis of expressed sequences reveals a conserved pattern of optimal codon usage in plants.

Codon usage bias is a ubiquitous phenomenon, which may be caused by mutational bias, selection, or both. The patterns of codon usage in plants are not well understood. Datasets of expressed sequence tags (ESTs) available for many plant species provide the resources for large-scale comparative analysis of codon usage patterns. We developed a computational approach to translate EST or assembled contig sequences, and then used the coding information for comparative analysis of codon usage in 12 plant species, including 6 eudicots, 5 monocots and the green alga Chlamydomonas reinhardtii. While codon nucleotide composition is highly conserved within eudicots or monocots, there is a significant difference between these two major taxonomic groups of higher plants. The third nucleotide position of codons is AU-rich in the eudicot genomes (35-42% of G+C content), but GC-rich in the monocot genomes (59-61% of G+C content). To identify optimal codons in these species, we used EST counts to estimate gene transcript levels. It was demonstrated that codon usage bias is correlated positively with gene transcript levels. Interestingly, the use of optimal codons appears to be well conserved between eudicots and monocots, and to a lesser degree between the higher plants and C. reinhardtii. Most of the optimal codons end with a C or G base, regardless of the different nucleotide composition in these genomes. The results suggest that plant codon usage is affected by translational selection, and the selective pressure appears to be conserved in the plant kingdom.

Animals↗

The WRKY transcription factor superfamily: its origin in eukaryotes and expansion in plants.

BACKGROUND: WRKY proteins are newly identified transcription factors involved in many plant processes including plant responses to biotic and abiotic stresses. To date, genes encoding WRKY proteins have been identified only from plants. Comprehensive search for WRKY genes in non-plant organisms and phylogenetic analysis would provide invaluable information about the origin and expansion of the WRKY family. RESULTS: We searched all publicly available sequence data for WRKY genes. A single copy of the WRKY gene encoding two WRKY domains was identified from Giardia lamblia, a primitive eukaryote, Dictyostelium discoideum, a slime mold closely related to the lineage of animals and fungi, and the green alga Chlamydomonas reinhardtii, an early branching of plants. This ancestral WRKY gene seems to have duplicated many times during the evolution of plants, resulting in a large family in evolutionarily advanced flowering plants. In rice, the WRKY gene family consists of over 100 members. Analyses suggest that the C-terminal domain of the two-WRKY-domain encoding gene appears to be the ancestor of the single-WRKY-domain encoding genes, and that the WRKY domains may be phylogenetically classified into five groups. We propose a model to explain the WRKY family's origin in eukaryotes and expansion in plants. CONCLUSIONS: WRKY genes seem to have originated in early eukaryotes and greatly expanded in plants. The elucidation of the evolution and duplicative expansion of the WRKY genes should provide valuable information on their functions.

Algorithms↗

Tall fescue EST-SSR markers with transferability across several grass species.

Tall fescue (Festuca arundinacea Schreb.) is a major cool season forage and turf grass in the temperate regions of the world. It is also a close relative of other important forage and turf grasses, including meadow fescue and the cultivated ryegrass species. Until now, no SSR markers have been developed from the tall fescue genome. We designed 157 EST-SSR primer pairs from tall fescue ESTs and tested them on 11 genotypes representing seven grass species. Nearly 92% of the primer pairs produced characteristic simple sequence repeat (SSR) bands in at least one species. A large proportion of the primer pairs produced clear reproducible bands in other grass species, with most success in the close taxonomic relatives of tall fescue. A high level of marker polymorphism was observed in the outcrossing species tall fescue and ryegrass (66%). The marker polymorphism in the self-pollinated species rice and wheat was low (43% and 38%, respectively). These SSR markers were useful in the evaluation of genetic relationships among the Festuca and Lolium species. Sequencing of selected PCR bands revealed that the nucleotide sequences of the forage grass genotypes were highly conserved. The two cereal species, particularly rice, had significantly different nucleotide sequences compared to the forage grasses. Our results indicate that the tall fescue EST-SSR markers are valuable genetic markers for the Festuca and Lolium genera. These are also potentially useful markers for comparative genomics among several grass species.

Base Sequence↗

Metabolomics spectral formatting, alignment and conversion tools (MSFACTs).

MOTIVATION: The amplified interest in metabolic profiling has generated the need for additional tools to assist in the rapid analysis of complex data sets. RESULTS: A new program; metabolomics spectral formatting, alignment and conversion tools, (MSFACTs) is described here for the automated import, reformatting, alignment, and export of large chromatographic data sets to allow more rapid visualization and interrogation of metabolomic data. MSFACTs incorporates two tools: one for the alignment of integrated chromatographic peak lists and another for extracting information from raw chromatographic ASCII formatted data files. MSFACTs is illustrated in the processing of GC/MS metabolomic data from different tissues of the model legume plant, Medicago truncatula. The results document that various tissues such as roots, stems, and leaves from the same plant can be easily differentiated based on metabolite profiles. Further, similar types of tissues within the same plant, such as the first to eleventh internodes of stems, could also be differentiated based on metabolite profiles. AVAILABILITY: Freely available upon request for academic and non-commercial use. Commercial use is available through licensing agreement http://www.noble.org/PlantBio/MS/MSFACTs/MSFACTs.html.

Cluster Analysis↗

Mapping the proteome of barrel medic (Medicago truncatula).

A survey of six organ-/tissue-specific proteomes of the model legume barrel medic (Medicago truncatula) was performed. Two-dimensional polyacrylamide gel electrophoresis reference maps of protein extracts from leaves, stems, roots, flowers, seed pods, and cell suspension cultures were obtained. Five hundred fifty-one proteins were excised and 304 proteins identified using peptide mass fingerprinting and matrix-assisted laser desorption ionization time-of-flight mass spectrometry. Nanoscale high-performance liquid chromatography coupled with tandem quadrupole time-of-flight mass spectrometry was used to validate marginal matrix-assisted laser desorption ionization time-of-flight mass spectrometry protein identifications. This dataset represents one of the most comprehensive plant proteome projects to date and provides a basis for future proteome comparison of genetic mutants, biotically and abiotically challenged plants, and/or environmentally challenged plants. Technical details concerning peptide mass fingerprinting, database queries, and protein identification success rates in the absence of a sequenced genome are reported and discussed. A summary of the identified proteins and their putative functions are presented. The tissue-specific expression of proteins and the levels of identified proteins are compared with their related transcript abundance as quantified through EST counting. It is estimated that approximately 50% of the proteins appear to be correlated with their corresponding mRNA levels.

Amino Acid Sequence↗