PubMed Health⌕ Search

Biomedical subjects

Dawei Lin

Publications and source records attributed to Dawei Lin.

8 recordsLinked to original sources

Reconstruction of ancient genome and gene order from complete microbial genome sequences.

Microbial genome sequences provide us with the fossil records for inferring their origination and evolution. Assuming that current microbial genomes are the evolutionary results of ancient genomes or fragments and the neighboring genes in ancient genomes are more likely neighbors in current genomes, in this paper we proposed a paleontological algorithm and assembled the orthologous gene groups from 66 complete and current microbial genome sequences into a pseudo-ancient genome, which consists of continuous fragments of various sizes. We performed bootstrap resampling and correlation analyses and the results showed that the assembled ancient genome and fragments are statistically significant and the genes of the same fragment are inherently related and likely derived from common ancestors. This method provides a new computational tool for studying microbial genome structure and evolution.

Algorithms↗

The high-throughput protein-to-structure pipeline at SECSG.

Using a high degree of automation, the crystallography core at the Southeast Collaboratory for Structural Genomics (SECSG) has developed a high-throughput protein-to-structure pipeline. Various robots and automation procedures have been adopted and integrated into a pipeline that is capable of screening 40 proteins for crystallization and solving four protein structures per week. This pipeline is composed of three major units: crystallization, structure determination/validation and crystallomics. Coupled with the protein-production cores at SECSG, the protein-to-structure pipeline provides a two-tiered approach for protein production at SECSG. In tier 1, all protein samples supplied by the protein-production cores pass through the pipeline using standard crystallization screening and optimization procedures. The protein targets that failed to yield diffraction-quality crystals (resolution better than 3.0 A) become tier 2 or salvaging targets. The goal of tier 2 target salvaging, carried out by the crystallomics core, is to produce the target proteins with increased purity and homogeneity, which would render them more likely to yield well diffracting crystals. This is performed by alternative purification procedures and/or the introduction of chemical modifications to the proteins (such as tag removal, methylation, surface mutagenesis, selenomethionine labelling etc.). Details of the various procedures in the pipeline for protein crystallization, target salvaging, data collection/processing and high-throughput structure determination/validation, as well as some examples, are described.

Crystallization↗

Parameter-space screening: a powerful tool for high-throughput crystal structure determination.

The determination of protein structures on a genomic scale requires both computing capacity and efficiency increases at many stages along the complex process. By combining bioinformatics workflow-management techniques, cluster-based computing and popular crystallographic structure-determination software packages, an efficient and powerful new tool for structural biology/genomics has been developed. Using the workflow manager and a simple web interface, the researcher can, in a few easy steps, set up hundreds of structure-determination jobs, each using a slightly different set of program input parameters, thus efficiently screening parameter space for the optimal input-parameter combination, i.e. a set of parameters that leads to a successful structure determination. Upon completion, results from the programs are harvested, analyzed, sorted based on success and presented to the user via the web interface. This approach has been applied with success in more than 30 cases. Examples of successful structure determinations based on single-wavelength scattering (SAS) are described and include cases where the 'rational' crystallographer-based selection of input parameters values had failed.

Computational Biology↗

Protein production and crystallization at SECSG -- an overview.

Using a high degree of automation, the Southeast Collaboratory for Structural Genomics (SECSG) has developed high throughput pipelines for protein production, and crystallization using a two-tiered approach. Primary, or tier-1, protein production focuses on producing proteins for members of large Pfam families that lack a representative structure in the Protein Data Bank. Target genomes are Pyrococcus furiosus and Caenorhabditis elegans. Selected human proteins are also under study. Tier-2 protein production, or target rescue, focuses on those tier-1 proteins, which either fail to crystallize or give poorly diffracting crystals. This two tier approach is more efficient since it allows the primary protein production groups to focus on the production of new targets while the tier-2 efforts focus on providing additional sample for further studies and modified protein for structure determination. Both efforts feed the SECSG high throughput crystallization pipeline, which is capable of screening over 40 proteins per week. Details of the various pipelines in use by the SECSG for protein production and crystallization, as well as some examples of target rescue are described.

Animals↗

Salvaging Pyrococcus furiosus protein targets at SECSG.

Proteins derived from the coding regions of Pyrococcus furiosus are targets for three-dimensional X-ray and NMR structure determination by the Southeast Collaboratory for Structural Genomics (SECSG). Of the 2,200 open reading frames (ORFs) in this organism, 220 protein targets were cloned and expressed in a high-throughput (HT) recombinant system for crystallographic studies. However, only 96 of the expressed proteins could be crystallized and, of these, only 15 have led to structures. To address this issue, SECSG has recently developed a two-tier approach to protein production and crystallization. In this approach, tier-1 efforts are focused on producing protein for new Pfu(italics?) targets using a high-throughput approach. Tier-2 protein production efforts support tier-1 activities by (1) producing additional protein for further crystallization trials, (2) producing modified protein (further purification, methylation, tag removal, selenium labeling, etc) as required and (3) serving as a salvaging pathway for failed tier-1 proteins. In a recent study using this two-tiered approach, nine structures were determined from a set of 50 Pfu proteins, which failed to produce crystals suitable for X-ray diffraction analysis. These results validate this approach and suggest that it has application to other HT crystal structure determination applications.

Archaeal Proteins↗

Life in the fast lane for protein crystallization and X-ray crystallography.

The common goal for structural genomic centers and consortiums is to decipher as quickly as possible the three-dimensional structures for a multitude of recombinant proteins derived from known genomic sequences. Since X-ray crystallography is the foremost method to acquire atomic resolution for macromolecules, the limiting step is obtaining protein crystals that can be useful of structure determination. High-throughput methods have been developed in recent years to clone, express, purify, crystallize and determine the three-dimensional structure of a protein gene product rapidly using automated devices, commercialized kits and consolidated protocols. However, the average number of protein structures obtained for most structural genomic groups has been very low compared to the total number of proteins purified. As more entire genomic sequences are obtained for different organisms from the three kingdoms of life, only the proteins that can be crystallized and whose structures can be obtained easily are studied. Consequently, an astonishing number of genomic proteins remain unexamined. In the era of high-throughput processes, traditional methods in molecular biology, protein chemistry and crystallization are eclipsed by automation and pipeline practices. The necessity for high-rate production of protein crystals and structures has prevented the usage of more intellectual strategies and creative approaches in experimental executions. Fundamental principles and personal experiences in protein chemistry and crystallization are minimally exploited only to obtain "low-hanging fruit" protein structures. We review the practical aspects of today's high-throughput manipulations and discuss the challenges in fast pace protein crystallization and tools for crystallography. Structural genomic pipelines can be improved with information gained from low-throughput tactics that may help us reach the higher-bearing fruits. Examples of recent developments in this area are reported from the efforts of the Southeast Collaboratory for Structural Genomics (SECSG).

Crystallization↗

Abundant type 10 17 beta-hydroxysteroid dehydrogenase in the hippocampus of mouse Alzheimer's disease model.

A full-length cDNA of mouse type 10 17 beta-hydroxysteroid dehydrogenase (17 beta-HSD10) was cloned from brain, representing the accurate nucleotide sequence information that rendered possible an accurate deduction of the amino acid sequence of the wild-type enzyme. A comparison of sequences and three-dimensional models of this enzyme revealed that structures previously reported by other groups carry either a truncated or mutated amino-terminal sequence. Fusion of the first 11 residues of the wild-type enzyme to the green fluorescent protein directed the reporter protein into mitochondria. Thus, the N-terminus was identified as a mitochondrial targeting signal that accounts for the intracellular localization of the mouse enzyme. This enzyme is normally associated with mitochondria, not with the endoplasmic reticulum as suggested by its trivial name 'endoplasmic reticulum-associated amyloid-beta biding protein (ERAB)'. After its C-terminal region was used to raise rabbit anti-17 betaHSD10 antibodies, immunogold electron microscopy showed that an abundance of this enzyme could be found in hippocampal synaptic mitochondria of betaAPP transgenic mice, but not in normal controls. High levels of this enzyme may disrupt steroid hormone homeostasis in synapses and contribute to synapse loss in the hippocampus of the mouse Alzheimer's disease model.

17-Hydroxysteroid Dehydrogenases↗

Improving solubility of Shewanella oneidensis MR-1 and Clostridium thermocellum JW-20 proteins expressed into Esherichia coli.

Low solubility of proteins overexpressed in E. coli is a frequent problem in high-throughput structural genomics. To improve solubility of proteins from mesophilic Shewanella oneidensis MR-1 and thermophilic Clostridium thermocellum JW20, an approach was attempted that included a fusion of the target protein to a maltose-binding protein (MBP) and a decrease of induction temperature. The MBP was selected as the most efficient solubilizing carrier when compared to a glutathione S-transferase and a Nus A protein. A tobacco etch virus (TEV) protease recognition site was introduced between fused proteins using a double polymerase-chain reaction and four primers. In this way, 79 S. oneidensis proteins have been expressed in one case with an N-terminal 30-residue tag and in another case as a fusion protein with MBP. A foreign tag might significantly affect the properties of the target polypeptide. At 37 degrees C and 18 degrees C induction temperatures, only 5 and 17 tagged proteins were soluble, respectively. In fusion with MBP 4, 34, and 38 proteins were soluble upon induction at 37 degrees, 28 degrees, and 18 degrees C, respectively. The MBP is assumed to increase stability and solubility of a target protein by changing both the mechanism and the cooperativity of folding/unfolding. The 66 C. thermocellum proteins were expressed as fusion proteins with MBP. Induction at 37 degrees, 28 degrees, and 18 degrees C produced 34, 57, and 60 soluble proteins, respectively. The higher solubility of C. thermocellum proteins in comparison with the S. oneidensis proteins under similar conditions of induction correlates with the thermophilicity of the host. The two-factor Wilkinson-Harrison statistical model was used to identify soluble and insoluble proteins. Theoretical and experimental data showed good agreement for S. oneidensis proteins; however, the model failed to identify soluble/insoluble Clostridium proteins. A suggestion has been made that the Wilkinson-Harrison model is not applicable to C. thermocellum proteins because it did not account for the peculiarities of protein sequences from thermophiles.

Amino Acid Sequence↗