PubMed Health⌕ Search

Biomedical subjects

Steven A Benner

Publications and source records attributed to Steven A Benner.

At least 19 recordsLinked to original sources

Molecular paleoscience: systems biology from the past.

Experimental paleomolecular biology, paleobiochemistry, and paleogenetics are closely related emerging fields that infer the sequences of ancient genes and proteins from now-extinct organisms, and then resurrect them for study in the laboratory. The goal of paleogenetics is to use information from natural history to solve the conundrum of modern genomics: How can we understand deeply the function of biomolecular structures uncovered and described by modern chemical biology? Reviewed here are the first 20 cases where biomolecular resurrections have been achieved. These show how paleogenetics can lead to an understanding of the function of biomolecules, analyze changing function, and put meaning to genomic sequences, all in ways that are not possible with traditional molecular biological studies.

Animals↗

2-Hydroxymethylboronate as a reagent to detect carbohydrates: application to the analysis of the formose reaction.

2-Hydroxymethylphenylboronate is described as a reagent that converts neutral 1,2-diols, as found in simple carbohydrates, into 1:1 anionic complexes that are easily detected by Fourier transform ion cyclotron resonance mass spectrometry. The value of this reagent was demonstrated through its application to analyze complex mixtures of carbohydrates formed in the formose process, often cited as a way that biologically significant carbohydrates might have been generated from formaldehyde under prebiotic conditions. Coupled with isotope studies, the reagent shows that the simplest autocatalytic cycle for the consumption of formaldehyde in this process cannot account for the bulk consumption of formaldehyde.

Boron Compounds↗

Artificially expanded genetic information system: a new base pair with an alternative hydrogen bonding pattern.

To support efforts to develop a 'synthetic biology' based on an artificially expanded genetic information system (AEGIS), we have developed a route to two components of a non-standard nucleobase pair, the pyrimidine analog 6-amino-5-nitro-3-(1'-beta-D-2'-deoxyribofuranosyl)-2(1H)-pyridone (dZ) and its Watson-Crick complement, the purine analog 2-amino-8-(1'-beta-D-2'-deoxyribofuranosyl)-imidazo[1,2-a]-1,3,5-triazin-4(8H)-one (dP). These implement the pyDDA:puAAD hydrogen bonding pattern (where 'py' indicates a pyrimidine analog and 'pu' indicates a purine analog, while A and D indicate the hydrogen bonding patterns of acceptor and donor groups presented to the complementary nucleobases, from the major to the minor groove). Also described is the synthesis of the triphosphates and protected phosphoramidites of these two nucleosides. We also describe the use of the protected phosphoramidites to synthesize DNA oligonucleotides containing these AEGIS components, verify the absence of epimerization of dZ in those oligonucleotides, and report some hybridization properties of the dZ:dP nucleobase pair, which is rather strong, and the ability of each to effectively discriminate against mismatches in short duplex DNA.

Base Pair Mismatch↗

Dynamic assembly of primers on nucleic acid templates.

A strategy is presented that uses dynamic equlibria to assemble in situ composite DNA polymerase primers, having lengths of 14 or 16 nt, from DNA fragments that are 6 or 8 nt in length. In this implementation, the fragments are transiently joined under conditions of dynamic equilibrium by an imine linker, which has a dissociation constant of approximately 1 muM. If a polymerase is able to extend the composite, but not the fragments, it is possible to prime the synthesis of a target DNA molecule under conditions where two useful specificities are combined: (i) single nucleotide discrimination that is characteristic of short oligonucleotide duplexes (four to six nucleobase pairs in length), which effectively excludes single mismatches, and (ii) an overall specificity of priming that is characteristic of long (14 to 16mers) oligonucleotides, potentially unique within a genome. We report here the screening of a series of polymerases that combine an ability not to accept short primer fragments with an ability to accept the long composite primer held together by an unnatural imine linkage. Several polymerases were found that achieve this combination, permitting the implementation of the dynamic combinatorial chemical strategy.

Base Pair Mismatch↗

Analysis of transitions at two-fold redundant sites in mammalian genomes. Transition redundant approach-to-equilibrium (TREx) distance metrics.

BACKGROUND: The exchange of nucleotides at synonymous sites in a gene encoding a protein is believed to have little impact on the fitness of a host organism. This should be especially true for synonymous transitions, where a pyrimidine nucleotide is replaced by another pyrimidine, or a purine is replaced by another purine. This suggests that transition redundant exchange (TREx) processes at the third position of conserved two-fold codon systems might offer the best approximation for a neutral molecular clock, serving to examine, within coding regions, theories that require neutrality, determine whether transition rate constants differ within genes in a single lineage, and correlate dates of events recorded in genomes with dates in the geological and paleontological records. To date, TREx analysis of the yeast genome has recognized correlated duplications that established a new metabolic strategies in fungi, and supported analyses of functional change in aromatases in pigs. TREx dating has limitations, however. Multiple transitions at synonymous sites may cause equilibration and loss of information. Further, to be useful to correlate events in the genomic record, different genes within a genome must suffer transitions at similar rates. RESULTS: A formalism to analyze divergence at two fold redundant codon systems is presented. This formalism exploits two-state approach-to-equilibrium kinetics from chemistry. This formalism captures, in a single equation, the possibility of multiple substitutions at individual sites, avoiding any need to "correct" for these. The formalism also connects specific rate constants for transitions to specific approximations in an underlying evolutionary model, including assumptions that transition rate constants are invariant at different sites, in different genes, in different lineages, and at different times. Therefore, the formalism supports analyses that evaluate these approximations. Transitions at synonymous sites within two-fold redundant coding systems were examined in the mouse, rat, and human genomes. The key metric (f2), the fraction of those sites that holds the same nucleotide, was measured for putative ortholog pairs. A transition redundant exchange (TREx) distance was calculated from f2 for these pairs. Pyrimidine-pyrimidine transitions at these sites occur approximately 14% faster than purine-purine transitions in various lineages. Transition rate constants were similar in different genes within the same lineages; within a set of orthologs, the f2 distribution is only modest overdispersed. No correlation between disparity and overdispersion is observed. In rodents, evidence was found for greater conservation of TREx sites in genes on the X chromosome, accounting for a small part of the overdispersion, however. CONCLUSION: The TREx metric is useful to analyze the history of transition rate constants within these mammals over the past 100 million years. The TREx metric estimates the extent to which silent nucleotide substitutions accumulate in different genes, on different chromosomes, with different compositions, in different lineages, and at different times.

Animals↗

Application of DETECTER, an evolutionary genomic tool to analyze genetic variation, to the cystic fibrosis gene family.

BACKGROUND: The medical community requires computational tools that distinguish missense genetic differences having phenotypic impact within the vast number of sense mutations that do not. Tools that do this will become increasingly important for those seeking to use human genome sequence data to predict disease, make prognoses, and customize therapy to individual patients. RESULTS: An approach, termed DETECTER, is proposed to identify sites in a protein sequence where amino acid replacements are likely to have a significant effect on phenotype, including causing genetic disease. This approach uses a model-dependent tool to estimate the normalized replacement rate at individual sites in a protein sequence, based on a history of those sites extracted from an evolutionary analysis of the corresponding protein family. This tool identifies sites that have higher-than-average, average, or lower-than-average rates of change in the lineage leading to the sequence in the population of interest. The rates are then combined with sequence data to determine the likelihoods that particular amino acids were present at individual sites in the evolutionary history of the gene family. These likelihoods are used to predict whether any specific amino acid replacements, if introduced at the site in a modern human population, would have a significant impact on fitness. The DETECTER tool is used to analyze the cystic fibrosis transmembrane conductance regulator (CFTR) gene family. CONCLUSION: In this system, DETECTER retrodicts amino acid replacements associated with the cystic fibrosis disease with greater accuracy than alternative approaches. While this result validates this approach for this particular family of proteins only, the approach may be applicable to the analysis of polymorphisms generally, including SNPs in a human population.

Amino Acid Substitution↗

Integrating protein structures and precomputed genealogies in the Magnum database: examples with cellular retinoid binding proteins.

BACKGROUND: When accurate models for the divergent evolution of protein sequences are integrated with complementary biological information, such as folded protein structures, analyses of the combined data often lead to new hypotheses about molecular physiology. This represents an excellent example of how bioinformatics can be used to guide experimental research. However, progress in this direction has been slowed by the lack of a publicly available resource suitable for general use. RESULTS: The precomputed Magnum database offers a solution to this problem for ca. 1,800 full-length protein families with at least one crystal structure. The Magnum deliverables include 1) multiple sequence alignments, 2) mapping of alignment sites to crystal structure sites, 3) phylogenetic trees, 4) inferred ancestral sequences at internal tree nodes, and 5) amino acid replacements along tree branches. Comprehensive evaluations revealed that the automated procedures used to construct Magnum produced accurate models of how proteins divergently evolve, or genealogies, and correctly integrated these with the structural data. To demonstrate Magnum's capabilities, we asked for amino acid replacements requiring three nucleotide substitutions, located at internal protein structure sites, and occurring on short phylogenetic tree branches. In the cellular retinoid binding protein family a site that potentially modulates ligand binding affinity was discovered. Recruitment of cellular retinol binding protein to function as a lens crystallin in the diurnal gecko afforded another opportunity to showcase the predictive value of a browsable database containing branch replacement patterns integrated with protein structures. CONCLUSION: We integrated two areas of protein science, evolution and structure, on a large scale and created a precomputed database, known as Magnum, which is the first freely available resource of its kind. Magnum provides evolutionary and structural bioinformatics resources that are useful for identifying experimentally testable hypotheses about the molecular basis of protein behaviors and functions, as illustrated with the examples from the cellular retinoid binding proteins.

Amino Acid Sequence↗

A review: synthesis of aryl C-glycosides via the heck coupling reaction.

In this article, we focus on the synthesis of aryl C-glycosides via Heck coupling. It is organized based on the type of structures used in the assembly of the C-glycosides (also called C-nucleosides) with the following subsections: pyrimidine C-nucleosides, purine C-nucleosides, and monocyclic, bicyclic, and tetracyclic C-nucleosides. The reagents and conditions used for conducting the Heck coupling reactions are discussed. The subsequent conversion of the Heck products to the corresponding target molecules and the application of the target molecules are also described.

Chemistry↗

Locked nucleic acid molecular beacons.

A novel LNA-MB (molecular beacon based on locked nucleic acid bases) has been designed and investigated. It exhibits very high melting temperature and is thermally stable, shows superior single base mismatch discrimination capability, and is stable against digestion by nuclease and has no binding with single-stranded DNA binding proteins. The LNA-MB will be widely useful in a variety of areas, especially for in vivo hybridization studies.

Fluorescence↗

The use of thymidine analogs to improve the replication of an extra DNA base pair: a synthetic biological system.

Synthetic biology based on a six-letter genetic alphabet that includes the two non-standard nucleobases isoguanine (isoG) and isocytosine (isoC), as well as the standard A, T, G and C, is known to suffer as a consequence of a minor tautomeric form of isoguanine that pairs with thymine, and therefore leads to infidelity during repeated cycles of the PCR. Reported here is a solution to this problem. The solution replaces thymidine triphosphate by 2-thiothymidine triphosphate (2-thioTTP). Because of the bulk and hydrogen bonding properties of the thione unit in 2-thioT, 2-thioT does not mispair effectively with the minor tautomer of isoG. To test whether this might allow PCR amplification of a six-letter artificially expanded genetic information system, we examined the relative rates of misincorporation of 2-thioTTP and TTP opposite isoG using affinity electrophoresis. The concentrations of isoCTP and 2-thioTTP were optimal to best support PCR amplification using thermostable polymerases of a six-letter alphabet that includes the isoC-isoG pair. The fidelity-per-round of amplification was found to be approximately 98% in trial PCRs with this six-letter DNA alphabet. The analogous PCR employing TTP had a fidelity-per-round of only approximately 93%. Thus, the A, 2-thioT, G, C, isoC, isoG alphabet is an artificial genetic system capable of Darwinian evolution.

Base Pairing↗

Resurrecting ancestral alcohol dehydrogenases from yeast.

Modern yeast living in fleshy fruits rapidly convert sugars into bulk ethanol through pyruvate. Pyruvate loses carbon dioxide to produce acetaldehyde, which is reduced by alcohol dehydrogenase 1 (Adh1) to ethanol, which accumulates. Yeast later consumes the accumulated ethanol, exploiting Adh2, an Adh1 homolog differing by 24 (of 348) amino acids. As many microorganisms cannot grow in ethanol, accumulated ethanol may help yeast defend resources in the fruit. We report here the resurrection of the last common ancestor of Adh1 and Adh2, called Adh(A). The kinetic behavior of Adh(A) suggests that the ancestor was optimized to make (not consume) ethanol. This is consistent with the hypothesis that before the Adh1-Adh2 duplication, yeast did not accumulate ethanol for later consumption but rather used Adh(A) to recycle NADH generated in the glycolytic pathway. Silent nucleotide dating suggests that the Adh1-Adh2 duplication occurred near the time of duplication of several other proteins involved in the accumulation of ethanol, possibly in the Cretaceous age when fleshy fruits arose. These results help to connect the chemical behavior of these enzymes through systems analysis to a time of global ecosystem change, a small but useful step towards a planetary systems biology.

Alcohol Dehydrogenase↗

Phylogenomic approaches to common problems encountered in the analysis of low copy repeats: the sulfotransferase 1A gene family example.

BACKGROUND: Blocks of duplicated genomic DNA sequence longer than 1000 base pairs are known as low copy repeats (LCRs). Identified by their sequence similarity, LCRs are abundant in the human genome, and are interesting because they may represent recent adaptive events, or potential future adaptive opportunities within the human lineage. Sequence analysis tools are needed, however, to decide whether these interpretations are likely, whether a particular set of LCRs represents nearly neutral drift creating junk DNA, or whether the appearance of LCRs reflects assembly error. Here we investigate an LCR family containing the sulfotransferase (SULT) 1A genes involved in drug metabolism, cancer, hormone regulation, and neurotransmitter biology as a first step for defining the problems that those tools must manage. RESULTS: Sequence analysis here identified a fourth sulfotransferase gene, which may be transcriptionally active, located on human chromosome 16. Four regions of genomic sequence containing the four human SULT1A paralogs defined a new LCR family. The stem hominoid SULT1A progenitor locus was identified by comparative genomics involving complete human and rodent genomes, and a draft chimpanzee genome. SULT1A expansion in hominoid genomes was followed by positive selection acting on specific protein sites. This episode of adaptive evolution appears to be responsible for the dopamine sulfonation function of some SULT enzymes. Each of the conclusions that this bioinformatic analysis generated using data that has uncertain reliability (such as that from the chimpanzee genome sequencing project) has been confirmed experimentally or by a "finished" chromosome 16 assembly, both of which were published after the submission of this manuscript. CONCLUSION: SULT1A genes expanded from one to four copies in hominoids during intra-chromosomal LCR duplications, including (apparently) one after the divergence of chimpanzees and humans. Thus, LCRs may provide a means for amplifying genes (and other genetic elements) that are adaptively useful. Being located on and among LCRs, however, could make the human SULT1A genes susceptible to further duplications or deletions resulting in 'genomic diseases' for some individuals. Pharmacogenomic studies of SULT1Asingle nucleotide polymorphisms, therefore, should also consider examining SULT1A copy number variability when searching for genotype-phenotype associations. The latest duplication is, however, only a substantiated hypothesis; an alternative explanation, disfavored by the majority of evidence, is that the duplication is an artifact of incorrect genome assembly.

Animals↗

Planetary systems biology.

Combining paleogenetics, protein engineering, synthetic biology, and metabolic modeling, a planetary biology perspective is brought to bear on adaptive evolutionary events in ancient bacteria.

Bacterial Physiological Phenomena↗

Synthetic biology.

Synthetic biologists come in two broad classes. One uses unnatural molecules to reproduce emergent behaviours from natural biology, with the goal of creating artificial life. The other seeks interchangeable parts from natural biology to assemble into systems that function unnaturally. Either way, a synthetic goal forces scientists to cross uncharted ground to encounter and solve problems that are not easily encountered through analysis. This drives the emergence of new paradigms in ways that analysis cannot easily do. Synthetic biology has generated diagnostic tools that improve the care of patients with infectious diseases, as well as devices that oscillate, creep and play tic-tac-toe.

Biology↗

Synthetic biology.

Chemistry is a broadly powerful discipline in contemporary science because it has the ability to create new forms of the matter that it studies. By doing so, chemistry can test models that connect molecular structure to behaviour without having to rely on what nature has provided. This creation, known as 'synthesis', began to be applied to living systems in the 1980s as recombinant DNA technologies allowed biologists to deliberately change the molecular structure of the microbes that they studied, and automated chemical synthesis of DNA became widely available to support these activities. The impact of the information that has emerged has made biologists aware of a truism that has long been known in chemistry: synthesis drives discovery and understanding in ways that analysis cannot. Synthetic biology is now setting an ambitious goal: to recreate in artificial systems the emergent properties found in natural biology. By doing so, it is advancing our understanding of the molecular basis of genetics in ways that analysis alone cannot. More practically, it has yielded artificial genetic systems that improve the healthcare of some 400,000 Americans annually. Synthetic biology is now set to take the next step, to create artificial Darwinian systems by direct construction. Supported by the National Science Foundation as part of its Chemical Bonding program, this work cannot help but generate clarity in our understanding of how biological systems work.

Animals↗

Quantitative analysis of a RNA-cleaving DNA catalyst obtained via in vitro selection.

In vitro selections performed in the presence of Mg(2+) generated DNA sequences capable of cleaving an internal ribonucleoside linkage. Several of these, surprisingly, displayed intermolecular catalysis and catalysis independent of Mg(2+), features that the selection protocol was not explicitly designed to select. A detailed physical organic analysis was applied to one of these DNAzymes, termed 614. First, the progress curve for the reaction was dissected to identify factors that prevented the molecule from displaying clean first-order transformation kinetics and 100% conversion. Several factors were identified and quantitated, including (a) competitive intra- and intermolecular rate processes, (b) alternative reactive and unreactive conformations, and (c) mutations within the catalyst. Other factors were excluded, including "approach to equilibrium" kinetics and product inhibition. The possibility of complementary strand inhibition was demonstrated but was shown to not be a factor under the conditions of these experiments. The rates of the intra- and intermolecular processes were compared, and saturation models for the intermolecular process were built. The rate-limiting step for the intermolecular reaction was found to be the association/folding of the enzyme with the substrate and not the cleavage step. The DNAzyme 614 is more active in trans than in cis and more active at temperatures below the selection temperature than at the selection temperature. Many of these properties have not been reported in similar systems; these results therefore expand the phenomenology known for this class of DNA-based catalysts. A brief survey of other catalysts arising from this selection found other Mg(2+)-independent DNAzymes and provided a preliminary view of the ruggedness of the landscape, relating function to structure in sequence space. Hypotheses are suggested to account for the fact that a selection in the presence of Mg(2+) did not exploit this Mg(2+). This study of a specific catalytically active DNAzyme is an example of studies that will be necessary generally to permit in vitro selection to help us understand the distribution of function in sequence space.

Bacteriophage lambda↗

Multiplexed genetic analysis using an expanded genetic alphabet.

BACKGROUND: All states require some kind of testing for newborns, but the policies are far from standardized. In some states, newborn screening may include genetic tests for a wide range of targets, but the costs and complexities of the newer genetic tests inhibit expansion of newborn screening. We describe the development and technical evaluation of a multiplex platform that may foster increased newborn genetic screening. METHODS: MultiCode PLx involves three major steps: PCR, target-specific extension, and liquid chip decoding. Each step is performed in the same reaction vessel, and the test is completed in approximately 3 h. For site-specific labeling and room-temperature decoding, we use an additional base pair constructed from isoguanosine and isocytidine. We used the method to test for mutations within the cystic fibrosis transmembrane conductance regulator (CFTR) gene. The developed test was performed manually and by automated liquid handling. Initially, 225 samples with a range of genotypes were tested retrospectively with the method. A prospective study used samples from >400 newborns. RESULTS: In the retrospective study, 99.1% of samples were correctly genotyped with no incorrect calls made. In the perspective study, 95% of the samples were correctly genotyped for all targets, and there were no incorrect calls. CONCLUSIONS: The unique genetic multiplexing platform was successfully able to test for 31 targets within the CFTR gene and provides accurate genotype assignments in a clinical setting.

Autoanalysis↗