PubMed Health⌕ Search

Biomedical subjects

G Brian Golding

Publications and source records attributed to G Brian Golding.

18 recordsLinked to original sources

Gene gain and gene loss in streptococcus: is it driven by habitat?

Bacterial genomes can evolve either by gene gain, gene loss, mutating existing genes, and/or by duplication of existing genes. Recent studies have clearly demonstrated that the acquisition of new genes by lateral gene transfer (LGT) is a predominant force in bacterial evolution. To better understand the significance of LGT, we employed a comparative genomics approach to model species-specific and intraspecies gene insertions/deletions (ins/del among 12 sequenced streptococcal genomes using a maximum likelihood method. This study indicates that the rate of gene ins/del is higher on the external branches and varies dramatically for each species. We have analyzed here some of the experimentally characterized species-specific genes that have been acquired by LGT and conclude that at least a portion of these genes have a role in adaptation.

Adaptation, Biological↗

An integrated approach to functional genomics: construction of a novel reporter gene fusion library for Sinorhizobium meliloti.

As a means of investigating gene function, we developed a robust transcription fusion reporter vector to measure gene expression in bacteria. The vector, pTH1522, was used to construct a random insert library for the Sinorhizobium meliloti genome. pTH1522 replicates in Escherichia coli and can be transferred to, but cannot replicate in, S. meliloti. Homologous recombination of the DNA fragments cloned in pTH1522 into the S. meliloti genome generates transcriptional fusions to either the reporter genes gfp(+) and lacZ or gusA and rfp, depending on the orientation of the cloned fragment. Over 12,000 fusion junctions in 6,298 clones were identified by DNA sequence analysis, and the plasmid clones were recombined into S. meliloti. Reporter enzyme activities following growth of these recombinants in complex medium (LBmc) and in minimal medium with glucose or succinate as the sole carbon source allowed the identification of genes highly expressed under one or more growth condition and those expressed at very low to background levels. In addition to generating reporter gene fusions, the vector allows Flp recombinase-directed deletion formation and gene disruption, depending on the nature of the cloned fragment. We report the identification of genes essential for growth on complex medium as deduced from an inability to recover recombinants from pTH1522 clones that carried fragments internal to gene or operon transcripts. A database containing all the gene expression activities together with a web interface showing the precise locations of reporter fusion junctions has been constructed (www.sinorhizobium.org).

Bacterial Proteins↗

Selection and slippage creating serine homopolymers.

Highly repetitive sequence within proteins is an abundant feature yet is considered by some to be the protein equivalent of "junk DNA." Homopolymer sequences, the most highly repetitive of this group, are typically encoded by trinucleotide repeats at the DNA level. It is thought that many of these sequences are produced by a replicative slippage mechanism. Recent studies suggest that these highly mutable regions within proteins may allow for rapid morphological evolution emerging from the increased variability afforded by such coding structures. However, in a homopolymer, it is difficult to determine if the repeated amino acid is due to slippage at the DNA level or due to selection at the protein level. Here we develop and test a model to detect cases for which the homopolymer tract has clearly been selected for, with no evidence of slippage at the DNA level. The polyserine tract within the phosphatidylserine receptor protein is used as an excellent example of one such case.

Amino Acid Sequence↗

Transcriptional profiling implicates novel interactions between abiotic stress and hormonal responses in Thellungiella, a close relative of Arabidopsis.

Thellungiella, an Arabidopsis (Arabidopsis thaliana)-related halophyte, is an emerging model species for studies designed to elucidate molecular mechanisms of abiotic stress tolerance. Using a cDNA microarray containing 3,628 unique sequences derived from previously described libraries of stress-induced cDNAs of the Yukon ecotype of Thellungiella salsuginea, we obtained transcript profiles of its response to cold, salinity, simulated drought, and rewatering after simulated drought. A total of 154 transcripts were differentially regulated under the conditions studied. Only six of these genes responded to all three stresses of drought, cold, and salinity, indicating a divergence among the end responses triggered by each of these stresses. Unlike in Arabidopsis, there were relatively few transcript changes in response to high salinity in this halophyte. Furthermore, the gene products represented among drought-responsive transcripts in Thellungiella associate a down-regulation of defense-related transcripts with exposure to water deficits. This antagonistic interaction between drought and biotic stress response may demonstrate Thellungiella's ability to respond precisely to environmental stresses, thereby conserving energy and resources and maximizing its survival potential. Intriguingly, changes of transcript abundance in response to cold implicate the involvement of jasmonic acid. While transcripts associated with photosynthetic processes were repressed by cold, physiological responses in plants developed at low temperature suggest a novel mechanism for photosynthetic acclimation. Taken together, our results provide useful starting points for more in-depth analyses of Thellungiella's extreme stress tolerance.

Arabidopsis↗

Asymmetrical evolution of cytochrome bd subunits.

Functionally linked genes generally evolve at similar rates and the knowledge of this particular feature of genomic evolution has been used as the basis for the phylogenetic profiling method. We illustrate here an exception to this rule in the evolution of the cytochrome bd complex. This is a two-component oxidase complex, with the subunits I and II known to be widely present in bacteria. The subunits within the cytochrome bd complex are under the same evolutionary pressure and most likely behave in the same evolutionary manner. However, the sequence similarity of genes encoding subunit II varies considerably across species. Genes encoding subunit II evolve 1.2 times faster on most of the branches of their phylogeny than subunit I genes. Furthermore, the genes encoding subunit II in Oceanobacillus iheyensis, Bacillus halodurans, and Staphylococcus species do not have detectable homologues within E. coli due to their large divergence. Together, the two subunits of cytochrome bd reveal an interesting example of an asymmetric pattern of evolutionary change.

Bacteria↗

The fate of laterally transferred genes: life in the fast lane to adaptation or death.

Large-scale genome arrangement plays an important role in bacterial genome evolution. A substantial number of genes can be inserted into, deleted from, or rearranged within genomes during evolution. Detecting or inferring gene insertions/deletions is of interest because such information provides insights into bacterial genome evolution and speciation. However, efficient inference of genome events is difficult because genome comparisons alone do not generally supply enough information to distinguish insertions, deletions, and other rearrangements. In this study, homologous genes from the complete genomes of 13 closely related bacteria were examined. The presence or absence of genes from each genome was cataloged, and a maximum likelihood method was used to infer insertion/deletion rates according to the phylogenetic history of the taxa. It was found that whole gene insertions/deletions in genomes occur at rates comparable to or greater than the rate of nucleotide substitution and that higher insertion/deletion rates are often inferred to be present at the tips of the phylogeny with lower rates on more ancient interior branches. Recently transferred genes are under faster and relaxed evolution compared with more ancient genes. Together, this implies that many of the lineage-specific insertions are lost quickly during evolution and that perhaps a few of the genes inserted by lateral transfer are niche specific.

Adaptation, Physiological↗

Lateral gene transfer in Mycobacterium avium subspecies paratuberculosis.

Lateral gene transfer is an integral part of genome evolution in most bacteria. Bacteria can readily change the contents of their genomes to increase adaptability to ever-changing surroundings and to generate evolutionary novelty. Here, we report instances of lateral gene transfer in Mycobacterium avium subsp. paratuberculosis, a pathogenic bacteria that causes Johne's disease in cattle. A set of 275 genes are identified that are likely to have been recently acquired by lateral gene transfer. The analysis indicated that 53 of the 275 genes were acquired after the divergence of M. avium subsp. paratuberculosis from M. avium subsp. avium, whereas the remaining 222 genes were possibly acquired by a common ancestor of M. avium subsp. paratuberculosis and M. avium subsp. avium after its divergence from the ancestor of M. tuberculosis complex. Many of the acquired genes were from proteobacteria or soil dwelling actinobacteria. Prominent among the predicted laterally transferred genes is the gene rsbR, a possible regulator of sigma factor, and the genes designated MAP3614 and MAP3757, which are similar to genes in eukaryotes. The results of this study suggest that like most other bacteria, lateral gene transfers seem to be a common feature in M. avium subsp. paratuberculosis and that the proteobacteria contribute most of these genetic exchanges.

Gene Transfer, Horizontal↗

The selective cause of an ancient adaptation.

Phylogenetic analysis reveals that the use of nicotinamide adenine dinucleotide phosphate (NADP) by prokaryotic isocitrate dehydrogenase (IDH) arose around the time eukaryotic mitochondria first appeared, about 3.5 billion years ago. We replaced the wild-type gene that encodes the NADP-dependent IDH of Escherichia coli with an engineered gene that possesses the ancestral NAD-dependent phenotype. The engineered enzyme is disfavored during competition for acetate. The selection intensifies in genetic backgrounds where other sources of reduced NADP have been removed. A survey of sequenced prokaryotic genomes reveals that those genomes that encode isocitrate lyase, which is essential for growth on acetate, always have an NADP-dependent IDH. Those with only an NAD-dependent IDH never have isocitrate lyase. Hence, the NADP dependence of prokaryotic IDH is an ancient adaptation to anabolic demand for reduced NADP during growth on acetate.

3-Isopropylmalate Dehydrogenase↗

Simple sequence in brain and nervous system specific proteins.

We examined sequences expressed in the brain and nervous system using EST data. A previous study including sequences thought to have neurological function found a deficiency of simple sequence within such sequences. This was despite many examples of neurodegenerative diseases, such as Huntington disease, which are thought to be caused by expansions of polyglutamine tracts within associated protein sequences. It may be that many of the sequences thought to have neurological function have other additional, non-neurological roles. For this reason, we examined sequences with specific expression in the brain and nervous system, using EST expression data to determine if they too are deficient of simple, repetitive sequences. Indeed, we find this class of sequences to be deficient. Unexpectedly, however, we find sequences expressed in the brain and nervous system to be consistently enriched for histidine-enriched simple sequence. Determining the function of these histidine-rich regions within brain-specific proteins requires more experimental data.

Amino Acid Sequence↗

A non-long terminal repeat retrotransposon family is restricted to the germ line micronucleus of the ciliated protozoan Tetrahymena thermophila.

The ciliated protozoan Tetrahymena thermophila undergoes extensive programmed DNA rearrangements during the development of a somatic macronucleus from the germ line micronucleus in its sexual cycle. To investigate the relationship between programmed DNA rearrangements and transposable elements, we identified several members of a family of non-long terminal repeat (LTR) retrotransposons (retroposons) in T. thermophila, the first characterized in the ciliated protozoa. This multiple-copy retrotransposon family is restricted to the micronucleus of T. thermophila. The REP (Tetrahymena non-LTR retroposon) elements encode an ORF2 typical of non-LTR elements that contains apurinic/apyrimidinic endonuclease (APE) and reverse transcriptase (RT) domains. Phylogenetic analysis of the RT and APE domains indicates that the element forms a deep-branching clade within the non-LTR retrotransposon family. Northern analysis with a probe to the conserved RT domain indicates that transcripts from the element are small and heterogeneous in length during early macronuclear development. The presence of a repeated transposable element in the genome is consistent with the model that programmed DNA deletion in T. thermophila evolved as a method of eliminating deleterious transposons from the somatic macronucleus.

3' Untranslated Regions↗

Neurological proteins are not enriched for repetitive sequences.

Proteins associated with disease and development of the nervous system are thought to contain repetitive, simple sequences. However, genome-wide surveys for simple sequences within proteins have revealed that repetitive peptide sequences are the most frequent shared peptide segments among eukaryotic proteins, including those of Saccharomyces cerevisiae, which has few to no specialized developmental and neurological proteins. It is therefore of interest to determine if these specialized proteins have an excess of simple sequences when compared to other sets of compositionally similar proteins. We have determined the relative abundance of simple sequences within neurological proteins and find no excess of repetitive simple sequence within this class. In fact, polyglutamine repeats that are associated with many neurodegenerative diseases are no more abundant within neurological specialized proteins than within nonneurological collections of proteins. We also examined the codon composition of serine homopolymers to determine what forces may play a role in the evolution of extended homopolymers. Codon type homogeneity tends to be favored, suggesting replicative slippage instead of selection as the main force responsible for producing these homopolymers.

Alanine↗

DNA and the revolutions of molecular evolution, computational biology, and bioinformatics.

The discovery of the structure of DNA was a necessary prerequisite for determining the sequence of DNA molecules. Technological advances have now made it possible to sequence DNA rapidly and has resulted in public databases with over 30 billion nucleotides of known sequence. The analysis of these data has lead to new fields of science and to amazing advances in our understanding of evolution.

Animals↗

A phylogenetic analysis of the pSymB replicon from the Sinorhizobium meliloti genome reveals a complex evolutionary history.

Microbial genomes are thought to be mosaic, making it difficult to decipher how these genomes have evolved. Whole-genome nearest-neighbor analysis was applied to the Sinorhizobium meliloti pSymB replicon to determine its origin, the degree of horizontal transfer, and the conservation of gene order. Prediction of the nearest neighbor based on contextual information, i.e., the nearest phylogenetic neighbor of adjacent genes, provided useful information for genes for which phylogenetic relationships could not be established. A large portion of pSymB genes are most closely related to genes in the Agrobacterium tumefaciens linear chromosome, including the rep and min genes. This suggests a common origin for these replicons. Genes with the nearest neighbor from the same species tend to be grouped in "patches". Gene order within these patches is conserved, but the content of the patches is not limited to operons. These data show that 13% of pSymB genes have nearest neighbors in species that are not members of the Rhizobiaceae family (including two archaea), and that these likely represent genes that have been involved in horizontal transfer.

Agrobacterium tumefaciens↗

Dinucleotide compositional analysis of Sinorhizobium meliloti using the genome signature: distinguishing chromosomes and plasmids.

The symbiotic N(2)-fixing alpha-proteobacterium Sinorhizobium meliloti has three replicons: a circular chromosome (3.7 Mb) and two smaller replicons, pSymA (1.4 Mb) and pSymB (1.7 Mb). Sequence analysis has revealed that an essential gene is carried on pSymB, which brings into question whether pSymB should be considered a chromosome or a plasmid. Based on the criterion that essential genes define a chromosome, several species have been shown to have multiple chromosomes. Many of these species are part of the alpha subdivision of the Proteobacteria family. Here, additional justification is presented for designating the pSymB replicon as a chromosome. It is shown that chromosomes within a species share a more similar dinucleotide composition, or genome signature, than plasmids do with the host chromosome(s). Dinucleotide signatures were determined for each of the S. meliloti replicons, and, consistent with the suggestion that pSymB is a chromosome, it is shown that the pSymB signature more closely resembles that of the S. meliloti chromosome, while the pSymA signature is typical of other alpha-proteobacterial plasmids.

Genome, Bacterial↗

Simple sequences are rare in the Protein Data Bank.

A simple sequence is abundant in the proteins that have been sequenced to date. But unusual protein features, such as a simple sequence, are not present in the same high frequency within structural databases. A subset of these simple sequences, a group with a highly repetitive nature has been shown to be abundant in eukaryotes but not in prokaryotes. In this study, an examination of the eukaryotic proteins in the Protein Data Bank (PDB) has revealed a large deficiency of low complexity, highly repetitive protein repeats. Through simulated databases of similar samples of eukaryotic proteins taken from the National Center for Biotechnology Information (NCBI) database, it is shown that the PDB contains a significantly less highly repetitive, simple sequence than artificial databases of similar composition randomly derived from NCBI. When the structural data for those few PDB sequences that did contain a highly repetitive simple sequence is examined in detail, it is found that in most cases the tertiary structure is unknown for the regions consisting of a simple sequence. This lack of a simple sequence both in the PDB database and in the structural information suggests that this type of simple sequence may produce disordered structures that make structural characterization difficult.

Amino Acid Sequence↗

Reconstructing the prior probabilities of allelic phylogenies.

In general when a phylogeny is reconstructed from DNA or protein sequence data, it makes use only of the probabilities of obtaining some phylogeny given a collection of data. It is also possible to determine the prior probabilities of different phylogenies. This information can be of use in analyzing the biological causes for the observed divergence of sampled taxa. Unusually "rare" topologies for a given data set may be indicative of different biological forces acting. A recursive algorithm is presented that calculates the prior probabilities of a phylogeny for different allelic samples and for different phylogenies. This method is a straightforward extension of Ewens' sample distribution. The probability of obtaining each possible sample according to Ewens' distribution is further subdivided into each of the possible phylogenetic topologies. These probabilities depend not only on the identity of the alleles and on 4N(mu) (four times the effective population size times the neutral mutation rate) but also on the phylogenetic relationships among the alleles. Illustrations of the algorithm are given to demonstrate how different phylogenies are favored under different conditions.

Algorithms↗

The pattern of amino acid replacements in alpha/beta-barrels.

The determinants of site-to-site variability in the rate of amino acid replacement in alpha/beta-barrel enzyme structures are investigated. Of 125 available alpha/beta-barrel structures, only 25 meet a variety of phylogenetic and statistical criteria necessary to ensure sufficient data for reliable analysis. These 25 enzyme structures (from a wide variety of taxa with diverse lifestyles in diverse habitats) differ greatly in size, number, and topology of domains in addition to the alpha/beta-barrel, quaternary structure, metabolic role, reaction catalyzed, presence of prosthetic groups, regulatory mechanisms, use of cofactors, and catalytic mechanisms. Yet, with the exception of ribulose-1,5-bisphosphate carboxylase, all structures have similar frequency distributions of amino acid replacement rates. Hence, site-specific variability in rates of evolution is largely independent of differences in biology, biochemistry, and molecular structure. A correlation between site-specific rate variation and (1) distance from the active site, (2) solvent accessibility, and (3) treating glycines in unusual main-chain conformations as a separate class, explains approximately half the causal variation. Secondary structure exerts little influence on the pattern and distribution of replacements. Additional domains and subunits, side-chain hydrogen bonds, unusual side-chain rotamers, nonplanar peptide bonds, strained main-chain conformations, and buried hydrophilic-charged residues contribute little to variability among sites because they are rare. Nonlinear models do not improve the fits. In several enzymes, deviations from the typical pattern of replacements suggest the possible action of natural selection. A statistical analysis shows that, in all cases, much of the remaining unexplained variation is not attributable to chance and that other, as yet unidentified, causal relations must exist.

Amino Acid Sequence↗