PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Purifying selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

Intracellular battlegrounds: conflict and cooperation between transposable elements.

Transposable elements (TEs) are genomic parasites that amplify their own representation on hosts' chromosomes by inserting into new positions. It is traditionally thought that their copy number is regulated by purifying selection that eliminates hosts with higher than average TE abundance. Here, we stress that selection due to beneficial or harmful interactions between TEs introduces a whole new dimension, with implications for TE evolutionary trajectories and TE loads on hosts. This framework poses new questions requiring conceptual and experimental advances. Considering primarily Drosophila data, we make a case for within host selection on TEs by thinking expansively about the lifecycle of several TE families.

Animals↗

One-pot glyco-affinity precipitation purification for enhanced proteomics: the flexible alignment of solution-phase capture/release and solid-phase separation.

A one-pot affinity precipitation purification of carbohydrate-binding protein was demonstrated by designing thermally responsive glyco-polypeptide polymers, which were synthesized by selective coupling of pendant carbohydrate groups to a recombinant elastin-like triblock protein copolymer (ELP). The thermally driven inverse transition temperature of the ELP-based triblock polymer is maintained upon incorporation of carbohydrate ligands, which was confirmed by differential scanning calorimetry and (1)H NMR spectroscopy experiments. As a test system, lactose derivatized ELP was used to selectively purify a galactose-specific binding lectin through simple temperature-triggered precipitation in a high level of efficiency. Potential opportunities might be provided for enhanced proteomic, cell isolation as well as pathogen detection applications.

Amino Acid Sequence↗

Sequence analysis of the polymerase domain of HIV-1 reverse transcriptase in naive and zidovudine-treated individuals reveals a higher polymorphism in alpha-helices as compared with beta-strands.

We report a statistical analysis of genetic heterogeneity of the reverse transcriptase (RT)-coding region of human immunodeficiency virus type 1. Both newly determined sequences and sequences contained in the data banks have been examined. For the calculations, the viral samples and the regions within the RT molecule were divided in two groups. The viral samples were split into those from patients not subjected to antiretroviral therapy and those from patients treated with zidovudine (AZT, 3'-azido-3'-deoxythymidine) alone or in combination with other RT inhibitors. The RT-coding region was divided into segments encoding beta-strands and segments encoding alpha-helices. A significantly lower heterogeneity was observed in beta-strands relative to the alpha-helix coding segments. Application of the D test of Tajima has provided evidence of operation of negative (or purifying) selection in sequences from viruses of patients not subjected to antiretroviral treatment as well as in treated patients. In the group of untreated individuals, regions encoding beta-strands are subjected to stronger negative selection than those encoding alpha-helices. It is likely that the observed differences reflect stronger functional constraints in beta-strands than in alpha-helices of RT.

Adolescent↗

Protein dispensability and rate of evolution.

If protein evolution is due in large part to slightly deleterious amino acid substitutions, then the rate of evolution should be greater in proteins that contribute less to individual fitness. The rationale for this prediction is that relatively dispensable proteins should be subject to weaker purifying selection, and should therefore accumulate mildly deleterious substitutions more rapidly. Although this argument was presented over twenty years ago, and is fundamental to many applications of evolutionary theory, the prediction has proved difficult to confirm. In fact, a recent study showed that essential mouse genes do not evolve more slowly than non-essential ones. Thus, although a variety of factors influencing the rate of protein evolution have been supported by extensive sequence analysis, the relationship between protein dispensability and evolutionary rate has remained unconfirmed. Here we use the results from a highly parallel growth assay of single gene deletions in yeast to assess protein dispensability, which we relate to evolutionary rate estimates that are based on comparisons of sequences drawn from twenty-one fully annotated genomes. Our analysis reveals a highly significant relationship between protein dispensability and evolutionary rate, and explains why this relationship is not detectable by categorical comparison of essential versus non-essential proteins. The relationship is highly conserved, so that protein dispensability in yeast is also predictive of evolutionary rate in a nematode worm.

Amino Acid Substitution↗

High intrinsic rate of DNA loss in Drosophila.

Pseudogenes are common in mammals but virtually absent in Drosophila. All putative Drosophila pseudogenes show patterns of molecular evolution that are inconsistent with the lack of functional constraints. The absence of bona fide pseudogenes is not only puzzling, it also hampers attempts to estimate rates and patterns of neutral DNA change. The estimation problem is especially acute in the case of deletions and insertions, which are likely to have large effects when they occur in functional genes and are therefore subject to strong purifying selection. We propose a solution to this problem by taking advantage of the propensity of retrotransposable elements without long terminal repeats (non-LTR) to create non-functional, 'dead-on-arrival' copies of themselves as a common by-product of their transpositional cycle. Phylogenetic analysis of a non-LTR element, Helena, demonstrates that copies lose DNA at an unusually high rate, suggesting that lack of pseudogenes in Drosophila is the product of rampant deletion of DNA in unconstrained regions. This finding has important implications for the study of genome evolution in general and the 'C-value paradox' in particular.

Animals↗

Conservation of Y-linked genes during human evolution revealed by comparative sequencing in chimpanzee.

The human Y chromosome, transmitted clonally through males, contains far fewer genes than the sexually recombining autosome from which it evolved. The enormity of this evolutionary decline has led to predictions that the Y chromosome will be completely bereft of functional genes within ten million years. Although recent evidence of gene conversion within massive Y-linked palindromes runs counter to this hypothesis, most unique Y-linked genes are not situated in palindromes and have no gene conversion partners. The 'impending demise' hypothesis thus rests on understanding the degree of conservation of these genes. Here we find, by systematically comparing the DNA sequences of unique, Y-linked genes in chimpanzee and human, which diverged about six million years ago, evidence that in the human lineage, all such genes were conserved through purifying selection. In the chimpanzee lineage, by contrast, several genes have sustained inactivating mutations. Gene decay in the chimpanzee lineage might be a consequence of positive selection focused elsewhere on the Y chromosome and driven by sperm competition.

Animals↗

Robustness-epistasis link shapes the fitness landscape of a randomly drifting protein.

The distribution of fitness effects of protein mutations is still unknown. Of particular interest is whether accumulating deleterious mutations interact, and how the resulting epistatic effects shape the protein's fitness landscape. Here we apply a model system in which bacterial fitness correlates with the enzymatic activity of TEM-1 beta-lactamase (antibiotic degradation). Subjecting TEM-1 to random mutational drift and purifying selection (to purge deleterious mutations) produced changes in its fitness landscape indicative of negative epistasis; that is, the combined deleterious effects of mutations were, on average, larger than expected from the multiplication of their individual effects. As observed in computational systems, negative epistasis was tightly associated with higher tolerance to mutations (robustness). Thus, under a low selection pressure, a large fraction of mutations was initially tolerated (high robustness), but as mutations accumulated, their fitness toll increased, resulting in the observed negative epistasis. These findings, supported by FoldX stability computations of the mutational effects, prompt a new model in which the mutational robustness (or neutrality) observed in proteins, and other biological systems, is due primarily to a stability margin, or threshold, that buffers the deleterious physico-chemical effects of mutations on fitness. Threshold robustness is inherently epistatic-once the stability threshold is exhausted, the deleterious effects of mutations become fully pronounced, thereby making proteins far less robust than generally assumed.

Epistasis, Genetic↗

Human T cell epitopes of Mycobacterium tuberculosis are evolutionarily hyperconserved.

Mycobacterium tuberculosis is an obligate human pathogen capable of persisting in individual hosts for decades. We sequenced the genomes of 21 strains representative of the global diversity and six major lineages of the M. tuberculosis complex (MTBC) at 40- to 90-fold coverage using Illumina next-generation DNA sequencing. We constructed a genome-wide phylogeny based on these genome sequences. Comparative analyses of the sequences showed, as expected, that essential genes in MTBC were more evolutionarily conserved than nonessential genes. Notably, however, most of the 491 experimentally confirmed human T cell epitopes showed little sequence variation and had a lower ratio of nonsynonymous to synonymous changes than seen in essential and nonessential genes. We confirmed these findings in an additional data set consisting of 16 antigens in 99 MTBC strains. These findings are consistent with strong purifying selection acting on these epitopes, implying that MTBC might benefit from recognition by human T cells.

Antigens, Bacterial↗

A high-resolution survey of deletion polymorphism in the human genome.

Recent work has shown that copy number polymorphism is an important class of genetic variation in human genomes. Here we report a new method that uses SNP genotype data from parent-offspring trios to identify polymorphic deletions. We applied this method to data from the International HapMap Project to produce the first high-resolution population surveys of deletion polymorphism. Approximately 100 of these deletions have been experimentally validated using comparative genome hybridization on tiling-resolution oligonucleotide microarrays. Our analysis identifies a total of 586 distinct regions that harbor deletion polymorphisms in one or more of the families. Notably, we estimate that typical individuals are hemizygous for roughly 30-50 deletions larger than 5 kb, totaling around 550-750 kb of euchromatic sequence across their genomes. The detected deletions span a total of 267 known and predicted genes. Overall, however, the deleted regions are relatively gene-poor, consistent with the action of purifying selection against deletions. Deletion polymorphisms may well have an important role in the genetics of complex traits; however, they are not directly observed in most current gene mapping studies. Our new method will permit the identification of deletion polymorphisms in high-density SNP surveys of trio or other family data.

Databases, Genetic↗

Pneumococcal within-host diversity during colonization, transmission and treatment.

Characterizing the genetic diversity of pathogens within the host promises to greatly improve surveillance and reconstruction of transmission chains. For bacteria, it also informs our understanding of inter-strain competition and how this shapes the distribution of resistant and sensitive bacteria. Here we study the genetic diversity of Streptococcus pneumoniae within 468 infants and 145 of their mothers by deep sequencing whole pneumococcal populations from 3,761 longitudinal nasopharyngeal samples. We demonstrate that deep sequencing has unsurpassed sensitivity for detecting multiple colonization, doubling the rate at which highly invasive serotype 1 bacteria were detected in carriage compared with gold-standard methods. The greater resolution identified an elevated rate of transmission from mothers to their children in the first year of the child's life. Comprehensive treatment data demonstrated that infants were at an elevated risk of both the acquisition and persistent colonization of a multidrug-resistant bacterium following antimicrobial treatment. Some alleles were enriched after antimicrobial treatment, suggesting that they aided persistence, but generally purifying selection dominated within-host evolution. Rates of co-colonization imply that in the absence of treatment, susceptible lineages outcompeted resistant lineages within the host. These results demonstrate the many benefits of deep sequencing for the genomic surveillance of bacterial pathogens.

Child↗

Chromosomal toxin-antitoxin systems in Pseudomonas putida are rather selfish than beneficial.

Chromosomal toxin-antitoxin (TA) systems are widespread genetic elements among bacteria, yet, despite extensive studies in the last decade, their biological importance remains ambivalent. The ability of TA-encoded toxins to affect stress tolerance when overexpressed supports the hypothesis of TA systems being associated with stress adaptation. However, the deletion of TA genes has usually no effects on stress tolerance, supporting the selfish elements hypothesis. Here, we aimed to evaluate the cost and benefits of chromosomal TA systems to Pseudomonas putida. We show that multiple TA systems do not confer fitness benefits to this bacterium as deletion of 13 TA loci does not influence stress tolerance, persistence or biofilm formation. Our results instead show that TA loci are costly and decrease the competitive fitness of P. putida. Still, the cost of multiple TA systems is low and detectable in certain conditions only. Construction of antitoxin deletion strains showed that only five TA systems code for toxic proteins, while other TA loci have evolved towards reduced toxicity and encode non-toxic or moderately potent proteins. Analysis of P. putida TA systems' homologs among fully sequenced Pseudomonads suggests that the TA loci have been subjected to purifying selection and that TA systems spread among bacteria by horizontal gene transfer.

Anti-Bacterial Agents↗

A summary statistic approach to sequence variation in noncoding regions of six schizophrenia-associated gene loci.

In order to explore the role of noncoding variants in the genetics of schizophrenia, we sequenced 27 kb of noncoding DNA from the gene loci RAC-alpha serine/threonine-protein kinase (AKT1), brain-derived neurotrophic factor (BDNF), dopamine receptor-3 (DRD3), dystrobrevin binding protein-1 (DTNBP1), neuregulin-1 (NRG1) and regulator of G-protein signaling-4 (RGS4) in 37 schizophrenia patients and 25 healthy controls. To compare the allele frequency spectrum between the two samples, we separately computed Tajima's D-value for each sample. The results showed a smaller Tajima's D-value in the case sample, pointing to an excess of rare variants as compared to the control sample. When randomly permuting the affection status of sequenced individuals, we observed a stronger decrease of Tajima's D in 2400 out of 100,000 permutations, corresponding to a P-value of 0.024 in a one-sided test. Thus, rare variants are significantly enriched in the schizophrenia sample, indicating the existence of disease-related sequence alterations. When categorizing the sequenced fragments according to their level of human-rodent conservation or according to their gene locus, we observed a wide range of diversity parameter estimates. Rare variants were enriched in conserved regions as compared to nonconserved regions in both samples. Nevertheless, rare variants remained more common among patients, suggesting an increased number of variants under purifying selection in this sample. Finally, we performed a heuristic search for the subset of gene loci, which jointly produces the strongest difference between controls and cases. This showed a more prominent role of variants from the loci AKT1, BDNF and RGS4. Taken together, our approach provides promising strategy to investigate the genetics of schizophrenia and related phenotypes.

Brain-Derived Neurotrophic Factor↗

Extent of mitochondrial DNA sequence variation in Atlantic cod from the Faroe Islands: a resolution of gene genealogy.

Variation in a 250 base pair (bp) fragment of the mitochondrial cytochrome b (cyt b) has been used extensively for population studies in Atlantic cod Gadus morhua. To study the shape of the gene genealogy and the nature of the polymorphism, sequences of another region of the cyt b gene and the TP intergenic spacer were added, making a total of 566 bp from 74 cod from the Faroe Islands. A total of 44 segregating sites defined 41 haplotypes, many at frequencies greater than 5%. Haplotype diversity was 0.97 and nucleotide diversity 0.73% per base. A topology referred to as a constellation gene genealogy was observed with four major haplotypes at high frequencies, from each of which a number of rare variants were derived. A young relative age of the haplotypes was gauged from the structure of the genealogy. The variation was mostly at synonymous sites within the coding region and thus likely to be neutral or under weak purifying selection. By comparative analysis this also applies to the TP spacer. Applying the locus to study population variation in the Faroe Islands by AMOVA revealed that the overall areas and localities within areas accounted for none of the variation, and all the variation was due to differences among individuals.

Alleles↗

Evolution of the acyl-CoA binding protein (ACBP).

Acyl-CoA-binding protein (ACBP) is a 10 kDa protein that binds C12-C22 acyl-CoA esters with high affinity. In vitro and in vivo experiments suggest that it is involved in multiple cellular tasks including modulation of fatty acid biosynthesis, enzyme regulation, regulation of the intracellular acyl-CoA pool size, donation of acyl-CoA esters for beta-oxidation, vesicular trafficking, complex lipid synthesis and gene regulation. In the present study, we delineate the evolutionary history of ACBP to get a complete picture of its evolution and distribution among species. ACBP homologues were identified in all four eukaryotic kingdoms, Animalia, Plantae, Fungi and Protista, and eleven eubacterial species. ACBP homologues were not detected in any other known bacterial species, or in archaea. Nearly all of the ACBP-containing bacteria are pathogenic to plants or animals, suggesting that an ACBP gene could have been acquired from a eukaryotic host by horizontal gene transfer. Many bacterial, fungal and higher eukaryotic species only harbour a single ACBP homologue. However, a number of species, ranging from protozoa to vertebrates, have evolved two to six lineage-specific paralogues through gene duplication and/or retrotransposition events. The ACBP protein is highly conserved across phylums, and the majority of ACBP genes are subjected to strong purifying selection. Experimental evidence indicates that the function of ACBP has been conserved from yeast to humans and that the multiple lineage-specific paralogues have evolved altered functions. The appearance of ACBP very early on in evolution points towards a fundamental role of ACBP in acyl-CoA metabolism, including ceramide synthesis and in signalling.

Acyl Coenzyme A↗

The 'rare allele phenomenon' in a ribosomal spacer.

We describe the increased frequency of a particular length variant of the internal transcribed spacer 1 (ITS-1) of the ribosomal DNA in a hybrid zone of the land snail Albinaria hippolyti. The phenomenon that normally rare alleles or other markers can increase in frequency in the centre of hybrid zones is not new. Under the term 'hybrizyme' or 'rare allele' phenomenon it has been recorded in many organisms and different genetic markers. However, this is the first time that it has been found in a multicopy locus. On the one hand, the pattern fits well with the view that purifying selection in hybrid populations works on many loci across the genome and should thus have its effect on many independent molecular markers. On the other hand, the results are puzzling, given that the multiple copies of rDNA are not expected to respond in unison. We suggest two possible explanations for these conflicting observations.

Alleles↗

Pangenome-wide identification and expression analysis of the chalcone synthase (CHS) gene family in five yellowhorn spp.

Chalcone synthase (CHS) is a pivotal enzyme in flavonoid biosynthesis involved in plant development, defense, and secondary metabolism. Xanthoceras sorbifolium (yellowhorn) is a medicinal and ornamental species with high resistance to environmental stresses, but its CHS gene family remains uncharacterized. We performed a pangenome-wide identification of CHS genes across five yellowhorn genomes (Xzs4, Xwf8, Xjg, Xg11, and Xzg2). Across the five yellowhorn genomes, 27 CHS genes were identified and classified into four core pangenes, present in all five genomes, and two dispensable genes, present only in a subset of genomes. Phylogenetic analysis grouped these genes into three major clades, and chromosomal mapping and duplication analyses identified four tandemly duplicated gene pairs under purifying selection. The analyses of conserved structural features, including protein motifs and exon-intron organization, together with promoter cis-regulatory elements and gene ontology annotation, further indicated the potential involvement of CHS genes in flavonoid biosynthesis and stress-responsive mechanisms. Gene expression profiling identified significant upregulation of Xg11_CHS1 and Xg11_CHS3 under cold and drought stress, with tissue-specific expression patterns. These findings provide valuable insights into the evolution, functional diversification, and stress-responsive roles of the CHS gene family, identifying candidate genes for future studies targeting stress tolerance and flavonoid biosynthesis in yellowhorn.

Acyltransferases↗

Comparisons of pollen coat genes across Brassicaceae species reveal rapid evolution by repeat expansion and diversification.

Reproductive genes and traits evolve rapidly in many organisms, including mollusks, algae, and primates. Previously we demonstrated that a family of glycine-rich pollen surface proteins (GRPs) from Arabidopsis thaliana and Brassica oleracea had diverged substantially, making identification of homologous genes impossible despite a separation of only 20 million years. Here we address the molecular genetic mechanisms behind these changes, sequencing the eight members of the GRP cluster, along with 11 neighboring genes in four related species, Arabidopsis arenosa, Olimarabidopsis pumila, Capsella rubella, and Sisymbrium irio. We found that GRP genes change more rapidly than their neighbors; they are more repetitive and have undergone substantially more insertion/deletion events while preserving repeat amino acid composition. Genes flanking the GRP cluster had an average K(a)/K(s) approximately 0.2, indicating strong purifying selection. This ratio rose to approximately 0.5 in the first GRP exon, indicating relaxed selective constraints. The repetitive nature of the second GRP exon makes alignment difficult; even so, K(a)/K(s) within the Arabidopsis genus demonstrated an increase that correlated with exon length. We conclude that rapid GRP evolution is primarily due to duplication, deletion, and divergence of repetitive sequences. GRPs may mediate pollen recognition and hydration by female cells, and divergence of these genes could correlate with or even promote speciation. We tested cross-species interactions, showing that the ability of A. arenosa stigmas to hydrate pollen correlated with GRP divergence and identifying A. arenosa as a model for future studies of pollen recognition.

Amino Acid Sequence↗

Long-term reinfection of the human genome by endogenous retroviruses.

Endogenous retrovirus (ERV) families are derived from their exogenous counterparts by means of a process of germ-line infection and proliferation within the host genome. Several families in the human and mouse genomes now consist of many hundreds of elements and, although several candidates have been proposed, the mechanism behind this proliferation has remained uncertain. To investigate this mechanism, we reconstructed the ratio of nonsynonymous to synonymous changes and the acquisition of stop codons during the evolution of the human ERV family HERV-K(HML2). We show that all genes, including the env gene, which is necessary only for movement between cells, have been under continuous purifying selection. This finding strongly suggests that the proliferation of this family has been almost entirely due to germ-line reinfection, rather than retrotransposition in cis or complementation in trans, and that an infectious pool of endogenous retroviruses has persisted within the primate lineage throughout the past 30 million years. Because many elements within this pool would have been unfixed, it is possible that the HERV-K(HML2) family still contains infectious elements at present, despite their apparent absence in the human genome sequence. Analysis of the env gene of eight other HERV families indicated that reinfection is likely to be the most common mechanism by which endogenous retroviruses proliferate in their hosts.

Endogenous Retroviruses↗