PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Purifying selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Genetic differences between human immunodeficiency virus type 1 subpopulations in faeces and serum.

To study human immunodeficiency virus type 1 (HIV-1) compartmentalization between intestine and blood, paired faecal and serum samples were collected from 204 HIV-1-infected persons. Direct sequencing of the gp120 V3 region obtained from 33 persons showed that faecal and serum sequences could be nearly homologous (0.3% different) or very dissimilar (11.3% different). Individual clones were obtained and sequenced from the faecal and serum samples of 13 persons. In 6 persons the HIV-1 subpopulations in faeces and serum were similar, whereas in 7 persons, distribution of V3 genotypes showed a marked difference. Genetic characterization of the HIV-1 subpopulations showed less heterogeneity in faecal subpopulations than in serum subpopulations in 5 of the 7 subjects. Furthermore, faecal and serum subpopulations differed predominantly by nonsynonymous nucleotide substitutions (in 6 of 7 persons). Comparison of the HIV-1 subpopulations in faeces and serum of these 7 persons, using resampling techniques, revealed a significant difference between faecal and serum subpopulations at an N-linked glycosylation site, C-terminal of the V3 loop (amino acids 331-333). Sequences from faecal subpopulations of all 7 persons contained a glycosylation site at amino acid position 331-333. Four of these 7 harboured serum variants lacking a glycosylation site at this position. The faecal subpopulations in these 4 persons showed limited nonsynonymous substitutions compared to synonymous substitutions, indicating that purifying selection is operational on these subpopulations.

Acquired Immunodeficiency Syndrome↗

Development and Validation of a Novel LC-MS/MS Based Proteomics Method for Quantitation of Retinol Binding Protein 4 (RBP4) and Transthyretin (TTR).

Retinol binding protein 4 (RBP4), the circulating carrier of retinol, complexes with transthyretin (TTR) and is a potential biomarker of cardiometabolic disease. However, RBP4 quantitation relies on immunoassays and western blots without retinol and TTR measurement. A liquid chromatography-tandem mass spectrometry (LC-MS/MS) method for simultaneous absolute quantitation of circulating RBP4 and TTR is critical to establishing their biomarker potential. Surrogate peptides with reproducible, linear LC-MS/MS response were selected. Purified proteins were used as quantitation standards and heavy-labelled peptides as internal standards. Matrix effects were evaluated. The validated method was applied to measure inter- and intra-individual variability in RBP4 and TTR concentrations in healthy individuals and patients with diabetic kidney disease. Quantitation was linear for the clinically relevant concentration ranges of RBP4 (0.5-6 &#x3bc;M) and TTR (5.8-69 &#x3bc;M). Assay inter-day variability was <12% and precision within 5%. The inter-individual variability for RBP4 and TTR concentrations was 18-26%, while intra-individual variability was similar to assay variability. RBP4 and TTR quantitation correlated with commercially available ELISA assays. The developed LC-MS/MS method enables simultaneous absolute quantitation of RBP4 and TTR in serum and plasma that can be applied to clinical biomarker studies and stoichiometric measurements of circulating RBP4, TTR, and retinol.

Retinol binding protein 4 (RBP4)↗

The genetic control of rapid genome content divergence in Arabidopsis thaliana.

Genome evolution in eukaryotes is predominantly driven by the dynamics of repetitive sequences, which vary widely in both copy number and sequence composition. Rates of repeat evolution differ between and within species and are likely modulated by both genetics and environment. To uncover factors shaping the rate of genome content evolution, we analyzed 1,142 resequenced Arabidopsis thaliana genomes using a novel K-mer based approach to characterize genome content variation and identify hypervariable regions underlying differences in repeat abundance. We next treated repeat abundance as a quantitative trait and performed genome-wide association analyses across more than 400 repeat families to identify the genetic basis of copy number variation. Integrating these results through a meta-GWAS approach revealed both cis-acting variants and more than 50 trans-acting loci that regulate repeat abundance genome-wide. Cis-acting variation was predominantly localized to pericentromeric and centromeric regions, whereas trans-acting loci were enriched for candidate genes involved in DNA replication, DNA repair, DNA methylation regulation. Finally, we found evidence that purifying selection acts against mutations that accelerate genome content divergence, favoring alleles that constrain repeat expansion. Together, these findings provide new insights into the genetic architecture and evolutionary forces shaping genome evolution in A. thaliana and establish a framework for investigating these processes in other plant species.

Journal Article↗

The mutation landscape of Daphnia obtusa reveals evolutionary forces shaping genome stability.

Spontaneous mutations are the primary source of genetic variation and play a central role in shaping evolutionary processes. To investigate mutational dynamics in Daphnia obtusa, we generated a chromosome-level genome assembly spanning 129.4 Mb across 12 chromosomes, encompassing 15,321 predicted protein-coding genes. Leveraging whole-genome sequencing of eight mutation accumulation (MA) lines propagated for an average of 482 generations (spanning over 20 years), we estimated a spontaneous single nucleotide mutation (SNM) rate of 2.23 &#xd7; 10-9 and an indel mutation rate of 2.75 &#xd7; 10-10 per site per generation. The SNM spectrum was strongly biased toward C:G > T:A transitions. Comparative analyses with natural population data revealed that exonic mutations observed in the MA lines were significantly less likely to be present in standing variation than intronic or intergenic mutations, suggesting that purifying selection in natural populations acts to remove deleterious alleles. We also identified 48 de novo loss-of-heterozygosity (LOH) events, comprising 8 heterozygous deletions and 40 gene conversion events. The genome-wide gene conversion rate was estimated at 2.62 &#xd7; 10-5 per heterozygous site per generation. These findings provide a comprehensive view of the mutation spectrum, selective pressures, and mechanisms underlying genome stability in D. obtusa.

Daphnia obtusa↗

Organization and evolution of a gene-rich region of the mouse genome: a 12.7-Mb region deleted in the Del(13)Svea36H mouse.

Del(13)Svea36H (Del36H) is a deletion of approximately 20% of mouse chromosome 13 showing conserved synteny with human chromosome 6p22.1-6p22.3/6p25. The human region is lost in some deletion syndromes and is the site of several disease loci. Heterozygous Del36H mice show numerous phenotypes and may model aspects of human genetic disease. We describe 12.7 Mb of finished, annotated sequence from Del36H. Del36H has a higher gene density than the draft mouse genome, reflecting high local densities of three gene families (vomeronasal receptors, serpins, and prolactins) which are greatly expanded relative to human. Transposable elements are concentrated near these gene families. We therefore suggest that their neighborhoods are gene factories, regions of frequent recombination in which gene duplication is more frequent. The gene families show different proportions of pseudogenes, likely reflecting different strengths of purifying selection and/or gene conversion. They are also associated with relatively low simple sequence concentrations, which vary across the region with a periodicity of approximately 5 Mb. Del36H contains numerous evolutionarily conserved regions (ECRs). Many lie in noncoding regions, are detectable in species as distant as Ciona intestinalis, and therefore are candidate regulatory sequences. This analysis will facilitate functional genomic analysis of Del36H and provides insights into mouse genome evolution.

Animals↗

Genome-wide regulatory complexity in yeast promoters: separation of functionally conserved and neutral sequence.

To gauge the complexity of gene regulation in yeast, it is essential to know how much promoter sequence is functional. Conservation across species can be a sensitive means of detecting functional sequences, provided that the significance of conservation can be accurately calibrated with the local neutral mutation rate. By analyzing yeast coding and promoter sequences, we find that neutral mutation rates in yeast are uniform genome-wide, in contrast to mammals, where neutral mutation rates vary along chromosomes. We develop an approach that uses this uniform rate to estimate the amount of promoter sequence under purifying selection. This amount is approximately 30%, corresponding to roughly 90 bp for a typical promoter. Furthermore, using a hidden Markov model, we are able to separate each promoter into distinct high and low conservation regions. Known regulatory motifs are strongly biased toward high conservation regions, while low conservation regions have mutation rates similar to that of the neutral background. Certain Gene Ontology groupings of genes (e.g., Carbohydrate Metabolism) have large amounts of high conservation sequence, suggesting complexity in their transcriptional regulation. Others (e.g., RNA Processing) have little high conservation sequence and are likely to be simply regulated. The separation of functionally conserved sequence from the neutral background allows us to estimate the complexity of cis-regulation on a genomic scale.

Base Sequence↗

Evaluation of regulatory potential and conservation scores for detecting cis-regulatory modules in aligned mammalian genome sequences.

Techniques of comparative genomics are being used to identify candidate functional DNA sequences, and objective evaluations are needed to assess their effectiveness. Different analytical methods score distinctive features of whole-genome alignments among human, mouse, and rat to predict functional regions. We evaluated three of these methods for their ability to identify the positions of known regulatory regions in the well-studied HBB gene complex. Two methods, multispecies conserved sequences and phastCons, quantify levels of conservation to estimate a likelihood that aligned DNA sequences are under purifying selection. A third function, regulatory potential (RP), measures the similarity of patterns in the alignments to those in known regulatory regions. The methods can correctly identify 50%-60% of noncoding positions in the HBB gene complex as regulatory or nonregulatory, with RP performing better than do other methods. When evaluated by the ability to discriminate genomic intervals, RP reaches a sensitivity of 0.78 and a true discovery rate of approximately 0.6. The performance is better on other reference sets; both phastCons and RP scores can capture almost all regulatory elements in those sets along with approximately 7% of the human genome.

Algorithms↗

Gene-balanced duplications, like tetraploidy, provide predictable drive to increase morphological complexity.

Controversy surrounds the apparent rising maximums of morphological complexity during eukaryotic evolution, with organisms increasing the number and nestedness of developmental areas as evidenced by morphological elaborations reflecting area boundaries. No "predictable drive" to increase this sort of complexity has been reported. Recent genetic data and theory in the general area of gene dosage effects has engendered a robust "gene balance hypothesis," with a theoretical base that makes specific predictions as to gene content changes following different types of gene duplication. Genomic data from both chordate and angiosperm genomes fit these predictions: Each type of duplication provides a one-way injection of a biased set of genes into the gene pool. Tetraploidies and balanced segments inject bias for those genes whose products are the subunits of the most complex biological machines or cascades, like transcription factors (TFs) and proteasome core proteins. Most duplicate genes are removed after tetraploidy. Genic balance is maintained by not removing those genes that are dose-sensitive, which tends to leave duplicate "functional modules" as the indirect products (spandrels) of purifying selection. Functional modules are the likely precursors of coadapted gene complexes, a unit of natural selection. The result is a predictable drive mechanism where "drive" is used rigorously, as in "meiotic drive." Rising morphological gain is expected given a supply of duplicate functional modules. All flowering plants have survived at least three large-scale duplications/diploidizations over the last 300 million years (Myr). An equivalent period of tetraploidy and body plan evolution may have ended for animals 500 million years ago (Mya). We argue that "balanced gene drive" is a sufficient explanation for the trend that the maximums of morphological complexity have gone up, and not down, in both plant and animal eukaryotic lineages.

Animals↗

Sequencing and analysis of 10,967 full-length cDNA clones from Xenopus laevis and Xenopus tropicalis reveals post-tetraploidization transcriptome remodeling.

Sequencing of full-insert clones from full-length cDNA libraries from both Xenopus laevis and Xenopus tropicalis has been ongoing as part of the Xenopus Gene Collection Initiative. Here we present 10,967 full ORF verified cDNA clones (8049 from X. laevis and 2918 from X. tropicalis) as a community resource. Because the genome of X. laevis, but not X. tropicalis, has undergone allotetraploidization, comparison of coding sequences from these two clawed (pipid) frogs provides a unique angle for exploring the molecular evolution of duplicate genes. Within our clone set, we have identified 445 gene trios, each comprised of an allotetraploidization-derived X. laevis gene pair and their shared X. tropicalis ortholog. Pairwise dN/dS, comparisons within trios show strong evidence for purifying selection acting on all three members. However, dN/dS ratios between X. laevis gene pairs are elevated relative to their X. tropicalis ortholog. This difference is highly significant and indicates an overall relaxation of selective pressures on duplicated gene pairs. We have found that the paralogs that have been lost since the tetraploidization event are enriched for several molecular functions, but have found no such enrichment in the extant paralogs. Approximately 14% of the paralogous pairs analyzed here also show differential expression indicative of subfunctionalization.

Animals↗

Characterization and predictive discovery of evolutionarily conserved mammalian alternative promoters.

Recent studies suggest that surprisingly many mammalian genes have alternative promoters (APs); however, their biological roles, and the characteristics that distinguish them from single promoters (SPs), remain poorly understood. We constructed a large data set of evolutionarily conserved promoters, and used it to identify sequence features, functional associations, and expression patterns that differ by promoter type. The four promoter categories CpG-rich APs, CpG-poor APs, CpG-rich SPs, and CpG-poor SPs each show characteristic strengths and patterns of sequence conservation, frequencies of putative transcription-related motifs, and tissue and developmental stage expression preferences. APs display substantially higher sequence conservation than SPs and CpG-poor promoters than CpG-rich promoters. Among CpG-poor promoters, APs and SPs show sharply contrasting developmental stage preferences and TATA box frequencies. We developed a discriminator to computationally predict promoter type, verified its accuracy through experimental tests that incorporate a novel method for deconvolving mixed sequence traces, and used it to find several new APs. The discriminator predicts that almost half of all mammalian genes have evolutionarily conserved APs. This high frequency of APs, together with the strong purifying selection maintaining them, implies a crucial role in expanding the expression diversity of the mammalian genome.

Alternative Splicing↗

Complex evolution of 7E olfactory receptor genes in segmental duplications.

Large segmental duplications (SDs) constitute at least 3.6% of the human genome and have increased its size, complexity, and diversity. SDs can mediate ectopic sequence exchange resulting in gross chromosomal rearrangements that could contribute to speciation and disease. We have identified and evaluated a subset of human SDs that harbor an 88-member subfamily of olfactory receptor (OR)-like genes called the 7Es. At least 92% of these genes appear to be pseudogenes when compared to other OR genes. The 7E-containing SDs (7E SDs) have duplicated to at least 35 regions of the genome via intra- and interchromosomal duplication events. In contrast to many human SDs, the 7E SDs are not biased towards pericentromeric or subtelomeric regions. We find evidence for gene conversion among 7E genes and larger sequence exchange between 7E SDs, supporting the hypothesis that long, highly similar stretches of DNA facilitate ectopic interactions. The complex structure and history of the 7E SDs necessitates extension of the current model of large-scale DNA duplication. Despite their appearance as pseudogenes, some 7E genes exhibit a signature of purifying selection, and at least one 7E gene is expressed.

Amino Acid Sequence↗

Essential genes are more evolutionarily conserved than are nonessential genes in bacteria.

The "knockout-rate" prediction holds that essential genes should be more evolutionarily conserved than are nonessential genes. This is because negative (purifying) selection acting on essential genes is expected to be more stringent than that for nonessential genes, which are more functionally dispensable and/or redundant. However, a recent survey of evolutionary distances between Saccharomyces cerevisiae and Caenorhabditis elegans proteins did not reveal any difference between the rates of evolution for essential and nonessential genes. An analysis of mouse and rat orthologous genes also found that essential and nonessential genes evolved at similar rates when genes thought to evolve under directional selection were excluded from the analysis. In the present study, we combine genomic sequence data with experimental knockout data to compare the rates of evolution and the levels of selection for essential versus nonessential bacterial genes. In contrast to the results obtained for eukaryotic genes, essential bacterial genes appear to be more conserved than are nonessential genes over both relatively short (microevolutionary) and longer (macroevolutionary) time scales.

Conserved Sequence↗

Recently duplicated maize R2R3 Myb genes provide evidence for distinct mechanisms of evolutionary divergence after duplication.

R2R3 Myb genes are widely distributed in the higher plants and comprise one of the largest known families of regulatory proteins. Here, we provide an evolutionary framework that helps explain the origin of the plant-specific R2R3 Myb genes from widely distributed R1R2R3 Myb genes, through a series of well-established steps. To understand the routes of sequence divergence that followed Myb gene duplication, we supplemented the information available on recently duplicated maize (Zea mays) R2R3 Myb genes (C1/Pl1 and P1/P2) by cloning and characterizing ZmMyb-IF35 and ZmMyb-IF25. These two genes correspond to the recently expanded P-to-A group of maize R2R3 Myb genes. Although the origins of C1/Pl1 and ZmMyb-IF35/ZmMyb-IF25 are associated with the segmental allotetraploid origin of the maize genome, other gene duplication events also shaped the P-to-A clade. Our analyses indicate that some recently duplicated Myb gene pairs display substantial differences in the numbers of synonymous substitutions that have accumulated in the conserved MYB domain and the divergent C-terminal regions. Thus, differences in the accumulation of substitutions during evolution can explain in part the rapid divergence of C-terminal regions for these proteins in some cases. Contrary to previous studies, we show that the divergent C termini of these R2R3 MYB proteins are subject to purifying selection. Our results provide an in-depth analysis of the sequence divergence for some recently duplicated R2R3 Myb genes, yielding important information on general patterns of evolution for this large family of plant regulatory genes.

Amino Acid Sequence↗

Genome organization of more than 300 defensin-like genes in Arabidopsis.

Defensins represent an ancient and diverse set of small, cysteine-rich, antimicrobial peptides in mammals, insects, and plants. According to published accounts, most species' genomes contain 15 to 50 defensins. Starting with a set of largely nodule-specific defensin-like sequences (DEFLs) from the model legume Medicago truncatula, we built motif models to search the near-complete Arabidopsis (Arabidopsis thaliana) genome. We identified 317 DEFLs, yet 80% were unannotated at The Arabidopsis Information Resource and had no prior evidence of expression. We demonstrate that many of these DEFL genes are clustered in the Arabidopsis genome and that individual clusters have evolved from successive rounds of gene duplication and divergent or purifying selection. Sequencing reverse transcription-PCR products from five DEFL clusters confirmed our gene predictions and verified expression. For four of the largest clusters of DEFLs, we present the first evidence of expression, most frequently in floral tissues. To determine the abundance of DEFLs in other plant families, we used our motif models to search The Institute for Genomic Research's gene indices and identified approximately 1,100 DEFLs. These expressed DEFLs were found mostly in reproductive tissues, consistent with our reverse transcription-PCR results. Sequence-based clustering of all identified DEFLs revealed separate tissue- or taxon-specific subgroups. Previously, we and others showed that more than 300 DEFL genes were expressed in M. truncatula nodules, organs not present in most plants. We have used this information to annotate the Arabidopsis genome and now provide evidence of a large DEFL superfamily present in expressed tissues of all sequenced plants.

Amino Acid Sequence↗

Mutational decay and age of chloroplast and mitochondrial genomes transferred recently to angiosperm nuclear chromosomes.

Transfers of organelle DNA to the nucleus established several thousand functional genes in eukaryotic chromosomes over evolutionary time. Recent transfers have also contributed nonfunctional plastid (pt)- and mitochondrion (mt)-derived DNA (termed nupts and numts, respectively) to plant nuclear genomes. The two largest transferred organelle genome copies are 131-kb nuptDNA in rice (Oryza sativa) and 262-kb numtDNA in Arabidopsis (Arabidopsis thaliana). These transferred copies were compared in detail with their bona fide organelle counterparts, to which they are 99.77% and 99.91% identical, respectively. No evidence for purifying selection was found in either nuclear integrant, indicating that they are nonfunctional. Mutations attributable to 5-methylcytosine hypermutation have occurred at a 6- to 10-fold higher rate than other point mutations in Arabidopsis numtDNA and rice nuptDNA, respectively, revealing this as a major mechanism of mutational decay for these transferred organelle sequences. Short indels occurred preferentially within homopolymeric stretches but were less frequent than point mutations. The 131-kb nuptDNA is absent in the O. sativa subsp. indica or Oryza rufipogon nuclear genome, suggesting that it was transferred within the O. sativa subsp. japonica lineage and, as revealed by sequence comparisons, after its divergence from the indica chloroplast lineage. The time of the transfer for the rice nupt was estimated as 148,000 (74,000--296,000) years ago and that for the Arabidopsis numtDNA as 88,000 (44,000--176,000) years ago. The results reveal transfer and integration of entire organelle genomes into the nucleus as an ongoing evolutionary process and uncover mutational mechanisms affecting organelle genomes recently transferred into a new mutational environment.

Arabidopsis↗

Recent proliferation and translocation of pollen group 1 allergen genes in the maize genome.

The dominant allergenic components of grass pollen are known by immunologists as group 1 allergens. These constitute a set of closely related proteins from the beta-expansin family and have been shown to have cell wall-loosening activity. Group 1 allergens may facilitate the penetration of pollen tubes through the grass stigma and style. In maize (Zea mays), group 1 allergens are divided into two classes, A and B. We have identified 15 genes encoding group 1 allergens in maize, 11 genes in class A and four genes in class B, as well as seven pseudogenes. The genes in class A can be divided by sequence relatedness into two complexes, whereas the genes in class B constitute a single complex. Most of the genes identified are represented in pollen-specific expressed sequence tag libraries and are under purifying selection, despite the presence of multiple copies that are nearly identical. Group 1 allergen genes are clustered in at least six different genomic locations. The single class B location and one of the class A locations show synteny with the rice (Oryza sativa) regions where orthologous genes are found. Both classes are expressed at high levels in mature pollen but at low levels in immature flowers. The set of genes encoding maize group 1 allergens is more complex than originally anticipated. If this situation is common in grasses, it may account for the large number of protein variants, or group 1 isoallergens, identified previously in turf grass pollen by immunologists.

Antigens, Plant↗

A calmodulin-sensitive interaction between microtubules and a higher plant homolog of elongation factor-1 alpha.

The microtubules (MTs) of higher plant cells are organized into arrays with essential functions in plant cell growth and differentiation; however, molecular mechanisms underlying the organization and regulation of these arrays remain largely unknown. We have approached this problem using tubulin affinity chromatography to isolate carrot proteins that interact with MTs. From these proteins, a 50-kD polypeptide was selectively purified by exploiting its Ca(2+)-dependent binding to calmodulin (CaM). This polypeptide was identified as a homolog of elongation factor-1 alpha (EF-1 alpha)--a highly conserved and ubiquitous protein translation factor. The carrot EF-1 alpha homolog bundles MTs in vitro, and moreover, this bundling is modulated by the addition of Ca2+ and CaM together (Ca2+/CaM). A direct binding between the EF-1 alpha homolog and MTs was demonstrated, providing novel evidence for such an interaction. Based on these findings, and others discussed herein, we propose that an EF-1 alpha homolog mediates the lateral association of MTs in plant cells by a Ca2+/CaM-sensitive mechanism.

Amino Acid Sequence↗

Phylogenetic and evolutionary analysis of the PLUNC gene family.

The PLUNC family of human proteins are candidate host defense proteins expressed in the upper airways. The family subdivides into short (SPLUNC) and long (LPLUNC) proteins, which contain domains predicted to be structurally similar to one or both of the domains of bactericidal/permeability-increasing protein (BPI), respectively. In this article we use analysis of the human, mouse, and rat genomes and other sequence data to examine the relationships between the PLUNC family proteins from humans and other species, and between these proteins and members of the BPI family. We show that PLUNC family clusters exist in the mouse and rat, with the most significant diversification in the locus occurring for the short PLUNC family proteins. Clear orthologous relationships are established for the majority of the proteins, and ambiguities are identified. Completion of the prediction of the LPLUNC4 proteins reveals that these proteins contain approximately a 150-residue insertion encoded by an additional exon. This insertion, which is predicted to be largely unstructured, replaces the structure homologous to the 40s hairpin of BPI. We show that the exon encoding this region is anomalously variable in size across the LPLUNC proteins, suggesting that this region is key to functional specificity. We further show that the mouse and human PLUNC family orthologs are evolving rapidly, which supports the hypothesis that these proteins are involved in host defense. Intriguingly, this rapid evolution between the human and mouse sequences is replaced by intense purifying selection in a large portion of the N-terminal domain of LPLUNC4. Our data provide a basis for future functional studies of this novel protein family.

Amino Acid Sequence↗