PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “evolutionary analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

An evolutionary trace method defines binding surfaces common to protein families.

X-ray or NMR structures of proteins are often derived without their ligands, and even when the structure of a full complex is available, the area of contact that is functionally and energetically significant may be a specialized subset of the geometric interface deduced from the spatial proximity between ligands. Thus, even after a structure is solved, it remains a major theoretical and experimental goal to localize protein functional interfaces and understand the role of their constituent residues. The evolutionary trace method is a systematic, transparent and novel predictive technique that identifies active sites and functional interfaces in proteins with known structure. It is based on the extraction of functionally important residues from sequence conservation patterns in homologous proteins, and on their mapping onto the protein surface to generate clusters identifying functional interfaces. The SH2 and SH3 modular signaling domains and the DNA binding domain of the nuclear hormone receptors provide tests for the accuracy and validity of our method. In each case, the evolutionary trace delineates the functional epitope and identifies residues critical to binding specificity. Based on mutational evolutionary analysis and on the structural homology of protein families, this simple and versatile approach should help focus site-directed mutagenesis studies of structure-function relationships in macromolecules, as well as studies of specificity in molecular recognition. More generally, it provides an evolutionary perspective for judging the functional or structural role of each residue in protein structure.

Amino Acid Sequence↗

Evolution of T-cell receptor gamma and delta constant region and other T-cell-related proteins in the human-rodent-artiodactyl triplet.

In this paper we report a detailed comparative and evolutionary analysis of the sequences of constant T-cell receptor (Tcr) C gamma delta genes of artiodactyls compared to the homologous sequences of rodents and primates. Because of the frequency and physiological distribution of gamma delta T-cells in different animals, rodents and humans are defined as "gamma delta low" species and ruminants as "gamma delta high" species. Such a characteristic seems to be due to an adaptive role of gamma delta T-cell function. By analyzing the ruminant gene phylogeny of Tcr C gamma we were able to estimate the distance between cattle and sheep at 18 million years ago, a time that is in agreement with other nonmolecular estimates. For Tcr C gamma delta genes a peculiar phylogenetic relationship was found, with human and mouse clustering together and leaving artiodactyls apart. By using appropriate outgroups, the same phylogenetic pattern was obtained with other T-cell related sequences: namely, Tcr C alpha chain, CD3 gamma and delta invariant subunits. Interleukin-2. Interleukin-2 receptor alpha chain and Interleukin-1 beta with the exception of Tcr C beta chain and Interleukin-1 alpha. In contrast, the analysis of all other T-cell nonrelated genes, available in primary databases reveals a different tree, where primates and artiodactyls are sister taxa and rodents are apart in accordance with the current view of mammalian phylogeny. These data are relevant to important evolutionary issues. They show how misleading a phylogeny based on a single or on a few homologous genes may be. In addition they demonstrate that genes with correlated functions may evolve in a lineage specific manner probably in relation to environmental conditions.

Amino Acid Sequence↗

Genetic evolution of structural region of hepatitis C virus in primary infection.

AIM: To investigate the dynamics of hepatitis C virus (HCV) variability through putative envelope genes during primary infection and the mechanism of viral genetic evolution in infected hosts. METHODS: Serial serum samples prospectively collected for 12 to 34 months from a cohort of acutely HCV-infected individuals were obtained, and a 1-kb fragment spanning E1 and the 5' half of E2, including Thirty-three cloned cDNAs representing each specimen were assessed by a method that combined a single-stranded conformational polymorphism (SSCP) and heteroduplex analysis (HDA) method to determine the number of clonotypes hypervariable region, was amplified by reverse transcriptase PCR and cloned. Nonsynonymous mutations per nonsynonymous site (dn), synonymous mutations per synonymous site (ds), dn/ds ratio and genetic distances within each sample were evaluated for intrahost evolutionary analysis. RESULTS: Quasispecies complexity and sequence diversity were lower in early samples and a further increase after seroconversion, although ds value in the envelope genes was higher than dn value during primary infection. The trend, pronounced in most of samples, toward lower ds values in the E1 than in the 5' portion of E2. Quasispecies complexity was higher and E2 dn/ds ratio was a trend toward higher value in later samples during persistent viremia. We also found individual features of HCV genetic evolution in different subjects who were infected with different HCV genotypes. CONCLUSION: Mutations of actively replicating virus arise stochastically with certain functional constaints. A complexity quasispecies exerted by a combination of either neutral evolution or selective forces shows clear differences in individuals, and associated with HCV persistence.

Adult↗

Evolution of cell lineage and pattern formation in the vulval equivalence group of rhabditid nematodes.

During the formation of the vulva in many nematode hermaphrodites or females, pattern formation, induction, and cell specification can readily be studied at a single-cell level. Nematodes thus allow an evolutionary analysis of developmental processes. We have analyzed cell lineages and pattern formation in the vulva equivalence group of six rhabditid nematodes of the genera Oscheius, Rhabditella, Rhabditoides, Pelodera, and Protorhabditis. The comparison of these species with four previously analyzed species of this family reveals evolutionary modification at several levels. The number of vulva precursor cells (VPCs) differ among species. Of the three particular cell lineages (1 degree, 2 degrees, and 3 degrees) generated by the vulva precursor cells in Caenorhabditis, two (2 degrees and 3 degrees) are altered, whereas the third lineage (1 degree) is conserved among the analyzed species. While most vulval lineages are invariant, we observe variability of the 3 degrees lineage in Pelodera with respect to the number of precursor cells adopting this fate and the number of progeny formed. In two species, the 3 degrees lineage generates an asymmetrical set of cells, oriented by the gonad. In Protorhabditis we frequently find animals with an additional or altered set of VPCs forming vulval tissue.

Animals↗

Molecular evolution of olfactomedin.

Olfactomedin is a secreted polymeric glycoprotein of unknown function, originally discovered at the mucociliary surface of the amphibian olfactory neuroepithelium and subsequently found throughout the mammalian brain. As a first step toward elucidating the function of olfactomedin, its phylogenetic history was examined to identify conserved structural motifs. Such conserved motifs may have functional significance and provide targets for future mutagenesis studies aimed at establishing the function of this protein. Previous studies revealed 33% amino acid sequence identity between rat and frog olfactomedins in their carboxyl terminal segments. Further analysis, however, reveals more extensive homologies throughout the molecule. Despite significant sequence divergence, cysteines essential for homopolymer formation such as the CXC motif near the amino terminus are conserved, as is the characteristic glycosylation pattern, suggesting that these posttranslational modifications are essential for function. Furthermore, evolutionary analysis of a region of 53 amino acids of fish, frog, rat, mouse, and human olfactomedins indicates that an ancestral olfactomedin gene arose before the evolution of terrestrial vertebrates and evolved independently in teleost, amphibian, and mammalian lineages. Indeed, a distant olfactomedin homolog was identified in Caenorhabditis elegans. Although the amino acid sequence of this invertebrate protein is longer and highly divergent compared with its vertebrate homologs, the protein from C. elegans shows remarkable similarities in terms of conserved motifs and posttranslational modification sites. Six universally conserved motifs were identified, and five of these are clustered in the carboxyl terminal half of the protein. Sequence comparisons indicate that evolution of the N-terminal half of the molecule involved extensive insertions and deletions; the C-terminal segment evolved mostly through point mutations, at least during vertebrate evolution. The widespread occurrence of olfactomedin among vertebrates and invertebrates underscores the notion that this protein has a function of universal importance. Furthermore, extensive modification of its N-terminal half and the acquisition of a C-terminal SDEL endoplasmic-reticulum-targeting sequence may have enabled olfactomedin to adopt new functions in the mammalian central nervous system.

Amino Acid Sequence↗

Understanding Mycobacterium tuberculosis through its genomic diversity and evolution.

Pathogen evolution and genomic diversity are shaped by specific host immune pressures and therapeutic interventions. Analysis of the extant genomes of circulating strains of Mycobacterium tuberculosis, a leading cause of infectious mortality that has co-evolved with humans for thousands of years, can provide new insights into host-pathogen interactions that underlie specific aspects of pathogenesis and onward transmission. With the explosion in the number of fully sequenced M. tuberculosis strains that are now paired with detailed clinical data, there are new opportunities to understand the evolutionary basis for and consequences of M. tuberculosis strain diversity. This review examines mechanistic findings that have emerged from pairing whole genome sequencing data and evolutionary analysis with functional dissection of specific bacterial variants. These include improved understanding of secreted effectors that modulate the properties and migratory behavior of infected macrophages as well as bacterial genetic alterations important for survival within hypoxic microenvironments. Genomic, evolutionary, and functional analyses across diverse M. tuberculosis strains will identify prominent bacterial adaptations to their human hosts and shape our understanding of TB disease biology and the host immune response.

Mycobacterium tuberculosis↗

Application of nucleotide sequence of RNA polymerase beta-subunit gene (rpoB) to molecular differentiation of serovars of Salmonella enterica subsp. enterica.

To establish a molecular differentiation method for Salmonella enterica subsp. enterica, a hyper-variable region of RNA polymerase beta-subunit (rpoB) of S. enterica subsp. enterica (I), serotype Typhimurium, and Escherichia coli were investigated through comparison of nucleotide sequence of the region. The hyper-variable region was identified at 612-937 of the gene. After PCR amplification of the region in the 17 serotypes and two biotypes of serotype Gallinarum of S. enterica subsp. enterica (I), the nucleotide sequences of the region were determined and compared. All serotypes were distantly related to E. coli with 82.8-84.7% identities in nucleotide sequence while showing 96.6-100% identities with each other. According to the phylogenetic analysis based on the sequenced region with the neighbor-joining method, relatedness of biotype Gallinarum to serotype Enteritidis and biotype Pullorum was determined. Biotype Gallinarum was more closely related to serotype Enteritidis than biotype Pullorum. These results suggested that the 612-937 variable region of rpoB might be useful for molecular evolutionary analysis of serotypes of S. enterica subsp. enterica (I).

Amino Acid Sequence↗

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning↗

The nop-1 gene of Neurospora crassa encodes a seven transmembrane helix retinal-binding protein homologous to archaeal rhodopsins.

Opsins are a class of retinal-binding, seven transmembrane helix proteins that function as light-responsive ion pumps or sensory receptors. Previously, genes encoding opsins had been identified in animals and the Archaea but not in fungi or other eukaryotic microorganisms. Here, we report the identification and mutational analysis of an opsin gene, nop-1, from the eukaryotic filamentous fungus Neurospora crassa. The nop-1 amino acid sequence predicts a protein that shares up to 81.8% amino acid identity with archaeal opsins in the 22 retinal binding pocket residues, including the conserved lysine residue that forms a Schiff base linkage with retinal. Evolutionary analysis revealed relatedness not only between NOP-1 and archaeal opsins but also between NOP-1 and several fungal opsin-related proteins that lack the Schiff base lysine residue. The results provide evidence for a eukaryotic opsin family homologous to the archaeal opsins, providing a plausible link between archaeal and visual opsins. Extensive analysis of Deltanop-1 strains did not reveal obvious defects in light-regulated processes under normal laboratory conditions. However, results from Northern analysis support light and conidiation-based regulation of nop-1 gene expression, and NOP-1 protein heterologously expressed in Pichia pastoris is labeled by using all-trans [3H]retinal, suggesting that NOP-1 functions as a rhodopsin in N. crassa photobiology.

Amino Acid Sequence↗

Controversies in the evolutionary social sciences: a guide for the perplexed.

It is 25 years since modern evolutionary ideas were first applied extensively to human behavior, jump-starting a field of study once known as 'sociobiology'. Over the years, distinct styles of evolutionary analysis have emerged within the social sciences. Although there is considerable complementarity between approaches that emphasize the study of psychological mechanisms and those that focus on adaptive fit to environments, there are also substantial theoretical and methodological differences. These differences have generated a recurrent debate that is now exacerbated by growing popular media attention to evolutionary human behavioral studies. Here, we provide a guide to current controversies surrounding evolutionary studies of human social behavior, emphasizing theoretical and methodological issues. We conclude that a greater use of formal models, measures of current fitness costs and benefits, and attention to adaptive tradeoffs, will enhance the power and reliability of evolutionary analyses of human social behavior.

Journal Article↗

Isolation of a cDNA encoding the B isozyme of human phosphoglycerate mutase (PGAM) and characterization of the PGAM gene family.

We previously reported the isolation of a full-length cDNA specifying the muscle-specific isozyme of human phosphoglycerate mutase (PGAM-M). We now report the isolation of a full-length cDNA specifying the non-muscle-specific, or brain (B), isozyme of human PGAM (PGAM-B). The PGAM-B cDNA encodes a deduced protein 254 amino acids long, 79% identical to PGAM-M, and contains a 913-nucleotide 3'-untranslated region, as compared to the unusually short 37-nucleotide 3'-untranslated region of PGAM-M. Northern analysis demonstrates the non-muscle-specific nature of PGAM-B transcription, while genomic Southern analysis implies the presence of a large PGAM family in the human genome. Most of the PGAM-hybridizing sequences in both the human and mouse genomes seem to be related to the B-isozyme gene; many members of the PGAM-B gene family in humans are apparently processed genes. These results agree with the evolutionary analysis, which indicates that the PGAM-B gene is the progenitor of the PGAM-M gene.

Amino Acid Sequence↗

Dynamic and non-additive gene regulation shapes maize responses to simultaneous salt and cold stress.

Salt and cold stresses often occur together in nature and severely impact crop productivity, yet their transcriptional regulation remains poorly understood. Here, we conducted a time-series transcriptomic analysis of maize under salt, cold, and their combination at 0, 6, 12, and 24 h. Differential expression analysis revealed dynamic, condition-specific gene responses grouped into eight distinct temporal patterns. Promoter motif analysis of genes within each pattern identified 5-39 significantly enriched motifs, with over 40% lacking known counterparts, suggesting the involvement of previously uncharacterized cis-regulatory elements in stress-responsive transcriptional regulation. By comparing combined stress responses to the sum of single-stress effects, we found that about 74% of DEGs showed non-additive patterns, suggesting that combined stress triggers a distinct transcriptional program. Evolutionary analysis showed that additive DEGs tend to be more recently evolved, subject to weaker purifying selection, and enriched in transposed duplications, contrasting with the stronger constraint observed in non-additive DEGs. WGCNA identified 24 co-expression modules, among which 65 hub DEGs were detected in modules significantly correlated with specific stress conditions. Furthermore, we reconstructed 228, 20, and 200 sequential transcription factor cascades spanning 6 h, 12 h, and 24 h under cold, salt, and combined stress, respectively, with no cascade shared across all three conditions. Together, these results reveal that maize responses to combined salt and cold stress are largely non-additive and temporally dynamic, with distinct evolutionary patterns underlying different response types, offering insights and candidate regulators for enhancing crop stress resilience.

Zea mays↗

Phylogenetic analysis of Ljungan virus and A-2 plaque virus, new members of the Picornaviridae.

In addition to the viruses belonging to the nine proposed genera of the Picornaviridae, Enterovirus, Rhinovirus, Cardiovirus, Aphtovirus, Hepatovirus, Parechovirus, Kobuvirus, Erbovirus and Teschovirus, two new members of this family have recently been discovered. Three strains of Ljungan virus (LV) were isolated from bank voles (Clethrionomys glareolus) and A-2 plaque virus (A-2) was isolated from human sera. To study the genetic relationship between these recently discovered viruses and the members of the family Picornaviridae, an evolutionary analysis has been carried out using the amino acid sequences of the two nonstructural proteins 2C and 3D. Phylogenetic analysis using prime members of the nine genera support the division of picornaviruses into the proposed genera. The study also supports a previous suggestion based on analysis of partial sequences of the structural proteins that LV is more related to the genus of Parechovirus than to other picornaviruses, but also shows that the three LV strains used in the comparison constitute a distinct monophyletic group, clearly separated from the parechoviruses. The analyses using the 2C and 3D sequences clearly showed that A-2 was related to the genera of Rhinovirus and Enterovirus, but it was not possible to group the A-2 with high confidence into one of the genera. Comparison using the VP1 protein sequences of Enterovirus and Rhinovirus showed that although the A-2 virus is positioned between the two genera, the virus is more related to the genus of Enterovirus than to Rhinovirus. Our analysis of the three LV strains based on the phylogenetic analysis of the 2C and 3D proteins suggests that the strains used in this study constitute a monophyletic group clearly related to Parechovirus of Picornaviridae. The taxonomic position of the A-2 virus is presently uncertain but available data indicate that this virus may be classified as a member of the genus of Enterovirus.

Animals↗

Primate evolution of an olfactory receptor cluster: diversification by gene conversion and recent emergence of pseudogenes.

The olfactory receptor (OR) subgenome harbors the largest known gene family in mammals, disposed in clusters on numerous chromosomes. We have carried out a comparative evolutionary analysis of the best characterized genomic OR gene cluster, on human chromosome 17p13. Fifteen orthologs from chimpanzee (localized to chromosome 19p15), as well as key OR counterparts from other primates, have been identified and sequenced. Comparison among orthologs and paralogs revealed a multiplicity of gene conversion events, which occurred exclusively within OR subfamilies. These appear to lead to segment shuffling in the odorant binding site, an evolutionary process reminiscent of somatic combinatorial diversification in the immune system. We also demonstrate that the functional mammalian OR repertoire has undergone a rapid decline in the past 10 million years: while for the common ancestor of all great apes an intact OR cluster is inferred, in present-day humans and great apes the cluster includes nearly 40% pseudogenes.

Animals↗

Drift, admixture, and selection in human evolution: a study with DNA polymorphisms.

Accuracy of evolutionary analysis of populations within a species requires the testing of a large number of genetic polymorphisms belonging to many loci. We report here a reconstruction of human differentiation based on 100 DNA polymorphisms tested in five populations from four continents. The results agree with earlier conclusions based on other classes of genetic markers but reveal that Europeans do not fit a simple model of independently evolving populations with equal evolutionary rates. Evolutionary models involving early admixture are compatible with the data. Taking one such model into account, we examined through simulation whether random genetic drift alone might explain the variation among gene frequencies across populations and genes. A measure of variation among populations was calculated for each polymorphism, and its distribution for the 100 polymorphisms was compared with that expected for a drift-only hypothesis. At least two-thirds of the polymorphisms appear to be selectively neutral, but there are significant deviations at the two ends of the observed distribution of the measure of variation: a slight excess of polymorphisms with low variation and a greater excess with high variation. This indicates that a few DNA polymorphisms are affected by natural selection, rarely heterotic, and more often disruptive, while most are selectively neutral.

Animals↗

A single-nucleus transcriptome atlas of soybean anthers.

Anther development is crucial for plant sexual reproduction. However, a high-resolution, cell-type-specific transcriptomic atlas of this process is lacking for the legume crop soybean (Glycine max). Here, we construct a comprehensive transcriptional atlas of developing soybean anthers using single-nucleus RNA sequencing (snRNA-seq). We identify and characterize nine distinct cell types spanning both somatic and reproductive lineages. Our analysis reveals robust transcriptional continuity across anther developmental stages and dynamic reprogramming during key transitions. Notably, the shift from diploid meiocytes to haploid unicellular microspores is marked by the induction of previously inactive genes, despite an overall reduction in transcript abundance. Subsequently, within bicellular microspores, generative and vegetative cell lineages exhibit sharply divergent transcriptional programs: generative cells specialize in mRNA export and turnover, whereas vegetative cells up-regulate translational machinery. Evolutionary analysis further indicates that generative-cell-specific genes are subject to more relaxed purifying selection compared to those specific to vegetative cells. Functional validation using mutants generated by CRISPR/Cas9-mediated genome editing and EMS mutagenesis reveals the essential roles of OSD1A and PKSA in pollen development and fertility. This high-resolution atlas provides fundamental insights into the transcriptional regulation of soybean anther development and serves as a valuable resource for manipulating male fertility to advance hybrid breeding programs. The data are available at https://databases.genedenovo.com/pollen.

Glycine max↗

Identification of rice DUF1719 gene family and analysis of alkaline tolerance function of OsDUF1719.8.

Alkaline stress severely constrains the physiological metabolism and growth and development of rice through high pH and ionic toxicity. Domains of unknown function (DUF) play significant roles in plant stress responses. However, the function of the DUF1719 family (PF08224) in rice has not been reported and further research is needed. This study systematically identified the OsDUF1719 gene family in rice and investigated the function of OsDUF1719.8 under alkaline stress. The results demonstrate that the rice DUF1719 family comprises 13 protein members, all containing the PF08224 domain. It is predicted that this domain may play a role in ATPase activation. Evolutionary analysis divided DUF1719 proteins from eight grass species into six subgroups, with highly conserved gene structures, motifs, and tertiary architectures within each subgroup. Promoter analysis indicated enrichment of stress- and hormone-responsive elements, implying broad involvement in stress regulation. Expression analysis revealed that several genes, including OsDUF1719.5 and OsDUF1719.8, were upregulated under multiple abiotic stresses. Notably, OsDUF1719.8 was strongly induced during early alkaline stress. Consequently, we further analyzed the function of OsDUF1719.8 in the rice alkaline stress response. The results demonstrate that overexpression of OsDUF1719.8 enhanced rice alkaline tolerance, whereas knockout mutants exhibited stress sensitivity. OsDUF1719.8 enhances rice tolerance to alkaline stress by coordinately regulating reactive oxygen species metabolism, promoting the accumulation of osmotic adjustment compounds, and modulating ion homeostasis. This study provides the first systematic identification of the DUF1719 family and elucidates the function of OsDUF1719.8 in positively regulating rice alkaline tolerance, offering a novel gene for alkali-tolerant molecular breeding of rice.

Oryza↗

Sequence analysis of hepatitis C virus genotypes 1 to 5 reveals multiple novel subtypes in the Benelux countries.

Hepatitis C virus (HCV) isolates from a cohort of 315 patients from the Benelux countries (Belgium, The Netherlands, Luxembourg) were genotyped by means of reverse hybridization Inno-LiPA (line probe assay). Genotypes 1a, 1b, 2a, 2b, 3a, 4a and 5a were detected. From the cohort, isolates representing all types and those showing an aberrant LiPA pattern were further analysed by sequencing parts of the 5' UTR, core (nt 1 to 326; aa residues 1 to 108) and core/E1 (nt 477 to 924; aa residues 159 to 308) regions. Molecular evolutionary analysis of the core and core/E1 regions allowed discrimination between known and additional subtypes, especially within types 2 and 4. The core region is not suitable for classification of new subtypes because of the relatively high level of conservation. The core/E1 region displays a higher level of sequence variation and allows much more distinct discrimination between subtypes. Genotypes 2 and 4 are particularly heterogeneous, with at least 7 and 10 subtypes, respectively. In contrast to previous reports from Europe, HCV isolates from the cohort constituted a highly heterogeneous population of virus variants, especially within genotypes 2 and 4.

Belgium↗