PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “evolutionary analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Application of nucleotide sequence of RNA polymerase beta-subunit gene (rpoB) to molecular differentiation of serovars of Salmonella enterica subsp. enterica.

To establish a molecular differentiation method for Salmonella enterica subsp. enterica, a hyper-variable region of RNA polymerase beta-subunit (rpoB) of S. enterica subsp. enterica (I), serotype Typhimurium, and Escherichia coli were investigated through comparison of nucleotide sequence of the region. The hyper-variable region was identified at 612-937 of the gene. After PCR amplification of the region in the 17 serotypes and two biotypes of serotype Gallinarum of S. enterica subsp. enterica (I), the nucleotide sequences of the region were determined and compared. All serotypes were distantly related to E. coli with 82.8-84.7% identities in nucleotide sequence while showing 96.6-100% identities with each other. According to the phylogenetic analysis based on the sequenced region with the neighbor-joining method, relatedness of biotype Gallinarum to serotype Enteritidis and biotype Pullorum was determined. Biotype Gallinarum was more closely related to serotype Enteritidis than biotype Pullorum. These results suggested that the 612-937 variable region of rpoB might be useful for molecular evolutionary analysis of serotypes of S. enterica subsp. enterica (I).

Amino Acid Sequence↗

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning↗

The nop-1 gene of Neurospora crassa encodes a seven transmembrane helix retinal-binding protein homologous to archaeal rhodopsins.

Opsins are a class of retinal-binding, seven transmembrane helix proteins that function as light-responsive ion pumps or sensory receptors. Previously, genes encoding opsins had been identified in animals and the Archaea but not in fungi or other eukaryotic microorganisms. Here, we report the identification and mutational analysis of an opsin gene, nop-1, from the eukaryotic filamentous fungus Neurospora crassa. The nop-1 amino acid sequence predicts a protein that shares up to 81.8% amino acid identity with archaeal opsins in the 22 retinal binding pocket residues, including the conserved lysine residue that forms a Schiff base linkage with retinal. Evolutionary analysis revealed relatedness not only between NOP-1 and archaeal opsins but also between NOP-1 and several fungal opsin-related proteins that lack the Schiff base lysine residue. The results provide evidence for a eukaryotic opsin family homologous to the archaeal opsins, providing a plausible link between archaeal and visual opsins. Extensive analysis of Deltanop-1 strains did not reveal obvious defects in light-regulated processes under normal laboratory conditions. However, results from Northern analysis support light and conidiation-based regulation of nop-1 gene expression, and NOP-1 protein heterologously expressed in Pichia pastoris is labeled by using all-trans [3H]retinal, suggesting that NOP-1 functions as a rhodopsin in N. crassa photobiology.

Amino Acid Sequence↗

Controversies in the evolutionary social sciences: a guide for the perplexed.

It is 25 years since modern evolutionary ideas were first applied extensively to human behavior, jump-starting a field of study once known as 'sociobiology'. Over the years, distinct styles of evolutionary analysis have emerged within the social sciences. Although there is considerable complementarity between approaches that emphasize the study of psychological mechanisms and those that focus on adaptive fit to environments, there are also substantial theoretical and methodological differences. These differences have generated a recurrent debate that is now exacerbated by growing popular media attention to evolutionary human behavioral studies. Here, we provide a guide to current controversies surrounding evolutionary studies of human social behavior, emphasizing theoretical and methodological issues. We conclude that a greater use of formal models, measures of current fitness costs and benefits, and attention to adaptive tradeoffs, will enhance the power and reliability of evolutionary analyses of human social behavior.

Journal Article↗

Isolation of a cDNA encoding the B isozyme of human phosphoglycerate mutase (PGAM) and characterization of the PGAM gene family.

We previously reported the isolation of a full-length cDNA specifying the muscle-specific isozyme of human phosphoglycerate mutase (PGAM-M). We now report the isolation of a full-length cDNA specifying the non-muscle-specific, or brain (B), isozyme of human PGAM (PGAM-B). The PGAM-B cDNA encodes a deduced protein 254 amino acids long, 79% identical to PGAM-M, and contains a 913-nucleotide 3'-untranslated region, as compared to the unusually short 37-nucleotide 3'-untranslated region of PGAM-M. Northern analysis demonstrates the non-muscle-specific nature of PGAM-B transcription, while genomic Southern analysis implies the presence of a large PGAM family in the human genome. Most of the PGAM-hybridizing sequences in both the human and mouse genomes seem to be related to the B-isozyme gene; many members of the PGAM-B gene family in humans are apparently processed genes. These results agree with the evolutionary analysis, which indicates that the PGAM-B gene is the progenitor of the PGAM-M gene.

Amino Acid Sequence↗

Dynamic and non-additive gene regulation shapes maize responses to simultaneous salt and cold stress.

Salt and cold stresses often occur together in nature and severely impact crop productivity, yet their transcriptional regulation remains poorly understood. Here, we conducted a time-series transcriptomic analysis of maize under salt, cold, and their combination at 0, 6, 12, and 24 h. Differential expression analysis revealed dynamic, condition-specific gene responses grouped into eight distinct temporal patterns. Promoter motif analysis of genes within each pattern identified 5-39 significantly enriched motifs, with over 40% lacking known counterparts, suggesting the involvement of previously uncharacterized cis-regulatory elements in stress-responsive transcriptional regulation. By comparing combined stress responses to the sum of single-stress effects, we found that about 74% of DEGs showed non-additive patterns, suggesting that combined stress triggers a distinct transcriptional program. Evolutionary analysis showed that additive DEGs tend to be more recently evolved, subject to weaker purifying selection, and enriched in transposed duplications, contrasting with the stronger constraint observed in non-additive DEGs. WGCNA identified 24 co-expression modules, among which 65 hub DEGs were detected in modules significantly correlated with specific stress conditions. Furthermore, we reconstructed 228, 20, and 200 sequential transcription factor cascades spanning 6 h, 12 h, and 24 h under cold, salt, and combined stress, respectively, with no cascade shared across all three conditions. Together, these results reveal that maize responses to combined salt and cold stress are largely non-additive and temporally dynamic, with distinct evolutionary patterns underlying different response types, offering insights and candidate regulators for enhancing crop stress resilience.

Zea mays↗

Primate evolution of an olfactory receptor cluster: diversification by gene conversion and recent emergence of pseudogenes.

The olfactory receptor (OR) subgenome harbors the largest known gene family in mammals, disposed in clusters on numerous chromosomes. We have carried out a comparative evolutionary analysis of the best characterized genomic OR gene cluster, on human chromosome 17p13. Fifteen orthologs from chimpanzee (localized to chromosome 19p15), as well as key OR counterparts from other primates, have been identified and sequenced. Comparison among orthologs and paralogs revealed a multiplicity of gene conversion events, which occurred exclusively within OR subfamilies. These appear to lead to segment shuffling in the odorant binding site, an evolutionary process reminiscent of somatic combinatorial diversification in the immune system. We also demonstrate that the functional mammalian OR repertoire has undergone a rapid decline in the past 10 million years: while for the common ancestor of all great apes an intact OR cluster is inferred, in present-day humans and great apes the cluster includes nearly 40% pseudogenes.

Animals↗

Drift, admixture, and selection in human evolution: a study with DNA polymorphisms.

Accuracy of evolutionary analysis of populations within a species requires the testing of a large number of genetic polymorphisms belonging to many loci. We report here a reconstruction of human differentiation based on 100 DNA polymorphisms tested in five populations from four continents. The results agree with earlier conclusions based on other classes of genetic markers but reveal that Europeans do not fit a simple model of independently evolving populations with equal evolutionary rates. Evolutionary models involving early admixture are compatible with the data. Taking one such model into account, we examined through simulation whether random genetic drift alone might explain the variation among gene frequencies across populations and genes. A measure of variation among populations was calculated for each polymorphism, and its distribution for the 100 polymorphisms was compared with that expected for a drift-only hypothesis. At least two-thirds of the polymorphisms appear to be selectively neutral, but there are significant deviations at the two ends of the observed distribution of the measure of variation: a slight excess of polymorphisms with low variation and a greater excess with high variation. This indicates that a few DNA polymorphisms are affected by natural selection, rarely heterotic, and more often disruptive, while most are selectively neutral.

Animals↗

A single-nucleus transcriptome atlas of soybean anthers.

Anther development is crucial for plant sexual reproduction. However, a high-resolution, cell-type-specific transcriptomic atlas of this process is lacking for the legume crop soybean (Glycine max). Here, we construct a comprehensive transcriptional atlas of developing soybean anthers using single-nucleus RNA sequencing (snRNA-seq). We identify and characterize nine distinct cell types spanning both somatic and reproductive lineages. Our analysis reveals robust transcriptional continuity across anther developmental stages and dynamic reprogramming during key transitions. Notably, the shift from diploid meiocytes to haploid unicellular microspores is marked by the induction of previously inactive genes, despite an overall reduction in transcript abundance. Subsequently, within bicellular microspores, generative and vegetative cell lineages exhibit sharply divergent transcriptional programs: generative cells specialize in mRNA export and turnover, whereas vegetative cells up-regulate translational machinery. Evolutionary analysis further indicates that generative-cell-specific genes are subject to more relaxed purifying selection compared to those specific to vegetative cells. Functional validation using mutants generated by CRISPR/Cas9-mediated genome editing and EMS mutagenesis reveals the essential roles of OSD1A and PKSA in pollen development and fertility. This high-resolution atlas provides fundamental insights into the transcriptional regulation of soybean anther development and serves as a valuable resource for manipulating male fertility to advance hybrid breeding programs. The data are available at https://databases.genedenovo.com/pollen.

Glycine max↗

Identification of rice DUF1719 gene family and analysis of alkaline tolerance function of OsDUF1719.8.

Alkaline stress severely constrains the physiological metabolism and growth and development of rice through high pH and ionic toxicity. Domains of unknown function (DUF) play significant roles in plant stress responses. However, the function of the DUF1719 family (PF08224) in rice has not been reported and further research is needed. This study systematically identified the OsDUF1719 gene family in rice and investigated the function of OsDUF1719.8 under alkaline stress. The results demonstrate that the rice DUF1719 family comprises 13 protein members, all containing the PF08224 domain. It is predicted that this domain may play a role in ATPase activation. Evolutionary analysis divided DUF1719 proteins from eight grass species into six subgroups, with highly conserved gene structures, motifs, and tertiary architectures within each subgroup. Promoter analysis indicated enrichment of stress- and hormone-responsive elements, implying broad involvement in stress regulation. Expression analysis revealed that several genes, including OsDUF1719.5 and OsDUF1719.8, were upregulated under multiple abiotic stresses. Notably, OsDUF1719.8 was strongly induced during early alkaline stress. Consequently, we further analyzed the function of OsDUF1719.8 in the rice alkaline stress response. The results demonstrate that overexpression of OsDUF1719.8 enhanced rice alkaline tolerance, whereas knockout mutants exhibited stress sensitivity. OsDUF1719.8 enhances rice tolerance to alkaline stress by coordinately regulating reactive oxygen species metabolism, promoting the accumulation of osmotic adjustment compounds, and modulating ion homeostasis. This study provides the first systematic identification of the DUF1719 family and elucidates the function of OsDUF1719.8 in positively regulating rice alkaline tolerance, offering a novel gene for alkali-tolerant molecular breeding of rice.

Oryza↗

Sequence analysis of hepatitis C virus genotypes 1 to 5 reveals multiple novel subtypes in the Benelux countries.

Hepatitis C virus (HCV) isolates from a cohort of 315 patients from the Benelux countries (Belgium, The Netherlands, Luxembourg) were genotyped by means of reverse hybridization Inno-LiPA (line probe assay). Genotypes 1a, 1b, 2a, 2b, 3a, 4a and 5a were detected. From the cohort, isolates representing all types and those showing an aberrant LiPA pattern were further analysed by sequencing parts of the 5' UTR, core (nt 1 to 326; aa residues 1 to 108) and core/E1 (nt 477 to 924; aa residues 159 to 308) regions. Molecular evolutionary analysis of the core and core/E1 regions allowed discrimination between known and additional subtypes, especially within types 2 and 4. The core region is not suitable for classification of new subtypes because of the relatively high level of conservation. The core/E1 region displays a higher level of sequence variation and allows much more distinct discrimination between subtypes. Genotypes 2 and 4 are particularly heterogeneous, with at least 7 and 10 subtypes, respectively. In contrast to previous reports from Europe, HCV isolates from the cohort constituted a highly heterogeneous population of virus variants, especially within genotypes 2 and 4.

Belgium↗

Expansion and molecular evolution of the interferon-induced 2'-5' oligoadenylate synthetase gene family.

The mammalian 2'-5' oligoadenylate synthetases (2'-5'OASs) are enzymes that are crucial in the interferon-induced antiviral response. They catalyze the polymerization of ATP into 2'-5'-linked oligoadenylates which activate a constitutively expressed latent endonuclease, RNaseL, to block viral replication at the level of mRNA degradation. A molecular evolutionary analysis of available OAS sequences suggests that the vertebrate genes are members of a multigene family with its roots in the early history of tetrapods. The modern mammalian 2'-5'OAS genes underwent successive gene duplication events resulting in three size classes of enzymes, containing one, two, or three homologous domains. Expansion of the OAS gene family occurred by whole-gene duplications to increase gene content and by domain couplings to produce the multidomain genes. Evolutionary analyses show that the 2'-5'OAS genes in rodents underwent gene duplications as recently as 11 MYA and predict the existence of additional undiscovered OAS genes in mammals.

2',5'-Oligoadenylate Synthetase↗

A novel measure of genetic distance for highly polymorphic tandem repeat loci.

Genetic distance measures are indicators of relatedness among populations or species and are useful for reconstructing the historic and phylogenetic relationships among such groups. Classical measures of genetic distance were developed to analyze biochemical and serological polymorphisms, systems which generally show limited variability. However, these traditional measures of genetic distance are inadequate for the analysis of certain classes of variable number tandem repeat (VNTR) loci, which have a larger number of alleles and higher levels of heterozygosity than traditional genetic markers. At the higher levels of heterozygosity observed at these loci, the standard measures of genetic distance are nonlinear and do not account for the mutational mechanisms of hypervariable loci. We have developed a measure of genetic distance, DSW, which is appropriate for the analysis of highly polymorphic DNA loci. Using computer simulations of diverging populations, we show that DSW conforms to linearity and that the variance is similar in magnitude to traditional measures of genetic distance. Comparisons of phylogenetic trees derived from the simulated divergence of human racial groups demonstrate that the branch lengths of trees prepared using DSW are more similar to the model tree than those generated using other measures. Finally, we demonstrate the applicability of DSW to evolutionary analysis by reconstructing the relationships among eight human populations using 14 microsatellite and STR loci. The phylogenetic trees generated using DSW are different from trees constructed with traditional measures and better reflect the well-documented ancient divergence of African and non-African populations.

Animals↗

The evolution of angiosperm seed proteins: a methionine-rich legumin subfamily present in lower angiosperm clades.

Analysis of legumin-encoding cDNAs from Dioscorea caucasica Lipsky (Dioscoreaceae) and from Asarum europaeum L. (Aristolochiaceae) shows that there is an especially methionine-rich legumin subfamily present in the lower angiosperm clades including the Monocotyledoneae. It is characterized by a methionine content of 3-4 mol% which is roughly triple the methionine proportion of most other legumins. These "MetR" legumins, if present, still have to be detected in the higher angiosperms including the important seed crops. Evolutionary analysis suggests that the MetR legumins are the result of a gene duplication allowing the differentiation of legumin genes according to their sulfur content. The duplication event must have taken place before the split into mono- and dicotyledonous plants but probably after the separation of angiosperms and gymnosperms.

Amino Acid Sequence↗

On the origin of self-incompatibility haplotypes: transition through self-compatible intermediates.

Self-incompatibility (SI) in flowering plants entails the inhibition of fertilization by pollen that express specificities in common with the pistil. In species of the Solanaceae, Rosaceae, and Scrophulariaceae, the inhibiting factor is an extracellular ribonuclease (S-RNase) secreted by stylar tissue. A distinct but as yet unknown gene (provisionally called pollen-S) appears to determine the specific S-RNase from which a pollen tube accepts inhibition. The S-RNase gene and pollen-S segregate with the classically defined S-locus. The origin of a new specificity appears to require, at minimum, mutations in both genes. We explore the conditions under which new specificities may arise from an intermediate state of loss of self-recognition. Our evolutionary analysis of mutations that affect either pistil or pollen specificity indicates that natural selection favors mutations in pollen-S that reduce the set of pistils from which the pollen accepts inhibition and disfavors mutations in the S-RNase gene that cause the nonreciprocal acceptance of pollen specificities. We describe the range of parameters (rate of receipt of self-pollen and relative viability of inbred offspring) that permits the generation of a succession of new specificities. This evolutionary pathway begins with the partial breakdown of SI upon the appearance of a mutation in pollen-S that frees pollen from inhibition by any S-RNase presently in the population and ends with the restoration of SI by a mutation in the S-RNase gene that enables pistils to reject the new pollen type.

Evolution, Molecular↗

Genetic, molecular and developmental analysis of the glutamine synthetase isozymes of Drosophila melanogaster.

The glutamine synthetase isozymes of Drosophila melanogaster offer an attractive model for the study of the molecular genetics and evolution of a small gene family encoding enzymatic isoforms that evolved to assume a variety of specific and sometimes essential biological functions. In Drosophila melanogaster two GS isozymes have been described which exhibit different cellular localisation and are coded by a two-member gene family. The mitochondrial GS structural gene resides at the 21B region of the second chromosome, the structural gene for the cytosolic isoform at the 10B region of the X chromosome. cDNA clones corresponding to the two genes have been isolated and sequenced. Evolutionary analysis data are in accord with the hypothesis that the two Drosophila glutamine synthetase genes are derived from a duplication event that occurred near the time of divergence between Insecta and Vertebrata. Both isoforms catalyse all reactions catalysed by other glutamine synthetases, but the different kinetic parameters and the different cellular compartmentalisation suggest strong functional specialisation. In fact, mutations of the mitochondrial GS gene produce embryo-lethal female sterility, defining a function of the gene product essential for the early stages of embryonic development. Preliminary results show strikingly distinct spatial and temporal patterns of expression of the two isoforms at later stages of development.

Animals↗

Comparison of phylogenetic metrics of transmission in symptomatic and asymptomatic tuberculosis.

BACKGROUND: Understanding drivers of Mycobacterium tuberculosis (Mtb) transmission remains a critical challenge in high-burden settings. Tuberculosis control efforts traditionally target symptomatic individuals, yet the role of asymptomatic cases in sustaining transmission is increasing recognized. METHODS: We conducted a genomic and epidemiological analysis of Mtb isolates collected in Mato Grosso do Sul, Brazil, between 2008 and 2024. From 2017 to 2022, active case finding was performed in three of the state's largest prisons, whereby sputum was collected from individuals irrespective of symptoms and tested by GeneXpert and culture. We evaluated several metrics of recent transmission from symptomatic and asymptomatic individuals, including phylogenetic clustering, Time-scaled Haplotype Density (THD), Local Branching Index (LBI), and transmission probabilities inferred using the Bayesian Reconstruction and Evolutionary Analysis of Transmission Histories (BREATH). FINDINGS: We sequenced 2,362 Mtb strains, of which 3.5% (115/2,362) were resistant to at least one drug, and 0.6% (16/2,362) were multi-drug resistant. Most strains were lineage 4, and 78.2% of all isolates were part of a genomic cluster. Among 2,362 individuals with tuberculosis, 1,137 were incarcerated at the time of diagnosis. Among these, 505 were identified through active case finding: 277 had symptomatic disease and 228 had asymptomatic tuberculosis. There was no significant difference in phylogenetic clustering proportion (77% vs. 85%; p= 0.816), THD (median 0.50 vs. 0.39; p = 0.120), or LBI (median 0.00863 vs. 0.00871; p = 0.086) between symptomatic and asymptomatic individuals. Bayesian transmission trees revealed no significant difference in the number of secondary infections inferred from symptomatic compared with asymptomatic individuals (p = 0.56). These findings were consistent across genomic clusters and robust to model assumptions. INTERPRETATION: We identified no differences in transmission from symptomatic compared with asymptomatic individuals, using several genomic measures of transmission, underscoring the substantial contribution that asymptomatic tuberculosis makes to transmission at the population level.

Asymptomatic↗

Transcriptome-wide analysis reveals sequence selection to avoid mRNA aggregation in E. coli.

The stability of RNA base pairing and its limited four-letter code create an intrinsic potential for promiscuous RNA-RNA interactions. In vitro, such interactions drive RNA to self-assemble into aggregates. This raises a fundamental unanswered question: within a confined cellular volume at physiological mRNA abundances, how much aggregation would arise from sequence-encoded chemistry alone? Here, we establish this baseline with large-scale kinetic simulations of the E. coli transcriptome. Our simulations reveal that sequence-encoded base-pairing energetics is sufficient to generate a dynamic network of large aggregates, organized by long, multivalent mRNA hubs. Strikingly, evolutionary analysis shows that native E. coli sequences exhibit clear signatures of selection to counteract this propensity: they fold more stably, minimize unstructured regions, and form weaker intermolecular contacts than dinucleotide-preserving controls. These findings demonstrate that maintaining transcriptome solubility has been a significant, previously unrecognized constraint shaping genome evolution, and provide a new lens to interpret cellular RNA management.

Biological Sciences (Biophysics and Computational ↗