PubMed HealthSearch

SEARCH · PubMed Health

Results for “evolutionary analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A second gene for the African green monkey poliovirus receptor that has no putative N-glycosylation site in the functional N-terminal immunoglobulin-like domain.

Using cDNA of the human poliovirus receptor (PVR) as a probe, two types of cDNA clones of the monkey homologs were isolated from a cDNA library prepared from an African green monkey kidney cell line. Either type of cDNA clone rendered mouse L cells permissive for poliovirus infection. Homologies of the amino acid sequences deduced from these cDNA sequences with that of human PVR were 90.2 and 86.4%, respectively. These two monkey PVRs were found to be encoded in two different loci of the genome. Evolutionary analysis suggested that duplication of the PVR gene in the monkey genome had occurred after the species differentiation between humans and monkeys. The NH2-terminal immunoglobulin-like domain, domain 1, of the second monkey PVR, which lacks a putative N-glycosylation site, mediated poliovirus infection. In addition, a human PVR mutant without N-glycosylation sites in domain 1 also promoted viral infection. These results suggest that domain 1 of the monkey receptor also harbors the binding site for poliovirus and that sugar moieties possibly attached to this domain of human PVR are dispensable for the virus-receptor interaction.

Amino Acid Sequence

Identification of a novel HIV-1 circulating recombinant form (CRF209_cpx) and its descendant unique recombinant form (URF) CRF209_cpx/B among MSM in Guangdong, southern China.

BACKGROUND: The epidemic of human immunodeficiency virus type 1 (HIV-1) continues to pose a significant global health challenge, with increasing genetic diversity. The co-circulation of multiple subtypes among the local population facilitates the emergence of unique or circulating recombinant forms (URFs or CRFs). In China, the predominant strains include CRF07_BC, CRF01_AE, CRF55_01B, and subtype B. This study characterizes a novel CRF209_cpx and its descendant recombinant CRF209_cpx/B among men who have sex with men (MSM) in Guangdong, southern China. METHODS: Individuals infected with URFs with similar genetic characteristics were recruited during routine surveillance of pretreatment drug resistance. Near full-length genomes (NFLGs) were amplified with two overlapping fragments using a serial dilution nested PCR approach after reverse transcription. We used SimPlot and IQ-TREE softwares to conduct recombination analyses and phylogenetic inferences. Time-scaled maximum clade credibility (MCC) phylogenetic trees were reconstructed using BEAST software to estimate evolutionary origins. Genotypic drug resistance mutations were interpreted via the Stanford HIV Database, and coreceptor usage was predicted using geno2pheno coreceptor 2.5 and the HIVcoPRED tool. RESULTS: Four NFLG sequences were obtained and identified as a novel CRF209_cpx, generated by recombination among CRF01_AE, CRF07_BC and subtype B. Phylogenetic analyses revealed that all the parental segments clustered with lineages prevalent among MSM in China. Bayesian evolutionary analysis estimated that the most recent common ancestor (tMRCA) of CRF209_cpx to have evolved between 2011 and 2013. The fifth strain was identified as a URF recombined from nascent CRF209_cpx and B. No transmitted drug resistance mutation was detected in these five sequences. The four CRF209_cpx sequences primarily utilized the CXCR4 coreceptor, while the URF exhibited R5/X4 dual tropism. CONCLUSIONS: The emergence of the complex CRF209_cpx and novel URF of CRF209_cpx/B highlights the active HIV-1 epidemic within the MSM population in Guangdong, underscoring the necessity for enhanced molecular surveillance and precise public health intervention in this key population.

HIV-1

The Myb DNA-binding domain is highly conserved in Dictyostelium discoideum.

The c-myb proto-oncogene encodes a protein that is highly conserved among birds and mammals. The amino-terminal domain of c-Myb contains three imperfect tandem repeats of approximately 50 amino acids each. This domain is required for DNA binding and has also been conserved to varying degrees in invertebrates, plants and yeast. Given that myb-related genes appear to control cellular differentiation in a variety of eucaryotic systems, the presence of a myb gene in the cellular slime mold Dictyostelium discoideum might provide a tractable system for studying the role of myb in differentiation. Degenerate oligonucleotide primers encoding regions that are highly conserved in the vertebrate and Drosophila Myb DNA-binding domains were used to amplify a related domain from Dictyostelium genomic DNA, which was then used to isolate a genomic clone. The putative DNA-binding domain of Dictyostelium Myb is as closely related to vertebrate c-Myb as is Drosophila Myb (65% identity), whereas the known Myb-related proteins of plants and yeast are more distantly related. The conserved domain of Dictyostelium Myb is capable of binding to the same DNA sequence as the vertebrate and Drosophila Myb proteins. The remainder of the deduced amino acid sequence of Dictyostelium Myb shows no homology to the divergent domains of the known animal, plant and yeast Myb-related proteins. Evolutionary analysis implies that the duplications that generated the repeats of the Myb DNA-binding domain began prior to the divergence of animals, plants, cellular slime molds and yeast.

Amino Acid Sequence

Evolutionary and Functional Analysis of Caspase-8 and ASC Interactions to Drive Lytic Cell Death, PANoptosis.

Caspases are evolutionarily conserved proteins essential for driving cell death in development and host defense. Caspase-8, a key member of the caspase family, is implicated in nonlytic apoptosis, as well as lytic forms of cell death. Recently, caspase-8 has been identified as an integral component of PANoptosomes, multiprotein complexes formed in response to innate immune sensor activation. Several innate immune sensors can nucleate caspase-8-containing PANoptosome complexes to drive inflammatory lytic cell death, PANoptosis. However, how the evolutionarily conserved and diverse functions of caspase-8 drive PANoptosis remains unclear. To address this, we performed evolutionary, sequence, structural, and functional analyses to decode caspase-8's complex-forming abilities and its interaction with the PANoptosome adaptor ASC. Our study distinguished distinct subgroups within the death domain superfamily based on their evolutionary and functional relationships, identified homotypic traits among subfamily members, and captured key events in caspase evolution. We also identified critical residues defining the heterotypic interaction between caspase-8's death effector domain and ASC's pyrin domain, validated through cross-species analyses, dynamic simulations, and in vitro experiments. Overall, our study elucidated recent evolutionary adaptations of caspase-8 that allowed it to interact with ASC, improving our understanding of critical molecular associations in PANoptosome complex formation and the underlying PANoptotic responses in host defense and inflammation. These findings have implications for understanding mammalian immune responses and developing new therapeutic strategies for inflammatory diseases.

Caspase 8

The nucleotide sequence of 3C proteinase region of the coxsackievirus A24 variant: comparison of the isolates in Taiwan in 1985-1988.

Acute hemorrhagic conjunctivitis caused by coxsackievirus A24 variant (CA24v) first appeared in Taiwan in October 1985, followed by two other sequential epidemics in 1986 and 1988. In order to know the evolutionary relationship of the CA24v strains isolated in Taiwan, we first determined the nucleotide sequence of the 3C proteinase (3Cpro) region of the prototype strain (EH24/70), isolated in Singapore in 1970, by molecular cloning. The nucleotide sequence of the 3Cpro region thus sequenced showed striking homology with polioviruses and coxsackievirus A21. Viral RNA of eight isolates obtained from the three epidemics was reverse transcribed, amplified by the polymerase chain reaction, and cloned into M13 phage for the production of ssDNA for nucleotide sequencing by the dideoxy chain termination method. When the number of nucleotide difference was taken as a genetic distance between isolates, all isolates showed a very similar distance from the EH24/70, the earliest isolate of CA24v, indicating that they evolved at a constant evolutionary rate. Phylogenetic analysis by the unweighted pairwise grouping method of arithmetic average (UPGMA) indicated that the six isolates collected in 1985 and 1986 were closely related, while two 1988 isolates were more distant from them. The branching time between these two groups was estimated to be May 1984, 18 months before the first recognition of the CA24v epidemic in Taiwan. This is the first report of the nucleotide sequence of CA24v genome RNA and of an evolutionary analysis of the virus using the nucleotide sequence.

3C Viral Proteases

Functional Prediction of Epitranscriptome.

N6-methyladenosine (m6A) is one of the most prevalent and well-studied RNA modifications, playing a pivotal role in many biological processes. With the recent advances in high-throughput sequencing technologies, tens of thousands of m6A sites have been reported. However, not all m6A sites are important or functionally significant, highlighting the need to distinguish biologically relevant m6As from non-functional or technically artefactual ones. Here, we describe ConsRM, which is a web-based resource that was designed to evaluate the importance of m6As from an evolutionary perspective. It introduced a novel scoring framework for quantifying the conservation degree of m6As in humans. Its web interface includes a database of 177998 distinct human m6A sites along with their calculated conservation score, and allows users to analyze their own data via the web server. ConsRM is freely accessible at: http://180.208.58.19/conservation/browser.html .

Humans

MaizeGDB Phylostrata Tool: exploring evolutionary origins of maize proteins.

MOTIVATION: Phylostratigraphic analysis identifies the evolutionary origins and level of conservation of proteins, facilitating research in evolutionary biology and comparative genomics. RESULTS: We developed the MaizeGDB Phylostrata Tool, a custom web application that enables users to explore the evolutionary origins of proteins in maize (Zea mays), a globally important crop and model organism. This tool features interactive visualizations and detailed gene pages incorporating subcellular localization, Gene Ontology (GO) terms, and links to resources for homologs, facilitating comparison of gene functions across evolutionary time. The tool also provides downloadable links for full-proteome phylostratigraphic results for 26 maize inbreds (B73 and the NAM founders). From these, we identified genome- and subgenome-wide trends, finding that more conserved proteins tended to be longer and more highly expressed. Finally, we provide code including updates to the "phylostratr" R package to make it more robust against taxonomic updates, as well as example scripts for phylostratigraphic analysis and web tool development for researchers and curators of other species. AVAILABILITY AND IMPLEMENTATION: The MaizeGDB Phylostrata Tool is freely available at https://phylostrata.maizegdb.org. Scripts used for the analysis and web tool are available at https://github.com/LTibbs/PhylostrataWebtool.

Journal Article

A nonstationarity test for the spectral analysis of physiological time series with an application to respiratory sinus arrhythmia.

The spectral analysis of time series requires the signal to be at least weakly stationary; i.e., the mean, (co-) variance, and spectrum of the time series should not vary from segment to segment. It is commonly assumed that psychophysiological time series are not stationary. This study introduces a nonstationarity test to the psychophysiological literature, which is derived from evolutionary spectral analysis. Basically, the test consists of a double window technique in both the time and frequency domains, leading to a two-way analysis of variance for times and frequencies. In the current study, the nonstationarity test is applied to heart rate data obtained in a typical psychophysiological setting. Heart rate and respiration were measured in four age groups under four conditions--rest, paced breathing, vigilance, and reaction time. The results indicate that only few physiological time series were completely stationary. However, for every subject, and in every condition stationary stretches could be found that were long enough to apply spectral analysis. Spectral measures (power, coherence, and phase spectra) were then compared for stationary parts of the data and the total data. This comparison indicated that nonstationarity affects all spectral measures. Most importantly, Stationarity x Task Condition x Frequency Band interactions were observed for coherence and phase spectra, and there were significant interactions with age for each of the spectral indices. These findings suggest that nonstationarity may result in biased outcomes of significance tests of the effects of task manipulations on the spectral indices of cardiac time series. Thus, it was concluded that the stationarity test should be routinely applied in the spectral analysis of physiological time series. In addition, it was suggested that the nonstationarity test has an even wider range of application that might be of interest to the psychophysiologist.

Adult

Genome-wide identification and analysis of paclobutrazol-resistance gene family in cotton and the positive role of GhPRE3 in salt stress and drought stress resistance.

Compared with other transcription factors, much less studies have been performed on paclobutrazol-resistance (PRE), a subgroup of the extensive bHLH transcription factor gene family, and the research in cotton was also limited. By utilizing the PRE genes and their conserved domains identified in Arabidopsis, a total of 23, 22, 11, and 12 PRE genes were identified from two major cultivated cotton species and their two ancestors, respectively. The cotton PRE gene family was categorized into three subgroups based on evolutionary tree analysis. Motif and intron analyses indicated that the PRE gene has remained highly conserved throughout evolution. Collinearity analysis indicated that gene duplication, particularly through fragment replication, has significantly contributed to the expansion of the cotton PRE family. An exploration of the conserved elements within the PRE gene family uncovered numerous elements associated with plant stress resistance. Additionally, cotton transcriptome and qRT-PCR analysis showed that PRE genes were associated with a variety of abiotic stresses, including salt, drought, and cold treatments. Subcellular localization experiments indicated that the GhPRE3 gene is associated with membrane proteins. Finally, we selected the GhPRE3 gene for a VIGS experiment, which revealed that under salt stress and drought stress conditions, the wilting of leaves in the GhPRE3-silenced plants was significantly more severe than that observed in the control group, with T-AOC levels notably lower and MDA levels significantly higher. Overexpression of GhPRE3 enhanced seed germination and root development in transgenic Arabidopsis thaliana under salt stress and drought stresses. This suggests that GhPRE3 plays a positive regulatory role in cotton tolerance to salt and drought stressed, providing a reference for molecular genetic breeding of cotton with salt and drought tolerance.

Gossypium

Identification of nucleotide substitutions necessary for trans-activation of mariner transposable elements in Drosophila: analysis of naturally occurring elements.

Six copies of the mariner element from the genomes of Drosophila mauritiana and Drosophila simulans were chosen at random for DNA sequencing and functional analysis and compared with the highly active element Mos1 and the inactive element peach. All elements were 1286 base pairs in length, but among them there were 18 nucleotide differences. As assayed in Drosophila melanogaster, three of the elements were apparently nonfunctional, two were marginally functional, and one had moderate activity that could be greatly increased depending on the position of the element in the genome. Both molecular (site-directed mutagenesis) and evolutionary (cladistic analysis) techniques were used to analyze the functional effects of nucleotide substitutions. The nucleotide sequence of the element is the primary determinant of function, though the activity level of elements is profoundly influenced by position effects. Cladistic analysis of the sequences has identified a T----A transversion at position 1203 (resulting in a Phe----Leu amino acid replacement in the putative transposase) as being primarily responsible for the low activity of the barely functional elements. Use of the sequences from the more distantly related species, Drosophila yakuba and Drosophila teissieri, as outside reference species, indicates that functional mariner elements are ancestral and argues against their origination by a novel mutation or by recombination among nonfunctional elements.

Animals

Genome-wide characterization of the tomato PERK gene family and its expression profiling under abiotic stresses.

UNLABELLED: This study presents the first systematic genome-wide characterization of the proline-rich extensin-like receptor kinases (PERK) gene family in tomato (Solanum lycopersicum) and their transcriptional responses under abiotic stresses. Using the latest SL4.0/ITAG4.0 genome assembly, we identified six SlPERK genes, all harboring the conserved Ser/Thr protein kinase domain. Evolutionary and structural analyses revealed strong purifying selection (Ka/Ks&#x2009;<&#x2009;1), distinct exon-intron organizations, and the presence of stress- and hormone-responsive cis-regulatory elements in their promoters. Furthermore, post-transcriptional regulation by 57 miRNAs and complex protein-protein interaction networks were predicted. To validate their stress-responsive roles, two tomato cultivars (GMOTL-1 and Roma) were subjected to cold, heat, and salinity treatments. Quantitative RT-PCR analysis revealed cultivar-specific expression dynamics: SlPERK4 exhibited strong transient induction under cold and heat stress, while SlPERK6 was highly responsive to salinity. Notably, the GMOTL-1 cultivar displayed significantly higher and broader stress-responsive expression profiles compared to Roma, indicating a potential role of these SlPERK genes in cultivar-specific stress tolerance. These findings provide a comprehensive genomic resource and establish a critical foundation for the functional validation and molecular breeding for stress-resilience tomato cultivars. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05044-y.

Abiotic stress

ERCnet: Phylogenomic Prediction of Interaction Networks in the Presence of Gene Duplication.

Assigning gene function from genome sequences is a rate-limiting step in molecular biology research. A protein's position within an interaction network can potentially provide insights into its molecular mechanisms. Phylogenetic analysis of evolutionary rate covariation (ERC) in protein sequence has been shown to be effective for large-scale prediction of functional relationships and interactions. However, gene duplication, gene loss, and other sources of phylogenetic incongruence are barriers for analyzing ERC on a genome-wide basis. Here, we developed ERCnet, a bioinformatic program designed to overcome these challenges, facilitating efficient all-versus-all ERC analyses for large protein sequence datasets. We simulated proteome datasets and found that ERCnet achieves combined false positive and negative error rates well below 10% and that our novel "branch-by-branch" length measurements outperforms "root-to-tip" approaches in most cases, offering a valuable new strategy for performing ERC. We also compiled a sample set of 35 angiosperm genomes to test the performance of ERCnet on empirical data, including its sensitivity to user-defined analysis parameters such as input dataset size and branch-length measurement strategy. We investigated the overlap between ERCnet runs with different species samples to understand how species number and composition affect predicted interactions and to identify the protein sets that consistently exhibit ERC across angiosperms. Our systematic exploration of the performance of ERCnet provides a roadmap for design of future ERC analyses to predict functional interactions in a wide array of genomic datasets. ERCnet code is freely available at https://github.com/EvanForsythe/ERCnet.

Gene Duplication

Genomic characteristics and tracing analysis of an acute gastroenteritis outbreak associated with rotavirus C in a boarding high school.

BACKGROUND: Rotaviruses are major pathogens of childhood acute gastroenteritis, dominated by rotavirus A (RVA). Outbreaks caused by human rotavirus C (RVC) are rarely reported, and relevant genomic data remain scarce. This genomic investigation of an RVC outbreak improves our understanding of viral diversity and transmission dynamics. METHODS: We performed epidemiological surveys, nucleic acid testing and whole-genome sequencing (WGS) on specimens from a 2025 RVC-associated gastroenteritis outbreak at a Chinese boarding high school. Sequence alignment, phylogenetic and molecular tracing analyses were conducted to explore RVC evolution via point mutation, segment reassortment and genomic recombination. RESULTS: This typical point-source campus outbreak was linked to an indoor student gathering matching the incubation period of RVC. Thirteen RVC FX strains were recovered from 11 rectal swabs and two vomitus samples. Their viral protein (VP) 4 and VP7 sequences shared high homology with Russian reference strains, carrying distinct amino acid variations. No segment reassortment or recombination was detected in VP4/VP7 genes. CONCLUSIONS: Dense, closed campus settings facilitate RVC clustered transmission. Limitations included absent screening of asymptomatic canteen staff. Rapid nucleic acid testing enabled timely pathogen identification for outbreak control. Greater attention should be paid to the public health risk of RVC. These whole-genome sequencing data enrich resources for studying RVC evolution and vaccine development.

Acute gastroenteritis outbreak

Genetic Variation and Evolutionary Characteristics of Coxsackievirus B1: F3 Subtype Associated With Hand, Foot and Mouth Disease in China.

Coxsackievirus B1 (CV-B1) is primarily associated with meningitis but can also cause localized outbreaks of hand, foot, and mouth disease (HFMD). This study analyzed the genetic diversity of the VP1 gene in 39 strains of the CVB1 virus isolated from HFMD children across 15 provinces in China between 2010 and 2024, as well as 179 strains from 17 countries. Based on the average nucleotide difference of VP1 gene, we classified CVB1 virus into six genotypes A to F, Notably, genotype F is newly classified. Since 2010, genotype F guadually replaced genotype E as the dominant genotype in China and has further subdivided into three subtypes: F1, F2, and F3, with F3 being the most prevalent subtype in China currently. We specifically study the mild and severe cases within the F3 subtype. Temperature-sensitivity experiments revealed no differences between mild and severe cases of the F3 subtype, and they all belong to temperature-sensitive strains. Interestingly, we found that mild cases of the F3 subtype did not involve recombination, whereas all severe cases of the F3 subtype showed recombination with Coxsackievirus B4 (CVB4). CVB4 has consistently been the primary pathogen responsible for severe neonatal illnesses, suggesting that recombination between the F3 subtype and CVB4 may be associated with the development of severe HFMD. These findings provide fundamental scientific data for further investigation into the epidemiology and genetic characteristics of variants of Coxsackievirus B1 in China.

Humans

Evolution of the primate lentiviruses: evidence from vpx and vpr.

The genomes of the four primate lentiviral groups are complex and contain several regulatory or accessory genes. Two of these genes, vpr and vpx, are found in various combinations within the four groups and encode proteins whose functions have yet to be elucidated. Comparison of the encoded protein sequences suggests that the vpx gene within the HIV-2 group arose by the duplication of an ancestral vpr gene within this group. Evolutionary distance analysis showed that both genes were well conserved when compared with viral regulatory genes, and indicated that the duplication occurred at approximately the same time as the HIV-2 group and the other primate lentivirus groups diverged from a common ancestor. Furthermore, although the SIVagm vpx proteins are homologous to the HIV-2 group vpx proteins, there are insufficient grounds from sequence analysis for classifying them as vpx proteins. Because of their similarity to the vpr proteins of other groups, we suggest reclassifying the SIVagm vpx gene as a vpr gene. This creates a simpler and more uniform picture of the genomic organization of the primate lentiviruses and allows the genomic organization of their common precursor to be defined; it probably contained five accessory genes: tat, rev, vif, nef and vpr.

Amino Acid Sequence

Application of genome sequence information in potyvirus taxonomy: an overview.

The application of protein and nucleic acid sequence analysis in evolutionary and phylogenetic studies is well established. Available sequence information for the 5' untranslated region of potyviruses including the fungus-transmitted barley yellow mosaic virus (BaYMV) RNA-1 suggests that a 12-nucleotide conserved sequence, the "potybox" is unique to this group. Various non-structural proteins of potyviruses share considerable "signature" sequence homology across a broad spectrum of unrelated viruses, which makes their value limited to "supergroup" or "superfamily" identity. However, in potyviruses, the coat-protein N-terminal sequences and 3' noncoding regions are variable among viruses, but similar among strains of the same virus. This suggests that these sequences may be an accurate marker of genetic relatedness. Until complete genome sequences from a large number of potyviruses become available and their value in systematics is tested, coat protein and 3' noncoding regions remain as the choice of taxonomic indicators. The reason being, that cloning and sequencing of the coat-protein gene and 3' noncoding regions are less complicated and time consuming and the sequences show significant differences among the virus species within the family Potyviridae.

Animals

Sequence and diversity of rhesus monkey T-cell receptor beta chain genes.

We have sequenced 23 rearranged T-cell receptor beta chain (Tcrb) cDNA clones derived from peripheral blood lymphocytes (PBL) of a rhesus monkey. All of the clones have a variable-diversity-joining-constant (V-D-J-C) rearrangement similar to that of humans. Two rhesus constant (C) region genes were found, each closely resembling human Cb 1 and 2. All of the rhesus J region sequences align well with ten of the 13 reported human J regions. 17 of the 23 rhesus V region sequences could be assigned to families homologous with eight different human families (Vb 1, 2, 6, 7, 8, 9, 13, and 14). The remaining six V region sequences are more distantly related to human Vb 1 and 13. Thus, the organization and sequences of studied rhesus Tcrb chains resemble human homologs. An evolutionary tree analysis revealed paralogous relationships between specific members of the rhesus and human V region families. Analysis of synonymous and nonsynonymous nucleotide sequence differences indicated that the evolution of the presumed major histocompatibility complex (MHC)-contact regions of the Tcrb chains is less constrained than that of the framework regions.

Amino Acid Sequence