PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “evolutionary analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Bovine seminal plasma proteins and their relatives: A new expanding superfamily in mammals.

BSP proteins represent three major proteins of bovine seminal plasma: BSP-A1/-A2, -A3 and -30 kDa. The BSP protein signature is characterized by two tandemly repeated fibronectin type 2 (Fn2) domains. Although classical affinity chromatography and protein sequencing have proven that the BSP protein homologs may be ubiquitous in mammals and functionally related to sperm capacitation, only the three bovine genes have been reported thus far. In this study, we report three new BSP protein-related genes from bovine, as well as other BSP protein-related DNA sequences from human, chimpanzee, mouse, rat, dog, horse and rabbit. Analysis of the relationships between all Fn2 domain-containing proteins revealed that the Fn2 domains found in BSP-related proteins have special features that distinguish them from non-BSP-related proteins. These features can be used to identify new BSP protein-related sequences. Further molecular evolutionary analysis of the BSP protein lineage revealed that all BSP proteins and their related sequences can be grouped into three subfamilies: BSPH4, BSPH5 and BSPH6, which indicates that the BSP protein family is much bigger than previously envisioned. More interestingly, the three BSP proteins in bovine within the BSPH4-subfamily were shown to evolve rapidly. The ratio of nonsynonymous to synonymous substitutions was higher than 1. The analysis also indicated that the rate of evolution was heterogeneous between the first and second Fn2 domains of the genes. These data may reflect that some amino acids in BSP proteins are under a strong positive selection after gene duplication and that each BSP protein evolves rapidly, possibly to acquire new functions.

Amino Acid Sequence↗

Understanding Mycobacterium tuberculosis through its genomic diversity and evolution.

Pathogen evolution and genomic diversity are shaped by specific host immune pressures and therapeutic interventions. Analysis of the extant genomes of circulating strains of Mycobacterium tuberculosis, a leading cause of infectious mortality that has co-evolved with humans for thousands of years, can provide new insights into host-pathogen interactions that underlie specific aspects of pathogenesis and onward transmission. With the explosion in the number of fully sequenced M. tuberculosis strains that are now paired with detailed clinical data, there are new opportunities to understand the evolutionary basis for and consequences of M. tuberculosis strain diversity. This review examines mechanistic findings that have emerged from pairing whole genome sequencing data and evolutionary analysis with functional dissection of specific bacterial variants. These include improved understanding of secreted effectors that modulate the properties and migratory behavior of infected macrophages as well as bacterial genetic alterations important for survival within hypoxic microenvironments. Genomic, evolutionary, and functional analyses across diverse M. tuberculosis strains will identify prominent bacterial adaptations to their human hosts and shape our understanding of TB disease biology and the host immune response.

Mycobacterium tuberculosis↗

Personality disorder as harmful dysfunction: DSM'S cultural deviance criterion reconsidered.

The DSM's general criteria for personality disorder (PD) attempt to define PD versus nondisordered personality conditions. If dimensionalization of PD occurs in the DSM-V (perhaps, it is suggested, with PD diagnosis moved to Axis I and overall personality assessment in Axis II, thus separating diagnosis from case formulation), general criteria likely will still be needed to prevent massive false positives. In this article, one of the general criteria, the cultural deviance requirement (CDR), is examined from the perspective of the evolution-based harmful-dysfunction analysis of disorder. The CDR is often assumed to express value relativity of harm in diagnosis, but cultural values are a designed feature of human social functioning that influence personality formation. The CDR is thus argued to be an indicator of whether an individual's personality organization is due to an evolutionary dysfunction. Value relativity and evolutionary analysis thus converge.

Comorbidity↗

Application of DETECTER, an evolutionary genomic tool to analyze genetic variation, to the cystic fibrosis gene family.

BACKGROUND: The medical community requires computational tools that distinguish missense genetic differences having phenotypic impact within the vast number of sense mutations that do not. Tools that do this will become increasingly important for those seeking to use human genome sequence data to predict disease, make prognoses, and customize therapy to individual patients. RESULTS: An approach, termed DETECTER, is proposed to identify sites in a protein sequence where amino acid replacements are likely to have a significant effect on phenotype, including causing genetic disease. This approach uses a model-dependent tool to estimate the normalized replacement rate at individual sites in a protein sequence, based on a history of those sites extracted from an evolutionary analysis of the corresponding protein family. This tool identifies sites that have higher-than-average, average, or lower-than-average rates of change in the lineage leading to the sequence in the population of interest. The rates are then combined with sequence data to determine the likelihoods that particular amino acids were present at individual sites in the evolutionary history of the gene family. These likelihoods are used to predict whether any specific amino acid replacements, if introduced at the site in a modern human population, would have a significant impact on fitness. The DETECTER tool is used to analyze the cystic fibrosis transmembrane conductance regulator (CFTR) gene family. CONCLUSION: In this system, DETECTER retrodicts amino acid replacements associated with the cystic fibrosis disease with greater accuracy than alternative approaches. While this result validates this approach for this particular family of proteins only, the approach may be applicable to the analysis of polymorphisms generally, including SNPs in a human population.

Amino Acid Substitution↗

Application of nucleotide sequence of RNA polymerase beta-subunit gene (rpoB) to molecular differentiation of serovars of Salmonella enterica subsp. enterica.

To establish a molecular differentiation method for Salmonella enterica subsp. enterica, a hyper-variable region of RNA polymerase beta-subunit (rpoB) of S. enterica subsp. enterica (I), serotype Typhimurium, and Escherichia coli were investigated through comparison of nucleotide sequence of the region. The hyper-variable region was identified at 612-937 of the gene. After PCR amplification of the region in the 17 serotypes and two biotypes of serotype Gallinarum of S. enterica subsp. enterica (I), the nucleotide sequences of the region were determined and compared. All serotypes were distantly related to E. coli with 82.8-84.7% identities in nucleotide sequence while showing 96.6-100% identities with each other. According to the phylogenetic analysis based on the sequenced region with the neighbor-joining method, relatedness of biotype Gallinarum to serotype Enteritidis and biotype Pullorum was determined. Biotype Gallinarum was more closely related to serotype Enteritidis than biotype Pullorum. These results suggested that the 612-937 variable region of rpoB might be useful for molecular evolutionary analysis of serotypes of S. enterica subsp. enterica (I).

Amino Acid Sequence↗

Fast, accurate construction of multiple sequence alignments from protein language embeddings.

Multiple sequence alignment (MSA) is a foundational task in computational biology, underpinning protein structure prediction, evolutionary analysis, and domain annotation. Traditional MSA algorithms rely on pairwise amino acid substitution matrices derived from conserved protein families. While effective for aligning closely related sequences, these scoring schemes struggle in the low-identity "twilight zone." Here, we present a new approach for constructing MSAs leveraging amino acid embeddings generated by protein language models (PLMs), which capture rich evolutionary and contextual information from massive and diverse sequence datasets. We introduce a windowed reciprocal-weighted embedding similarity metric that is surprisingly effective in identifying corresponding amino acids across sequences. Building on this metric, we develop ARIES (Alignment via RecIprocal Embedding Similarity), an algorithm that constructs a PLM-generated template embedding and aligns each sequence to this template via dynamic time warping in order to build a global MSA. Across diverse benchmark datasets, ARIES achieves higher accuracies than existing state-of-the-art approaches, especially in low-identity regimes where traditional methods degrade, while scaling almost linearly with the number of sequences to be aligned. Together, these results provide the first large-scale demonstration of the power of PLMs for accurate and scalable MSA construction across protein families of varying sizes and levels of similarity, highlighting the potential of PLMs to transform comparative sequence analysis.

Deep Learning↗

The nop-1 gene of Neurospora crassa encodes a seven transmembrane helix retinal-binding protein homologous to archaeal rhodopsins.

Opsins are a class of retinal-binding, seven transmembrane helix proteins that function as light-responsive ion pumps or sensory receptors. Previously, genes encoding opsins had been identified in animals and the Archaea but not in fungi or other eukaryotic microorganisms. Here, we report the identification and mutational analysis of an opsin gene, nop-1, from the eukaryotic filamentous fungus Neurospora crassa. The nop-1 amino acid sequence predicts a protein that shares up to 81.8% amino acid identity with archaeal opsins in the 22 retinal binding pocket residues, including the conserved lysine residue that forms a Schiff base linkage with retinal. Evolutionary analysis revealed relatedness not only between NOP-1 and archaeal opsins but also between NOP-1 and several fungal opsin-related proteins that lack the Schiff base lysine residue. The results provide evidence for a eukaryotic opsin family homologous to the archaeal opsins, providing a plausible link between archaeal and visual opsins. Extensive analysis of Deltanop-1 strains did not reveal obvious defects in light-regulated processes under normal laboratory conditions. However, results from Northern analysis support light and conidiation-based regulation of nop-1 gene expression, and NOP-1 protein heterologously expressed in Pichia pastoris is labeled by using all-trans [3H]retinal, suggesting that NOP-1 functions as a rhodopsin in N. crassa photobiology.

Amino Acid Sequence↗

Controversies in the evolutionary social sciences: a guide for the perplexed.

It is 25 years since modern evolutionary ideas were first applied extensively to human behavior, jump-starting a field of study once known as 'sociobiology'. Over the years, distinct styles of evolutionary analysis have emerged within the social sciences. Although there is considerable complementarity between approaches that emphasize the study of psychological mechanisms and those that focus on adaptive fit to environments, there are also substantial theoretical and methodological differences. These differences have generated a recurrent debate that is now exacerbated by growing popular media attention to evolutionary human behavioral studies. Here, we provide a guide to current controversies surrounding evolutionary studies of human social behavior, emphasizing theoretical and methodological issues. We conclude that a greater use of formal models, measures of current fitness costs and benefits, and attention to adaptive tradeoffs, will enhance the power and reliability of evolutionary analyses of human social behavior.

Journal Article↗

The murine orthologue of human antichymotrypsin: a structural paradigm for clade A3 serpins.

Antichymotrypsin (SERPINA3) is a widely expressed member of the serpin superfamily, required for the regulation of leukocyte proteases released during an inflammatory response and with a permissive role in the development of amyloid encephalopathy. Despite its biological significance, there is at present no available structure of this serpin in its native, inhibitory state. We present here the first fully refined structure of a murine antichymotrypsin orthologue to 2.1 A, which we propose as a template for other antichymotrypsin-like serpins. A most unexpected feature of the structure of murine serpina3n is that it reveals the reactive center loop (RCL) to be partially inserted into the A beta-sheet, a structural motif associated with ligand-dependent activation in other serpins. The RCL is, in addition, stabilized by salt bridges, and its plane is oriented at 90 degrees to the RCL of antitrypsin. A biochemical and biophysical analysis of this serpin demonstrates that it is a fast and efficient inhibitor of human leukocyte elastase (ka: 4 +/- 0.9 x 10(6) m(-1) s(-)1) and cathepsin G (ka: 7.9 +/- 0.9 x 10(5) m(-1) s(-)1) giving a spectrum of activity intermediate between that of human antichymotrypsin and human antitrypsin. An evolutionary analysis reveals that residues subject to positive selection and that have contributed to the diversity of sequences in this sub-branch (A3) of the serpin superfamily are essentially restricted to the P4-P6' region of the RCL, the distal hinge, and the loop between strands 4B and 5B.

Amino Acid Sequence↗

Evolution of dispersal in a structured metapopulation model in discrete time.

In this article, a structured metapopulation model in discrete time with catastrophes and density-dependent local growth is introduced. The fitness of a rare mutant in an environment set by the resident is defined, and an efficient method to calculate fitness is presented. With this fitness measure evolutionary analysis of this model becomes feasible. This article concentrates on the evolution of dispersal. The effect of catastrophes, dispersal cost, and local dynamics on the evolution of dispersal is investigated. It is proved that without catastrophes, if all population-dynamical attractors are fixed points, there will be selection for no dispersal. A new mechanism for evolutionary branching is also found: Even though local population sizes approach fixed points, catastrophes can cause enough temporal variability, so that evolutionary branching becomes possible.

Algorithms↗

Isolation of a cDNA encoding the B isozyme of human phosphoglycerate mutase (PGAM) and characterization of the PGAM gene family.

We previously reported the isolation of a full-length cDNA specifying the muscle-specific isozyme of human phosphoglycerate mutase (PGAM-M). We now report the isolation of a full-length cDNA specifying the non-muscle-specific, or brain (B), isozyme of human PGAM (PGAM-B). The PGAM-B cDNA encodes a deduced protein 254 amino acids long, 79% identical to PGAM-M, and contains a 913-nucleotide 3'-untranslated region, as compared to the unusually short 37-nucleotide 3'-untranslated region of PGAM-M. Northern analysis demonstrates the non-muscle-specific nature of PGAM-B transcription, while genomic Southern analysis implies the presence of a large PGAM family in the human genome. Most of the PGAM-hybridizing sequences in both the human and mouse genomes seem to be related to the B-isozyme gene; many members of the PGAM-B gene family in humans are apparently processed genes. These results agree with the evolutionary analysis, which indicates that the PGAM-B gene is the progenitor of the PGAM-M gene.

Amino Acid Sequence↗

Dynamic and non-additive gene regulation shapes maize responses to simultaneous salt and cold stress.

Salt and cold stresses often occur together in nature and severely impact crop productivity, yet their transcriptional regulation remains poorly understood. Here, we conducted a time-series transcriptomic analysis of maize under salt, cold, and their combination at 0, 6, 12, and 24 h. Differential expression analysis revealed dynamic, condition-specific gene responses grouped into eight distinct temporal patterns. Promoter motif analysis of genes within each pattern identified 5-39 significantly enriched motifs, with over 40% lacking known counterparts, suggesting the involvement of previously uncharacterized cis-regulatory elements in stress-responsive transcriptional regulation. By comparing combined stress responses to the sum of single-stress effects, we found that about 74% of DEGs showed non-additive patterns, suggesting that combined stress triggers a distinct transcriptional program. Evolutionary analysis showed that additive DEGs tend to be more recently evolved, subject to weaker purifying selection, and enriched in transposed duplications, contrasting with the stronger constraint observed in non-additive DEGs. WGCNA identified 24 co-expression modules, among which 65 hub DEGs were detected in modules significantly correlated with specific stress conditions. Furthermore, we reconstructed 228, 20, and 200 sequential transcription factor cascades spanning 6 h, 12 h, and 24 h under cold, salt, and combined stress, respectively, with no cascade shared across all three conditions. Together, these results reveal that maize responses to combined salt and cold stress are largely non-additive and temporally dynamic, with distinct evolutionary patterns underlying different response types, offering insights and candidate regulators for enhancing crop stress resilience.

Zea mays↗

Thromboxane synthase (TBXAS1) polymorphisms in African-American and Caucasian populations: evidence for selective pressure.

Thromboxane synthase (TBXAS1), a cytochrome P450 enzyme, converts prostaglandin H2 into thromboxane A2, a potent vasoconstrictor and inducer of platelet aggregation. Thromboxane A2 has been implicated in modulating cell cytotoxicity and in tumor growth and metastasis. Twelve coding-region variants were identified in the human TBXAS1 gene in 48 African-American and 46 Caucasian individuals, of which eight were amino-acid substitutions. The latter were confirmed in an independent Caucasian population (n=94 unrelated individuals). We performed an evolutionary analysis of patterns of nucleotide diversity and identified patterns of amino acid replacement in human-mouse comparisons consistent with purifying selection on an inter-species time scale using the McDonald-Kreitman test. We also observed patterns of nucleotide diversity within humans consistent with purifying selection acting on existing polymorphism using Tajima's D within coding regions. These evolutionary tests suggest that some of the rare coding variations observed in the human population are deleterious. We used two sequence-homology-based software programs and molecular modeling to predict the potential impact of these polymorphisms on TBXAS1 function. The c.772C>T (p.Lys258Glu), c.1249C>G (p.Gln417Glu), and c.1348G>A (p.Glu450Lys) substitutions are predicted as most likely to alter protein function; another, c.1352C>A (p.Thr451Asn), may also affect function. Given the evolutionary evidence, these variants may be functional and therefore of relevance for disease endpoints related to inflammation and angiogenesis, as well as for the pharmacogenetics of non-steroidal anti-inflammatory drugs.

Black or African American↗

Molecular modeling and characterization of Vibrio cholerae transcription regulator HlyU.

BACKGROUND: The SmtB/ArsR family of prokaryotic metal-regulatory transcriptional repressors represses the expression of operons linked to stress-inducing concentrations of heavy metal ions, while derepression results from direct binding of metal ions by these 'metal-sensor' proteins. The HlyU protein from Vibrio cholerae is the positive regulator of haemolysin gene, it also plays important role in the regulation of expression of the virulence genes. Despite the understanding of biochemical properties, its structure and relationship to other protein families remain unknown. RESULTS: We find that HlyU exhibits structural features common to the SmtB/ArsR family of transcriptional repressors. Analysis of the modeled structure of HlyU reveals that it does not have the key metal-sensing residues which are unique to the SmtB/ArsR family of repressors, yet the tertiary structure is very similar to the family members. HlyU is the only member that has a positive control on transcription, while all the other members in the family are repressors. An evolutionary analysis with other SmtB/ArsR family members suggests that during evolution HlyU probably occurred by gene duplication and mutational events that led to the emergence of this protein from ancestral transcriptional repressor by the loss of the metal-binding sites. CONCLUSION: The study indicates that the same protein family can contain both the positive regulator of transcription and repressors--the exact function being controlled by the absence or the presence of metal-binding sites.

Bacterial Proteins↗

New methods for inferring population dynamics from microbial sequences.

The reduced cost of high throughput sequencing, increasing automation, and the amenability of sequence data for evolutionary analysis are making DNA data (or the corresponding amino acid sequences) the molecular marker of choice for studying microbial population genetics and phylogenetics. Concomitantly, due to the ever-increasing computational power, new, more accurate (and sometimes faster), sequence-based analytical approaches are being developed and applied to these new data. Here we review some commonly used, recently improved, and newly developed methodologies for inferring population dynamics and evolutionary relationships using nucleotide and amino acid sequence data, including: alignment, model selection, bifurcating and network phylogenetic approaches, and methods for estimating demographic history, population structure, and population parameters (recombination, genetic diversity, growth, and natural selection). Because of the extensive literature published on these topics this review cannot be comprehensive in its scope. Instead, for all the methods discussed we introduce the approaches we think are particularly useful for analyses of microbial sequences and where possible, include references to recent and more inclusive reviews.

Bacteria↗

Phylogenetic analysis of Ljungan virus and A-2 plaque virus, new members of the Picornaviridae.

In addition to the viruses belonging to the nine proposed genera of the Picornaviridae, Enterovirus, Rhinovirus, Cardiovirus, Aphtovirus, Hepatovirus, Parechovirus, Kobuvirus, Erbovirus and Teschovirus, two new members of this family have recently been discovered. Three strains of Ljungan virus (LV) were isolated from bank voles (Clethrionomys glareolus) and A-2 plaque virus (A-2) was isolated from human sera. To study the genetic relationship between these recently discovered viruses and the members of the family Picornaviridae, an evolutionary analysis has been carried out using the amino acid sequences of the two nonstructural proteins 2C and 3D. Phylogenetic analysis using prime members of the nine genera support the division of picornaviruses into the proposed genera. The study also supports a previous suggestion based on analysis of partial sequences of the structural proteins that LV is more related to the genus of Parechovirus than to other picornaviruses, but also shows that the three LV strains used in the comparison constitute a distinct monophyletic group, clearly separated from the parechoviruses. The analyses using the 2C and 3D sequences clearly showed that A-2 was related to the genera of Rhinovirus and Enterovirus, but it was not possible to group the A-2 with high confidence into one of the genera. Comparison using the VP1 protein sequences of Enterovirus and Rhinovirus showed that although the A-2 virus is positioned between the two genera, the virus is more related to the genus of Enterovirus than to Rhinovirus. Our analysis of the three LV strains based on the phylogenetic analysis of the 2C and 3D proteins suggests that the strains used in this study constitute a monophyletic group clearly related to Parechovirus of Picornaviridae. The taxonomic position of the A-2 virus is presently uncertain but available data indicate that this virus may be classified as a member of the genus of Enterovirus.

Animals↗

MilkProtChip--a microarray of SNPs in candidate genes associated with milk protein biosynthesis--development and validation.

MilkProtChip is an oligonucleotide microarray based on the arrayed primer extension (APEX) technique, allowing genotyping of single nucleotide polymorphisms (SNPs) in genes of interest for bovine milk protein biosynthesis. APEX consists of a sequencing reaction primed by an oligonucleotide anchored with its 5'end to a glass slide and terminating one nucleotide before the polymorphic site. The extension with one fluorescently labeled dideoxy nucleotide complementary to the template reveals the polymorphism. A total of 75 SNPs were selected among those associated directly or potentially with milk protein content. Among the 75 SNPs, 4 did not produce a positive signal. Most of the remaining SNPs produced a signal for both strands, except for 4 (one strand). In the validation step, 12 Polish Holstein bulls, 1 Polish Red bull, 1 bison (Bison bonasus), 11 Jersey cows and 25 Polish Holstein cows were screened to validate SNPs. Among the 71 selected SNPs--26 were found monoallelic, the rest showing at least two genotypes for the entire population under study. All the animals were earlier genotyped for 2-5 SNPs by PCR-RFLP and PCR sequencing and all showed complete concordance with APEX genotyping. APEX reactions showed relatively high signal frequencies: more than 0.9, 0.9-0.8 and below 0.8, for 65, 4 and 2 DNA samples, respectively. The primary application of the MilkProtChip is the simultaneous genotyping of dozens of SNPs to reveal and clarify the genetic background of milk protein biosynthesis. The chip may possibly be used for dairy cattle identification and paternity analysis, evolutionary studies, the evaluation of genetic distances between wild and domestic cattle breeds and the domestication history of bovine species.

Animals↗

Primate evolution of an olfactory receptor cluster: diversification by gene conversion and recent emergence of pseudogenes.

The olfactory receptor (OR) subgenome harbors the largest known gene family in mammals, disposed in clusters on numerous chromosomes. We have carried out a comparative evolutionary analysis of the best characterized genomic OR gene cluster, on human chromosome 17p13. Fifteen orthologs from chimpanzee (localized to chromosome 19p15), as well as key OR counterparts from other primates, have been identified and sequenced. Comparison among orthologs and paralogs revealed a multiplicity of gene conversion events, which occurred exclusively within OR subfamilies. These appear to lead to segment shuffling in the odorant binding site, an evolutionary process reminiscent of somatic combinatorial diversification in the immune system. We also demonstrate that the functional mammalian OR repertoire has undergone a rapid decline in the past 10 million years: while for the common ancestor of all great apes an intact OR cluster is inferred, in present-day humans and great apes the cluster includes nearly 40% pseudogenes.

Animals↗