PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Purifying selection”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Dynamic and non-additive gene regulation shapes maize responses to simultaneous salt and cold stress.

Salt and cold stresses often occur together in nature and severely impact crop productivity, yet their transcriptional regulation remains poorly understood. Here, we conducted a time-series transcriptomic analysis of maize under salt, cold, and their combination at 0, 6, 12, and 24 h. Differential expression analysis revealed dynamic, condition-specific gene responses grouped into eight distinct temporal patterns. Promoter motif analysis of genes within each pattern identified 5-39 significantly enriched motifs, with over 40% lacking known counterparts, suggesting the involvement of previously uncharacterized cis-regulatory elements in stress-responsive transcriptional regulation. By comparing combined stress responses to the sum of single-stress effects, we found that about 74% of DEGs showed non-additive patterns, suggesting that combined stress triggers a distinct transcriptional program. Evolutionary analysis showed that additive DEGs tend to be more recently evolved, subject to weaker purifying selection, and enriched in transposed duplications, contrasting with the stronger constraint observed in non-additive DEGs. WGCNA identified 24 co-expression modules, among which 65 hub DEGs were detected in modules significantly correlated with specific stress conditions. Furthermore, we reconstructed 228, 20, and 200 sequential transcription factor cascades spanning 6 h, 12 h, and 24 h under cold, salt, and combined stress, respectively, with no cascade shared across all three conditions. Together, these results reveal that maize responses to combined salt and cold stress are largely non-additive and temporally dynamic, with distinct evolutionary patterns underlying different response types, offering insights and candidate regulators for enhancing crop stress resilience.

Zea mays↗

Assembly and Characterization of the First Complete Mitochondrial Genome of Tussilago farfara L.: Insights into Biological Functions and Phylogenetic Relationships within the Asteraceae Family.

Tussilago farfara L., a member of the Asteraceae family, is an economically valuable species due to its edible and medicinal properties. To elucidate the structural characteristics, genetic mechanisms, and evolutionary pathways of the organelle genomes of T. farfara, we sequenced, assembled, and annotated its mitochondrial genome for the first time. The complete mitochondrial genome of T. farfara spans 306,024 bp and contains 33 mitochondrial protein-coding genes (PCGs), 3 rRNAs, and 22 tRNAs. Analysis of the nucleotide substitution rate and genetic diversity revealed that most mitochondrial genome genes may have undergone purifying selection, indicating a slow evolutionary rate and a relatively conserved genomic structure. We further identified 13 fragments of chloroplast-derived DNA integrated into the mitochondrial genome, evidencing intracellular gene transfer. Collinearity analysis showed that Arctium lappa shares the most extensive mitochondrial homologous sequences and the highest sequence similarity with T. farfara. Phylogenetic analysis based on the mitochondrial genome helped to clarify the evolutionary and taxonomic position of T. farfara within the Asteraceae family. The mitochondrial genome sequence of T. farfara provides a valuable genomic resource for species identification and for evolutionary studies within the Asteraceae family.

Genome, Mitochondrial↗

Sexually dimorphic expression and hormonal responsiveness of steroidogenic Cyp genes during gonadal differentiation in mandarin fish.

Steroid hormones play a pivotal role in fish sex differentiation, yet the dynamic expression patterns of key steroidogenic enzymes during this process remain incompletely characterized. Here, we combined genome-wide identification, time series transcriptomes spanning gonadal development (5-360 days post-hatch), and multiple hormone treatment experiments (17α-methyltestosterone, estrone, and etonogestrel) to investigate the Cyp11, Cyp17, Cyp19, and Cyp21 subfamilies in mandarin fish (Siniperca chuatsi). Seven steroidogenic Cyp genes were identified, showing teleost-specific expansion, with one duplicated pair (cyp17a2 and cyp2u1) exhibiting strong purifying selection. Expression profiling revealed pronounced sexually dimorphic and stage-specific patterns: During female differentiation (20-30 days), cyp19a1a and associated genes were highly expressed, coinciding with ovarian differentiation; during male differentiation (30-60 days), cyp17a2 and related genes were upregulated, aligning with testicular development. Exogenous hormone treatments further demonstrated that these genes are dynamically responsive: cyp19a1a and cyp17a2 were highly responsive to androgenic and progestogenic treatments, and their expression changes correlated closely with gonadal sex reversal phenotypes observed histologically. Collectively, this study provides a comprehensive expression atlas of steroidogenic Cyp genes during gonadal differentiation and identifies key hormonally responsive candidates for sex control in aquaculture.

Animals↗

Mitogenomic and phylogenomic analyses identify a cohesive Western Atlantic lineage within the Narcine complex (Torpediniformes: Narcinidae).

BACKGROUND: Accurate species delimitation within electric rays of the genus Narcine has been hindered by overlapping morphological characters and limited molecular resolution in previous single-locus studies. This study aims to evaluate phylogenetic relationships and species boundaries within the Narcine species complex across the Western Atlantic using complete mitochondrial genomes. METHODS AND RESULTS: Seven complete mitogenomes were newly assembled from individuals representing distinct morphotypes sampled across geographically widespread Western Atlantic localities and analyzed together with publicly available reference sequences. Mitochondrial protein-coding genes (PCGs) were examined using concatenated nucleotide and amino acid datasets under partitioned maximum-likelihood frameworks. Both approaches recovered highly congruent topologies, consistently supporting a single, well-defined western Atlantic mitochondrial lineage with low internal divergence (0.04-2.13%). Species delimitation analyses based on multiple methods yielded partially congruent results but consistently identified a dominant lineage encompassing all Atlantic samples. In contrast, two Colombian reference mitogenomes formed a separate and highly divergent lineage relative to the Atlantic group, despite showing moderate divergence between them. Comparative mitogenomic analyses revealed conserved genome organization, nucleotide composition bias, codon usage, and transfer RNA (tRNA) structures. All PCGs evolved under strong purifying selection, with Ka/Ks ratios well below unity. CONCLUSIONS: These results support mitochondrial genetic continuity across the Western Atlantic Narcine populations and do not provide mitochondrial evidence for multiple evolutionary lineages within the Western Atlantic. The marked mitochondrial divergence of Colombian reference mitogenomes highlights potential issues in sequence attribution and underscores the importance of data curation. Overall, complete mitochondrial genomes provide a robust framework for species delimitation and future integrative taxonomic assessments within Narcine.

Animals↗

Comparative sequence analysis of the phytochrome C gene and its upstream region in allohexaploid wheat reveals new data on the evolution of its three constituent genomes.

Bread wheat is an allohexaploid with genome composition AABBDD. Phytochrome C is a gene involved in photomorphogenesis that has been used extensively for phylogenetic analyses. In wheat, the PhyC genes are single copy in each of the three homoeologous genomes and map to orthologous positions on the long arms of the group 5 chromosomes. Comparative sequence analysis of the three homoeologous copies of the wheat PhyC gene and of some 5 kb of upstream region has demonstrated a high level of conservation of PhyC, but frequent interruption of the upstream regions by the insertion of retroelements and other repeats. One of the repeats in the region under investigation appeared to have inserted before the divergence of the diploid wheat genomes, but was degraded to the extent that similarity between the A and D copies could only be observed at the amino acid level. Evidence was found for the differential presence of a foldback element and a miniature inverted-repeat transposable element (MITE) 5' to PhyC in different wheat cultivars. The latter may represent the first example of an active MITE family in the wheat genome. Several conserved non-coding sequences were also identified that may represent functional regulatory elements. The level of sequence divergence (Ks) between the three wheat PhyC homoeologs suggests that the divergence of the diploid wheat ancestors occurred some 6.9 Mya, which is considerably earlier than the previously estimated 2.5-4.5 Mya. Ka/Ks ratios were <0.15 indicating that all three homoeologs are under purifying selection and presumably represent functional PhyC genes. RT-PCR confirmed expression of the A, B and D copies. The discrepancy in evolutionary age of the wheat genomes estimated using sequences from different parts of the genome may reflect a mosaic origin of some of the Triticeae genomes.

5' Flanking Region↗

Pan-genome characterization of the maize 4CL gene family and its dynamic responses to abiotic stress.

1.Pan-genome analysis across 26 maize inbred lines identified 13&#xa0;Zm4CL&#xa0;genes (nine core and four near-core) classified into three evolutionary clades.2.Structural variations (SVs) are significantly associated with the expression and altered conserved protein domains of key&#xa0;Zm4CL&#xa0;genes.3.Zm4CL&#xa0;genes exhibit distinct tissue-specific expression patterns and dynamic enzymatic and transcriptional responses to stresses, particularly cold and drought.4-Coumarate:CoA ligase (4CL) is a key enzyme in the phenylpropanoid pathway and plays important roles in plant growth, development, and responses to environmental stresses. However, a comprehensive pan-genome analysis of the 4CL gene family in maize is still lacking. In this study, 13 Zm4CL genes were identified from a maize pan-genome comprising 26 diverse inbred lines, including nine core genes and four near-core genes. Phylogenetic analysis classified these genes into three evolutionary clades, while Ka/Ks analysis indicated that most members have been maintained under purifying selection, although several genes exhibited greater evolutionary divergence and relatively relaxed evolutionary constraints. Structural variation (SV) analysis revealed significant associations between SVs and the expression of Zm4CL2 and Zm4CL3, while sequence comparisons suggested that SVs were also associated with alterations in conserved protein domains in some genotypes. Transcriptome analyses revealed distinct tissue-specific expression patterns and diverse transcriptional responses to abiotic and biotic stresses. Enzyme activity assays showed that cold stress significantly increased 4CL activity at 12&#xa0;h, whereas heat, salt, and alkali stresses caused an initial decrease followed by recovery, while drought had no significant effect. Time-course RT-qPCR further validated dynamic expression changes of representative Zm4CL genes under cold and drought stresses. Overall, this study provides a comprehensive pan-genome framework for understanding the evolutionary conservation, regulatory diversification, and stress-responsive characteristics of the maize Zm4CL gene family, providing valuable resources for future functional studies and the genetic improvement of stress tolerance in maize.

Zea mays↗

Loss of linkage disequilibrium and accelerated protein divergence in duplicated cytomegalovirus chemokine genes.

Human CMV (hCMV) encodes several captured chemokine ligand and chemokine receptor genes that may play a role in immune evasion. The adjacent viral alpha-chemokine genes UL146 and UL147 appear to have duplicated subsequent to a recent gene capture event. Sequence data from multiple hCMV isolates suggest accelerated protein evolution in one of the paralogues, UL146. Extensive sequence variation was noted throughout the more rapidly evolving paralogue, although significant variation was also observed within the more slowly evolving gene, especially within a region corresponding to a possible signal peptide. In contrast to the haplotype structure observed for other hCMV genes, the distribution of nucleotide variants indicates a marked loss of linkage disequilibrium within UL146 and to a lesser extent UL147. Despite evidence of accelerated protein evolution, the rate of nonsynonymous to synonymous substitutions (d(N)/d(S)) in the more rapidly evolving paralogue was not indicative of neutral evolution, but of moderate purifying selection. The data presented here provides a unique opportunity to study the mechanisms by which a recently duplicated pair of genes has diverged and suggests a role for recombination.

Animals↗

Genome-wide characterization of the tomato PERK gene family and its expression profiling under abiotic stresses.

UNLABELLED: This study presents the first systematic genome-wide characterization of the proline-rich extensin-like receptor kinases (PERK) gene family in tomato (Solanum lycopersicum) and their transcriptional responses under abiotic stresses. Using the latest SL4.0/ITAG4.0 genome assembly, we identified six SlPERK genes, all harboring the conserved Ser/Thr protein kinase domain. Evolutionary and structural analyses revealed strong purifying selection (Ka/Ks&#x2009;<&#x2009;1), distinct exon-intron organizations, and the presence of stress- and hormone-responsive cis-regulatory elements in their promoters. Furthermore, post-transcriptional regulation by 57 miRNAs and complex protein-protein interaction networks were predicted. To validate their stress-responsive roles, two tomato cultivars (GMOTL-1 and Roma) were subjected to cold, heat, and salinity treatments. Quantitative RT-PCR analysis revealed cultivar-specific expression dynamics: SlPERK4 exhibited strong transient induction under cold and heat stress, while SlPERK6 was highly responsive to salinity. Notably, the GMOTL-1 cultivar displayed significantly higher and broader stress-responsive expression profiles compared to Roma, indicating a potential role of these SlPERK genes in cultivar-specific stress tolerance. These findings provide a comprehensive genomic resource and establish a critical foundation for the functional validation and molecular breeding for stress-resilience tomato cultivars. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at https://doi.org/10.1007/s13205-026-05044-y.

Abiotic stress↗

Mining the sHSP20 (small heat-shock protein) gene family in finger millet (Eleusine coracana (L.) Gaertn.): structural, evolutionary and predicted abiotic-stress-responsive insights.

Small heat-shock proteins (sHSPs, the HSP20 family) are ATP-independent molecular chaperones that hold partially unfolded substrates and protect the proteome during heat and other abiotic stresses; every member is defined by a conserved &#x3b1;-crystallin domain (ACD). Finger millet (Eleusine coracana) is a climate-resilient, calcium-rich allotetraploid cereal of the semi-arid tropics whose HSP20 repertoire had not been catalogued. The present study is an entirely computational (in silico) analysis of the chromosome-scale reference genome of finger millet (NCBI GenBank assembly GCA_032690845.1, cultivar KNE 796-S). Mining the predicted proteome with the ACD profile (Pfam PF00011) and confirming every candidate by NCBI CD-search recovered 76 non-redundant ACD-bearing HSP20 genes (EcHSP20-1-EcHSP20-76). Based on phylogeny and TargetP-predicted localization, the members were classified into ten subfamilies: seven cytosolic/nuclear classes (C-I to C-VII, 60 members) together with chloroplastic (11), mitochondrial (3) and endoplasmic-reticulum (2) groups. The proteins ranged from 110 to 355 amino acids (12.1-39.2&#xa0;kDa) with theoretical pI of 4.85-9.69. The 76 loci were distributed over 14 of the 18 chromosomes and were conspicuously absent from chromosomes 8&#xa0;A, 8B, 9&#xa0;A and 9B, with pronounced clustering on chromosomes 1, 2, 3 and 6. Duplication analysis detected 149 paralogous pairs (49 homoeologous, 80 segmental/dispersed and 18 tandem); 147 of 148 pairs for which substitution rates could be calculated returned Ka/Ks&#x2009;<&#x2009;1 (mean 0.20), indicating strong purifying selection consistent with retention after whole-genome/allopolyploid duplication. Promoter analysis (PlantCARE) revealed enrichment of abscisic-acid-responsive (ABRE), MYB/MYC drought-related, STRE, DRE, low-temperature (LTR) and methyl-jasmonate/salicylic-acid elements, whereas canonical heat-shock elements (HSE) were not recovered. Expression profiling against a public drought transcriptome (SRP081350) showed that about half of the genes (39 of 76) are transcribed in leaf tissue, the expressed fraction being dominated by the cytosolic class C-I. This first finger-millet HSP20 catalogue provides a verified, reproducible framework and nominates computationally predicted candidate genes for future functional work on thermotolerance in cereals.

Allotetraploid↗

Identification of peptides containing tryptophan, tyrosine, and phenylalanine using photodiode-array spectrophotometry.

The characteristic absorption spectra of aromatic amino acids between 240 and 310 nm were used to identify tryptophan, tyrosine, and phenylalanine-containing peptides. In acidic solution, the absorption spectra of these amino acids exhibit minima or maxima at 255, 270, and 286 nm. Based on these characteristics, the content of the aromatic amino acid in peptide can be estimated. For this study, 2 nmol of tryptic peptides from human apolipoprotein A-1 was separated by high-performance liquid chromatography using a reverse-phase column. The peptide fragments were monitored by a photodiode-array spectrophotometer. This new approach offers a rapid, simple, sensitive, and direct identification of peptides containing aromatic amino acids. Those containing Trp, which may be of interest for DNA sequencing and important in sequence analysis of proteins, can be selectively purified using this technique.

Apolipoprotein A-I↗

Physiological correlation between glycyrrhizin, glycyrrhizin-binding lipoxygenase and casein kinase II.

By means of glycyrrhizin (GL)-affinity column chromatography, a GL-binding lipoxygenase (gbLOX) was selectively purified from the partially purified soybean LOX-1 fraction. Polypeptide analysis of the purified gbLOX by SDS-PAGE detected two distinct polypeptides (p96 and p94), which were identical to LOX-3 as determined by their partial N-terminal amino acid sequences. Moreover, it was found that (i) phosphorylation of gpLOX by casein kinase II (CK-II) is significantly stimulated by 3 microM GL, but inhibited by 30 microM GL or 10 microM oGA; and (ii) gbLOX activity is enhanced when the enzyme is phosphorylated by CK-II in the presence of 3 microM GL. These results suggest that (i) CK-II is a kinase responsible for the activation of gbLOX through its specific phosphorylation; and (ii) GL is one of the regulatory substances for specific phosphorylation of gbLOX (LOX-3) by CK-II in plant cells.

Amino Acid Sequence↗

Mobile phase effects on membrane protein elution during immobilized artificial membrane chromatography.

The eluotropic strength of different mobile phases for eluting membrane proteins from immobilized artificial membrane (IAM) chromatography surfaces was studied. Two protein mixtures containing bovine pancreatic PLA2 were used in this study. Protein mixture I was PLA2 obtained from Sigma which contained approximately 5-10 major protein bands in electrophoretic gels. Protein mixture II was obtained from flesh bovine pancreatic tissue and contained > 100 proteins including the target protein, PLA2. After adsorbing Sigma PLA2 to IAM columns, the elution conditions common to conventional chromatographic methods were evaluated for their ability to selectively purify PLA2. Elution conditions tested were (i) detergent gradients, (ii) salt gradients used during ion-exchange chromatography, (iii) salt conditions used during hydrophobic interaction chromatography, (iv) acetonitrile gradients used during reversed-phase chromatography, and (v) a two-step gradient consisting of first a detergent gradient followed by an acetonitrile gradient. Based on silver-stained electrophoretic protein gels. PLA2 from protein mixture I was purified to electrophoretic homogeneity with 417-fold increase in specific activity in one step using elution condition (v), and PLA2 from protein mixture II was purified in one step (660-fold increase in specific activity) using elution condition (iv). Total protein recovery from IAM columns is 70-100%.

Animals↗

Isolation and nucleotide sequence analysis of the beta-type globin pseudogene from human, gorilla and chimpanzee.

The beta-globin gene cluster of human, gorilla and chimpanzee contain the same number and organization of beta-type globin genes: 5'-epsilon (embryonic)-G gamma and A gamma (fetal)-psi beta (inactive)-delta and beta (adult)-3'. We have isolated the psi beta-globin gene regions from the three species and determined their nucleotide sequences. These three pseudogenes each share the same substitutions in the initiator codon (ATG----GTA), a substitution in codon 15 which generates a termination signal TGG----TGA, nucleotide deletion in codon 20 and the resulting frame shift which yields many termination signals in exons 2 and 3. The basic structure of these psi beta-globin genes, however, remains consistent with that found for functional beta-globin genes: their coding regions are split by two introns, IVS 1 (which splits codon 30, 121 base-pairs in length) and IVS 2 (which splits codon 104, 840 to 844 base-pairs in length). These introns retain the normal splice junctions found in other eukaryotic split genes. The three hominoid psi beta-globin genes show a high degree of sequence correspondence, with the number of differences found among them being only about one-third of that predicted for DNA sites evolving at the neutral rate (i.e. for sites evolving in the absence of purifying selection). Thus, there appears to be a deceleration in the rate of evolution of the psi beta-globin locus in higher primates.

Animals↗

Sequence specificity of 125I-labelled Hoechst 33258 in intact human cells.

Using polyacrylamide/urea DNA sequencing gels, the DNA sequence selectivity of 125I-labelled Hoechst 33258 damage has been determined in intact human cells to the exact base-pair. This was accomplished using a novel procedure with human alpha RI-DNA as the target DNA sequence. In this procedure, after size fractionation, the alpha RI-DNA is selectively purified by hybridization to a single-stranded M13 clone containing an alpha RI-DNA insert. The sequence specificity of [125I]Hoechst 33258 was indistinguishable in intact cells from purified high molecular weight DNA; and this is surprising considering the more complex environment of DNA in the nucleus where DNA is bound to nucleosomes and other DNA binding proteins. The ligand preferentially binds to DNA sequences which have four or more consecutive A.T base-pairs. The extent of damage was measured with a densitometer and, relative to the damage hotspot at base-pair 94, the extent of damage was similar in both purified high molecular weight DNA and intact cells. [125I]Hoechst 33258 causes only double-strand breaks, since single-strand breaks or base damage were not detected. These experiments represent the first occasion that the sequence specificity of a DNA damaging agent, which causes only double-strand breaks, has been determined to the exact base-pair in intact cells.

Autoradiography↗

Detection of human exposure to carcinogens by measurement of alkyl-DNA adducts using immunoaffinity clean-up in combination with gas chromatography-mass spectrometry and other methods of quantitation.

A brief overview is given of recent developments from our laboratory in the use of immunoaffinity clean-up in the determination of alkyl-DNA adducts. Compound- and group-specific antibodies have been prepared against 7-alkylguanines and 3-alkyladenines. The antibodies were attached to solid supports to make immunoaffinity columns which could then be used to selectively purify either single adducts or groups of adducts prior to quantitation by various methods. In the case of methyl adducts quantitation was achieved by ELISA (3-methyl-adenine, using a monoclonal antibody) and HPLC-electrochemical detection (7-methylguanine). For groups of adducts, quantitation of the individual compounds was effected by gas chromatography-mass spectrometry (3-alkyladenines, using deuterated analogues of each adduct as an internal standard) and HPLC-fluorescence detection (7-alkylguanines). In all of these cases efficient purification of adducts from urine or DNA hydrolysates could be easily carried out. Using these techniques human exposure to alkylating agents in tobacco smoke and from cancer chemotherapy has been studied.

Alkylation↗

Immunochemical analysis of a recombinant, genetically engineered, secreted HLA-A2/Q10b fusion protein.

We engineered a fusion gene which encodes the alpha 1 and alpha 2 domains of HLA-A2 with the alpha 3 and truncated transmembrane domains of the murine class I-like protein Q10b, and transferred it into mouse L cells along with the gene for human beta 2-microglobulin (beta 2m). The secreted rA2/Q10b gene product consisted of a single heavy chain of molecular weight 42 kd that was noncovalently associated with the human beta 2m light chain. Native detergent-solubilized HLA-A2 and secreted rA2/Q10b proteins were found to be similar by: (a) the binding to mouse monoclonal anti-HLA antibodies in an ELISA; (b) the blocking of lysis of HLA-A2+ cells by human anti-HLA-A2,-B17, anti-HLA-A2,9,28, and anti-HLA-A2,28 cross-reactive group (CREG) antisera in a complement-dependent cytotoxicity assay; and (c) the ability when coupled to Sepharose to selectively purify HLA-A2,9,28 and HLA-A2,28 CREG-specific antibodies. Mouse L cells expressing rA2/Q10b produced as much as 2.5 micrograms protein per 10(6) cells/day, or 50- to 100-fold more antigen on a per cell basis than the level of HLA-A2 expressed by B-lymphoblastoid cell line or spleen cells. Thus rA2/Q10b represents a viable alternative to detergent-solubilized HLA-A2 for purification of anti-HLA-A2 antibodies and analysis of anti-HLA-A2 immune responses.

Animals↗

Duplicated cytoglobin genes in teleost fishes.

Cytoglobin is a recently discovered myoglobin-related O2-binding protein of vertebrates with uncertain function. It occurs as single-copy gene in mammals. Here, we demonstrate the presence of two paralogous cytoglobin genes (Cygb-1 and Cygb-2) in the teleost fishes Danio rerio, Oryzias latipes, Tetraodon nigroviridis, and Takifugu rubripes. The globin-typical introns at positions B12.2 and G7.0 are conserved in both genes, whereas the C-terminal exon found in mammalian cytoglobin is absent in the fish genes. Phylogenetic analyses show that the two cytoglobin genes diverged early in teleost evolution. This is confirmed by gene synteny analyses, which suggest a large-scale duplication event. Although both cytoglobin genes are highly conserved and have evolved under purifying selection, substitution rates are significantly higher in Cygb-1 than in Cygb-2. Similar to their mammalian ortholog, both fish cytoglobins are expressed in a broad range of tissues. However, Cygb-2 is more than 250-fold stronger expressed in neuronal tissues, suggesting a subfunctionalization of the two cytoglobin paralogs after gene duplication.

Amino Acid Sequence↗

Comparative genomic and proteomic analysis reveals orthogroup structured evolution of tick protease inhibitors.

Protease inhibitors (PIs) play central roles in regulating endogenous proteolysis and host-parasite interactions in ticks. However, the evolutionary architecture underlying their diversification across tick lineages remains insufficiently resolved. Here, we performed a genome-wide comparative analysis of predicted proteomes from 14 tick species to systematically characterize PI repertoires. In total, 4931 putative PIs were identified and grouped into 20 families using the MEROPS classification system. Further, PI families such as Antistasin, WAP-type, and Pacifastin, which have not previously been systematically reported in tick genomes, were classified. Orthogroup inference demonstrated that PI expansion is structured at the level of evolutionary lineages rather than uniformly across families. By stratifying orthogroups according to duplication burden and taxonomic conservation, we identified a broadly conserved single-copy core under strong purifying selection. Motif level analysis of serpin reactive center loops further revealed conservation of inhibitory specificity within single copy orthogroups and diversification of key functional residues in duplication-associated lineages. Integration of secretion prediction and tissue-resolved proteomics from Hyalomma anatolicum and Rhipicephalus microplus demonstrated that evolutionary stratification is reflected at the protein level. Together, these findings provide an orthogroup-resolved evolutionary framework linking duplication dynamics, molecular evolution, and tissue-level protein deployment. This integrative approach offers a systematic basis for prioritizing conserved and diversified PI lineages for future functional and anti-tick intervention studies.

Animals↗