PubMed HealthSearch

SEARCH · PubMed Health

Results for “Conserved Sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

FUSE-PhyloTree: linking functions and sequence conservation modules of a protein family through phylogenomic analysis.

SUMMARY: FUSE-PhyloTree is a phylogenomic analysis software for identifying local sequence conservation associated with the different functions of a multi-functional (e.g. paralogous or multi-domain) protein family. FUSE-PhyloTree introduces an original approach that combines advanced sequence analysis with phylogenetic methods. First, local sequence conservation modules within the family are identified using partial local multiple sequence alignment. Next, the evolution of the detected modules and known protein functions is inferred within the family's phylogenetic tree using three-level phylogenetic reconciliation and ancestral state reconstruction. As a result, FUSE-PhyloTree provides a gene tree annotated with both predicted sequence modules and ancestral gene functions, enabling the association of functions with specific sequence regions based on their co-emergence. AVAILABILITY AND IMPLEMENTATION: FUSE-PhyloTree is provided as Docker and Singularity images including all the required software tools. Images, source code, test data, and documentation are available at https://github.com/OcMalde/fuse-phylotree and https://zenodo.org/records/15855068.

Phylogeny

Whole-Genome Conservation Analysis for the Specific and Accurate Detection of Influenza A and B Viruses and Respiratory Syncytial Virus by Quadruplex RT-qPCR.

Influenza virus (Flu) and respiratory syncytial virus (RSV) are the primary pathogens responsible for acute respiratory infections. Both viruses are prone to mutations due to the seasonal epidemic, leading to an increasing rate of false-negative results. In this study, comprehensive meta-analyses of the genomes focusing on most conserved fragments have been performed for the four seasonal influenza viruses (two subtypes of Flu A: H1N1 and H3N2; two subtypes of Flu B: Yamagata and Victoria) and the two types of RSV: RSVA and RSVB), respectively. The most conserved sequences of 200 bp were identified as targets of the designed primer/probe sets for RT-qPCR were screened and optimized. Good sensitivities of the optimized primer/probe sets were obtained with the limits of detections of 2.95, 2.82, 1.57, 2.8, 1.19, and 2.12 copies/reaction for H1N1, H3N2, Yamagata, Victoria, RSVA and RSVB, respectively. Eventually, quadruplex qPCR using the four designed primer/probe sets can achieve simultaneous screening of the four viruses at a single tube. Furthermore, the assay's good performance in detecting target viruses from clinical throat swab samples demonstrated its potential for diagnosis of these viruses. The method, based on the identified conserved sequences and primer/probe sets, can effectively reduce false-negative results and rapidly respond to these viruses during respiratory disease outbreaks, or even before their widespread emergence, which aid in preventing outbreaks and guiding clinical treatment.

Humans

Exploring the Effect of Whole-Genome Duplication on Salmonid LincRNA Repertoire.

Long intergenic non-coding RNAs (lincRNAs) are key epigenetic regulators of genome function, yet their evolutionary dynamics following whole-genome duplication (WGD) events remain poorly understood. Salmonids, which underwent a lineage-specific autotetraploidization (salmonid-specific WGD, ~88-100 million years ago), provide an excellent model to investigate the retention, divergence, and functional potential of recently duplicated non-coding elements. LincRNA repertoires were compared across five genome-annotated salmonids (Oncorhynchus tshawytscha, O. kisutch, O. mykiss, Salmo salar, and S. trutta) and their closest non-duplicated relative, northern pike (Esox lucius). LincRNAs represented ~5-7% of annotated genes in all salmonids except S. salar (18%). Sequence conservation was low relative to coding genes, with only 11-68 highly similar (e-value < 1 &#xd7; 10-30; similarity > 70% and alignments > 100 nucleotides) putative orthologues shared between salmonids and northern pike, and 161-338 among salmonids alone. Synteny conservation was modest in lincRNAs, with lower conservation in putative orthologues (8-16%) compared to putative ohnologues (8-33%). Secondary structure conservation was associated with sequence similarity (&#x3c1; = -0.45; p = 2.2 &#xd7; 10-16), and the association was stronger among WGD ohnologues than orthologues. In S. salar and O. mykiss, lincRNA putative ohnologues showed weaker expression correlations than coding genes, suggesting widespread regulatory divergence, possibly through neo- and subfunctionalisation. Conserved salmonid lincRNAs showed enriched predicted interactions with miRNAs involved in tumour suppression, brain, bone, and muscle development (e.g., miR-455, miR-365, miR124, miR-133a, miR-140, and miR-9), a finding supported by limited transcriptomic data. Although salmonid WGD expanded lincRNA repertoires, lincRNAs have undergone rapid sequence and transcriptional divergence, with limited conservation across species based on sequence similarity, chromosomal position, synteny, and secondary structure. A subset of conserved lincRNAs retains structural features and regulatory signatures consistent with roles as miRNA sponges in brain, skeletal, and muscle development and tumour suppression, potentially acting within conserved regulatory networks. These findings provide new insights into lincRNA evolution following genome duplication and highlight the need for experimental validation of their regulatory functions.

Animals

Screening, optimization and artificial recombination of dsRNA fragments for RNAi-mediated pest resistance in Apolygus lucorum.

RNA interference (RNAi) is an eco-friendly strategy for pest management, with double-stranded RNA (dsRNA) as the core functional component. In this study, three RNAi target genes (Ubx, wupA and Dpp) with strong lethal effects on Apolygus lucorum were screened via microinjection. The 7-day cumulative mortalities were 56.67 &#xb1; 3.33% for dsUbx, 94.44 &#xb1; 1.11% for dswupA and 92.22 &#xb1; 1.11% for dsDpp. We optimized dsRNA sequences by removing conserved sequences in non-target organisms based on homology alignment and off-target risk analysis. The optimized fragments dswupA-OTE and dsDpp-OTE still exhibited high insecticidal activity, with 7-day cumulative mortalities of 77.78 &#xb1; 2.94% and 70.00 &#xb1; 1.93%, respectively. We also evaluated the effects of dsRNA length and target sites on RNAi efficiency and screened potent short dsRNA fragments. Novel artificially recombinant dsRNAs were constructed by assembling effective short fragments from different genes, which retained strong insecticidal activity despite shorter sequence length. This study verifies the feasibility of multi-target recombinant dsRNA for pest control and provides a theoretical basis for developing multi-gene RNAi technologies against A. lucorum.

Apolygus lucorum

Identification of lineage-associated polymorphisms in the ROP18 3' flanking region and development of molecular assays for differentiation of Toxoplasma gondii lineages.

BACKGROUND: Toxoplasma gondii (T. gondii) exhibits substantial genetic diversity, and different parasite lineages are associated with distinct epidemiological distributions and biological characteristics. Accurate molecular characterization of T. gondii strains is important for understanding parasite population structure and transmission patterns. However, existing genotyping approaches often require multiple loci, extensive experimental procedures, or complex data analysis. Therefore, simplified and reliable molecular markers for rapid lineage differentiation are still needed. METHODS: In this study, comparative genomic analysis was performed using representative T. gondii strains with well-defined genetic backgrounds and virulence phenotypes. The ROP18 genomic region, including partial genomic sequences, 5' flanking regions, coding sequence (CDS), and 3' flanking regions, was analyzed to identify informative polymorphic signatures. A short conserved sequence region containing lineage-associated polymorphic sites was identified within the ROP18 3' flanking region. Based on these sequence signatures, HRM-PCR and TaqMan MGB probe-based real-time PCR assays were developed and evaluated using plasmid standards and representative T. gondii genomic DNA samples. RESULTS: Phylogenetic analyses based on different ROP18 genomic regions demonstrated distinct clustering patterns among analyzed strains. Although the ROP18 3' flanking region was highly conserved, a short conserved sequence region containing informative polymorphic sites was identified, and the combination of these sites generated three distinct lineage-associated ROP18 patterns. Analysis of publicly available genomic datasets further demonstrated that individual strains contained one of these defined patterns rather than multiple patterns simultaneously. The developed HRM-PCR assay successfully discriminated the three ROP18-associated patterns based on distinct melting profiles with good reproducibility. Furthermore, the TaqMan MGB probe-based assay enabled specific identification of different ROP18-associated patterns through defined probe-recognition combinations and showed good analytical performance. CONCLUSION: This study identifies novel lineage-associated molecular signatures within the ROP18 3' flanking region and establishes complementary HRM-PCR and TaqMan MGB probe-based approaches for rapid molecular differentiation of T. gondii strains. These findings highlight the potential of conserved non-coding regions adjacent to functionally important genes as informative targets for parasite genotyping and provide a practical complementary tool for epidemiological surveillance and strain characterization.

HRM-PCR

Strong phylogenetic signal from chloroplast genomes of three Barringtonia species provides the first genomic resources for their conservation.

BACKGROUND: The genus Barringtonia (Lecythidaceae) is a vital component of tropical coastal forests and mangrove ecosystems. Among its members, B. racemosa and B. fusicarpa are classified as Endangered and Vulnerable, respectively, due to habitat degradation and anthropogenic pressures, underscoring the urgent need for genetic studies to guide conservation. Chloroplast (cp.) genomes serve as essential resources for phylogenetic reconstruction and conservation genetics. However, the scarcity of cp. genome data for Barringtonia has limited comprehensive evolutionary and conservation-oriented investigations. RESULTS: We assembled and annotated the first complete cp. genomes of B. racemosa, B. fusicarpa, and B. acutangula. All three genomes exhibit the typical quadripartite structure, ranging from 158,959&#xa0;bp (B. racemosa) to 159,837&#xa0;bp (B. acutangula), and contain 132 genes (87 protein-coding, 37 tRNA, 8 rRNA) with a GC content of 36.68%-36.86%. Collinearity and IR boundary analyses revealed high structural conservation without large-scale rearrangements. Interspecific sequence-level variations were detected in simple sequence repeats (SSRs) and long repeats. Nucleotide diversity (&#x3c0;) analysis identified highly polymorphic regions, including rpl20 (&#x3c0;&#x2009;=&#x2009;0.080), rpoA (&#x3c0;&#x2009;=&#x2009;0.064), rps3 (&#x3c0;&#x2009;=&#x2009;0.063), and ndhF (&#x3c0;&#x2009;=&#x2009;0.060), which represent promising molecular markers for population genetics within the genus. Codon-based selection analyses (Ka/Ks) showed that all protein-coding genes are under strong purifying selection (mean Ka/Ks 0.32-0.37), with no evidence of positive selection. Pairwise genetic distances (p-distances) among Barringtonia species are extremely low (mean 0.0046), while distances to the related genus Bertholletia are ~&#x2009;6-fold higher, supporting their generic distinction. CONCLUSIONS: Phylogenetic analysis robustly supports Barringtonia as a monophyletic clade (bootstrap&#x2009;=&#x2009;100%), with B. racemosa and B. fusicarpa forming a sister lineage to B. acutangula. This study provides the first high-quality cp. genome resources for the two threatened Barringtonia species, revealing strong structural and sequence conservation but no direct chloroplast genomic correlates of endangerment. The identified polymorphic regions and repeat markers lay a foundation for future population genetics, phylogeographic studies, and conservation-oriented genetic management of these ecologically important coastal plants.

Genome, Chloroplast

Conserved protein folds underpin the diversification of secreted proteins in a fungal pathogen.

BACKGROUND: During host colonization, fungal plant pathogens secrete effector-like proteins that alter host cell physiology and target plant-associated microbes. However, rapid evolution and low sequence conservation hinder the study and characterization of these proteins. The fungus Zymoseptoria passerinii infects Hordeum spp. and includes lineages adapted to wild and domesticated barley. To date, the evolution of effector-like proteins in this species has not been addressed. RESULTS: We combined multiple structure-based and network analyses to unravel the secretome of Z. passerinii. We first compared AlphaFold2 and ESMFold predictions to establish the baseline for structural analyses. We identified 72 structural clusters in the secretome, revealing fold-level relationships across divergent sequences. We showed that effector-like proteins with predicted host immune-interfering functions evolved from a limited group of protein folds, whereas proteins with predicted antimicrobial properties were distributed across fold groups. Physicochemical comparisons indicate that putative antimicrobial effectors predominantly emerged through amino acid replacements on common effector-enriched scaffolds in Z. passerinii, reconfiguring surface charge and electrostatics. We analyzed intra- and interspecific variation in selected effector-enriched families by comparing Z. passerinii proteins and homologs across the genus Zymoseptoria. We describe constrained core folds, with local variation in loop and surface-exposed regions, consistent with fold stability while still enabling protein diversification. We further report that putative antimicrobial effector homologs are broadly distributed across the genus despite sequence divergence. CONCLUSIONS: The secretome of Z. passerinii is organized around common structural folds that support diverse biological roles, including host manipulation and host-associated microbial interactions. Conserved scaffolds combined with surface and physicochemical variation likely contribute to rapid adaptive evolution of effector-like proteins in Z. passerinii.

Fungal Proteins

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569&#x2009;bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A&#x2009;+&#x2009;T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU >&#x2009;1) with a pronounced preference for A/C-ending codons. The &#x394;RSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean&#x2009;=&#x2009;1.209&#x2009;&#xb1;&#x2009;0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae

Transition of Staphylococcus aureus tetracycline resistance plasmid pT181 from independent multicopy replicon to predominantly integrated chromosomal element over 65 years.

Mobile genetic elements (MGEs), including plasmids, phages and genome islands, are major sources of bacterial genetic diversity. The small plasmid pT181 confers tetracycline resistance in bacterial pathogen Staphylococcus aureus via an efflux pump, TetK. pT181 was one of the earliest sequenced S. aureus plasmids, and has been isolated in both clinical and livestock-associated strains for decades, both as an independent replicon and integrated in the chromosome as part of staphylococcal cassette chromosome mec (SCCmec). Bacterial genome analysis tools and high-quality sequences with metadata are publicly available, but these resources remain underleveraged for examining historical data, especially when studying the spread of MGEs across a species and over time. Using publicly available reads and metadata, we explored the evolution of pT181 over almost seven decades of samples to identify temporal trends in sequence evolution, copy number changes, and spread across S. aureus and beyond. pT181 was prevalent across S. aureus (found in 9.5% of 83,366 genomes tested), with a conserved sequence outside of three hypervariable regions. The history of pT181 since 1954 is characterized by spread across strains, significant variation in plasmid copy number of the independent replicon, and increasing frequency of integration of the plasmid into the S. aureus chromosome. We have identified multiple chromosomal integration locations of the plasmid, including outside of the previously characterized SCCmec. We find that pT181 has been transferred across staphylococcaceae and into a Gram-negative species. The repeated integration of pT181 into the chromosome may indicate co-evolution of the plasmid and the host, potentially to facilitate increased antibiotic resistance.

Journal Article

Comparative and Subtractive Genomics Analysis of Multidrug-Resistant Klebsiella pneumoniae Strains for Novel Target Identification and Drug Repurposing Strategies.

The rapid rise of multidrug-resistant (MDR) Klebsiella pneumoniae has created a major global health challenge due to the limited availability of conserved therapeutic targets effective across diverse resistant strains. In this study, an integrative computational target-discovery and drug-repurposing framework was applied to six clinically relevant K. pneumoniae strains. Comparative genomic analysis identified 3012 conserved genes, which were subsequently filtered to nine essential, non-host homologous proteins. Among these, three conserved cytoplasmic proteins (accD, cpxR, and mraZ) were prioritized for functional analysis, with acetyl-CoA carboxylase subunit beta (accD) emerging as the most promising therapeutic target based on sequence conservation, predicted essentiality, subcellular localization, and pathway association. Structural assessment supported the reliability of the predicted accD model, whereas consensus binding-site analysis identified key residues suitable for ligand interaction. Virtual screening of FDA-approved drugs followed by molecular docking identified several compounds with favorable binding profiles toward accD. Subsequent molecular dynamics simulations, including root mean square deviation (RMSD), root mean square fluctuation (RMSF), radius of gyration (Rg), hydrogen-bond occupancy, principal component analysis (PCA), and PCA-based free energy landscape (FEL) analyses, consistently identified tenapanor, micafungin, deferoxamine, and cobicistat as the most stable protein-ligand complexes, with tenapanor exhibiting the most favorable overall structural and thermodynamic stability profile. These findings identify accD as a promising therapeutic target in MDR K. pneumoniae and suggest several FDA-approved compounds as potential candidates for drug repurposing. Although experimental validation is needed to confirm their biological activity and therapeutic potential, this study demonstrates the potential of integrating comparative genomics with molecular dynamics analyses to support antimicrobial target identification and drug repurposing against MDR bacterial pathogens.

Klebsiella pneumoniae

The emergence and diversification of the DUX gene family across placental mammals.

The DUX gene family encodes transcription factors with paired homeodomains. It has critical roles in embryogenesis and disease, including facioscapulohumeral muscular dystrophy (FSHD) and cancer. This study conducts a comparative analysis of the DUX gene family-DUXA, DUXB (including DUXBL), and DUXC (including DUX4 and Dux)-across placental mammals, highlighting their structural diversity within macrosatellite repeat contexts. Using long-read genomes, we explore gene distribution, array patterns, and phylogenetic relationships in various vertebrate species. Our analysis reveals that DUXA and DUXB are highly conserved, with intriguing variations such as intronless forms likely arising from ancestral retrotransposition events. While DUXBL is inconsistently retained across clades, its locus-which in non-placental mammals harbors the ancestral single-homeodomain sDUX gene-served as an evolutionary hub for diversification, giving rise to DUXA, DUXB and DUXC, as well as macrosatellite tandem array structures. Sequence conservation and syntenic analyses demonstrate array adaptability, exemplified by higher-order repeats in orangutans and disrupted patterns of concerted evolution in elephants. Furthermore, analysis of human pseudo-DUX4 arrays indicates their potential role in disease mechanisms, including as possible contributors to rare cases of FSHD, warranting further investigation. This study thus provides insights into DUX-family gene evolution, offering a foundation for future research into developmental roles and disease implications.

Animals

[Genetic analysis of a male with Multiple morphological abnormalities of sperm flagella combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene and a literature review].

OBJECTIVE: To explore the clinical phenotype and genetic etiology of a male with Multiple morphological abnormalities of sperm flagella (MMAF) combined with sperm head abnormalities due to compound heterozygous variants of DNAH1 gene, with an aim to provide guidance for assisted reproductive technology in his family. METHODS: A man with MMAF combined with sperm head abnormalities who visited Women and Children's Hospital of Ningbo University in October 2024 was selected as study subject. Clinical data of the patient's family were retrospectively collected. Peripheral blood samples were collected from the patient and his spouse, and G-banding karyotyping and whole exome sequencing (WES) were carried out. Candidate variants were validated by Sanger sequencing. Conservation of the DNAH1 protein was queried on the UCSC website. The difference between wild type and variant DNAH1 proteins were analyzed using AlphaFold v3.0.1 and PyMOL v2.5.6. The pathogenicity of variant was rated based on the guidelines from American College of Medical Genetics and Genomics (ACMG). Previous literature was searched using keywords "DNAH1 gene" and "multiple morphological abnormalities of the sperm flagella" on CNKI, Wanfang Data Knowledge Service Platform, and PubMed database to identify cases of MMAF attributed to biallelic DNAH1 gene variants. The retrieval period was set from the establishment of the databases to December 31, 2025. The genotypes and clinical phenotypes of patients with biallelic DNAH1 mutations were analyzed. This study was approved by the Medical Ethics Committee of the hospital (Ethics No.: EC2023-094). RESULTS: The 30-year-old patient and his 30-year-old wife had infertility for 2 years. Semen analysis revealed no motile sperm and a 99.0% abnormal morphology rate. Typical MMAF was observed with phase-contrast microscopy. Sperm morphology analysis revealed abnormalities of the head, neck, and tail with an approximate ratio of 9:5:1. The patient's karyotype was 46,XY, and his wife's karyotype was 45,X[4]/47,XXX[1]/46,XX[84]. WES and Sanger sequencing revealed that the patient harbored compound heterozygous variants of the DNAH1 gene, namely c.1435_1444+3del and c.12204_12206del (p.Asn4069del), but their origin remained unidentified. UCSC genome browser query results showed that the amino acid residue at position 4 069 of the DNAH1 protein is highly conserved across various species. Protein structure prediction reveals that, in the wild-type DNAH1 protein, the Asparagine at position 4 069 (Asn4069) can form hydrogen bonds with the Leucine on the main chain at position 4 086 (Leu4086) and the Serine on the side chain at position 4 087 (Ser4087). The c.12204_12206del variant, resulting in deletion of Asn4069, disrupts these hydrogen bonds and does not generate any compensatory interactions. Based on the ACMG guidelines, the c.1435_1444+3del variant was predicted to be likely pathogenic (PM2_Supporting+PVS1), and the c.12204_12206del(p.Asn4069del) variant was rated as likely pathogenic (PM2_Supporting+PM4+PM3+PP4). The couple had elected for in vitro fertilization using donor sperm. During this cycle, 12 oocytes were retrieved, 10 oocytes were successfully fertilized, 1 embryo and 6 blastocysts were obtained. Following the first transfer of a frozen-thawed blastocyst, implantation of an empty gestational sac occurred, which led to a miscarriage. After the second transfer of a high-quality blastocyst, the embryo split into twins following implantation, and the spouse had selected fetal reduction. The gestational age was 33+3 weeks on June 1, 2026. Literature review identified three studies reporting biallelic mutations of the DNAH1 gene in association with MMAF combined with sperm head abnormalities. Together with the patient from this study, a total of 20 patients were included in the analysis. The rate of sperm flagellar abnormalities in these patients was above 80.0%, while the rate of sperm head abnormalities has ranged from 12.0% to 100.0%. In four patients, the genetic basis was unknown. In the remaining 16 patients, 35 mutations were detected, with c.8626-1G>A being the most common (22.9%, 8/35). CONCLUSION: This patient showed MMAF with frequent sperm head defects. Compound heterozygous variants of the DNAH1 gene probably underlay these abnormalities, which in turn has led to his primary infertility. This study revealed the phenotypic variability of MMAF and broadened the mutational spectrum of the DNAH1 gene.

Humans

The auxin gatekeepers: Evolution and diversification of the YUCCA family.

The critically important YUCCA (YUC) gene family is highly conserved and specific to the plant kingdom, primarily responsible for the final and rate-limiting step for indole-3-acetic acid (IAA) biosynthesis. IAA is an essential phytohormone, involved in virtually all aspects of plant growth and development. In addition, IAA is involved in fine-tuning plant responses to biotic and abiotic interactions and stresses. While the YUC gene family has significantly expanded throughout the plant kingdom, a detailed analysis of the evolutionary patterns driving this diversification has not been performed. Here, we present a comprehensive phylogenetic analysis of the YUC family, combining YUCs from species representing key evolutionary plant lineages. The evolutionary history of YUCs is complex and suggests multiple recruitment events via horizontal gene transfer from bacteria. We identify and hierarchically classify the YUC family into an early diverging grade, five distinct classes and 41 subclasses. Angiosperm YUC diversity and expansion are explained in the context of protein sequence conservation, as well as spatial and gene expression patterns. The presented YUC gene landscape offers new perspectives on the distribution and evolutionary trends of this crucial family, which facilitates further YUC characterization within plant development and response to environmental change.

Indoleacetic Acids

Comparative analysis of conserved non-coding elements identifies gene regulatory networks rewired during the water-to-land transition in vertebrates.

The conquest of land by vertebrates has been a pivotal moment in evolutionary history. Adapting to the new habitats necessitated numerous changes in vertebrate anatomy and physiology, creating an enduring imprint on the developmental gene regulatory networks (GRNs) of tetrapods. The increase of high-quality genomic resources over the past decade has made it possible to study the genomic legacy of the water-to-land transition. While much attention has been given to the highly conserved non-coding elements (CNEs) of the genome that share high levels of similarity across evolutionarily diverged clades, recent evidence suggests that perhaps comparable attention should be given to "missing" CNE-s, conserved sequence patches present in extant stem gnathostomes and actinopterygian fishes that have become undetectable in tetrapods during the adaptation to terrestrial life, whether through true sequence loss or divergence beyond alignability. These sequences could help us reveal the relaxation of certain developmental constraints, related to the aquatic lifestyle, that made reaching new adaptive peaks in the developmental landscape possible. In this paper, we search for such CNEs and characterize them in comparison with pan-Gnathostome CNEs, using the zebrafish (Danio rerio) genome as a reference. Our results suggest that the rewiring of developmental networks related to pigmentation and muscle structure formation has left the largest genomic imprint. We also find that components of canonical Wnt and Hedgehog signalling, are enriched among CNEs retained in fish.

cis-regulatory evolution

DescribePROT Database of Residue-Level Protein Structure and Function Annotations.

DescribePROT is a freely available online database of structural and functional descriptors of proteins at the amino acid level. It provides access to 13 diverse descriptors that include sequence conservation, putative secondary structure, solvent accessibility, intrinsic disorder, and signal peptides, and putative annotations of residues that interact with proteins, peptides and nucleic acids. These data can be used to elucidate protein functions, to support efforts to develop therapeutics, and to develop and evaluate future predictors of protein structure and function. DescribePROT includes 7.8&#xa0;billion predictions for 1.4&#xa0;million proteins from 83 complete proteomes of popular model organisms. This information can be downloaded at multiple levels of scope (entire database, specific organisms, and individual proteins) and can be interacted with using a graphical interface that simultaneously displays data on multiple descriptors. We describe the contents of this resource, provide directions on how to use its interface, and offer instructions on how to obtain and interact with the underlying data. Moreover, we briefly discuss plans for a future expansion of this database. DescribePROT is available at http://biomine.cs.vcu.edu/servers/DESCRIBEPROT/ .

Databases, Protein

Genomic evolution of EGF-CFC genes in deuterostomes.

BACKGROUND: EGF-CFC proteins are a bilaterian innovation, but they are best known for their roles in Nodal signaling during gastrulation and left-right patterning in vertebrates. Species with multiple family members show evidence of functional specialization. For example, in mouse, Cripto is required for gastrulation, whereas CFC1 is involved in left-right patterning. However, members of the EGF-CFC family across model organisms exhibit limited sequence conservation beyond the EGF-CFC domain, posing challenges for determining their evolutionary history and functional conservation. RESULTS: In this study, we describe the evolutionary history of the EGF-CFC family of proteins across several branches of deuterostomes, with a particular focus on vertebrates. We trace the EGF-CFC gene family from a single gene in the deuterostome ancestor through its expansion and functional specialization in tetrapods, and subsequent gene loss and translocation in eutherian mammals. Mouse Cripto and CFC1, zebrafish Tdgf1, and each Xenopus EGF-CFC gene (Tdgf1, Tdgf1.2 and Cripto.3) are all descendants of the ancestral deuterostome Tdgf1 gene. CONCLUSIONS: We propose that subsequent to EGF-CFC family expansion in tetrapods, Tdgf1B (Xenopus Tdgf1.2) acquired specialization in the left-right patterning cascade, and then after its translocation in eutherians to a different chromosomal location, CFC1 has maintained that specialization.

Animals

Mutation rate heterogeneity biases variant effect prediction and reveals genuine mutational robustness.

Variant effect predictors (VEPs) are widely used to interpret the functional consequences of human genetic variation. Because most methods rely on sequence conservation, they implicitly treat conservation as evidence of functional constraint. However, substitution patterns across a phylogeny reflect not only selection but also differences in underlying mutation rates. Here, we show that this creates a systematic confounding: most VEPs capture mutation rate variation and misinterpret it as variation in functional importance. Widely used conservation metrics exhibit a related bias; in particular, phyloP scores correlate strongly with mutation rate even at putatively neutral sites. Consequently, variants at low-mutation-rate sites tend to be predicted as more damaging, and variants at highly mutable sites as more tolerated, than warranted by their true functional impact. We also identify a distinct biological signal in experimental measurements of mutational effects on protein stability: amino acid substitutions that are more likely to arise are, on average, less destabilizing than rarer substitutions. This provides empirical support for mutational robustness in the context of protein stability. However, this relationship is insufficient to explain the mutation-rate dependence observed in current VEP outputs. Together, our findings show that mutation rate heterogeneity systematically biases current variant effect prediction frameworks, highlight the need to model mutation probabilities explicitly in future VEPs, and reveal a genuine biological signal of mutational robustness.

conservation scores

Melioribacter sulfuriphilus sp. nov., facultatively anaerobic thermophilic sulfur- and thiosulfate-respiring bacterium from Karmadon hot springs of North Ossetia (Russian Federation).

Novel facultatively anaerobic moderately thermophilic bacteria, strains OK-6-MeT and OK-1-Me, were isolated from the hot springs of Karmadon (North Ossetia, Russian Federation). Gram-stain-negative, motile rods were present singly, in rosettes, and formed biofilms. Both strains grew optimally at 55&#xa0;&#xb0;C, pH&#xa0;7.0 and did not require sodium chloride. They were chemoorganoheterotrophs, growing on mono-, di- and polysaccharides (cellulose, xylan, lichenan, xyloglucan, mannan, locust bean gum, pectin) as well as proteinaceous substrates (gelatin, casein). Growth under anaerobic conditions was observed both in the presence and absence of external electron acceptors (sulfur, thiosulfate, nitrite, arsenate, Fe-citrate, ferrihydrite). Major cellular fatty acids of both strains were iso-C15:0, anteiso-C15:0, and anteiso-C17:0. The size of the genomes were 3.3 and 3.2&#xa0;Mb for strain OK-6-MeT and OK-1-Me, respectively. Genomic DNA G&#xa0;+&#xa0;C content was 37% for both strains. According to the 16S rRNA gene sequence and conserved protein sequences phylogenies, the strains represented a new species of the genus Melioribacter of family Melioribacteraceae within the class Ignavibacteria, for which the name Melioribacter sulfuriphilus sp. nov. is proposed, with type strain OK-6-MeT (= B-3972T&#xa0;=&#xa0;CGMCC 1.18264 T&#xa0;=&#xa0;BIM B-2154T&#xa0;=&#xa0;UQM 41932T). Analysis of OK-1 and OK-6 metagenomes revealed presence of various genes involved in carbon (CO2 fixation, carbohydrate hydrolysis, hydrocarbons degradation, fermentation), nitrogen (nitrate, nitrite, NO and N2O reduction) and sulfur cycles (sulfate reduction, sulfur or thiosulfate reduction, oxidation of sulfur compounds). MAGs OK-1-035 and OK-6-024 almost identical to genomes of strain OK-1-Me and OK-6-MeT presumably are integral part of these complex trophic chains.

Facultative anaerobe