PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

FimX, a multidomain protein connecting environmental signals to twitching motility in Pseudomonas aeruginosa.

Twitching motility is a form of surface translocation mediated by the extension, tethering, and retraction of type IV pili. Three independent Tn5-B21 mutations of Pseudomonas aeruginosa with reduced twitching motility were identified in a new locus which encodes a predicted protein of unknown function annotated PA4959 in the P. aeruginosa genome sequence. Complementation of these mutants with the wild-type PA4959 gene, which we designated fimX, restored normal twitching motility. fimX mutants were found to express normal levels of pilin and remained sensitive to pilus-specific bacteriophages, but they exhibited very low levels of surface pili, suggesting that normal pilus function was impaired. The fimX gene product has a molecular weight of 76,000 and contains four predicted domains that are commonly found in signal transduction proteins: a putative response regulator (CheY-like) domain, a PAS-PAC domain (commonly involved in environmental sensing), and DUF1 (or GGDEF) and DUF2 (or EAL) domains, which are thought to be involved in cyclic di-GMP metabolism. Red fluorescent protein fusion experiments showed that FimX is located at one pole of the cell via sequences adjacent to its CheY-like domain. Twitching motility in fimX mutants was found to respond relatively normally to a range of environmental factors but could not be stimulated by tryptone and mucin. These data suggest that fimX is involved in the regulation of twitching motility in response to environmental cues.

Amino Acid Sequence↗

DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing.

SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).

Alleles↗

Complex Genetics and Regulatory Drivers of Hypermobile Ehlers-Danlos Syndrome: Insights from Genome-Wide Association Study Meta-analysis.

BACKGROUND: Hypermobile Ehlers-Danlos syndrome (hEDS) is the most common subtype of EDS, a group of heritable connective tissue disorders. Clinically, hEDS is defined by generalized joint hypermobility and chronic musculoskeletal pain, but its impact extends beyond the musculoskeletal system. Affected individuals frequently experience autonomic, gastrointestinal, immune, and neuropsychiatric involvement, highlighting both the multisystemic nature of the condition and challenges of diagnosis. In contrast to other EDS subtypes with defined genetic causes, the molecular basis of hEDS has remained elusive. METHODS: We conducted a genome-wide association study (GWAS) of hEDS across three case controls studies, including 1,815 cases and 5,008 ancestry-matched controls. Fixed-effects meta-analysis of 6.2 million variants was complemented with LDAK gene-based association testing, transcriptome-wide association studies, and integrative annotation across multiple tissues and cell types including eQTLs, enhancer marks and open chromatin accessibility profiles, supported by luciferase assays on one candidate variant. LD-score genetic correlations were assessed between hEDS and 19 frequently reported comorbid conditions. RESULTS: Two loci reached genome-wide significance, including a regulatory region near the atypical chemokine receptor 3 gene (ACKR3) on chromosome 2. Functional annotation supports ACKR3 risk alleles colocalize with eQTLs in tibial nerve, alter enhancer activity, and generate a de novo AHR transcription factor regulatory site, implicating neuroimmune and pain signaling pathways. Gene-based and transcriptome-wide analyses identified common variants in a locus containing multiple candidates, including SLC39A13, a zinc transporter critical for connective tissue development previously implicated in a rare form of EDS, and PSMC3, a gene involved in central nervous system development. LD-score regression revealed significant genetic correlations between hEDS and joint hypermobility, myalgic encephalomyelitis/chronic fatigue syndrome, fibromyalgia, depression, anxiety, autism spectrum disorder, migraine, and gastrointestinal diseases. CONCLUSIONS: These results establish the first evidence of common variant contributions to hEDS, supporting a complex, multisystem model involving neuroimmune-stromal dysregulation. Our findings add novel indications to hEDS pathogenesis and provide solid foundations for future molecular definition and therapeutic discovery.

Genome-wide association study↗

Metagenomic Analysis of Gut Microbiome of Persistent Pulmonary Hypertension of the Newborn.

Persistent pulmonary hypertension of the newborn (PPHN) is one of the most common diseases in the neonatal intensive care unit which severely affects neonatal survival. Gut microbes play an increasingly important role in human health, but there are rarely reported how gut microbiota contribute to PPHN. In our study, the metagenomic sequencing of feces from 12 PPHN's neonates and 8 controls were performed to expose the relation between neonatal gut microbes and PPHN disease. Firstly, we found that the abundance of Actinobacteria, Proteobacteria, Bacteroidetes were significantly increased in PPHN compared with controls, but the Firmicutes components was reduced. And some pathogenic strains (like Vibrio metschnikovii) were significantly enriched in the PPHN compared with controls. Secondly, functional annotation of genes found that PPHN up-regulated transmembrane transport, but down-regulated ribosome and ATP binding. Lastly, microbial metabolic pathway enrichment analysis indicated that some metabolic pathway in PPHN were conflicting and contradictory, showed that an abnormally increased metabolism, disturbed protein synthesis and genomic instability in the PPHN neonate. Our results contribute to understanding the changes in the species and function of gut microbiota in PPHN, thus providing a theoretical basis for the explanation and treatment of PPHN.

Gastrointestinal Microbiome↗

Mutation accumulation in a hybrid parthenogenetic vertebrate.

Asexual lineages are thought to experience elevated extinction rates compared with sexual species, yet direct evidence for the underlying genetic causes remains scarce. Muller's ratchet predicts that the absence of recombination in asexual organisms facilitates the accumulation of deleterious mutations, thereby reducing long-term fitness. Here, we test this hypothesis in the hybrid-origin, parthenogenetic whiptail lizard Aspidoscelis tesselatus by integrating short-read RNAseq and long-read IsoSeq data from both the asexual lineage and its parental sexual species. We reconstructed phased transcripts for A. tesselatus to quantify mutation accumulation relative to the parental sexual species. Comparative analyses revealed elevated ω ratios in both parental genomic complements (subgenomes) of the parthenogenetic lineage, consistent with accelerated accumulation of nonsynonymous mutations. Structural variant analyses identified multiple indels in expressed transcripts predicted to disrupt protein domains. Functional annotation indicated that genes affected by both single-nucleotide variants and indels were enriched for roles in chromatin organization, apoptosis regulation, and transcriptional control. While both parental subgenomes showed similar evolutionary patterns, the maternal complement exhibited more structural and missense mutations than the paternal complement. Together, these results provide evidence that mutations accumulate in asexual A. tesselatus in genes involved in core cellular functions, supporting theoretical predictions that Muller's ratchet contributes to mutation accumulation in asexual lineages.

Animals↗

Genome-resolved analysis reveals disruption of gut microbial vitamin B and K2 biosynthesis during Toxoplasma gondii infection in mice.

UNLABELLED: Toxoplasma gondii infection remodels the gut microbiome, yet its impact on microbial vitamin biosynthetic potential and host redox metabolism remains unclear. Here, we integrated mouse gut metagenomes with publicly available metagenome-assembled genomes (MAGs) to construct a genome-resolved atlas of B-vitamin and vitamin K2 biosynthesis. From 45,697 MAGs, we curated 4,771 representative genomes, of which 2,682 met high-quality criteria (completeness &#x2265;90%, contamination <5%). Functional annotation identified 229,717 vitamin-related genes corresponding to 177 Kyoto Encyclopedia of Genes and Genomes (KEGG) orthologs across de novo pathways for eight B vitamins, thiamine (B1), riboflavin (B2), niacin (B3), pantothenate (B5), pyridoxine (B6), biotin (B7), folate (B9), cobalamin (B12), and vitamin K2. Among the high-quality genomes, 1,665 encoded complete de novo pathways for at least one vitamin, highlighting functional specialization and community-level complementarity. Transcripts per million-normalized metagenomic read counts revealed significant differences in KEGG ortholog abundances across six of the nine vitamin pathways. Reanalysis of metagenomic data from infected mice (acute, chronic, and control; n = 10 per group) revealed a stage-dependent reduction in &#x3b1;-diversity of vitamin biosynthesis pathways during acute infection, and a clear &#x3b2;-diversity separation from chronic and control groups. Core niacin biosynthesis genes (nadB, nadA, nadC) displayed phylum-specific redistribution, indicating selective remodeling of microbial NAD+ precursor production under infection-induced metabolic stress. These results suggest that T. gondii infection disrupts cooperative vitamin biosynthetic networks while specifically modulating niacin pathways linked to host NAD+ metabolism. IMPORTANCE: Gut microbes can synthesize essential vitamins, but how infection alters this function is poorly understood. By integrating mouse gut metagenomes with genome-resolved microbial data, we show that Toxoplasma gondii infection reshapes the vitamin biosynthetic potential of the gut microbiome in a stage-dependent manner. Acute infection reduces the diversity of vitamin biosynthesis pathways and shifts the taxonomic distribution of key niacin biosynthesis genes involved in microbial NAD+ precursor production. These findings identify vitamin metabolism, especially niacin-related pathways, as a sensitive functional axis of microbiome remodeling during infection. Our work links microbial taxonomic changes to functional metabolic consequences and suggests that microbiome-mediated regulation of NAD+-related metabolism may contribute to host redox adaptation during T. gondii infection.

B vitamins↗

dictyBase: a new Dictyostelium discoideum genome database.

Dictyostelium discoideum is a powerful and genetically tractable model system used for the study of numerous cellular molecular mechanisms including chemotaxis, phagocytosis and signal transduction. The past 2 years have seen a significant expansion in the scope and accessibility of online resources for Dictyostelium. Recent advances have focused on the development of a new comprehensive online resource called dictyBase (http://dictybase.org). This database not only provides access to genomic data including functional annotation of genes, gene products and chromosomal mapping, but also to extensive biological information such as mutant phenotypes and corresponding reference material. In conjunction with additional sites (http://genome. imb-jena.de/dictyostelium/, http://dictyensembl. bioch.bcm.tmc.edu and http://www.sanger.ac.uk/Projects/D_discoideum/) from the genome sequencing and assembly centers, these improvements have expanded the scope of the Dictyostelium databases making them accessible and useful to any researcher interested in comparative and functional genomics in metazoan organisms.

Animals↗

Conserved codon composition of ribosomal protein coding genes in Escherichia coli, Mycobacterium tuberculosis and Saccharomyces cerevisiae: lessons from supervised machine learning in functional genomics.

Genomics projects have resulted in a flood of sequence data. Functional annotation currently relies almost exclusively on inter-species sequence comparison and is restricted in cases of limited data from related species and widely divergent sequences with no known homologs. Here, we demonstrate that codon composition, a fusion of codon usage bias and amino acid composition signals, can accurately discriminate, in the absence of sequence homology information, cytoplasmic ribosomal protein genes from all other genes of known function in Saccharomyces cerevisiae, Escherichia coli and Mycobacterium tuberculosis using an implementation of support vector machines, SVM(light). Analysis of these codon composition signals is instructive in determining features that confer individuality to ribosomal protein genes. Each of the sets of positively charged, negatively charged and small hydrophobic residues, as well as codon bias, contribute to their distinctive codon composition profile. The representation of all these signals is sensitively detected, combined and augmented by the SVMs to perform an accurate classification. Of special mention is an obvious outlier, yeast gene RPL22B, highly homologous to RPL22A but employing very different codon usage, perhaps indicating a non-ribosomal function. Finally, we propose that codon composition be used in combination with other attributes in gene/protein classification by supervised machine learning algorithms.

Algorithms↗

TAIR: a resource for integrated Arabidopsis data.

The Arabidopsis Information Resource (TAIR; http://arabidopsis.org) provides an integrated view of genomic data for Arabidopsis thaliana. The information is obtained from a battery of sources, including the Arabidopsis user community, the literature, and the major genome centers. Currently TAIR provides information about genes, markers, polymorphisms, maps, sequences, clones, DNA and seed stocks, gene families and proteins. In addition, users can find Arabidopsis publications and information about Arabidopsis researchers. Our emphasis is now on incorporating functional annotations of genes and gene products, genome-wide expression, and biochemical pathway data. Among the tools developed at TAIR, the most notable is the Sequence Viewer, which displays gene annotation, clones, transcripts, markers and polymorphisms on the Arabidopsis genome, and allows zooming in to the nucleotide level. A tool recently released is AraCyc, which is designed for visualization of biochemical pathways. We are also developing tools to extract information from the literature in a systematic way, and building controlled vocabularies to describe biological concepts in collaboration with other database groups. A significant new feature is the integration of the ABRC database functions and stock ordering system, which allows users to place orders for seed and DNA stocks directly from the TAIR site.

Arabidopsis↗

Beyond synexpression relationships: local clustering of time-shifted and inverted gene expression profiles identifies new, biologically relevant interactions.

The complexity of biological systems provides for a great diversity of relationships between genes. The current analysis of whole-genome expression data focuses on relationships based on global correlation over a whole time-course, identifying clusters of genes whose expression levels simultaneously rise and fall. There are, of course, other potential relationships between genes, which are missed by such global clustering. These include activation, where one expects a time-delay between related expression profiles, and inhibition, where one expects an inverted relationship. Here, we propose a new method, which we call local clustering, for identifying these time-delayed and inverted relationships. It is related to conventional gene-expression clustering in a fashion analogous to the way local sequence alignment (the Smith-Waterman algorithm) is derived from global alignment (Needleman-Wunsch). An integral part of our method is the use of random score distributions to assess the statistical significance of each cluster. We applied our method to the yeast cell-cycle expression dataset and were able to detect a considerable number of additional biological relationships between genes, beyond those resulting from conventional correlation. We related these new relationships between genes to their similarity in function (as determined from the MIPS scheme) or their having known protein-protein interactions (as determined from the large-scale two-hybrid experiment); we found that genes strongly related by local clustering were considerably more likely than random to have a known interaction or a similar cellular role. This suggests that local clustering may be useful in functional annotation of uncharacterized genes. We examined many of the new relationships in detail. Some of them were already well-documented examples of inhibition or activation, which provide corroboration for our results. For instance, we found an inverted expression profile relationship between genes YME1 and YNT20, where the latter has been experimentally documented as a bypass suppressor of the former. We also found new relationships involving uncharacterized yeast genes and were able to suggest functions for many of them. In particular, we found a time-delayed expression relationship between J0544 (which has not yet been functionally characterized) and four genes associated with the mitochondria. This suggests that J0544 may be involved in the control or activation of mitochondrial genes. We have also looked at other, less extensive datasets than the yeast cell-cycle and found further interesting relationships. Our clustering program and a detailed website of clustering results is available at http://www.bioinfo.mbb.yale.edu/expression/cluster (or http://www.genecensus.org/expression/cluster).

Algorithms↗

Ab initio protein structure prediction.

Steady progress has been made in the field of ab initio protein folding. A variety of methods now allow the prediction of low-resolution structures of small proteins or protein fragments up to approximately 100 amino acid residues in length. Such low-resolution structures may be sufficient for the functional annotation of protein sequences on a genome-wide scale. Although no consistently reliable algorithm is currently available, the essential challenges to developing a general theory or approach to protein structure prediction are better understood. The energy landscapes resulting from the structure prediction algorithms are only partially funneled to the native state of the protein. This review focuses on two areas of recent advances in ab initio structure prediction-improvements in the energy functions and strategies to search the caldera region of the energy landscapes.

Chemistry, Physical↗

Yeast genomic expression studies using DNA microarrays.

The exploration and characterization of yeast genomic expression programs is providing a wealth of information about yeast biology, as well as other organisms. The intriguing biology of yeast species invites characterization of genomic expression patterns to illuminate the details of cellular physiology. In addition to its value as an interesting organism, yeast maintains its role as an excellent model in which to characterize genomic expression programs. Microarray studies are quickly spreading to plant, animal, and microbial organisms that remain in the early stages of characterization. The extensive knowledge of yeast biology, as well as the relative ease with which yeast studies can be performed and controlled, facilitates interpretation of the genomic expression data. Importantly, existing information about yeast biology, including functional annotations for each gene, is captured and efficiently presented in databases such as the Saccharomyces Genome Database (SGD), the Munich Information Center Yeast Genome Database (MIPS), the Yeast and Pombe Protein Databases (YPD and PPD, respectively), and others. A number of databases also allow the exploration of published genomic expression studies, including the "Expression Connection" at SGD and the Microarray Global Viewer (yMGV) organized by Marc et al. Consulting these databases to retrieve known details about gene function and regulation vastly facilitates interpretation of the genomic expression data, allowing biological hypotheses to be formulated and tested. These hypotheses can be applied to other organisms that may execute genomic expression programs similar to those seen in yeast. Furthermore, as more genomic expression studies in multiple organisms emerge, large-scale data comparisons can be conducted, within and across organisms. Incorporating the results of yeast studies into such comparisons is certain to increase our understanding about the function, regulation, and evolution of genomic expression programs.

Carbocyanines↗

Annotation transfer for genomics: measuring functional divergence in multi-domain proteins.

Annotation transfer is a principal process in genome annotation. It involves "transferring" structural and functional annotation to uncharacterized open reading frames (ORFs) in a newly completed genome from experimentally characterized proteins similar in sequence. To prevent errors in genome annotation, it is important that this process be robust and statistically well-characterized, especially with regard to how it depends on the degree of sequence similarity. Previously, we and others have analyzed annotation transfer in single-domain proteins. Multi-domain proteins, which make up the bulk of the ORFs in eukaryotic genomes, present more complex issues in functional conservation. Here we present a large-scale survey of annotation transfer in these proteins, using scop superfamilies to define domain folds and a thesaurus based on SWISS-PROT keywords to define functional categories. Our survey reveals that multi-domain proteins have significantly less functional conservation than single-domain ones, except when they share the exact same combination of domain folds. In particular, we find that for multi-domain proteins, approximate function can be accurately transferred with only 35% certainty for pairs of proteins sharing one structural superfamily. In contrast, this value is 67% for pairs of single-domain proteins sharing the same structural superfamily. On the other hand, if two multi-domain proteins contain the same combination of two structural superfamilies the probability of their sharing the same function increases to 80% in the case of complete coverage along the full length of both proteins, this value increases further to > 90%. Moreover, we found that only 70 of the current total of 455 structural superfamilies are found in both single and multi-domain proteins and only 14 of these were associated with the same function in both categories of proteins. We also investigated the degree to which function could be transferred between pairs of multi-domain proteins with respect to the degree of sequence similarity between them, finding that functional divergence at a given amount of sequence similarity is always about two-fold greater for pairs of multi-domain proteins (sharing similarity over a single domain) in comparison to pairs of single-domain ones, though the overall shape of the relationship is quite similar. Further information is available at http://partslist.org/func or http://bioinfo.mbb.yale.edu/partslist/func.

Computational Biology↗

Genomes of the ex-type strains of Elsino&#xeb; mangiferae and E. perseae, the causal agents of scab on mango and avocado.

Elsino&#xeb; species are slow-growing, hemibiotrophic to necrotrophic fungi that cause scab diseases on economically important fruit crops. Genome resources for many host-specific species remain limited. We report high-quality draft genome assemblies for the ex-type strains of Elsino&#xeb; mangiferae (CBS 226.50) and E. perseae (CBS 406.34), causal agents of mango and avocado scab, respectively. Among 5 approaches tested, a Nanopore-only NextDenovo assembly produced the most contiguous genomes, yielding 24.5 Mb (E. mangiferae) and 25.1 Mb (E. perseae) assemblies with 13 and 18 contigs, respectively, BUSCO completeness scores of &#x223c;94%, and multiple putative telomere-to-telomere chromosomes. Gene prediction identified 9,134 and 9,243 genes, respectively. Functional annotation revealed enrichment of metabolic and regulatory pathways, including those involved in posttranslational modification, protein transport, and secondary metabolism. Carbohydrate-active enzyme repertoires were small but conserved, consistent with stealth pathogenicity strategies and low plant cell wall degradation. Both genomes encoded large secretomes (>850 proteins), diverse protease repertoires (>300 proteins), Ecp2-like effector proteins, and multiple biosynthetic gene clusters, including clusters with similarity to those associated with elsinochrome and ACT-toxin II biosynthesis, some of which may contribute to host-pathogen interactions and disease development. A large fraction of genes lacked functional characterization, suggesting incomplete databases and/or the presence of lineage-specific genes potentially involved in virulence or host adaptation. These genome resources fill critical gaps for underrepresented Elsino&#xeb; species and provide taxonomically anchored references essential for diagnostics, comparative genomics, and research into the molecular basis of host specificity and pathogenicity in scab-causing fungi.

Persea↗

B cell pathways implicate shared genetic architecture between schizophrenia and immune-mediated diseases.

BACKGROUND: Schizophrenia and immune-mediated diseases are globally prevalent and highly heritable conditions that frequently co-occur, posing major public health burdens. However, their shared genetic architecture remains poorly understood. METHODS: We applied the bivariate causal mixture model (MiXeR) to investigate the polygenic overlap between schizophrenia and eight common immune-mediated diseases, using genome-wide association study summary statistics comprising 2,489 to 67,323 cases and 9,066 to 497,622 controls. Shared loci were identified through conditional/conjunctional false discovery rate (cond/conjFDR), local genetic correlation (LAVA), and colocalization analyses. Subsequently, gene mapping, functional annotation, expression-trait association, and drug-gene interaction analyses were performed to explore shared genes and enriched pathways, and genetic risk scores (GRS) from the UK Biobank were used to validate the findings. RESULTS: MiXeR estimated substantial polygenic overlap between schizophrenia and immune-mediated diseases, and conjFDR identified 133 shared loci, with eight prioritized through local genetic correlation and colocalization signals. These eight loci were mapped to 85 protein-coding genes enriched in pathways essential for B cell function. Among them, S-PrediXcan analyses identified 14 genes whose expression in brain tissues or blood was associated with both diseases. These genes also interact with immunomodulatory or antihypertensive drugs. Additionally, 11 of the 14 genes were linked to innate immunity and/or cognitive traits. Using UK Biobank data, we further confirmed that overall, shared gene, and B cell activation and receptor signaling pathway&#x2013;specific genetic risk for schizophrenia is associated with immune-mediated disease susceptibility. CONCLUSIONS: These findings underscore the shared genetic architecture of schizophrenia and immune-mediated diseases, advancing insights at the interface of psychiatric genetics and immunology.

Schizophrenia↗

Identifying fundamental gaps in functional metagenomics: a step towards unlocking microbiome research potential.

Incomplete functional annotation limits biological interpretation in microbiome studies and their translational potential. Poor annotation arises from multiple causes, with incomplete gene-protein-reaction mapping being one tractable yet under-examined contributor. We address this gap by developing a comprehensive hierarchical framework that systematically integrates gene families in UniRef, proteins in UniProt, and metabolic reactions in MetaCyc and BioCyc through UniProtKB accession, EC number, and Pfam-domain matching. Applied to a human gut metagenome dataset via HUMAnN3, our MetaCyc-based mapping recovers up to 2.3-fold more unique reaction identifiers than the default pipeline and increases reaction prevalence across samples from &#x2248;32% to 52% core reactions, addressing the data sparsity that limits statistical and machine-learning applications in microbiome research. Biological plausibility for the tested functions was supported by positive and negative controls: gut-microbial hormone-metabolism reactions previously linked to this dataset were recovered, while vertebrate-specific hormone-metabolism reactions remained correctly undetected. These gains derive from systematic database integration alone, without predictive algorithms, indicating that a tractable, mapping-related component of functional dark matter and data sparsity in microbiome studies is directly addressable. Because Pfam- and BioCyc-derived mappings trade specificity for coverage, confidence in any individual reaction assignment depends on the supporting evidence tier and source database.

Humans↗

Chromosomal-level genome assembly of minute pirate bug Orius nagaii Yasunaga, 1993 (Hemiptera: Anthocoridae).

Species of the genus Orius, diminutive predatory insects that act as natural enemies of other arthropods, are frequently employed in agricultural pest management for controlling various pests, such as thrips, mites, aphids, whiteflies, etc. However, the scarcity of high-quality genomic resources for these predators hinders our comprehension of their population evolution and predation ecology. Consequently, we assembled and annotated a chromosomal-scale genome of Orius nagaii by collating PacBio and Illumina sequencing and Hi-C genomic analysis techniques. The final genome assembly size 152.62&#x2009;Mb, with scaffold and contig N50 lengths of 11.53 and 2.39&#x2009;Mb, respectively. It is organized into 12 pairs of autosomes and a pair of XY sex chromosomes. The quality assessment of the genomic data with BUSCO revealed a completeness of 98.5% (n&#x2009;=&#x2009;1,367). Also, 11,917 protein-coding genes were discovered, with 94.28% of them having functional annotations. The high-quality genome of O. nagaii produced serves as a valuable resource for comprehending the interactions between predatory natural enemies and hosts, along with their evolutionary trajectories.

Animals↗

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗