PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Chromosomal-level genome assembly of minute pirate bug Orius nagaii Yasunaga, 1993 (Hemiptera: Anthocoridae).

Species of the genus Orius, diminutive predatory insects that act as natural enemies of other arthropods, are frequently employed in agricultural pest management for controlling various pests, such as thrips, mites, aphids, whiteflies, etc. However, the scarcity of high-quality genomic resources for these predators hinders our comprehension of their population evolution and predation ecology. Consequently, we assembled and annotated a chromosomal-scale genome of Orius nagaii by collating PacBio and Illumina sequencing and Hi-C genomic analysis techniques. The final genome assembly size 152.62 Mb, with scaffold and contig N50 lengths of 11.53 and 2.39 Mb, respectively. It is organized into 12 pairs of autosomes and a pair of XY sex chromosomes. The quality assessment of the genomic data with BUSCO revealed a completeness of 98.5% (n = 1,367). Also, 11,917 protein-coding genes were discovered, with 94.28% of them having functional annotations. The high-quality genome of O. nagaii produced serves as a valuable resource for comprehending the interactions between predatory natural enemies and hosts, along with their evolutionary trajectories.

Animals↗

EucaMOD: a comprehensive multi-omics database for functional genomics research and molecular breeding of fast-growing eucalyptus trees.

Eucalyptus, one of the most widely planted plantation tree species globally, is primarily found in tropical and subtropical regions and contributes significantly to economic and social benefits. With advances in sequencing technologies, there is an increasing demand for the systematic analysis of multi-omics data among Eucalyptus species to enhance genetic breeding efforts. Although several early genomic databases have been established for eucalyptus, they have not been updated in a timely manner and lack recent multi-omics data, rendering them insufficient for current research needs. To address this gap, we developed the eucalyptus multi-omics database (EucaMOD, http://eucalyptusggd.net/eucamod), a comprehensive resource for cross-omics studies. In this study, we functionally annotated 45 eucalyptus genomes and structurally annotated 15, conducting comparative genomics and pan-proteomics analyses across all genomes. Additionally, we analyzed eucalyptus transcriptome, epigenome, and variome data through standardized workflows, enabling the in-depth mining and reanalysis of multi-omics datasets. EucaMOD is the most comprehensive multi-omics database for eucalyptus to date and includes data from 45 genomes (39 species), 870 mRNA-seq samples, 17 miRNA-seq samples, 52 epigenomic datasets (histone modifications and transcription factor binding), and genetic variation data from 1219 samples. To support functional genomics and molecular breeding research, the database is organized into the following 11 modules: Home, Species, Genomics, Comparative genomics, Pan-proteomics, Transcriptomics, Epigenetics, Variomics, Tools, Download, and Help. EucaMOD also offers online analysis tools for data mining, providing free public services to aid eucalyptus gene function and genetic engineering studies.

Eucalyptus↗

Microbial and functional shifts between flare and remission in a single-center cohort of children with inflammatory bowel disease.

BACKGROUND: Gut microbial dysbiosis is central to the pathogenesis of inflammatory bowel disease (IBD). While gut microbiome differences between patients with and without IBD are well established, microbiome changes associated with disease activity and remission remain limited, particularly in paediatric populations. AIM: To examine intra-individual taxonomic and functional gut microbiome changes during transition from active flare to remission under maintenance immunosuppression in a pilot single-center Singapore cohort of children with IBD. METHODS: Paired stool samples and clinical data were collected from seven patients with paediatric IBD [5 Crohn's disease (CD), 2 ulcerative colitis; &#x2264; 18 years] during active disease/flare (visit 1; Pediatric CD Activity Index/Pediatric Ulcerative Colitis Activity Index &#x2265; 10) and subsequent clinical remission (visit 2; Pediatric CD Activity Index/Pediatric Ulcerative Colitis Activity Index < 10). Samples underwent shotgun metagenomic sequencing for high-resolution taxonomic profiling and functional annotation of Kyoto Encyclopaedia of Genes and Genomes pathways. RESULTS: Gut microbial diversity was reduced during flare compared to remission, with Actinobacteria abundance significantly higher in remission. Two distinct microbial clusters differentiated flare and remission states: The remission cluster was enriched with Bifidobacterium adolescentis, Bifidobacterium dentium, Lactobacillus gasseri, Faecalibacterium prausnitzii, while the flare state showed increased Klebsiella pneumoniae. Remission was further characterized by a downregulation of pathogenic microbes and an upregulation of beneficial microbes including a higher abundance of the butyrate producer Anaerostipes hadrus (P = 0.046). Microbial functional genes enriched in remission were predominantly associated with metabolic pathways including vitamin and cofactor biosynthesis, as well as carbohydrate, amino acid, and lipid metabolism. CONCLUSION: The transition from flare to remission in Singaporean children with IBD is characterized by functional remodeling of the gut microbiome, which may contribute to recovery processes related to intestinal barrier integrity, cellular maintenance, and tissue repair. Targeted modulation of the gut microbiome may help sustain remission in paediatric IBD.

Functional shift↗

Dissecting cancer pathways and vulnerabilities with RNAi.

The latest generation of molecular-targeted cancer therapeutics has bolstered the notion that a better understanding of the networks governing cancer pathogenesis can be translated into substantial clinical benefits. However, functional annotation exists for only a small proportion of genes in the human genome, raising the likelihood that many cancer-relevant genes and potential drug targets await identification. Unbiased genetic screens in invertebrate organisms have provided substantial insights into signaling networks underlying many cellular and organismal processes. However, such approaches in mammalian cells have been limited by the lack of genetic tools. The emergence of RNA interference (RNAi) as a mechanism to suppress gene expression has revolutionized genetics in mammalian cells and has begun to facilitate decoding of gene functions on a genome scale. Here, we discuss the application of such RNAi-based genetic approaches to elucidating cancer-signaling networks and uncovering cancer vulnerabilities.

Animals↗

Transcription factor binding element detection using functional clustering of mutant expression data.

As a powerful tool to reveal gene functions, gene mutation has been used extensively in molecular biology studies. With high throughput technologies, such as DNA microarray, genome-wide gene expression changes can be monitored in mutants. Here we present a simple approach to detect the transcription-factor-binding motif using microarray expression data from a mutant in which the relevant transcription factor is deleted. A core part of our approach is clustering of differentially expressed genes based on functional annotations, such as Gene Ontology (GO). We tested our method with eight microarray data sets from the Rosetta Compendium and were able to detect canonical binding motifs for at least four transcription factors. With the support of chromatin IP chip data, we also predict a possible variant of the Swi4 binding motif and recover a core motif for Arg80. Our approach should be readily applicable to microarray experiments using other types of molecular biology techniques, such as conditional knockout/overexpression or RNAi-mediated 'knockdown', to perturb the expression of a transcription factor. Functional clustering included in our approach may also provide new insights into the function of the relevant transcription factor.

Base Sequence↗

Information assessment on predicting protein-protein interactions.

BACKGROUND: Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. RESULTS: Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. CONCLUSIONS: In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.

Artificial Intelligence↗

An enzyme that regulates ether lipid signaling pathways in cancer annotated by multidimensional profiling.

Hundreds, if not thousands, of uncharacterized enzymes currently populate the human proteome. Assembly of these proteins into the metabolic and signaling pathways that govern cell physiology and pathology constitutes a grand experimental challenge. Here, we address this problem by using a multidimensional profiling strategy that combines activity-based proteomics and metabolomics. This approach determined that KIAA1363, an uncharacterized enzyme highly elevated in aggressive cancer cells, serves as a central node in an ether lipid signaling network that bridges platelet-activating factor and lysophosphatidic acid. Biochemical studies confirmed that KIAA1363 regulates this pathway by hydrolyzing the metabolic intermediate 2-acetyl monoalkylglycerol. Inactivation of KIAA1363 disrupted ether lipid metabolism in cancer cells and impaired cell migration and tumor growth in vivo. The integrated molecular profiling method described herein should facilitate the functional annotation of metabolic enzymes in any living system.

Carbamates↗

TransportDB: a comprehensive database resource for cytoplasmic membrane transport systems and outer membrane channels.

TransportDB (http://www.membranetransport.org/) is a comprehensive database resource of information on cytoplasmic membrane transporters and outer membrane channels in organisms whose complete genome sequences are available. The complete set of membrane transport systems and outer membrane channels of each organism are annotated based on a series of experimental and bioinformatic evidence and classified into different types and families according to their mode of transport, bioenergetics, molecular phylogeny and substrate specificities. User-friendly web interfaces are designed for easy access, query and download of the data. Features of the TransportDB website include text-based and BLAST search tools against known transporter and outer membrane channel proteins; comparison of transporter and outer membrane channel contents from different organisms; known 3D structures of transporters, and phylogenetic trees of transporter families. On individual protein pages, users can find detailed functional annotation, supporting bioinformatic evidence, protein/DNA sequences, publications and cross-referenced external online resource links. TransportDB has now been in existence for over 10 years and continues to be regularly updated with new evidence and data from newly sequenced genomes, as well as having new features added periodically.

Bacterial Outer Membrane Proteins↗

Application of a Translational Research Platform to Unveil Efficacy Signals and Mechanisms of Resistance of FGFR Inhibitors in Multiple FGFR-Altered Solid Tumors.

PURPOSE: The predictive value of fibroblast growth factor receptor (FGFR) amplifications (amp) and the role of FGFR mutations (mut) beyond known activating variants remain unclear. We aimed to establish a translational research platform to characterize FGFR alterations (alt) and explore their potential as predictive biomarkers for FGFR-targeted agents. EXPERIMENTAL DESIGN: This ambispective study included a retrospective analysis of patients with FGFR-alt tumors treated with selective FGFR inhibitors (FGFRi) and a prospective collection of longitudinal tumor samples. Patient-derived xenografts (PDX) were generated to investigate FGFRi mechanisms of action and resistance. Molecular characterization included genomic, transcriptomic, proteomic, and functional analyses using the Functional Annotation for Cancer Treatment (FACT) assay. RESULTS: Among 36 retrospectively analyzed patients, clinical benefit from FGFRis was observed in cases with FGFR mRNA overexpression or FGFR2/11q co-amp, but no association was found with the amplification levels. In archival tumor samples, exploratory proteomic analysis showed FGFR1-4 protein expression in 78% of FGFR1/2-amp tumors detected by fluorescence in situ hybridization. RNA sequencing identified a higher prevalence of FGFR mRNA overexpression than proteomic analysis. Among patients harboring FGFR-mut, only one bladder cancer with an FGFR3-mut S249C derived benefit. FACT assay supported the functional activity of selected variants, including FGFR3 T689M, and suggested potential resistance mechanisms involving PI3K/PTEN and MAPK pathway co-alterations. A prospective FGFR-alt PDX biorepository enabled exploratory biomarker analyses, supporting the hypothesis that FGFR1-4 mRNA expression may better reflect FGFR dependency than genomic alterations alone. CONCLUSIONS: These findings highlight the complexity of FGFR-driven oncogenesis and support integrative molecular approaches to refine patient selection for FGFR-targeted therapies.

Humans↗

Structure modeling of all identified G protein-coupled receptors in the human genome.

G protein-coupled receptors (GPCRs), encoded by about 5% of human genes, comprise the largest family of integral membrane proteins and act as cell surface receptors responsible for the transduction of endogenous signal into a cellular response. Although tertiary structural information is crucial for function annotation and drug design, there are few experimentally determined GPCR structures. To address this issue, we employ the recently developed threading assembly refinement (TASSER) method to generate structure predictions for all 907 putative GPCRs in the human genome. Unlike traditional homology modeling approaches, TASSER modeling does not require solved homologous template structures; moreover, it often refines the structures closer to native. These features are essential for the comprehensive modeling of all human GPCRs when close homologous templates are absent. Based on a benchmarked confidence score, approximately 820 predicted models should have the correct folds. The majority of GPCR models share the characteristic seven-transmembrane helix topology, but 45 ORFs are predicted to have different structures. This is due to GPCR fragments that are predominantly from extracellular or intracellular domains as well as database annotation errors. Our preliminary validation includes the automated modeling of bovine rhodopsin, the only solved GPCR in the Protein Data Bank. With homologous templates excluded, the final model built by TASSER has a global C(alpha) root-mean-squared deviation from native of 4.6 angstroms, with a root-mean-squared deviation in the transmembrane helix region of 2.1 angstroms. Models of several representative GPCRs are compared with mutagenesis and affinity labeling data, and consistent agreement is demonstrated. Structure clustering of the predicted models shows that GPCRs with similar structures tend to belong to a similar functional class even when their sequences are diverse. These results demonstrate the usefulness and robustness of the in silico models for GPCR functional analysis. All predicted GPCR models are freely available for noncommercial users on our Web site (http://www.bioinformatics.buffalo.edu/GPCR).

Algorithms↗

Subfunction partitioning, the teleost radiation and the annotation of the human genome.

Half of all vertebrate species are teleost fish. What accounts for this explosion of biodiversity? Recent evidence and advances in evolutionary theory suggest that genomic features could have played a significant role in the teleost radiation. This review examines evidence for an ancient whole-genome duplication (tetraploidization) event that probably occurred just before the teleost radiation. The partitioning of ancestral subfunctions between gene copies arising from this duplication could have contributed to the genetic isolation of populations, to lineage-specific diversification of developmental programs, and ultimately to phenotypic variation among teleost fish. Beyond its importance for understanding mechanisms that generate biodiversity, the partitioning of subfunctions between teleost co-orthologs of human genes can facilitate the identification of tissue-specific conserved noncoding regions and can simplify the analysis of ancestral gene functions obscured by pleiotropy or haploinsufficiency. Applying these principles on a genomic scale can accelerate the functional annotation of the human genome and understanding of the roles of human genes in health and disease.

Animals↗

Chromosome-level genome assembly and annotation of the porcupine fish (Diodon hystrix).

The porcupinefish (Diodon hystrix), a coral reef teleost, is widely distributed in tropical/subtropical waters of the Pacific, Atlantic, Indian Oceans, and Mediterranean Sea. It shares easily recognizable features with pufferfish, such as body inflation and spines. Additionally, its culinary value makes D. hystrix a highly desirable species in many tropical coastal regions, with considerable market potential. However, lack of a high-quality genome hindered further studies on its reproduction, molecular biology, and genomic improvement. Here, we assembled the chromosome-scale genome using PacBio HiFi, ultra-long reads, and Hi-C. Of the 713.62&#x2009;Mb genome, 98.63% anchored to 23 chromosomes (scaffold N50: 31.52&#x2009;Mb) with 39.82% repetitive sequences. The assembled genome achieved a BUSCO completeness score of 97.7%, with 23,171 protein-coding genes predicted, 22,221 of which were functionally annotated. Phylogenetic analysis identified D. hystrix's evolutionary relationships with other species in the Tetraodontiformes. In summary, the high-quality genome of D. hystrix sheds light on valuable insights into genome size evolution, and provides a valuable resource for exploiting genomic study and breeding applications in this species.

Animals↗

Chromosome-level genome assembly and annotation of Spinibarbus caldwelli.

Spinibarbus caldwelli is an economically important freshwater species within the Cyprinidae family, abundant in the middle and lower reaches of the Yangtze River and its adjacent basins. As a promising species suitable for aquaculture in southern China, the lack of genomic resources has hampered the genetic breeding and conservation. Here, we release a chromosome-level genome assembly for S. caldwelli using PacBio HiFi long-reads, Illumina short-reads, and Hi-C sequencing data. The final genome assembly is 1.77&#x2009;Gb in size, with a contig N50 of 24.27&#x2009;Mb. Using Hi-C scaffolding, 99.14% of the contigs were successfully anchored to 50 chromosomes, resulting in a scaffold N50 of 35.29&#x2009;Mb. The final genome assembly shows a BUSCO completeness of 98.27%. The assembled genome contains 49.41% repetitive sequences and 51,505 predicted genes, 90.83% of which have been functionally annotated. This genome provides a genetic basis for S. caldwelli, facilitating the exploration of Cyprinid phylogeny, genetic improvement, and conservation efforts.

Animals↗

Large-scale statistical analysis of secondary xylem ESTs in pine.

A computational analysis of pine transcripts was conducted to contribute to the functional annotation of conifer sequences. A statistical analysis of expressed sequential tags(ESTs) belonging the 7732 contigs in the TIGR Pinus Gene Index (PGI1.0) identified 260 differentially represented gene sequences across six cDNA libraries from loblolly pine secondary xylem. Cluster analysis of this subset of contigs resulted in five groups representing genes preferentially represented in one of the xylem samples (compression wood, plannings, root xylem, latewood) and one group containing mostly genes simultaneously present in compression and side wood libraries. To complement the sequence annotation, 27 cDNA clones representing selected transcripts were completely sequenced. Several genes were identified that could represent putative markers for xylem from different organs, at different stages of development. Several sequences encoding regulatory proteins were over-represented in root xylem as opposed to the other xylem samples. Some of them belonged to known families of plant transcription factors, but two genes were previously uncharacterized in plants. One transcript was homologous to the gene encoding the Smad4 interacting factor, a key co-activator in TGFbeta (transforming growth factor) signalling in animals. Thus, the digital analysis of pine ESTs highlighted a putative gene function of potentially broad interest but that has yet to be investigated in plants. More generally, this study showed that the application of numerical approaches to EST databases should be helpful in establishing priorities among genes to consider for targeted functional studies. Thus, we illustrated the potential of extracting information from conifer sequences already accessible through well-structured public databases.

Amino Acid Sequence↗

From masking repeats to identifying functional repeats in the mouse transcriptome.

The back-to-back release of the mouse genome and the functionally annotated RIKEN mouse full-length cDNA collection was an important milestone in mammalian genomics. Yet much of the data remain to be explored in terms of biological effects and mechanisms. For example, interspersed repeats account for 39 per cent of the mouse genome sequence and 11 per cent of representative transcripts. A considerable number of transposable repeat elements are still active and propagating in mouse compared with human. While existing repeat databases and tools assist the classification of repeats or identification of new repeats, there is little bioinformatic support towards exploring the extent and role of repeats in transcriptional variation, modulation of protein function, or gene regulatory events. Since the mouse is used as a model organism to study human genes and their disease associations, this review focuses on information extraction and collation that captures the functional context of repeats in mouse transcripts to facilitate the biological interpretation and extrapolation of findings to the human.

Animals↗

A new domain family in the superfamily of alkaline phosphatases.

During the course of our large-scale genome analysis a conserved domain, currently detectable only in the genomes of Drosophila melanogaster, Caenorhabditis elegans and Anopheles gambiae, has been identified. The function of this domain is currently unknown and no function annotation is provided for this domain in the publicly available genomic, protein family and sequence databases. The search for the homologues of this domain in the non-redundant sequence database using PSI-BLAST, resulted in identification of distant relationship between this family and the alkaline phosphatase-like superfamily, which includes families of aryl sulfatase, N-acetylgalactosomine-4-sulfatase, alkaline phosphatase and 2,3-bisphosphoglycerate-independent phosphoglycerate mutase (iPGM). The fold recognition procedures showed that this new domain could adopt a similar 3-D fold as for this superfamily. Most of the phosphatases and sulfatases of this superfamily are characterized by functional residues Ser and Cys respectively in the topologically equivalent positions. This functionally important site aligns with Ser/Thr in the members of the new family. Additionally, set of residues responsible for a metal binding site in phosphatases and sulphtases are conserved in the new family. The in-depth analysis suggests that the new family could possess phosphatase activity.

Alkaline Phosphatase↗

Transcriptomic insights into the coordinated regulation of signaling, apoptosis, immunity, and metabolism during Sinonovacula constricta larval metamorphosis.

Metamorphosis is a critical ontogenetic transition for marine bivalves, marking the shift from planktonic to benthic lifestyles, where successful transformation dictates survival. The razor clam Sinonovacula constricta is economically important; however, low larval metamorphosis rates remain a major bottleneck in seedling production. To elucidate the mechanisms governing this process, we performed a comparative transcriptome analysis of S. constricta larvae at pre- and post-metamorphosis stages using Illumina sequencing. A total of 3701 differentially expressed genes (DEGs) were identified, including 3254 up-regulated and 447 down-regulated genes. Functional annotation of the respective top 20 significantly up-regulated and down-regulated DEGs indicated their potential pivotal roles in signal transduction (e.g., up-regulated: CAV1, CHRNA2; down-regulated: APP, NOTCH1), cellular proliferation and differentiation (e.g., up-regulated: TUBA, EGF1; down-regulated: KIF23, TTC25), transcriptional and epigenetic regulation (e.g., up-regulated: NFIL3; down-regulated: OVO, HMX1), substance transport (e.g., up-regulated: LRP2, LRP1B; down-regulated: SLC51A, Slc33a1), substance metabolism (e.g., up-regulated: CPK3, CYP26A1; down-regulated: RDMT1, ADAC), immunomodulation (e.g., up-regulated: CPN2, CRISP2), and protein homeostasis (e.g., up-regulated: HSP27, NAS-27). Functional enrichment analysis further revealed that DEGs were significantly enriched in pathways related to signal transduction and developmental regulation (e.g., Ras, TNF), cell death and homeostasis (e.g., apoptosis), immune responses (e.g., Toll-like receptor), energy metabolism (e.g., lipid), cardiovascular related (e.g., Fluid shear stress), cell junction and architecture (e.g., Tight junction), and infectious disease (e.g., measles). These results suggest a synergistic interplay between signaling, apoptosis, immunity, and metabolism during S. constricta metamorphosis. This study advances our understanding of marine bivalve metamorphosis and offers candidate genes for further mechanistic studies.

Animals↗

Meta-QTL Analysis Reveals Consensus Genomic Regions and Candidate Genes for Resistance to Sudden Death Syndrome in Soybean.

Sudden death syndrome (SDS), caused by Fusarium virguliforme, is one of the most economically important diseases limiting soybean production worldwide. Although numerous quantitative trait loci (QTL) associated with SDS resistance have been reported, inconsistencies among mapping populations, marker systems, and experimental conditions have hindered the identification of robust resistance loci for soybean improvement. In this study, a comprehensive meta-analysis was conducted to integrate published QTL and identify stable consensus genomic regions associated with SDS resistance. After a systematic literature survey and data curation, 153 QTL derived from 14 linkage-mapping studies were analyzed using a custom R-based workflow, resulting in the identification of 23 consensus meta-QTL (MQTL) distributed across 17 chromosomes. Several MQTL, particularly those located on chromosomes 6, 8, 18, and 20, were supported by multiple independent studies and represented major genomic hotspots for SDS resistance. Physical localization and functional annotation of these MQTL identified 217 candidate genes, including genes predicted to be involved in plant defense, signal transduction, transcriptional regulation, and secondary metabolism. Gene Ontology enrichment analysis identified response to salicylic acid as the only biological process that remained significant after FDR correction, whereas Kyoto Encyclopedia of Genes and Genomes pathway analysis did not identify significantly enriched pathways. Independent support using five published genome-wide association studies further supported several MQTL, especially those on chromosomes 6, 18, and 20, thereby increasing confidence in these genomic regions. The identified MQTL and prioritized candidate genes provide potential genomic resources for future marker development, improvement applications, and functional validation aimed at improving soybean resistance to SDS.

Fusarium virguliforme↗