PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “functional annotations”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

Comparative genomics reveals population structure and functional differentiation in Limosilactobacillus fermentum.

Limosilactobacillus fermentum is a widely distributed lactic acid bacterium frequently detected in fermented foods and host-associated microbiota, yet its global genomic diversity and functional variability remain insufficiently characterized. Here, we performed a large-scale comparative genomic analysis of 336 high-quality L. fermentum genomes curated from public databases. Species identity was validated using average nucleotide identity (ANI), and population structure was examined using pairwise ANI comparisons together with Mash-based phylogenetic reconstruction. Clustering at ≥ 99% ANI resolved the dataset into 15 genomic clusters, with four dominant lineages comprising the majority of genomes. Pangenome reconstruction identified 5,853 gene clusters, including 1,325 core genes (22.6%) and a large accessory component dominated by low-frequency genes. Heap's law modeling (λ = 0.19) indicated a weakly open pangenome, suggesting ongoing gene acquisition as additional genomes are sampled. Functional annotation revealed that core genes were primarily associated with essential cellular processes, whereas accessory genes were enriched in carbohydrate metabolism, membrane-associated functions, and defense-related systems. Variation in carbohydrate-active enzymes (CAZymes), transport systems, and stress-response genes was observed across lineages, indicating strain-level functional diversity. Although genomes from human and food sources were broadly distributed across phylogenetic lineages, multivariate analysis showed that gene-content variation was more strongly associated with genomic lineage than with isolation source. These results provide a population genomic framework for understanding genomic diversity and functional potential in L. fermentum.

Phylogeny↗

Functional bioinformatics for Arabidopsis thaliana.

MOTIVATION: The genome of Arabidopsis thaliana, which has the best understood plant genome, still has approximately one-third of its genes with no functional annotation at all from either MIPS or TAIR. We have applied our Data Mining Prediction (DMP) method to the problem of predicting the functional classes of these protein sequences. This method is based on using a hybrid machine-learning/data-mining method to identify patterns in the bioinformatic data about sequences that are predictive of function. We use data about sequence, predicted secondary structure, predicted structural domain, InterPro patterns, sequence similarity profile and expressions data. RESULTS: We predicted the functional class of a high percentage of the Arabidopsis genes with currently unknown function. These predictions are interpretable and have good test accuracies. We describe in detail seven of the rules produced.

Algorithms↗

Proteomic analysis of rat hippocampal plasma membrane: characterization of potential neuronal-specific plasma membrane proteins.

The hippocampus is a distinct brain structure that is crucial in memory storage and retrieval. To identify comprehensively proteins of hippocampal plasma membrane (PM) and detect the neuronal-specific PM proteins, we performed a proteomic analysis of rat hippocampus PM using the following three technical strategies. First, proteins of the PM were purified by differential and density-gradient centrifugation from hippocampal tissue and separated by one-dimensional electophoresis, digested with trypsin and analyzed by electrospray ionization (ESI) quadrupole time-of-flight (Q-TOF) tandem mass spectrometry (MS/MS). Second, the tryptic peptide mixture from PMs purified from hippocampal tissue using the centrifugation method was analyzed by liquid chromatography ion-trap ESI-MS/MS. Finally, the PM proteins from primary hippocampal neurons purified by a biotin-directed affinity technique were separated by one-dimensional electrophoresis, digested with trypsin and analyzed by ESI-Q-TOF-MS/MS. A total of 345, 452 and 336 non-redundant proteins were identified by each technical procedure respectively. There was a total of 867 non-redundant protein entries, of which 64.9% are integral membrane or membrane-associated proteins. One hundred and eighty-one proteins were detected only in the primary neurons and could be regarded as neuronal PM marker candidates. We also found some hypothetical proteins with no functional annotations that were first found in the hippocampal PM. This work will pave the way for further elucidation of the mechanisms of hippocampal function.

Animals↗

Discovery of 342 putative new genes from the analysis of 5'-end-sequenced full-length-enriched cDNA human transcripts.

In this work we describe the process that, starting with the production of human full-length-enriched cDNA libraries using the CAP-Trapper method, led us to the discovery of 342 putative new human genes. Twenty-three thousand full-length-enriched clones, obtained from various cell lines and tissues in different developmental stages, were 5'-end sequenced, allowing the identification of a pool of 5300 unique cDNAs. By comparing these sequences to various human and vertebrate nucleotide databases we found that about 40% of our clones extended previously annotated 5' ends, 662 clones were likely to represent splice variants of known genes, and finally 342 clones remained unknown, with no or poor functional annotation. cDNA-microarray gene expression analysis showed that 260 of 342 unknown clones are expressed in at least one cell line and/or tissue. Further analysis of their sequences and the corresponding genomic locations allowed us to conclude that most of them represent potential novel genes, with only a small fraction having protein-coding potential.

5' Flanking Region↗

Chromosome-level genome assembly of Sinocyclocheilus jii based on PacBio HiFi and Hi-C sequencing.

Sinocyclocheilus jii, a cavefish species endemic to China, belongs to the genus Sinocyclocheilus within the family Cyprinidae. Species within this genus exhibit significant morphological differentiation, making it not only the most species-rich genus within Cyprinidae in China but also the most diverse group of cavefishes worldwide. However, the limited availability of genomic resources has limited investigations into the genetic basis of trait variations, phylogenetic relationships, and adaptive evolution in this genus. In this study, we assembled a chromosome-level reference genome for S. jii by integrating PacBio HiFi long reads, Illumina short reads, and Hi-C sequencing data. Flow cytometry was used to estimate the genome size prior to assembly, providing a key step in technical validation. The final genome assembly spans 1.75 Gb with a contig N50 of 35.0 Mb. Using Hi-C sequencing data, the assembled scaffolds were successfully anchored to 50 chromosomes. The completeness of the chromosome-level assembly was estimated at 98.9% by BUSCO analysis. Genome annotation identified 855.5 Mb of repetitive sequences and predicted a total of 52,867 protein-coding genes, of which 51,932 genes were functionally annotated. This study presents a high-quality chromosome-level genome assembly and annotation of S. jii, providing a fundamental genomic resource for future phylogenetic and evolutionary studies.

Animals↗

The genome-wide transcriptional responses of Saccharomyces cerevisiae grown on glucose in aerobic chemostat cultures limited for carbon, nitrogen, phosphorus, or sulfur.

Profiles of genome-wide transcriptional events for a given environmental condition can be of importance in the diagnosis of poorly defined environments. To identify clusters of genes constituting such diagnostic profiles, we characterized the specific transcriptional responses of Saccharomyces cerevisiae to growth limitation by carbon, nitrogen, phosphorus, or sulfur. Microarray experiments were performed using cells growing in steady-state conditions in chemostat cultures at the same dilution rate. This enabled us to study the effects of one particular limitation while other growth parameters (pH, temperature, dissolved oxygen tension) remained constant. Furthermore, the composition of the media fed to the cultures was altered so that the concentrations of excess nutrients were comparable between experimental conditions. In total, 1881 transcripts (31% of the annotated genome) were significantly changed between at least two growth conditions. Of those, 484 were significantly higher or lower in one limitation only. The functional annotations of these genes indicated cellular metabolism was altered to meet the growth requirements for nutrient-limited growth. Furthermore, we identified responses for several active transcription factors with a role in nutrient assimilation. Finally, 51 genes were identified that showed 10-fold higher or lower expression in a single condition only. The transcription of these genes can be used as indicators for the characterization of nutrient-limited growth conditions and provide information for metabolic engineering strategies.

Carbon↗

Correlated fragile site expression allows the identification of candidate fragile genes involved in immunity and associated with carcinogenesis.

BACKGROUND: Common fragile sites (cfs) are specific regions in the human genome that are particularly prone to genomic instability under conditions of replicative stress. Several investigations support the view that common fragile sites play a role in carcinogenesis. We discuss a genome-wide approach based on graph theory and Gene Ontology vocabulary for the functional characterization of common fragile sites and for the identification of genes that contribute to tumour cell biology. RESULTS: Common fragile sites were assembled in a network based on a simple measure of correlation among common fragile site patterns of expression. By applying robust measurements to capture in quantitative terms the non triviality of the network, we identified several topological features clearly indicating departure from the Erdos-Renyi random graph model. The most important outcome was the presence of an unexpected large connected component far below the percolation threshold. Most of the best characterized common fragile sites belonged to this connected component. By filtering this connected component with Gene Ontology, statistically significant shared functional features were detected. Common fragile sites were found to be enriched for genes associated to the immune response and to mechanisms involved in tumour progression such as extracellular space remodeling and angiogenesis. Moreover we showed how the internal organization of the graph in communities and even in very simple subgraphs can be a starting point for the identification of new factors of instability at common fragile sites. CONCLUSION: We developed a computational method addressing the fundamental issue of studying the functional content of common fragile sites. Our analysis integrated two different approaches. First, data on common fragile site expression were analyzed in a complex networks framework. Second, outcomes of the network statistical description served as sources for the functional annotation of genes at common fragile sites by means of the Gene Ontology vocabulary. Our results support the hypothesis that fragile sites serve a function; we propose that fragility is linked to a coordinated regulation of fragile genes expression.

Cells, Cultured↗

Combining gene annotations and gene expression data in model-based clustering: weighted method.

It has been increasingly recognized that incorporating prior knowledge into cluster analysis can result in more reliable and meaningful clusters. In contrast to the standard modelbased clustering with a global mixture model, which does not use any prior information, a stratified mixture model was recently proposed to incorporate gene functions or biological pathways as priors in model-based clustering of gene expression profiles: various gene functional groups form the strata in a stratified mixture model. Albeit useful, the stratified method may be less efficient than the global analysis if the strata are non-informative to clustering. We propose a weighted method that aims to strike a balance between a stratified analysis and a global analysis: it weights between the clustering results of the stratified analysis and that of the global analysis; the weight is determined by data. More generally, the weighted method can take advantage of the hierarchical structure of most existing gene functional annotation systems, such as MIPS and Gene Ontology (GO), and facilitate choosing appropriate gene functional groups as priors. We use simulated data and real data to demonstrate the feasibility and advantages of the proposed method.

Algorithms↗

De novo assembly of transcriptomes of six Hua species (Semisulcospiridae, Cerithioidea, Gastropoda).

Species in Semisulcospiridae are important in freshwater ecology and have great research value, yet their genomic resources remain very limited. Here, we present de novo assembled transcriptomes from six species of Hua in Semisulcospiridae, including Hua textrix (Heude, 1888), H. yangi L.-N. Du, J.-X. Yang & Chen, 2023, H. wujiangensis L.-N. Du, J.-X. Yang & Chen, 2023, and three undescribed species. Assembly was performed using Trinity, resulting in average contig lengths ranging from 716.6 to 883.3 bp and transcript numbers ranging from 147,147 to 268,741. Benchmarking Universal Single-Copy Ortholog (BUSCO) analysis was used to assess the transcriptome completeness. The functional annotation of transcripts for each species had over 18,000 BLAST hits, 17,000 GO terms, 15,000 KEGG pathways, 8,000 Pfam accessions, and 140 COG functional categories. This study provides valuable transcriptomic resources for the six Hua species, which can be used for various research of Semisulcospiridae, including biodiversity, phylogeny, and comparative genomics.

Transcriptome↗

Annotating proteins by mining protein interaction networks.

MOTIVATION: In general, most accurate gene/protein annotations are provided by curators. Despite having lesser evidence strengths, it is inevitable to use computational methods for fast and a priori discovery of protein function annotations. This paper considers the problem of assigning Gene Ontology (GO) annotations to partially annotated or newly discovered proteins. RESULTS: We present a data mining technique that computes the probabilistic relationships between GO annotations of proteins on protein-protein interaction data, and assigns highly correlated GO terms of annotated proteins to non-annotated proteins in the target set. In comparison with other techniques, probabilistic suffix tree and correlation mining techniques produce the highest prediction accuracy of 81% precision with the recall at 45%. AVAILABILITY: Code is available upon request. Results and used materials are available online at http://kirac.case.edu/PROTAN.

Amino Acid Sequence↗

The TIGR Rice Genome Annotation Resource: improvements and new features.

In The Institute for Genomic Research Rice Genome Annotation project (http://rice.tigr.org), we have continued to update the rice genome sequence with new data and improve the quality of the annotation. In our current release of annotation (Release 4.0; January 12, 2006), we have identified 42,653 non-transposable element-related genes encoding 49,472 gene models as a result of the detection of alternative splicing. We have refined our identification methods for transposable element-related genes resulting in 13,237 genes that are related to transposable elements. Through incorporation of multiple transcript and proteomic expression data sets, we have been able to annotate 24 799 genes (31,739 gene models), representing approximately 50% of the total gene models, as expressed in the rice genome. All structural and functional annotation is viewable through our Rice Genome Browser which currently supports 59 tracks. Enhanced data access is available through web interfaces, FTP downloads and a Data Extractor tool developed in order to support discrete dataset downloads.

DNA Transposable Elements↗

Impact of commensal microbiota on murine gastrointestinal tract gene ontologies.

The gastrointestinal tract (GIT) of eukaryotes is colonized by a vast number of bacteria, where the commensal microbiota play an important role in defining the healthy gut. To investigate the influence of commensal bacteria on multiple regions of the host GIT transcriptome, the gene expression profiles of the corpus, jejunum, descending colon, and rectum of conventional (n = 3) and germ-free mice (n = 3) were examined using the Affymetrix Mu74Av2 GeneChip. Differentially regulated genes were identified using the global error assessment model, and a novel method of Gene Ontology (GO) clustering was used to identify significantly modulated biological functions. The microbiota modify the greatest number of genes in the jejunum (267 genes with an alpha < 0.001) and the fewest in the rectum (137 genes with an alpha < 0.001). Clustering genes by GO biological process and molecular function annotations revealed that, despite the large number of differentially regulated genes, the residential microbiota most significantly modified genes involved in such biological processes as immune function and water transport all along the length of the mouse GIT. Additionally, region-specific communication between the host and microbiota were identified in the corpus and jejunum, where tissue kallikrein and apoptosis regulator activities were modulated, respectively. These findings identify important interactions between the microbiota and the mouse gut tissue transcriptome and, furthermore, suggest that interactions between the microbial population and host GIT are implicated in the coordination of region-specific functions.

Animals↗

Gene array analysis of bone morphogenetic protein type I receptor-induced osteoblast differentiation.

UNLABELLED: The genomic response to BMP was investigated by ectopic expression of activated BMP type I receptors in C2C12 myoblast using cDNA microarrays. Novel BMP receptor target genes with possible roles in inhibition of myoblast differentiation and stimulation of osteoblast differentiation were identified. INTRODUCTION: Bone morphogenetic proteins (BMPs) have an important role in controlling mesenchymal cell fate and mediate these effects by regulating gene expression. BMPs signal through three distinct specific BMP type I receptors (also termed activin receptor-like kinases) and their downstream nuclear effectors, termed Smads. The critical target genes by which activated BMP receptors mediate change cell fate are poorly characterized. MATERIALS AND METHODS: We performed transcriptional profiling of C2C12 myoblasts differentiation into osteoblast-like cells by ectopic expression of three distinct constitutively active (ca)BMP type I receptors using adenoviral gene transfer. Cells were harvested 48 h after infection, which allowed detection of both early and late response genes. Expression analysis was performed using the mouse GEM1 microarray, which is comprised of approximately 8700 unique sequences. Hybridizations were performed in duplicate with a reverse fluor labeling. Genes were considered to be significantly regulated if the p value for differential expression was less than 0.01 and inverted expression ratios per duplicate successful reciprocal hybridizations differed by less than 25%. RESULTS AND CONCLUSIONS: Each of the three caBMP type I receptors stimulated equal levels of R-Smad phosphorylation and alkaline phosphatase activity, an early marker for osteoblast differentiation. Interestingly, all three type I receptors induced identical transcriptional profiles; 97 genes were significantly upregulated and 103 genes were downregulated. Many extracellular matrix genes were upregulated, muscle-related genes downregulated, and transcription factors/signaling components modulated. In addition to 41 expressed sequence tags without known function and a number of known BMP target genes, including PPAR-gamma and fibromodulin, a large number of novel BMP target genes with an annotated function were identified, including transcription factors HesR1, ITF-2, and ICSBP, apoptosis mediators DRP-1 death kinase and ZIP kinase, IkappaB alpha, Edg-2, ZO-1, and E3 ligase Dactylin. These target genes, some of them unexpected, offer new insights into how BMPs elicit biological effects, in particular into the mechanism of inhibition of myoblast differentiation and stimulation of osteoblast differentiation.

Animals↗

The TIGR Gene Indices: clustering and assembling EST and known genes and integration with eukaryotic genomes.

Although the list of completed genome sequencing projects has expanded rapidly, sequencing and analysis of expressed sequence tags (ESTs) remain a primary tool for discovery of novel genes in many eukaryotes and a key element in genome annotation. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi) are a collection of 77 species-specific databases that use a highly refined protocol to analyze gene and EST sequences in an attempt to identify and characterize expressed transcripts and to present them on the Web in a user-friendly, consistent fashion. A Gene Index database is constructed for each selected organism by first clustering, then assembling EST and annotated cDNA and gene sequences from GenBank. This process produces a set of unique, high-fidelity virtual transcripts, or tentative consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to genetic and physical maps, to provide links to orthologous and paralogous genes, and as a resource for comparative and functional genomic analysis.

Animals↗

Automatic extraction of gene/protein biological functions from biomedical text.

MOTIVATION: With the rapid advancement of biomedical science and the development of high-throughput analysis methods, the extraction of various types of information from biomedical text has become critical. Since automatic functional annotations of genes are quite useful for interpreting large amounts of high-throughput data efficiently, the demand for automatic extraction of information related to gene functions from text has been increasing. RESULTS: We have developed a method for automatically extracting the biological process functions of genes/protein/families based on Gene Ontology (GO) from text using a shallow parser and sentence structure analysis techniques. When the gene/protein/family names and their functions are described in ACTOR (doer of action) and OBJECT (receiver of action) relationships, the corresponding GO-IDs are assigned to the genes/proteins/families. The gene/protein/family names are recognized using the gene/protein/family name dictionaries developed by our group. To achieve wide recognition of the gene/protein/family functions, we semi-automatically gather functional terms based on GO using co-occurrence, collocation similarities and rule-based techniques. A preliminary experiment demonstrated that our method has an estimated recall of 54-64% with a precision of 91-94% for actually described functions in abstracts. When applied to the PUBMED, it extracted over 190 000 gene-GO relationships and 150 000 family-GO relationships for major eukaryotes.

Abstracting and Indexing↗

Operon prediction without a training set.

MOTIVATION: Annotation of operons in a bacterial genome is an important step in determining an organism's transcriptional regulatory program. While extensive studies of operon structure have been carried out in a few species such as Escherichia coli, fewer resources exist to inform operon prediction in newly sequenced genomes. In particular, many extant operon finders require a large body of training examples to learn the properties of operons in the target organism. For newly sequenced genomes, such examples are generally not available; moreover, a model of operons trained on one species may not reflect the properties of other, distantly related organisms. We encountered these issues in the course of predicting operons in the genome of Bacteroides thetaiotaomicron (B.theta), a common anaerobe that is a prominent component of the normal adult human intestinal microbial community. RESULTS: We describe an operon predictor designed to work without extensive training data. We rely on a small set of a priori assumptions about the properties of the genome being annotated that permit estimation of the probability that two adjacent genes lie in a common operon. Predictions integrate several sources of information, including intergenic distance, common functional annotation and a novel formulation of conserved gene order. We validate our predictor both on the known operons of E.coli and on the genome of B.theta, using expression data to evaluate our predictions in the latter.

Algorithms↗

The Legume Information System (LIS): an integrated information resource for comparative legume biology.

The Legume Information System (LIS) (http://www.comparative-legumes.org), developed by the National Center for Genome Resources in cooperation with the USDA Agricultural Research Service (ARS), is a comparative legume resource that integrates genetic and molecular data from multiple legume species enabling cross-species genomic and transcript comparisons. The LIS virtual plant interface allows simplified and intuitive navigation of transcript data from Medicago truncatula, Lotus japonicus, Glycine max and Arabidopsis thaliana. Transcript libraries are represented as images of plant organs in different developmental stages, which are selected to query the analyzed and annotated data. Complex queries can be accomplished by adding modifiers, keywords and sequence names. The LIS also contains annotated genomic data featuring transcript alignments to validate gene predictions as well as motif and similarity analyses. The genomic browser supports comparative analysis via novel dynamic functional annotation comparisons. CMap, developed as part of the GMOD project (http://www.gmod.org/cmap/index.shtml), has been incorporated to support comparative analyses of community linkage and physical map data. LIS is being expanded to incorporate gene expression and biochemical pathways which will be seamlessly integrated forming a knowledge discovery framework.

Arabidopsis↗

Use of a large-scale Triticeae expressed sequence tag resource to reveal gene expression profiles in hexaploid wheat (Triticum aestivum L.).

The US Wheat Genome Project, funded by the National Science Foundation, developed the first large public Triticeae expressed sequence tag (EST) resource. Altogether, 116,272 ESTs were produced, comprising 100,674 5' ESTs and 15 598 3' ESTs. These ESTs were derived from 42 cDNA libraries, which were created from hexaploid bread wheat (Triticum aestivum L.) and its close relatives, including diploid wheat (T. monococcum L. and Aegilops speltoides L.), tetraploid wheat (T. turgidum L.), and rye (Secale cereale L.), using tissues collected from various stages of plant growth and development and under diverse regimes of abiotic and biotic stress treatments. ESTs were assembled into 18,876 contigs and 23,034 singletons, or 41,910 wheat unigenes. Over 90% of the contigs contained fewer than 10 EST members, implying that the ESTs represented a diverse selection of genes and that genes expressed at low and moderate to high levels were well sampled. Statistical methods were used to study the correlation of gene expression patterns, based on the ESTs clustered in the 1536 contigs that contained at least 10 5' EST members and thus representing the most abundant genes expressed in wheat. Analysis further identified genes in wheat that were significantly upregulated (p < 0.05) in tissues under various abiotic stresses when compared with control tissues. Though the function annotation cannot be assigned for many of these genes, it is likely that they play a role associated with the stress response. This study predicted the possible functionality for 4% of total wheat unigenes, which leaves the remaining 96% with their functional roles and expression patterns largely unknown. Nonetheless, the EST data generated in this project provide a diverse and rich source for gene discovery in wheat.

Cluster Analysis↗