PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “genomic data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Finding function: evaluation methods for functional genomic data.

BACKGROUND: Accurate evaluation of the quality of genomic or proteomic data and computational methods is vital to our ability to use them for formulating novel biological hypotheses and directing further experiments. There is currently no standard approach to evaluation in functional genomics. Our analysis of existing approaches shows that they are inconsistent and contain substantial functional biases that render the resulting evaluations misleading both quantitatively and qualitatively. These problems make it essentially impossible to compare computational methods or large-scale experimental datasets and also result in conclusions that generalize poorly in most biological applications. RESULTS: We reveal issues with current evaluation methods here and suggest new approaches to evaluation that facilitate accurate and representative characterization of genomic methods and data. Specifically, we describe a functional genomics gold standard based on curation by expert biologists and demonstrate its use as an effective means of evaluation of genomic approaches. Our evaluation framework and gold standard are freely available to the community through our website. CONCLUSION: Proper methods for evaluating genomic data and computational approaches will determine how much we, as a community, are able to learn from the wealth of available data. We propose one possible solution to this problem here but emphasize that this topic warrants broader community discussion.

Algorithms↗

Clustering of diverse genomic data using information fusion.

MOTIVATION: Genome sequencing projects and high-through-put technologies like DNA and Protein arrays have resulted in a very large amount of information-rich data. Microarray experimental data are a valuable, but limited source for inferring gene regulation mechanisms on a genomic scale. Additional information such as promoter sequences of genes/DNA binding motifs, gene ontologies, and location data, when combined with gene expression analysis can increase the statistical significance of the finding. This paper introduces a machine learning approach to information fusion for combining heterogeneous genomic data. The algorithm uses an unsupervised joint learning mechanism that identifies clusters of genes using the combined data. RESULTS: The correlation between gene expression time-series patterns obtained from different experimental conditions and the presence of several distinct and repeated motifs in their upstream sequences is examined here using publicly available yeast cell-cycle data. The results show that the combined learning approach taken here identifies correlated genes effectively. The algorithm provides an automated clustering method, but allows the user to specify apriori the influence of each data type on the final clustering using probabilities. AVAILABILITY: Software code is available by request from the first author. CONTACT: jkasturi@cse.psu.edu.

Algorithms↗

Mass spectrometric genomic data mining: Novel insights into bioenergetic pathways in Chlamydomonas reinhardtii.

A new high-throughput computational strategy was established that improves genomic data mining from MS experiments. The MS/MS data were analyzed by the SEQUEST search algorithm and a combination of de novo amino acid sequencing in conjunction with an error-tolerant database search tool, operating on a 256 processor computer cluster. The error-tolerant search tool, previously established as GenomicPeptideFinder (GPF), enables detection of intron-split and/or alternatively spliced peptides from MS/MS data when deduced from genomic DNA. Isolated thylakoid membranes from the eukaryotic green alga Chlamydomonas reinhardtii were separated by 1-D SDS gel electrophoresis, protein bands were excised from the gel, digested in-gel with trypsin and analyzed by coupling nano-flow LC with MS/MS. The concerted action of SEQUEST and GPF allowed identification of 2622 distinct peptides. In total 448 peptides were identified by GPF analysis alone, including 98 intron-split peptides, resulting in the identification of novel proteins, improved annotation of gene models, and evidence of alternative splicing.

Algorithms↗

AI-HOPE: an AI-driven conversational agent for enhanced clinical and genomic data integration in precision medicine research.

MOTIVATION: The growing complexity of clinical cancer research has fueled a surge in demand for automated bioinformatics tools capable of integrating clinical and genomic data to accelerate discovery efforts. RESULTS: We present the Artificial Intelligence Agent for High-Optimization and Precision Medicine (AI-HOPE), an AI-driven system that enables domain experts to conduct integrative data analyses through natural language interactions. Powered by Large Language Models, AI-HOPE interprets user instructions, converts them into executable code, and autonomously analyzes locally stored data. It supports flexible association studies, subset comparisons, clinical prevalence assessments and survival analyses. In addition, AI-HOPE enables global variable scans to identify features significantly associated with a user-defined outcome, making a powerful and intuitive tool for advancing precision medicine research. Importantly, its closed-system design prevents clinical data leakage. To demonstrate its utility, AI-HOPE was applied to The Cancer Genome Atlas data to address two clinical questions. First, it identified significant enrichment of TP53 mutations in late-stage colorectal cancer compared to early-stage cases. Second, it uncovered a strong association between KRAS mutations and poorer progression-free survival in FOLFOX-treated patients. These findings align with established literature and demonstrate AI-HOPE's ability to generate meaningful insights independently, without prior assumptions. By removing programming barriers and simplifying complex analyses, AI-HOPE bridges the gap between data complexity and research needs. With its scalable and adaptable framework, AI-HOPE has the potential to support diverse biomedical research fields, driving innovation and efficiency in translational studies. AVAILABILITY AND IMPLEMENTATION: The AI-HOPE software and demonstration data is available at https://github.com/Velazquez-Villarreal-Lab/AI-HOPE.

Precision Medicine↗

Implementing a training resource for large-scale genomic data analysis in the All of Us Researcher Workbench.

A lack of representation in genomic research and limited access to computational training create barriers for many researchers seeking to analyze large-scale genetic datasets. The All of Us Research Program provides an unprecedented opportunity to address these gaps by offering genomic data from a broad range of participants, but its impact depends on equipping researchers with the necessary skills to use it effectively. The All of Us Biomedical Researcher (BR) Scholars Program at Baylor College of Medicine aims to break down these barriers by providing early-career researchers with hands-on training in computational genomics through the All of Us Evenings with Genetics Research Program. The year-long program begins with the faculty summit, an in-person computational boot camp that introduces scholars to foundational skills for using the All of Us dataset via a cloud-based research environment. The genomics tutorials focus on genome-wide association studies (GWASs), utilizing Jupyter Notebooks and the Hail computing framework to provide an accessible and scalable approach to large-scale data analysis. Scholars engage in hands-on exercises covering data preparation, quality control, association testing, and result interpretation. By the end of the summit, participants will have successfully conducted a GWAS, visualized key findings, and gained confidence in computational resource management. This initiative expands access to genomic research by equipping early-career researchers from a variety of backgrounds with the tools and knowledge to analyze All of Us data. By lowering barriers to entry and promoting the study of representative populations, the program fosters innovation in precision medicine and advances equity in genomic research.

Humans↗

Hairpins in a Haystack: recognizing microRNA precursors in comparative genomics data.

UNLABELLED: Recently, genome-wide surveys for non-coding RNAs have provided evidence for tens of thousands of previously undescribed evolutionary conserved RNAs with distinctive secondary structures. The annotation of these putative ncRNAs, however, remains a difficult problem. Here we describe an SVM-based approach that, in conjunction with a non-stringent filter for consensus secondary structures, is capable of efficiently recognizing microRNA precursors in multiple sequence alignments. The software was applied to recent genome-wide RNAz surveys of mammals, urochordates, and nematodes. AVAILABILITY: The program RNAmicro is available as source code and can be downloaded from http://www.bioinf.uni-leipzig/Software/RNAmicro.

Algorithms↗

Integrating genomic data to predict transcription factor binding.

Transcription factor binding sites (TFBS) in gene promoter regions are often predicted by using position specific scoring matrices (PSSMs), which summarize sequence patterns of experimentally determined TF binding sites. Although PSSMs are more reliable than simple consensus string matching in predicting a true binding site, they generally result in high numbers of false positive hits. This study attempts to reduce the number of false positive matches and generate new predictions by integrating various types of genomic data by two methods: a Bayesian allocation procedure, and support vector machine classification. Several methods will be explored to strengthen the prediction of a true TFBS in the Saccharomyces cerevisiae genome: binding site degeneracy, binding site conservation, phylogenetic profiling, TF binding site clustering, gene expression profiles, GO functional annotation, and k-mer counts in promoter regions. Binding site degeneracy (or redundancy) refers to the number of times a particular transcription factor's binding motif is discovered in the upstream region of a gene. Phylogenetic conservation takes into account the number of orthologous upstream regions in other genomes that contain a particular binding site. Phylogenetic profiling refers to the presence or absence of a gene across a large set of genomes. Binding site clusters are statistically significant clusters of TF binding sites detected by the algorithm ClusterBuster. Gene expression takes into account the idea that when the gene expression profiles of a transcription factor and a potential target gene are correlated, then it is more likely that the gene is a genuine target. Also, genes with highly correlated expression profiles are often regulated by the same TF(s). The GO annotation data takes advantage of the idea that common transcription targets often have related function. Finally, the distribution of the counts of all k-mers of length 4, 5, and 6 in gene's promoter region were examined as means to predict TF binding. In each case the data are compared to known true positives taken from ChIP-chip data, Transfac, and the Saccharomyces Genome Database. First, degeneracy, conservation, expression, and binding site clusters were examined independently and in combination via Bayesian allocation. Then, binding sites were predicted with a support vector machine (SVM) using all methods alone and in combination. The SVM works best when all genomic data are combined, but can also identify which methods contribute the most to accurate classification. On average, a support vector machine can classify binding sites with high sensitivity and an accuracy of almost 80%.

Algorithms↗

Proteomics-based validation of genomic data: applications in colorectal cancer diagnosis.

Multiple factors are involved in the translation of functional genomic results into proteins for proteome research and target validation on tumoral tissues. In this report, genes were selected by using DNA microarrays on a panel of colorectal cancer (CRC) paired samples. A large number of up-regulated genes in colorectal cancer patients were investigated for cellular location, and those corresponding to membrane or extracellular proteins were used for a non-biased expression in Escherichia coli. We investigated different sources of cDNA clones for protein expression as well as the influence of the protein size and the different tags with respect to protein expression levels and solubility in E. coli. From 29 selected genes, 21 distinct proteins were finally expressed as soluble proteins with, at least, one different fusion protein. In addition, seven of these potential markers (ANXA3, BMP4, LCN2, SPARC, SPP1, MMP7, and MMP11) were tested for antibody production and/or validation. Six of the seven proteins (all except SPP1) were confirmed to be overexpressed in colorectal tumoral tissues by using immunoblotting and tissue microarray analysis. Although none of them could be associated to early stages of the tumor, two of them (LCN2 and MMP11) were clearly overexpressed in late Dukes' stages (B and C). This proteomic study reveals novel clues for the assembly of a robust and highly efficient high throughput system for the validation of genomic data. Moreover it illustrates the different difficulties and bottlenecks encountered for performing a quick conversion of genomic results into clinically useful proteins.

Acute-Phase Proteins↗

Genome data: what do we learn?

Genome sequence information has continued to accumulate at a spectacular pace during the past year. Details of the sequence and gene content of human chromosome 22 were published. The sequencing and annotation of the first two Arabidopsis thaliana chromosomes was completed. The sequence of chromosome 3 from Plasmodium falciparum, the second sequenced malaria chromosome, was reported, as was that of chromosome 1 from Leishmania major. The complete genomic sequences of five microbes were reported. Approaches to using data from completely sequenced microbial genomes in phylogenetic studies are being explored, as is the application of microarrays to whole genome expression analysis.

Animals↗

Patenting nonassociated polymeric structures (NAPS): implications for structural genomic data release.

The intellectual property laws that govern patent rights should provide a reasonable balance between the competing concerns of open access and exclusivity. Open access can facilitate knowledge dissemination and collaboration in furthering science. On the other hand, exclusivity can ensure interest and financial investment in scientific research and development. In recent days, the appropriate balance between open access and exclusivity has been a focus of public debate, particularly with regard to genomic inventions and their applications. In seeking to reconcile the timing of structural genomic data release with certain efforts to secure intellectual property rights, the International Structural Genomics Organisation joins others confronting this controversy. This paper seeks to inform the discussion with an overview of the U.S. standards for patenting nonassociated polymeric structures (NAPS), which include polynucleotides or polypeptides of unknown biological significance, and their corresponding structural data. In the United States, the present ability to obtain patent rights to these discoveries appears problematic given the requirement of specific, substantial and credible utility, among other things. Without demonstrable utility, NAPS and NAPS-related data likely will not be entitled to patent protection, whether the U.S. Patent & Trademark Office rejects NAPS claims as unpatentable in the first instance, or the U.S. federal courts invalidate NAPS claims in later patent litigation. As such, the improbability of obtaining enforceable patent rights to NAPS might undermine the rationale for delaying structural genomic data release to allow for the filing of patent applications in this regard.

Expressed Sequence Tags↗

Reconstruction of metabolic networks from genome data and analysis of their global structure for various organisms.

MOTIVATION: Information from fully sequenced genomes makes it possible to reconstruct strain-specific global metabolic network for structural and functional studies. These networks are often very large and complex. To properly understand and analyze the global properties of metabolic networks, methods for rationally representing and quantitatively analyzing their structure are needed. RESULTS: In this work, the metabolic networks of 80 fully sequenced organisms are in silico reconstructed from genome data and an extensively revised bioreaction database. The networks are represented as directed graphs and analyzed by using the 'breadth first searching algorithm to identify the shortest pathway (path length) between any pair of the metabolites. The average path length of the networks are then calculated and compared for all the organisms. Different from previous studies the connections through current metabolites and cofactors are deleted to make the path length analysis physiologically more meaningful. The distribution of the connection degree of these networks is shown to follow the power law, indicating that the overall structure of all the metabolic networks has the characteristics of a small world network. However, clear differences exist in the network structure of the three domains of organisms. Eukaryotes and archaea have a longer average path length than bacteria. AVAILABILITY: The reaction database in excel format and the programs in VBA (Visual Basic for Applications) are available upon request. SUPPLEMENTARY MATERIAL: Bioinformatics Online.

Archaea↗

Accessing genomic data through XML-based remote procedure calls.

As the amount of data in public genomic databases grows, interoperability among them is becoming an increasingly critical feature. The ability for automated systems to mine and integrate data will be crucial to extracting knowledge from sources of data whose volume far exceeds the capabilities of human researchers. The currently dominant paradigm of presenting information as Web pages and using hyperlinks to describe relationships between pieces of information favors usability, but makes interoperability and automated data exchange more difficult. In this paper we describe how SNPper, a web-based system for the retrieval and analysis of Single Nucleotide Polymorphisms (SNPs), was augmented with a Remote Procedure Call interface, allowing client applications to query our program for SNP data and to receive the response as an XML document. Data represented in this form can be easily parsed by the requesting program, and thus reused for other applications. In this paper we describe the implementation of the interface and we show examples of its usage in a number of existing applications.

Databases, Genetic↗

SynView: a GBrowse-compatible approach to visualizing comparative genome data.

UNLABELLED: We present SynView, a simple and generic approach to dynamically visualize multi-species comparative genome data. It is a light-weight application based on the popular and configurable web-based GBrowse framework. It can be used with a variety of databases and provides the user with a high degree of interactivity. The tool is written in Perl and runs on top of the GBrowse framework. It is in use in the PlasmoDB (http://www.PlasmoDB.org) and the CryptoDB (http://www.CryptoDB.org) projects and can be easily integrated into other cross-species comparative genome projects. AVAILABILITY: The program and instructions are freely available at http://www.ApiDB.org/apps/SynView/ CONTACT: jkissing@uga.edu.

Algorithms↗

Workgroup report: Review of genomics data based on experience with mock submissions--view of the CDER Pharmacology Toxicology Nonclinical Pharmacogenomics Subcommittee.

Over the past few years, both the U.S. Food and Drug Administration (FDA) and the pharmaceutical industry have recognized the potential importance of pharmacogenomics and toxicogenomics to drug development. To resolve the uncertainties surrounding the use of microarray technology and the presentation of genomics data for regulatory purposes, several pharmaceutical companies and genomics technology providers have provided the FDA with reports of genomics studies that included supporting toxicology data (e.g., serum chemistry, histopathology). These studies were not associated with any active drug application and were exploratory or hypothesis generating in nature. For training purposes, these reports were reviewed by the Nonclinical Pharmacogenomics Subcommittee consisting of the Center for Drug Evaluation and Research pharmacology and toxicology researchers and reviewers. In this article, we describe some of these submissions and report on our assessment of data content, format, and quality control metrics that were useful for evaluating these nonclinical genomics submissions, specifically in relation to the proposed MIAME/MINTox (minimum information about a microarray experiment/minimum information needed for a toxicology experiment) recommendations. These genomics submissions allowed both researchers and regulators to gain experience in the process of reviewing and analyzing toxicogenomics data. The experience will allow development of recommendations for the submission and review of these data as the state of the science evolves.

Animals↗

Associative clustering for exploring dependencies between functional genomics data sets.

High-throughput genomic measurements, interpreted as cooccurring data samples from multiple sources, open up a fresh problem for machine learning: What is in common in the different data sets, that is, what kind of statistical dependencies are there between the paired samples from the different sets? We introduce a clustering algorithm for exploring the dependencies. Samples within each data set are grouped such that the dependencies between groups of different sets capture as much of pairwise dependencies between the samples as possible. We formalize this problem in a novel probabilistic way, as optimization of a Bayes factor. The method is applied to reveal commonalities and exceptions in gene expression between organisms and to suggest regulatory interactions in the form of dependencies between gene expression profiles and regulator binding patterns.

Algorithms↗

Hidden likelihood support in genomic data: can forty-five wrongs make a right?

Combined analysis of multiple phylogenetic data sets can reveal emergent character support that is not evident in separate analyses of individual data sets. Previous parsimony analyses have shown that this hidden support often accounts for a large percentage of the overall phylogenetic signal in cladistic studies. Here, reanalysis of a large comparative genomic data set for yeast (genus Saccharomyces) demonstrates that hidden support can be an important factor in maximum likelihood analyses of multiple data sets as well. Emergent signal in a concatenation of 106 genes was responsible for up to 64% of the likelihood support at a particular node (the difference in log likelihood scores between optimal topologies that included and excluded a supported clade). A grouping of four yeast species (S. cerevisiae, S. paradoxus, S. mikatae, and S. kudriavzevii) was robustly supported by combined analysis of all 106 genes, but separate analyses of individual genes suggested numerous conflicts. Forty-eight genes strictly contradicted S. cerevisiae + S. paradoxus + S. mikatae + S. kudriavzevii in separate analyses, but combined likelihood analyses that included up to 45 of the "wrong" data sets supported this group. Extensive hidden support also emerged in a combined likelihood analysis of 41 genes that each recovered the exact same topology in separate analyses of the individual genes. These results show that isolated analyses of individual data sets can mask congruence and distort interpretations of clade stability, even in strictly model-based phylogenetic methods. Consensus and supertree procedures that ignore hidden phylogenetic signals are, at best, incomplete.

Classification↗

Benchmarking ortholog identification methods using functional genomics data.

BACKGROUND: The transfer of functional annotations from model organism proteins to human proteins is one of the main applications of comparative genomics. Various methods are used to analyze cross-species orthologous relationships according to an operational definition of orthology. Often the definition of orthology is incorrectly interpreted as a prediction of proteins that are functionally equivalent across species, while in fact it only defines the existence of a common ancestor for a gene in different species. However, it has been demonstrated that orthologs often reveal significant functional similarity. Therefore, the quality of the orthology prediction is an important factor in the transfer of functional annotations (and other related information). To identify protein pairs with the highest possible functional similarity, it is important to qualify ortholog identification methods. RESULTS: To measure the similarity in function of proteins from different species we used functional genomics data, such as expression data and protein interaction data. We tested several of the most popular ortholog identification methods. In general, we observed a sensitivity/selectivity trade-off: the functional similarity scores per orthologous pair of sequences become higher when the number of proteins included in the ortholog groups decreases. CONCLUSION: By combining the sensitivity and the selectivity into an overall score, we show that the InParanoid program is the best ortholog identification method in terms of identifying functionally equivalent proteins.

Algorithms↗

Toward metabolic phenomics: analysis of genomic data using flux balances.

Small genome sequencing and annotations are leading to the definition of metabolic genotypes in an increasing number of organisms. Proteomics is beginning to give insights into the use of the metabolic genotype under given growth conditions. These data sets give the basis for systemically studying the genotype-phenotype relationship. Methods of systems science need to be employed to analyze, interpret, and predict this complex relationship. These endeavors will lead to the development of a new field, tentatively named phenomics. This article illustrates how the metabolic characteristics of annotated small genomes can be analyzed using flux balance analysis (FBA). A general algorithm for the formulation of in silico metabolic genotypes is described. Illustrative analyses of the in silico Escherichia coli K-12 metabolic genotypes are used to show how FBA can be used to study the capabilities of this strain.

Biotechnology↗