PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “genomic data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

ChemGenXplore: an interactive tool for exploring and analysing chemical genomic data.

MOTIVATION: Chemical genomics is a powerful high-throughput approach to systematically link phenotypes to genotypes. However, the vast datasets generated remain challenging to explore due to the lack of integrated, interactive tools for visualization and analysis. Existing workflows often require multiple independent software tools, limiting data accessibility and collaboration. Therefore, we created a user-friendly platform that enables efficient exploration and sharing of chemical genomics data. RESULTS: We developed ChemGenXplore, a web-based Shiny application designed to streamline the visualization and analysis of chemical genomic screens. It offers two primary functionalities: one for exploring pre-implemented datasets and another for analysing user-uploaded datasets. ChemGenXplore enables users to visualize phenotypic profiles, assess gene-gene and condition-condition correlations, perform GO and KEGG enrichment analysis, and generate customizable, interactive heatmaps. To further support collaborative research, ChemGenXplore also facilitates the comparative analysis of chemical genomic and other omics datasets. By consolidating these features into a single interactive and accessible tool, ChemGenXplore facilitates data sharing, enhances reproducibility, and promotes collaboration within the research community. AVAILABILITY AND IMPLEMENTATION: ChemGenXplore is freely accessible as a web application at https://chemgenxplore.kaust.edu.sa/. Source code and documentation, including instructions for local installation, are provided on GitHub (https://github.com/Hudaahmadd/ChemGenXplore). A Docker image is also available on DockerHub (https://hub.docker.com/r/hudaahmad/chemgenxplore) to ensure reproducibility and simplify installation.

Software↗

AWGE-ESPCA: An edge sparse PCA model based on adaptive noise elimination regularization and weighted gene network for Hermetia illucens genomic data analysis.

Hermetia illucens is an important insect resource. Studies have shown that exploring the effects of Cu2+-stressed on the growth and development of the Hermetia illucens genome holds significant scientific importance. There are three major challenges in the current studies of Hermetia illucens genomic data analysis: firstly, the lack of available genomic data which limits researchers in Hermetia illucens genomic data analysis. Secondly, to the best of our knowledge, there are no Artificial Intelligence (AI) feature selection models designed specifically for Hermetia illucens genome. Unlike human genomic data, noise in Hermetia illucens data is a more serious problem. Third, how to choose those genes located in the pathway enrichment region. Existing models assume that each gene probe has the same priori weight. However, researchers usually pay more attention to gene probes which are in the pathway enrichment region. Based on the above challenges, we initially construct experiments and establish a new Cu2+-stressed Hermetia illucens growth genome dataset. Subsequently, we propose AWGE-ESPCA: an edge Sparse PCA model based on adaptive noise elimination regularization and weighted gene network. The AWGE-ESPCA model innovatively proposes an adaptive noise elimination regularization method, effectively addressing the noise challenge in Hermetia illucens genomic data. We also integrate the known gene-pathway quantitative information into the Sparse PCA(SPCA) framework as a priori knowledge, which allows the model to filter out the gene probes in pathway-rich regions as much as possible. Ultimately, this study conducts five independent experiments and compared four latest Sparse PCA models as well as representative supervised and unsupervised baseline models to validate the model performance. The experimental results demonstrate the superior pathway and gene selection capabilities of the AWGE-ESPCA model. Ablation experiments validate the role of the adaptive regularizer and network weighting module. To summarize, this paper presents an innovative unsupervised model for Hermetia illucens genome analysis, which can effectively help researchers identify potential biomarkers. In addition, we also provide a working AWGE - ESPCA model code in the address: https://github.com/yhyresearcher/AWGE_ESPCA.

Animals↗

Environmentally responsible human genomic data governance: points for consideration.

We introduce five points for integrating environmental ethics into human genomic data governance: (i) recognizing the ethical imperative to consider environmental impacts of human genomic data; (ii) fostering collective responsibility for environmental harms; (iii) prospectively assessing benefits and harms; (iv) anticipating barriers to integration of environmental ethics into genomic data governance; and (v) meaningfully engaging all interest-holders. These points will be useful to all involved in the genomic data ecosystem.

Letter↗

A Sociotechnical Approach to Genomic Data Privacy: A Comparative Analysis.

The sharing of genomic data across international borders presents significant privacy law challenges.Secured computed environments on smartphones allow the storing and processing of sensitive data without the underlying data being shared with processors.A novel technology, described here, to process genomic data within a secured computing environment seems to comport with EU and US privacy laws, despite their differing aims and rules.This technology suggests there may be technological solutions to privacy law fragmentation across jurisdictions, so long as data subjects socially trust the technology and have control over their data.

genome↗

Embracing the complexity of genomic data for personalized medicine.

Numerous recent studies have demonstrated the use of genomic data, particularly gene expression signatures, as clinical prognostic factors in cancer and other complex diseases. Such studies herald the future of genomic medicine and the opportunity for personalized prognosis in a variety of clinical contexts that utilizes genome-scale molecular information. The scale, complexity, and information content of high-throughput gene expression data, as one example of complex genomic information, is often under-appreciated as many analyses continue to focus on defining individual rather than multiplex biomarkers for patient stratification. Indeed, this complexity of genomic data is often--rather paradoxically--viewed as a barrier to its utility. To the contrary, the complexity and scale of global genomic data, as representing the many dimensions of biology, must be embraced for the development of more precise clinical prognostics. The need is for integrated analyses--approaches that embrace the complexity of genomic data, including multiple forms of genomic data, and aim to explore and understand multiple, interacting, and potentially conflicting predictors of risk, rather than continuing on the current and traditional path that oversimplifies and ignores the information content in the complexity. All forms of potentially relevant data should be examined, with particular emphasis on understanding the interactions, complementarities, and possible conflicts among gene expression, genetic, and clinical markers of risk.

Breast Neoplasms↗

Inference of Gene Flow between Species from Genomic Data When the Mode, Direction, and Lineages are Misspecified.

Thanks to genomic data, interspecific gene flow is increasingly recognized as a major evolutionary force that shapes biodiversity. Two models have been developed in the multispecies coalescent (MSC) framework to infer gene flow from genomic data, assuming either constant-rate continuous migration (MSC-M) or discrete introgression/hybridization (MSC-I). The extreme simplicity of these models raises concerns about their usefulness as they represent misspecified models when applied to real data. Here, we study inference of gene flow under the MSC-M model, considering mis-assignment of gene flow onto incorrect parental or daughter lineages, misspecification of the direction of gene flow, and misspecification of the mode of gene flow. Mis-assignment of gene flow to an incorrect lineage causes large biases in the estimated rates. The Bayesian test has high power for inferring both recent and ancient gene flow, between either sister lineages or nonsister lineages, although misspecification of the direction of gene flow may make it hard to distinguish early divergence with gene flow from recent complete isolation. Misspecification of the mode of gene flow (MSC-I versus MSC-M) has small local effects, and gene flow is detected with high power despite the misspecification. We analyze a genomic dataset from the purple cone spruce (Picea spp., Pinaceae), which putatively arose through homoploid hybrid speciation, to demonstrate practical implications of our theoretical analyses. Overall, we find that the extremely idealized models of gene flow (in particular the discrete MSC-I model) are very effective for extracting information about species divergence and gene flow from genomic data.

Gene Flow↗

Integration of genomic data in Electronic Health Records--opportunities and dilemmas.

OBJECTIVES: In this paper we give an overview about the challenge the postgenomic era poses on biomedical informaticists. The occurrence of new (genomic) data types necessitates new data models, new viewing metaphors and methods to deal with the disclosure of genomic data. We discuss integration issues when inferring phenotype and genotype data. Another challenge is to find the right phenotype to genotype data in order to get appropriate case numbers for sound clinical genotype-phenotype inference studies. METHODS: Genomic data could be integrated in an Electronic Health Record (EHR) in several ways. We describe patient-centered and pointer-based integration strategies and the corresponding data types and data models. The inference mechanisms for the interpretation of row data contain different agents. We describe vertical, horizontal and temporal agents. RESULTS: We have to deal with several new data types, not being standardized for EHR integration. Genomic data tends to be more structured than phenotype data. Beyond the development of new data models, vertical, horizontal and temporal agents have to be developed in order to link genotype and phenotype. As the genomic EHR will contain very sensitive data, confidentiality and privacy concerns have to be addressed. CONCLUSIONS: Given the necessity to capture both environment and genomic state of a patient and their interaction, clinical information systems have to be redesigned. While genotyping seems to be automatable easily, this is not the case for clinical information. More integration work on terminologies and ontologies has to be done.

Computational Biology↗

Integration of tools and resources for display and analysis of genomic data for protozoan parasites.

Centralisation of tools for analysis of genomic data is paramount in ensuring that research is always carried out on the latest currently available data. As such, World Wide Web sites providing a range of online analyses and displays of data can play a crucial role in guaranteeing consistency of in silico work. In this respect, the protozoan parasite research community is served by several resources, either focussing on data and tools for one species or taking a broader view and providing tools for analysis of data from many species, thereby facilitating comparative studies. In this paper, we give a broad overview of the online resources available. We then focus on the GeneDB project, detailing the features and tools currently available through it. Finally, we discuss data curation and its importance in keeping genomic data 'relevant' to the research community.

Animals↗

The GDB Human Genome Data Base anno 1994.

In 1991 the Genome Data Base at Johns Hopkins University School of Medicine was selected as the central repository for mapping data from the Human Genome Project, and was funded by NIH and DOE under a three year award. GDB has now finished 28 months of Federally funded operation. During this period a great deal of progress and many internal changes have taken place. In addition, many changes have also occurred in the external environment, and GDB has adapted its strategies to play an appropriate role in those changes as well. Recognizing the central role of mapping information in the genome project, it is important that GDB respond aggressively to the increasing demands of genomic researchers, as well as formulate a program of response to a number of long standing, but still unmet, needs of that community. It is even more important that GDB provide leadership in the genome informatics enterprise. Three themes described here are dominant in our future plans and represent the essence of the major changes made in the past year. They include: enhanced data acquisition, better map representation, and full integration into the collection of genomic databases.

Computer Communication Networks↗

Assessing bioequivalence using genomic data.

For approval of a generic drug product, the assessment of bioequivalence in drug absorption is usually considered as a surrogate for evaluation of drug efficacy and safety in clinical studies. For some drug products, the United States Food and Drug Administration indicates that the assessment of similarity between dissolution profiles may be used as a surrogate for assessment of bioequivalence. Along this line, we propose assessing bioequivalence using genomic data collected from the same individuals, assuming that there is an established relationship between pharmacokinetic and genomic data. Because there may be a bias in the prediction of pharmacokinetic data using genomic data and the variations in these two types of data are different, we propose to assess bioequivalence based on sensitivity analysis of prediction bias and variation difference within some predetermined limits. Our methods are derived for average, population, and individual bioequivalence.

Acyclovir↗

Computational cluster validation in post-genomic data analysis.

MOTIVATION: The discovery of novel biological knowledge from the ab initio analysis of post-genomic data relies upon the use of unsupervised processing methods, in particular clustering techniques. Much recent research in bioinformatics has therefore been focused on the transfer of clustering methods introduced in other scientific fields and on the development of novel algorithms specifically designed to tackle the challenges posed by post-genomic data. The partitions returned by a clustering algorithm are commonly validated using visual inspection and concordance with prior biological knowledge--whether the clusters actually correspond to the real structure in the data is somewhat less frequently considered. Suitable computational cluster validation techniques are available in the general data-mining literature, but have been given only a fraction of the same attention in bioinformatics. RESULTS: This review paper aims to familiarize the reader with the battery of techniques available for the validation of clustering results, with a particular focus on their application to post-genomic data analysis. Synthetic and real biological datasets are used to demonstrate the benefits, and also some of the perils, of analytical clustervalidation. AVAILABILITY: The software used in the experiments is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/. SUPPLEMENTARY INFORMATION: Enlarged colour plots are provided in the Supplementary Material, which is available at http://dbkweb.ch.umist.ac.uk/handl/clustervalidation/.

Algorithms↗

The chimeric mapping problem: algorithmic strategies and performance evaluation on synthetic genomic data.

The Human Genome Project requires better software for the creation of physical maps of chromosomes. Current mapping techniques involve breaking large segments of DNA into smaller, more-manageable pieces, gathering information on all the small pieces, and then constructing a map of the original large piece from the information about the small pieces. Unfortunately, in the process of breaking up the DNA some information is lost and noise of various types is introduced; in particular, the order of the pieces is not preserved. Thus, the map maker must solve a combinatorial problem in order to reconstruct the map. Good software is indispensable for quick, accurate reconstruction. The reconstruction is complicated by various experimental errors. A major source of difficulty--which seems to be inherent to the recombination technology--is the presence of chimeric DNA clones. It is fairly common for two disjoint DNA pieces to form a chimera, i.e., a fusion of two pieces which appears as a single piece. Attempts to order chimera will fail unless they are algorithmically divided into their constituent pieces. Despite consensus within the genomic mapping community of the critical importance of correcting chimerism, algorithms for solving the chimeric clone problem have received only passing attention in the literature. Based on a model proposed by Lander (1992a, b) this paper presents the first algorithms for analyzing chimerism. We construct physical maps in the presence of chimerism by creating optimization functions which have minimizations which correlate with map quality. Despite the fact that these optimization functions are invariably NP-complete our algorithms are guaranteed to produce solutions which are close to the optimum. The practical import of using these algorithms depends on the strength of the correlation of the function to the map quality as well as on the accuracy of the approximations. We employ two fundamentally different optimization functions as a means of avoiding biases likely to decorrelate the solutions from the desired map. Experiments on simulated data show that both our algorithm which minimizes the number of chimeric fragments in a solution and our algorithm which minimizes the maximum number of fragments per clone in a solution do, in fact, correlate to high quality solutions. Furthermore, tests on simulated data using parameters set to mimic real experiments show that that the algorithms have the potential to find high quality solutions with real data. We plan to test our software against real data from the Whitehead Institute and from Los Alamos Genomic Research Center in the near future.

Algorithms↗

How (not) to protect genomic data privacy in a distributed network: using trail re-identification to evaluate and design anonymity protection systems.

The increasing integration of patient-specific genomic data into clinical practice and research raises serious privacy concerns. Various systems have been proposed that protect privacy by removing or encrypting explicitly identifying information, such as name or social security number, into pseudonyms. Though these systems claim to protect identity from being disclosed, they lack formal proofs. In this paper, we study the erosion of privacy when genomic data, either pseudonymous or data believed to be anonymous, are released into a distributed healthcare environment. Several algorithms are introduced, collectively called RE-Identification of Data In Trails (REIDIT), which link genomic data to named individuals in publicly available records by leveraging unique features in patient-location visit patterns. Algorithmic proofs of re-identification are developed and we demonstrate, with experiments on real-world data, that susceptibility to re-identification is neither trivial nor the result of bizarre isolated occurrences. We propose that such techniques can be applied as system tests of privacy protection capabilities.

Algorithms↗

Improvements to the GDB Human Genome Data Base.

Version 6.0 of the Human Genome Data Base introduces a number of significant improvements over previous releases of GDB. The most important of these are revised data representations for genes and genomic maps and a new curatorial model for the database. GDB 6.0 is the first major genomic database to provide read/write access directly to the scientific community, including capabilities for third-party annotation. The revised database can represent all major categories of genetic and physical maps, along with the underlying order and distance information used to construct them. The improved representation permits more sophisticated map queries to be posed and supports the graphical display of maps. In addition the new GDB has a richer model for gene information, better suited for supporting cross-references to databases describing gene function, structure, products, expression and associated phenotypes.

Animals↗

Restructuring the genome data base: a model for a federation of biological databases.

The creation of a federation of public biological databases has been proposed. Formerly independent systems will need to be modified to interoperate better within this federation. This will enable the federated system to provide biologists with an integrated view of biological data. The GDB Human Genome Data Base is being restructured to participate in the proposed federation. GDB itself will be organized into a collection of related data sets in support of human gene mapping. The techniques that will be used to link these data sets will be applicable to the federation as a whole. Links will be based on stable accession numbers that have no inherent information content and are guaranteed always to be recognized. Improvements will be made in the links between GDB and the nucleotide sequence databases to test this approach further.

Amino Acid Sequence↗

Interordinal relationships and timescale of eutherian evolution as inferred from mitochondrial genome data.

Extensive phylogenetic analyses of the updated sequence data of mammalian mitochondrial genomes were carried out using the maximum likelihood method in order to resolve deep branchings in eutherian evolution. The divergence times in the mammalian tree were estimated by a relaxed molecular clock of the mitochondrial proteins calibrated with multiple references. A Chiroptera/Eulipotyphla (i.e. bat/mole) clade and a close relationship of this clade to Fereuungulata (Carnivora+Perissodactyla+Cetartiodactyla) were reconfirmed with high statistical significance. However, a support for a monophyly of Fereuungulata relative to the Chiroptera/Eulipotyphla clade was fragile, and we suggest that the three branchings among Carnivora, Perissodactyla, Cetartiodactyla and Chiroptera/Eulipotyphla occurred successively in a short time period, estimated to be approximately 77Myr BP. The Chiroptera/Eulipotyphla divergence was estimated to roughly coincide with the Cretaceous-Tertiary boundary (65Myr BP). The monophyly of Rodentia, the Lagomorpha/Rodentia clade (traditionally called Glires), and the Afrotheria/Xenarthra clade were preferred over alternative relationships, but the supports of these clades were not strong enough to exclude other possibilities. Although several super-order taxa of eutherians were strongly supported by the analyses of the mitochondrial genome data, the branching order in the deepest part of the eutherian tree remained ambiguous from the data presently available.

Animals↗

The recognition of protein structure and function from sequence: adding value to genome data.

The explosion of DNA sequence data from genome projects presents many challenges. For instance, we must extend our current knowledge of protein structure and function so that it can be applied to these new sequences. The derivation of rules for the relationships between sequence and structure allow us to recognize a common fold by the use of tertiary templates. New techniques enable us to begin to meet the challenge of rule-based modelling of distantly related proteins. This paper describes an integrated and knowledge-based approach to the prediction of protein structure and function which can maximize the value of sequence information.

Amino Acid Sequence↗