PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data accessibility”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

A task framework for the web interface W2H.

SUMMARY: The W3H task framework allows the execution of compound jobs utilizing the description of work and data flows in a heterogeneous bioinformatics environment using meta-data information. By means of these descriptions, the task system can schedule the necessary execution of applications available in the environment, depending on rules specified in the meta-data. By integrating this task framework into the publicly available web interface W2H, similarly based on meta-data, web access and data management are immediately available for each task description. Authors of task descriptions can base their work on the underlying classes and objects to be able to describe dependency rules between previously independent applications. The result of a compound task is given as XML data that is translated according to XSLT data into web pages or plain text to report the result of the task to the user. AVAILABILITY: Within the HUSAR environment at DKFZ http://genome.dkfz-heidelberg.de/

Database Management Systems↗

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata↗

ELViS: an R package for estimating copy number levels of viral genomic segments at base-resolution.

MOTIVATION: Tumor viruses account for ∼10% of cancer diagnoses. Virally induced tumorigenesis is understood as direct signaling through oncogenes such as E6 and E7 genes in the case of human papillomavirus. Furthermore, pathogen characteristics such as viral oncogene dose may impact the disease course. To our knowledge, no tool has been proposed to assess the intra-viral copy number alterations that define the gene dose of viral oncogenes and associated suppressive pathways native to the pathogen's normal life cycle. RESULTS: We propose an R package, "ELViS," that analyzes viral copy number changes from DNA sequencing of whole viral genomes. The method adjusts for viral load with 2D transformation and segmentation to offer the relative viral gene doses. AVAILABILITY AND IMPLEMENTATION: The ELViS R package is available from https://bioconductor.org/packages/ELViS. This article used controlled access data from dbGaP (phs001713.v1.p1).

Software↗

GandrKB--ontological microarray annotation and visualization.

SUMMARY: The Gandr (gene annotation data representation) knowledgebase is an ontological framework for laboratory-specific gene annotation. Gandr uses Protege 2000 for editing, querying and visualizing microarray data and annotations. Genes can be annotated with provided, newly created or imported ontological concepts. Annotated genes can inherit assigned concept properties and can be related to each other. The resulting knowledgebase can be visualized as interactive network of nodes and edges representing genes and their functional relationships. This allows for immediate and associative gene context exploration. Ontological query techniques allow for powerful data access.

Algorithms↗

Age differences in decision making: a process methodology for examining strategic information processing.

This study explored the use of process tracing techniques in examining the decision-making processes of older and younger adults. Thirty-six college-age and thirty-six retirement-age participants decided which one of six cars they would purchase on the basis of computer-accessed data. They provided information search protocols. Results indicate that total time to reach a decision did not differ according to age. However, retirement-age participants used less information, spent more time viewing, and re-viewed fewer bits of information than college-age participants. Information search patterns differed markedly between age groups. Patterns of retirement-age adults indicated their use of noncompensatory decision rules which, according to decision-making literature (Payne, 1976), reduce cognitive processing demands. The patterns of the college-age adults indicated their use of compensatory decision rules, which have higher processing demands.

Adolescent↗

MedImg: An Integrated Database for Public Medical Images.

The advancements in deep learning algorithms for medical image analysis have garnered significant attention in recent years. While several studies have shown promising results, with models achieving or even surpassing human performance, translating these advancements into clinical practice is still accompanied by various challenges. A primary obstacle lies in the availability of large-scale, well-characterized datasets for validating the generalization of approaches. To address this challenge, we curated a diverse collection of medical image datasets from multiple public sources, containing 105 datasets and a total of 1,995,671 images. These images span 14 modalities, including X-ray, computed tomography, magnetic resonance imaging, optical coherence tomography, ultrasound, and endoscopy, and originate from 13 organs, such as the lung, brain, eye, and heart. Subsequently, we constructed an online database, MedImg, which incorporates and systematically organizes these medical images to facilitate data accessibility. MedImg serves as an intuitive and open-access platform for facilitating research in deep learning-based medical image analysis, accessible at https://www.cuilab.cn/medimg/.

Humans↗

The natural history of prostatism: the effects of non-response bias.

BACKGROUND: In epidemiological studies, non-response may raise the question of generalizability to the target population. Most investigations have not been able to access data that could provide information about the potential impact of non-response bias. METHODS: A 55% response rate was realized at baseline for a prospective cohort investigation of the natural history of benign prostatic hyperplasia in Olmsted County, Minnesota, during 1989-1991 (the Olmsted County Study of Urinary Symptoms and Health Status Among Men). This prompted a preliminary study of potential non-response bias among full participants, partial participants and complete non-responders. The medical diagnostic index maintained by the Rochester Epidemiology Project was used to ascertain the prevalence of specific conditions in the 9 years prior to study inception. RESULTS: The age-adjusted period prevalence rate for benign prostatic hyperplasia (%) was 9.6 (95% confidence interval [CI]: 8.1-11.0) for full participants, 8.2 (95% CI: 5.8-10.6) for partial participants and 5.3 (95% CI: 3.6-6.9) for complete non-responders. Other urologic diagnoses followed the same pattern. However, age-adjusted prevalence rates for general medical examination history and major non-urologic morbidities were decidedly similar across response groups. CONCLUSIONS: These data suggest response may have been driven, in part, by concerns about urologic disease. However, the similarity in non-urologic diagnoses and general medical examinations provide some preliminary reassurance that the 55% response rate did not necessarily compromise generalizability.

Adult↗

Object oriented Transcription Factors Database (ooTFD).

ooTFD is an object-oriented database for the representation of information pertaining to transcription factors, the proteins and biochemical entities which play a central role in the regulation of gene expression. Given the recent explosion of genome sequence information, and that a large percentage of proteins encoded by fully sequenced genomes fall into this category, information pertaining to this class of molecules may become an essential aspect of biology and of genomics in the 21st century. In the past year, there was a small increase in the size of this database, and a number of new tools to facilitate data access and analysis have been added at the MIRAGE (Molecular Informatics Resource for the Analysis of Gene Expression) web site. ooTFD and associated tools and resources can be accessed at http://www.ifti.org/

Databases, Factual↗

Recent additions and improvements to the Onto-Tools.

The Onto-Tools suite is composed of an annotation database and six seamlessly integrated, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner and Pathway-Express. The Onto-Tools database has been expanded to include various types of data from 12 new databases. Our database now integrates different types of genomic data from 19 sequence, gene, protein and annotation databases. Additionally, our database is also expanded to include complete Gene Ontology (GO) annotations. Using the enhanced database and GO annotations, Onto-Express now allows functional profiling for 24 organisms and supports 17 different types of input IDs. Onto-Translate is also enhanced to fully utilize the capabilities of the new Onto-Tools database with an ultimate goal of providing the users with a non-redundant and complete mapping from any type of identification system to any other type. Currently, Onto-Translate allows arbitrary mappings between 29 types of IDs. Pathway-Express is a new tool that helps the users find the most interesting pathways for their input list of genes. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Animals↗

New Onto-Tools: Promoter-Express, nsSNPCounter and Onto-Translate.

The Onto-Tools suite is composed of an annotation database and eight complementary, web-accessible data mining tools: Onto-Express, Onto-Compare, Onto-Design, Onto-Translate, Onto-Miner, Pathway-Express, Promoter-Express and nsSNPCounter. Promoter-Express is a new tool added to the Onto-Tools ensemble that facilitates the identification of transcription factor binding sites active in specific conditions. nsSNPCounter is another new tool that allows computation and analysis of synonymous and non-synonymous codon substitutions for studying evolutionary rates of protein coding genes. Onto-Translate has also been enhanced to expand its scope and accuracy by fully utilizing the capabilities of the Onto-Tools database. Currently, Onto-Translate allows arbitrary mappings between 28 types of IDs for 53 organisms. Onto-Tools are freely available at http://vortex.cs.wayne.edu/Projects.html.

Binding Sites↗

ForestTreeDB: a database dedicated to the mining of tree transcriptomes.

ForestTreeDB is intended as a resource that centralizes large-scale expressed sequence tag (EST) sequencing results from several tree species (http://foresttree.org/ftdb). It currently encompasses 344,878 quality sequences from 68 libraries, from diverse organs of conifer and hybrid poplar trees. It utilizes the Nimbus data model to provide a hosting system for multiple projects, and uses object-relational mapping APIs in Java and Perl for data accesses within an Oracle database designed to be scalable, maintainable and extendable. Transcriptome builds or unigene sets occupy the focal point of the system. Several of the five current species-specific unigenes were used to design microarrays and SNP resources. The ForestTreeDB web application provides the means for multiple combination database queries. It presents the user with a list of discrete queries to retrieve and download large EST datasets or sequences from precompiled unigene assemblies. Functional annotation assignment is not trivial in conifers which are distantly related to angiosperm model plants. Optimal annotations are achieved through database queries that integrate results from several procedures based open-source tools. ForestTreeDB aims to facilitate sequence mining of coherent annotations in multiple species to support comparative genomic approaches. We plan to continuously enrich ForestTreeDB with other resources through collaborations with other genomic projects.

Databases, Nucleic Acid↗

VectorBase: a home for invertebrate vectors of human pathogens.

VectorBase (http://www.vectorbase.org/) is a web-accessible data repository for information about invertebrate vectors of human pathogens. VectorBase annotates and maintains vector genomes providing an integrated resource for the research community. Currently, VectorBase contains genome information for two organisms: Anopheles gambiae, a vector for the Plasmodium protozoan agent causing malaria, and Aedes aegypti, a vector for the flaviviral agents causing Yellow fever and Dengue fever.

Aedes↗

Current problems that are likely to affect the future of epidemiology.

The current discussion focuses on criticism as a positive force for improving epidemiologic practice through periodic reexamination of the basic approach to the discipline and the strategy for meeting the future educational needs of students and practicing epidemiologists. The types of epidemiologic research conducted and the settings within which the research will be conducted are also discussed. Epidemiology can be expected to play a major role in new areas of research that are created by changes in the medical care system and the development of large data systems associated with these approaches to health care delivery. This paper also discusses the growing threat to data access, the problems of communicating epidemiologic research findings to the public through the media, and the expanding interface between epidemiologic research and the legal system. The role of epidemiologic organizations in helping to shape the discipline's response to these issues and the opportunities these issues or problems present for improving epidemiologic research are also discussed.

Databases, Factual↗

Macromolecular query language (MMQL): prototype data model and implementation.

Macromolecular query language (MMQL) is an extensible interpretive language in which to pose questions concerning the experimental or derived features of the 3-D structure of biological macromolecules. MMQL portends to be intuitive with a simple syntax, so that from a user's perspective complex queries are easily written. A number of basic queries and a more complex query--determination of structures containing a five-strand Greek key motif--are presented to illustrate the strengths and weaknesses of the language. The predominant features of MMQL are a filter and pattern grammar which are combined to express a wide range of interesting biological queries. Filters permit the selection of object attributes, for example, compound name and resolution, whereas the patterns currently implemented query primary sequence, close contacts, hydrogen bonding, secondary structure, conformation and amino acid properties (volume, polarity, isoelectric point, hydrophobicity and different forms of exposure). MMQL queries are processed by MMQLlib; a C++ class library, to which new query methods and pattern types are easily added. The prototype implementation described uses PDBlib, another C(++)-based class library from representing the features of biological macromolecules at the level of detail parsable from a PDB file. Since PDBlib can represent data stored in relational and object-oriented databases, as well as PDB files, once these data are loaded they too can be queried by MMQL. Performance metrics are given for queries of PDB files for which all derived data are calculated at run time and compared to a preliminary version of OOPDB, a prototype object-oriented database with a schema based on a persistent version of PDBlib which offers more efficient data access and the potential to maintain derived information. MMQLlib, PDBlib and associated software are available via anonymous ftp from cuhhca.hhmi.columbia.edu.

Amino Acid Sequence↗

Census of genes expressed in porcine embryos and reproductive tissues by mining an expressed sequence tag database based on human genes.

A total of 98,898 expressed sequence tags (ESTs) derived from embryos and reproductive tissues in pigs were identified in the GenBank "est_others" database. Pig embryos were collected at 11, 12, 13, 14, 15, 20, 30, and 45 days after gestation. The reproductive tissues were sampled from testis, ovary, endometrium, hypothalamus, anterior pituitary, uterus, and placenta. A gene-oriented approach was developed to annotate these porcine EST sequences to census the genes expressed from these sources. Of the 33 308 mRNA sequences from the human genes used as references (data accessed on 1 November 2002), 9410 had the porcine EST homologs expressed in embryos and 11 795 had the EST homologs expressed in reproductive tissues. The entire genome contributes at least 28.3% of its genes to embryo development and 35.4% of its genes to reproduction. Using the EST entry numbers as indicators of gene expression, we determined that the gene expression patterns differ significantly between embryos and reproductive tissues in pigs. The basic active genes were identified for each source, but most of them are not coexpressed abundantly. Few genes were expressed on the Y chromosome (P < 0.01), but they may represent counterparts of the double-dose genes that remain active in an inactivated X chromosome in females but are needed for proper development and growth. The census provides a panel of transcripts in a broad sense that can be used as targets to study the mechanisms involved in embryo development and reproduction in pigs and other mammals, including humans.

Animals↗

Homicide offenders in Auckland, New Zealand.

In the 14-year period from 1976 to 1989, 183 homicide offenders have been recorded in Auckland, New Zealand. Data accessed from police files show that Maori and Polynesians made up the majority of the offenders. Almost all offenders were male, and the largest proportion was between 20 and 24 years of age. Racial differences were noted in the methods used to commit the homicide, in the offender-victim relationship, and in whether the offender acted alone. The majority of offenders were not convicted of murder, but were convicted on lesser charges.

Adolescent↗

Introducing information technologies into medical education: activities of the AAMC.

Previous articles in this column have discussed how new information technologies are revolutionizing medical education. In this article, two staff members from the Association of American Medical College's Division of Medical Education discuss how the Association (the AAMC) is working both to support the introduction of new technologies into medical education and to facilitate dialogue on information technology and curriculum issues among AAMC constituents and staff. The authors describe six AAMC initiatives related to computing in medical education: the Medical School Objectives Project, the National Curriculum Database Project, the Information Technology and Medical Education Project, a professional development program for chief information officers, the AAMC ACCESS Data Collection and Dissemination System, and the internal Staff Interest Group on Medical Informatics and Medical Education.

Curriculum↗

The air-kerma rate constant of 192Ir.

The air-kerma rate constant gamma delta (and its precursors), as one of the basic radiation characteristics of 192Ir, was determined by many authors. Analysis of accessible data on this quantity led us to the conclusion that published data strongly disagree. That is the reason we calculated this quantity on the basis of our and many other authors' gamma-ray spectral data and the latest data for mass energy-transfer coefficients for air. In this way, a value was obtained for gamma delta of 30.0 +/- 0.9 a Gy m2 s-1 Bq-1 for an unshielded 192Ir source and 27.8 +/- 0.9 a Gy m2s -1Bq-1 for a standard packaged radioactive source taking into account attenuation of gamma rays in the platinum source wall.

Iridium Radioisotopes↗