PubMed HealthSearch

SEARCH · PubMed Health

Results for “Information Storage and Retrieval”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Analysis of human factors in aircraft accidents.

This paper describes our current approach and accomplishments in the analysis of human factors aspects of aircraft accidents. Emphasis has been placed upon methods of analysis of Boards of Inquiry and human factors information storage and retrieval methods.

Accidents, Aviation

The tachistoscopic recognition of letters under whole and partial report procedures as related to intelligence.

Investigations of the short-term memory task performance of retarded individuals have indicated that these individuals demonstrate a deficit in the mechanisms necessary for the acquisition, storage and/or retrieval of information. The present study examined the tachistoscopic letter recognition task performance of retarded and non-retarded individuals under a partial report and a whole report procedure. The results revealed that the retarded subjects did significantly more poorly relative to the non-retarded subjects under both procedures. The data were interpreted as indicating that the retarded subjects were inefficient in their strategy to make the simultaneous imput task manageable. Further, the data provided no support for the suggestion that a visual-to-auditory encoding process exists between iconic and short-term memory.

Adult

OmniExtract: an automatic data extraction tool based on large language model and prompt engineering.

Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recent advances in large language models (LLMs) have demonstrated strong capabilities in language understanding, and a number of LLM-based tools have been developed for extraction-oriented tasks. However, it's still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files that can adapt to various data extraction tasks. OmniExtract employs a prompt optimization method to refine task-specific prompts and achieve high extraction performance. It also supports comprehensive data extraction from both documents and tables, making it applicable to a broad range of data sources. Evaluation results show that OmniExtract obtains a high accuracy ~90% for three datasets. Furthermore, two additional data extraction applications of OmniExtract in real-world scenarios have been presented, achieving an accuracy of 92.21% and ~90% precision and recall, respectively. Specifically, OmniExtract can handle tabular files of various sizes and formats, and achieve over 99% precision and recall on table information extraction tasks. The data reliability performance shows that OmniExtract is a valuable tool for database updating. An online testing service is available at https://ngdc.cncb.ac.cn/omniextract/. The service can be deployed locally with the code in https://github.com/wyb39/OmniExtract.

Large Language Models

Storing covariance with nonlinearly interacting neurons.

A time-dependent, nonlinear model of neuronal interaction which was probabilistically analyzed in a previous article is shown here to be a natural generalization of the Hartline-Ratliff model of the Limulus retina. Although the primary physical variables in the model are the membrane potentials of neurons, the equations which govern the means and covariances of the membrane potentials are coupled through the average firing rates; as a consequence, the average firing rates control the selective storage and retrieval of covariance information. Motor learning in the cerebellar cortex is treated as a problem of covariance storage, and a predicition is made for the underlying synaptic plasticity: the change in synaptic strength between a parallel fiber and a Purkinje cell should be proportional to the covariance between discharges in the parallel fiber and the climbing fiber. Unlike previous proposals for synaptic plasticity, this prediction requires both facilitation and depression to occur (under different conditions) at the same synapse.

Cerebellum

MetaServe: a lightweight, metadata-aware governance and delivery layer for pre-publication research omics data.

BACKGROUND: Institutional research teams and core facilities routinely manage pre-publication omics datasets that span heterogeneous file types, nested project structures, and multiple downstream uses. Public repositories mainly support post-publication dissemination, while workflow systems and enterprise data platforms do not directly provide a lightweight governance and delivery layer for internal research assets. RESULTS: We present MetaServe, an open-source governance and delivery layer for pre-publication research assets in institutional multi-omics settings. MetaServe registers and delivers heterogeneous assets, including sequencing files, processed matrices, imaging data, analysis-ready objects, tabular files, and documents, without requiring repository-grade standardization. Its metadata-aware design combines file-type recognition, partial automatic extraction for selected formats, manually supplied project and biological annotations, and indexed faceted retrieval. MetaServe supports authenticated web download, viewer-oriented handoff for compatible services such as cellxgene, and path-manifest export for downstream workflows under shared-storage assumptions. The current implementation combines role-based controls, explicit file-level sharing, path-constrained delivery, and operational traceability to support controlled institutional access. MetaServe has been deployed at the Chinese Institutes for Medical Research (CIMR) as part of an institutional multi-omics data-management system. CONCLUSIONS: MetaServe provides a practical layer between institutional storage and downstream analytical platforms for pre-publication research data. Its contribution is the integration of lightweight metadata-aware registration, permission-aware retrieval, and controlled delivery for heterogeneous institutional omics assets. Rather than replacing workflow engines, public repositories, or enterprise-scale research data platforms, MetaServe offers a deployable governance layer for core facilities and collaborative teams that need structured discovery and traceable delivery before public deposition or manuscript release.

Metadata

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic

Pre-Meta: priors-augmented retrieval for LLM-based metadata generation.

MOTIVATION: While high-throughput sequencing technologies have dramatically accelerated genomic data generation, the manual processes required for dataset annotation and metadata creation impede the efficient discovery and publication of these resources across disparate public repositories. Large language models (LLMs) have the potential to streamline dataset profiling and discovery. However, their current limitations in generalizing across specialized knowledge domains, particularly in fields such as biomedical genomics, prevent them from fully realizing this potential. This article presents Pre-Meta, an LLM-agnostic and domain-independent data annotation pipeline with an enriched retrieval procedure that leverages related priors-such as pre-generated metadata tags and ontologies-as auxiliary information to improve the accuracy of automated metadata generation. RESULTS: Validated using five selected metadata fields sampled across 1500 papers, the Pre-Meta assisted annotation experiment-without finetuning and prompt optimization-demonstrates a systemic improvement in the annotation task: shown through a 23%, 72%, and 75% accuracy gain from conventional RAG adoptions of GPT-4o mini, Llama 8B, and Mistral 7B respectively. AVAILABILITY AND IMPLEMENTATION: The code, data access, and scripts are available at: https://github.com/SINTEF-SE/LLMDap.

Metadata

Imagery as an aid to retrieval for Korsakoff patients.

Six Korsakoff patients and six alcoholic controls learned a five item P-A task under each of the following three learning conditions; Rote, Imagery, and Cued learning. Under all conditions the Korsakoff patients took more trials to learn than did the control patients. However, both imagery learning and cued learning were easier than rote learning for the Korsakoff patients when recall was used as the learning index. When a recognition measure was used instead of the recall, imagery learning proved easiest with no difference existing between cued and rote learning. In a second experiment, the patients were given the cue (a mediating link) during presentation, but not during retrieval. Under this condition the Korsakoff patients learned no more rapidly than they did by rote regardless which response measure was required. It was concluded that imagery can aid both storage and retrieval of verbal information for Korsakoff patients, while cuing aids only the retrieval process.

Alcohol Amnestic Disorder

Sharing and community curation of mass spectrometry data with Global Natural Products Social Molecular Networking.

The potential of the diverse chemistries present in natural products (NP) for biotechnology and medicine remains untapped because NP databases are not searchable with raw data and the NP community has no way to share data other than in published papers. Although mass spectrometry (MS) techniques are well-suited to high-throughput characterization of NP, there is a pressing need for an infrastructure to enable sharing and curation of data. We present Global Natural Products Social Molecular Networking (GNPS; http://gnps.ucsd.edu), an open-access knowledge base for community-wide organization and sharing of raw, processed or identified tandem mass (MS/MS) spectrometry data. In GNPS, crowdsourced curation of freely available community-wide reference MS libraries will underpin improved annotations. Data-driven social-networking should facilitate identification of spectra and foster collaborations. We also introduce the concept of 'living data' through continuous reanalysis of deposited data.

Biological Products

A computer-based reporting system for bone marrow evaluation.

A computerized system of reporting results of bone marrow examination has been implemented at The University of Texas System Cancer Center, M. D. Anderson Hospital and Tumor Institute, since 1975. All results are entered via the keyboard of cathode-ray tube (CRT) consoles and are instantly available to the attending physician in CRTs strategically located throughout the institution in patient-related areas. Differential counts are available within two to three hours after aspiration. Description of clot sections and smear, diagnosis, and differential diagnosis are completed within 24 hours. Permanent reports are also printed by the computer at regular intervals during the day. The physician can ask the computer for results on a given day or for a cumulative summary displayed in tabular or plot form. This system has proven efficient and rapid, not only for direct reporting of bone marrow examination results but also for storage and retrieval of patient information.

Bone Marrow Examination

WormBase as an integrated platform for the C. elegans ORFeome.

The ORFeome project has validated and corrected a large number of predicted gene models in the nematode C. elegans, and has provided an enormous resource for proteome-scale studies. To make the resource useful to the research and teaching community, it needs to be integrated with other large-scale data sets, including the C. elegans genome, cell lineage, neurological wiring diagram, transcriptome, and gene expression map. This integration is also critical because the ORFeome data sets, like other 'omics' data sets, have significant false-positive and false-negative rates, and comparison to related data is necessary to make confidence judgments in any given data point. WormBase, the central data repository for information about C. elegans and related nematodes, provides such a platform for integration. In this report, we will describe how C. elegans ORFeome data are deposited in the database, how they are used to correct gene models, how they are integrated and displayed in the context of other data sets at the WormBase Web site, and how WormBase establishes connection with the reagent-based resources at the ORFeome project Web site.

Animals

Automated Extraction of Tumor Staging and Diagnosis Information From Surgical Pathology Reports.

PURPOSE: Typically stored as unstructured notes, surgical pathology reports contain data elements valuable to cancer research that require labor-intensive manual extraction. Although studies have described natural language processing (NLP) of surgical pathology reports to automate information extraction, efforts have focused on specific cancer subtypes rather than across multiple oncologic domains. To address this gap, we developed and evaluated an NLP method to extract tumor staging and diagnosis information across multiple cancer subtypes. METHODS: The NLP pipeline was implemented on an open-source framework called Leo. We used a total of 555,681 surgical pathology reports of 329,076 patients to develop the pipeline and evaluated our approach on subsets of reports from patients with breast, prostate, colorectal, and randomly selected cancer subtypes. RESULTS: Averaged across all four cancer subtypes, the NLP pipeline achieved an accuracy of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.89 for T staging, 0.90 for N staging, and 0.97 for M staging. It achieved an F1 score of 1.00 for International Classification of Diseases, Tenth Revision codes, 0.88 for T staging, 0.90 for N staging, and 0.24 for M staging. CONCLUSION: The NLP pipeline was developed to extract tumor staging and diagnosis information across multiple cancer subtypes to support the research enterprise in our institution. Although it was not possible to demonstrate generalizability of our NLP pipeline to other institutions, other institutions may find value in adopting a similar NLP approach-and reusing code available at GitHub-to support the oncology research enterprise with elements extracted from surgical pathology reports.

Humans

Computer processing of audiological and vestibular data. II. A further note.

This paper provides an addendum to an earlier paper describing the development of a computer record system for patient data. The specific problems addressed pertain to the storage and retrieval of historical information, physical signs and diagnosis. Some preliminary comparisons of audiological and vestibular test results are given for groups of patients with diagnoses of acoustic neuroma, Ménière's disease, temporal bone fracture and vestibular neuronitis.

Diagnosis, Computer-Assisted

Pithos - a scalable and secure data container for FAIR-compliant research data management in life sciences.

Modern research techniques have led to exponential growth in the volume and complexity of scientific data. Consequently, managing these volumes securely and efficiently has become a major challenge. While all research domains face these challenges, life science research is particularly affected because current approaches often rely on a large set of different file formats, with metadata stored in separated databases or spreadsheets. This leads to fragmented datasets, orphaned data, and compromised research reproducibility. Traditional solutions also force researchers to choose between security and accessibility, with encrypted files preventing selective access and indexed formats lacking adequate security for sensitive data. These limitations are particularly problematic in large-scale genomic studies where researchers must decompress multi-gigabyte files to access specific regions, creating computational bottlenecks and inefficient network usage when working with cloud-stored datasets. We introduce Pithos, a next-generation file format specifically designed for scientific data management in distributed cloud environments. The format uses content-defined chunking to enable efficient deduplication across distributed storage systems, thereby reducing storage costs and bandwidth requirements. The append-only structure ensures data immutability and allows for incremental updates without compromising content. Benchmark results show that Pithos outperforms existing solutions in read and write performance, with comparable or improved storage efficiency.

Biological Science Disciplines

[The design of a personal computerized database].

The development of a personal data base about informatics and nursing for use on a microcomputer is presented. The advantages, effectiveness and problems involved with the construction of personal library, as well as the employed software and the references classification problems are discussed.

Computer Systems