PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Data Storage And Retrieval”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 775 records · Page 43Linked to original sources

Classification and regression tree analysis of 1000 consecutive patients with unknown primary carcinoma.

The clinical features and survival times of patients with unknown primary carcinoma (UPC) are heterogeneous. Therefore, the goals of this study were to apply a novel analytical method to UPC patients to: (a) identify novel prognostic factors; (b) explore the interactions between clinical variables and their impact on survival; and (c) illustrate explicitly how the covariates interact. The 1000 patients analyzed were referred to the University of Texas M. D. Anderson Cancer Center from January 1, 1987 through November 30, 1994. Clinical data from these patients were entered into a computerized database for storage, retrieval, and analysis. Multivariate analyses of survival were performed using recursive partitioning referred to as classification and regression tree (CART) analysis. The median survival for all 1000 consecutive UPC patients was 11 months. CART was performed with an initial split on liver involvement, and 10 terminal subgroups were formed. Median survival of the 10 subgroups ranged from 40 months (95% confidence interval, 22-66 months) for UPC patients with one or two metastatic organ sites, with nonadenocarcinoma histology, and without liver, bone, adrenal, or pleural metastases to 5 months (95% confidence interval, 4-7 months) in UPC patients with liver metastases, tumor histologies other than neuroendocrine carcinoma, age >61.5 years, and a small subgroup of patients with adrenal metastases. Two additional trees were also explored. These analyses demonstrated that important prognostic variables were consistently applied by the CART program and effectively segregated patients into groups with similar clinical features and survival. CART also identified previously unappreciated patient subsets and is a useful method for dissecting complex clinical situations and identifying homogeneous patient populations for future clinical trials.

Adolescent↗

Bioinformatics methods to predict protein structure and function. A practical approach.

Protein structure prediction by using bioinformatics can involve sequence similarity searches, multiple sequence alignments, identification and characterization of domains, secondary structure prediction, solvent accessibility prediction, automatic protein fold recognition, constructing three-dimensional models to atomic detail, and model validation. Not all protein structure prediction projects involve the use of all these techniques. A central part of a typical protein structure prediction is the identification of a suitable structural target from which to extrapolate three-dimensional information for a query sequence. The way in which this is done defines three types of projects. The first involves the use of standard and well-understood techniques. If a structural template remains elusive, a second approach using nontrivial methods is required. If a target fold cannot be reliably identified because inconsistent results have been obtained from nontrivial data analyses, the project falls into the third type of project and will be virtually impossible to complete with any degree of reliability. In this article, a set of protocols to predict protein structure from sequence is presented and distinctions among the three types of project are given. These methods, if used appropriately, can provide valuable indicators of protein structure and function.

Algorithms↗

Information management of a department of diagnostic imaging.

It is well-known that while RIS allows the management of all input and output data of a Radiology service, PACS plays a major role in the management of all radiologic images. However, the two systems should be closely integrated: scheduling of a radiologic exam requires direct automated integration with the system of image management for retrieval of previous exams and storage of the exam just completed. A modern information system of integration of data and radiologic images should be based on an automated work flow management in al its components, being at the same time flexible and compatible with the ward organization to support and computerize each stage of the working process. Similarly, standard protocols (DICOM 3.0, HL7) defined for interfacing with the Diagnostic Imaging (D.I.) department and the other components of modules of a modern HIS, should be used. They ensure the system to be expandable and accessible to ensure share and integration of information with HIS, emergency service or wards. Correct RIS/PACS integration allows a marked improvement in the efficiency of a modern D.I. department with a positive impact on the daily activity, prompt availability of previous data and images with sophisticated handling of diagnostic images to enhance the reporting quality. The increased diffusion of internet and intranet technology predicts future developments still to be discovered.

Appointments and Schedules↗

New concepts in the design of a clinical laboratory information system (LIS).

Laboratory computers have expanded in potential and are now capable of handling a communication system to support patient care. New technics and interactive software to expand this function are described and were developed because of serious deficiencies discovered in the hospital laboratory operation. New concepts in ordering of tests, collection and identification of samples, reporting of results, and retrieval of data are described.

Clinical Laboratory Information Systems↗

Electronic signatures for long-lasting storage purposes in electronic archives.

Communication and co-operation in healthcare and welfare require a certain set of trusted third party (TTP) services describing both status and relation of communicating principals as well as their corresponding keys and attributes. Additional TTP services are needed to provide trustworthy information about dynamic issues of communication and co-operation such as time and location of processes, workflow relations, and system behaviour. Legal and ethical requirements demand securely stored patient information and well-defined access rights. Among others, electronic signatures based on asymmetric cryptography are important means for securing the integrity of a message or file as well as for accountability purposes including non-repudiation of both origin and receipt. Electronic signatures along with certified time stamps or time signatures are especially important for electronic archives in general, electronic health records (EHR) in particular, and especially for typical purposes of long-lasting storage. Apart from technical storage problems (e.g. lifetime of the storage devices, interoperability of retrieval and presentation software), this paper identifies mechanisms of e.g. re-signing and re-stamping of data items, files, messages, sets of archived items or documents, archive structures, and even whole archives.

Computer Security↗

Software tools for high-throughput analysis and archiving of immunohistochemistry staining data obtained with tissue microarrays.

The creation of tissue microarrays (TMAs) allows for the rapid immunohistochemical analysis of thousands of tissue samples, with numerous different antibodies per sample. This technical development has created a need for tools to aid in the analysis and archival storage of the large amounts of data generated. We have developed a comprehensive system for high-throughput analysis and storage of TMA immunostaining data, using a combination of commercially available systems and novel software applications developed in our laboratory specifically for this purpose. Staining results are recorded directly into an Excel worksheet and are reformatted by a novel program (TMA-Deconvoluter) into a format suitable for hierarchical clustering analysis or other statistical analysis. Hierarchical clustering analysis is a powerful means of assessing relatedness within groups of tumors, based on their immunostaining with a panel of antibodies. Other analyses, such as generation of survival curves, construction of Cox regression models, or assessment of intra- or interobserver variation, can also be done readily on the reformatted data. Finally, the immunoprofile of a specific case can be rapidly retrieved from the archives and reviewed through the use of Stainfinder, a novel web-based program that creates a direct link between the clustered data and a digital image database. An on-line demonstration of this system is available at http://genome-www.stanford.edu/TMA/explore.shtml.

Cluster Analysis↗

Graph kernels for chemical informatics.

Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MinMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed.

Anticarcinogenic Agents↗

PoPS: a computational tool for modeling and predicting protease specificity.

Proteases play a fundamental role in the control of intra- and extracellular processes by binding and cleaving specific amino acid sequences. Identifying these targets is extremely challenging. Current computational attempts to predict cleavage sites are limited, representing these amino acid sequences as patterns or frequency matrices. Here we present PoPS, a publicly accessible bioinformatics tool (http://pops.csse.monash.edu.au/) which provides a novel method for building computational models of protease specificity that, while still being based on these amino acid sequences, can be built from any experimental data or expert knowledge available to the user. PoPS specificity models can be used to predict and rank likely cleavages within a single substrate, and within entire proteomes. Other factors, such as the secondary or tertiary structure of the substrate, can be used to screen unlikely sites. Furthermore, the tool also provides facilities to infer, compare and test models, and to store them in a publicly accessible database.

Amino Acid Sequence↗

PGAGENE: integrating quantitative gene-specific results from the NHLBI programs for genomic applications.

SUMMARY: PGAGENE is a web-based gene-specific genomic data search engine, which allows users to search over 5.9 million pieces of collective genetic and genomic data from the NHLBI supported Programs for Genomic Applications. This data includes microarray measurements, SNPs, and mutations, and data may be found using symbols, parts of gene names or products, Affymetrix probe IDs, GenBank accession numbers, UniGene IDs, dbSNP IDs, and others. The PGAGENE indexing agent periodically maps all publicly available gene-specific PGA data onto LocusLink using dynamically generated cross-referencing tables.

Base Sequence↗

eMelanoBase: an online locus-specific variant database for familial melanoma.

A proportion of melanoma-prone individuals in both familial and non-familial contexts has been shown to carry inactivating mutations in either CDKN2A or, rarely, CDK4. CDKN2A is a complex locus that encodes two unrelated proteins from alternately spliced transcripts that are read in different frames. The alpha transcript (exons 1alpha, 2, and 3) produces the p16INK4A cyclin-dependent kinase inhibitor, while the beta transcript (exons 1beta and 2) is translated as p14ARF, a stabilizing factor of p53 levels through binding to MDM2. Mutations in exon 2 can impair both polypeptides and insertions and deletions in exons 1alpha, 1beta, and 2, which can theoretically generate p16INK4A-p14ARF fusion proteins. No online database currently takes into account all the consequences of these genotypes, a situation compounded by some problematic previous annotations of CDKN2A-related sequences and descriptions of their mutations. As an initiative of the international Melanoma Genetics Consortium, we have therefore established a database of germline variants observed in all loci implicated in familial melanoma susceptibility. Such a comprehensive, publicly accessible database is an essential foundation for research on melanoma susceptibility and its clinical application. Our database serves two types of data as defined by HUGO. The core dataset includes the nucleotide variants on the genomic and transcript levels, amino acid variants, and citation. The ancillary dataset includes keyword description of events at the transcription and translation levels and epidemiological data. The application that handles users' queries was designed in the model-view-controller architecture and was implemented in Java. The object-relational database schema was deduced using functional dependency analysis. We hereby present our first functional prototype of eMelanoBase. The service is accessible via the URL www.wmi.usyd.edu.au:8080/melanoma.html.

Computer Security↗

SSHSuite: an integrated software package for analysis of large-scale suppression subtractive hybridization data.

Suppression subtractive hybridization (SSH) is a widely used technique for the identification of differentially expressed genes. SSH as well as other types of sequencing projects generate large amounts of anonymous sequences. SSHSuite automates the handling and storage of these sequences and enables identification through similarity searches. SSHSuite also offers analysis tools for the retrieval and comparison of the resulting similarity data. SSHSuite consists of four programs: SSHHandler, SSHOverview, SSHAnalysis, and SSHCompare.

Algorithms↗

Improving sensitivity in shotgun proteomics using a peptide-centric database with reduced complexity: protease cleavage and SCX elution rules from data mining of MS/MS spectra.

Correct identification of a peptide sequence from MS/MS data is still a challenging research problem, particularly in proteomic analyses of higher eukaryotes where protein databases are large. The scoring methods of search programs often generate cases where incorrect peptide sequences score higher than correct peptide sequences (referred to as distraction). Because smaller databases yield less distraction and better discrimination between correct and incorrect assignments, we developed a method for editing a peptide-centric database (PC-DB) to remove unlikely sequences and strategies for enabling search programs to utilize this peptide database. Rules for unlikely missed cleavage and nontryptic proteolysis products were identified by data mining 11 849 high-confidence peptide assignments. We also evaluated ion exchange chromatographic behavior as an editing criterion to generate subset databases. When used to search a well-annotated test data set of MS/MS spectra, we found no loss of critical information using PC-DBs, validating the methods for generating and searching against the databases. On the other hand, improved confidence in peptide assignments was achieved for tryptic peptides, measured by changes in DeltaCN and RSP. Decreased distraction was also achieved, consistent with the 3-9-fold decrease in database size. Data mining identified a major class of common nonspecific proteolytic products corresponding to leucine aminopeptidase (LAP) cleavages. Large improvements in identifying LAP products were achieved using the PC-DB approach when compared with conventional searches against protein databases. These results demonstrate that peptide properties can be used to reduce database size, yielding improved accuracy and information capture due to reduced distraction, but with little loss of information compared to conventional protein database searches.

Amino Acid Sequence↗

AluGene: a database of Alu elements incorporated within protein-coding genes.

Alu elements are short interspersed elements (SINEs) approximately 300 nucleotides in length. More than 1 million Alus are found in the human genome. Despite their being genetically functionless, recent findings suggest that Alu elements may have a broad evolutionary impact by affecting gene structures, protein sequences, splicing motifs and expression patterns. Because of these effects, compiling a genomic database of Alu sequences that reside within protein-coding genes seemed a useful enterprise. Presently, such data are limited since the structural and positional information on genes and Alu sequences are scattered throughout incompatible and unconnected databases. AluGene (http://Alugene.tau.ac.il/) provides easy access to a complete Alu map of the human genome, as well as Alu-associated information. The Alu elements are annotated with respect to coding region and exon/intron location. This design facilitates queries on Alu sequences, locations, as well as motifs and compositional properties via a one-stop search page.

Alu Elements↗

A multiple-pattern biosequence analysis method for diverse source association mining.

BACKGROUND: In order to understand the intricacy of biomolecules more comprehensively, significant patterns extracted from related data collected from diverse sources must be integrated. These data sources may be local or distributed, possibly with different representation schemes. Often, related data from different sources correspond only with respect to some of their values. METHODS: In biological sequence analysis, a goal is to identify new, previously unknown, relevant patterns, to obtain additional insights into the biomolecule. This is known as a pattern discovery task, rather than a pattern matching task. In this research, we present a method to tackle this problem typically found in molecular sequence analysis when the alignment of the sequences is represented as a relation. In this article, we propose an information measure to select attribute values that reflect multiple patterns of significant interdependence information. Based on these selected values, the patterns are evaluated with data values from other sources. RESULTS: In the experiments, a cancer-suppressor gene known as TP53 (encoding tumour protein p53) is analysed with the mutation records of patients. The experiments identify previously unknown points in the molecule that have patterns negatively associated with the occurrence of cancer. CONCLUSION: Since the evaluated interdependence pattern is a global property of the molecule, we conjecture that the identified points might also be a reflection of the molecule's cancer-suppressor characteristics. The experiments also confirm the usefulness of the proposed method.

Amino Acid Sequence↗

Probability-based protein identification by searching sequence databases using mass spectrometry data.

Several algorithms have been described in the literature for protein identification by searching a sequence database using mass spectrometry data. In some approaches, the experimental data are peptide molecular weights from the digestion of a protein by an enzyme. Other approaches use tandem mass spectrometry (MS/MS) data from one or more peptides. Still others combine mass data with amino acid sequence data. We present results from a new computer program, Mascot, which integrates all three types of search. The scoring algorithm is probability based, which has a number of advantages: (i) A simple rule can be used to judge whether a result is significant or not. This is particularly useful in guarding against false positives. (ii) Scores can be compared with those from other types of search, such as sequence homology. (iii) Search parameters can be readily optimised by iteration. The strengths and limitations of probability-based scoring are discussed, particularly in the context of high throughput, fully automated protein identification.

Amino Acid Sequence↗

Facilitating the work of a meta-analyst.

Meta-analysis facilitates the transfer of knowledge from nurse researchers to clinicians. In this article, the benefits and criticisms of meta-analysis for nursing are identified along with the specific problems a meta-analyst may encounter in conducting a quantitative analysis and synthesis of the literature. Problems in data retrieval from the primary studies for a quantitative literature review can plague a meta-analyst. These problems can include insufficient data such as inexact p-values, incomplete descriptions of samples or experimental and control groups, and errors in tables in research reports. Suggestions for removing some of these roadblocks are addressed along with recommendations to authors, editors, and manuscript reviewers. An example of a suggestion for authors is to focus on reporting exact statistical values and p-levels. Also, sample characteristics and methodological variables are deserving of detailed descriptions in research reports. Editors can consider including a summary table of basic descriptive and inferential statistics. Another recommendation for editors is to develop a reviewer's checklist to ensure the author has included all relevant statistical information and study characteristics a meta-analyst needs.

Authorship↗

Moving time window aggregates over patient histories.

Moving window concepts are used in temporal query languages for aggregate functions over the dimension 'time'. In the medical domain, aggregation of patient data over time windows builds a powerful mechanism within clinical database queries to satisfy a class of typical medical question formulations. Contrary to other fields, like the business domain for example, there is the additional need to synchronize time windows with the individual course of diseases rather than with the calendar system only. In this paper, we present several variants of shifting time windows over patient histories and suggest a set of essential options for a moving window clause. The proposed parameters for window creation as well as suitable default settings are discussed in the context of retrieving data from medical records.

Databases as Topic↗

Covariation of amino acid positions in HIV-1 protease.

We have examined patterns of sequence variability for evidence of linked sequence changes in HIV-1 subtype B protease using translated sequences from protease inhibitor (PI) treated and untreated subjects downloaded from the Stanford HIV RT and Protease Sequence Database (http://hivdb.stanford.edu). The final data set size was 648 sequences from untreated subjects (notx) and 531 for PI-treated subjects (tx). Each subject was uniquely represented by a single sequence. Mutual information was calculated for all pairwise comparisons of positions with nonconsensus amino acids in at least 5% of sequences; significance of pairwise association was assessed using permutation tests. In addition pairs of positions were assessed for linkage by comparing the observed occurrences of amino acid combinations to expected values. The mutual information statistic indicated linkage between nine pairs of sites in the untreated data set (10:93, 12:19, 35:38, 37:41, 62:71, 63:64, 71:77, 71:93, 77:93). Strong statistical support for linkage in the treated data set was seen for 32 pairs, eight involving position 10:7 involving position 71, with the rest being 12:19, 15:77, 20:36, 30:88, 35:36, 35:37, 36:62, 36:77, 46:82, 46:84, 48:54, 48:82, 54:82, 63:64, 63:90, 73:90, 77:93, and 84:90. Most associations were positive, although negative associations were seen for five pairs of interactions. Structural proximity suggests that numerous pairs may interact within a local environment. These interactions include two distinct clusters around 36/77 and 71/93. While some of these interactions may reflect fortuitous linkage in heavily treated subjects with many resistance mutations, others will likely represent important cooperative interactions that are amenable to experimental validation.

Amino Acid Sequence↗