PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data fragmentation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

FragMatch--a program for the analysis of DNA fragment data.

FragMatch is a user-friendly Java-supported program that automates the identification of taxa present in mixed samples by comparing community DNA fragment data against a database of reference patterns for known species. The program has a user-friendly Windows interface and was primarily designed for the analysis of fragment data derived from terminal restriction fragment length polymorphism analysis of ectomycorrhizal fungal communities, but may be adapted for other applications such as microsatellite analyses. The program uses a simple algorithm to check for the presence of reference fragments within sample files that can be directly imported, and the results appear in a clear summary table that also details the parameters that were used for the analysis. This program is significantly more flexible than earlier programs designed for matching RFLP patterns as it allows default or user-defined parameters to be used in the analysis and has an unlimited database size in terms of both the number of reference species/individuals and the number of diagnostic fragments per database entry. Although the program has been developed with mycorrhizal fungi in mind, it can be used to analyse any DNA fragment data regardless of biological origin. FragMatch, along with a full description and users guide, is freely available to download from the Aberdeen Mycorrhiza Group web page (http://www.aberdeenmycorrhizas.com).

Algorithms↗

Creation and characterization of a new, non-redundant fragment data bank.

The success achieved for protein structure prediction of loop regions with insertions and deletions by knowledge-based methods depends on the quality of the underlying information, i.e. a fragment data bank as complete as possible is needed. However, the greater the number of proteins contributing to the data base the more redundant information is included, which leads to structurally similar proposals in loop predictions and to longer times for extracting fragments. So it is not only necessary to increase the number of proteins for building the loop data base but also to cluster the resulting fragments according to their structural similarities in order to remove redundancy. Here, a new, non-redundant fragment data bank is described, which is based on all proteins in the Brookhaven Protein Data Bank (release 7/95) with a resolution > or = 2.0 A and which can be updated easily by including new information from structures to be solved in the future. In the clustering process presented, the resulting clusters are optimized in several cycles until self-consistency. In this way all redundant information is removed without loosing any significantly different fragments. Finally the resulting fragment data bank is analysed with respect to its completeness.

Algorithms↗

Development of a mass fingerprinting tool for automated interpretation of oligosaccharide fragmentation data.

The bioinformatic tool GlycosidIQ was developed for computerized interpretation of oligosaccharide mass spectrometric fragmentation based on matching experimental data with theoretically fragmented oligosaccharides generated from the database GlycoSuiteDB. This use of the software for glycofragment mass fingerprinting obviates a large part of the manual, labor intensive, and technically challenging interpretation of oligosaccharide fragmentation. Using 130 negative ion electrospray ionization-tandem mass spectrometry fragment spectra from identified oligosaccharide structures, it was shown that the GlycosidIQ scoring algorithms were able to correctly identify oligosaccharides in the great majority of cases (correct structure top ranked in 78% of the cases and an additional 17% were ranked second highest in the sample set).

Algorithms↗

Identifying proteins using matrix-assisted laser desorption/ionization in-source fragmentation data combined with database searching.

Metastable ion decay in matrix-assisted laser desorption/ionization (MALDI) has become a routine method for obtaining primary structures of peptides. Significant fragmentation occurs in the MALDI ion source and can be observed via delayed ion extraction TOF-MS. In-source decay (ISD) can provide C- and N-terminal primary sequence data for even moderate-sized peptides (< 5000 Da). The unique cn series fragmentation that occurs in ISD has been exploited to obtain partial C-terminal sequences for proteins as large as human apotransferrin (75 kDa). Two approaches for combining this ISD MALDI-generated partial sequence information with protein database searching techniques are presented. In one approach, cyanogen bromide is used to cleave relatively large peptide fragments from a sample of human apotransferrin. One of the larger cleavage products (6034.84 Da) was isolated by HPLC and subjected to ISD MALDI analysis. An easily identified cn fragment ion series allowed two noncontiguous segments of the peptide's sequence to be determined (about 55% of the total sequence). This partial sequence information was used to search protein and oligonucleotide sequence databases. In addition to uniquely identifying human apotransferrin in a protein sequence database, an example of the use of this ISD MALDI-determined partial sequence information to search expressed sequence tag databases is presented. Such searches have the potential for rapidly identifying new genes that code for target proteins. An alternate approach for obtaining partial sequence information on proteins is also demonstrated that utilizes ISD MALDI fragmentation of the intact protein to generate partial sequence information. This approach is shown to generate about 5-7% of a protein's sequence, usually near the C-terminus of the protein. Examples of the ISD MALDI fragmentation data obtained from intact (reduced) human apotransferrin and intact (nonreduced) bovine serum albumin (66 kDa) proteins are presented.

Animals↗

Neural network classification of mutagens using structural fragment data.

A neural network was applied to a large, structurally heterogeneous data set of mutagens and non-mutagens to investigate structure-property relationships. Substructural data comprising a total of 1280 fragments were used as inputs. The training of the back-propagation networks was directed by an algorithm which selected an optimal subset of fragments in order to maximize their discriminating power, and a good predictive network. The system comprised three programs: the first used a keyfile of 100 fragments to generate training and test files, the second was the network itself and a procedure for ranking the effectiveness of these fragments and the third randomly replaced the lowest fragments. This cycle was then repeated. After running on a 386/33 PC several networks produced approximately 11% failures in the test set and 6% in the training set. By simplifying the output of the hidden layer it was possible to describe the hidden layer states in terms of clusters of mutagens and non-mutagens. Some of these clusters were structurally homogeneous and contained known mutagenic and non-mutagenic structural classes. This analysis provided a useful means of demonstrating how the network was classifying the data.

Cluster Analysis↗

CHASE, a charge-assisted sequencing algorithm for automated homology-based protein identifications with matrix-assisted laser desorption/ionization time-of-flight post-source decay fragmentation data.

We describe CHASE, a novel algorithm for automated de novo sequencing based on the mass spectrometric (MS) fragmentation analysis of tryptic peptides. This algorithm is used for protein identification from sequence similarity criteria and consists of four steps: (1) derivatization of tryptic peptides at the N-terminus with a negatively charged reagent; (2) post-source decay (PSD) fragmentation analysis of peptides; (3) interpretation of the mass peaks with the CHASE algorithm and reconstruction of the amino acid sequence; (4) transfer of these data to software for protein identifications based on sequence homology (Basic Local Alignment Search Tool, BLAST). This procedure deduced the correct amino acid sequence of tryptic peptide samples and also was able to deduce the correct sequence from difficult mass patterns and identify the amino acid sequence. This allows complete automation of the process starting from MS fragmentation of complex peptide mixtures at low concentration (e.g. from silver-stained gel bands) to identification of the protein. We also show that if PSD data are collected in a single spectrum (instead of the segmented mode offered by conventional matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) instrumentation), the complete workflow from MS-PSD data acquisition to similarity-based identification can be completely automated. This strategy may be applied to proteomic studies for protein identification based on automated de novo sequencing instead of MS or tandem MS patterns. We describe the Charge Assisted Sequencing Engine (CHASE) algorithm, the working protocol, the performance of the algorithm on spectra from MALDI-TOFMS and the data comparison between a TOF and a TOF-TOF instrument.

Algorithms↗

Study of the mass spectrometric fragmentation of pseudouridine: comparison of fragmentation data obtained by matrix-assisted laser desorption/ionisation post-source decay, electrospray ion trap multistage mass spectrometry, and by a method utilising electrospray quadrupole time-of-flight tandem mass spectrometry and in-source fragmentation.

Many nucleosides and their modified forms have been studied by mass spectrometry elaborating the detailed fragmentation pathways under MS2 and MS(n) conditions. Although the C-nucleoside pseudouridine has been fragmented and studied briefly, usually amongst many other nucleosides, it has not been investigated to the same extent as other nucleosides. In this report a number of different mass spectrometric techniques are applied to obtain a fuller picture of pseudouridine fragmentation. At the same time this study is used to compare different tandem mass spectrometric techniques, including a novel methodology utilising a quadrupole time-of-flight (Q-ToF) instrument for MS(n) analysis comparable with that available with an ion trap mass spectrometer.

Pseudouridine↗

Effect of bleaching on the width and index of refraction of goldfish rod and cone outer segment fragments.

Data on goldfish rod and cone fragments were obtained before and after a strong bleach by using a Zeiss Jamin-Lebedeff infrared interference microscope and computer image processing techniques. The receptor fragments were fractured in the inner segment and were immersed in goldfish aqueous humor medium. On bleaching we found that: (1) there is a small increase in rod diameter. This effect is of the same magnitude reported earlier in Rana pipiens outer segments (Vision Res 1973; 13:171). (2) There is an inferred decrease in rod refractive index. (3) There is a decrease in cone width. (4) There is a slight increase in inferred cone refractive index. These data are presented.

Animals↗

Determination of three-dimensional protein structures from nuclear magnetic resonance data using fragments of known structures.

A method to build a three-dimensional protein model from nuclear magnetic resonance (NMR) data using fragments from a data base of crystallographically determined protein structures is presented. The interproton distances derived from the nuclear Overhauser effect (NOE) data are compared to the precalculated distances in the known protein structures. An efficient search algorithm is used, which arranges the distances in matrices akin to a C alpha diagonal distance plot, and compares the NOE distance matrices for short sequential zones of the protein to the data base matrices. After cluster analysis of the fragments found in this way, the structure is built by aligning fragments in overlapping zones. The sequentially long-range NOEs cannot be used in the initial fragments search but are vital to discriminate between several possible combinations of different groups of fragments. The method has been tested on one simulated NOE data set derived from a crystal structure and one experimental NMR data set. The method produces models that have good local structure, but may contain larger global errors. These models can be used as the starting point for further refinement, e.g., by restrained molecular dynamics or interactive graphics.

Crystallography↗

Relative efficiencies of the maximum-parsimony and distance-matrix methods of phylogeny construction for restriction data.

The relative efficiencies of the maximum-parsimony (MP), UPGMA, and neighbor-joining (NJ) methods in obtaining the correct tree (topology) for restriction-site and restriction-fragment data were studied by computer simulation. In this simulation, six DNA sequences of 16,000 nucleotides were assumed to evolve following a given model tree. The recognition sequences of 20 different six-base restriction enzymes were used to identify the restriction sites of the DNA sequences generated. The restriction-site data and restriction-fragment data thus obtained were used to reconstruct a phylogenetic tree, and the tree obtained was compared with the model tree. This process was repeated 300 times. The results obtained indicate that when the rate of nucleotide substitution is constant the probability of obtaining the correct tree (Pc) is generally higher in the NJ method than in the MP method. However, if we use the average topological deviation from the model tree (dT) as the criterion of comparison, the NJ and MP methods are nearly equally efficient. When the rate of nucleotide substitution varies with evolutionary lineage, the NJ method is better than the MP method, whether Pc or dT is used as the criterion of comparison. With 500 nucleotides and when the number of nucleotide substitutions per site was very small, restriction-site data were, contrary to our expectation, more useful than sequence data. Restriction-fragment data were less useful than restriction-site data, except when the sequence divergence was very small. UPGMA seems to be useful only when the rate of nucleotide substitution is constant and sequence divergence is high.

Computer Simulation↗

Relationship between daunorubicin concentration and apoptosis induction in leukemic cells.

Aiming to determine if a concentration window exists in which apoptosis induction by daunorubicin (DNR) is optimal, we studied the relationship between DNR concentration and apoptosis induction in HL60 and K562 cells and in peripheral leukemic cells isolated from three patients with acute myelogenous leukemia (AML). Cells were incubated for 2hr with increasing DNR concentrations and thereafter for 22hr in drug-free medium. Apoptosis was measured by detection of caspase-3-like activity and DNA fragmentation assayed by propidium iodide and flow cytometry. High DNR concentrations initiated faster apoptosis in HL60 cells and in AML cells, as shown by caspase-3 and DNA fragmentation data. DNA fragmentation into small fragments was preceded by the formation of a narrow peak on the left side of the G1 peak, most likely large DNA fragments, but further studies are required for unequivocal confirmation. This peak could easily be misinterpreted as a G1 peak without careful time monitoring. In K562 cells, no left peak was detected, apoptosis was slow and not related to concentration. In AML cells, large interindividual variations were observed in the time course of DNA fragmentation at 0.25microg DNR/mL. In conclusion, our findings support the concept of dose intensification for optimal apoptosis induction as higher doses correlate with earlier and more rapid caspase-3 induction and DNA fragmentation in leukemic cells. The DNA fragmentation assay may be a valuable tool to determine leukemic cells' chemosensitivity to apoptosis.

Antibiotics, Antineoplastic↗

The application of an automated allele concordance analysis system (CompareCalls) to ensure the accuracy of single-source STR DNA profiles.

A powerful method for validating a scientific result is to confirm specific results utilizing independent methodologies and processing pathways. Thus, we have designed, developed and validated an automated allele concordance analysis system (CompareCalls, patent pending) that performs comparisons between two independent DNA analysis platforms to ensure the highest accuracy for allele calls. Application of this system in a quality assurance role has shown the potential to eliminate greater than 90% of the STR analysis required of a DNA data analyst. While this system is broadly applicable for use with any two independent STR analysis programs, either prior to or following human data review, we are presenting its application to data generated with the ABI Prism Genotyper software system versus data generated with the SurelockID system. With the automated allele concordance analysis system, the GeneScan DNA fragment data generated from an ABI 377 gel image are analyzed in two independent pathways. In one analysis pathway, the GeneScan data are imported into Genotyper software where STR labels are assigned to the fragment data based upon the criteria of the Kazam 20% macro. The "Kazam" macro provided with the Genotyper program works by labeling all peaks in a category (or locus) and then filtering (or removing) the labels from peaks, such as those in stutter positions, that meet predefined criteria. In the second pathway, the GeneScan data are imported into the SurelockID analysis platform where STR labels and error messages are assigned to the fragment data based upon hard-coded allele calling criteria and quality parameters. The resulting STR allele calls for each analysis platform are then compared, utilizing the automated allele concordance analysis system. Any differences in the STR allele calls between the two systems are flagged in a discordance report for further review by a qualified DNA data analyst. The automated allele concordance analysis system guides the DNA data analyst to the discordant data generated by either analysis platform. Additionally, the analyst is also directed to data that are of less than pristine quality which may have an increased potential for errors in interpretation by either analysis platform or by a human DNA data analyst. Implementation of an automated allele concordance analysis system will yield high-quality data for CODIS and free the human DNA data analyst to perform other critical duties within the laboratory.

Alleles↗

Identifying citrus limonoid aglycones by HPLC-EI/MS and HPLC-APCI/MS techniques.

HPLC coupled with normal phase electron ionisation (EI) and atmospheric pressure chemical ionisation (APCI)/ mass spectrometry methods has been applied to identify 17 known neutral limonoid aglycones from Citrus sources. The HPLC-MS data from the known limonoids provided chromatographic characteristics, APCI-derived molecular weight data and EI fragmentation data for each limonoid. EI fragmentation patterns for the limonoids were correlated with structural characteristics. The EI fragmentation patterns coupled with APCI-derived molecular weights were utilised as a potential method by which to discern the structural character of unknown citrus limonoids.

Chromatography, High Pressure Liquid↗

Uncovering hidden protein modifications with native top-down mass spectrometry.

Protein modifications drive dynamic cellular processes by modulating biomolecular interactions, yet capturing these modifications within their native structural context remains a significant challenge. Native top-down mass spectrometry promises to preserve the critical link between modifications and interactions. However, current methods often fail to detect uncharacterized or low-abundance modifications, limiting insights into proteoform diversity. To address this gap, we introduce precise and accurate Identification Of Native proteoforms (precisION), an interactive end-to-end software package that leverages a robust, data-driven fragment-level open search to detect, localize and quantify 'hidden' modifications within intact protein complexes. Applying precisION to four therapeutically relevant targets-PDE6, ACE2, osteopontin (SPP1) and a GABA transporter (GAT1)-we discover undocumented phosphorylation, glycosylation and lipidation, and resolve previously uninterpretable density in an electron cryo-microscopy map of GAT1. As an open-source software package, precisION offers an intuitive means for interpreting complex protein fragmentation data. This tool will empower the community to unlock the potential of native top-down mass spectrometry, advancing integrative structural biology, molecular pathology and drug development.

Mass Spectrometry↗