PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data fragmentation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Data structures for DNA sequence manipulation.

Two data structures designated Fragment and Construct are described. The Fragment data structure defines a continuous nucleic acid sequence from a unique genetic origin. The Construct defines a continuous sequence composed of sequences from multiple genetic origins. These data structures are manipulated by a set of software tools to simulate the construction of mosaic recombinant DNA molecules. They are also used as an interface between sequence data banks and analytical programs.

Base Sequence↗

Multi-isotopic modelling of mass spectra: a procedure for verification of the fragmentation hypothesis for the organometallic and coordination compounds.

Interpretations of mass spectra usually depend on explanations of the molecular structure by determination of the fragmentation ions structures. Therefore, identification of fragmentation ions formulas is an initial step of such investigations. It seems that multi-isotopic modelling of all ions proposed in fragmentation data should lead to the theoretical mass spectrum. Good adjustment of experimental and predicted spectra proves the validity of the fragmentation table, whose contents are recognised as fragmentation hypothesis. This means that multi-isotopic modelling is a helpful tool for verification of proposed fragmentation data. Applications of the method are presented for 1,1',2,2',3,3'-hexachloro-ferrocene, C10H4Cl6Fe, and bis(dibutyldithiophosphate)-zinc(II), [(C4H9O)2PS2]2Zn.

Journal Article↗

Using fragment chemistry data mining and probabilistic neural networks in screening chemicals for acute toxicity to the fathead minnow.

The paper is illustrating how the general data mining methodology may be adapted to provide solutions to the problem of high throughput virtual screening of organic chemicals for possible acute toxicity to the fathead minnow fish. The present approach involves mining fragment information from chemical structures and is using probabilistic neural networks to model the relationship between structure and toxicity. Probabilistic neural networks implement a special class of multivariate non-linear Bayesian statistical models. The mathematical principles supporting their use for value prediction purposes are clarified and their peculiarities discussed. As part of the research phase of the data mining process, a dataset consisting of 800 structures and associated fathead minnow (Pimephales promelas) 96-h LC50 acute toxicity endpoint information is used for both the purpose of identifying an advantageous combination of fragment descriptors and for training the neural networks. As a result, two powerful models are generated. Model 1 implements the basic PNN with Gaussian kernel (statistical corrections included) while Model 2 implements the PNN with Gaussian kernel and separated variables. External validation is performed using a separate dataset consisting of 86 structures and associated toxicity information. Both learning and generalization capabilities of the two models are investigated and their limitations discussed.

Animals↗

Prospects for estimating nucleotide divergence with RAPDs.

The technique of random amplification of polymorphic DNA (RAPD), which is simply polymerase chain reaction (PCR) amplification of genomic DNA by a single short oligonucleotide primer, produces complex patterns of anonymous polymorphic DNA fragments. The information provided by these banding patterns has proved to be of great utility for mapping and for verification of identity of bacterial strains. Here we consider whether the degree of similarity of the banding patterns can be used to estimate nucleotide diversity and nucleotide divergence. With haploid data, fragments generated by RAPD-PCR can be treated in a fashion very similar to that for restriction-fragment data. Amplification of diploid samples, on the other hand, requires consideration of the fact that presence of a band is dominant to absence of the band. After describing a method for estimating nucleotide divergence on the basis of diploid samples, we summarize the restrictions and criteria that must be met when RAPD data are used for estimating population genetic parameters.

Computer Simulation↗

High-throughput direct-infusion ion trap mass spectrometry: a new method for metabolomics.

A fast method was developed to directly infuse raw plant extracts into a linear ion trap mass spectrometer, using the ion trap to isolate and fragment as many ions as possible from the extract. The full mass spectra can be analysed by multivariate statistics to determine discriminating ions, and the fragmentation data allows rapid classification or identification of these ions. The methodology was used to screen a wide range of strains of endophytic fungi in perennial ryegrass seeds for differences in metabolic profiles. The results show that this newly developed methodology is able to determine discriminating ions that can be present in very low concentrations. It also yielded sufficient fragmentation data to classify or identify the discriminating ions.

Ascomycota↗

New SRM data dependent exclusion (MS)n measurement for structural determination of drug metabolites using LC/ESI/Ion trap MS.

New SRM (selected reaction monitoring) data dependent exclusion (MS)(n) measurement makes it possible to obtain MS(3) fragmentation data for all MS(2) fragments, useful for structural determination of drug metabolites using ESI ion trap. MS(2) fragments are produced by cleavage of all protonated molecules at the lone electron pairs of heteroatoms or the pi electrons of double and triple bonds, benzene rings and hetero-rings of drugs. Usually, data dependent MS(3) measurement cleaves only MS(2) fragment of highest intensity, that normally does not contain important metabolic sites. Fragmentation data from all parts of drug metabolites is required to determine structure. In addition to the usual basic measurement of protonated molecules and (MS)(n) fragmentation of drug metabolites, we demonstrate the use of SRM data dependent (MS)(n) measurement, plus new SRM data dependent exclusion (MS)(n) measurement for structural determination of metabolites.

Journal Article↗

Differential molecular connectivity in data-base fragment searching.

A general scheme is described in which molecular fragments are coded from molecular connectivity values. Specifically a fragment is described by the difference between a simple connectivity index of a certain order and the valence connectivity index of the same order. This numerical value is then used to search for that particular fragment among stored fragment values associated with a molecular connectivity calculation. Examples illustrate the method.

Chemical Phenomena↗

An analytical and structural database provides a strategy for sequencing O-glycans from microgram quantities of glycoproteins.

A sensitive, rapid, quantitative strategy has been developed for O-glycan analysis. A structural database has been constructed that currently contains analytical parameters for more than 50 glycans, enabling identification of O-glycans at the subpicomole level. The database contains the structure, molecular weight, and both normal and reversed-phase HPLC elution positions for each glycan. These observed parameters reflect the mass, three-dimensional shape, and hydrophobicity of the glycans and, therefore, provide information relating to linkage and arm specificity as well as monosaccharide composition. Initially the database was constructed by analyzing glycans released by mild hydrazinolysis from bovine serum fetuin, synthetic glycopeptides, human glycophorin A, and serum IgA1. The structures of the fluorescently labeled sugars were determined from a combination of HPLC data, mass spectrometric composition and mass fragmentation data, and exoglycosidase digestions. This approach was then applied to human neutrophil gelatinase B and secretory IgA, where 18 and 25 O-glycans were identified, respectively, and the parameters of these glycans were added to the database. This approach provides a basis for the analysis of subpicomole quantities of O-glycans from normal levels of natural glycoproteins.

Animals↗

AMASS: a structured pattern matching approach to shotgun sequence assembly.

In this paper, we propose an efficient, reliable shotgun sequence assembly algorithm based on a fingerprinting scheme that is robust to both noise and repetitive sequences in the data, two primary roadblocks to effective whole-genome shotgun sequencing. Our algorithm uses exact matches of short patterns randomly selected from fragment data to identify fragment overlaps, construct an overlap map, and deliver a consensus sequence. We show how statistical clues made explicit in our approach can easily be exploited to correctly assemble results even in the presence of extensive repetitive sequences. Our approach is both accurate and exceptionally fast in practice: e.g., we have correctly assembled the whole Mycoplasma genitalium genome (approximately 580 kbp) is roughly 8 minutes of 64MB 200MHz Pentium Pro CPU time from real shotgun data, where most existing algorithms can be expected to run for several hours to a day on the same data. Moreover, experiments with artificially-shotgunned data prepared from real DNA sequences from a wide range of organisms (including human DNA) and containing complex repeating regions demonstrate our algorithm's robustness to input noise and the presence of repetitive sequences. For example, we have correctly assembled a 238-kbp human DNA sequence in less than 3 min of 64-MB 200-MHz Pentium Pro CPU time.

Algorithms↗

Proteomic analysis of phytopathogenic fungus Botrytis cinerea as a potential tool for identifying pathogenicity factors, therapeutic targets and for basic research.

Botrytis cinerea is a phytopathogenic fungus causing disease in a substantial number of economically important crops. In an attempt to identify putative fungal virulence factors, the two-dimensional gel electrophoresis (2-DE) protein profile from two B. cinerea strains differing in virulence and toxin production were compared. Protein extracts from fungal mycelium obtained by tissue homogenization were analyzed. The mycelial 2-DE protein profile revealed the existence of qualitative and quantitative differences between the analyzed strains. The lack of genomic data from B. cinerea required the use of peptide fragmentation data from MALDI-TOF/TOF and ESI ion trap for protein identification, resulting in the identification of 27 protein spots. A significant number of spots were identified as malate dehydrogenase (MDH) and glyceraldehyde-3-phosphate dehydrogenase (GAPDH). The different expression patterns revealed by some of the identified proteins could be ascribed to differences in virulence between strains. Our results indicate that proteomic analysis are becoming an important tool to be used as a starting point for identifying new pathogenicity factors, therapeutic targets and for basic research on this plant pathogen in the postgenomic era.

Botrytis↗

[Electrophoretic fractions and NH2-terminal amino acids of high-molecular tryptic fragments of fibrinogen and fibrin].

The electrophoretic and NH2-terminal analyses were performed for D and E fragments obtained by tryptic digestion of bovine fibrinogen and fibrin under various conditions. The preparations of fragment D were heterogeneous, they were separated by polyacrylamide gel electrophoresis into a number of electrophoretic components with molecular weight ranging from 65 000 to 85 000. NH2-terminal analysis revealed from 6 to 8 NH2-terminal amino acids: Ser, Gly, Thr, Asp, Gly, Lys, Gln, Asn. Their composition and quantitative ratios were found to vary depending on the conditions of the fragment D production. The electrophoregrams showed that with Ca++ in the medium the enzymatic splitting of fibrinogen was limited. Fragment D obtained from fibrinogen without Ca++ was electrophoretically rather similar to that obtained from fibrinogen at Ca++ optimal concentration. Taking into consideration a very high specific anti-coagulational activity of these two fragment D preparations, one may conclude that both the polymerized state of protein molecules and the presence of Ca++ stabilize the specific molecular structure, that favorus the preservation of specific inhibitory activity in fragment D. According to the NH2-terminal analysis data, fragment E derived from fibrinogen hydrolyzed in the presence of Ca++ optimal concentration is also heterogeneous. The following amino acids were found: Tyr, Lys, His (main ones) and Gly, Ser, Thr, Val (minor ones). As to molecular weight determined by electrophoresis, fragment E appears to be homogeneous.

Animals↗

Phylogenetic signal in AFLP data sets.

AFLP markers provide a potential source of phylogenetic information for molecular systematic studies. However, there are properties of restriction fragment data that limit phylogenetic interpretation of AFLPs. These are (a) possible nonindependence of fragments, (b) problems of homology assignment of fragments, (c) asymmetry in the probability of losing and gaining fragments, and (d) problems in distinguishing heterozygote from homozygote bands. In the present study, AFLP data sets of Lactuca s.l. were examined for the presence of phylogenetic signal. An indication of this signal was provided by carrying out tree length distribution skewness (g1) tests, permutation tail probability (PTP) tests, and relative apparent synapomorphy analysis (RASA). A measure of the support for internal branches in the optimal parsimony tree (MPT) was made using bootstrap, jackknife, and decay analysis. Finally, the extent of congruence in MPTs for AFLP and internal transcribed spacer (ITS)-1 data sets for the same taxa was made using the partition homogeneity test (PHT) and the Templeton test. These analytical studies suggested the presence of phylogenetic signal in the AFLP data sets, although some incongruence was found between AFLP and ITS MPTs. An extensive literature survey undertaken indicated that authors report a general congruence of AFLP and ITS tree topologies across a wide range of taxonomic groups, suggesting that the present results and conclusions have a general bearing. In these earlier studies and those for Lactuca s.l., AFLP markers have been found to be informative at somewhat lower taxonomic levels than ITS sequences. Tentative estimates are suggested for the levels of ITS sequence divergence over which AFLP profiles are likely to be phylogenetically informative.

Classification↗

PMP1 18-38, a yeast plasma membrane protein fragment, binds phosphatidylserine from bilayer mixtures with phosphatidylcholine: a (2)H-NMR study.

PMP1 is a 38-residue plasma membrane protein of the yeast Saccharomyces cerevisiae that regulates the activity of the H(+)-ATPase. The cytoplasmic domain conformation results in a specific interfacial distribution of five basic side chains, thought to strongly interact with anionic phospholipids. We have used the PMP1 18-38 fragment to carry out a deuterium nuclear magnetic resonance ((2)H-NMR) study for investigating the interactions between the PMP1 cytoplasmic domain and phosphatidylserines. For this purpose, mixed bilayers of 1-palmitoyl, 2-oleoyl-sn-glycero-3-phosphocholine (POPC) and 1-palmitoyl, 2-oleoyl-sn-glycero-3-phosphoserine (POPS) were used as model membranes (POPC/POPS 5:1, m/m). Spectra of headgroup- and chain-deuterated POPC and POPS phospholipids, POPC-d4, POPC-d31, POPS-d3, and POPS-d31, were recorded at different temperatures and for various concentrations of the PMP1 fragment. Data obtained from POPS deuterons revealed the formation of specific peptide-POPS complexes giving rise to a slow exchange between free and bound PS lipids, scarcely observed in solid-state NMR studies of lipid-peptide/protein interactions. The stoichiometry of the complex (8 POPS per peptide) was determined and its significance is discussed. The data obtained with headgroup-deuterated POPC were rationalized with a model that integrates the electrostatic perturbation induced by the cationic peptide on the negatively charged membrane interface, and a "spacer" effect due to the intercalation of POPS/PMP1f complexes between choline headgroups.

Amino Acid Sequence↗

Synthesis of novel anti-inflammatory peptides derived from the amino-acid sequence of the bioactive protein SV-IV.

SV-IV is a basic, thermostable, secretory protein of low Mr (9758) that is synthesized by rat seminal vesicle (SV) epithelium under strict androgen transcriptional control. This protein is of obvious pharmacological interest because it has potent nonspecies-specific immunomodulatory, anti-inflammatory, and pro-coagulant activities. In evaluating the clinical relevance and the possible use in medicine of SV-IV, we became interested in the study of its structure-function relationships and aimed to identify in its polypeptide chain specific peptide fragments possessing the marked anti-inflammatory properties of the protein not associated with other biological activities (pro-coagulation and immunomodulation) typical of this molecule. By using two different experimental approaches (the fragmentation of the protein into peptide derivatives by chemical methods and the organic synthesis on solid phase of selected peptide fragments), data were obtained showing that in this protein: (a) the immunomodulatory activity is related to the structural integrity of the whole molecule; (b) the anti-inflammatory activity is located in the N-terminal region of the molecule, the 8-16 peptide fragment being the most active; (c) the identified anti-inflammatory peptide derivatives do not seem to possess pro-coagulant activity, even though this particular function has been located in the 1-70 segment of the molecule.

Amino Acid Sequence↗

Fragmentation reactions of alkylphenylammonium ions

The fragmentation reactions of a variety of alkylphenylammonium ions, C(6)H(5)NH(3 -n)R(n)(+) (n >/= 1, R = CH(3), C(2)H(5), i-C(3)H(7), n-C(4)H(9)) were studied by energy-resolved mass spectrometry. Ionization was by fast atom bombardment (FAB) or electrospray ionization. Energy-resolved fragmentation data were obtained by low-energy collision-induced dissociation (CID) in the quadrupole cell of a hybrid sector/quadrupole instrument following FAB ionization and by cone-voltage CID in the interface region of the electrospray/quadrupole instrument. A comparison of the two methods of obtaining energy-resolved data showed that very similar results are obtained by the two methods. The fragmentation reactions of the alkylphenylammonium ions are rationalized in terms of competitive formation of an [R(+)-NC(6)H(5)H(3-n)R(n-1)] complex or a [C(6)H(5)H(3-n)R(n-1)N(+.)-(.)R] complex. The former complex fragments by internal proton transfer to yield C(6)H(5)H(3 -n)R(n -1)NH(+) and [R -H] whereas the latter complex fragments to form C(6)H(5)H(3 -n)R(n -1)N(+) and an alkyl radical. Alkane elimination, which is very prominent for tetraalkylammonium ions, most likely involves sequential elimination of an alkyl radical and either an H atom or an alkyl radical for the phenyl-substituted ammonium ions. Copyright 1999 John Wiley & Sons, Ltd.

Journal Article↗

Laser-induced dissociation/high-energy collision-induced dissociation fragmentation using MALDI-TOF/TOF-MS instrumentation for the analysis of neutral and acidic oligosaccharides.

Producing detailed mass spectrometric fragmentation data of native oligosaccharides for the purpose of basic structure elucidation has become a readily accessible tool since the availability of enhanced technical equipment. In this report, high-energy collision-induced dissociation (heCID) in combination with MALDI-TOF/TOF technology for analysis of native neutral and acidic oligosaccharides is described. By providing complementary data, heCID-MALDI-TOF/TOF adds a variety of valuable cross-ring fragmentation information to the information of glycosidic fragmentation obtained preferably by laser-induced dissociation (LID). We examined parameters influencing fragmentation behavior of both-acidic and neutral-compounds. Results show a dependency of the fragmentation pattern for the employed matrix as well as the laser intensity provided for the ionization of the analytes and the complexity of the analytes. Due to instrument-specific settings, protonated glycosidic ion series within spectra of sodiated compounds could also be identified. Furthermore, acquired spectra could be readily used to identify compounds by comparison to existing glycan databases such as GlycoSuiteDB and GlycosciencesDB. The results show a better scoring of heCID data sets in comparison to LID-derived data. heCID-MALDI-TOF/TOF analysis in combination with database search algorithms is demonstrated to be suitable for an initial identification/classification of carbohydrates.

Carbohydrate Sequence↗

Characterisation of global protein expression by two-dimensional electrophoresis and mass spectrometry: proteomics of Toxoplasma gondii.

The development of tools for the analysis of global gene expression is vital for the optimal exploitation of the data on parasite genomes that are now being generated in abundance. Recent advances in two-dimensional electrophoresis (2-DE), mass spectrometry and bioinformatics have greatly enhanced the possibilities for mapping and characterisation of protein populations. We have employed these developments in a proteomics approach for the analysis of proteins expressed in the tachyzoite stage of Toxoplasma gondii. Over 1000 polypeptides were reproducibly separated by high-resolution 2-DE using the pH ranges 4-7 and 6-11. Further separations using narrow range gels suggest that at least 3000-4000 polypeptides should be resolvable by 2-DE using multiple single pH unit gels. Mass spectrometry was used to characterise a variety of protein spots on the 2-DE gels. Peptide mass fingerprints, acquired by matrix-assisted laser desorption/ionisation-(MALDI) mass spectrometry, enabled unambiguous protein identifications to be made where full gene sequence information was available. However, interpretation of peptide mass fingerprint data using the T. gondii expressed sequence tag (EST) database was less reliable. Peptide fragmentation data, acquired by post-source decay mass spectrometry, proved a more successful strategy for the putative identification of proteins using the T. gondii EST database and protein databases from other organisms. In some instances, several protein spots appeared to be encoded by the same gene, indicating that post-translational modification and/or alternative splicing events may be a common feature of functional gene expression in T. gondii. The data demonstrate that proteomic analyses are now viable for T. gondii and other protozoa for which there are good EST databases, even in the absence of complete genome sequence. Moreover, proteomics is of great value in interpreting and annotating EST databases.

Animals↗