PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “data fragmentation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Calcium binding, hydroxylation, and glycosylation of the precursor epidermal growth factor-like domains of fibrillin-1, the Marfan gene protein.

The extracellular matrix protein fibrillin-1 is a major component of elastic microfibrils, which are complex assemblies of several proteins and are found in most connective tissues, frequently associated with elastin. Fibrillin-1 contains 43 precursor epidermal growth factor-like (pEGF) domains that have a consensus sequence for calcium binding. The calcium binding potential of a fibrillin-1 pepsin fragment (PF2) was quantitatively analyzed using microvolume equilibrium dialysis. Peptide sequence data and pepsin fragment size determination indicate that PF2 contains seven pEGF domains, each with the calcium binding consensus sequence. Scatchard plot analysis of the calcium binding data shows that PF2 has six to seven high affinity binding sites with a Kd = 250 microM at pH 7.5. There is a second overlapping consensus sequence in the pEGF domains for beta-hydroxylation of a specific Asp/Asn residue. Five partially hydroxylated Asn residues have been identified by protein sequence analysis of fibrillin-1 fragments. This is the first demonstration of this modification in a connective tissue protein. The calcium binding consensus sequence also contains a conserved Ser residue with an apparently novel modification, which causes the Ser residue to behave like an Asp residue during protein sequencing. Marfan syndrome, a heritable disorder of connective tissue, is known to be associated with mutations in the FBN1 gene. Most of these mutations have been found in pEGF domains, frequently substituting Cys for another amino acid, destroying the pEGF motif secondary structure along with its calcium binding potential. Other mutations cause the substitution of single amino acids in the calcium binding consensus sequence, which could affect calcium binding but also the hydroxylation of Asp/Asn residues or the modification of Ser residues.

Amino Acid Sequence↗

[Nucleotide sequence analysis of the minimal replicon of the Streptomyces plasmid pSGL1].

The high-copy-number plasmid pSGL1(7.4 kb) was isolated from Streptomyces globisporus. Deletion experiments showed its minimal replicon is located on the 2.0 kb Sau3 AI fragment. This fragment was subcloned. DNA sequence data analysis showed this fragment is a new sequence. Only an open reading frame with high coding probability is located on the minimal replicon. The deduced protein contains motifs characteristic of replicase for rolling-circle replication.

Amino Acid Sequence↗

Comparative evaluation of the predictive power of calculation procedures for molecular lipophilicity.

The predictive power of four calculation procedures for molecular lipophilicity is checked by comparing with experimental data (log P and chromatographical RMw) taken from the literature. Two sets of test compounds are used: the first comprises simple organic molecules and the second consists of more complicated drug molecules. Our comparative evaluation leads us to conclude that the predictive power is significantly better for not too complicated organic molecules than for drugs with complicated structural pattern. The four investigated calculation procedures should be arranged in two groups with significantly differing predictive power: (a) Rekker and Hansch/Leo and (b) Ghose/Crippen and Suzuki/Kudo. This conclusion is based on a statistical control using log P and RMw as the independent parameters. Correlations have in common: (1) slopes in correlations with calculated data based on fragmental methods are not significantly different from 1; calculations with data from atom-based procedures show up in most cases with slopes below 1. (2) The accompanying overall statistics underline the superiority of the fragmental methods. We think that all four tested calculation procedures have their own restrictions; for future development we would advise a thorough reconsideration of structural effects not fully (or even not at all) incorporated in the data sets. Special attention will have to be paid to the conformational aspects of lipophilic behavior.

Models, Chemical↗

Presenilin 1 cleavage is a universal event in human organs.

A panel of antibodies raised against various regions of human presenilin 1(PS1)--the amino-terminal domain, the domain between the transmembrane domains 1 and 2, the cleavage-site, loop domains, or carboxyl-terminal domain--was prepared to analyze PS1 in human tissues. We observed the predominance of two fragments (28-kDa NH2 and 18-kDa COOH fragments) in various tissues, including cerebral cortices. In addition to these two fragments, we found a previously unidentified amino-terminal fragment of PS1 with Mr 14 kDa in the lungs, spleen, pancreas, and testes. Using a sensitive ELISA for PS1, we measured the amount of PS1 species in tissues and found high contents of PS1 fragment in the testes. Our data show that common and unique processing pathways of PS1 occur in a tissue-dependent manner. It is likely that cleavage at the loop structure of PS1 to produce a functional form is a common event in human organs.

Alzheimer Disease↗

Increased protein identification capabilities through novel tandem MS calibration strategies.

High mass measurement accuracy is critical for confident protein identification and characterization in proteomics research. Fourier transform ion cyclotron resonance (FTICR) mass spectrometry is a unique technique which can provide unparalleled mass accuracy and resolving power. However, the mass measurement accuracy of FTICR-MS can be affected by space charge effects. Here, we present a novel internal calibrant-free calibration method that corrects for space charge-induced frequency shifts in FTICR fragment spectra called Calibration Optimization on Fragment Ions (COFI). This new strategy utilizes the information from fixed mass differences between two neighboring peptide fragment ions (such as y(1) and y(2)) to correct the frequency shift after data collection. COFI has been successfully applied to LC-FTICR fragmentation data. Mascot MS/MS ion search data demonstrate that most of the fragments from BSA tryptic digested peptides can be identified using a much lower mass tolerance window after applying COFI to LC-FTICR-MS/MS of BSA tryptic digest. Furthermore, COFI has been used for multiplexed LC-CID-FTICR-MS which is an attractive technique because of its increased duty cycle and dynamic range. After the application of COFI to a multiplexed LC-CID-FTICR-MS of BSA tryptic digest, we achieved an average measured mass accuracy of 2.49 ppm for all the identified BSA fragments.

Algorithms↗

Differential display of messenger RNA expressed in bronchoalveolar lavage cells in pulmonary sarcoidosis patients.

Sarcoidosis is a systemic granulomatous disease of unknown origin. To clarify its pathogenesis, we searched for known or unknown genes which are specifically expressed in sarcoidosis. Bronchoalveolar lavage (BAL) cells from 18 patients with sarcoidosis and 8 patients with various lung diseases were analyzed by differential display method. mRNA was extracted from BAL cells and reverse transcribed with 12 kinds of anchored primer, which theoretically cover all mRNAs, followed by polymerase chain reaction (PCR) with the anchored primer and a 10-mer arbitrary primer. PCR products were displayed on a polyacrylamide gel and fragments showing characteristic alterations in intensity between sarcoidosis and other patients were extracted, sequenced, and compared against Genbank and EMBL DNA data bases. One fragment was detected with specifically increased intensity and another disappeared in patients with sarcoidosis. These fragments were likely derived from unknown genes. CD44 and tumor necrosis factor (TNF-alpha) cDNA sequences were also detected as fragments commonly expressed in sarcoidosis. The cloned fragments with specifically increased or decreased intensity in sarcoidosis may provide important information on the pathogenesis of sarcoidosis, and the display pattern implies the potential usefulness of this method as a tool for diagnosis of the disease.

Adolescent↗

Anisotropic rotation in nucleic acid fragments: significance for determination of structures from NMR data.

Proton-proton relaxation rate constants depend on the angle between the internuclear vector and the principal axis of rotation in symmetric top molecules. It is possible to determine to rotational correlation times of the equivalent ellipsoid for DNA fragments from a knowledge of the axial ratio and the cross-relaxation rate constant for the cytosine H6-H5 vectors. The cross-relaxation rate constants for the cytosine H6-H5 vectors have been measured in the 14-base-pair sequence dGCTGTTGACAATTA.dTAATTGTCAACAGC at four temperatures. The results, along with literature data for DNA fragments ranging from 6 to 20 base pairs can be accounted for by a simple hydrodynamic equation based on the formalism of Woessner (1962). The measured cross-relaxation rate constant is independent of position in the sequence and is consistent with the absence of large amplitude internal motions on the Larmor time scale. All the data can be described by a simple hydrodynamic model, which accounts for the rotational anisotropy of the DNA fragments and allows the correlation time for end-over-end tumbling to be determined if the approximate rise per base pair is known. This is the correlation time that dominates the spectral density functions for internucleotide vectors and is significantly different from that calculated for a sphere of the same hydrodynamic volume for fragments containing more than about 14 base pairs. This method therefore allows NOE intensities used for structure calculation of nucleic acids to be treated more rigorously.

Base Sequence↗

Molecular cloning of the c locus of Zea mays: a locus regulating the anthocyanin pathway.

The c locus of Zea mays, involved in the regulation of anthocyanin biosynthesis, has been cloned by transposon tagging. A clone (# 18En) containing a full size En1 element was initially isolated from the En element-induced mutable allele c-m668655. Sequences of clone # 18En flanking the En1 element were used to clone other c mutants, whose structure was predicted genetically. Clone #23En (isolated from c-m668613) contained a full size En1 element, clone #3Ds (isolated from c-m2) a Ds element and clone # 5 (isolated from c+) had no element on the cloned fragment. From these data we conclude that the clones obtained contain at least part of the c locus. Preliminary data on transcript analysis using a 1-kb DNA fragment from wild-type clone # 5 showed that at least three transcripts are encoded by that part of the locus, indicating that c is a complex locus.

Alleles↗

Amino acid sequence analysis of fragments generated by partial proteolysis from large simian virus 40 tumor antigen.

Large simian virus 40 tumor antigen was bound as immune complex to protein A-Sepharose and then subjected to limited proteolysis which yielded several discrete fragments. Primary structures near the cleavage sites were determined by radiosequencing techniques. Experimental data for five fragments matched an amino acid sequence predicted from a nucleotide sequence at 0.51 map unit of the viral genome. We have thus identified the reading frame of translation beyond the intervening sequence at 0.60 to 0.53 map units. A cleavage map of tumor antigen was established on the basis of the sequence data and of the apparent molecular weights of the fragments. The bond most susceptible to cleavage by trypsin was between arginine-130 and lysine-131 in a cluster of five basis amino acids. Other cleavage sites were located in the COOH-terminal half of tumor antigen. Each fragment was analyzed by complete tryptic proteolysis and peptide mapping on an ion exchange column. Peaks occurring in the peptide map of large tumor antigen could thus be assigned to different segments of the protein. Two specific regions of tumor antigen were shown to be phosphorylated.

Amino Acid Sequence↗

Providing easy access to distributed medical data.

Many hospitals are fragmented along departmental boundaries, leading to islands of information about patients. This makes data integration difficult, and therefore can increase hospital costs and reduce patient care. This paper presents an architecture to provide uniform and transparent access to computerized data and functions available in this kind of heterogeneous computer environment.

Computer Communication Networks↗

Role of a bulged A residue in a specific RNA-protein interaction.

The translational operator of the R17 replicase gene contains a bulged A residue that is essential for the specific binding to R17 coat protein. A large number of operator variants have been synthesized to more precisely examine the role of the bulged A residue on this specific protein-RNA interaction. By use of RNA ligase and transcription of synthetic DNA templates by T7 RNA polymerase, 14 different nucleotides were introduced to the bulged A position of three different coat protein binding fragments. The affinity between coat protein and each fragment was determined by a nitrocellulose filter binding assay. The data indicate that while functional groups on N1, C2, C6, N7, and 2'OH of the bulged A can be substituted without greatly changing protein binding, bulky substituents cannot be tolerated at these positions. Data from additional fragments that have base-pair changes adjacent to the bulged A suggest that the propensity of the bulged A to intercalate into the helix can affect protein binding.

Base Sequence↗

Out of Anatolia: longitudinal gradients in genetic diversity support an eastern origin for a circum-Mediterranean oak gallwasp Andricus quercustozae.

Many studies have addressed the latitudinal gradients in intraspecific genetic diversity of European taxa generated during postglacial range expansion from southern refugia. Although Asia Minor is known to be a centre of diversity for many taxa, relatively few studies have considered its potential role as a Pleistocene refugium or a potential source for more ancient westward range expansion into Europe. Here we address these issues for an oak gallwasp, Andricus quercustozae (Hymenoptera: Cynipidae), whose distribution extends from Morocco along the northern coast of the Mediterranean through Turkey to Iran. We use sequence data for a fragment of the mitochondrial gene cytochrome b and allele frequency data for 12 polymorphic allozyme loci to answer the following questions: (1). which regions represent current centres of genetic diversity for A. quercustozae? Do eastern populations represent one refuge or several discrete glacial refugia? (2). Can we infer the timescale and sequence of the colonization processes linking current centres of diversity? Our results suggest that A. quercustozae was present in five distinct refugia (Iberia, Italy, the Balkans, southwestern Turkey and northeastern Turkey) with recent genetic exchange between Italy and Hungary. Genetic diversity is greatest in the Turkish refugia, suggesting that European populations are either (a). derived from Asia Minor, or (b). subject to more frequent population bottlenecks. Although Iberian populations show the lowest diversity for putatively selectively neutral markers, they have colonized a new oak host and represent a genetically and biologically discrete entity within the species.

Animals↗

Polymerase chain reaction-restriction fragment length polymorphism analysis: a simple method for species identification in food.

The polymerase chain reaction (PCR) technique was applied to meat species identification in marinated and heat-treated or fermented products and to the differentiation of closely related species. DNA was isolated from meat samples by using a DNA-binding resin and was subjected to PCR analysis. Primers used were complementary to conserved areas of the vertebrate mitochondrial cytochrome b (cytb) gene and yielded a 359 base-pair (bp) fragment, including a variable 307 bp region. Restriction endonuclease analysis based on sequence data of those fragments was used for differentiation among species. Restriction fragment length polymorphisms (RFLPs) were detected when pig, cattle, wild boar, buffalo, sheep, goat, horse, chicken, and turkey amplicons were cut with AluI, RsaI, TaqI, and HinfI. Analysis of sausages indicates the applicability of this approach to food products containing meat from 3 different species. The PCR-RFLP analytical method detected pork in heated meat mixtures with beef at levels below 1%, and the method was confirmed with porcine- and bovine-specific PCR assays by amplifying fragments of their growth hormone genes. Inter- and intraspecific differences of more than 22 animal species with nearly unknown cytb DNA sequences, including hoofed mammals (ungulates), and poultry were determined with PCR-RFLP typing by using 20 different endonucleases. This typing method allowed the discrimination of game meats, including stag, roe deer, chamois, moose, reindeer, kangaroo, springbok, and other antelopes in marinated and heat-treated products.

Animals↗

Primary and tertiary structure studies of p-hydroxybenzoate hydroxylase from Pseudomonas fluorescens. Isolation and alignment of the CNBr peptides; interactions of the protein with flavin adenine dinucleotide.

p-Hydroxybenzoate hydroxylase from Pseudomonas fluorescens contains six methionine residues, one of which is N-terminal. After CNBr cleavage five peptides, ranging from 13 to 158 residues in length, and free homoserine were isolated and purified by repeated gel filtration. The alignment of the CNBr fragments was deduced from a 0.25-nm electron density map and sequence data. The isolated fragments account for the entire polypeptide chain. The amino acid sequence of the N-terminal quarter of the polypeptide chain was determined. The X-ray results together with the sequence data yielded details of the binding of FAD. The AMP moiety was bound to a beta alpha beta unit resembling that found in the dehydrogenases. Hydrogen bonds were present between the protein and the ribityl residue and the isoalloxazine ring. Furthermore, a homology was found between the N-terminal amino acid sequence of p-hydroxybenzoate hydroxylase and another enzyme containing FAD, viz. D-amino acid oxidase. This finding suggests the presence of a mononucleotide binding fold at the N terminus of the latter.

4-Hydroxybenzoate-3-Monooxygenase↗

A new hierarchical parallelization scheme: generalized distributed data interface (GDDI), and an application to the fragment molecular orbital method (FMO).

A two-level hierarchical scheme, generalized distributed data interface (GDDI), implemented into GAMESS is presented. Parallelization is accomplished first at the upper level by assigning computational tasks to groups. Then each group does parallelization at the lower level, by dividing its task into smaller work loads. The types of computations that can be used with this scheme are limited to those for which nearly independent tasks and subtasks can be assigned. Typical examples implemented, tested, and analyzed in this work are numeric derivatives and the fragment molecular orbital method (FMO) that is used to compute large molecules quantum mechanically by dividing them into fragments. Numeric derivatives can be used for algorithms based on them, such as geometry optimizations, saddle-point searches, frequency analyses, etc. This new hierarchical scheme is found to be a flexible tool easily utilizing network topology and delivering excellent performance even on slow networks. In one of the typical tests, on 16 nodes the scalability of GDDI is 1.7 times better than that of the standard parallelization scheme DDI and on 128 nodes GDDI is 93 times faster than DDI (on a multihub Fast Ethernet network). FMO delivered scalability of 80-90% on 128 nodes, depending on the molecular system (water clusters and a protein). A numerical gradient calculation for a water cluster achieved a scalability of 70% on 128 nodes. It is expected that GDDI will become a preferred tool on massively parallel computers for appropriate computational tasks.

Journal Article↗

Receptor specificity of influenza viruses from birds and mammals: new data on involvement of the inner fragments of the carbohydrate chain.

We studied receptor-binding properties of influenza virus isolates from birds and mammals using polymeric conjugates of sialooligosaccharides terminated with common Neu5Ac alpha2-3Gal beta fragment but differing by the structure of the inner part of carbohydrate chain. Viruses isolated from distinct avian species differed by their recognition of the inner part of oligosaccharide receptor. Duck viruses displayed high affinity for receptors having beta1-3 rather than beta1-4 linkage between Neu5Ac alpha2-3Gal-disaccharide and penultimate N-acetylhexosamine residue. Fucose and sulfate substituents at this residue had negative and low effect, respectively, on saccharide binding to duck viruses. By contrast, gull viruses preferentially bound to receptors bearing fucose at N-acetylglucosamine residue, whereas chicken and mammalian viruses demonstrated increased affinity for oligosaccharides that harbored sulfo group at position 6 of (beta1-4)-linked GlcNAc. These data suggest that although all avian influenza viruses preferentially bind to Neu5Ac alpha2-3Gal-terminated receptors, the fine receptor specificity of the viruses varies depending on the avian species. Further studies are required to determine whether observed host-dependent differences in the receptor specificity of avian viruses can affect their ability to infect humans.

Animals↗

Sequence of an oligonucleotide derived from the 3' end of each of the four brome mosaic viral RNAs.

A 3'-terminal oligonucleotide fragment, 161 bases long, can be obtained from each of the four brome mosaic virus RNAs by means of nuclease digestion. Like the four intact brome mosaic virus RNAs, each fragment accepts tyrosine in a reaction catalyzed by wheat germ aminoacyl-tRNA synthetase. The complete nucleotide sequence of the RNA 4 fragment has been determined by use of standard radiochemical methods. Comparative data for the fragments from RNAs 1, 2, and 3 show that they have nearly the same sequence as the RNA 4 fragment. The eight bases adjacent to the 3' terminus of the RNA 4 fragment are identical in sequence to the eight terminal bases of tyrosine tRNA from Torula utilis and eleven interior bases are identical in sequence to eleven bases encompassing the anticodon region of tyrosine tRNA from Saccharomyces cerevisiae, T. utilis, and Escherichia coli. Nevertheless, reasonable base-pairing schemes yield, at best, a distorted cloverleaf secondary structure.

Base Sequence↗

Escherichia coli molecular phylogeny using the incongruence length difference test.

Molecular phylogeny of the species Escherichia coli using the E. coli reference (ECOR) collection strains has been hampered by (1) the absence of rooting in the commonly used phenogram obtained from multilocus enzyme electrophoresis (MLEE) data and (2) the existence of recombination events between strains that scramble phylogenetic trees reconstructed from the nucleotide sequences of genes. We attempted to determine the phylogeny for E. coli based on the ECOR strain data by extracting from GenBank the nucleotide sequences of 11 chromosomal structural and 2 plasmid genes for which the Salmonella enterica homologous gene sequences were available. For each of the 13 DNA data sets studied, incongruence with a nonnucleotide whole-genome data set including MLEE, random amplified polymorphic DNA, and rrn restriction fragment length polymorphism data was measured using the incongruence length difference (ILD) test of Farris et al. As previously reported, the incongruence observed between the gnd and plasmid gene data and the whole-genome data was multiple, indicating numerous horizontal transfer and/or recombination events. In five cases, the incongruence detected by the ILD test was punctual, and the donor group was identified. Congruence was not rejected for the remaining data sets. The strains responsible for incongruences with the whole-genome data set were removed, leading to a "prior-agreement" approach, i.e., the determination of a phylogeny for E. coli based on several genes, excluding (1) the genes with multiple incongruences with the whole genome data, (2) the strains responsible for punctual incongruences, and (3) the genes incongruent with each other. The obtained phylogeny shows that the most basal group of E. coli strains is the B2 group rather than the A group, as generally thought. The D group then emerges as the sister group of the rest. Finally, the A and B1 groups are sister groups. Interestingly, the most primitive taxon within E. coli in terms of branching pattern, i.e., the B2 group, includes highly virulent extraintestinal strains with derived characters (extraintestinal virulence determinants) occurring on its own branch.

Enzymes↗