PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reduced representation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Cycloheximide treatment of cotton ovules alters the abundance of specific classes of mRNAs and generates novel ESTs for microarray expression profiling.

Fibres of cotton (Gossypium hirsutum L.) are single elongated epidermal cells that start to develop on the outer surface of cotton ovules on the day of anthesis. Little is known about the control of fibre initiation and development. As a first step towards discovering important genes involved in fibre initiation and development using a genomics approach, we report technical advances aimed at reducing redundancy and increasing coverage for anonymous cDNA microarrays in this study. Cotton ovule cDNA libraries (both normalised and un-normalised) from around the time of fibre initial formation have been prepared and partially characterised by sequencing. Re-association-based normalisation partially reduced library redundancy and increased representation of novel sequences. However, another library generated from in vitro cultured cotton ovules treated with the protein synthesis inhibitor, cycloheximide, showed a significantly altered gene representation including a greater proportion of protein phosphorylation genes, transport genes and transcription factors and a much reduced proportion of protein synthesis genes than were identified in the conventional types of libraries. Over 10,000 expressed sequence tag (EST) clones randomly selected from the three libraries were printed on microarray slides and used to assess gene expression in tissue cultured ovules with and without cycloheximide treatment. The microarray results showed that cycloheximide had a dramatic effect in modifying the pattern of the gene expression in cultured ovules, affecting the same types of genes identified in the preliminary analysis on relative EST abundance in the different ovule cDNA libraries. Cycloheximide clearly provided a simple and useful method for enriching novel gene sequences for genomic studies.

Base Sequence↗

ESPERR: learning strong and weak signals in genomic sequence alignments to identify functional elements.

Genomic sequence signals - such as base composition, presence of particular motifs, or evolutionary constraint - have been used effectively to identify functional elements. However, approaches based only on specific signals known to correlate with function can be quite limiting. When training data are available, application of computational learning algorithms to multispecies alignments has the potential to capture broader and more informative sequence and evolutionary patterns that better characterize a class of elements. However, effective exploitation of patterns in multispecies alignments is impeded by the vast number of possible alignment columns and by a limited understanding of which particular strings of columns may characterize a given class. We have developed a computational method, called ESPERR (evolutionary and sequence pattern extraction through reduced representations), which uses training examples to learn encodings of multispecies alignments into reduced forms tailored for the prediction of chosen classes of functional elements. ESPERR produces a greatly improved Regulatory Potential score, which can discriminate regulatory regions from neutral sites with excellent accuracy ( approximately 94%). This score captures strong signals (GC content and conservation), as well as subtler signals (with small contributions from many different alignment patterns) that characterize the regulatory elements in our training set. ESPERR is also effective for predicting other classes of functional elements, as we show for DNaseI hypersensitive sites and highly conserved regions with developmental enhancer activity. Our software, training data, and genome-wide predictions are available from our Web site (http://www.bx.psu.edu/projects/esperr).

Algorithms↗

Leafing through the genomes of our major crop plants: strategies for capturing unique information.

Crop plants not only have economic significance, but also comprise important botanical models for evolution and development. This is reflected by the recent increase in the percentage of publicly available sequence data that are derived from angiosperms. Further genome sequencing of the major crop plants will offer new learning opportunities, but their large, repetitive, and often polyploid genomes present challenges. Reduced-representation approaches - such as EST sequencing, methyl filtration and Cot-based cloning and sequencing - provide increased efficiency in extracting key information from crop genomes without full-genome sequencing. Combining these methods with phylogenetically stratified sampling to allow comparative genomic approaches has the potential to further accelerate progress in angiosperm genomics.

Crops, Agricultural↗

A computational approach to simplifying the protein folding alphabet.

What is the minimal number of residue types required to form a structured protein? This question is important for understanding protein modeling and design. Recently, an experimental finding by Baker and coworkers suggested a five-residue solution to this problem. We were motivated by their results and by the arguments of Wolynes to study reductions of protein representation based on the concept of mismatch between a reduced interaction matrix and the Miyazawa and Jernigan (MJ) matrix. We find several possible simplified schemes from the relationship of minimized mismatch versus the number of residue types (N = approximately 2-20). As a specific case, an optimal reduction with five types of residues has the same form as the simplified palette of Baker and coworkers. Statistical and kinetic features of a number of sequences are tested. Comparison of results from sequences with 20 residue types and their reduced representations indicates that the reduction by mismatch minimization is successful. For example, sequences with five types of residues have good folding ability and kinetic accessibility in model studies.

Algorithms↗

Conservation of amphipathic conformations in multiple protein structural alignments.

Protein amphipathic conformations, mainly alpha-helices and beta-strands, are believed to play an important role in protein folding, stability and function. The most popular method for characterizing such structures is the hydrophobic moment. We have analyzed the distribution of hydrophobic moment characteristics (peak magnitude, amphipathic indices and characteristic frequency) in a data bank containing several families of distant sequences multiply aligned by structural superposition. Sequence fragments were classified according to alpha-helix, beta-strand, non-alpha and non-beta conformations. This data bank provided an enhanced sample space compared with those previously reported in the literature. Precautions were taken to reduce over-representation of homologous sequences. Approximately 50% of all individual alpha-helices showed a hydrophobic moment peak in the expected position of the periodicity spectrum while only 38% of individual beta-strands fell in the expected range. False positives account for a surprisingly large 14 and 36% of the non-alpha and non-beta samples respectively. Conservation of hydrophobic moment characteristics and mainly the hydrophobic peak position in the expected periodicity range was examined in the multiple alignments of the distant sequences. Helices tend to conserve more frequently their hydrophobic moment than any other conformation and yet only 13% of all helical segments display such conservation in three-quarters or more of the familial sequences; the similar observation for beta-strands was even lower at 9%. Nonetheless, strongly hydrophobic positions within the structural segments were more conserved than expected.

Amino Acid Sequence↗

Prediction of protein--protein interaction sites in heterocomplexes with neural networks.

In this paper we address the problem of extracting features relevant for predicting protein--protein interaction sites from the three-dimensional structures of protein complexes. Our approach is based on information about evolutionary conservation and surface disposition. We implement a neural network based system, which uses a cross validation procedure and allows the correct detection of 73% of the residues involved in protein interactions in a selected database comprising 226 heterodimers. Our analysis confirms that the chemico-physical properties of interacting surfaces are difficult to distinguish from those of the whole protein surface. However neural networks trained with a reduced representation of the interacting patch and sequence profile are sufficient to generalize over the different features of the contact patches and to predict whether a residue in the protein surface is or is not in contact. By using a blind test, we report the prediction of the surface interacting sites of three structural components of the Dnak molecular chaperone system, and find close agreement with previously published experimental results. We propose that the predictor can significantly complement results from structural and functional proteomics.

Animals↗

Spectral analysis of alpha rhythm during Schultz's autogenic training. A tentative approach to rapid visualization.

The records from experimental sessions including autogenic standard exercises and other conscious states, in two long-term trainees, were treated by using a frequency analysis programme. Starting from a synoptic reduced representation of the data and selecting sequences for the building of histograms of averaged energy levels in limited frequency bands, the authors describe the progressive spread of a dominant alpha band all over the scalp during autogenic training.

Adult↗

Representation of DNA sequences in recombinant DNA libraries prepared by restriction enzyme partial digestion.

We present a theoretical study of the fraction of sequences incorporated in a recombinant DNA partial digest library as a function of the size of the library. The fraction incorporated depends on the degree of restriction enzyme partial digestion. If all restriction sites in the target DNA can be cleaved with the same rate, optimum incorporation of sequences is observed when the number average length of the digested DNA equals the desired average length of the cloned insert. Overdigestion severely reduces the fraction of sequences present in a sample of clones. Heterogeneity in restriction enzyme cleavage rates also reduces the fraction incorporated, and underdigestion improves sequence representation in the face of cleavage rate heterogeneity. Practical methods for determining the number average length of partially digested DNAs are also presented.

Base Composition↗

Evolution and similarity evaluation of protein structures in contact map space.

Prediction of fold from amino acid sequence of a protein has been an active area of research in the past few years, but the limited accuracy of existing techniques emphasizes the need to develop newer approaches to tackle this task. In this study, we use contact map prediction as an intermediate step in fold prediction from sequence. Contact map is a reduced graph-theoretic representation of proteins that models the local and global inter-residue contacts in the structure. We start with a population of random contact maps for the protein sequence and "evolve" the population to a "high-feasibility" configuration using a genetic algorithm. A neural network is employed to assess the feasibility of contact maps based on their 4 physically relevant properties. We also introduce 5 parameters, based on algebraic graph theory and physical considerations, that can be used to judge the structural similarity between proteins through contact maps. To predict the fold of a given amino acid sequence, we predict a contact map that will sufficiently approximate the structure of the corresponding protein. Then we assess the similarity of this contact map with the representative contact map of each fold; the fold that corresponds to the closest match is our predicted fold for the input sequence. We have found that our feasibility measure is able to differentiate between feasible and infeasible contact maps. Further, this novel approach is able to predict the folds from sequences significantly better than a random predictor.

Amino Acid Sequence↗

Normalization and subtraction: two approaches to facilitate gene discovery.

Large-scale sequencing of cDNAs randomly picked from libraries has proven to be a very powerful approach to discover (putatively) expressed sequences that, in turn, once mapped, may greatly expedite the process involved in the identification and cloning of human disease genes. However, the integrity of the data and the pace at which novel sequences can be identified depends to a great extent on the cDNA libraries that are used. Because altogether, in a typical cell, the mRNAs of the prevalent and intermediate frequency classes comprise as much as 50-65% of the total mRNA mass, but represent no more than 1000-2000 different mRNAs, redundant identification of mRNAs of these two frequency classes is destined to become overwhelming relatively early in any such random gene discovery programs, thus seriously compromising their cost-effectiveness. With the goal of facilitating such efforts, previously we developed a method to construct directionally cloned normalized cDNA libraries and applied it to generate infant brain (INIB) and fetal liver/spleen (INFLS) libraries, from which a total of 45,192 and 86,088 expressed sequence tags, respectively, have been derived. While improving the representation of the longest cDNAs in our libraries, we developed three additional methods to normalize cDNA libraries and generated over 35 libraries, most of which have been contributed to our integrated Molecular Analysis of Genomes and Their Expression (IMAGE) Consortium and thus distributed widely and used for sequencing and mapping. In an attempt to facilitate the process of gene discovery further, we have also developed a subtractive hybridization approach designed specifically to eliminate (or reduce significantly the representation of) large pools of arrayed and (mostly) sequenced clones from normalized libraries yet to be (or just partly) surveyed. Here we present a detailed description and a comparative analysis of four methods that we developed and used to generate normalize cDNA libraries from human (15), mouse (3), rat (2), as well as the parasite Schistosoma mansoni (1). In addition, we describe the construction and preliminary characterization of a subtracted liver/spleen library (INFLS-SI) that resulted from the elimination (or reduction of representation) of -5000 INFLS-IMAGE clones from the INFLS library.

Adult↗

Adaptation of a fast Fourier transform-based docking algorithm for protein design.

Designing proteins with novel protein/protein binding properties can be achieved by combining the tools that have been developed independently for protein docking and protein design. We describe here the sequence-independent generation of protein dimer orientations by protein docking for use as scaffolds in protein sequence design algorithms. To dock monomers into sequence-independent dimer conformations, we use a reduced representation in which the side chains are approximated by spheres with atomic radii derived from known C2 symmetry-related homodimers. The interfaces of C2-related homodimers are usually more hydrophobic and protein core-like than the interfaces of heterodimers; we parameterize the radii for docking against this feature to capture and recreate the spatial characteristics of a hydrophobic interface. A fast Fourier transform-based geometric recognition algorithm is used for docking the reduced representation protein models. The resulting docking algorithm successfully predicted the wild-type homodimer orientations in 65 out of 121 dimer test cases. The success rate increases to approximately 70% for the subset of molecules with large surface area burial in the interface relative to their chain length. Forty-five of the predictions exhibited less than 1 A C(alpha) RMSD compared to the native X-ray structures. The reduced protein representation therefore appears to be a reasonable approximation and can be used to position protein backbones in plausible orientations for homodimer design.

Algorithms↗

Differential methylation of genes and retrotransposons facilitates shotgun sequencing of the maize genome.

The genomes of higher plants and animals are highly differentiated, and are composed of a relatively small number of genes and a large fraction of repetitive DNA. The bulk of this repetitive DNA constitutes transposable, and especially retrotransposable, elements. It has been hypothesized that most of these elements are heavily methylated relative to genes, but the evidence for this is controversial. We show here that repeat sequences in maize are largely excluded from genomic shotgun libraries by the selection of an appropriate host strain because of their sensitivity to bacterial restriction-modification systems. In contrast, unmethylated genic regions are preserved in these genetically filtered libraries if the insert size is less than the average size of genes. The representation of unique maize sequences not found in plant reference genomes is also greatly enriched. This demonstrates that repeats, and not genes, are the primary targets of methylation in maize. The use of restrictive libraries in genome shotgun sequencing in plant genomes should allow significant representation of genes, reducing the number of reactions required.

Cloning, Molecular↗

Visualization of viral candidate cDNAs in infectious brain fractions from Creutzfeldt-Jakob disease by representational difference analysis.

Creutzfeldt-Jakob Disease (CJD), a neurodegenerative and dementing disease of later life, is caused by a viruslike entity that is incompletely characterized. As in scrapie, all more purified infectious brain preparations contain nucleic acids. However, it has not been possible to visualize unique bands that may derive from a viral genome. We here used a subtractive strategy known as representational difference analysis (RDA) to uncover such sequences. To reduce the complexity of starting target nucleic acids, sucrose gradients were first used to select nuclease resistant particles with a defined 120S size. In CJD this single 120S gradient peak is highly enriched for infectivity, and contains reduced amounts of PrP (Proc. Natl. Acad. Sci. 92, 5124-8, 1995). Parallel 120S fractions from uninfected brain were made to generate subtractor sequences. 120S particles were lysed in GdnSCN, and ng amounts of released RNA were purified for random-primed cDNA synthesis. To capture representative fragments of 100-500 bp, cDNAs were cleaved with Mbo I for adaptor ligation and amplification. In the first experiment with moderate RDA selection, it was possible to visualize clones from CJD cDNA that did not hybridize to control cDNA. In the second experiment, more exhaustive subtractions yielded a discrete set of CJD derived gel bands. Competitive hybridization showed a subset of these bands were not present in either the control 120S cDNA or in the hamster genome. This represents the first demonstration of apparently CJD-specific nucleic acid bands in more purified infectious preparations. Although exhaustive cloning, sequencing and correlative titration studies need to be done, it is encouraging that most of the viral candidates selected thus far have no significant homology with any previously described sequence in the database.

Animals↗

Suppression of cortical representation through backward conditioning.

Temporal stimulus reinforcement sequences have been shown to determine the directions of synaptic plasticity and behavioral learning. Here, we examined whether they also control the direction of cortical reorganization. Pairing ventral tegmental area stimulation with a sound in a backward conditioning paradigm specifically reduced representations of the paired sound in the primary auditory cortex (AI). This temporal sequence-dependent bidirectional cortical plasticity modulated by dopamine release hypothetically serves to prevent the over-representation of frequently occurring stimuli resulting from their random pairing with unrelated rewards.

Adaptation, Physiological↗

Numerical characterization and similarity analysis of DNA sequences based on 2-D graphical representation of the characteristic sequences.

Based on the classifications of the four nucleic acid bases, He and Wang reduced a DNA sequence to three binary sequences, which are called the characteristic sequences (J. Chem. Inf. Comput. Sci. 42 (2002) 1080). In this paper, we associate each characteristic sequence with a (b)L / (b)L matrix by giving a 2-D 'two horizontal lines' graphical representation, and thus obtain a 3-component vector with entries being the sums of the maximal and minimal eigenvalues of the (b)L / (b)L matrices. The introduced vector results in more simple characterizations and comparisons among the coding sequences of exon 1 of beta-globin gene of eleven different species.

Animals↗

A method for predicting protein structure from sequence.

BACKGROUND: The ability to predict the native conformation of a globular protein from its amino-acid sequence is an important unsolved problem of molecular biology. We have previously reported a method in which reduced representations of proteins are folded on a lattice by Monte Carlo simulation, using statistically-derived potentials. When applied to sequences designed to fold into four-helix bundles, this method generated predicted conformations closely resembling the real ones. RESULTS: We now report a hierarchical approach to protein-structure prediction, in which two cycles of the above-mentioned lattice method (the second on a finer lattice) are followed by a full-atom molecular dynamics simulation. The end product of the simulations is thus a full-atom representation of the predicted structure. The application of this procedure to the 60 residue, B domain of staphylococcal protein A predicts a three-helix bundle with a backbone root mean square (rms) deviation of 2.25-3 A from the experimentally determined structure. Further application to a designed, 120 residue monomeric protein, mROP, based on the dimeric ROP protein of Escherichia coli, predicts a left turning, four-helix bundle native state. Although the ultimate assessment of the quality of this prediction awaits the experimental determination of the mROP structure, a comparison of this structure with the set of equivalent residues in the ROP dime- crystal structure indicates that they have a rms deviation of approximately 3.6-4.2 A. CONCLUSION: Thus, for a set of helical proteins that have simple native topologies, the native folds of the proteins can be predicted with reasonable accuracy from their sequences alone. Our approach suggest a direction for future work addressing the protein-folding problem.

Journal Article↗

Reduced representation approach to protein tertiary structure prediction: statistical potential and simulated annealing.

A reduced representation model has been developed and used to predict the folded structures of proteins from their primary sequences and random starting conformations. The molecular structure of each protein is reduced to its backbone atoms (with ideal fixed bond lengths and valence angles) and each side chain approximated by a single virtual united atom. The co-ordinate variables are the backbone dihedral angles phi and psi. A statistical potential function, which includes local and non-local interactions and is computed from known X-ray elucidated protein structures, is used in the structure minimization. Simulated annealing method of energy minimization is employed to search for folded conformations. Simulations using the reduced representation model reproduce many structural features of the studied proteins.

Computer Simulation↗

Classifier ensembles for protein structural class prediction with varying homology.

Structural class characterizes the overall folding type of a protein or its domain. A number of computational methods have been proposed to predict structural class based on primary sequences; however, the accuracy of these methods is strongly affected by sequence homology. This paper proposes, an ensemble classification method and a compact feature-based sequence representation. This method improves prediction accuracy for the four main structural classes compared to competing methods, and provides highly accurate predictions for sequences of widely varying homologies. The experimental evaluation of the proposed method shows superior results across sequences that are characterized by entire homology spectrum, ranging from 25% to 90% homology. The error rates were reduced by over 20% when compared with using individual prediction methods and most commonly used composition vector representation of protein sequences. Comparisons with competing methods on three large benchmark datasets consistently show the superiority of the proposed method.

Algorithms↗