PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reduced representation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A subtractive hybridisation method for the enrichment of moderately induced sequences.

Moderately induced genes often escape detection in conventional subtraction hybridisation cloning. Here a modification of a phagemid subtraction protocol is described that overcomes this problem. The protocol uses low ratio hybridisation of driver to target sequences to allow enrichment of the sequences of interest, and back-hybridisation of the subtracted sequences with induced sequences to reduce the accumulation of false positive clones. The procedure takes advantage of the quantitative representation of cellular RNA populations in cDNA libraries, therefore, they may serve not only as renewable sources of driver and target sequences, but also as sources of population cRNAs used in northern blots and differential Southern blots.

Animals↗

Crystal chemistry of zirconosilicates and their analogs: topological classification of MT frameworks and suprapolyhedral invariants.

The first attempt is undertaken to consider systematically topological structures of zirconosilicates and their analogs (60 minerals and 34 synthetic phases), where the simplest structure units are MO(6) octahedra and TO(4) tetrahedra united by vertices ([TO(4)]:[MO(6)] = 1:1-6:1). A method of analysis and classification of mixed three-dimensional MT frameworks by topological types with coordination sequences [N(k)] is developed, which is based on the representation of crystal structure as a finite "reduced" graph. The method is optimized for the frameworks of any composition and complexity and implemented within the TOPOS3.2 program package. A procedure of hierarchical analysis of MT-framework structure organization is proposed, which is based on the concept of polyhedral microensemble (PME) being a geometrical interpretation of coordination sequences of M and T nodes. All 12 theoretically possible PMEs of MT(6) polyhedral composition are considered where T is a separate and/or connected tetrahedron. Using this methodology the MT frameworks in crystal structures of zirconosilicates and their analogs were analyzed within the first 12 coordination spheres of M and T nodes and related to 41 topological types. The structural correlations were revealed between rosenbuschite, lavenite, hiortdahlite, woehlerite, siedozerite and the minerals of the eudialyte family.

Journal Article↗

Reduced representation model of protein structure prediction: statistical potential and genetic algorithms.

A reduced representation model, which has been described in previous reports, was used to predict the folded structures of proteins from their primary sequences and random starting conformations. The molecular structure of each protein has been reduced to its backbone atoms (with ideal fixed bond lengths and valence angles) and each side chain approximated by a single virtual united-atom. The coordinate variables were the backbone dihedral angles phi and psi. A statistical potential function, which included local and nonlocal interactions and was computed from known protein structures, was used in the structure minimization. A novel approach, employing the concepts of genetic algorithms, has been developed to simultaneously optimize a population of conformations. With the information of primary sequence and the radius of gyration of the crystal structure only, and starting from randomly generated initial conformations, I have been able to fold melittin, a protein of 26 residues, with high computational convergence. The computed structures have a root mean square error of 1.66 A (distance matrix error = 0.99 A) on average to the crystal structure. Similar results for avian pancreatic polypeptide inhibitor, a protein of 36 residues, are obtained. Application of the method to apamin, an 18-residue polypeptide with two disulfide bonds, shows that it folds apamin to native-like conformations with the correct disulfide bonds formed.

Algorithms↗

Chunking in task sequences modulates task inhibition.

In a study of the formation of representations of task sequences and its influence on task inhibition, participants first performed tasks in a predictable sequence (e.g., ABACBC) and then performed the tasks in a random sequence. Half of the participants were explicitly instructed about the predictable sequence, whereas the other participants did not receive these instructions. Task-sequence learning was inferred from shorter reaction times (RTs) in predictable relative to random sequences. Persisting inhibition of competing tasks was indicated by increased RTs in n- 2 task repetitions (e.g., ABA) compared with n- 2 nonrepetitions (e.g., CBA). The results show task-sequence learning for both groups. However, task inhibition was reduced in predictable relative to random sequences among instructed-learning participants who formed an explicit representation of the task sequence, whereas sequence learning and task inhibition were independent in the noninstructed group. We hypothesize that the explicit instructions led to chunking of the task sequence, and that n- 2 repetitions served as chunk points (ABA-CBC), so that within-chunk facilitation modulated the inhibition effect.

Adult↗

Sequence representation and prediction of protein secondary structure for structural motifs in twilight zone proteins.

Characterizing and classifying regularities in protein structure is an important element in uncovering the mechanisms that regulate protein structure, function and evolution. Recent research concentrates on analysis of structural motifs that can be used to describe larger, fold-sized structures based on homologous primary sequences. At the same time, accuracy of secondary protein structure prediction based on multiple sequence alignment drops significantly when low homology (twilight zone) sequences are considered. To this end, this paper addresses a problem of providing an alternative sequences representation that would improve ability to distinguish secondary structures for the twilight zone sequences without using alignment. We consider a novel classification problem, in which, structural motifs, referred to as structural fragments (SFs) are defined as uniform strand, helix and coil fragments. Classification of SFs allows to design novel sequence representations, and to investigate which other factors and prediction algorithms may result in the improved discrimination. Comprehensive experimental results show that statistically significant improvement in classification accuracy can be achieved by: (1) improving sequence representations, and (2) removing possible noise on the terminal residues in the SFs. Combining these two approaches reduces the error rate on average by 15% when compared to classification using standard representation and noisy information on the terminal residues, bringing the classification accuracy to over 70%. Finally, we show that certain prediction algorithms, such as neural networks and boosted decision trees, are superior to other algorithms.

Algorithms↗

NotI subtraction and NotI-specific microarrays to detect copy number and methylation changes in whole genomes.

Methylation, deletions, and amplifications of cancer genes constitute important mechanisms in carcinogenesis. For genome-wide analysis of these changes, we propose the use of NotI clone microarrays and genomic subtraction, because NotI recognition sites are closely associated with CpG islands and genes. We show here that the CODE (Cloning Of DEleted sequences) genomic subtraction procedure can be adapted to NotI flanking sequences and to CpG islands. Because the sequence complexity of this procedure is greatly reduced, only two cycles of subtraction are required. A NotI-CODE procedure can be used to prepare NotI representations (NRs) containing 0.1-0.5% of the total DNA. The NRs contain, on average, 10-fold less repetitive sequences than the whole human genome and can be used as probes for hybridization to NotI microarrays. These microarrays, when probed with NRs, can simultaneously detect copy number changes and methylation. NotI microarrays offer a powerful tool with which to study carcinogenesis.

Animals↗

A general diffusion model for analyzing the efficacy of synaptic input to threshold neurons.

We describe a general diffusion model for analyzing the efficacy of individual synaptic inputs to threshold neurons. A formal expression is obtained for the system propagator which, when given an arbitrary initial state for the cell, yields the conditional probability distribution for the state at all later times. The propagator for a cell with a finite threshold is written as a series expansion, such that each term in the series depends only on the infinite threshold propagator, which in the diffusion limit reduces to a Gaussian form. This procedure admits a graphical representation in terms of an infinite sequence of diagrams. To connect the theory to experiment, we construct an analytical expression for the primary correlation kernel (PCK) which profiles the change in the instantaneous firing rate produced by a single postsynaptic potential (PSP). Explicit solutions are obtained in the diffusion limit to first order in perturbation theory. Our approximate expression resembles the PCK obtained by computer simulation, with the accuracy depending strongly on the mode of firing. The theory is most accurate when the synaptic input drives the membrane potential to a mean level more than one standard deviation below the firing threshold, making such cells highly sensitive to synchronous synaptic input.

Computer Simulation↗

Very short patch repair: reducing the cost of cytosine methylation.

In Escherichia coli and related bacteria, the product of gene dcm methylates the second cytosine of 5'-CCWGG sequences (where W is A or T). Deamination of 5-methylcytosine (5meC) results in C to T mutations. The mutagenic potential of 5meC is reduced by a system called very short patch (VSP) repair, which replaces T with C. T:G and U:G mispairs in the methylatable sequence and in related sequences are recognized by the product of vsr, a gene adjacent to dcm. Vsr creates a nick just 5' of the mispaired pyrimidine to initiate the repair. Additional products known to be required for VSP repair are DNA polymerase I and DNA ligase. MutS and MutL have a stimulatory role but are not required. The ability of Vsr to recognize T:G mispairs in sequences related to CCWGG is probably responsible for over- and under-representation of certain tetranucleotides in the E. coli genome. Although VSP repair reduces spontaneous mutations at 5meCs in replicating bacteria, mutation hot-spots persist at these sites. Under conditions that more accurately mimic the natural environment of E. coli, VSP repair appears to be effective in preventing mutation at 5meC.

Amino Acid Sequence↗

Conversion of nucleotides sequences into genomic signals.

An original tetrahedral representation of the Genetic Code (GC) that better describes its structure, degeneration and evolution trends is defined. The possibility to reduce the dimension of the representation by projecting the GC tetrahedron on an adequately oriented plane is also analyzed, leading to some equivalent complex representations of the GC. On these bases, optimal symbolic-to-digital mappings of the linear, nucleic acid strands into real or complex genomic signals are derived at nucleotide, codon and amino acid levels. By converting the sequences of nucleotides and polypeptides into digital genomic signals, this approach offers the possibility to use a large variety of signal processing methods for their handling and analysis. It is also shown that some essential features of the nucleotide sequences can be better extracted using this representation. Specifically, the paper reports for the first time the existence of a global helicoidal wrapping of the complex representations of the bases along DNA sequences, a large scale trend of genomic signals. New tools for genomic signal analysis, including the use of phase, aggregated phase, unwrapped phase, sequence path, stem representation of components' relative frequencies, as well as analysis of the transitions are introduced at the nucleotide, codon and amino acid levels, and in a multiresolution approach.

Amino Acid Sequence↗

Three-dimensional visualization of the mandible: a new method for presenting the periodontal status and diseases.

A three-dimensional representation of an ideal human jaw was reconstructed from a series of 178 digitized photographic cross-sections. The two-dimensional images were taken from the slices of an artificial skull and have been digitized by means of a high resolution scanner with a spatial resolution of 860 dots per inch. On the basis of that sequence of cross-sections, a semi-automatic segmentation algorithm was developed to reduce the information to a quadriliteral surface-representation of the teeth and the bone. An algorithm was developed which simulates the individual pathology of a patient both on the basis of the findings stored in this patient's dental record and by using the representation of the reference jaw. The results of the modeling module were automatically prepared for rendering with visualization tools compatible to the RenderMan standard. This new method of presenting periodontal situations in especially helpful for the diagnostic support of periodontologists and for dental educational purposes.

Gingiva↗

Coordination in childhood: modifications of visuomotor representations in 6- to 11-year-old children.

This research investigated the development of visuomotor coordination in childhood, more specifically the conversion of visual information into motor sequences. Three groups of children (aged 6, 8 and 11 years) and a group of adults performed pointing movements without direct feedback from their arm displacements. Visual information, provided by a video camera, was disturbed by rotations of 0 degree, 45 degrees, 90 degrees, 135 degrees or 180 degrees. Six-year-old children showed poor accuracy for 180 degrees rotations. These results suggest that the youngest children use unidirectional representations to convert visual information into motor sequences. At 8 years of age, children showed a shift from unidirectional to bidirectional representations, as reflected by reduced errors for 180 degrees rotations. Eleven-year-old children and adults showed the same type of representations, i.e., bidirectional. However, as reflected by their slower movement time and slower modifications in temporal accuracy across trials, the oldest children have not yet reached maturation in their adaptive process when compared to adults.

Adaptation, Physiological↗

Spectrum of X-ray-induced mutations in the human hprt gene.

We have characterized the molecular spectrum of mutations in 116 X-ray-induced and 78 spontaneous, HPRT- mutants derived from the human B lymphoblastoid cell line TK6. Multiplex PCR analysis demonstrated that the overall representation of large deletions was not significantly different in the two spectra. However, highly significant differences were observed for specific deletion types. Total gene deletions represented 41/78 (0.53) X-ray-induced, but only 7/43 (0.16) spontaneous deletions (P < 0.0001). In contrast, 5' terminal deletions were significantly more common among spontaneous (17/43, 0.40) than X-ray-induced (13/78, 0.17) large deletions (P = 0.0079). The types of point mutations induced by X-ray exposure were very diverse including all classes of transitions and transversions, tandem base substitutions, frameshifts, small deletions and a deletion/insertion compound mutation. Compared to spontaneous data, radiation-induced point mutations exhibited a reduced number of transitions and an increased representation of small deletions. Small deletions were uniformly surrounded by direct sequence repeats. The distribution of point mutations was characterized by a cluster within the 5' portion of exon 8. Thirteen HPRT- point mutations exhibited aberrant splicing. Four of these were attributable to coding sequence alterations in exons 4 and 8. These results suggest that it may be possible to identify hallmark mutations associated with X-ray exposure of human cells.

B-Lymphocytes↗

Distinguishing features of 16S rDNA gene for five dominating bacterial genus observed in bioremediation.

Defining a microbial community and identifying bacteria, at least at the genus level, is a first step in predicting the behavior of a microbial community in bioremediation. In biological treatment systems, the most dominating groups observed are Pseudomonas, Moraxella, Acinetobactor, Burkholderia, and Alcaligenes. Our interest lies in identifying the distinguishing features of these bacterial groups based on their 16S rDNA sequence data, which could be used further for generating genus-specific probes. Accordingly, 20 sequences representing different species from each genus above were retrieved, which constituted a training set. A 16-dimensional feature vector comprised of transition probabilities of nucleotides was considered and each sampled sequence was expressed in terms of these features. A stepwise feature selection method was used to identify features that are distinct across the species of these five groups. Wilk's lambda selection criterion was used and resulted in a subset with six distinguishing features. The discriminating efficacy of this subset was tested through multiple group discriminant analysis. Two linear composites, as a function of these features, could discriminate the test set of forty-five sequences from these groups with 95% accuracy, thereby ascertaining the relevance of the identified features. The geometric representation of feature correlation in the reduced discriminant space demonstrated the dominance of identified features in specific groups. These features independently or in combination could be used to generate genus-specific patterns to design probes, so as to develop a tracking tool for the selected group of bacteria.

Bacteria↗

Improved in situ hybridization to HIV with RNA probes derived from PCR products.

These experiments tested the hypothesis that a pool of PCR-derived RNA probes with defined length and even representation of the target sequences could produce more specific and intense in situ hybridization signals than randomly size-reduced, plasmid-derived RNA probes. In situ hybridization was performed with sense and anti-sense HIV-1 RNA probes that were derived from PCR products tailed with the T7 RNA polymerase promoter or from plasmid DNA. In situ hybridization using a pool of seven anti-sense or sense PCR-derived RNA probes (1805 nucleotides of HIV sequence, 257 nucleotides average probe length) was compared with hybridization using anti-sense or sense RNA probes made from a plasmid representing the HIV-1 env gene (3151 nucleotides of HIV-1 target). The pooled PCR-derived probes resulted in stronger in situ hybridization signals and less background than those produced with plasmid-derived RNA probes. This method for creating PCR-derived RNA probes improves the feasibility of synthesizing multiple, discrete RNA probes for studies of specific mRNA expression because it does not require the subcloning steps used to construct plasmids. PCR-derived RNA probes may provide a viable alternative to the use of plasmid-derived RNA probes for in situ hybridization.

Cells, Cultured↗

A graph-topological approach to recognition of pattern and similarity in RNA secondary structures.

Secondary and tertiary RNA structures play an important role in many biological processes. Therefore the necessity arises to find similar higher-order structures for different but functionally homologous RNA sequences. We propose here a graph-topological approach to the problem, which shows two main features: simplified graph representation which allows the recognition of similarity of RNA secondary structures with the same branching look despite minor differences. This allows comparison among foldings from different sequences, and "pruning" of the secondary structures not shared by all the sequences since the early stages of the search. (b) The graph representation is encoded by the Randić topological index, and the search for the folding similarity is reduced to checking the identity of single numbers. These characteristics make this approach significantly different, less depending on empirical criteria, and less computationally heavy then previous methods, where the folding consensus has been measured by an alignment procedure or correlation of strings representing the secondary structures. Some U2 snRNA and viroid sequences are studied by this approach, which is imbedded in our previous search method based on genetic algorithms.

Algorithms↗

A unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification.

The rapid growth of genomic sequencing demands fast, accurate, and scalable analysis methods. In viral genomic classification, expanding labeled reference collections can make supervised models costly to update and dependent on fixed label sets, motivating retrieval-based genomic classification as a simpler, more flexible alternative. We present a unified benchmark of supervised and retrieval-based methods for viral genomic sequence classification across three viral classification tasks: hepatitis C virus (HCV) genotyping, COVID-19 discrimination, and human papillomavirus (HPV) genotyping. We compare standard sequence encodings (one-hot, k-mers, FCGR) with dense embeddings (dna2vec, DNABERT). For each representation, we evaluate supervised classifiers (Random Forest, Decision Tree, XGBoost) and retrieval-based classification, where sequence vectors are indexed with FAISS and labels are assigned via similarity-weighted k-NN. Furthermore, we benchmark multiple FAISS index types (Flat, IVF, HNSW, IVFPQ, OPQ) to characterize accuracy-speed-memory trade-offs at scale. The results show that XGBoost and retrieval using Flat or IVF indexes achieve strong classification performance under different computational profiles. Compressed indexes such as IVFPQ and OPQ substantially reduce memory usage, although their accuracy loss depends on the dataset and representation. Overall, supervised XGBoost provides a favorable accuracy-size trade-off, while retrieval-based classification remains competitive and allows labeled reference sequences to be incorporated without retraining a global classifier. This benchmark provides practical guidance for selecting sequence representations, classifiers, and vector-search indexes under different accuracy, memory, and update requirements.

Genome, Viral↗

Multiple sequence alignment using partial order graphs.

MOTIVATION: Progressive Multiple Sequence Alignment (MSA) methods depend on reducing an MSA to a linear profile for each alignment step. However, this leads to loss of information needed for accurate alignment, and gap scoring artifacts. RESULTS: We present a graph representation of an MSA that can itself be aligned directly by pairwise dynamic programming, eliminating the need to reduce the MSA to a profile. This enables our algorithm (Partial Order Alignment (POA)) to guarantee that the optimal alignment of each new sequence versus each sequence in the MSA will be considered. Moreover, this algorithm introduces a new edit operator, homologous recombination, important for multidomain sequences. The algorithm has improved speed (linear time complexity) over existing MSA algorithms, enabling construction of massive and complex alignments (e.g. an alignment of 5000 sequences in 4 h on a Pentium II). We demonstrate the utility of this algorithm on a family of multidomain SH2 proteins, and on EST assemblies containing alternative splicing and polymorphism. AVAILABILITY: The partial order alignment program POA is available at http://www.bioinformatics.ucla.edu/poa.

Algorithms↗

Visual learning by imitation with motor representations.

We propose a general architecture for action (mimicking) and program (gesture) level visual imitation. Action-level imitation involves two modules. The viewpoint Transformation (VPT) performs a "rotation" to align the demonstrator's body to that of the learner. The Visuo-Motor Map (VMM) maps this visual information to motor data. For program-level (gesture) imitation, there is an additional module that allows the system to recognize and generate its own interpretation of observed gestures to produce similar gestures/goals at a later stage. Besides the holistic approach to the problem, our approach differs from traditional work in i) the use of motor information for gesture recognition; ii) usage of context (e.g., object affordances) to focus the attention of the recognition system and reduce ambiguities, and iii) use iconic image representations for the hand, as opposed to fitting kinematic models to the video sequence. This approach is motivated by the finding of visuomotor neurons in the F5 area of the macaque brain that suggest that gesture recognition/imitation is performed in motor terms (mirror) and rely on the use of object affordances (canonical) to handle ambiguous actions. Our results show that this approach can outperform more conventional (e.g., pure visual) methods.

Algorithms↗