PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reduced representation sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

When practice does not make perfect: well-practiced handwriting interferes with the consolidation phase gains in learning a movement sequence.

Practice on a novel sequence of movements can lead to two expressions of procedural memory consolidation: delayed performance gains evolving hours after training, and a decrease in the susceptibility of the training-related gains to interference by subsequent experience. It has been assumed that behavioral interference occurs only if a critical overlap between the representations of the two tasks exists, and that such overlap is more likely when the two tasks are novel, competing for general resources for their execution. We investigated whether the delayed gains in the simple finger-opposition sequence (FOS) learning task are more prone to interference by well practiced than by less practiced complex hand movements. Participants were trained on the FOS task in a baseline (no interference) and an interference training condition. In the Interference condition, after FOS practice, participants wrote Hebrew common words in Hebrew (native script) or a Latin script (Heblatin). Native script writing but not the less practiced Heblatin, interfered with FOS learning, with significantly reduced delayed gains. Our results show that interference can occur even when two tasks share little or no kinematic or dynamic features and indicate that the representation of complex but well-practiced movement sequences may overlap with the representation of simpler ones. This result is in line with the notion that well-practiced complex movement sequences come to be represented as simpler ones in long-term motor memory.

Adult↗

Improving the quality of automatic DNA sequence assembly using fluorescent trace-data classifications.

Virtually all large-scale sequencing projects use automatic sequence-assembly programs to aid in the determination of DNA sequences. The computer-generated assemblies required substantial hand-editing to transform them into submissions for GenBank. As the size of sequencing projects increases, it becomes essential to improve the quality of the automated assemblies so that this time consuming hand-editing may be reduced. Current ABI sequencing technology uses base calls made from fluorescently-labeled DNA fragments run on gels. We present a new representation for the fluorescent trace data associated with individual base calls. This representation can be used before, during, and after fragment assembly to improve the quality of assemblies. We demonstrate one such use-end-trimming of sub-optimal data-that results in a significant improvement in the quality of subsequent assemblies.

Algorithms↗

A strategy for genome-wide gene analysis: integrated procedure for gene identification.

We have developed a technique called the Integrated Procedure for Gene Identification that modifies and integrates parts from several existing techniques to increase the efficiency for genome-wide gene identification. The procedure has the following features: (i) Only the 3' portion of the expressed templates is used to ensure a match to 3' expressed sequence tag (EST) sequences; (ii) the 3' portion of the cDNA is poly dA/poly dT minus, which maintains complete representation of the expressed copies, particularly the rare copies, which otherwise would be lost heavily because of random poly dA/poly dT hybridization in the subtraction reaction; (iii) redundancy is decreased substantially by the subtraction reaction to reduce the effort for sequencing analysis; (iv) the nonsubtracted templates that largely contain the rare copies are amplified selectively with suppression PCR and are sequenced directly or through serial analysis of gene expression (SAGE); and (v) the identified sequences are matched to databases to determine whether they are cloned genes, ESTs, or novel sequences. Using this procedure in a model system, we showed that the redundant copies were largely removed, and the rates of EST matches and the novel sequence identification were significantly increased. Most of the plasmids containing the matched EST are readily available from the IMAGE consortium. This technique can be used to index genome-wide expressed genes and to identify differentially expressed genes in different cells. Compared with the existing techniques, this procedure is relatively efficient, simple, less expensive, and labor intensive. It is especially useful for standard molecular laboratories to perform genome-wide studies.

Base Sequence↗

Pan-genomics and multi-omics for deciphering genetic variation and accelerating genetic improvement in ruminant livestock.

Livestock reference genomes have transformed the discovery of variants associated with production, reproduction, health, and environmental adaptation. Nevertheless, a single linear reference represents only one mosaic haplotype and incompletely captures sequence diversity within a species, particularly structural variants, copy-number changes, repeat-rich regions, and breed-specific sequences. Pangenomes address this limitation by integrating multiple high-quality assemblies or population-scale variants into a unified sequence or graph representation. Concurrently, multi-omics approaches connect genomic variation with transcriptomic, epigenomic, manuscriptproteomic, metabolomic, and microbiome responses, thereby improving biological interpretation of genotype-phenotype relationships. This review synthesizes recent progress in livestock pangenomics and multi-omics, with emphasis on cattle, goats, sheep, water buffalo, and chickens. It describes advances in long-read and haplotype-resolved sequencing, graph construction, structural-variant discovery and genotyping, functional annotation, and integrative analysis. Recent pangenome studies have uncovered substantial non-reference sequence, reduced reference bias, identified breed- and population-specific structural variants, and resolved candidate variants underlying pigmentation, body size, tail morphology, cashmere production, altitude adaptation, and other economically relevant traits. However, translation into routine breeding remains constrained by uneven population representation, inconsistent structural-variant definitions, limited functional annotation, computational demands, and insufficient validation across environments. Future progress will depend on diverse near-complete assemblies, graph-aware imputation and genomic prediction, long-read transcriptomics, single-cell and spatial omics, rigorous causal validation, and open, interoperable resources. Together, these developments can support more accurate, resilient, and biologically informed livestock improvement. Importantly, current dairy-cattle evidence indicates that pangenome-derived structural variants can substantially improve variant discovery and functional interpretation while yielding only marginal average gains in routine genomic prediction, favoring targeted augmentation rather than wholesale replacement of established SNP-based evaluations.

Animals↗

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans↗

A new method to find a set of energetically optimal RNA secondary structures.

We present a computer method to determine nucleic acid secondary structures. It is based on three steps: 1) the search for all possible helical regions relied on a mathematical approach derived from the convolution theorem; it uses a tetradimensional complex vector representation of the bases along the sequence; 2) a 'tree' search for a set of minimum free energy structures, by the aid of an approximate energy evaluation to reduce the computer time requirements; 3) the exact calculation and refinement of the energies. A method to introduce the experimental data and reach an arrangement between them and the free energy minimization criterion is shown. In order to demonstrate the confidence of the program a test on four RNA sequences is performed. The method has computer time requirement proportional to N2, where N is the length of the sequence and retrieves a set of optimal free energy structures.

Base Composition↗

Reducing haystacks to needles - ViralClust: A Nextflow pipeline to cluster viral sequences.

BACKGROUND: The rapid accumulation of viral genome sequences presents major challenges for downstream analysis tools, including tools for multiple sequence alignments, phylogeny, and genome/alignment visualization, due to computational constraints and sampling biases caused by outbreak-driven over-representation. Selecting representative genomes through clustering offers a principled alternative to random subsampling, yet choosing appropriate clustering strategies remains non-trivial and context-dependent. RESULTS: Here, we present ViralClust, a modular Nextflow pipeline for bias-aware representative selection from large viral genome datasets. ViralClust integrates five distinct clustering algorithms (CD-HIT-EST, SUMACLUST, VSEARCH, MMSeqs2, and HDBSCAN) within a unified workflow, enabling direct comparison of clustering outcomes and flexible adaptation to diverse biological questions, considering a balanced phylogenetic distribution of the selected sequences. We evaluated ViralClust on six RNA and DNA virus datasets ranging from 632 to 156,586 sequences and spanning genome lengths from 890 to 197,185 nucleotides. Across all datasets, clustering reduced dataset size by ~95 % or more while preserving genetic diversity across species, genera, and families, and effectively mitigating biases introduced by outbreaks, partial genomes, and sequence orientation artifacts. CONCLUSIONS: By supporting whole-genome clustering and scalable representative selection, ViralClust enables efficient and reproducible downstream analyses that would otherwise be computationally infeasible. Rather than offering a prescriptive, guided analysis engine, our framework functions as a flexible comparative collection of complementary strategies, allowing users to empirically evaluate trade-offs and choose the ideal method tailored to their specific analytical endpoints.

Bioinformatics↗

[A method of selective PCR-amplification of genomic DNA fragments (SAGF method)].

A method for separating into definite sets of a complex mixture of fragments obtained by DNA cleavage with IIS- or IIN-types of restriction endonucleases producing single-stranded termini of different sequences at the fragment ends has been developed. The method is based on the ligation of short double-stranded adapters with single-stranded termini complementary to the termini of a selected set of fragments followed by PCR-amplification with the primer which represents a strand of the adapters. Using endonucleases BcoKI and Bli7361 recognizing sequences CTCTTC and GGTCTC and producing three- and four-nucleotide 5'-termini, respectively, it has been shown that amplification of a set of fragments occurs only when the adapters are attached to DNA fragments with DNA-ligase. Several applications of the SAGF-method are suggested: for obtaining individual bands in DNA fingerprinting; for reducing the kinetic complexity of DNA in the representational difference analysis (RDA method) of complex genomes; for cataloguing DNA fragments, and for constructing physical genomic maps.

Animals↗

Visual imagery in hemianopic patients.

In this article we report some findings about visual imagery in patients with stable homonymous hemianopia compared to healthy control subjects. These findings were obtained by analyzing the gaze control through recording of eye movements in different phases of viewing and imagery. We used six different visual stimuli for the consecutive viewing and imagery phases. With infrared oculography, we recorded eye movements during this presentation phase and in three subsequent imagery phases in absence of the stimulus. Analyzing the basic parameters of the gaze sequences (known as "scanpaths"), we discovered distinct characteristics of the "viewing scanpaths" and the "imagery scanpaths" in both groups, which suggests a reduced extent of the image within the cognitive representation. We applied different similarity measures (string/vector string editing, Markov analysis). We found a "progressive consistency of imagery," shown through raising similarity values for the comparison of the late imagery scanpaths. This result suggests a strong top-down component in picture exploration: In both groups, healthy subjects and hemianopic patients, a mental model of the viewed picture must evolve very soon and substantially determine the eye movements. As our hemianopic patients showed analogous results to the normal subjects, we conclude that these patients are well adjusted to their deficit and, despite their perceptual defect, have a preserved cognitive representation, which follows the same top-down vision strategies in the process of visual imagery.

Adult↗

Using information content and base frequencies to distinguish mutations from genetic polymorphisms in splice junction recognition sites.

Predicting the effects of nucleotide substitutions in human splice sites has been based on analysis of consensus sequences. We used a graphic representation of sequence conservation and base frequency, the sequence logo, to demonstrate that a change in a splice acceptor of hMSH2 (a gene associated with familial nonpolyposis colon cancer) probably does not reduce splicing efficiency. This confirms a population genetic study that suggested that this substitution is a genetic polymorphism. The information theory-based sequence logo is quantitative and more sensitive than the corresponding splice acceptor consensus sequence for detection of true mutations. Information analysis may potentially be used to distinguish polymorphisms from mutations in other types of transcriptional, translational, or protein-coding motifs.

Base Sequence↗

Breast volume denoising and noise characterization by 3D wavelet transform.

Breast imaging through cone-beam computed tomography provides a digital breast volume, with which the three-dimensional (3D) breast tissues can be analyzed. Data denoising, as a preprocessing step for subsequent volumetric breast segmentation is always needed. In this paper, we report a volumetric denoising technique by a separable 3D wavelet transform (WT), i.e. a '2D WT plus 1D WT' scheme. Specifically, the scheme performs two-dimensional (2D) wavelet denoising on a stack of slice images of the breast volume, followed by one-dimensional (1D) wavelet denoising along the stacking direction. The denoising is achieved by wavelet decomposition, high-pass subband attenuation, and wavelet synthesis. A one-level 3D WT produces eight subbands occupying the octants of the 3D wavelet space. Multilevel WT also provides a multiresolution representation of breast volume, i.e. a sequence of low-pass subbands. In general, most noise and irregularity features are imparted into the high-pass subbands, which are removed or reduced for denoising purpose. Meanwhile, the information in a subband can be characterized in terms of energy, variance, and entropy. Through 3D visualization, the spatial structure in a subband can also be visually perceived. Experimental demonstration with the breast volume reconstructed from a specimen is provided.

Mammography↗

Fatty acid-oxidizing consortia along a nutrient gradient in the Florida Everglades.

The Florida Everglades is one of the largest freshwater marshes in North America and has been subject to eutrophication for decades. A gradient in P concentrations extends for several kilometers into the interior of the northern regions of the marsh, and the structure and function of soil microbial communities vary along the gradient. In this study, stable isotope probing was employed to investigate the fate of carbon from the fermentation products propionate and butyrate in soils from three sites along the nutrient gradient. For propionate microcosms, 16S rRNA gene clone libraries from eutrophic and transition sites were dominated by sequences related to previously described propionate oxidizers, such as Pelotomaculum spp. and Syntrophobacter spp. Significant representation was also observed for sequences related to Smithella propionica, which dismutates propionate to butyrate. Sequences of dominant phylotypes from oligotrophic samples did not cluster with known syntrophs but with sulfate-reducing prokaryotes (SRP) and Pelobacter spp. In butyrate microcosms, sequences clustering with Syntrophospora spp. and Syntrophomonas spp. dominated eutrophic microcosms, and sequences related to Pelospora dominated the transition microcosm. Sequences related to Pelospora spp. and SRP dominated clone libraries from oligotrophic microcosms. Sequences from diverse bacterial phyla and primary fermenters were also present in most libraries. Archaeal sequences from eutrophic microcosms included sequences characteristic of Methanomicrobiaceae, Methanospirillaceae, and Methanosaetaceae. Oligotrophic microcosms were dominated by acetotrophs, including sequences related to Methanosarcina, suggesting accumulation of acetate.

Bacteria↗

Robust remote homology detection by feature based Profile Hidden Markov Models.

The detection of remote homologies is of major importance for molecular biology applications like drug discovery. The problem is still very challenging even for state-of-the-art probabilistic models of protein families, namely Profile HMMs. In order to improve remote homology detection we propose feature based semi-continuous Profile HMMs. Based on a richer sequence representation consisting of features which capture the biochemical properties of residues in their local context, family specific semi-continuous models are estimated completely data-driven. Additionally, for substantially reducing the number of false predictions an explicit rejection model is estimated. Both the family specific semi-continuous Profile HMM and the non-target model are competitively evaluated. In the experimental evaluation of superfamily based screening of the SCOP database we demonstrate that semi-continuous Profile HMMs significantly outperform their discrete counterparts. Using the rejection model the number of false positive predictions could be reduced substantially which is an important prerequisite for target identification applications.

Journal Article↗

Information, intelligence, and interface: the pillars of a successful medical information system.

This paper addresses three key issues facing developers of clinical and/or research medical information systems. 1. INFORMATION. The basic function of every database is to store information about the phenomenon under investigation. There are many ways to organize information in a computer; however only a few will prove optimal for any real life situation. Computer Science theory has developed several approaches to database structure, with relational theory leading in popularity among end users [8]. Strict conformance to the rules of relational database design rewards the user with consistent data and flexible access to that data. A properly defined database structure minimizes redundancy i.e.,multiple storage of the same information. Redundancy introduces problems when updating a database, since the repeated value has to be updated in all locations--missing even a single value corrupts the whole database, and incorrect reports are produced [8]. To avoid such problems, relational theory offers a formal mechanism for determining the number and content of data files. These files not only preserve the conceptual schema of the application domain, but allow a virtually unlimited number of reports to be efficiently generated. 2. INTELLIGENCE. Flexible access enables the user to harvest additional value from collected data. This value is usually gained via reports defined at the time of database design. Although these reports are indispensable, with proper tools more information can be extracted from the database. For example, machine learning, a sub-discipline of artificial intelligence, has been successfully used to extract knowledge from databases of varying size by uncovering a correlation among fields and records[1-6, 9]. This knowledge, represented in the form of decision trees, production rules, and probabilistic networks, clearly adds a flavor of intelligence to the data collection and manipulation system. 3. INTERFACE. Despite the obvious importance of collecting data and extracting knowledge, current systems often impede these processes. Problems stem from the lack of user friendliness and functionality. To overcome these problems, several features of a successful human-computer interface have been identified [7], including the following "golden" rules of dialog design [7]: consistency, use of shortcuts for frequent users, informative feedback, organized sequence of actions, simple error handling, easy reversal of actions, user-oriented focus of control, and reduced short-term memory load. To this list of rules, we added visual representation of both data and query results, since our experience has demonstrated that users react much more positively to visual rather than textual information. In our design of the Orthopaedic Trauma Registry--under development at the Carolinas Medical Center--we have made every effort to follow the above rules. The results were rewarding--the end users actually not only want to use the product, but also to participate in its development.

Artificial Intelligence↗

Gene function and expression level influence the insertion/fixation dynamics of distinct transposon families in mammalian introns.

BACKGROUND: Transposable elements (TEs) represent more than 45% of the human and mouse genomes. Both parasitic and mutualistic features have been shown to apply to the host-TE relationship but a comprehensive scenario of the forces driving TE fixation within mammalian genes is still missing. RESULTS: We show that intronic multispecies conserved sequences (MCSs) have been affecting TE integration frequency over time. We verify that a selective economizing pressure has been acting on TEs to decrease their frequency in highly expressed genes. After correcting for GC content, MCS density and intron size, we identified TE-enriched and TE-depleted gene categories. In addition to developmental regulators and transcription factors, TE-depleted regions encompass loci that might require subtle regulation of transcript levels or precise activation timing, such as growth factors, cytokines, hormones, and genes involved in the immune response. The latter, despite having reduced frequencies of most TE types, are significantly enriched in mammalian-wide interspersed repeats (MIRs). Analysis of orthologous genes indicated that MIR over-representation also occurs in dog and opossum immune response genes, suggesting, given the partially independent origin of MIR sequences in eutheria and metatheria, the evolutionary conservation of a specific function for MIRs located in these loci. Consistently, the core MIR sequence is over-represented in defense response genes compared to the background intronic frequency. CONCLUSION: Our data indicate that gene function, expression level, and sequence conservation influence TE insertion/fixation in mammalian introns. Moreover, we provide the first report showing that a specific TE family is evolutionarily associated with a gene function category.

Animals↗

Correction of CSF motion artifact on MR images of the brain and spine by pulse sequence modification: clinical evaluation.

A modification of the standard spin-echo pulse sequence designed to suppress motion artifacts was clinically evaluated on T2-weighted MR images of the cervicocranial region. A retrospective study involving 40 patients, half of whom were examined with a standard T2-weighted multislice spin-echo sequence and half of whom were examined with a gradient waveform modification of the same sequence, uniformly demonstrated restoration of CSF signal intensity on images obtained with the gradient modified sequence. The cervical subarachnoid spaces, cisterna magna, medullary cistern, pontine cistern, fourth ventricle, and aqueduct were more consistently and brightly represented. However, the phase-encoding artifacts arising from CSF motion were not significantly reduced by using the gradient waveform modified pulse sequence. Digital subtraction of an image obtained with the standard sequence from an image of the same slice with the gradient modified sequence provides a direct image representation of CSF flow.

Brain↗

A probabilistic algorithm for interactive huge genome comparison.

We designed a new probabilistic algorithm, named PAGEC (probabilistic algorithm for genome comparison), which allowed a highly interactive study of long genomic strings. The comparison between two nucleic acid sequences is based on the creation of multiple index tables, which drastically reduces processing time for huge genomes, e.g. 13 min for a 4 Mb/4 Mb comparison. PAGEC lowered the need for memory when compared with other types of algorithm and took into account the low resolution of the final representation (paper or computer screen). Considering that standard printers permit a 300 d.p.i. resolution, the loss of computed information due to the probabilistic conception of the algorithm was not usually noticeable in the present study, mainly due to increased genomic sizes. Refinement was possible through an interactive zooming system, which enabled the visualization of the lexical base sequences of a considered part of both of the studied genomes. Biological examples of computation based on yeast and animal nucleic acid sequences presented in this paper reveal the flexibility of the PAGEC program, which is a valuable tool for genetic studies as it offers a solution to an important problem that will become even more important as time passes.

Algorithms↗

PCR-SSCP comparison of 16S rDNA sequence diversity in soil DNA obtained using different isolation and purification methods.

This study compared different methods of direct DNA extraction and purification from a silt loam soil and investigated the relationship between DNA quantity and sequence diversity. Five extraction methods and four purification techniques were investigated. Quantities of DNA extracted were between 3.4+/-0.55 and 54.3+/-8.18 &mgr;g g(-1) (dry wt) of soil with OD(260)/OD(230) purity ratios between 0.80 and 1.15. Analysis of sequence diversity in all extracts was conducted using PCR-single strand conformation polymorphism (SSCP). Profiles generated using universal 16S rDNA primers (Com1/Com2) were found to be identical when used to amplify 16S rDNA extracted directly from soil. The genus Pseudomonas was targeted in order to reduce profile complexity, which was apparent when using universal 16S rDNA primers, and which hindered direct comparison of sequence diversity. A Pseudomonas culture library and non-cultured Pseudomonas 16S rDNA genes were used to provide a background count of Pseudomonas operational taxonomic units present in the soil. Cloning and sequencing of amplicons generated using a Pseudomonas-specific (Ps-for) and a universal 16S rDNA (Com2) primer, coupled with nested amplification (Com1/Com2 amplification from Ps-for/Ps-rev amplicons), used in conjunction with SSCP, revealed that environmental contaminants co-extracted with DNA, such as humic acid, significantly reduced primer specificity. SSCP was sensitive enough to reveal template bias in different primer sets. PCR-restriction fragment length-SSCP of Pseudomonas 16S rDNA amplified from soil-extracted DNA revealed distinct differences in sequence representation between extraction methods and showed that greater DNA yield is not synonymous with higher sequence diversity. We, therefore, suggest that DNA extractions from soil should be evaluated not only in terms of quantity and purity, but also in terms of the sequence diversity present. SSCP proved to be a valuable tool for the assessment of the methodologies commonly used in PCR-mediated microbial ecology studies.

Journal Article↗