PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Purification of genomic sequences from bacteriophage libraries by recombination and selection in vivo.

Cloned genes have been purified from recombinant DNA bacteriophage libraries by a method exploiting homologous reciprocal recombination in vivo. In this method 'probe' sequences are inserted in a very small plasmid vector and introduced into recombination-proficient bacterial cells. Genomic bacteriophage libraries are propagated on the cells, and phage bearing sequences homologous to the probe acquire an integrated copy of the plasmid by reciprocal recombination. Phage bearing integrated plasmids can be purified from the larger pool of phage lacking plasmid integrates by growth under the appropriate selective conditions.

Animals↗

Random sequencing of cDNA library derived from partially-fed adult female Haemaphysalis longicornis salivary gland.

A cDNA library was constructed from salivary glands of partially-fed adult female Haemaphysalis longicornis (hard tick). Randomly selected clones were sequenced and a total of 633 sequences were analyzed by bioinformatic programs. The sequences were grouped into 213 clusters, with each cluster being considered to be composed of mRNAs derived from the same gene or closely related genes. About 36% of the mRNA sequences showed significant similarity to known proteins in the non-redundant protein database by the NCBI blastx program and appeared to be coding for functional predicted proteins, whereas the remaining 64% had no similar sequences. Two thirds of the predicted proteins were annotated as basic cellular proteins (housekeeping proteins). Among the functional predicted protein sequences, other than the housekeeping proteins, several protease inhibitors including anticoagulants, two metalloproteases and a potential immunosuppressive protein could be identified. These proteins may play important roles during tick feeding and could be novel anti-tick vaccine candidates.

Amino Acid Sequence↗

Gene capture prediction and overlap estimation in EST sequencing from one or multiple libraries.

BACKGROUND: In expressed sequence tag (EST) sequencing, we are often interested in how many genes we can capture in an EST sample of a targeted size. This information provides insights to sequencing efficiency in experimental design, as well as clues to the diversity of expressed genes in the tissue from which the library was constructed. RESULTS: We propose a compound Poisson process model that can accurately predict the gene capture in a future EST sample based on an initial EST sample. It also allows estimation of the number of expressed genes in one cDNA library or co-expressed in two cDNA libraries. The superior performance of the new prediction method over an existing approach is established by a simulation study. Our analysis of four Arabidopsis thaliana EST sets suggests that the number of expressed genes present in four different cDNA libraries of Arabidopsis thaliana varies from 9155 (root) to 12005 (silique). An observed fraction of co-expressed genes in two different EST sets as low as 25% can correspond to an actual overlap fraction greater than 65%. CONCLUSION: The proposed method provides a convenient tool for gene capture prediction and cDNA library property diagnosis in EST sequencing.

Algorithms↗

Expressed sequence tags and chromosomal localization of cDNA clones from a subtracted retinal pigment epithelium library.

Expressed sequence tags (ESTs) provide useful molecular landmarks for physical mapping and identify the position of an expressed region in the genome. The use of subtracted cDNA libraries enriched for tissue-specific genes as a source of ESTs should reduce the repetitive isolation of constitutively expressed sequences. We report here the sequence tags from the 3'-end region of 58 new directionally cloned cDNAs from a subtracted human retinal pigment epithelium (RPE) cell line library. Eight of the cDNAs have been assigned to human chromosomes using PCR-based EST assays. Chromosomal mapping of subtracted RPE cDNA clones may also help in identifying candidate genes for inherited eye diseases.

Base Sequence↗

Construction of libraries enriched for sequence repeats and jumping clones, and hybridization selection for region-specific markers.

We describe a simple and rapid method for constructing small-insert genomic libraries highly enriched for dimeric, trimeric, and tetrameric nucleotide repeat motifs. The approach involves use of DNA inserts recovered by PCR amplification of a small-insert sonicated genomic phage library or by a single-primer PCR amplification of Mbo I-digested and adaptor-ligated genomic DNA. The genomic DNA inserts are heat denatured and hybridized to a biotinylated oligonucleotide. The biotinylated hybrids are retained on a Vectrex-avidin matrix and eluted specifically. The eluate is PCR amplified and cloned. More than 90% of the clones in a library enriched for (CA)n microsatellites with this approach contained clones with inserts containing CA repeats. We have also used this protocol for enrichment of (CAG)n and (AGAT)n sequence repeats and for Not I jumping clones. We have used the enriched libraries with an adaptation of the cDNA selection method to enrich for repeat motifs encoded in yeast artificial chromosomes.

Base Sequence↗

FastGroup: a program to dereplicate libraries of 16S rDNA sequences.

BACKGROUND: Ribosomal 16S DNA sequences are an essential tool for identifying and classifying microbes. High-throughput DNA sequencing now makes it economically possible to produce very large datasets of 16S rDNA sequences in short time periods, necessitating new computer tools for analyses. Here we describe FastGroup, a Java program designed to dereplicate libraries of 16S rDNA sequences. By dereplication we mean to: 1) compare all the sequences in a data set to each other, 2) group similar sequences together, and 3) output a representative sequence from each group. In this way, duplicate sequences are removed from a library. RESULTS: FastGroup was tested using a library of single-pass, bacterial 16S rDNA sequences cloned from coral-associated bacteria. We found that the optimal strategy for dereplicating these sequences was to: 1) trim ambiguous bases from the 5' end of the sequences and all sequence 3' of the conserved Bact517 site, 2) match the sequences from the 3' end, and 3) group sequences > or =97% identical to each other. CONCLUSIONS: The FastGroup program simplifies the dereplication of 16S rDNA sequence libraries and prepares the raw sequences for subsequent analyses.

Algorithms↗

Mass spectrometric sequencing of individual peptides from combinatorial libraries via specific generation of chain-terminated sequences.

Combinatorial peptide libraries are a versatile tool for drug discovery. On-bead assays identify reactive peptides by enzyme-catalyzed staining and, usually, sequencing by Edman degradation. Unfortunately, the latter method is expensive and time-consuming and requires free N termini of the peptides. A method of rapid and unambiguous peptide sequencing by utilizing synthesis-implemented generation of termination sequences with subsequent matrix-assisted laser desorption ionization time of flight (MALDI-TOF) mass spectrometric analysis is introduced here. The required capped sequences are determined and optimized for a specific peptide library by a computer algorithm implemented in the program Biblio. A total of 99.7% of the sequences of a heptapeptide library sample could be decoded utilizing a single bead for each spectrum. To synthesize these libraries, an optimized capping approach has been introduced.

Algorithms↗

Construction and characterization of phage libraries displaying artificial proteins with random sequences.

Three phage libraries, PL1, PL2, and PL3, displaying artificial proteins with random sequences were constructed. The artificial proteins, which are model of ancestral proteins, are derivatives of the 25 kinds of random proteins with about 140 amino acid residues produced via random mutagenesis and combinatorial recombination. The random proteins were displayed on the surface of filamentous bacteriophage as fusion protein with the pIII coat protein at an estimated average number on the phage particles in PL1, PL2, and PL3 of 0.32, 0.32, and 0.08, respectively. Each library was shown to express 10(5) to 10(6) kinds of random proteins. With the phage libraries displaying long random peptides, we now have an effective selection system to observe in vitro evolution of new functional proteins from artificial proteins with random sequences.

Journal Article↗

[Isolation and analysis of brain-specific sequences from cDNA libraries for various segments of the human brain].

The cDNA libraries in gt10 were constructed from total poly(A)+RNA of human forebrain cortex, cerebellar cortex and medulla oblongata. We selected the clones which gave hybridization signal with brain cDNA only, or gave no signal from these libraries. Expression pattern and structure of two brain-specific clones Hfb1 from forebrain library and Hmob3 from medulla oblongata library were analyzed in detail. Hfb1 hybridized to two different transcripts (about 5 and 2 kb) from frontal cortex, but to a single (longest) from cerebellum. Hfb1 sequence includes 958 nucleotides. Comparison of Hfb1 with the Gene Bank revealed no homology with the sequences present in the Bank. At 3'-end there is poly(A) tail of 24 bases, there is the AATCAA sequence 55 nucleotides upstream which probably serves as a polyadenylation signal. However, AATCAA directs polyadenylation in vitro with very low efficiency. We found no open reading frame in the clone and this is in agreement with the data indicating that brain-specific RNAs has extremely long 3'-untranslated regions. Hmob3 was partially sequences. We compared its primary structure with the sequences from the Gene Bank and revealed no homology. Hmob3 expresses in different parts of human brain and in sceletal muscle but does not express in other tissues.

Base Sequence↗

Generation of cDNA expression libraries enriched for in-frame sequences.

Bacterial cDNA expression libraries are made to reproduce protein sequences present in the mRNA source tissue. However, there is no control over which frame of the cDNA is translated, because translation of the cDNA must be initiated on vector sequence. In a library of nondirectionally cloned cDNAs, only some 8% of the protein sequences produced are expected to be correct. Directional cloning can increase this by a factor of two, but it does not solve the frame problem. We have therefore developed and tested a library construction methodology using a novel vector, pKE-1, with which translation in the correct reading frame confers kanamycin resistance on the host. Following kanamycin selection, the cDNA libraries contained 60-80% open, in-frame clones. These, compared with unselected libraries, showed a 10-fold increase in the number of matches between the cDNA-encoded proteins made by the bacteria and database protein sequences. cDNA sequencing programs will benefit from the enrichment for correct coding sequences, and screening methods requiring protein expression will benefit from the enrichment for authentic translation products.

Base Sequence↗

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning↗

Random AT library: autonomously replicating sequence (ARS) activity of chemically synthesized random sequences for transformation of nonconventional yeast species.

In a search for sequences that confer on bacterial plasmids the capacity of autonomous replication in yeast cells, we chemically synthesized polynucleotides 80 bp in length from an equimolar mixture of A and T. The random AT-polymer population, W80, was inserted into the plasmid YIp5-Kan1 (which carries the markers URA3 and G418(R), but does not replicate in yeast) and amplified in Escherichia coli. This library, representing 10 000 different AT sequences, was transformed into three species of yeast: Saccharomyces cerevisiae, Kluyveromyces lactis and Torulaspora delbrueckii. The aim was to evaluate the frequency, if any, of autonomously replicating sequences (ARSs) in the random sequences. A large number of transformants were obtained from each species. Many of them showed a stable transformed phenotype. Several W80 sequences were found many times for a given species, suggesting that each species preferred particular sequences for ARS function, although they are diverse in their primary sequence. In view of the high frequency and stability of the replicative plasmids found in the different hosts, this small random AT library may be conveniently used as a source of replicative gene vectors for genetic manipulation of many nonconventional yeast species, in place of searching for species-specific chromosomal ARSs.

Ascomycota↗

Venn analysis as part of a bioinformatic approach to prioritize expressed sequence tags from cardiac libraries.

OBJECTIVES: We needed to sort expressed sequence tags (ESTs) from human cardiac expression libraries. DESIGN AND METHODS: We annotated DNA sequence text files of 35,152 cardiac ESTs using our search and annotation tool called Multiblast.pl. We generated lists of the most prevalent ESTs in each library, and using a novel Venn tool, we grouped ESTs that were common to all or exclusive to particular libraries. RESULTS: Hypothetical protein KIAA0553 was expressed 120 times among 917 ESTs from an adult cardiac library (13.1%) compared only once among 8075 ESTs from fetal cardiac libraries (P < 10(-114)), this was confirmed using Northern analysis. We collated biochemical features of KIAA0553 and determined DNA polymorphism frequencies. We also used the Venn tool to specify genes that were uniquely expressed in hypertrophic cardiomyocytes. CONCLUSIONS: Annotating ESTs and sorting them using Venn analysis can help specify new candidate disease genes from the current lists of "hypothetical proteins".

Amino Acid Sequence↗

Discovering distinct genes represented in 29,570 clones from infant brain cDNA libraries by applying sequencing by hybridization methodology.

To discover all distinct human genes and to determine their patterns of expression across different cell types, developmental stages, and physiological conditions, a procedure is needed for fast, mutual comparison of hundreds of thousands (and perhaps millions) of clones from cDNA libraries, as well as their comparison against data bases of sequenced DNA. In a pilot study, 29,570 clones in duplicate from both original and normalized, directional, infant brain cDNA libraries were hybridized with 107-215 heptamer oligonucleotide probes to obtain oligonucleotide sequence signatures (OSSs). The OSSs were compared and clustered based on mutual similarity into 16,741 clusters, each corresponding to a distinct cDNA. A number of distinct cDNAs were successfully recognized by matching their 107-probe OSSs against GenBank entries, indicating the possibility of sequence recognition with only a few hundred randomly chosen oligomers.

Animals↗

Creating multiple-crossover DNA libraries independent of sequence identity.

We have developed, experimentally implemented, and modeled in silico a methodology named SCRATCHY that enables the combinatorial engineering of target proteins, independent of sequence identity. The approach combines two methods for recombining genes: incremental truncation for the creation of hybrid enzymes and DNA shuffling. First, incremental truncation for the creation of hybrid enzymes is used to create a comprehensive set of fusions between fragments of genes in a DNA homology-independent fashion. This artificial family is then subjected to a DNA-shuffling step to augment the number of crossovers. SCRATCHY libraries were created from the glycinamide-ribonucleotide formyltransferase (GART) genes from Escherichia coli (purN) and human (hGART). The developed modeling framework eSCRATCHY provides insight into the effect of sequence identity and fragmentation length on crossover statistics and draws contrast with DNA shuffling. Sequence analysis of the naive shuffled library identified members with up to three crossovers, and modeling predictions are in good agreement with the experimental findings. Subsequent in vivo selection in an auxotrophic E. coli host yielded functional hybrid enzymes containing multiple crossovers.

Algorithms↗

Genomic characterization of Helicobacter hepaticus: ordered cosmid library and comparative sequence analysis.

Helicobacter hepaticus is an important pathogen in laboratory mice and induces the development of liver tumors and gastrointestinal disease in susceptible strains of mice. In this study, a miniset of 36 cosmid clones from a genomic library of H. hepaticus was ordered and grouped into four large contigs representing approximately 1 Mb of the H. hepaticus genome using PCR, DNA sequencing, Southern and dot-blot hybridization and pulsed-field gel electrophoresis. From the 200-300 terminal nucleotide sequences of 38 cosmid clones, 56 coding regions were predicted, of which 51 were found to have orthologs in the public databases and five appeared to be unique to H. hepaticus. Of these 51 genes, 36 have orthologs in Helicobacter pylori and 25 display the highest sequence similarity to H. pylori. However, chromosomal positions of these genes are not conserved between these two helicobacters. In addition, 10 H. hepaticus genes had the highest sequence similarity to orthologs in Campylobacter jejuni. The GC content in a randomly selected 21-kb H. hepaticus genomic sequence was 35.8%, which approximates the average between H. pylori (39%) and C. jejuni (30.6%). These results demonstrate that: (1) H. hepaticus is more closely related to H. pylori than C. jejuni; (2) significant genomic alterations exist between H. hepaticus and H. pylori, including gene organization, protein sequences and GC content, probably in part due to specific adaptation to distinct ecological niches.

Animals↗

Fragment walking for long DNA sequencing by using a library as small as 16 primers.

A DNA sequence can be rapidly and efficiently determined by digesting it into segments small enough to be sequenced at one time, then assembling them into contigs by searching for overlaps. Fragment walking achieves this without subcloning or preparing many kinds of primers; fragments obtained by digesting a template DNA are sequenced in parallel directly from the fragment mixture by using a set of 16 primers. Since the sequence adjacent to each cutting site is determined together with the fragment sequence, the contiguous fragment can be easily determined. The complete template DNA sequence can thus efficiently determined. The sequencing of pUC19 (2.7 kb) from both sides was done in one step by using only 15 fragments. The redundancy was only about 1.3, greatly reducing base-reading redundancy. Automating this strategy would increase the speed and efficiency of large-scale DNA sequencing.

Base Sequence↗

An expressed sequence tag (EST) library from developing fruits of an Hawaiian endemic mint (Stenogyne rugosa, Lamiaceae): characterization and microsatellite markers.

BACKGROUND: The endemic Hawaiian mints represent a major island radiation that likely originated from hybridization between two North American polyploid lineages. In contrast with the extensive morphological and ecological diversity among taxa, ribosomal DNA sequence variation has been found to be remarkably low. In the past few years, expressed sequence tag (EST) projects on plant species have generated a vast amount of publicly available sequence data that can be mined for simple sequence repeats (SSRs). However, these EST projects have largely focused on crop or otherwise economically important plants, and so far only few studies have been published on the use of intragenic SSRs in natural plant populations. We constructed an EST library from developing fleshy nutlets of Stenogyne rugosa principally to identify genetic markers for the Hawaiian endemic mints. RESULTS: The Stenogyne fruit EST library consisted of 628 unique transcripts derived from 942 high quality ESTs, with 68% of unigenes matching Arabidopsis genes. Relative frequencies of Gene Ontology functional categories were broadly representative of the Arabidopsis proteome. Many unigenes were identified as putative homologs of genes that are active during plant reproductive development. A comparison between unigenes from Stenogyne and tomato (both asterid angiosperms) revealed many homologs that may be relevant for fruit development. Among the 628 unigenes, a total of 44 potentially useful microsatellite loci were predicted. Several of these were successfully tested for cross-transferability to other Hawaiian mint species, and at least five of these demonstrated interesting patterns of polymorphism across a large sample of Hawaiian mints as well as close North American relatives in the genus Stachys. CONCLUSION: Analysis of this relatively small EST library illustrated a broad GO functional representation. Many unigenes could be annotated to involvement in reproductive development. Furthermore, first tests of microsatellite primer pairs have proven promising for the use of Stenogyne rugosa EST SSRs for evolutionary and phylogeographic studies of the Hawaiian endemic mints and their close relatives. Given that allelic repeat length variation in developmental genes of other organisms has been linked with morphological evolution, these SSRs may also prove useful for analyses of phenotypic differences among Hawaiian mints.

5' Untranslated Regions↗