PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Large-scale gene discovery in human airway epithelia reveals novel transcripts.

The airway epithelium represents an important barrier between the host and the environment. It is a first site of contact with pathogens, particulates, and other stimuli, and has evolved the means to dynamically respond to these challenges. In an effort to define the transcript profile of airway epithelia, we created and sequenced cDNA libraries from cystic fibrosis (CF) and non-CF epithelia and from human lung tissue. Sequencing of these libraries produced approximately 53,000 3'-expressed sequence tags (3'-ESTs). From these, a nonredundant UniGene set of more than 19,000 sequences was generated. Despite the relatively small contribution of airway epithelia to the total mass of the lung, focused gene discovery in this tissue yielded novel results. The ESTs included several thousand transcripts (6,416) not previously identified from cDNA sequences as expressed in the lung. Among the abundant transcripts were several genes involved in host defense. Most importantly, the set also included 879 3'-ESTs that appear to be novel sequences not previously represented in the National Center for Biotechnology Information UniGene collection. This UniGene set should be useful for studies of pulmonary diseases involving the airway epithelium including cystic fibrosis, respiratory infections and asthma. It also provides a reagent for large-scale expression profiling.

Adult↗

Detailed characterization of a human 8q24.1 microdissection library and generation of "sequence-tagged sites".

A total of 533 clones from a human 8q24.1 microdissection library was analyzed by automatic DNA sequencing. Three hundred and thirty-seven different insert sequences were found. The insert size ranged from 45 to 376 bp (mean, 170 bp). Eighty-six percent (291/337) of these sequences were free of repetitive DNA. Each of 19 clones tested was successfully translated into a sequence-tagged site. We conclude that the microdissection library is a rich source of DNA markers for 8q24.1.

Base Sequence↗

Genes associated with the end of dormancy in grapes.

A grape bud EST library was constructed and 4270 ESTs sequenced. The library clones were arrayed for the purpose of investigating the level of gene expression over time, particularly leading up to the buds' release from dormancy. The arrays were hybridized with P(33)-labeled probes produced from samples of buds collected at weekly intervals. These probes covered the time from 9 weeks prior to bud burst until just after the emergence of the shoots. Expression patterns from these genes have been examined. It was found that 74% of the genes in the data set were homologous to known proteins. Genes were then assigned to functional categories according to their primary BLAST match. Of these 13% were involved with photosynthesis, 13% with disease resistance and defense, 5% energy, 12% metabolism, 20% protein production and processing, 25% cell structure and plant growth and the remaining 12% were unclassified The expression pattern of a selection of "candidate" genes retrieved from literature previously reporting an association with dormancy changes was assessed. On closer examination most of these genes relate to the oxidative processes and stress responses within the cell. The results of this study show that even in the dormant state, gene expression in the buds is high.

Databases, Factual↗

Analysis of region-specific library constructed by sequence-independent amplification of microdissected fragments surrounding weaver (wv) gene on mouse chromosome 16.

The C3-C4 region of mouse chromosome 16 was microdissected and amplified directly by sequence-independent amplification (SIA). The SIA product was proved to originate from the microdissected region by fluorescence in situ hybridization (FISH) and was cloned into the PCR II vector (mean insert size 506 bp). Colony hybridization showed that about 59% of the clones contained either unique or low copy number sequences. Southern blot analysis of 100 unique clones demonstrated that 50 clones hybridized with single (33 clones) or multiple (17 clones) bands on blots of DNA from a hamster-mouse hybrid cell line that contains mouse chromosome 16, 13 clones hybridized with mouse but not with the hamster-mouse hybrid DNA, 19 clones contained repetitive sequences, and the remaining 18 clones failed to yield bands. One third of the 100 unique clones hybridized to human genomic DNA. Thirty-three clones were sequenced. None of them was found in GenBank. Our results demonstrate that this relatively simple method of microdissection and cloning can produce a library of good quality.

Animals↗

Tabu search algorithm for DNA sequencing by hybridization with isothermic libraries.

In this paper, a problem of isothermic DNA sequencing by hybridization (SBH) is considered. In isothermic SBH a new type of oligonucleotide libraries is used. The library consists of oligonucleotides of different lengths depending on an oligonucleotide content. It is assumed that every oligonucleotide in such a library has an equal melting temperature. Each nucleotide adds its increment to the oligonucleotide temperature and it is assumed that A and T add 2 degrees C and C and G add 4 degrees C. The hybridization experiment using isothermic libraries should provide data with a lower number of errors due to an expected similarity of melting temperatures. From the computational point of view the problem of isothermic DNA sequencing with errors is hard, similarly like its classical counterpart. Hence, there is a need for developing heuristic algorithms that construct good suboptimal solutions. The aim of the paper is to propose a heuristic algorithm based on tabu search approach. The algorithm solves the problem with both positive and negative errors. Results of an extensive computational experiment are presented, which prove the high quality of the proposed method.

Algorithms↗

Construction and evaluation of a hncDNA library of human 12p transcribed sequences derived from a somatic cell hybrid.

An arrayed library of human heterogeneous nuclear complementary (hnc) DNA was constructed from a somatic cell hybrid (M28) containing an i(12p) marker as the sole human chromosome. Heterogeneous nuclear (hn) RNA of M28 was used to synthesize first-strand hncDNA with a primer (RT) containing a random hexanucleotide at its 3' end. Specific amplification of human sequences from this hncDNA was performed using Alu primers in combination with the RT primer. The products were directionally cloned and an arrayed library was constructed. Experiments indicated that all clones were derived from transcribed sequences. A number of randomly isolated clones were evaluated by Southern and Northern experiments, sequence analysis, and PCR. At least 80% of these clones were of human 12p origin. The number of independent clones in the library was estimated to be approximately 550. Using 60 hncDNA clones as probes, 6 showed positive signals on Northern blots. For 3 of these, the corresponding cDNAs were isolated: clone CD60A1 codes for the cation-dependent mannose 6-phosphate receptor, clone CC6 is a human homologue of the bovine 39-kDa nuclear-encoded NADH:ubiquinone oxidoreductase subunit, and CD18 belongs to the family of tumor necrosis factor receptor proteins. Southern experiments showed the 3 cDNAs to map to human chromosome 12p as expected. Taken together these results show that the generation of a hncDNA library is a useful tool for the isolation of unknown genes located on a human chromosome (fragment) present in a somatic cell hybrid.

Animals↗

Statistical theory for protein ensembles with designed energy landscapes.

Combinatorial protein libraries provide a promising route to investigate the determinants and features of protein folding and to identify novel folding amino acid sequences. A library of sequences based on a pool of different monomer types are screened for folding molecules, consistent with a particular foldability criterion. The number of sequences grows exponentially with the length of the polymer, making both experimental and computational tabulations of sequences infeasible. Herein a statistical theory is extended to specify the properties of sequences having particular values of global energetic quantities that specify their energy landscape. The theory yields the site-specific monomer probabilities. A foldability criterion is derived that characterizes the properties of sequences by quantifying the energetic separation of the target state from low-energy states in the unfolded ensemble and the fluctuations of the energies in the unfolded state ensemble. For a simple lattice model of proteins, excellent agreement is observed between the theory and the results of exact enumeration. The theory may be used to provide a quantitative framework for the design and interpretation of combinatorial experiments.

Models, Statistical↗

Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.

MOTIVATION: Computer-based selection of entries from sequence databases with respect to a related functional description, e.g. with respect to a common cellular localization or contributing to the same phenotypic function, is a difficult task. Automatic semantic analysis of annotations is not only hampered by incomplete functional assignments. A major problem is that annotations are written in a rich, non-formalized language and are meant for reading by a human expert. This person can extract from the text considerably more information than is immediately apparent due to his extended biological background knowledge and logical reasoning. APPROACH: A technique of automated annotation evaluation based on a combination of lexical analysis and the usage of biological rule libraries has been developed. The proposed algorithm generates new functional descriptors from the annotation of a given entry using the semantic units of the annotation as prepositions for implications executed in accordance with the rule library. RESULTS: The prototype of a software system, the Meta_A(nnotator) program, is described and the results of its application to sequence attribute assignment and sequence selection problems, such as cellular localization and sequence domain annotation of SWISS-PROT entries, are presented. The current software version assigns useful subcellular localization qualifiers to approximately 88% of all SWISS-PROT entries. As shown by demonstrative examples, the combination of sequence and annotation analysis is a powerful approach for the detection of mutual annotation/sequence inconsistencies. AVAILABILITY: Results for the cellular localization assignment can be viewed at the URL http://www.bork. embl-heidelberg.de/CELL_LOC/CELL_LOC.html.

Algorithms↗

Model for a transcript map of human chromosome 21: isolation of new coding sequences from exon and enriched cDNA libraries.

The construction of a transcriptional map for human chromosome 21 requires the generation of a specific catalogue of genes, together with corresponding mapping information. Towards this goal, we conducted a pilot study on a pool of random chromosome 21 cosmids representing 2 Mb of non-contiguous DNA. Exon-amplification and cDNA selection methods were used in combination to extract the coding content from these cosmids, and to derive expressed sequences libraries. These libraries and the source cosmid library were arrayed at high density for hybridisation screening. A strategy was used which related data obtained by multiple hybridisations of clones originating from one library, screened against the other libraries. In this way, it was possible to integrate the information with the physical map and to compare the gene recovery rate of each technique. cDNAs and exons were grouped into bins delineated by EcoRI cosmid fragments, and a subset of 91 cDNAs and 29 exons have been sequenced. These sequences defined 79 non-overlapping potential coding segments distributed in 24 transcriptional units, which were mapped along 21q. Northern blot analysis performed for a subset of cDNAs indicated the existence of a cognate transcript. Comparison to databases indicated three segments matching to known chromosome 21 genes: PFKL, COL6A1 and S100B and six segments matching to unmapped anonymous expressed sequence tags (ESTs). At the translated nucleotide level, strong homologies to known proteins were found with ATP-binding transporters of the ABC family and the dihydroorotase domain of pyrimidine synthetases. These data strongly suggest that bona fide partial genes have been isolated. Several of the newly isolated transcriptional units map to clinically important regions, in particular those involved in Down's syndrome, progressive myoclonus epilepsia and auto-immune polyglandular disease. The study presented here illustrates the complementarity of exon-amplification and cDNA selection techniques for generating a large resource of new expressed landmarks, which contribute to the construction of a chromosome 21 transcript map.

Chromosome Mapping↗

Cloning and expression of the dihydroorotate dehydrogenase from Toxoplasma gondii.

A full-length dihydroorotate dehydrogenase (DHODase) sequence was cloned from a Toxoplasma gondii tachyzoite cDNA library. The sequence was most similar to family 2 DHODases, and had a calculated molecular mass of 65.1 kDa. The full-length and two N-terminally truncated T. gondii DHODase sequences were expressed as recombinant proteins. One of the truncated sequences complemented a DHODase-deficient bacterial host.

Amino Acid Sequence↗

Detecting DNA-binding helix-turn-helix structural motifs using sequence and structure information.

In this work, we analyse the potential for using structural knowledge to improve the detection of the DNA-binding helix-turn-helix (HTH) motif from sequence. Starting from a set of DNA-binding protein structures that include a functional HTH motif and have no apparent sequence similarity to each other, two different libraries of hidden Markov models (HMMs) were built. One library included sequence models of whole DNA-binding domains, which incorporate the HTH motif, the second library included shorter models of 'partial' domains, representing only the fraction of the domain that corresponds to the functionally relevant HTH motif itself. The libraries were scanned against a dataset of protein sequences, some containing the HTH motifs, others not. HMM predictions were compared with the results obtained from a previously published structure-based method and subsequently combined with it. The combined method proved more effective than either of the single-featured approaches, showing that information carried by motif sequences and motif structures are to some extent complementary and can successfully be used together for the detection of DNA-binding HTHs in proteins of unknown function.

Amino Acid Sequence↗

Genomic and cDNA sequence tags of the hyperthermophilic archaeon Pyrobaculum aerophilum.

The hyperthermophilic archaeum, Pyrobaculum aerophilum, grows optimally at 100 degrees C with a doubling time of 180 min. It is a member of the phylogenetically ancient Thermoproteales order, but differs significantly from all other members by its facultatively aerobic metabolism. Due to its simple cultivation requirements and its nearly 100% plating efficiency, it was chosen as a model organism for studying the genome organization of hyperthermophilic ancient archaea. By a G+C content of the DNA of 52 mol%, sequence analysis was easily possible. At least some of the mRNA of P. aerophilum carried poly-A tails facilitating the construction of a cDNA library. 245 sequence tags of a poly-A primed cDNA library and 55 sequence tags from a 1-2 kb Sau3AI-fragment containing genomic library were analyzed and the corresponding amino acid sequences compared with protein sequences from databases. Fourteen percent of the cDNA and >9% of genomic DNA sequence tags revealed significant similarities to proteins in the databases. Matches were obtained to proteins from archaeal, bacterial and eukaryal sources. Some sequences showed greatest similarity to eukaryal rather than to bacterial versions of proteins, other matches were found to proteins which had previously only been found in eukaryotes.

Archaea↗

Identification of anti-TNFalpha peptides with consensus sequence.

Phage displayed peptide library was used to select tumor necrosis factor alpha (TNFalpha) binding peptides. After three sequential rounds of biopanning, some linear TNFalpha-binding peptides were identified from a 12-mer peptide library. A consensus sequence (L/M)HEL(Y/F)(L/M)X(W/Y/F), where X might be variable residue, was deduced from sequences of these peptides. The phages bearing these peptides showed specific binding to immobilized TNFalpha, with over 80% of phages bound being competitively eluted by free TNFalpha. To confirm the binding activity and to explore further functional properties, three peptides with typical structure were selected and expressed as GST-fused protein. These recombinant peptides effectively competed for [125I]TNFalpha binding to TNFR1 in a dose-dependent manner, with IC(50) from 10 to 160 microM. Furthermore, the GST-fused derivatives showed inhibitory effects on TNFalpha-induced cytotoxicity. Taken together, these data demonstrate that the TNFalpha-binding peptides are effective antagonists of TNFalpha and the deduced motif might be useful in development of novel low molecular weight anti-TNFalpha drugs.

Bacteriophages↗

Cloning and structure of cDNA encoding alpha-latrotoxin from black widow spider venom.

cDNA encoding the putative alpha-latrotoxin precursor was isolated from spider venom glands cDNA library and sequenced. The cDNA contained the 4203 base-pair open reading frame corresponding to the 156,855-Da protein composed of 1401 amino acids. Computer analysis of the deduced primary structure revealed the presence of various internal imperfect repeats mainly in its central and C-terminal regions.

Amino Acid Sequence↗

Mapping of transcription start sites in Saccharomyces cerevisiae using 5' SAGE.

A minimally addressed area in Saccharomyces cerevisiae research is the mapping of transcription start sites (TSS). Mapping of TSS in S.cerevisiae has the potential to contribute to our understanding of gene regulation, transcription, mRNA stability and aspects of RNA biology. Here, we use 5' SAGE to map 5' TSS in S.cerevisiae. Tags identifying the first 15-17 bases of the transcripts are created, ligated to form ditags, amplified, concatemerized and ligated into a vector to create a library. Each clone sequenced from this library identifies 10-20 TSS. We have identified 13,746 unique, unambiguous sequence tags from 2231 S.cerevisiae genes. TSS identified in this study are consistent with published results, with primer extension results described here, and are consistent with expectations based on previous work on transcription initiation. We have aligned the sequence flanking 4637 TSS to identify the consensus sequence A(A(rich))5NPyA(A/T)NN(A(rich))6, which confirms and expands the previous reported PyA(A/T)Pu consensus pattern. The TSS data allowed the identification of a previously unrecognized gene, uncovered errors in previous annotation, and identified potential regulatory RNAs and upstream open reading frames in 5'-untranslated region.

Base Sequence↗

Discovery of active proteins directly from combinatorial randomized protein libraries without display, purification or sequencing: identification of novel zinc finger proteins.

We have successfully linked protein library screening directly with the identification of active proteins, without the need for individual purification, display technologies or physical linkage between the protein and its encoding sequence. By using 'MAX' randomization we have rapidly constructed 60 overlapping gene libraries that encode zinc finger proteins, randomized variously at the three principal DNA-contacting residues. Expression and screening of the libraries against five possible target DNA sequences generated data points covering a potential 40,000 individual interactions. Comparative analysis of the resulting data enabled direct identification of active proteins. Accuracy of this library analysis methodology was confirmed by both in vitro and in vivo analyses of identified proteins to yield novel zinc finger proteins that bind to their target sequences with high affinity, as indicated by low nanomolar apparent dissociation constants.

Binding Sites↗

Recovery of novel bacterial diversity from a forested wetland impacted by reject coal.

Sulphide mineral mining together with improperly contained sulphur-rich coal represents a significant environmental problem caused by leaching of toxic material. The Savannah River Site's D-area harbours a 22-year-old exposed reject coal pile (RCP) from which acidic, metal rich, saline runoff has impacted an adjacent forested wetland. In order to assess the bacterial community composition of this region, composite sediment samples were collected at three points along a contamination gradient (high, middle and no contamination) and processed for generation of bacterial and archaeal 16S rDNA clone libraries. Little sequence overlap occurred between the contaminated (RCP samples) and unimpacted sites, indicating that the majority of 16S rDNAs retrieved from the former represent organisms selected by the acidic runoff. Archaeal diversity within the RCP samples consisted mainly of sequences related to the genus Thermoplasma and to sequences of a novel type. Bacterial RCP libraries contained 16S rRNA genes related to isolates (Acidiphilium sp., Acidobacterium capsulatum, Ferromicrobium acidophilium and Leptospirillum ferrooxidans) and environmental clones previously retrieved from acidic habitats, including ones phylogenetically associated with organisms capable of sulphur and iron metabolism. These libraries also exhibited particularly novel 16S rDNA types not retrieved from other acid mine drainage habitats, indicating that significant diversity remains to be detected in acid mine drainage-type systems.

Archaea↗

Random exploration of the Kluyveromyces lactis genome and comparison with that of Saccharomyces cerevisiae.

The genome of the yeast Kluyveromyces lactis was explored by sequencing 588 short tags from two random genomic libraries (random sequenced tags, or RSTs), representing altogether 1.3% of the K. lactis genome. After systematic translation of the RSTs in all six possible frames and comparison with the complete set of proteins predicted from the Saccharomyces cerevisiae genomic sequence using an internally standardized threshold, 296 K.lactis genes were identified of which 292 are new. This corresponds to approximately 5% of the estimated genes of this organism and triples the total number of identified genes in this species. Of the novel K.lactis genes, 169 (58%) are homologous to S.cerevisiae genes of known or assigned functions, allowing tentative functional assignment, but 59 others (20%) correspond to S.cerevisiae genes of unknown function and previously without homolog among all completely sequenced genomes. Interestingly, a lower degree of sequence conservation is observed in this latter class. In nearly all instances in which the novel K.lactis genes have homologs in different species, sequence conservation is higher with their S.cerevisiae counterparts than with any of the other organisms examined. Conserved gene order relationships (synteny) between the two yeast species are also observed for half of the cases studied.

Base Sequence↗