PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 739 records · Page 41Linked to original sources

Human antibodies with specificity for the C2 domain of factor VIII are derived from VH1 germline genes.

A serious complication in hemophilia care is the development of factor VIII (FVIII) neutralizing antibodies (inhibitors). The authors used V gene phage display technology to define human anti-FVIII antibodies at the molecular level. The IgG4-specific, variable, heavy-chain gene repertoire of a patient with acquired hemophilia was combined with a nonimmune, variable, light-chain gene repertoire for display as single-chain variable domain antibody fragments (scFv) on filamentous phage. ScFv were selected by 4 rounds of panning on immobilized FVIII light chain. Sequence analysis revealed that isolated scFv were characterized by V(H) domains encoded by germline genes DP-10, DP-14, and DP-88, all belonging to the V(H)1 gene family. All clones displayed extensive hypermutation and were characterized by unusually long CDR3 sequences of 20 to 23 amino acids. Immunoprecipitation revealed that all scFv examined bound to the C2 domain of FVIII. Furthermore, isolated scFv competed with an inhibitory murine monoclonal antibody for binding to the C2 domain. Even though scFv bound FVIII with high affinity, they did not inhibit FVIII activity. Interestingly, the addition of scFv diminished the inhibitory potential of patient-derived antibodies with C2 domain specificity. These results suggest that the epitope of a significant portion of anti-C2 domain antibodies overlaps with that of the scFv isolated in this study. (Blood. 2000;95:558-563)

Adult↗

Microarray-based genetics of cardiac malformations.

One of the most revolutionary approaches in human genomics is DNA microarray technology. Latest developments have brought this technology to a widespread use. In this paper we discuss its usefulness especially for the study of the genetic component in congenital heart disease as a model of multifactorial disease and the possible clinical applications in the near future. Malformations of the heart and blood vessels account for the largest number of human birth defects. The susceptibility of the heart to developmental anomalies reflects the complexity of the morphogenetic events responsible for the heart formation. The genetics of congenital heart disease points to the existence of powerful disease modifiers. Tissue analysis of gene expression with cDNA microarrays provides a measure of transcriptional or posttranscriptional regulation. Large-scale partial sequencing of cDNA libraries generating expressed sequence tags is an effective means of discovering novel genes and characterizing transcription patterns in different organs and tissues. The qualitative and quantitative analysis of genes expressed in cardiac tissue by means of comparison of expression patterns related to the normal and to the pathological tissue may be of great importance for the study of cardiac pathologies. The variation in phenotypic penetrance and severity suggests that if we can identify high-risk individuals, a reduction in infant morbidity might be possible by altering environmental or maternal factors.

Gene Expression Profiling↗

Robust accurate identification of peptides (RAId): deciphering MS2 data using a structured library search with de novo based statistics.

MOTIVATION: The key to MS -based proteomics is peptide sequencing. The major challenge in peptide sequencing, whether library search or de novo, is to better infer statistical significance and better attain noise reduction. Since the noise in a spectrum depends on experimental conditions, the instrument used and many other factors, it cannot be predicted even if the peptide sequence is known. The characteristics of the noise can only be uncovered once a spectrum is given. We wish to overcome such issues. RESULTS: We designed RAId to identify peptides from their associated tandem mass spectrometry data. RAId performs a novel de novo sequencing followed by a search in a peptide library that we created. Through de novo sequencing, we establish the spectrum-specific background score statistics for the library search. When the database search fails to return significant hits, the top-ranking de novo sequences become potential candidates for new peptides that are not yet in the database. The use of spectrum-specific background statistics seems to enable RAId to perform well even when the spectral quality is marginal. Other important features of RAId include its potential in de novo sequencing alone and the ease of incorporating post-translational modifications.

Algorithms↗

Systematic performance evaluation and application validation of an end-to-end NGS workstation.

Next-generation sequencing (NGS) library preparation is a core component of precision genomics, but it is commonly constrained by inefficiency, variability, and low throughput of manual protocols. To address these limitations, we developed and systematically evaluated a fully automated NGS workstations and further validated its performance across representative application scenarios. The automated system reduced total processing time from 8 to 10 to 4–6 h. At the same time, it maintained similar performance in pre-library metric, including DNA yield and fragment size, as well as post-capture sequencing metrics (Q30 > 90%, mapping rates > 95%, on-target rates 85–90%). The duplication rate was reduced to 5–8%, compared with 10–15% for manual methods, indicating increased library complexity. Bioinformatic evaluation of inter-species read mapping showed minimal cross-contamination, with a maximum contamination ratio of 0.0003%, indicating effective sample isolation in the automated workflow. High concordance in variant detection was observed between automated and manual workflows. Overall, this automated workstation provides a standardized and reproducible workflow that supports scalable precision genomics applications.

High-Throughput Nucleotide Sequencing↗

Isolation of rapidly evolving genomic sequences: construction of a differential library and identification of a human DNA fragment that does not hybridize to chimpanzee DNA.

A differential library enriched in rapidly evolving human genomic sequences was obtained by phenol-enhanced hybridization of human genomic DNA with an excess of chimpanzee DNA. A DNA fragment 110 bp in length that did not hybridize to either chimpanzee or other primate DNA was identified in this library. It was shown to be a substantially diverged member of the human beta satellite family of tandem repeats. The genomic sequences homologous to the fragment were located on the short arms of human acrocentric chromosomes by in situ hybridization. The human-specific fragment failed to hybridize with RNA from different human tissues. The human-specific fragment exhibits a remarkable level of DNA polymorphism in humans and may be used in the identification of human tissue samples, in the selection of human/rodent somatic cell hybrids containing human acrocentric chromosomes, and in the mapping of these chromosomes.

Animals↗

Nucleic acid library construction using synthetic DNA constructs.

This chapter outlines seven synthetic and molecular biology techniques that allow the controlled synthesis of nucleic acid libraries. Specifically: (1) The high-diversity chemical synthesis of point mutations; (2) the high-diversity chemical synthesis of point deletions; (3) the split-bead approach for constructing point mutation or deletion libraries with limited sequence diversity; (4) pool deprotection, gel purification, and quality-control techniques; (5) large-scale polymerase chain reaction amplification for the generation of high-diversity double-stranded deoxyribonucleic acid libraries; (6) type II restriction enzyme digestion techniques for the construction of long-sequence libraries containing minimal fixed sequence; and (7) extension techniques for the rapid synthesis of long, low-diversity oligonucleotide sequences.

DNA↗

Construction and analysis of cDNA library of Necator americanus third stage larvae.

OBJECTIVE: To obtain the genetic information on Necator americanus and to search for the purpose genes. METHODS: mRNA was isolated from the third stage larvae of Necator americanus maintained in hamsters. Double strand cDNA was synthesized and ligated to lambda ZAPII vector to construct the cDNA library. Expressed sequence tages (ESTs) were obtained by single pass sequencing of randomly isolated cDNA clones from the established library. RESULTS: A cDNA library of N. americanus was successfully constructed with high recombinant efficiency. The titer of unamplified library was 1 x 10(7). The insert size was about 750-3,000 bp. Of 11 ESTs obtained from the library, 7 have a significant homology with certain functional genes. CONCLUSION: A high quality and high representative cDNA library of N. americanus was constructed at the first time and some functional genes were identified from the library by ESTs.

Animals↗

Construction of small-insert genomic DNA libraries highly enriched for microsatellite repeat sequences.

We describe an efficient method for the construction of small-insert genomic libraries enriched for highly polymorphic, simple sequence repeats. With this approach, libraries in which 40-50% of the members contain (CA)n repeats are produced, representing an approximately 50-fold enrichment over conventional small-insert genomic DNA libraries. Briefly, a genomic library with an average insert size of less than 500 base pairs was constructed in a phagemid vector. Amplification of this library in a dut ung strain of Escherichia coli allowed the recovery of the library as closed circular single-stranded DNA with uracil frequently incorporated in place of thymine. This DNA was used as a template for second-strand DNA synthesis, primed with (CA)n or (TG)n oligonucleotides, at elevated temperatures by a thermostable DNA polymerase. Transformation of this mixture into wild-type E. coli strains resulted in the recovery of primer-extended products as a consequence of the strong genetic selection against single-stranded uracil-containing DNA molecules. In this manner, a library highly enriched for the targeted microsatellite-containing clones was recovered. This approach is widely applicable and can be used to generate marker-selected libraries bearing any simple sequence repeat from cDNAs, whole genomes, single chromosomes, or more restricted chromosomal regions of interest.

Animals↗

Removal of polyA tails from full-length cDNA libraries for high-efficiency sequencing.

We have developed a method to overcome sequencing problems caused by the presence of homopolymer stretches, such as polyA/T, in cDNA libraries. PolyA tails are shortened by cleaving before cDNA cloning with type IIS restriction enzymes, such as GsuI, placed next to the oligo-dT used to prime the polyA tails of mRNAs. We constructed four rice Cap-Trapper-selected, full-length normalized cDNA libraries, of which the average residual polyA tail was 4 bases or shorter in most of the clones analyzed Because of the removal of homopolymeric stretches, libraries prepared with this method can be used for direct sequencing and transcriptional sequencing without the slippage observed for libraries prepared with currently available methods, thus improving sequencing accuracy, operations, and throughput.

Animals↗

Identification of eukaryotic open reading frames in metagenomic cDNA libraries made from environmental samples.

Here we describe the application of metagenomic technologies to construct cDNA libraries from RNA isolated from environmental samples. RNAlater (Ambion) was shown to stabilize RNA in environmental samples for periods of at least 3 months at -20 degrees C. Protocols for library construction were established on total RNA extracted from Acanthamoeba polyphaga trophozoites. The methodology was then used on algal mats from geothermal hot springs in Tengchong county, Yunnan Province, People's Republic of China, and activated sludge from a sewage treatment plant in Leicestershire, United Kingdom. The Tenchong libraries were dominated by RNA from prokaryotes, reflecting the mainly prokaryote microbial composition. The majority of these clones resulted from rRNA; only a few appeared to be derived from mRNA. In contrast, many clones from the activated sludge library had significant similarity to eukaryote mRNA-encoded protein sequences. A library was also made using polyadenylated RNA isolated from total RNA from activated sludge; many more clones in this library were related to eukaryotic mRNA sequences and proteins. Open reading frames (ORFs) up to 378 amino acids in size could be identified. Some resembled known proteins over their full length, e.g., 36% match to cystatin, 49% match to ribosomal protein L32, 63% match to ribosomal protein S16, 70% to CPC2 protein. The methodology described here permits the polyadenylated transcriptome to be isolated from environmental samples with no knowledge of the identity of the microorganisms in the sample or the necessity to culture them. It has many uses, including the identification of novel eukaryotic ORFs encoding proteins and enzymes.

Acanthamoeba↗

Polyclonal Fab phage display libraries with a high percentage of diverse clones to Cryptosporidium parvum glycoproteins.

The protozoan parasite Cryptosporidium parvum is regarded as a major public health problem world-wide, especially for immunocompromised individuals. Although no effective therapy is presently available, specific immune responses prevent or terminate cryptosporidiosis and passively administered antibodies have been found to reduce the severity of infection. Therefore, as an immunotherapeutic approach against cryptosporidiosis, we set out to develop C. parvum-specific polyclonal antibody libraries, standardised, perpetual mixtures of polyclonal antibodies, for which the genes are available. A combinatorial Fab phage display library was generated from the antibody variable region gene repertoire of mice immunised with C. parvum surface and apical complex glycoproteins which are believed to be involved in mediating C. parvum attachment and invasion. The variable region genes used to construct this starting library were shown to be diverse by nucleotide sequencing. The library was subjected to one round of antigen selection on C. parvum glycoproteins or a C. parvum oocyst/sporozoite preparation. The two selected libraries showed specific reactivity to the glycoproteins as well as to the oocyst/sporozoite preparation, with 50-73% antigen-reactive members. Fingerprint analysis of individual clones from the two antigen-selected libraries showed high diversity, confirming the polyclonality of the selected libraries. Furthermore, immunoblot analysis on the oocyst/sporozoite and glycoprotein preparations with selected library phage showed reactivity to multiple bands, indicating diversity at the antigen level. These C. parvum-specific polyclonal Fab phage display libraries will be converted to libraries of polyclonal full-length antibodies by mass transfer of the selected heavy and light chain variable region gene pairs to a mammalian expression vector. Such polyclonal antibody libraries would be expected to mediate effector functions and provide optimal passive immunity against cryptosporidiosis.

Animals↗

Human complement factor H: molecular cloning and cDNA expression reveals variability in the factor H-related mRNA species of 1.4 kb.

Previously, we have shown that three different mRNA species of 4.3 kb, 1.8 kb and 1.4 kb, related to human complement factor H, are constitutively expressed in the human liver. Probing with our cDNA clone H-46 which represents 920 bp of the 3' end of the 4.3 kb mRNA of factor H on human liver RNA, we always detected the 4.3 kb and the additional, abundantly expressed mRNA species of 1.4 kb in length, indicating that the 1.4 kb transcript is highly homologous to the 3' end of the classical factor H mRNA of 4.3 kb. Using H-46 as a probe, several cDNA clones were isolated from a liver cDNA library and sequenced. The open reading frame of the novel mRNA species encodes a peptide consisting of five internal short consensus repeat motifs (SCR), identifying the translational product to be a member of the SCR family. Sequence comparison with cDNA clones derived from liver RNA of a different donor provided evidence for variability in the factor H related proteins encoded by the 1.4 kb mRNA species. Interestingly, this variability was found to be restricted to the three carboxyterminal SCR domains. Expression data indicate that our variant is not recognized by the monoclonal antibody 3D11.

Amino Acid Sequence↗

Identification of 4370 expressed sequence tags from a 3'-end-specific cDNA library of human skeletal muscle by DNA sequencing and filter hybridization.

A systematic study on the mRNA species expressed in the human skeletal muscle is presented in this paper. To carry on this study, a new method has been developed for the construction of unbiased cDNA libraries specially designed for the production of ESTs corresponding to the 3'-end portion of the mRNAs. The method has been applied to human skeletal muscle, where the analysis of the transcription profile is particularly difficult for the presence of several very abundant transcripts. To detect and quantify high-level mRNAs, the first 1054 ESTs were obtained from randomly selected clones. The 10 most abundant transcripts accounted for > 45% of the clones. Subsequently, these transcripts were identified by filter hybridization, thus making DNA sequencing more productive. Overall, 4370 clones were identified: 3372 by DNA sequencing and 998 by filter hybridization. The number of groups of sequences identifying individual transcripts was relatively low compared with other tissues, resulting in a total of 934 groups out of 4370 ESTs. Of these, 719 groups were represented by only one sequence.

Cloning, Molecular↗

Homology of the root adhesin of Pseudomonas fluorescens OE 28.3 with porin F of P. aeruginosa and P. syringae.

The gene encoding the root adhesin from the outer membrane of Pseudomonas fluorescens OE 28.3 was isolated from a genomic lambda EMBL3 library and sequenced. The deduced protein (32104 daltons) displayed strong homology with the amino- and carboxyterminal parts of porin F (OprF) from P. aeruginosa and P. syringae. Significant homology was also found within the C-terminal domain of the OmpA proteins from Enterobacteria and major outer membrane proteins from Neisseria species. However, a cysteine-rich domain present in the OprFs of P. aeruginosa and P. syringae is absent from the adhesin of P. fluorescens. Instead, it contains a shorter sequence with eight alternating proline residues.

Adhesins, Bacterial↗

Peroxisomal isocitrate lyase of the n-alkane-assimilating yeast Candida tropicalis: gene analysis and characterization.

A genomic DNA clone encoding isocitrate lyase, a key enzyme of the glyoxylate cycle and a peroxisomal enzyme of the n-alkane-assimilating yeast Candida tropicalis has been isolated with a cDNA probe from the yeast lambda EMBL library. Nucleotide sequence analysis of the genomic DNA clone disclosed that the region coding isocitrate lyase had a length of 1,650 base pairs, corresponding to 550 amino acids (61,602 Da). RNA blot analysis demonstrated that only one kind of mRNA (2 kb) supposed to be transcribed from this gene was present in the cells. A comparison of the amino acid sequences was made with the isocitrate lyase of castor bean, one of the glyoxysomal enzymes, and the enzyme of E. coli. The isocitrate lyases of C. tropicalis and castor bean had high homology, and the presence of some amino acid stretches conserved in all three enzymes suggests that these might be involved in the catalysis of the common reaction. There was an insertion common to the isocitrate lyases of C. tropicalis and castor bean, which is of interest concerning their evolution. In the C-terminal region, a characteristic sequence similar to that previously proposed as the import signal to peroxisomes was present.

Amino Acid Sequence↗

Computer selection of oligonucleotide probes from amino acid sequences for use in gene library screening.

We present a computer program, FINPROBE, which utilizes known amino acid sequence data to deduce minimum redundancy oligonucleotide probes for use in screening cDNA or genomic libraries or in primer extension. The user enters the amino acid sequence of interest, the desired probe length, the number of probes sought, and the constraints on oligonucleotide synthesis. The computer generates a table of possible probes listed in increasing order of redundancy and provides the location of each probe in the protein and mRNA coding sequence. Activation of a next function provides the amino acid and mRNA sequences of each probe of interest as well as the complementary sequence and the minimum dissociation temperature of the probe. A final routine prints out the amino acid sequence of the protein in parallel with the mRNA sequence listing all possible codons for each amino acid.

Amino Acid Sequence↗

celA, another gene coding for a multidomain cellulase from the extreme thermophile Caldocellum saccharolyticum.

Caldocellum saccharolyticum is an extremely thermophilic anaerobic bacterium capable of growth on cellulose and hemicellulose as sole carbon sources. Cellulase and hemicellulase genes have been found clustered together on its genome. The gene for one of the cellulases (celA) was isolated on a lambda genomic library clone, sequenced and found to comprise a large open-reading frame of 5253 base pairs that could be translated into a peptide of 1751 amino acids. To date, it is the largest cellulase gene sequenced. The translated product is a multidomain structure composed of two catalytic domains and two cellulose-binding domains linked by proline-threonine-rich regions (PT linkers). The N-terminal domain of celA encodes for an endoglucanase activity on carboxymethylcellulose, consistent with its high homology to the sequences of several other endo-1, 4-beta-D-glucanases. The carboxyterminal domain shows sequence homology with a cellulase from Clostridium thermocellum (CelS), which is known to act synergistically with a second component to hydrolyze crystalline cellulose. In the absence of a Caldocellum homologue for this second protein, we can detect no activity from this domain.

Amino Acid Sequence↗