PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Construction of rare restriction site (NotI, SacII and ClaI) linking libraries and sequence-tagged site analysis of single-copy clones from a human chromosome-3-specific library.

We have constructed rare restriction-site (NotI, SacII and ClaI) chromosome 3 (Chr 3)-specific linking libraries in a plasmid-based vector by mass transfer of a lambda phage human Chr-3-specific library (LA03NS01-ATCC57717) into pUC18. Total plasmid DNA isolated from the plasmid-based Chr-3-specific library was digested with either ClaI, NotI or SacII. Linear molecules were separated from undigested circles by pulsed-field polyacrylamide-gel electrophoresis. Purified linear molecules were circularized with T4 DNA ligase and transformed into bacteria. The resulting clones were greatly enriched for sequences recognized by the original restriction endonuclease used for digestion (83 to 95%). These sublibraries are composed of 600 (NotI) 1000 (SacII) or 30,000 (ClaI) clones. Thus, this procedure allows for easy isolation of Chr-3-specific DNA clones containing a variety of rare restriction sites. Sequence-tagged site (STS) data are also presented for five site-specifically mapped Chr-3-specific DNA clones. These studies may facilitate the construction of region specific linking libraries for mapping of various disease-specific loci on Chr 3.

Animals↗

Expressed sequence tags for the chicken genome from a normalized 10-day-old white leghorn whole-embryo cDNA library. 3. DNA sequence analysis of genetic variation in commercial chicken populations.

Single nucleotide polymorphisms (SNPs) have emerged as a major class of DNA markers with the advantage of permitting the development of high-density genetic maps adequate for quantitative trait loci (QTL) identification by linkage-disequilibrium analysis. Here we describe results of a relatively high-depth survey of chicken broiler and layer populations for SNPs in targeted genomic regions of chicken expressed sequence tag (EST) sites. The sequences scanned, representing the composite sequence of 12 amplified fragments for a total of 6489 bp, were randomly distributed, occurring on six different chromosomes or linkage groups in the chicken genome. Although one of the genomic DNA sequences did not match the reference cDNA sequence, another contained an intron that separated two putative exons. The number of SNPs observed within each of the 12 EST-targeted genomic regions ranged from 0 to 10 for a total of 44 and a frequency of 0.7%. About 70% of the polymorphisms were shared between layer and broiler populations. The average heterozygosity within the populations ranged from 0.15 to 0.48, with the layer populations showing the higher heterozygosity. SNPs and oligonucleotides described will provide a resource for genetic analysis in commercial chicken populations. The data appear to indicate that the relative frequency of SNPs in the targeted regions scanned is higher than the frequency reported for any of the other regions scanned to date in other eukaryotic genomes. Additionally, the results suggest that the use of DNA pools may offer an efficient approach to SNP detection in chickens, as has been shown in other vertebrates.

Animals↗

Microarray profiling of gene expression during trypomastigote to amastigote transition in Trypanosoma cruzi.

Trypanosoma cruzi, the causative agent of Chagas disease, remains a significant public health concern throughout South and Central America. Although much is known about immune control of T. cruzi and in particular the importance of recognition of parasite-infected cells, relatively little is known about the target antigens of these protective immune responses. For instance, few of the genes expressed in the intracellular amastigote stage have been identified. To gain insight into the molecular events, at the level of mRNA abundance, involved in this critical point in the parasite life-cycle, we used DNA microarrays of 4400 sequences from T. cruzi ORF-selected and random, genomic sequencing libraries to determine relative mRNA abundances in trypomastigotes and developing amastigotes. Results from six hybridizations using independently generated parasite samples consistently identified 60 probes that detected genes upregulated within 2h after extracellular trypomastigotes were induced, in vitro, to differentiate into amastigotes. Sequence analysis from these 60 probes identified 14 known and 25 novel T. cruzi genes. The general direction of regulation was confirmed by quantitative RT-PCR for seven of the array-identified, amastigote upregulated, known genes. This work demonstrates the feasibility of computational and microarray approaches to gene discovery in T. cruzi, an organism for which a fully assembled and annotated genome sequence is not yet available and in which control of transcription initiation is believed to be absent. Moreover, this work is the first report of amastigote up regulation for 38 genes, thus expanding considerably the pool of genes known to be upregulated in this important yet poorly-studied stage of the T. cruzi life-cycle.

Animals↗

RNA recognition site of PP7 coat protein.

The coat proteins of different single-strand RNA phages use a common protein tertiary structural framework to recognize different RNA hairpins and thus offer a natural model for understanding the molecular basis of RNA-binding specificity. Here we describe the RNA structural requirements for binding to the coat protein of bacteriophage PP7, an RNA phage of Pseudomonas. Its recognition specificity differs substantially from those of the coat proteins of its previously characterized relatives such as the coliphages MS2 and Qbeta. Using designed variants of the wild-type RNA, and selection of binding-competent sequences from random RNA sequence libraries (i.e. SELEX) we find that tight binding to PP7 coat protein is favored by the existence of an 8 bp hairpin with a bulged purine on its 5' side separated by 4 bp from a 6 nt loop having the sequence Pu-U-A-G/U-G-Pu. However, another structural class possessing only some of these features is capable of binding almost as tightly.

Base Sequence↗

Microarray gene expression profiles in dilated and hypertrophic cardiomyopathic end-stage heart failure.

Despite similar clinical endpoints, heart failure resulting from dilated cardiomyopathy (DCM) or hypertrophic cardiomyopathy (HCM) appears to develop through different remodeling and molecular pathways. Current understanding of heart failure has been facilitated by microarray technology. We constructed an in-house spotted cDNA microarray using 10,272 unique clones from various cardiovascular cDNA libraries sequenced and annotated in our laboratory. RNA samples were obtained from left ventricular tissues of precardiac transplantation DCM and HCM patients and were hybridized against normal adult heart reference RNA. After filtering, differentially expressed genes were determined using novel analyzing software. We demonstrated that normalization for cDNA microarray data is slide-dependent and nonlinear. The feasibility of this model was validated by quantitative real-time reverse transcription-PCR, and the accuracy rate depended on the fold change and statistical significance level. Our results showed that 192 genes were highly expressed in both DCM and HCM (e.g., atrial natriuretic peptide, CD59, decorin, elongation factor 2, and heat shock protein 90), and 51 genes were downregulated in both conditions (e.g., elastin, sarcoplasmic/endoplasmic reticulum Ca2+-ATPase). We also identified several genes differentially expressed between DCM and HCM (e.g., alphaB-crystallin, antagonizer of myc transcriptional activity, beta-dystrobrevin, calsequestrin, lipocortin, and lumican). Microarray technology provides us with a genomic approach to explore the genetic markers and molecular mechanisms leading to heart failure.

Adult↗

Brain-specific expression of transthyretin mRNA as revealed by cDNA cloning from brain.

cDNAs for rat transthyretin mRNA were cloned from a brain cDNA library. Sequencing analyses showed the presence of an additional 5' sequence that had not been reported for the liver mRNA corresponding to the flanking promoter region of the gene. This additional sequence was expressed only in the brain, suggesting the presence of a brain-specific promoter.

Animals↗

Structural and biochemical characterization of DSL ribozyme.

We recently reported on the molecular design and synthesis of a new RNA ligase ribozyme (DSL), whose active site was selected from a sequence library consisting of 30 random nucleotides set on a defined 3D structure of a designed RNA scaffold. In this study, we report on the structural and biochemical analyses of DSL. Structural analysis indicates that the active site, which consists of the selected sequence, attaches to the folded scaffold as designed. To see whether DSL resembles known ribozymes, a biochemical assay was performed. Metal-dependent kinetic studies suggest that the ligase requires Mg2+ ions. The replacement of Mg2+ with Co(NH3)6(3+) prohibits the reaction, indicating that DSL requires innersphere coordination of Mg2+ for a ligation reaction. The results show that DSL has requirements similar to those of previously reported catalytic RNAs.

Base Sequence↗

Flexible sequence similarity searching with the FASTA3 program package.

The FASTA3 and FASTA2 packages provide a flexible set of sequence-comparison programs that are particularly valuable because of their accurate statistical estimates and high-quality alignments. Traditionally, sequence similarity searches have sought to ask one question: "Is my query sequence homologous to anything in the database?" Both FASTA and BLAST can provide reliable answers to this question with their statistical estimates; if the expectation value E is < 0.001-0.01 and you are not doing hundreds of searches a day, the answer is probably yes. In general, the most effective search strategies follow these rules: 1. Whenever possible, compare at the amino acid level, rather than the nucleotide level. Search first with protein sequences (blastp, fasta3, and ssearch3), then with translated DNA sequences (fastx, blastx), and only at the DNA level as a last resort (Table 5). 2. Search the smallest database that is likely to contain the sequence of interest (but it must contain many unrelated sequences for accurate statistical estimates). 3. Use sequence statistics, rather than percent identity or percent similarity, as your primary criterion for sequence homology. 4. Check that the statistics are likely to be accurate by looking for the highest-scoring unrelated sequence, using prss3 to confirm the expectation, and searching with shuffled copies of the query sequence [randseq, searches with shuffled sequences should have E approx 1.0]. 5. Consider searches with different gap penalties and other scoring matrices. Searches with long query sequences against full-length sequence libraries will not change dramatically when BLOSUM62 is used instead of BLOSUM50 (20), or a gap penalty of -14/-2 is used in place of -12/-2. However, shallower or more stringent scoring matrices are more effective at uncovering relationships in partial sequences (3,18), and they can be used to sharpen dramatically the scope of the similarity search. However, as illustrated in the last section, the E value is only the first step in characterizing a sequence relationship. Once one has confidence that the sequences are homologous, one should look at the sequence alignments and percent identities, particularly when searching with lower quality sequences. When sequence alignments are very short, the alignment should become more significant when a shallower scoring matrix is used, e.g., BLOSUM62 rather than BLOSUM50 (remember to change the gap penalties). Homology can be reliably inferred from statistically significant similarity. Whereas homology implies common three-dimensional structure, homology need not imply common function. Orthologous sequences usually have similar functions, but paralogous sequences often acquire very different functional roles. Motif databases, such as PROSITE (21), can provide evidence for the conservation of critical functional residues. However, motif identity in the absence of overall sequence similarity is not a reliable indicator of homology.

Amino Acid Sequence↗

Mouse tetranectin: cDNA sequence, tissue-specific expression, and chromosomal mapping.

Tetranectin is a plasminogen-binding tetrameric protein originally isolated from plasma. Expression of tetranectin appears ubiquitous, although particularly high expression is noted in the stroma of malignant tumors and during mineralization. To dissect the molecular basis of tetranectin gene regulation, mouse tetranectin cDNA was cloned from a 16-day-old mouse embryo library. Sequence analysis revealed a 992-bp cDNA with an open reading frame of 606 bp, which is identical in length to the human tetranectin cDNA. The deduced amino acid sequence showed high homology to the human cDNA with 76% identity and 87% similarity at the amino acid level. Sequence comparisons between mouse and human tetranectin and some C-type lectins confirmed a complete conservation in the position of six cysteines as well as numerous other amino acid residues, indicating an essential structure for potential function(s) of tetranectin. The sequence analysis revealed a difference in both sequence and size of the noncoding regions between mouse and human cDNAs. Northern analysis of the various tissues from mouse, rat, and cow showed the major transcript(s) to be approximately 1 kb, which is similar in size to that observed in human. Although additional minor bands of 1.5 and 3.3 kb were found in Northern blots, RT-PCR (reverse transcription polymerase chain reaction) analysis failed to provide evidence that these minor bands are products of the tetranectin gene. Finally, the genetic map location for this gene, Tna, was determined to be on distal mouse Chromosome (Chr) 9 by analysis of two sets of multilocus crosses.

Amino Acid Sequence↗

[Expressed genes related to adolescent idiopathic scoliosis in spinal facet].

OBJECTIVE: To construct a subtractive cDNA library from the spinal facets of adolescent idiopathic scoliosis (AIS) patient with suppression subtractive hybridization (SSH) technique and clone differentially expressed genes related to AIS in spinal facet. METHODS: mRNA was isolated separately from the convex side and concave side apical facets of an adolescent idiopathic scoliosis patient. Moreover, single-strand (ss) and double-strand (ds) cDNAs were synthesized in turn using SMART PCR cDNA synthesis technology. (ds) cDNAs then were digested with Rsa I and divided into two groups, and ligated to the specific adaptor 1 and adaptor 2R, respectively. After cDNAs were hybridized with each other twice and underwent two rounds of nested PCR, the PCR products were ligated with pDrive Cloning Vectors to set up the subtractive library. Sequence analysis was performed and the acquired data were aligned against the GeneBank nucleotide database. Furthermore, the spinal facets of 11 IS patients were collected. The techniques of immunohistochemistry and in situ hybridization were adopted in order to approve the gene expression difference. The images of immunohistochemistry and in situ hybridization were input to the image analysis system and were analyzed semi-quantitatively. RESULTS: A cDNA subtractive library of AIS in spinal facet was set up successfully with high subtractive efficiency. Beta2 Microglobulin (beta 2M) and calgranulin A (S100A8) were identified as highly differentially expressed genes in this study. CONCLUSION: All results confirm the effectiveness and sensitivity of SSH in detecting differentially expressed genes from a small amount of clinical samples. Information about such alterations in gene expression can be useful for elucidating the genetic events in the development of AIS.

Adolescent↗

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20&#xd7;2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30&#xd7; coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N&#xa0;=&#xa0;75), stool (N&#xa0;=&#xa0;242), and saliva (N&#xa0;=&#xa0;190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics↗

Identification of ABCC6 pseudogenes on human chromosome 16p: implications for mutation detection in pseudoxanthoma elasticum.

Pseudoxanthoma elasticum (PXE), a heritable disorder affecting the skin, eyes, and the cardiovascular system, has recently been linked to mutations in the ABCC6 gene on chromosome 16p13.1. The original mutation detection strategy employed by us consisted of the amplification of each exon of the ABCC6 gene with primer pairs placed on the flanking introns, followed by heteroduplex scanning and direct nucleotide sequencing. However, this approach suggested the presence of multiple copies of the 5'-region of the gene when total genomic DNA was used as a template. In this study, we have identified two pseudogenes containing sequences highly homologous to the 5'-end of ABCC6. First, by the use of allele-specific polymerase chain reaction (PCR), two bacterial artificial chromosome (BAC) clones containing a putative pseudogene of ABCC6, designated as ABCC6-psi 1, were isolated from the human BAC library. Sequence analysis of ABCC6-psi 1 revealed it to be a truncated copy of ABCC6, which contains the upstream region and exon 1 through intron 9 of the gene. Secondly, a homology search of a high-throughput sequence database revealed the presence of another truncated copy of ABCC6, which was designated as ABCC6-psi 2, and which was shown to harbor upstream sequences and a segment spanning exon 1 through intron 4 of ABCC6. In addition to several nucleotide differences in the flanking introns and the upstream region, both pseudogenes contain several nucleotide changes in the exonic sequences, including stop codon mutations, which complicate mutation analysis in patients with PXE. Nucleotide differences in flanking introns between these two pseudogenes and ABCC6 allowed us to design allele-specific primers that eliminated the amplification of both pseudogene sequences by PCR and provided reliable amplification of ABCC6-specific sequences only. The use of allele-specific PCR has revealed, thus far, two novel 5'-end PXE mutations, 179del9 and T364R in exons 2 and 9, respectively, and several polymorphisms within the upstream region and exons 1-9 of ABCC6. These strategies facilitate comprehensive analysis of ABCC6 for mutations in PXE.

Alleles↗

Structural and functional characterization of the mouse hepatocyte growth factor gene promoter.

To understand the molecular mechanisms underlying the regulation of hepatocyte growth factor (HGF) gene expression and to define the DNA sequences essential for its cell-type specific and inducible expression, we have isolated and characterized the 5'-flanking region of the HGF gene. A genomic clone containing 2.8 kilobases of the 5'-flanking region of the HGF gene has been isolated from a mouse liver genomic library. Sequence analysis showed that the promoter region of the mouse HGF gene contains a noncanonical TATA box (ATAAA). Further analysis of the 5'-flanking region revealed a number of putative regulatory elements, such as four interleukin-6 response elements (IL-6 RE), two potential binding sites for NF-IL6, a TGF-beta inhibitory element (TIE), a cAMP response element (CRE), two estrogen response elements (ERE) including one located in the first intron, a potential vitamin D response element (VDRE) which overlaps a chicken ovalbumin upstream promoter (COUP) transcription factor binding element, two liver-specific transcription factor (C/EBP) binding sites, and a B cell- and macrophage-specific transcriptional factor binding site (PU.1/ETS). To determine the location of sites that may be critical for the function of the HGF promoter, we constructed a series of chimeric genes containing variable regions of the 5'-flanking sequence of HGF gene and the coding region for chloramphenicol acetyltransferase (CAT). Transient transfection of chimeric plasmids demonstrated that the mouse HGF gene promoter containing 70 base pairs of the 5'-flanking sequences were active in mouse fibroblast NIH 3T3 cells and in human endometrial carcinoma RL95-2 cells. This basal transcription activity of the HGF promoter was modulated in NIH 3T3 and RL95-2 cells by multiple upstream elements. Three positive elements were identified at positions -2848 to -2674, -1386 to -1231, and -699 to -274, and three negative candidate elements were mapped to positions -1652 to -1386, -964 to -699, and -274 to -70, respectively. By the combination of a series of 5'-end deletion and internal deletion, a cell type-specific negative regulatory element in RL95-2 cells was localized to the nucleotide position -964 to -699. Moreover, the reporter plasmid containing interleukin 6 (IL-6) response element was responsive to IL-6 stimulation in stably transfected NIH 3T3 cells. Our findings revealed a complex pattern of transcriptional regulation of the mouse HGF gene expression.

Animals↗

Cloning and sequence analysis of cDNAs encoding human placental tissue protein 17 (PP17) variants.

Using monospecific anti-PP17 serum with chemiluminescence Western-blot analysis, we detected different molecular-mass variants of human soluble placental tissue protein 17 (PP17) in different normal adult and fetal human tissues besides term placenta. 13 cDNAs with three different insert lengths encoding PP17 variants were isolated by screening a human placental cDNA library. Sequence analysis of the shortest clones showed that the inserts contain the same open reading frame encoding PP17a variant (28,129 kDa) consisting of 251 residues, which is identical to the previously isolated and characterised PP17 antigen described in 1983. The ubiquitous PP17b variant is encoded by longer clones and contains 434 residues with a predicted molecular mass of 47,208 kDa. Compared to normal conditions, these newly discovered PP17 variants are overexpressed in cervix carcinoma tissue, as are their three different-size messenger RNAs in HeLa cell line. Increased amounts of PP17b are secreted into the circulation in cervix carcinoma patients. We also observed a typical elevation in serum levels of PP17 variants during healthy pregnancy. An alignment search of the protein databank showed that PP17a and PP17b are homologous to adipose tissue differentiation and lipid-droplet-associated proteins: human adipophilin, mouse adipose differentiation-related protein and rat perilipin A and B.

Amino Acid Sequence↗

Physical characterization and sequence identification of the ovary maturating parsin. A new neurohormone purified from the nervous corpora cardiaca of the African locust (Locusta migratoria migratorioides.

A novel neurohormone, which anticipates ovarian maturation, was recently purified using liquid chromatography from the African locust nervous corpora cardiaca. Both its function and production by the pars intercerebralis of Locusta migratoria lead to its name, the ovary maturating parsin (Lom OMP). In this study, the Lom OMP was physically and chemically characterized. Its multiply charged ion spectrum was interpreted as two peaks of quite equal size having molecular masses of 6923.4 Da (major peak) and 6907.3 Da. The Lom OMP presented no periodic secondary structure according to the far ultraviolet circular dichroism spectrum obtained. It is composed of 65 amino acids and included a high concentration of alanine but is devoid of cysteine, isoleucine, methionine, lysine and threonine. The amino acid sequence indicated only one microheterogeneity, observed at position 26, consisted in the replacement of serine by alanine. The calculated Mr of the two acidic isoforms (calculated pHi = 4.87) were found to be in agreement with mass spectrometry measurements. When compared to the sequence libraries, the Lom OMP, the first insect gonadotropic neurohormone, was revealed as an unique protein.

Amino Acid Sequence↗

Clostridium perfringens type A enterotoxin: characterization of the amino-terminal region.

The amino-terminal region of the enterotoxin of Clostridium perfringens was investigated by automated sequence analysis. The primary structure results revealed that the enterotoxin is composed of a single polypeptide amino acid sequence. Computer comparison of a 20-residue sequence with a sequence library of reported proteins revealed no significant chemical similarities, indicating that the enterotoxin represents a unique polypeptide primary structure.

Amino Acid Sequence↗

[Mapping and human genome sequence program].

Until recently, human genome programs focused primarily on establishing maps that would provide signposts to researchers seeking to identify genes responsible for inherited diseases, as well as a basis for genome sequencing studies. Preestablished gene mapping goals have been reached. The over 7,000 microsatellite markers identified to date provide a map of sufficient density to allow localization of the gene of a monogenic disease with a precision of 1 to 2 million base pairs. The physical map, based on systematically arranged overlapping sets of artificial yeast chromosomes (YACs), has also made considerable headway during the last few years. The most recently published map covers more than 90% of the genome. However, currently available physical maps cannot be used for sequencing studies because multiple rearrangements occur in YACs. The recently developed sets of radioinduced hybrids are extremely useful for incorporating genes into existing maps. A network of American and European laboratories has successfully used these radioinduced hybrids to map 15,000 gene tags from large-scale cDNA library sequencing programs. There are increasingly pressing reasons for initiating large scale human genome sequencing studies.

Chromosome Mapping↗