PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

The 30-kilodalton subunit of bovine mitochondrial complex I is homologous to a protein coded in chloroplast DNA.

In cattle, 7 of the 30 or more subunits of the respiratory enzyme NADH:ubiquinone reductase (complex I) are encoded in mitochondrial DNA, and potential genes (open reading frames, orfs) for related proteins are found in the chloroplast genomes of Marchantia polymorpha and Nicotiana tabacum. Homologues of the nuclear-coded 49- and 23-kDa subunits are also coded in chloroplast DNA, and these orfs are clustered with four of the homologues of the mammalian mitochondrial genes. These findings have been taken to indicate that chloroplasts contain a relative of complex I. The present work provides further support. The 30-kDa subunit of the bovine enzyme is a component of the iron-sulfur protein fraction. Partial protein sequences have been determined, and synthetic oligonucleotide mixtures based on them have been employed as hybridization probes to identify cognate cDNA clones from a bovine library. Their sequences encode the mitochondrial import precursor of the 30-kDa subunit. The mature protein of 228 amino acids contains a segment of 57 amino acids which is closely related to parts of proteins encoded in orfs 169 and 158 in the chloroplast genomes of M. polymorpha and N. tabacum. Moreover, the chloroplast orfs are found near homologues of the mammalian mitochondrial genes for subunit ND3. Therefore, the plant chloroplast genomes have at least two separate clusters of potential genes encoding homologues of subunits of mitochondrial complex I. The bovine 30-kDa subunit has no extensive sequences of hydrophobic amino acids that could be folded into membrane-spanning alpha-helices, and although it contains two cysteine residues, there is no clear evidence in the sequence that it is an iron-sulfur protein.

Amino Acid Sequence↗

Cloning and sequence analysis of cDNA encoding the bovine testis-derived male-enhanced antigen (Mea).

A full-length cDNA encoding the bovine male enhanced antigen (Mea) has been cloned from a bovine testicular cDNA library and sequenced. The primary structure of the bovine Mea peptide deduced from this nucleotide sequence has 174 amino acid residues and is highly homologous to human (95.9%, 165/172) and mouse (92.5%, 161/174) Mea gene products. It is located on an autosome, and is expressed highly in the testes.

Amino Acid Sequence↗

A set of 99 cattle microsatellites: characterization, synteny mapping, and polymorphism.

Cattle microsatellite clones (136) were isolated from cosmid (10) and plasmid (126) libraries and sequenced. The dinucleotide repeats were studied in each of these sequences and compared with dinucleotide repeats found in other vertebrate species where information was available. The distribution in cattle was similar to that described for other mammals, such as rat, mouse, pig, or human. A major difference resides in the number of sequences present in the bovine genome, which seemed at best one-third as large as in other species. Oligonucleotide primers (117 pairs) were synthesized, and a PCR product of expected size was obtained for 88 microsatellite sequences (75%). Synteny or chromosome assignment was searched for each locus with PCR amplification on a panel of 36 hamster/bovine somatic cell hybrids. Of our bovine microsatellites, eighty-six could be assigned to synteny groups of chromosomes. In addition, 10 other microsatellites--HEL 5, 6, 9, 11, 12, 13 (Kaukinen and Varvio 1993), HEL 4, 7, 14, 15--as well as the microsatellite found in the kappa-casein gene (Fries et al. 1990) were mapped on the hybrids. Microsatellite polymorphism was checked on at least 30 unrelated animals of different breeds. Almost all the autosomal and X Chr microsatellites displayed polymorphism, with the number of alleles varying between two and 44. We assume that these microsatellites could be very helpful in the construction of a primary public linkage map of the bovine genome, with an aim of finding markers for Economic Trait Loci (ETL) in cattle.

Animals↗

Cloning of a novel ubiquitin-conjugating enzyme (E2) gene from the ciliate Paramecium tetraurelia.

We isolated a 1.7 kb gene (UbcP1) for a ubiquitin-conjugating enzyme from a P. tetraurelia cDNA library and sequenced it. Its deduced polypeptide sequence consists of 425 amino acid residues (48 kDa). The UbcP1 protein contains novel N- and C-terminal extensions in addition to a UBC domain, and within the UBC domain it shares low identity with sequences of other known E2s. A constructed phylogenetic tree suggests that the UbcP1 protein may represent a member of a distinct subfamily of E2s. Southern blot analysis showed that the N-terminal extension of the UbcP1 is conserved in P. multimicronucleatum.

Amino Acid Sequence↗

Microdissection of the fragile X region.

We have microdissected and cloned the region around the fragile site at Xq27.3 on the human X chromosome. All of the clones tested map to the Xq27-Xq28 region, and detailed mapping on a panel of somatic cell hybrids indicates that the microdissected library contains sequences derived from both sides of the fragile X mutation. Some of these clones give signals in rodent DNA. This library demonstrates the power of microdissection for the identification of potential coding sequences near a disease locus and provides a promising resource for the identification of the fragile X mutation.

Animals↗

Genomic organization of the mouse OSF-1 gene.

The mouse OSF-1 protein (also known as pleiotrophin, HB-GAM, HBGF-8, or HBNF) gene was isolated from a mouse genomic library and sequenced. OSF-1 is a 15-kD secreted protein specifically expressed in bone and brain, and is believed to play a role in brain development and osteogenesis. The mouse OSF-1 gene consists of at least 5 exons and 4 introns and spans > 32 kb. Computer analysis of approximately 4 kb of 5'-flanking sequence of the OSF-1 gene revealed two candidate promoter regions. One candidate promoter contains a thyroid hormone/retinoic acid-responsive element and the other contains two glucocorticoid-responsive elements. DNA sequence analysis of novel OSF-1 cDNA clones indicates that two promoters can be utilized in MC3T3-E1 osteoblastic cells. The overall organization of the mouse OSF-1 gene is similar and the locations of the three exon-intron junctions within the coding region are identical to the mouse gene encoding the differentiation-related factor midkine (MK). Based on this similarity and on the high degree of nucleotide sequence homology (approximately 55%) of mouse OSF-1 and mouse MK, we conclude that OSF-1 and MK are generated from a common ancestral gene and are members of a family of structurally and probably functionally related proteins.

Amino Acid Sequence↗

[Screening and identification of novel genes involved in biosynthesis of ginsenoside in Panax ginseng plant].

The root of Panax ginseng plant undergoes a specific developmental process to become a biosynthesis and accumulation organ for ginsenosides. To identify and analyze genes involved in the biosynthesis of ginsenoside, suppression subtractive hybridization (SSH) between mRNAs of 4- and 1-year-old root tissues was performed, and a subtracted cDNA library specific to 4-year-old roots was constructed. Forty cDNA clones selected randomly from the subtracted cDNA library were sequenced. Sequence information of all clones was evaluated by Nucleotide Blast analysis in GenBank/DDBJ/EMBL. The results showed that six subtracted cDNA clones represented the novel genes (ESTs), because no sequence homology with any known sequences was found in the database. Expression in 4-year-old P. ginseng root tissues was verified by reverse Northern dot hybridization for the six clones. These six novel genes were named GBR1, GBR2, GBR3, GBR4, GBR5, and GBR6, and their Accession numbers of GenBank are AF485334, AF485335, AF485336, AF485337, AF485332, and AF485333, respectively. Finally, Northern blot analysis and semi-quantitative reverse transcription polymerase chain reaction (RT-PCR) confirmed that these six novel genes were differentially expressed in the defined development stage of P. ginseng plant roots. It is possible that their overexpression may play an important role in the ginsenoside biosynthesis. In addition, most of transcripts of all genes could also be detected in other P. ginseng plant tissues such as stem, leaf and seed. Our results provided a basis for obtaining the full-length cDNA sequences of such six novel genes, and for identifying their function involved in the biosynthesis of ginsenoside.

Base Sequence↗

Using Chromosome Conformation Capture Combined with Deep Sequencing (Hi-C) to Study Genome Organization in Bacteria.

Genome organization is fundamental to all living organisms. Long DNA molecules are organized in hierarchical orders to be accommodated into eukaryotic nuclei or bacterial cells, which are thousands of folds shorter. Over the past two decades, chromosome conformation capture (3C) techniques substantially advanced our understanding of genome folding inside cells. 3C involves crosslinking and proximity ligation, and quantifies the physical contacts between two DNA regions within the genome. Coupled with high-throughput sequencing, 3C-seq and Hi-C techniques detect genome-wide DNA interactions, providing a comprehensive view of global genome organization. Here, we describe a detailed method to prepare Hi-C libraries using Bacillus subtilis, which includes procedures of crosslinking chromatin, digesting the crosslinked genome, labeling DNA ends with biotin, ligating DNA, and preparing the DNA library for sequencing using an Illumina platform.

High-Throughput Nucleotide Sequencing↗

Development of 90 new bovine microsatellite loci.

Polymorphic genetic markers are important tools in the construction of comprehensive genetic maps. This paper reports on the isolation and characterization of 90 new bovine microsatellite (ms) loci from enriched genomic libraries. The sequence of one clone (locus MNB-85) showed significant similarity to an intron of the human promyelocytic leukemia zinc finger protein (PLZF) gene. Screening of bovine and porcine somatic cell panels places the bovine PLZF homolog on BTA-15 and the porcine PLZF homolog on SSC-9. The 90 new microsatellite loci increase the number of microsatellites available for cattle by >5%.

Animals↗

Identification of peptides that neutralize bacterial endotoxins using beta-hairpin conformationally restricted libraries.

Bacterial endotoxins are the major mediator of septic shock; therefore, endotoxin-neutralizing molecules could have biomedical applications. The septic shock cascade relies in a series of molecular recognition processes. The large contact-surface described for the interacting macromolecules, in most cases, prevents the identification of small molecules that could modulate such recognition events. Here we report on a beta-hairpin conformationally restricted combinatorial library that has been generated and screened towards the identification of new peptides that neutralize bacterial endotoxins. Starting with a de novo designed linear peptide that shows a beta-hairpin structure population of around 30%, (Ramirez-Alvarado, M., Blanco, F. J. and Serrano, L. Nat. Struc. Biol., 7, 604-612 (1996)), we selected four positions to build up a combinatorial library of 20(4) sequences. Deconvolution of the library reduced such a sequence complexity to 8 defined sequences. The newly identified peptides have a biological activity equivalent to that reported for peptides derived from natural endotoxin-binding proteins.

Amino Acid Sequence↗

Green fluorescent protein as a scaffold for intracellular presentation of peptides.

Peptide aptamers provide probes for biological processes and adjuncts for development of novel pharmaceutical molecules. Such aptamers are analogous to compounds derived from combinatorial chemical libraries which have specific binding or inhibitory activities. Much as it is generally difficult to determine the composition of combinatorial chemical libraries in a quantitative manner, determining the quality and characteristics of peptide libraries displayed in vivo is problematical. To help address these issues we have adapted green fluorescent protein (GFP) as a scaffold for display of conformationally constrained peptides. The GFP-peptide libraries permit analysis of library diversity and expression levels in cells and allow enrichment of the libraries for sequences with predetermined characteristics, such as high expression of correctly folded protein, by selection for high fluorescence.

Amino Acid Sequence↗

A single gene encodes membrane-bound and free forms of GP-2, the major glycoprotein in pancreatic secretory (zymogen) granule membranes.

GP-2, a 75-kDa glycoprotein, was isolated from dog pancreatic zymogen granule membranes (ZGMs). In a carbohydrate-shift strategy, N-terminal and internal peptide sequences were obtained on glycosylated and deglycosylated forms of GP-2, respectively, by gas-phase sequencing. Sets of mixed oligonucleotides and the polymerase chain reaction were used to obtain a double-stranded cDNA probe, which was used to isolate overlapping cDNA clones from a dog pancreatic cDNA library. The sequence of these clones revealed an open reading frame that encodes a protein of 509 amino acids, eight N-linked oligosaccharide attachment sites, and an N-terminal signal sequence absent from the mature form of GP-2 associated with ZGMs. The C terminus shows a 20-residue hydrophobic transmembrane domain preceded by a decapeptide containing potential phosphatidylinositol-glycan attachment sites. GP-2 completely released from ZGMs by exogenous phospholipase C showed similar immunochemical properties and electrophoretic mobilities compared to the form associated with ZGMs. A similar form of GP-2 was released from zymogen granules permeabilized with saponin and incubated in the absence of added phospholipase C. Kinetic analysis of GP-2 release at 0 degrees C and 37 degrees C suggested the presence of a granule enzyme responsible for endogenous release of GP-2 to granule contents and into the apical medium. The data indicate that GP-2 is a phosphatidylinositol-glycan-linked membrane protein released from the membrane of mature zymogen granules by an enzymatic mechanism. The cDNA structure presented here thus encodes both membrane-bound and free forms of GP-2.

Amino Acid Sequence↗

The coding sequence of the hemolytically inactive C4A6 allotype of human complement component C4 reveals that a single arginine to tryptophan substitution at beta-chain residue 458 is the likely cause of the defect.

The C4A6 allotype of the human complement component C4 is known to be defective in C5 binding within the C5 convertase. To characterize the position and nature of the molecular defect in the C4A6 allotype we have isolated the C4A6 gene from a cosmid genomic DNA library. Direct sequencing of a 4.4-kb region of the gene covering exons 17 to 31 and encoding the C4d fragment and most of the rest of the alpha chain of C4 revealed that the C4A6 allele encodes the A isotypic residues Pro Cys-Leu Asp at positions 1101, 1102, 1105, and 1106 and the same residues as the C4A3 alpha gene at the polymorphic positions 1054 (Asp), 1157 (Asn), 1182 (Thr), 1188 (Val), 1191 (Leu) and 1267 (Ala). In addition the C4A6 allele was shown to encode a Pro at the previously characterized polymorphic position 707 in the C4a peptide where the C4A3 alpha allele encodes a Leu. The remaining 26 exons of the C4A6 gene were analyzed by detecting nucleotide mismatches in C4A6/C4A3 and C4A6/C4B1 DNA heteroduplexes using the chemical cleavage of mismatch technique. The regions around detected mismatches were sequenced. In total seven nucleotide differences were defined on comparison of the C4A6 and other C4 sequences, of which three were present in exons. Two of these resulted in amino acid changes. One of the amino acid differences is a known polymorphism in C4, a Tyr/Ser substitution at position 328 in the beta-chain. The second amino acid difference caused by a C to T transition in the first base of the codon for amino acid residue 458 was the only one shown to be specific to the C4A6 allotype. The C4A6 allotype contains a Trp residue at this position in the beta-chain instead of the Arg residue found in all other C4A and C4B allotypes so far characterized. We propose that this Arg to Trp substitution at beta-chain residue 458 is responsible for the inability of C4A6 to bind C5 in the C5 convertase.

Alleles↗

Irreversible repression of DNA synthesis in Fanconi anemia cells is alleviated by the product of a novel cyclin-related gene.

Primary fibroblasts from patients with the genetic disease Fanconi anemia, which are hypersensitive to cross-linking agents, were used to screen a cDNA library for sequences involved in their abnormal cellular response to a cross-linking challenge. By using library partition and microinjection of in vitro-transcribed RNA, a cDNA clone, pSPHAR (S-phase response), which is able to correct the permanent repression of semiconservative DNA synthesis rates characteristic of these cells, was isolated. Wild-type SPHAR mRNA is expressed in all fibroblasts so far analyzed, including those of Fanconi anemia patients. Correction of the abnormal response in these cells appears therefore to be due to overexpression after cDNA transfer rather than to genetic complementation. The cDNA contains an open reading frame coding for a polypeptide of 7.5 kDa. Rabbit antiserum directed against a SPHAR peptide detects a protein of 7.9 kDa in Western blots (immunoblots) of whole-cell extracts from proliferating, but not resting, fibroblasts. The deduced amino acid sequence of SPHAR contains a motif found in the cyclins, and it is proposed that SPHAR acts within the injected cell by interfering with the cyclin-controlled maintenance of S phase. In agreement with this proposal, normal cells transfected with an antisense SPHAR expression vector have a significantly reduced rate of DNA synthesis during S phase and a prolonged G2 phase, reflecting the need for postreplicative DNA processing before entry into mitosis.

Amino Acid Sequence↗

Human immunodeficiency virus type 1-neutralizing monoclonal antibody 2F5 is multispecific for sequences flanking the DKW core epitope.

Human monoclonal antibody 2F5 is one of a few human antibodies that neutralize a broad range of HIV-1 primary isolates. The 2F5 epitope on gp41 includes the sequence ELDKWA, with the core residues, DKW, being critical for antibody binding. HIV-neutralizing antibodies have never been elicited by immunization with peptides bearing ELDKWA, suggesting that important part(s) of the 2F5 paratope remain unidentified. The use of longer peptides extending beyond ELDKWA has resulted in increased epitope antigenicity, but neutralizing antibodies have not been generated. We sought to develop peptides that bind to 2F5, and that function as specific probes of the 2F5 paratope. Thus, we used 2F5 to screen a set of phage-displayed, random peptide libraries. Tight-binding clones from the random peptide libraries displayed sequence variability in the regions flanking the DKW motif. To further reveal flanking regions involved in 2F5 binding, two semi-defined libraries were constructed having 12 variegated residues either N-terminal or C-terminal to the DKW core (X(12)-AADKW and AADKW-X(12), respectively). Three clones isolated from the AADKW-X(12) library had similar high affinities, despite a lack of sequence homology among them, or with gp41. The contribution of each residue of these clones to 2F5 binding was evaluated by Ala substitution and amino acid deletion studies, and revealed that each clone bound 2F5 by a different mechanism. These results suggest that the 2F5 paratope is formed by at least two functionally distinct regions: one that displays specificity for the DKW core epitope, and another that is multispecific for sequences C-terminal to the core epitope. The implications of this second, multispecific region of the 2F5 paratope for its unique biological function are discussed.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of hamster CYP2E1.

Two cDNA clones, pHSj31 and pHSj3, coding for CYP2E1 were isolated from a hamster liver cDNA library. The sequence analyses revealed that they encoded the same polypeptide of 493 amino acid residues (M(r) = 56,616) and differed from the length of their 3'-untranslated region. The deduced amino acid sequence of hamster CYP2E1 showed approx. 90% identities with those of the rats and mice, and approx. 80% identities with those of the rabbits, monkeys and humans. The NH2-terminal 35 deduced amino acid sequences of hamster CYP2E1 were completely identical with purified protein, ha P-450j. The Northern blot analysis showed that CYP2E1 was expressed in livers and to lesser extents in kidneys and lungs.

Amino Acid Sequence↗

The human lysozyme gene. Sequence organization and chromosomal localization.

We have isolated two overlapping recombinant lambda-phage clones from a genomic lambda-EMBL3 library containing 25 kb of the human lysozyme gene region. Furthermore a full-lenght human lysozyme cDNA clone of 1.5 kb was isolated from a human placenta cDNA library. Nucleotide sequences of the entire structural gene and the cDNA clone were determined. The human lysozyme gene spans 5856 bp and its sequence organization with four exons and three introns is homologous to the chicken lysozyme gene and the human alpha-lactalbumin gene. Human and chicken lysozyme genes differ mainly in the size of their introns and 3' non-coding region. Four Alu repetitive elements were found in the human lysozyme gene, one in each intron and one on the fourth exon. Lysozyme transcripts of 1.6 kb and 0.6 kb in size were detected in human myeloid cell lines U-937, HL-60 and THP-1 and surprisingly in human hepatoma cell lines HepG2 and Hep3B. The lysozyme gene locus was assigned to human chromosome 12 by hybridization to a panel of DNAs from human-rodent somatic cell hybrids.

Animals↗

HTSinfer: inferring metadata from bulk Illumina RNA-Seq libraries.

SUMMARY: The Sequencing Read Archive is one of the largest and fastest-growing repositories of sequencing data, containing tens of petabytes of sequenced reads. Its data is used by a wide scientific community, often beyond the primary study that generated them. Such analyses rely on accurate metadata concerning the type of experiment and library, as well as the organism from which the sequenced reads were derived. These metadata are typically entered manually by contributors in an error-prone process, and are frequently incomplete. In addition, easy-to-use computational tools that verify the consistency and completeness of metadata describing the libraries to facilitate data reuse, are largely unavailable. Here, we introduce HTSinfer, a Python-based tool to infer metadata directly and solely from bulk RNA-sequencing data generated on Illumina platforms. HTSinfer leverages genome sequence information and diagnostic genes to rapidly and accurately infer the library source and library type, as well as the relative read orientation, 3' adapter sequence and read length statistics. HTSinfer is written in a modular manner, published under a permissible free and open-source license and encourages contributions by the community, enabling easy addition of new functionalities, e.g. for the inference of additional metrics, or the support of different experiment types or sequencing platforms. AVAILABILITY AND IMPLEMENTATION: HTSinfer is released under the Apache License 2.0. Latest code is available via GitHub at https://github.com/zavolanlab/htsinfer, while releases are published on Bioconda. A snapshot of the HTSinfer version described in this article was deposited at Zenodo at 10.5281/zenodo.13985958.

Metadata↗