PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

WindowMasker: window-based masker for sequenced genomes.

MOTIVATION: Matches to repetitive sequences are usually undesirable in the output of DNA database searches. Repetitive sequences need not be matched to a query, if they can be masked in the database. RepeatMasker/Maskeraid (RM), currently the most widely used software for DNA sequence masking, is slow and requires a library of repetitive template sequences, such as a manually curated RepBase library, that may not exist for newly sequenced genomes. RESULTS: We have developed a software tool called WindowMasker (WM) that identifies and masks highly repetitive DNA sequences in a genome, using only the sequence of the genome itself. WM is orders of magnitude faster than RM because WM uses a few linear-time scans of the genome sequence, rather than local alignment methods that compare each library sequence with each piece of the genome. We validate WM by comparing BLAST outputs from large sets of queries applied to two versions of the same genome, one masked by WM, and the other masked by RM. Even for genomes such as the human genome, where a good RepBase library is available, searching the database as masked with WM yields more matches that are apparently non-repetitive and fewer matches to repetitive sequences. We show that these results hold for transcribed regions as well. WM also performs well on genomes for which much of the sequence was in draft form at the time of the analysis. AVAILABILITY: WM is included in the NCBI C++ toolkit. The source code for the entire toolkit is available at ftp://ftp.ncbi.nih.gov/toolbox/ncbi_tools++/CURRENT/. Once the toolkit source is unpacked, the instructions for building WindowMasker application in the UNIX environment can be found in file src/app/winmasker/README.build. SUPPLEMENTARY INFORMATION: Supplementary data are available at ftp://ftp.ncbi.nlm.nih.gov/pub/agarwala/windowmasker/windowmasker_suppl.pdf

Algorithms↗

Profile of human macrophage transcripts: insights into macrophage biology and identification of novel chemokines.

High throughput partial sequencing of randomly selected cDNA clones has proven to be a powerful tool for examining the relative abundance of mRNAs and for the identification of novel gene products. Because of the important role played by macrophages in immune and inflammatory responses, we sequenced over 3000 randomly selected cDNA clones from a human macrophage library. These sequences represent a molecular inventory of mRNAs from macrophages and provide a catalog of highly expressed transcripts. Two of the most abundant clones encode recently identified CC chemokines. Macrophage-derived chemokine (MDC) plays a complex role in immunoregulation and is a potent chemoattractant for dendritic cells, T cells, and natural killer cells. The chemokine receptor CCR4 binds MDC with high affinity and also responds by calcium flux and chemotaxis. CCR4 has been shown to be expressed by Th2 type T cells. Recent studies also implicate MDC as a major component of the host defense against human immunodeficiency virus.

Base Sequence↗

Phylogenetic relationships among hypotrichous ciliates determined with the macronuclear gene encoding the large, catalytic subunit of DNA polymerase alpha.

The complete macronuclear DNA polymerase alpha gene, previously sequenced in Oxytricha nova, has been cloned from a genomic macronuclear library and sequenced for the hypotrich O. trifallax. Macronuclear DNA clones of DNA polymerase alpha encoding approximately 1000 amino acids, or approximately two-thirds of the open reading frame, have been obtained by PCR and sequenced for Halteria grandinella, Holosticha species, Paraurostyla viridis, Pleurotricha lanceolata, Stylonychia lemnae Teller, Sty. mytilus, Uroleptus gallina, and Urostyla grandis. Phylogenetic relationships inferred from DNA polymerase alpha amino acid sequences have been used to clarify taxonomic relationships previously determined by morphology of the cell cortex. Hypotrich phylogenies based on DNA polymerase alpha amino acid sequences are incongruent with morphological and other molecular phylogenies. Based upon these data, we assert that, contrary to morphological data, O. nova and O. trifallax are different species, and we propose that the oligotrich Halteria grandinella be reclassified as a hypotrich. This work also extends the available data base of eukaryotic DNA polymerase alpha sequences, and suggests new amino acid sequence targets for mutagenesis experiments to continue the functional dissection of DNA pol alpha biochemistry at the molecular level.

Amino Acid Sequence↗

Isolation and identification of partial cDNA clones for endoplasmin, the major glycoprotein of mammalian endoplasmic reticulum.

The amino acid sequences of peptides isolated from murine endoplasmin showed significant homology (approximately 50%) with sequences in the heat-shock proteins 90 and 83 of yeast and Drosophila, respectively, indicating that they are related proteins. Mixed oligonucleotide probes, deduced from the peptide sequences, were used to isolate cDNAs from a murine liver cDNA library. DNA sequencing confirmed the presence of a coding sequence for one of the endoplasmin peptides, formally establishing the authenticity of the cDNA. The identity of the murine and hamster endoplasmin sequences suggests a level of sequence conservation associated with proteins that perform a structural role in cells.

Amino Acid Sequence↗

EST-based identification of genes expressed in the liver of adult seabass (Dicentrarchus labrax, L.).

The scarcity of the genomic resources for some fish species, in spite of their commercial interest, could retard the positive effects that modern biotechnology can offer to aquaculture industry. Then an effort should be made to reduce, as far as it concerns genomic resources, the gap that separates farming species from "model organisms". In this paper, we present an EST project in which we performed single pass sequencing on 1229 randomly selected clones from a sea bass cDNA library. The sequences are deposited in the NCBI database with the following accession numbers: from , from , from and from . EST cataloguing and profiling of seabass will set the basis for functional genomic research in this species, but will also serve for comparative and environmental genomics, for the identification of polymorphic markers useful, for example, to survey the disease resistance of fish, for the discovery of new molecular markers of exposure and for the production of micro- and macro-arrays.

Animals↗

Analysis of ESTs from Lutzomyia longipalpis sand flies and their contribution toward understanding the insect-parasite relationship.

An expressed sequence tag library has been generated from a sand fly vector of visceral leishmaniasis, Lutzomyia longipalpis. A normalized cDNA library was constructed from whole adults and 16,608 clones were sequenced from both ends and assembled into 10,203 contigs and singlets. Of these 58% showed significant similarity to known genes from other organisms, <4% were identical to described sand fly genes, and 42% had no match to any database sequence. Our analyses revealed putative proteins involved in the barrier function of the gut (peritrophins, microvillar proteins, glutamine synthase), digestive physiology (secreted and membrane-anchored hydrolytic enzymes), and the immune response (gram-negative binding proteins, thioester proteins, scavenger receptors, galectins, signaling pathway factors, caspases, serpins, and peroxidases). Sequence analysis of this transcriptome dataset has provided new insights into genes that might be associated with the response of the vector to the development of Leishmania.

Animals↗

Midgut carboxypeptidase from Helicoverpa armigera (Lepidoptera: Noctuidae) larvae: enzyme characterisation, cDNA cloning and expression.

Using synthetic substrates we have characterised carboxypeptidase activity in gut extracts from Helicoverpa armigera larvae. Carboxypeptidase A activity predominates, with only low levels of carboxypeptidase B activity present. Maximum carboxypeptidase A activity occurs over a broad pH range and is inhibited by phenanthroline and potato carboxypeptidase inhibitor. A cDNA clone encoding carboxypeptidase (the first such sequence from a lepidopteran insect) was isolated from a larval gut library. The sequence predicts a secreted polypeptide of Mr 46.6 k with homology to metallocarboxypeptidases from mammalian and invertebrate species. The presence of a serine residue at the active site suggests carboxypeptidase A activity. To further characterise the gene product, the complete cDNA sequence was expressed in insect cells using the baculovirus system. Culture supernatant from these cells contained carboxypeptidase A activity, with no activity against a carboxypeptidase B substrate; the carboxypeptidase B activity in gut extracts must thus be due to a separate enzyme. In agreement with this conclusion, the expressed carboxypeptidase cDNA is a member of a small multigene family. Chronic ingestion of soybean Kunitz trypsin inhibitor by H. armigera larvae results in increased accumulation of carboxypeptidase mRNA in the midgut cells, and an increase in carboxypeptidase A activity detected in gut extract.

Amino Acid Sequence↗

Isolation and chromosomal localization of unique DNA sequences from a human genomic library.

Recombinant bacteriophage lambda from a human genomic library were screened to indentify human DNA inserts having only unique sequences. Unique human inserts were found in about 1% of the phage screened. One recombinant phage, P3-2, was studied in detail. It contains a human insert of 14.7 kilobases with four internal EcoRI cleavage sites. A restriction map was constructed for EcoRI and BamHI sites. Hybridization of the 32P-labeled P3-2 probe to a Southern blot of EcoRI-digested total human DNA yielded distinct bands at positions corresponding to the human insert fragments contained in P3-2. By using a series of human-Chinese hamster somatic cell hybrids containing unique combinations of human chromosomes, the human DNA segment in phage P3-2 was assigned to human chromosome 22 by blot hybridization and synteny analysis. In addition, another human DNA segment, 11.4 kilobases, in phage P3-10 was assigned to human chromosome 10 by similar procedures. With this approach, more unique DNA sequences can be isolated, assigned to specific human chromosomes, and used as genetic markers for gene mapping and linkage, polymorphism, and other genetic studies in the human genome.

Animals↗

Toward cataloguing all rice genes: large-scale sequencing of randomly chosen rice cDNAs from a callus cDNA library.

A large-scale sequence analysis of rice cDNA was performed for a library from rice callus cultured in a medium containing 1 p.p.m. of 2,4-dichlorophenoxyacetic acid. Random sequencing of 2778 cDNA clones generated 2259 non-redundant expressed sequence tags (ESTs). The strategy of sequencing cDNAs can yield quickly a large number of novel genes. After translation, 690 sequences showed a significant amino acid sequence similarity to sequences already known from PIR. The source of known proteins ranged from bacteria to human. In this report, the non-redundant set of 280 identified ESTs is analyzed in detail.

Animals↗

A cDNA encoding a muscle-type tropomyosin cloned from a human epithelial cell line: identity with human fibroblast tropomyosin TM1.

Tropomyosins (TM) expressed by human epithelial cells have only recently been characterized, and no sequence data for them has appeared. We cloned a cDNA encoding a high molecular weight, muscle-type TM from a LS174T human colon carcinoma epithelial cell cDNA library. On sequence analysis this cDNA (TMe1) was virtually identical to the previously reported sequence for human fibroblast TM1 encoded by the hTM beta gene. Expression of TM1/TMe1 mRNA and protein are low in epithelial cells compared with fibroblasts. The results indicate that cells of different developmental lineages (entodermal and mesodermal) can produce identical TM beta gene splice products while regulating expression of those transcripts in a lineage-specific way.

Base Sequence↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗

Expression of carbonic anhydrase IV mRNA in rabbit kidney: stimulation by metabolic acidosis.

The renal carbonic anhydrases, CA II (cytosolic) and CA IV (membrane bound), are believed to facilitate renal acid secretion. We have recently shown that renal cortical sodium dodecyl sulfate (SDS)-resistant hydratase (presumably CA IV) activity was stimulated 241% during chronic metabolic acidosis (CMA). In the present study, we examined the expression and regulation of CA IV mRNA in kidneys from control and acidotic rabbits. To obtain a CA IV probe, we reverse transcribed rabbit kidney total RNA and amplified a approximately 780-base pair (bp) DNA product using primers derived from the human CA IV sequence. Using this product, we screened one-half of a kidney cortex cDNA library and sequenced a 1,194-bp cDNA, which contained the entire open-reading frame of rabbit CA IV. The cDNA was 78% identical to human and 71% to rat CA IV. The deduced amino acid sequence projected an active zinc binding site and two glycosylation sites. Northern analysis yielded a single transcript of approximately 1,600 bp in size expressed more abundantly in cortex and inner medulla than in outer medulla. CA IV mRNA was also expressed abundantly in lung but not in liver or spleen. The high abundance of CA IV mRNA in inner medulla was localized by in situ hybridization to medullary collecting duct cells. Rabbits exposed to CMA showed significant upregulation of CA IV mRNA expression in kidney cortex and outer medulla. Despite a numerical increase, excessive variability precluded statistical significance in the inner medulla. Thus CA IV mRNA was expressed abundantly in kidney and stimulated by CMA, similar to what has been previously observed for SDS-resistant hydratase (presumed CA IV) activity. It is likely that the regulation of CA IV mRNA and activity is relevant to the kidney's adaptation to CMA.

Acidosis↗

Screening and sequence determination of a cDNA encoding the human brain 4-aminobutyrate aminotransferase.

A human brain cDNA library constructed in the lambda ZAP II vector was screened using a fragment of pig brain cDNA encoding 4-aminobutyrate aminotransferase (pGaba-t). A cDNA that encodes the human brain Gaba-t (hGaba-t) has been isolated from the library and sequenced. Using the GenBank and EMBL databases, comparison of the predicted amino-acid sequence of hGaba-t with the pig enzyme revealed 95.4% homology.

4-Aminobutyrate Transaminase↗

Ion-extraction ladder sequencing from bead-based libraries.

Ion-extraction mass spectrometry of ladders of mixtures of isotopically labeled compounds from single beads allows the unambiguous sequencing of bead-based peptides and offers significant advantages over traditional methods of library analysis.

Cluster Analysis↗

Genetic characterization of the bovine leukaemia inhibitory factor (LIF) gene: isolation and sequencing, chromosome assignment and microsatellite analysis.

The bovine leukaemia inhibitory factor was isolated from a phage library and sequences for the gene, in addition to 1213 bp of 5' and 432 bp of 3' sequences, were obtained and compared with other mammalian leukaemia inhibitory factor genes. Comparisons indicated amino acid homologies ranging from 89.6% to 77.2% with the human and mouse homologues, respectively. Analysis of 500 bp of 5' regulatory regions indicated homologies ranging from 83.6% to 74.4% with the corresponding human and sheep sequences, respectively. Additionally, bovine leukaemia inhibitory factor-specific primers were prepared, and a panel of bovine x hamster somatic cell lines were analysed by the polymerase chain reaction (PCR). Data indicated 93% concordance of leukaemia inhibitory factor with aldehyde dehydrogenase 2 located on bovine chromosome 17, and concordance of 81% with myelin basic protein situated on bovine chromosome 24. Southern analysis of selected hybrids confirmed the PCR results, thus conclusively assigning the bovine leukaemia inhibitory factor gene to chromosome 17. Sequence analysis also revealed a microsatellite in intron 2 of the bovine leukaemia inhibitory factor. Analysis of this region by PCR in 22 unrelated Bos taurus and 19 unrelated Bos indicus cattle detected nine different alleles. Polymorphic information content values were 0.53 and 0.80 in B. taurus and B. indicus, respectively. Additionally, the same leukaemia inhibitory factor primers successfully detected allelic variants at this locus in Bos javanicus, Bos guarus and Bison bison but not in Odocoileus virginianus.

Amino Acid Sequence↗

Similarities in the immunoglobulin response and VH gene usage in rhesus monkeys and humans exposed to porcine hepatocytes.

BACKGROUND: The use of porcine cells and organs as a source of xenografts for human patients would vastly increase the donor pool; however, both humans and Old World primates vigorously reject pig tissues due to xenoantibodies that react with the polysaccharide galactose alpha (1,3) galactose (alphaGal) present on the surface of many porcine cells. We previously examined the xenoantibody response in patients exposed to porcine hepatocytes via treatment(s) with bioartficial liver devices (BALs), composed of porcine cells in a support matrix. We determined that xenoantibodies in BAL-treated patients are predominantly directed at porcine alphaGal carbohydrate epitopes, and are encoded by a small number of germline heavy chain variable region (VH) immunoglobulin genes. The studies described in this manuscript were designed to identify whether the xenoantibody responses and the IgVH genes encoding antibodies to porcine hepatocytes in non-human primates used as preclinical models are similar to those in humans. Adult non-immunosuppressed rhesus monkeys (Macaca mulatta) were injected intra-portally with porcine hepatocytes or heterotopically transplanted with a porcine liver lobe. Peripheral blood leukocytes and serum were obtained prior to and at multiple time points after exposure, and the immune response was characterized, using ELISA to evaluate the levels and specificities of circulating xenoantibodies, and the production of cDNA libraries to determine the genes used by B cells to encode those antibodies. RESULTS: Xenoantibodies produced following exposure to isolated hepatocytes and solid organ liver grafts were predominantly encoded by genes in the VH3 family, with a minor contribution from the VH4 family. Immunoglobulin heavy-chain gene (VH) cDNA library screening and gene sequencing of IgM libraries identified the genes as most closely-related to the IGHV3-11 and IGHV4-59 germline progenitors. One of the genes most similar to IGHV3-11, VH3-11cyno, has not been previously identified, and encodes xenoantibodies at later time points post-transplant. Sequencing of IgG clones revealed increased usage of the monkey germline progenitor most similar to human IGHV3-11 and the onset of mutations. CONCLUSION: The small number of IGVH genes encoding xenoantibodies to porcine hepatocytes in non-human primates and humans is highly conserved. Rhesus monkeys are an appropriate preclinical model for testing novel reagents such as those developed using structure-based drug design to target and deplete antibodies to porcine xenografts.

Amino Acid Sequence↗

Isolation and sequence analysis of carp gonadotropin beta-subunit gene.

Using the cDNA encoding the beta subunit of carp gonadotropin (cGTH-beta) as a probe, 14 clones containing cGTH-beta gene have been isolated from a carp genomic library. Nucleotide sequence analysis indicated that the transcriptional unit of the cGTH-beta gene is 1.2 Kb. Similar to mammalian GTH-beta genes, cGTH-beta gene contains three exons and two introns. The locations of the exon/intron junctions also correspond to those of mammalian GTH-beta gene. Using the primer extension assay, the start site of transcription was determined to be 35 or 37 bp upstream from the translation initiation codon. The TATAA box is present in the 5' flanking region of the gene, 21 bp upstream from the start site of transcription. Three polyadenylation signals, AATAAA, are located in the 3' noncoding region, 111, 430, and 442 bp downstream from the stop codon of translation, respectively.

Amino Acid Sequence↗

Identification of the 19S regulatory particle subunits from the rice 26S proteasome.

The 26S proteasome, a protein complex consisting of a 20S proteasome and a pair of 19S regulatory particles (RP), is involved in ATP-dependent proteolysis in eukaryotes. In yeast, the RP contains six different ATPase subunits and, at least, 11 non-ATPase subunits. In this study, we identified the rice homologs of yeast RP subunit genes from the rice expressed sequence tag (EST) library. The complete nucleotide sequences of the homologs for five ATPase subunits, OsRpt1, OsRpt2, OsRpt4, OsRpt5 and OsRpt6, and five non-ATPase subunits, OsRpn7, OsRpn8, OsRpn10, OsRpn11 and OsRpn12, and the partial sequences of one ATPase subunit, OsRpt3, and six non-ATPase subunits, OsRpn1, OsRpn2, OsRpn3, OsRpn5, OsRpn6 and OsRpn9, were determined. Gene homologs of four ATPase subunits, OsRpt1, OsRpt2, OsRpt4 and OsRpt5, and three non- ATPase subunits, OsRpn1, OsRpn2 and OsRpn9, were found to be encoded by duplicated genes. The rice RP was purified by immunoaffinity chromatography with a Protein A column immobilized antibody against rice 20S proteasome, and the subunit composition was determined. The homologs obtained from the rice EST library were identified as genes encoding subunits of RP purified from rice, including the both products of duplicated genes by using electrospray ionization quadrupole time-of-flight mass spectrometry. Post-translational modifications and processing in rice RP subunits were also identified. Various types of RP complex with different subunit compositions are present in rice cells, suggesting the multiple functions of rice proteasome.

Adenosine Triphosphatases↗