PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Complementary DNA clones of chicken proto-oncogene c-ets: sequence divergence from the viral oncogene v-ets.

The avian acute leukemia virus E26 induces erythroblastosis and myeloblastosis in chickens. The oncogene of this virus includes sequences derived from the cellular gene designated c-ets, which is normally expressed in lymphoid cells and whose product is a protein of apparent molecular weight ca. 54,000 daltons. Complementary DNA clones representing the major transcript of the chicken c-ets proto-oncogene were isolated from a spleen cell library. Sequence analysis of the cDNA revealed that it contains an open reading frame encoding a polypeptide of 441 amino acids with a molecular weight of 49,932 daltons. This open reading frame can be transcribed and translated in vitro into a 50 kd protein that is specifically immunoprecipitated with antiserum to the v-ets oncogene product. Within the central region of homology between c-ets and v-ets, there are only 5 nucleotide substitutions resulting in 4 amino acid changes. However, coding sequences at the 5' and 3' ends of the v-ets oncogene and the chicken c-ets cDNA differ from one another. These changes may be responsible for the differential functions of c-ets and v-ets in cells of different hematopoietic lineages and may account for the pathogenic properties of the v-ets oncogene.

Amino Acid Sequence

Identification of a cDNA clone that contains the complete coding sequence for a 140-kD rat NCAM polypeptide.

Neural cell adhesion molecules (NCAMs) are cell surface glycoproteins that appear to mediate cell-cell adhesion. In vertebrates NCAMs exist in at least three different polypeptide forms of apparent molecular masses 180, 140, and 120 kD. The 180- and 140-kD forms span the plasma membrane whereas the 120-kD form lacks a transmembrane region. In this study, we report the isolation of NCAM clones from an adult rat brain cDNA library. Sequence analysis indicated that the longest isolate, pR18, contains a 2,574 nucleotide open reading frame flanked by 208 bases of 5' and 409 bases of 3' untranslated sequence. The predicted polypeptide encoded by clone pR18 contains a single membrane-spanning region and a small cytoplasmic domain (120 amino acids), suggesting that it codes for a full-length 140-kD NCAM form. In Northern analysis, probes derived from 5' sequences of pR18, which presumably code for extracellular portions of the molecule hybridized to five discrete mRNA size classes (7.4, 6.7, 5.2, 4.3, and 2.9 kb) in adult rat brain but not to liver or muscle RNA. However, the 5.2- and 2.9-kb mRNA size classes did not hybridize to either a large restriction fragment or three oligonucleotides derived from the putative transmembrane coding region and regions that lie 3' to it. The 3' probes did hybridize to the 7.4-, 6.7-, and 4.3-kb message size classes. These combined results indicate that clone pR18 is derived from either the 7.4-, 6.7-, or 4.3-kb adult rat brain RNA size class. Comparison with chicken and mouse NCAM cDNA sequences suggests that pR18 represents the amino acid coding region of the 6.7- or 4.3-kb mRNA. The isolation of pR18, the first cDNA that contains the complete coding sequence of an NCAM polypeptide, unambiguously demonstrates the predicted linear amino acid sequence of this probable rat 140-kD polypeptide. This cDNA also contains a 30-base pair segment not found in NCAM cDNAs isolated from other species. The significance of this segment and other structural features of the 140-kD form of NCAM can now be studied.

Amino Acid Sequence

Computer generation and statistical analysis of a data bank of protein sequences translated from GenBank.

We describe PGtrans, a new and freely available protein sequence databank (2625 sequences, 554198 amino-acids). This data bank is routinely produced by automatic computer translation of the nucleotide sequence library GenBank. The information needed for the translation process (transcriptional orientation, location of coding regions, splice sites and pertinent genetic code) is gathered by the translation program through an "intelligent" scanning of the documentary field of each GenBank entry. Inconsistencies resulting in unexpected termination codons are detected and reported thus allowing the correction of data bank errors. PGtrans is intended as a tool for protein similarity searches. Its reasonable overall size (2 Moctets) makes it suitable for micro-computer environments. Up to date amino-acid composition data and relative abundances of di-, tri-, and tetra-peptides in proteins of known sequences are presented and discussed.

Amino Acid Sequence

Tracking-seq: a universal off-target detection approach for CRISPR-Cas genome editing.

Tracking-seq is a highly sensitive method for genome-wide detection of off-target effects in cells edited with diverse genome editing modalities, including Cas9, cytosine base editors, adenine base editors and prime editors. Since most genome editors induce DNA repair pathways and generate single-stranded DNA (ssDNA) intermediates, Tracking-seq leverages this process by tracking replication protein A-a key protein that binds and protects ssDNA-to identify on-target and off-target events. Here we provide a detailed protocol for Tracking-seq, covering genome editing of cells, extraction of replication protein A-bound ssDNA, sequencing library construction and data analysis using our custom computational tool Offtracker. Tracking-seq is applicable to various genome editing scenarios with low cell input, delivering high-performance results. The entire workflow, from genome editing to data analysis, can be completed within 1-2 weeks, making it a rapid solution for assessing genome-wide off-target activity.

CRISPR-Cas Systems

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Cryptic simplicity in DNA is a major source of genetic variation.

DNA regions which are composed of a single or relatively few short sequence motifs usually in tandem ('pure simple sequences') have been reported in the genomes of diverse species, and have been implicated in a range of functions including gene regulation, signals for gene conversion and recombination, and the replication of telomeres. They are thought to accumulate by DNA slippage and mispairing during replication and recombination or extension of single-strand ends. In order to systematize the range of DNA simplicity and the genetic nature of the regions that are simple, we have undertaken an extensive computer search of the DNA sequence library of the European Molecular Biology Laboratory (EMBL). We show here that nearly all possible simple motifs occur 5-10 times more frequently than equivalent random motifs. Furthermore, a new computer algorithm reveals the widespread occurrence of significantly high levels of a new type of 'cryptic simplicity' in both coding and noncoding DNA. Cryptically simple regions are biased in nucleotide composition and consist of scrambled arrangements of repetitive motifs which differ within and between species. The universal existence of DNA simplicity from monotonous arrays of single motifs to variable permutations of relatively short-lived motifs suggests that ubiquitous slippage-like mechanisms are a major source of genetic variation in all regions of the genome, not predictable by the classical mutation process.

Animals

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans

Atherosclerotic plaque fibroblasts derive from adventitial and medial Pdgfra-lineage-positive cells and predominantly maintain fibroblast identity.

AIMS: Fibroblasts are mesenchymal cells in the healthy vascular adventitia. In atherosclerosis, single-cell sequencing datasets suggest fibroblasts are abundant in plaques. However, their identity, origin, and fate during plaque progression remain unclear, which we aim to unravel here. APPROACH AND RESULTS: To robustly define fibroblast identity, origin, and fate, we employed meta-analyses of 54 single-cell RNA sequencing libraries, including murine smooth muscle cell (Myh11) and endothelial cell (EC) (Cdh5) lineage reporter mice with and without atherosclerosis; human control and atherosclerotic arteries; and murine adventitia and atherosclerotic plaques processed separately from low-density lipoprotein (LDL) receptor knockout (Ldlr-/-) mice. These meta-analyses showed that murine and human plaque fibroblast identity was robustly defined by Pdgfra, Pi16, Cygb, and Serpinf1 mRNA. Ninety-five percent of plaque fibroblasts do not derive from the Myh11 lineage, while no Cdh5-lineage-positive cells were present in the fibroblast cluster. We identified five murine arterial fibroblast subsets in atherosclerotic murine aorta: progenitor fibroblasts, matrix fibroblasts, inflammatory fibroblasts, an EC-like fibroblast subset, detected in both adventitia and plaques, and Col5a3+ fibroblasts, unique to the adventitia. We next studied fibroblast identity, origin, and fate using pseudotime analysis and Pdgfra-CreERT2/tdTomato lineage reporter mice (Pdgfra Lin+). Healthy Pdgfra Lin+ reporter mice showed predominant adventitial tdTomato expression, and infrequent medial and intimal Pdgfra Lin+ cells co-expressing MYH11 and PECAM1, respectively. The Pdgfra Lin+ plaque area increased with diet duration. Pdgfra Lin+ cells largely maintain fibroblast identity in the plaque, while <10% co-express SMC markers (MYH11, SM22&#x3b1;), or contribute to ACTA2+ cap cells. ECs gaining mesenchymal markers are transcriptionally distinct from Cdh5-lineage-negative fibroblasts gaining EC markers. Plaque-resident EC-like fibroblasts displayed a mesenchymal-to-endothelial transition transcriptome, which was induced in human primary fibroblasts in vitro by starvation, and dampened or reversed by IL1B, TGFB1, TGFB3, and oxidized LDL. Cross-species integration showed that all murine plaque fibroblasts were conserved in human atherosclerosis, with one additional subset partially resembling murine subsets, and three human-specific subsets. Importantly, human fibroblast subsets differentially correlated to human plaque traits, with EC-like fibroblasts correlating to plaque instability. CONCLUSION: Our results indicate that 95% of plaque-residing fibroblasts are Myh11 Lin- Plaque fibroblasts have a dual origin, predominantly adventitial Pdgfra Lin+ progenitor fibroblasts, with a minor contribution from medial Pdgfra Lin+ &#xa0;Myh11+ SMCs. Most plaque fibroblasts maintain fibroblast identity. Murine plaque fibroblast subsets were conserved in human atherosclerosis. EC-like fibroblasts are linked to human plaque instability. Intervening in progenitor-to-specific fibroblast transitions could present a new avenue to promote plaque stability in atherosclerosis.

Atherosclerosis

Isoproterenol response following transfection of the mouse beta 2-adrenergic receptor gene into Y1 cells.

The beta 2-adrenergic receptor (beta 2AR) gene was isolated from a mouse genomic library, sequenced and shown to share 93% identity with the hamster beta 2AR cDNA at the amino acid level. This mouse beta 2AR genomic clone was transfected into the Y1 mouse adrenal cortex tumor cell line. Northern blot and S1 nuclease analysis showed that the beta 2AR-transfected cells expressed an mRNA of the appropriate size to encode the receptor. Membrane receptor number and affinities for various beta-adrenergic agonists demonstrated that the transfected clone encoded a beta 2AR protein product. Incubation of the transfected Y1 cells, which do not normally possess beta 2AR, with the beta 2AR agonist, isoproterenol, resulted in an increase in the rate of steroid secretion by these cells as well as a rapid change in cell morphology. This response was fully blocked by the beta 2AR antagonist, propranolol. Prolonged incubation of the cells with isoproterenol resulted in agonist insensitivity and an 80% reduction in membrane receptor number.

Amino Acid Sequence

Cell-cycle-dependent expression of human ornithine decarboxylase.

A human ornithine decarboxylase (ODC) gene probe has been isolated from a Jurkat T-cell cDNA expression library, sequenced, and used to analyze ODC mRNA levels in untransformed human lymphocytes and fibroblasts stimulated to proliferate by various mitogens. The partial cDNA sequence is 86% homologous to the mouse ODC cDNA, and Northern blots indicate that the human and mouse mRNA species are similar in size. ODC mRNA is barely detectable in quiescent human T lymphocytes and undetectable in density-arrested W138 fibroblasts. Following stimulation of T-lymphocyte proliferation with phytohemagglutinin, the ODC mRNA level rises to a peak around mid G1 phase and decreases as the cells enter S phase. Serum stimulation of density-arrested fibroblasts results in an elevation of the ODC mRNA level which persists throughout the cell cycle. Epidermal growth factor (20 ng/ml) but not insulin (10 mg/ml) or dexamethasone (55 ng/ml) stimulates ODC expression in quiescent W138 fibroblasts. Southern blots suggest that human cells have a single copy of the ODC gene.

Amino Acid Sequence

Isolation and characterization of a fruit-specific cDNA and the corresponding genomic clone from tomato.

Differential screening of a cDNA bank constructed from ripe tomato fruit mRNA allowed the isolation of cDNA clone 2A11 which is entirely fruit-specific, is expressed at steadily increasing levels from anthesis to breaker, and accounts for approximately 1% of the messenger RNA in mature tomato fruit. A genomic clone corresponding to the 2A11 cDNA was isolated from a tomato genomic library. Sequence comparison of the cDNA clone with the genomic clone shows they are identical over the shared region with the genomic clone possessing a single large intron near the 5' end of the message. The open reading frame of 2A11 would encode a sulfur-rich polypeptide 96 amino acids in length. The identity of the putative protein is unknown. In situ hybridization shows that the 2A11 message is found throughout the pericarp cells in a tomato fruit. In contrast, in situ hybridization of early ripening stages with a polygalacturonase probe shows higher mRNA levels in cells of the outer pericarp and cells surrounding the vascular regions of the pericarp.

Amino Acid Sequence

One of two different ADP-glucose pyrophosphorylase genes from potato responds strongly to elevated levels of sucrose.

The key regulatory step in starch biosynthesis is catalyzed by the tetrameric enzyme ADP-glucose pyrophosphorylase (AGPase). In leaf and storage tissue, the enzyme catalyzes the synthesis of ADP-glucose from glucose-1-phosphate and ATP. Using heterologous probes from maize, two sets (B and S) of cDNA clones encoding potato AGPase were isolated from a tuberspecific cDNA library. Sequence analysis revealed homology to other plant and bacterial sequences. Transcript sizes are 1.9 kb (AGPase B) and 2.1 kb (AGPase S). Northern blot experiments show that the two genes differ in their expression patterns in different organs. Furthermore, one of the genes (AGPase S) is strongly inducible by metabolizable carbohydrates (e.g. sucrose) at the RNA level. The accumulation of AGPase S mRNA was always found to be accompanied by an increase in starch content. This suggests a link between AGPase S expression and the status of a tissue as either a sink for or a source of carbohydrates. By contrast, expression of AGPase B is much less variable under various experimental conditions.

Amino Acid Sequence

Uncovering host transcriptional responses to tilapia lake virus (TiLV) through De novo RNA-seq assembly in Nile tilapia, Oreochromis niloticus.

Tilapia lake virus (TiLV) has emerged as an important pathogen that negatively impacts tilapia farming globally. Using RNA sequencing technology, this study investigated the liver transcriptomic profile of apparently healthy and TiLV-infected Oreochromis niloticus from wild. RNA sequence libraries generated 3,356 differentially expressed genes (DEGs), with 1,726 genes that were upregulated. Kyoto Encyclopedia of Genes and Genomes (KEGG) analysis identified 680 different pathways with differential regulation of the metabolic and immune-related pathways indicating that TiLV may interfere in host metabolism and replicate to establish the infection. This study provides transcriptomic insights into the liver responses of naturally TiLV-infected wild O. niloticus and highlights key immune and metabolic pathways associated with viral infection.

Animals

Isolation of a putative collagen-like gene from the sea urchin Paracentrotus lividus.

Using a Caenorhabditis elegans collagen probe we have isolated a 17.6 kb clone from a Paracentrotus lividus genomic library. Sequencing of nearly 2.6 kb identified five open reading frames flanked at both sides by splice site consensus sequences and coding for ninety-five uninterrupted Gly-X-Y repeats. Interestingly, three of the putative exons exhibit sizes which are identical to those featured by vertebrate fibrillar collagen genes, namely 54 bp and 99 bp. Hybridization of the Gly-X-Y encoding sequences to RNA extracted from different developmental stages identified a specific 6 kb transcript, which appears first at mid-gastrula, greatly increases at prism and then progressively accumulates until pluteus stage. Based on these data, we conclude that the genomic clone is likely to code for a developmentally regulated mRNA whose expression coincides with the reported time of appearance of collagenous molecules in the sea urchin embryo.

Amino Acid Sequence

Plasmodium berghei: cloning of the circumsporozoite protein gene.

A DNA fragment encoding the carboxy terminal 80% of the Plasmodium berghei circumsporozoite protein was selected from a genomic DNA expression library. Sequencing revealed that the P. berghei circumsporozoite protein was similar in overall structure to circumsporozoite proteins from other malaria species, although the central repeat region was unique in comprising two different blocks of tandem peptide repeats: 11 eight amino acid repeats with predominant sequence DPAPPNAN were followed by 16 two amino repeats, predominantly PQ. The P. berghei circumsporozoite protein exhibited limited, but about equal amino acid homology to circumsporozoite proteins from P. knowlesi, P. vivax, and P. falciparum, indicating that P. berghei is not closely related to any of these other malaria species. Cloning of the P. berghei circumsporozoite protein gene will allow direct testing of sporozoite vaccines in mice.

Amino Acid Sequence

Isolation and characterization of a previously unrecognized myosin heavy chain gene present in the Syrian hamster.

A full length (25,000 base-pair) myosin heavy chain gene completely contained within a single cosmid clone was isolated from a Syrian hamster cosmid genomic library. Sequence comparison of the 3' untranslated region indicated the presence of a 75% homology with the rat embryonic myosin heavy chain gene. Extensive 5' flanking region regulatory element conservation was also found when the sequence was compared to the rat myosin heavy chain gene. S1 nuclease digestion analysis, however, indicated that the Syrian hamster myosin heavy chain gene exhibited expression in adult Syrian hamster ventricular tissue, as well as the adult vastus medialis, a fast twitch skeletal muscle. Expression also appears to be enhanced in myopathic relative to control hearts. This myosin heavy chain gene is neither the alpha nor beta cardiac myosin heavy chain gene, but is a unique, previously unrecognized, myosin heavy chain gene present in both myocardial and skeletal muscle tissues.

Animals

Recruitment to the cytoplasm of a cellular lamin-like protein from the nucleus during a poxvirus infection.

Monoclonal antibodies (Mabs) directed against core proteins of rabbit poxvirus (RPV) have proven effective in the identification of host cell proteins such as RNA polymerase II (Pol II) that may play a role in the infectious process (D. K. Morrison and R. W. Moyer, 1986, Cell 44, 587-596). In this article we describe a Mab that has allowed the detection and characterization of a lamin-like protein derived from the nucleus of the infected cell, which like Pol II is recruited to the cytoplasm following RPV infection. A portion of the gene encoding this protein has been isolated through the screening of a lambda gt11 expression vector library. Sequence analysis of the gene shows it to be derived from a member of the HindIII 1.9-kb repetitive element, a family of mammalian repetitive sequences that are highly conserved. Immunoblot analysis and sequence analysis of the open reading frame show divergent relatedness to certain nuclear lamins. The protein is not, however, one of the three principal lamins characterized to date, but instead appears to be a perinuclear protein related to the highly conserved nuclear lamins that is recruited to the cytoplasm during the infectious process.

Amino Acid Sequence