PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

Ruminococcin A, a new lantibiotic produced by a Ruminococcus gnavus strain isolated from human feces.

When cultivated in the presence of trypsin, the Ruminococcus gnavus E1 strain, isolated from a human fecal sample, was able to produce an antibacterial substance that accumulated in the supernatant. This substance, called ruminococcin A, was purified to homogeneity by reverse-phase chromatography. It was shown to be a 2,675-Da bacteriocin harboring a lanthionine structure. The utilization of Edman degradation and tandem mass spectrometry techniques, followed by DNA sequencing of part of the structural gene, allowed the identification of 21 amino acid residues. Similarity to other bacteriocins present in sequence libraries strongly suggested that ruminococcin A belonged to class IIA of the lantibiotics. The purified ruminococcin A was active against various pathogenic clostridia and bacteria phylogenetically related to R. gnavus. This is the first report on the characterization of a bacteriocin produced by a strictly anaerobic bacterium from human fecal microbiota.

Amino Acid Sequence↗

Gnotobiotic zebrafish reveal evolutionarily conserved responses to the gut microbiota.

Animals have developed the means for supporting complex and dynamic consortia of microorganisms during their life cycle. A transcendent view of vertebrate biology therefore requires an understanding of the contributions of these indigenous microbial communities to host development and adult physiology. These contributions are most obvious in the gut, where studies of gnotobiotic mice have disclosed that the microbiota affects a wide range of biological processes, including nutrient processing and absorption, development of the mucosal immune system, angiogenesis, and epithelial renewal. The zebrafish (Danio rerio) provides an opportunity to investigate the molecular mechanisms underlying these interactions through genetic and chemical screens that take advantage of its transparency during larval and juvenile stages. Therefore, we developed methods for producing and rearing germ-free zebrafish through late juvenile stages. DNA microarray comparisons of gene expression in the digestive tracts of 6 days post fertilization germ-free, conventionalized, and conventionally raised zebrafish revealed 212 genes regulated by the microbiota, and 59 responses that are conserved in the mouse intestine, including those involved in stimulation of epithelial proliferation, promotion of nutrient metabolism, and innate immune responses. The microbial ecology of the digestive tracts of conventionally raised and conventionalized zebrafish was characterized by sequencing libraries of bacterial 16S rDNA amplicons. Colonization of germ-free zebrafish with individual members of its microbiota revealed the bacterial species specificity of selected host responses. Together, these studies establish gnotobiotic zebrafish as a useful model for dissecting the molecular foundations of host-microbial interactions in the vertebrate digestive tract.

Air Sacs↗

Capture of a recombination activating sequence from mammalian cells.

We have developed a genetic trap for identifying sequences that promote homologous DNA recombination. The trap employs a retroviral vector that normally disables itself after one round of replication. Insertion of defined DNA sequences into the vector induced the repair of a 300 base pair deletion, which restored its ability to replicate. Tests of random sequence libraries made in the vector revealed a putative recombination signal (CCCACCC). When this heptamer or an abbreviated form (CCCACC) were reinserted into the vector, they stimulated vector repair and other DNA rearrangements. Mutant forms of these oligomers (eg CCCAACC or CCWACWS) did not. Our data suggest that the recombination events occurred within 48 h after transfection.

Animals↗

The fidelity of template-directed oligonucleotide ligation and the inevitability of polymerase function.

The first living systems may have employed template-directed oligonucleotide ligation for replication. The utility of oligonucleotide ligation as a mechanism for the origin and evolution of life is in part dependent on its fidelity. We have devised a method for evaluating ligation fidelity in which ligation substrates are selected from random sequence libraries. The fidelities of chemical and enzymatic ligation are compared under a variety of conditions. While reaction conditions can be found that promote high fidelity copying, departure from these conditions leads to error-prone copying. In particular, ligation reactions with shorter oligonucleotide substrates are less efficient but more faithful. These results support a model for origins in which there was selective pressure for template-directed oligonucleotide ligation to be gradually supplanted by mononucleotide polymerization.

Base Sequence↗

Tracking-seq: a universal off-target detection approach for CRISPR-Cas genome editing.

Tracking-seq is a highly sensitive method for genome-wide detection of off-target effects in cells edited with diverse genome editing modalities, including Cas9, cytosine base editors, adenine base editors and prime editors. Since most genome editors induce DNA repair pathways and generate single-stranded DNA (ssDNA) intermediates, Tracking-seq leverages this process by tracking replication protein A-a key protein that binds and protects ssDNA-to identify on-target and off-target events. Here we provide a detailed protocol for Tracking-seq, covering genome editing of cells, extraction of replication protein A-bound ssDNA, sequencing library construction and data analysis using our custom computational tool Offtracker. Tracking-seq is applicable to various genome editing scenarios with low cell input, delivering high-performance results. The entire workflow, from genome editing to data analysis, can be completed within 1-2 weeks, making it a rapid solution for assessing genome-wide off-target activity.

CRISPR-Cas Systems↗

Amplified restriction fragment length polymorphism-based mRNA fingerprinting using a single restriction enzyme that recognizes a 4-bp sequence.

Using amplified restriction fragment length polymorphism (AFLP) technology, we have developed a new protocol for the fingerprinting of mRNA that allows systematic comparison of the differential expression of genes between mRNA samples. The major advantage of our protocol is the use of only a single restriction enzyme that recognizes a 4-bp sequence but allows screening of large numbers of different cDNAs. Using this new protocol, we compared mRNA samples obtained from the flower buds of two lines of the common morning glory (Ipomoea purpurea) with red and white flowers, respectively. Approximately 50 bands were observed in each lane of a denaturing polyacrylamide gel and the results were highly reproducible, as indicated by the results of analysis of two sets of independent mRNA samples. Two cDNA fragments, which were differentially amplified in the samples from the two lines, were shown to have been derived from a single gene that was actively expressed in the buds of red flowers but not in those of white flowers. A full-length cDNA of this gene was cloned from a bud cDNA library. Sequence analysis showed that this cDNA carries a sequence highly homologous to the chalcone synthase (CHS) genes, the key enzyme in the flavonoid biosynthetic pathway.

Amino Acid Sequence↗

Genome wide profiling of human embryonic stem cells (hESCs), their derivatives and embryonal carcinoma cells to develop base profiles of U.S. Federal government approved hESC lines.

BACKGROUND: In order to compare the gene expression profiles of human embryonic stem cell (hESC) lines and their differentiated progeny and to monitor feeder contaminations, we have examined gene expression in seven hESC lines and human fibroblast feeder cells using Illumina bead arrays that contain probes for 24,131 transcript probes. RESULTS: A total of 48 different samples (including duplicates) grown in multiple laboratories under different conditions were analyzed and pairwise comparisons were performed in all groups. Hierarchical clustering showed that blinded duplicates were correctly identified as the closest related samples. hESC lines clustered together irrespective of the laboratory in which they were maintained. hESCs could be readily distinguished from embryoid bodies (EB) differentiated from them and the karyotypically abnormal hESC line BG01V. The embryonal carcinoma (EC) line NTera2 is a useful model for evaluating characteristics of hESCs. Expression of subsets of individual genes was validated by comparing with published databases, MPSS (Massively Parallel Signature Sequencing) libraries, and parallel analysis by microarray and RT-PCR. CONCLUSION: we show that Illumina's bead array platform is a reliable, reproducible and robust method for developing base global profiles of cells and identifying similarities and differences in large number of samples.

Carcinoma, Embryonal↗

Characterization and cDNA cloning of a hemoprotein in the salivary glands of the blood-sucking insect, Rhodnius prolixus.

Three major red hemoproteins, named RpSG I, II (identical with prolixin-S) and III, in the salivary glands of the blood-sucking insect, Rhonius prolixus, show homology in N-terminal amino acid (AA) sequences, and are immunologically related. We focussed on one of these proteins, RpSG-I, in this paper. RpSG-I in fresh salivary gland extract was separated into two components (Ia and Ib) by isoelectric focussing gel electrophoresis. Absorption spectra of RpSG-Ia and Ib showed Soret peaks at 400 nm and 420 nm, respectively, suggesting that they are nitric oxide (NO)-unbound and -bound hemoproteins and function as NO-carriers. RpSG-I is stage-specific in appearance, being absent in 3rd and 4th instar nymphs, appearing and increasing gradually in 5th (last) instar nymphs after engorgement, and present in the adult stage. We purified RpSG-I from salivary gland extract by size exclusion and ion exchange HPLCs. It is a single electrophoretic band with an absorption peak at 400 nm, representing the NO-unbound molecule. Full-size cDNA of RpSG-I was cloned by screening with a specific polyclonal antibody from a salivary gland cDNA library. Sequence analysis of RpSG-I cDNA showed an open reading frame encoding a signal peptide (23 AA) and mature protein (179 AA) of 19,778 daltons. The deduced N-terminal AA sequence of the RpSG-I was identical with that of the hemoprotein reported as nitrophorin-3 (Champagne et al., 1995).

Amino Acid Sequence↗

Evolutionary optimization of a nonbiological ATP binding protein for improved folding stability.

Structural comparison of in vitro evolved proteins with biological proteins will help determine the extent to which biological proteins sample the structural diversity available in protein sequence space. We have previously isolated a family of nonbiological ATP binding proteins from an unconstrained random sequence library. One of these proteins was further optimized for high-affinity binding to ATP, but biophysical characterization proved impossible due to poor solubility. To determine if such nonbiological proteins can be optimized for improved folding stability, we performed multiple rounds of mRNA-display selection under increasingly denaturing conditions. Starting from a pool of protein variants, we evolved a population of proteins capable of binding ATP in 3 M guanidine hydrochloride. One protein was chosen for further characterization. Circular dichroism, tryptophan fluorescence, and (1)H-(15)N correlation NMR studies show that this protein has a unique folded structure.

Adenosine Triphosphate↗

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans↗

Interacting RNA species identified by combinatorial selection.

RNA molecules were selected from a random sequence library for their ability to bind to an RNA stem-loop target. Oligonucleotides with extensive Watson-Crick complementarity to the RNA ligand were selected against by inclusion of a blocking oligodeoxynucleotide in the binding phase of the selection protocol. After 18 generations of SELEX (systematic evolution of ligands by exponential enrichment) a single RNA family was predominant in the binding population. The winning aptamer RNA bound the target RNA with an apparent Kd = 70 nM. Structural mapping and Fe(II)-EDTA protection indicated that the target RNA interacted with small unpaired loops in the aptamer structure.

Base Composition↗

Combining expression and comparative evolutionary analysis. The COBRA gene family.

Plant cell shape is achieved through a combination of oriented cell division and cell expansion and is defined by the cell wall. One of the genes identified to influence cell expansion in the Arabidopsis (Arabidopsis thaliana) root is the COBRA (COB) gene that belongs to a multigene family. Three members of the AtCOB gene family have been shown to play a role in specific types of cell expansion or cell wall biosynthesis. Functional orthologs of one of these genes have been identified in maize (Zea mays) and rice (Oryza sativa; Schindelman et al., 2001; Li et al., 2003; Brown et al., 2005; Persson et al., 2005; Ching et al., 2006; Jones et al., 2006). We present the maize counterpart of the COB gene family and the COB gene superfamily phylogeny. Most of the genes belong to a family with two main clades as previously identified by analysis of the Arabidopsis family alone. Within these clades, however, clear differences between monocot and eudicot family members exist, and these are analyzed in the context of Type I and Type II cell walls in eudicots and monocots. In addition to changes at the sequence level, gene regulation of this family in a eudicot, Arabidopsis, and a monocot, maize, is also characterized. Gene expression is analyzed in a multivariate approach, using data from a number of sources, including massively parallel signature sequencing libraries, transcriptional reporter fusions, and microarray data. This analysis has revealed that the expression of Arabidopsis and maize COB gene family members is highly developmentally and spatially regulated at the tissue and cell type-specific level, that gene superfamily members show overlapping and unique expression patterns, and that only a subset of gene superfamily members act in response to environmental stimuli. Regulation of expression of the Arabidopsis COB gene family members has highly diversified in comparison to that of the maize COB gene superfamily members. We also identify BRITTLE STALK 2-LIKE 3 as a putative ortholog of AtCOB.

Amino Acid Sequence↗

Molecular detection of marine bacterial populations on beaches contaminated by the Nakhodka tanker oil-spill accident.

In January 1997, the tanker Nakhodka sank in the Japan Sea, and more than 5000 tons of heavy oil leaked. The released oil contaminated more than 500 km of the coastline, and some still remained even by June 1999. To investigate the long-term influence of the Nakhodka oil spill on marine bacterial populations, sea water and residual oil were sampled from the oil-contaminated zones 10, 18, 22 and 29 months after the accident, and the bacterial populations in these samples were analysed by denaturing gradient gel electrophoresis (DGGE) of PCR-amplified 16S rDNA fragments. The dominant DGGE bands were sequenced, and the sequences were compared with those in DNA sequence libraries. Most of the bacteria in the sea water samples were classified as the Cytophaga-Flavobacterium-Bacteroides phylum, alpha-Proteobacteria or cyanobacteria. The bacteria detected in the oil paste samples were different from those detected in the sea water samples; they were types related to hydrocarbon degraders, exemplified by strains closely related to Sphingomonas subarctica and Alcanivorax borkumensis. The sizes of the major bacterial populations in the oil paste samples ranged from 3.4 x 10(5) to 1.6 x 10(6) bacteria per gram of oil paste, these low numbers explaining the slow rate of natural attenuation.

Accidents↗

Gene expression changes in MDBK cells infected with genotype 2 bovine viral diarrhoea virus.

Bovine viral diarrhoea viruses (BVDVs) are ubiquitous viral pathogens of cattle. These viruses exist as one of two biotypes, cytopathic and noncytopathic, based on the ability to induce cytopathic effect in cell culture. The noncytopathic biotypes are able to establish inapparent, persistent infections in both cell culture and in bovine foetuses of less than 150 days gestation. Interactions with the host cell and the mechanism by which viral tolerance is established are unknown. To examine the changes in gene expression that occur following infection of host cells with BVDV, serial analysis of gene expression (SAGE), a global gene expression technology was used. SAGE allows quantitation of virtually every transcript in a cell type without prior sequence information. Transcript expression levels and identities are determined by sequencing libraries composed of concatamers of 14 base DNA fragments (tags) derived from the 3'-end of each cellular mRNA transcript. Comparison of data obtained from uninfected and BVDV genotype 2-infected cell libraries revealed changes in gene expression associated with distinct biochemical pathways or functions. Isotypes of both alpha- and beta-tubulins were down-regulated, indicating possible dysfunction in cell division and other functions where microtubules play a major role. Expression of genes encoding proteins involved in energy metabolism were expressed at essentially equivalent levels in both infected and uninfected cells. Genes encoding proteins involved in protein translation and post-translational modifications, functions necessary for viral replication, were generally up-regulated. These data indicate that following infection with BVDV, changes in gene expression occur that are beneficial for virus replication while having only minor changes in energy metabolism.

Animals↗

Cryptic simplicity in DNA is a major source of genetic variation.

DNA regions which are composed of a single or relatively few short sequence motifs usually in tandem ('pure simple sequences') have been reported in the genomes of diverse species, and have been implicated in a range of functions including gene regulation, signals for gene conversion and recombination, and the replication of telomeres. They are thought to accumulate by DNA slippage and mispairing during replication and recombination or extension of single-strand ends. In order to systematize the range of DNA simplicity and the genetic nature of the regions that are simple, we have undertaken an extensive computer search of the DNA sequence library of the European Molecular Biology Laboratory (EMBL). We show here that nearly all possible simple motifs occur 5-10 times more frequently than equivalent random motifs. Furthermore, a new computer algorithm reveals the widespread occurrence of significantly high levels of a new type of 'cryptic simplicity' in both coding and noncoding DNA. Cryptically simple regions are biased in nucleotide composition and consist of scrambled arrangements of repetitive motifs which differ within and between species. The universal existence of DNA simplicity from monotonous arrays of single motifs to variable permutations of relatively short-lived motifs suggests that ubiquitous slippage-like mechanisms are a major source of genetic variation in all regions of the genome, not predictable by the classical mutation process.

Animals↗

[Screening differentially expressed genes in human bone marrow stromal cells at defined stage of differentiation.].

To screen differentially expressed genes involved in osteogenic differentiation of human bone marrow stromal cells (BMSCs) at defined stages, subtractive cDNA library was established by means of suppression subtractive hybridization. The BMSCs cultured for 12 and 21 d were used as driver and tester, respectively. A subtract library was successfully constructed and five positive clones were selected from the library. Sequencing analysis and homology comparison showed that the five clones differentially expressed in BMSCs cultured for 21 d were at least 90% homologous with the known genes in human GenBank. It was interestingly found that the osteogenic BMSCs cultured for 21 d differentially expressed decorin and Bax inhibitor 1. RT-PCR was performed to confirm the differentially expressed genes. The results showed that the expression of Bax inhibitor 1 was significantly higher in the cells of 21-day than that of 12-day, while the expression of decorin was only detected in the cells of 21-day.

Apoptosis Regulatory Proteins↗

The CATH database: an extended protein family resource for structural and functional genomics.

The CATH database of protein domain structures (http://www.biochem.ucl.ac.uk/bsm/cath_new) currently contains 34 287 domain structures classified into 1383 superfamilies and 3285 sequence families. Each structural family is expanded with domain sequence relatives recruited from GenBank using a variety of efficient sequence search protocols and reliable thresholds. This extended resource, known as the CATH-protein family database (CATH-PFDB) contains a total of 310 000 domain sequences classified into 26 812 sequence families. New sequence search protocols have been designed, based on these intermediate sequence libraries, to allow more regular updating of the classification. Further developments include the adaptation of a recently developed method for rapid structure comparison, based on secondary structure matching, for domain boundary assignment. The philosophy behind CATHEDRAL is the recognition of recurrent folds already classified in CATH. Benchmarking of CATHEDRAL, using manually validated domain assignments, demonstrated that 43% of domains boundaries could be completely automatically assigned. This is an improvement on a previous consensus approach for which only 10-20% of domains could be reliably processed in a completely automated fashion. Since domain boundary assignment is a significant bottleneck in the classification of new structures, CATHEDRAL will also help to increase the frequency of CATH updates.

Animals↗

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans↗