PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Novel transfer RNAs that are active in Escherichia coli.

Many of the mammalian mitochondrial tRNAs contain significant nucleotide deletions in the dihydrouridine (D) stem or T psi C stem, so that they cannot fold into the canonical cloverleaf structure. This suggests that alternative forms and shapes are possible for a mitochondrial tRNA that functions in the specialized translational apparatus of the mammalian mitochondria. The question of whether significant structural alterations may be accommodated by a bacterial protein synthesis machinery, such as in Escherichia coli, is unanswered. In this work, all but ten positions in the gene for the 76-nucleotide coding sequence of an E. coli amber suppressor tRNA were permuted and screened for biological activity in vivo. Sequence analysis of a collection of biologically active variants established that many have unusual structures that include base-pair mismatches in helical stems, substitutions of normally conserved bases, and deletions. Independent mutations were obtained that weaken base pairs or tertiary interactions that normally stabilize the coaxial stacking of the D and anticodon stems, suggesting that the translational apparatus can accommodate considerable flexibility in this part of the molecule. The results demonstrate the capacity of the bacterial protein synthetic apparatus to accommodate altered tRNA structures that are not represented by any naturally occurring tRNAs.

Base Sequence↗

Cooperative computer system for genome sequence analysis.

Analysis of the huge volumes of data generated by large scale sequencing projects clearly requires the construction of new sophisticated computer systems. These systems should be able to handle the biological data as well as the results of the analysis of this data. They should also help the user to choose the most appropriate method for a simple task and to string together the methods needed to solve a global analysis task. In this paper we present the prototype of a software system that provides an environment for the analysis of large-scale sequence data. In a first approach this environment has been put to the test within the B. subtilis sequencing project. This system integrates both a descriptive knowledge of the entities involved (genes, regulatory signals etc.) and the methodological knowledge concerning an extendable set of analytical methods (i.e. how to solve a sequence analysis problem through task decomposition and method selection). A knowledge representation based on two existing object-oriented models, named Shirka and SCARP, is used to implement this integrated system. In addition, the present prototype provides a suitable user interface for both displaying the results generated by several methods and interacting with the objects. We present in this paper an overview of the knowledge-based models used to build this integrated system, and a description of the way in which biological entities and sequence analysis tasks are represented. We give illustrations of the co-operation between user and system during the problem solving process. Such a system constitutes a computer workbench for molecular biologists studying the genetic programs of living organisms.

Bacillus subtilis↗

A special-purpose processor for gene sequence analysis.

Advances in computational biology have occurred primarily in the areas of software and algorithm development; new designs of hardware to support biological computing are extremely scarce. This is due, we believe, to the presence of a non-trivial knowledge gap between molecular biologists and computer designers. The existence of this gap is unfortunate, as it has long been known that for certain problems, special-purpose computers can achieve significant cost/performance gains over general-purpose machines. We describe one such computer here: a custom accelerator for gene sequence analysis. The accelerator implements a version of the Needleman-Wunsch algorithm for nucleotide sequence alignment. Sequence lengths are constrained only by available memory; the product of sequence lengths in the current implementation can be up to 2(22). The machine is implemented as two NuBus boards connected to a Mac IIf/x, using a mixture of TTL and FPGA technology clocked at 10 MHz. The boards are completely functional, and yield a 15-fold performance improvement over an unassisted host.

Algorithms↗

The chemistry and biology of thymosin. II. Amino acid sequence analysis of thymosin alpha1 and polypeptide beta1.

The amino acid sequences of two polypeptide components of thymosin Fraction 5 termed thymosin alpha1 and polypeptide beta1 have been established. The sequences were determined by automatic Edman degradation of the intact molecules as well as by manual sequence analysis of the enzymatic cleavage products. Thymosin alpha1, an immunologically active polypeptide, is highly acidic with an isoelectric point of 4.2. This molecule is composed of 28 amino acid residues with acetylserine as the NH2 terminus. A chemically synthesized molecule of thymosin alpha1 has been found to be as active as the natural molecule in our bioassay systems. Polypeptide beta1 is a molecule consisting of 74 amino acid residues and has an isoelectric point of 6.7. This peptide is not biologically active in our assay systems, suggesting that it is not involved in thymic hormone action. The sequence of beta1 was found to be identical with ubiquitin and a portion of protein A24, a nuclear chromosomal protein. The relationships among these proteins are discussed.

Amino Acid Sequence↗

FpvB, an alternative type I ferripyoverdine receptor of Pseudomonas aeruginosa.

Under conditions of iron limitation, Pseudomonas aeruginosa secretes a high-affinity siderophore pyoverdine to scavenge Fe(III) in the extracellular environment and shuttle it into the cell. Uptake of the pyoverdine-Fe(III) complex is mediated by a specific outer-membrane receptor protein, FpvA (ferripyoverdine receptor). Three P. aeruginosa siderovars can be distinguished, each producing a different pyoverdine (type I-III) and a cognate FpvA receptor. Growth of an fpvA mutant of P. aeruginosa PAO1 (type I) under iron-limiting conditions can still be stimulated by its cognate pyoverdine, suggesting the presence of an alternative uptake route for type I ferripyoverdine. In silico analysis of the PAO1 genome revealed that the product of gene PA4168 has a high similarity with FpvA. Inactivation of PA4168 (termed fpvB) in an fpvA mutant totally abolished the capacity to utilize type I pyoverdine. The expression of fpvB is induced by iron limitation in Casamino acids (CAA) and in M9-glucose medium, but, unlike fpvA, not in a complex deferrated medium containing glycerol as carbon source. The fpvB gene was also detected in other P. aeruginosa isolates, including strains producing type II and type III pyoverdines. Inactivation of the fpvB homologues in these strains impaired their capacity to utilize type I ferripyoverdine as a source of iron. Accordingly, introduction of fpvB in trans restored the capacity to utilize type I ferripyoverdine.

Amino Acid Sequence↗

Comparative sequence analysis of the coat proteins of biologically distinct citrus tristeza closterovirus isolates.

The genome of citrus tristeza closterovirus (CTV) consists of a 20 kb single-stranded RNA encapsidated in a 2000 nm long, flexuous particle. Double-stranded (replicative form) RNAs purified from CTV-infected tissue were used to prepare complementary DNA libraries that involved initial first-strand cDNA synthesis followed by selective amplification of the coat protein gene. CTV-specific antisera were used to select clones expressing the coat protein. The coat protein genes of seven Florida and four exotic isolates that differ in their biological properties were cloned and sequenced. The gene is 669 base pairs long and encodes a 223 amino acid protein. There was a greater than 80% homology at both nucleotide and amino acid levels among all the isolates examined. However, comparisons showed that each isolate was found to have several unique amino acid residues. Several blocks of amino acid residues were conserved among all the isolates. A cluster dendrogram showed greater similarities among groups of mild and severe Florida isolates that differed significantly from those of the geographically distinct, exotic CTV isolates.

Amino Acid Sequence↗

Recent developments in laboratory automation using magnetic particles for genome analysis.

The majority of research for genome analysis has shifted from nucleic acid sequencing to the biological functional analysis of each gene. Based on past success, it may not be long before genome diagnostics becomes a widespread tool in human, veterinary and botany research fields. Genome analysis involves the processes of nucleic acid purification, amplification, labeling and signal detection (specific reaction, separation and signal counting). Except for the purification of nucleic acids, the other processes cannot be achieved without instruments, resulting in the advancement of automation processes. Since purification of nucleic acids can be done manually, automating this process has been delayed. However, because the purification of nucleic acids using magnetic particles is suitable for automation, its development has also been accelerated. The need for full automation for other processes is not as great because the majority of genome analysis is to identify the nucleic acid sequence and analyze genome expression. However, once useful diagnostic tools are generated, the desire for full automation will significantly increase. In order to develop realistic and practical automation, various technologies developed for each process in genome analysis have to be evaluated and only a few technologies, useful for automation, selected. The other key factor in automation is the development of methods to manage reagents and reaction mixtures precisely without any risks specifically related to genome handling, such as cross-contamination. Methods using magnetic particles, which have been used for the automation of nucleic acid purification and immunoassay, appear to be the most promising way to automate processes used in biological research.

Animals↗

Predicting and comparing transcription start sites in single cell populations.

The advent of 5' single-cell RNA sequencing (scRNA-seq) technologies offers unique opportunities to identify and analyze transcription start sites (TSSs) at a single-cell resolution. These technologies have the potential to uncover the complexities of transcription initiation and alternative TSS usage across different cell types and conditions. Despite the emergence of computational methods designed to analyze 5' RNA sequencing data, current methods often lack comparative evaluations in single-cell contexts and are predominantly tailored for paired-end data, neglecting the potential of single-end data. This study introduces scTSS, a computational pipeline developed to bridge this gap by accommodating both paired-end and single-end 5' scRNA-seq data. scTSS enables joint analysis of multiple single-cell samples, starting with TSS cluster prediction and quantification, followed by differential TSS usage analysis. It employs a Binomial generalized linear mixed model to accurately and efficiently detect differential TSS usage. We demonstrate the utility of scTSS through its application in analyzing transcriptional initiation from single-cell data of two distinct diseases. The results illustrate scTSS's ability to discern alternative TSS usage between different cell types or biological conditions and to identify cell subpopulations characterized by unique TSS-level expression profiles.

Transcription Initiation Site↗

Type III effectors orchestrate a complex interplay between transcriptional networks to modify basal defence responses during pathogenesis and resistance.

To successfully infect a plant, bacterial pathogens inject a collection of Type III effector proteins (TTEs) directly into the plant cell that function to overcome basal defences and redirect host metabolism for nutrition and growth. We examined (i) the transcriptional dynamics of basal defence responses between Arabidopsis thaliana and Pseudomonas syringae and (ii) how basal defence is subsequently modulated by virulence factors during compatible interactions. A set of 96 genes displaying an early, sustained induction during basal defence was identified. These were also universally co-regulated following other bacterial basal resistance and non-host responses or following elicitor challenges. Eight hundred and eighty genes were conservatively identified as being modulated by TTEs within 12 h post-inoculation (hpi), 20% of which represented transcripts previously induced by the bacteria at 2 hpi. Significant over-representation of co-regulated transcripts encoding leucine rich repeat receptor proteins and protein phosphatases were, respectively, suppressed and induced 12 hpi. These data support a model in which the pathogen avoids detection through diminution of extracellular receptors and attenuation of kinase signalling pathways. Transcripts associated with several metabolic pathways, particularly plastid based primary carbon metabolism, pigment biosynthesis and aromatic amino acid metabolism, were significantly modified by the bacterial challenge at 12 hpi. Superimposed upon this basal response, virulence factors (most likely TTEs) targeted genes involved in phenylpropanoid biosynthesis, consistent with the abrogation of lignin deposition and other wall modifications likely to restrict the passage of nutrients and water to the invading bacteria. In contrast, some pathways associated with stress tolerance are transcriptionally induced at 12 hpi by TTEs.

Amino Acids, Aromatic↗

The Schistosoma mansoni gene index: gene discovery and biology by reconstruction and analysis of expressed gene sequences.

Expressed sequence tag (EST) sequencing and analysis is a primary research tool to identify and characterize the Schistosoma mansoni transcriptome. As part of our gene discovery effort, a total of 5,793 ESTs have been generated from clones selected randomly from complementary DNA (cDNA) libraries constructed from male and female adult worms. Assembly analysis of all the 16,813 public S. mansoni ESTs has identified 1,920 distinct tentative consensus sequences (TCs) and 5,571 nonoverlapping ESTs (singletons). Of these, 376 TCs (20%) and 1,449 singletons (26%) are unique to the SUNY/TIGR sequencing effort. Tentative consensus sequences and singletons were distributed into various categories of biological roles associated with cell structure, metabolism, protein fate, signal transduction, transcription, protein synthesis, transporters, and cell growth. The TCs and singletons represent transcripts that can be used as a resource for functional annotation of genomic sequence data, comparative sequence analysis, and cDNA clone selection for microarray projects. The utility of EST analysis is demonstrated by identifying new protease genes, which may be involved in hemoglobin degradation.

Amino Acid Sequence↗

Isolation of a kidney-specific peptide recognized by alloreactive HLA-A3-restricted human CTL.

The molecular nature of tissue-specific Ags involved in MHC-restricted CTL responses is as yet undefined. To determine the specificity of these peptides, their function, and their possible relationship to allograft rejection, we have utilized human kidney-specific CD8+ CTL clones to screen reversed-phase HPLC (RP-HPLC)-separated self peptides presented by allo-class I molecules. One of these clones is HLA-A3-restricted and the other HLA-B62-restricted, lysing human kidney cell lines but not MHC identical B lymphoblastoid cells which express the appropriate HLA molecules. We have identified a biologically active RP-HPLC fraction containing self peptides eluted from affinity-purified MHC molecules from HLA-A3+ kidney. This peptide is not expressed in HLA-A3+ spleen. Similarly, a HLA-B62-associated peptide fraction was identified in kidney but not in spleen using the HLA-B62-restricted CTL clone. Sequence analysis of the biologically active fraction from HLA-A3 kidney revealed multiple peptides. Because of the ambiguity of the peptide sequence, a mixed peptide library corresponding to this sequence was synthesized that included the HLA-A3 binding motif. The biologically active peptide library was RP-HPLC fractionated and the fraction containing HLA-A3-restricted CTL activity was sequenced. The resulting sequence of the alloreactive HLA-A3-restricted peptide epitope is GPPGVTIVK. By using this unique strategy, we describe the successful isolation and sequencing of an antigenic peptide that is recognized by a human alloreactive kidney-specific CTL clone.

Amino Acid Sequence↗

Light-generated oligonucleotide arrays for rapid DNA sequence analysis.

In many areas of molecular biology there is a need to rapidly extract and analyze genetic information; however, current technologies for DNA sequence analysis are slow and labor intensive. We report here how modern photolithographic techniques can be used to facilitate sequence analysis by generating miniaturized arrays of densely packed oligonucleotide probes. These probe arrays, or DNA chips, can then be applied to parallel DNA hybridization analysis, directly yielding sequence information. In a preliminary experiment, a 1.28 x 1.28 cm array of 256 different octanucleotides was produced in 16 chemical reaction cycles, requiring 4 hr to complete. The hybridization pattern of fluorescently labeled oligonucleotide targets was then detected by epifluorescence microscopy. The fluorescence signals from complementary probes were 5-35 times stronger than those with single or double base-pair hybridization mismatches, demonstrating specificity in the identification of complementary sequences. This method should prove to be a powerful tool for rapid investigations in human genetics and diagnostics, pathogen detection, and DNA molecular recognition.

Base Sequence↗

EPIPDLF: a pretrained deep learning framework for predicting enhancer-promoter interactions.

MOTIVATION: Enhancers and promoters, as regulatory DNA elements, play pivotal roles in gene expression, homeostasis, and disease development across various biological processes. With advancing research, it has been uncovered that distal enhancers may engage with nearby promoters to modulate the expression of target genes. This discovery holds significant implications for deepening our comprehension of various biological mechanisms. In recent years, numerous high-throughput wet-lab techniques have been created to detect possible interactions between enhancers and promoters. However, these experimental methods are often time-intensive and costly. RESULTS: To tackle this issue, we have created an innovative deep learning approach, EPIPDLF, which utilizes advanced deep learning techniques to predict EPIs based solely on genomic sequences in an interpretable manner. Comparative evaluations across six benchmark datasets demonstrate that EPIPDLF consistently exhibits superior performance in EPI prediction. Additionally, by incorporating interpretable analysis mechanisms, our model enables the elucidation of learned features, aiding in the identification and biological analysis of important sequences. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at: https://github.com/xzc196/EPIPDLF.

Deep Learning↗

Transcription factor binding element detection using functional clustering of mutant expression data.

As a powerful tool to reveal gene functions, gene mutation has been used extensively in molecular biology studies. With high throughput technologies, such as DNA microarray, genome-wide gene expression changes can be monitored in mutants. Here we present a simple approach to detect the transcription-factor-binding motif using microarray expression data from a mutant in which the relevant transcription factor is deleted. A core part of our approach is clustering of differentially expressed genes based on functional annotations, such as Gene Ontology (GO). We tested our method with eight microarray data sets from the Rosetta Compendium and were able to detect canonical binding motifs for at least four transcription factors. With the support of chromatin IP chip data, we also predict a possible variant of the Swi4 binding motif and recover a core motif for Arg80. Our approach should be readily applicable to microarray experiments using other types of molecular biology techniques, such as conditional knockout/overexpression or RNAi-mediated 'knockdown', to perturb the expression of a transcription factor. Functional clustering included in our approach may also provide new insights into the function of the relevant transcription factor.

Base Sequence↗

Strategy for the sequence analysis of heparin.

The versatile biological activities of proteoglycans are mainly mediated by their glycosaminoglycan (GAG) components. Unlike proteins and nucleic acids, no satisfactory method for sequencing GAGs has been developed. This paper describes a strategy to sequence the GAG chains of heparin. Heparin, prepared from animal tissue, and processed by proteinases and endoglucuronidases, is 90% GAG heparin and 10% peptidoglycan heparin (containing small remnants of core protein). Raw porcine mucosal heparin was labelled on the amino termini of these core protein remnants with a hydrophobic, fluorescent tag [N-4-(6-dimethylamino-2-benzofuranyl) phenyl (NDBP)-isothiocyanate]. Enrichment of the NDBP-heparin using phenyl-Sepharose chromatography, followed by treatment with a mixture of heparin lyase I and III, resulted in a single NDBP-linkage region tetrasaccharide, which was characterized as deltaUAp(1-->3)-beta-D-Galp(1-->3)-beta-D-Galp(1-->4)-beta-Xylp -(1-->O-Ser-NDBP (deltaUAp is 4-deoxy-alpha-L-threo-hex-4-enopyranosyl uronic acid). Several NDBP-octasaccharides were isolated when NDBP-heparin was treated with only heparin lyase I. The structure of one of these NDBP-octasaccharides, deltaUAp2S(1-->4)-alpha-D-GlcNpAc(1-->4)-alpha-L-IdoAp (1-->4)-alpha-D-GlcNpAc6S(1-->4)-beta-D-GlcAp(1-->3)-beta-D- Galp(1-->3)-beta-D-Galp(1-->4)-beta-Xylp-(1-->O-Ser NDBP (S is sulphate, Ac is acetate), was determined by 1H-NMR and enzymatic methods. Enriched NDBP-heparin was treated with lithium hydroxide to release heparin, and the GAG chain was then labelled at xylose with 7-amino-1,3-naphthalene disulphonic acid (AGA). The resulting AGA-Xyl-heparin was sequenced on gradient PAGE using heparin lyase I and heparin lyase III. A predominant sequence in heparin at the protein core attachment site was deduced to be -D-GlcNp2S6S(or 6OH)(1-->4)-alpha-L-IdoAp2S-(1-->4)-alpha-D-GlcNp2S6S (or60H) (1-->4)-alpha-L-IdoAp2S(1-->4)-alpha-D-GlcNp2S6S( or 6OH)(1-->4)-alpha-L-IdoAp2S(1-->4)-alpha-D-GlcNpAc (1- ->4)-alpha-L-IdoAp(1-->4)-alpha-D-GlcNpAc6S(1-->4)-beta-D-++ +GlcAp(1-->3)-beta-D-Galp(1-->3)-beta-D-Galp(1-->4)-beta-Xyl-AGA.

Animals↗

v-maf, a viral oncogene that encodes a "leucine zipper" motif.

We have molecularly cloned the provirus of the avian musculoaponeurotic fibrosarcoma virus AS42. Nucleotide sequence analysis of a biologically active clone of AS42 showed that this virus encodes a viral oncogene, maf. The deduced amino acid sequence of the v-maf gene product contains a "leucine zipper" motif similar to that found in a number of DNA binding proteins, including the gene products of the fos, jun, and myc oncogenes. However, unlike these oncogenes, the cellular maf gene was not transcriptionally activated by growth stimulation of cultured cells.

Amino Acid Sequence↗