PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Multichannel genomic recording of biological information with ENGRAM.

Molecular recording is an emerging paradigm for measuring biology over time. Enhancer-mediated genomic recording of activity in multiplex (ENGRAM) is a recently described synthetic biology circuit architecture that converts the transient activity of cis-regulatory elements (CREs) into stable genomic records that can be retrospectively recovered via DNA sequencing. Here we provide a step-by-step protocol for conducting ENGRAM experiments and analyzing the resulting data. We also describe key design considerations for ENGRAM recorders, summarize the strengths and limitations of ENGRAM, and highlight applications, including multiplex signal recording and high-throughput CRE screening. In contrast to other systems for DNA-based recording in mammalian systems, ENGRAM relies on prime editing-mediated insertions to record the activity of a given CRE, such that it is inherently multiplexable-for example, four-base-pair insertions can represent the activities of up to 256 distinct CREs. A further contrast lies with ENGRAM's compatibility with DNA Typewriter, which facilitates the capture of signal order. For users with basic skills in molecular biology, mammalian cell culture and DNA sequencing analysis, ENGRAM experiments can typically be completed within 5-6 weeks.

Genomics↗

Object-oriented parsing of biological databases with Python.

MOTIVATION: While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way. RESULTS: We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.

Databases, Factual↗

Analysis of glycosaminoglycan-derived disaccharides in biologic samples by capillary electrophoresis and protocol for sequencing glycosaminoglycans.

Glycosaminoglycans are biologically significant carbohydrates which either as free chains (hyaluronan) or constituents of proteoglycans (chondroitin/dermatan sulfates, heparin, heparan sulfate and keratan sulfate) participate and regulate several cellular events and (patho)physiological processes. Capillary electrophoresis, due to its high resolving power and sensitivity, has been successfully used for the analysis of glycosaminoglycans. Determination of compositional characteristics, such as disaccharide sulfation pattern, is a useful prerequisite for elucidating the interactions of glycosaminoglycans with matrix effective molecules and, therefore, essential in understanding the biological functions of proteoglycans. The interest in the field of characterization of such biologically important carbohydrates is soaring and advances in this field will signal a new revolution in the area of glycomics equivalent to that of genomics and proteomics. This review focuses on the capillary electrophoresis methods used to determine the disaccharide pattern of glycosaminoglycans in various biologic samples as well as advances in the sequence analysis of glycosaminoglycans using both chromatographic and electrophoretic techniques.

Carbohydrate Sequence↗

Complexities in ETS-domain transcription factor function and regulation: lessons from the TCF (ternary complex factor) subfamily. The Colworth Medal Lecture.

The ETS-domain transcription factor family can be divided into a series of subfamilies. Elk-1 represents the founding member of the ternary complex factor (TCF) subfamily. By focusing on the TCF subfamily, we can demonstrate the complexities that exist in the function and regulation of ETS-domain transcription factors. This article focuses on Elk-1 in detail and summarizes the functions of other TCFs. The key themes covered include the domain structure of the TCFs, the mechanisms of complex formation with serum response factor, regulation of TCFs by mitogen-activated protein kinase cascades, and transcriptional regulatory properties of the TCFs. Finally, the emerging role of the TCFs in vivo is discussed. A picture is developing indicating that, while these proteins exhibit significant sequence and functional conservation, key differences in their structure and regulation are being identified which may relate to unique functions of these proteins in vivo.

Amino Acid Sequence↗

AU-rich transient response transcripts in the human genome: expressed sequence tag clustering and gene discovery approach.

Transient response genes regulate critical biological responses that include cell proliferation, signal transduction events, and responses to exogenous agents such as inflammatory stimuli, microbes, and radiation. An important feature that ensures a timely response is the short half-life of the messenger RNA (mRNA), which is thought to be predominantly mediated by adenylate uridylate-rich sequence elements (AREs) in the 3' untranslated region (3' UTR). The repertoire and extent of transient response genes in the human genome are not known. We used a computational approach to delineate those genes that code for transient ARE mRNAs. We utilized a 3' UTR-specific ARE motif to retrieve and cluster 3'-end ESTs using a refined extraction protocol. With the availability of the entire human genome, we were able to utilize ARE EST clusters for further mining and computational prediction of ARE genes. The described approaches led to the finding of more than 1500 ARE genes in the human genome. In particular, "hidden" ARE mRNAs and alternative forms due to 3'UTR completeness, variant polyadenylation, and splicing were uncovered.

3' Untranslated Regions↗

A single nucleotide substitution in the transcription start signal of the M2 gene of respiratory syncytial virus vaccine candidate cpts248/404 is the major determinant of the temperature-sensitive and attenuation phenotypes.

Respiratory syncytial virus (RSV) cpts248/404 is a live-attenuated, temperature-sensitive (ts) vaccine candidate derived from cole-passaged cpRSV by two rounds of chemical mutagenesis and biological selection. Previous sequence analysis showed that these two steps introduced three single nucleotide substitutions into the cpRSV parent. Two of these occurred with the coding sequence for the L protein, and each resulted in a single amino acid substitution: Gin-831-Leu (248 mutation) and Asp-1183-Glu (404-L mutation). The third mutation resulted in a nucleotide substitution at position 9 of the c/s-acting gene start signal of the M2 gene (404-M2 mutation). In the present study, the genetic basis of attenuation of cpts248/404 was defined by the introduction of each of these mutations (singly or in combination) into a full-length cDNA clone of cpRSV. Recombinant RSV derived from each mutant cDNA was analyzed to determine the contribution of each mutation to the ts and attenuation phenotypes of the virus. This analysis showed that the 248 mutation specifies a significant reduction of plaque formation at 38 degrees and is responsible for an intermediate level of attenuation in mice. In contrast, the 404-L mutation did not contribute to the ts or attenuation phenotype alone or in combination with other mutations and is thus an incidental change. unexpectedly, the 404-M2 mutation alone specified complete restriction of plaque formation at 37 degrees C an a high level of attenuation in mice. This indicates that the level of temperature sensitivity and attenuation of cpts248/404 can be attributed primarily to the 404-M2 mutation. Thus the cpts248/404 virus contains a set of ts and non-ts attenuating mutations, which likely accounts for its genetic stability. The recombinant version of this virus, rA2cp248/404, was phenotypically indistinguishable from cpts248/404 and represents a background into which additional mutations can be introduced as needed to obtain the desired level of attenuation for successful immunization of the very young human infant.

Animals↗

Eclosion hormone of the silkworm Bombyx mori. Expression in Escherichia coli and location of disulfide bonds.

A gene encoding eclosion hormone (EH) from the silkworm, Bombyx mori was chemically synthesized, inserted into a secretion vector and expressed in Escherichia coli, leading to the production of biologically active EH. Sequence analysis of cystine-containing peptides in a thermolysin digest of this EH established the locations of 3 disulfide bonds in the molecule. Evidence was also obtained that the 6 residues at the NH2-terminal are dispensable but 4 residues at the COOH-terminal play an important role in EH activity.

Amino Acid Sequence↗

Analyzing and comparing nucleic acid sequences by hybridization to arrays of oligonucleotides: evaluation using experimental models.

An efficient method was developed for making complete sets of oligonucleotides of defined length, covalently attached to the surface of a glass plate, by synthesizing them in situ. A device carrying all octapurine sequences was used to explore factors affecting molecular hybridization of the tethered oligonucleotides, to develop computer-aided methods for analyzing the data, and to test the feasibility of using the method for sequence analysis. Further development is needed before the method can be used routinely, but our work shows that it has a number of potential advantages over gel-based methods: it should be easy to automate; the quality of the sequence results can be evaluated statistically; it provides a powerful way of comparing related sequences and detecting mutation; it can be applied to both DNA and RNA; and specific motifs can be incorporated into all sequences of the array to focus analysis on sequences of biological interest.

Base Sequence↗

Expression and characterization of recombinant TGF-beta 2 proteins produced in mammalian cells.

Recombinant DNA plasmids coding for transforming growth factor beta 2 (TGF-beta 2) precursor and a hybrid TGF-beta 1(NH2)/beta 2(COOH) molecule consisting of the amino-terminal precursor portion of transforming growth factor-beta 1 (TGF-beta 1) linked in phase to the carboxyl terminus of mature TGF-beta 2 were constructed and transfected into COS cells. Both plasmids directed the synthesis of active TGF-beta 2 which was secreted into the supernatants of transfected cells. The TGF-beta 2 was secreted in a latent form, as an acidification step was required to demonstrate optimal biological activity. Using site-specific anti-peptide antibodies, we show that precursor and mature forms of TGF-beta 2 are produced. A stable Chinese hamster ovary (CHO) cell line expressing the hybrid TGF-beta 1(NH2)/beta 2(COOH) protein was isolated. This cell line secreted both precursor and mature forms of TGF-beta 1(NH2)/beta 2(COOH); acidification was required to demonstrate biological activity. Protein sequence analysis of recombinant TGF-beta 2 produced by this CHO clone demonstrated that correct proteolytic cleavage had occurred, suggesting that the processing signals contained within the TGF-beta 1 amino portion can function in producing mature TGF-beta 2. Receptor binding studies showed that TGF-beta 2 specifically bound predominantly to type III receptors on the surface of human palatal mesenchymal cells. The availability of active TGF-beta 2 should aid in determining its potential therapeutic use as a growth modulator.

Animals↗

Structural diversity of monoclonal CD4 antibodies and their capacity to block the HIV gp120/CD4 interaction.

A number of monoclonal antibodies have been raised against CD4, the receptor on T cells for the HIV envelope glycoprotein gp120. In the present paper we describe biological activities and sequence analysis of seven CD4 MAb. Five of these MAb preparations compete with HIV/gp120 for CD4 binding. The sequences of the variable regions for these MAb were determined in order to ascertain any correlation with selective V gene usage. A relationship was found between the expressed variable region genes and the CD4 recognition pattern. The VH genes that are used can be subdivided into two major groups expressing either a VH gene belonging to the J558 family or to the VGam family. The usage of the VL genes varies, indicating that the epitope specificity is predominantly determined by the rearranged VH genes. The distinct cross-reactivity pattern of these MAb also correlates with their capacity to block binding of recombinant gp120 to CD4 in vitro. Although five of these MAb were able to block gp120 binding none of the CDR sequences shows a relevant homology to the gp120 sequence. This indicates a steric hinderence mechanism for blocking gp120 binding and not a direct interaction with the receptor binding site on CD4. The data also confirm the failure of these MAb as a potential target for receptor mimicry.

Amino Acid Sequence↗

Clustering of DNA sequences in human promoters.

We have determined the distribution of each of the 65,536 DNA sequences that are eight bases long (8-mer) in a set of 13,010 human genomic promoter sequences aligned relative to the putative transcription start site (TSS). A limited number of 8-mers have peaks in their distribution (cluster), and most cluster within 100 bp of the TSS. The 156 DNA sequences exhibiting the greatest statistically significant clustering near the TSS can be placed into nine groups of related sequences. Each group is defined by a consensus sequence, and seven of these consensus sequences are known binding sites for the transcription factors (TFs) SP1, NF-Y, ETS, CREB, TBP, USF, and NRF-1. One sequence, which we named Clus1, is not a known TF binding site. The ninth sequence group is composed of the strand-specific Kozak sequence that clusters downstream of the TSS. An examination of the co-occurrence of these TF consensus sequences indicates a positive correlation for most of them except for sequences bound by TBP (the TATA box). Human mRNA expression data from 29 tissues indicate that the ETS, NRF-1, and Clus1 sequences that cluster are predominantly found in the promoters of housekeeping genes (e.g., ribosomal genes). In contrast, TATA is more abundant in the promoters of tissue-specific genes. This analysis identified eight DNA sequences in 5082 promoters that we suggest are important for regulating gene expression.

Base Sequence↗

Protein secondary structure modelling with probabilistic networks.

In this paper we study the performance of probabilistic networks in the context of protein sequence analysis in molecular biology. Specifically, we report the results of our initial experiments applying this framework to the problem of protein secondary structure prediction. One of the main advantages of the probabilistic approach we describe here is our ability to perform detailed experiments where we can experiment with different models. We can easily perform local substitutions (mutations) and measure (probabilistically) their effect on the global structure. Window-based methods do not support such experimentation as readily. Our method is efficient both during training and during prediction, which is important in order to be able to perform many experiments with different networks. We believe that probabilistic methods are comparable to other methods in prediction quality. In addition, the predictions generated by our methods have precise quantitative semantics which is not shared by other classification methods. Specifically, all the causal and statistical independence assumptions are made explicit in our networks thereby allowing biologists to study and experiment with different causal models in a convenient manner.

Algorithms↗

[The correlation between occurrence of intracranial germ cell tumors and p53 tumor suppressor gene mutations].

Using PCR-SSCP molecular biological techniques and sequencing analysis, an investigation on intracranial germ cell tumors for a correlation between the p53 tumor suppressor gene mutations and the occurrence of these tumors was performed. The results were as follows: 1. The p53 gene mutations were closely related with the development of intracranial germ cell tumors; 2. Among the germ cell tumors, different types of tumors had different p53 tumor gene mutations.

Adolescent↗

Specific modelling of regulatory units in DNA sequences.

Transcriptional control regions are usually composed of a complex arrangement of individual transcriptional elements like protein binding sites. This modular structure allows generation of enormous functional diversity of regulatory regions with a limited set of individual elements. We implemented simple formal representations of these general features of regulatory regions into an algorithm capable of developing complex models reflecting both the element composition and the functional organization of individual elements. Our method (ModelGenerator) requires a training set of at least 10 sequences containing the regulatory regions to be modelled and a very simple initial model which may consist of just two characteristic transcription factor binding sites. We show the capability of our algorithm to expand the initial model solely by comparative sequence analysis leading to complex, biologically meaningful models. A second program (ModelInspector) is capable to scan new sequence data for matches to models defined by ModelGenerator. We show two models for retroviral transcriptional control regions to be highly specific. A search against GenBank using one of the models is shown to be free of false negatives and to produce less than 2 false positives/million nucleotides. Thus, our algorithms appear to be useful tools for the analysis of extremely long genomic sequences which are now becoming available as results of various genome sequencing projects.

Base Sequence↗

A new family of powerful multivariate statistical sequence analysis techniques.

A novel multivariate statistical approach is presented for extracting and exploiting intrinsic information present in our ever-growing sequence data banks. The information extraction from the sequences avoids the pitfalls of intersequence alignment by analyzing secondary invariant functions derived from the sequences in the data bank rather than the sequences themselves. Such typical invariant function is a 20 x 20 histogram of occurrences of amino acid pairs in a given sequence or fragment thereof. To illustrate the potential of the approach an analysis of 10,000 protein sequences from the National Biomedical Research Foundation Protein Identification Resource is presented, whose analysis already reveals great biological detail. For example, zeta-hemoglobin is found to lie close to amphibian and fish chi-hemoglobin which, in turn, is an important clue to the physiological function of this mammalian early embryonic hemoglobin. The multivariate statistical framework presented unifies such apparently unrelated issues as phylogenetic comparisons between a set of sequences and distance matrices between the constituents of the biological sequences. The Multivariate Statistical Sequence Analysis (MSSA) principles can be used for a wide spectrum of sequence analysis problems such as: assignment of family memberships to new sequences, validation of new incoming sequences to be entered into the database, prediction of structure from sequence, discrimination of coding from non-coding DNA regions, and automatic generation of an atlas of protein or DNA sequences. The MSSA techniques represent a self-contained approach to learning continuously and automatically from the growing stream of new sequences. The MSSA approach is particularly likely to play a significant role in major sequencing efforts such as the human genome project.

Amino Acid Sequence↗

Cloning and sequencing of a cDNA encoding a copper-zinc superoxide dismutase enzyme from the marine yeast Debaryomyces hansenii.

Cu-Zn superoxide dismutase (SOD-1) is a ubiquitously occurring eukaryotic enzyme with a variety of important effects on respiring organisms. A gene (dhsod-1) encoding a Cu-Zn superoxide dismutase of the marine yeast Debaryomyces hansenii was cloned using mRNA by the RT-PCR technique. The deduced amino-acid sequence shows approximately 70% homology with that of cytosolic superoxide dismutase from Saccharomyces cerevisiae and Neurospora crassa, as well as lower homologies (between 55 and 65%) with the corresponding enzyme of other eukaryotic organisms, including human. The gene sequence encodes a protein of 153 amino acids with a calculated molecular mass of 15-92 kDa, in agreement with the observed characteristics of the purified protein from D. hansenii.

Amino Acid Sequence↗