PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Variable window binding for mutually exclusive alternative splicing.

BACKGROUND: Genes of advanced organisms undergo alternative splicing, which can be mutually exclusive, in the sense that only one exon is included in the mature mRNA out of a cluster of alternative choices, often arranged in a tandem array. In many cases, however, the details of the underlying biologic mechanisms are unknown. RESULTS: We describe 'variable window binding'--a mechanism used for mutually exclusive alternative splicing by which a segment ('window') of a conserved nucleotide 'anchor' sequence upstream of the exon 6 cluster in the pre-mRNA of the fruitfly Dscam gene binds to one of the introns, thereby activating selection of the exon directly downstream from the binding site. This mechanism is supported by the fact that the anchor sequence can be inferred solely from a comparison of the intron sequences using a genetic algorithm. Because the window location varies for each exon choice, regulation can be achieved by obstructing part of that sequence. We also describe a related mechanism based on competing pre-mRNA stem-loop structures that could explain the mutually exclusive choice of exon 17 of the Dscam gene. CONCLUSION: On the basis of comparative sequence analysis, we propose efficient biologic mechanisms of alternative splicing of the Drosophila Dscam gene that rely on the inherent structure of the pre-mRNA. Related mechanisms employing 'locus control regions' could be involved on other occasions of mutually exclusive choices of exons or genes.

Algorithms↗

Identification and mapping of resistance gene analogs (RGAs) in Prunus: a resistance map for Prunus.

The genetically anchored physical map of peach is a valuable tool for identifying loci controlling economically important traits in Prunus. Breeding for disease resistance is a key component of most breeding programs. The identification of loci for pathogen resistance in peach provides information about resistance loci, the organization of resistance genes throughout the genome, and permits comparison of resistance regions among other genomes in the Rosaceae. This information will facilitate the breeding of resistant species of Prunus. A candidate gene approach was implemented for locating resistance loci in the genome of peach. Candidate genes representing NBS-LRR, kinase, transmembrane domain classes, as well as, pathogen response (PR) proteins and resistance-associated transcription factors were hybridized to a peach BAC library and mapped by using the peach physical map database and the Genome Database for Rosaceae (GDR). A resistance map for Prunus was generated and currently contains 42 map locations for putative resistance regions distributed among 7 of the 8 linkage groups.

Amino Acid Sequence↗

Impact of cultivation on characterisation of species composition of soil bacterial communities.

The species composition of culturable bacteria in Scottish grassland soils was investigated using a combination of Biolog and 16S rDNA analysis for characterisation of isolates. The inclusion of a molecular approach allowed direct comparison of sequences from culturable bacteria with sequences obtained during analysis of DNA extracted directly from the same soil samples. Bacterial strains were isolated on Pseudomonas isolation agar (PIA), a selective medium, and on tryptone soya agar (TSA), a general laboratory medium. In total, 12 and 21 morphologically different bacterial cultures were isolated on PIA and TSA, respectively. Biolog and sequencing placed PIA isolates in the same taxonomic groups, the majority of cultures belonging to the Pseudomonas (sensu stricto) group. However, analysis of 16S rDNA sequences proved more efficient than Biolog for characterising TSA isolates due to limitations of the Microlog database for identifying environmental bacteria. In general, 16S rDNA sequences from TSA isolates showed high similarities to cultured species represented in sequence databases, although TSA-8 showed only 92.5% similarity to the nearest relative, Bacillus insolitus. In general, there was very little overlap between the culturable and uncultured bacterial communities, although two sequences, PIA-2 and TSA-13, showed >99% similarity to soil clones. A cloning step was included prior to sequence analysis of two isolates, TSA-5 and TSA-14, and analysis of several clones confirmed that these cultures comprised at least four and three sequence types, respectively. All isolate clones were most closely related to uncultured bacteria, with clone TSA-5.1 showing 99.8% similarity to a sequence amplified directly from the same soil sample. Interestingly, one clone, TSA-5.4, clustered within a novel group comprising only uncultured sequences. This group, which is associated with the novel, deep-branching Acidobacterium capsulatum lineage, also included clones isolated during direct analysis of the same soil and from a wide range of other sample types studied elsewhere. The study demonstrates the value of fine-scale molecular analysis for identification of laboratory isolates and indicates the culturability of approximately 1% of the total population but under a restricted range of media and cultivation conditions.

Journal Article↗

Amino acid sequence of CAP2b, an insect cardioacceleratory peptide from the tobacco hawkmoth Manduca sexta.

The primary structure of a novel insect neuropeptide, Cardioacceleratory Peptide 2b (CAP2b), from the tobacco hawkmoth Manduca sexta has been established using a combination of mass spectroscopy, Edman degradation microsequencing, amino acid analysis, and biological assays. The sequence of CAP2b, pyroGlu-Leu-Tyr-Ala-Phe-Pro-Arg-Val-amide, has a molecular weight of 974.6 and is blocked at both the amino and carboxyl ends. Examination of several national computer protein data bases failed to reveal other peptides or proteins with any sequence homology to CAP2b indicating that this is likely to be a novel insect neuropeptide. This peptide may be a general activator of insect viscera since it causes an increase in heart rate in Manduca and in Drosophila, and has also been implicated in the regulation of fluid secretion by the Malphigian tubules of Drosophila.

Amino Acid Sequence↗

Biosequence exegesis.

Annotation of large-scale gene sequence data will benefit from comprehensive and consistent application of well-documented, standard analysis methods and from progressive and vigilant efforts to ensure quality and utility and to keep the annotation up to date. However, it is imperative to learn how to apply information derived from functional genomics and proteomics technologies to conceptualize and explain the behaviors of biological systems. Quantitative and dynamical models of systems behaviors will supersede the limited and static forms of single-gene annotation that are now the norm. Molecular biological epistemology will increasingly encompass both teleological and causal explanations.

Animals↗

Morphological, phylogenetic and biological characteristics of Ectropis obliqua single-nucleocapsid nucleopolyhedrovirus.

The tea looper caterpillar, Ectropis obliqua, is one of the major pests of tea bushes. E. obliqua single-nucleocapsid nucleopolyhedrovirus (EcobSNPV) has been used as a commercial pesticide for biocontrol of this insect. However only limited genetic analysis for this important virus has been done up to now. EcobSNPV was characterized in this study. Electron microscopy analysis of the occlusion body showed polyhedra of 0.7 to 1.7 mum in diameter containing a single nucleocapsid per envelope of the virion. A 15.5 kb genomic fragment containing EcoRI-L, EcoRI-N and HindIII-F fragments, was sequenced. Analysis of the sequence revealed that the fragment contained eleven potential open reading frames (ORFs): lef-1, egt, 38.7k, rr1, polyhedrin, orf1629, pk-1, hoar and homologues to Spodoptera exigua multicapsid NPV (SeMNPV) ORFs 15, 28, and 29. Gene arrangement and phylogeny analysis suggest that EcobSNPV is closely related to the previously described Group II NPV. Bioassays on lethal concentration (LC(50) and LC(90)) and lethal time (LT(50) and LT(90)) were conducted to test the susceptibility of E. obliqua larvae to the virus.

Animals↗

Improved K-means clustering algorithm for exploring local protein sequence motifs representing common structural property.

Information about local protein sequence motifs is very important to the analysis of biologically significant conserved regions of protein sequences. These conserved regions can potentially determine the diverse conformation and activities of proteins. In this work, recurring sequence motifs of proteins are explored with an improved K-means clustering algorithm on a new dataset. The structural similarity of these recurring sequence clusters to produce sequence motifs is studied in order to evaluate the relationship between sequence motifs and their structures. To the best of our knowledge, the dataset used by our research is the most updated dataset among similar studies for sequence motifs. A new greedy initialization method for the K-means algorithm is proposed to improve traditional K-means clustering techniques. The new initialization method tries to choose suitable initial points, which are well separated and have the potential to form high-quality clusters. Our experiments indicate that the improved K-means algorithm satisfactorily increases the percentage of sequence segments belonging to clusters with high structural similarity. Careful comparison of sequence motifs obtained by the improved and traditional algorithms also suggests that the improved K-means clustering algorithm may discover some relatively weak and subtle sequence motifs, which are undetectable by the traditional K-means algorithms. Many biochemical tests reported in the literature show that these sequence motifs are biologically meaningful. Experimental results also indicate that the improved K-means algorithm generates more detailed sequence motifs representing common structures than previous research. Furthermore, these motifs are universally conserved sequence patterns across protein families, overcoming some weak points of other popular sequence motifs. The satisfactory result of the experiment suggests that this new K-means algorithm may be applied to other areas of bioinformatics research in order to explore the underlying relationships between data samples more effectively.

Algorithms↗

Peptidomics: the comprehensive analysis of peptides in complex biological mixtures.

Progress in the sequencing of genomes has resulted in an increasing demand for a functional analysis of gene products in order to understand the underlying physiology. Proteomics has established itself as a highly valuable technology for producing functionally related data in an unparalleled fashion, but is methodologically restricted to the analysis of proteins with higher molecular masses (>10 kDa). The development of a technology which covers peptides with low molecular weight and small proteins (0.5 to 15 kDa) was necessary, since peptides, amongst them families of hormones, cytokines and growth factors, play a central role in many biological processes. To summarise the technologies used for this approach the term "peptidomics" is introduced. In this article, we present the rationale and first results of a novel, universal peptide display approach for the analysis and visualisation of peptides and small proteins from biological samples. Special attention is given to samples derived from extracellular fluids such as blood plasma and cerebrospinal fluid. Additionally, a high throughput identification procedure for the analysis of peptides in their native and processed molecular form is outlined.

Combinatorial Chemistry Techniques↗

New bioinformatics tools for viral genome analyses at Viral Bioinformatics--Canada.

Viruses are much smaller than prokaryotes and eukaryotes, and it is now practical to sequence closely related members of virus families, strains, or even different isolates recovered during the course of an outbreak. However, comparative analysis of viral genomes requires the development of novel bioinformatics tools that allow us to align, edit, compare and interact with these genomes at all levels, from whole genome, to gene family, to single nucleotide polymorphisms. Comparative viral genomics can lead to the identification of the core characteristics that define a virus family, as well as the unique properties of viral species or isolates that contribute to variations in pathogenesis. This paper describes a number of tools, mainly developed for Viral Bioinformatics--Canada, that can be used for annotation and comparative genomic analysis of poxviruses. Nonetheless, these tools are also broadly applicable to other virus families.

Base Sequence↗

The prepro vasoactive intestinal contractor (VIC)/endothelin-2 gene (EDN2): structure, evolution, production, and embryonic expression.

Murine vasoactive intestinal contractor (VIC) and its human analog endothelin-2 (ET2) are potent vasoactive hormones composed of 21 amino acids. To study the structural characteristics of the VIC/ET2 gene (HGMW-approved symbol EDN2), we isolated the full length of the mouse VIC gene. Sequence analysis indicates that a biologically active mature VIC peptide is produced from a 175-residue precursor protein; preproVIC (PPVIC). Several remarkable similarities of the PPVIC gene to the human preproendothelin-1 gene strongly suggest that the two genes have arisen from a common progenitor by gene duplication. Transfection of ACHN adenocarcinoma cells with the cDNA resulted in the production of VIC peptide. VIC production was increased by the deletion of the 3'-untranslated region, which contains an AU-rich mRNA destabilizing sequence. Increased PPVIC gene expression during the late embryonic stage suggests an important function in development. This study provides the basis for disruption and regulation analysis of the gene, which may lead to a better understanding of VIC/ET2's physiological significance.

Amino Acid Sequence↗

Evaluating distance functions for clustering tandem repeats.

Tandem repeats are an important class of DNA repeats and much research has focused on their efficient identification, their use in DNA typing and fingerprinting, and their causative role in trinucleotide repeat diseases such as Huntington Disease, myotonic dystrophy, and Fragile-X mental retardation. We are interested in clustering tandem repeats into groups or families based on sequence similarity so that their biological importance may be further explored. To cluster tandem repeats we need a notion of pairwise distance which we obtain by alignment. In this paper we evaluate five distance functions used to produce those alignments--Consensus, Euclidean, Jensen-Shannon Divergence, Entropy-Surface, and Entropy-weighted. It is important to analyze and compare these functions because the choice of distance metric forms the core of any clustering algorithm. We employ a novel method to compare alignments and thereby compare the distance functions themselves. We rank the distance functions based on the cluster validation techniques--Average Cluster Density and Average Silhouette Width. Finally, we propose a multi-phase clustering method which produces good-quality clusters. In this study, we analyze clusters of tandem repeats from five sequences: Human Chromosomes 3, 5, 10 and X and C. elegans Chromosome III.

Algorithms↗

A scale-independent signal processing method for sequence analysis.

In this paper, we present methods to detect and localize patterns in biologically related protein sequences (family). The patterns common to the sequences of the family are detected by using Fourier analysis. No previous scales (codes) are needed, they are actually produced as a result of the analysis procedure, together with the frequencies of the Fourier decompositions. Characteristic features of the family are thus expressed as (code-frequency) pairs. Various tools are proposed in order to localize the patterns, to compare the codes, and to evaluate the proximity of an arbitrary sequence to the investigated family. The general strategy is illustrated on a family composed of calcium-binding proteins.

Amino Acid Sequence↗

New insights into the nature and phylogeny of prasinophyte antenna proteins: Ostreococcus tauri, a case study.

The basal position of the Mamiellales (Prasinophyceae) within the green lineage makes these unicellular organisms key to elucidating early stages in the evolution of chlorophyll a/b-binding light-harvesting complexes (LHCs). Here, we unveil the complete and unexpected diversity of Lhc proteins in Ostreococcus tauri, a member of the Mamiellales order, based on results from complete genome sequencing. Like Mantoniella squamata, O. tauri possesses a number of genes encoding an unusual prasinophyte-specific Lhc protein type herein designated "Lhcp". Biochemical characterization of the complexes revealed that these polypeptides, which bind chlorophylls a, b, and a chlorophyll c-like pigment (Mg-2,4-divinyl-phaeoporphyrin a5 monomethyl ester) as well as a number of unusual carotenoids, are likely predominant. They are retrieved to some extent in both reaction center I (RCI)- and RCII-enriched fractions, suggesting a possible association to both photosystems. However, in sharp contrast to previous reports on LHCs of M. squamata, O. tauri also possesses other LHC subpopulations, including LHCI proteins (encoded by five distinct Lhca genes) and the minor LHCII polypeptides, CP26 and CP29. Using an antibody against plant Lhca2, we unambiguously show that LHCI proteins are present not only in O. tauri, in which they are likely associated to RCI, but also in other Mamiellales, including M. squamata. With the exception of Lhcp genes, all the identified Lhc genes are present in single copy only. Overall, the discovery of LHCI proteins in these prasinophytes, combined with the lack of the major LHCII polypeptides found in higher plants or other green algae, supports the hypothesis that the latter proteins appeared subsequent to LHCI proteins. The major LHC of prasinophytes might have arisen prior to the LHCII of other chlorophyll a/b-containing organisms, possibly by divergence of a LHCI gene precursor. However, the discovery in O. tauri of CP26-like proteins, phylogenetically placed at the base of the major LHCII protein clades, yields new insight to the origin of these antenna proteins, which have evolved separately in higher plants and green algae. Its diverse but numerically limited suite of Lhc genes renders O. tauri an exceptional model system for future research on the evolution and function of LHC components.

Amino Acid Sequence↗

Intradialytic cytokine gene expression.

Along with the numerous technological improvements in molecular biology, polymerase chain reaction, which permits analysis of sequences of a very small amount of biological material, enables evaluation of hemodialysis-induced gene transcription of inflammatory cytokines. Blood samples drawn from 22 hemodialysis patients, treated with cellulose-derived or synthetic membranes, were collected at 0 and 15 min of hemodialysis. Total RNA, purified from mononuclear cells, was reverse transcribed and cDNA amplified by polymerase chain reaction primed with specific oligomers in order to determine tumor necrosis factor alpha (TNF alpha), interleukin (IL) 1 beta and IL6 gene expression. Plasma samples were collected at 0 and 180 min for detection of mature cytokines by enzyme immunoassay with plates pre-coated with monoclonal antibodies to TNF alpha, IL1 beta and IL6. A significant increase in TNF alpha mRNA was detected at 15 min of hemodialysis in 12 of 22 patients: 5 of 9 treated with cuprophan; 3 of 3 with cellulose triacetate; 3 of 5 with polysulfone, and only 1 of 5 treated with polymethyl-methacrylate membranes. A parallel increase in IL1 beta or IL6 mRNA was detected, and significant relationships were found between TNF alpha and IL1 beta (p < 0.001), and IL1 beta and IL6 gene expression (p < 0.05). Increased levels of mature TNF alpha and IL1 beta molecules in plasma were detected in the majority of patients showing an increased cytokine gene expression. However, the absolute amount of cytokine mRNA transcription at 15 min did not predict the levels of mature molecules reached in plasma at 180 min. Cytokine mRNA transcription is quite common at the beginning of a dialysis run. Possibly due to intracellular degradation of critical sequences of cytokine mRNA, gene expression does not necessarily imply translation into mature protein. It is suggested that mechanisms related to cell-to-cell interaction, which may possibly involve procytokine biology, are needed to drive phenomena of cytokine activation to clinical effectiveness.

Adult↗

Multiple sequence alignment.

Multiple sequence alignments are an essential tool for protein structure and function prediction, phylogeny inference and other common tasks in sequence analysis. Recently developed systems have advanced the state of the art with respect to accuracy, ability to scale to thousands of proteins and flexibility in comparing proteins that do not share the same domain architecture. New multiple alignment benchmark databases include PREFAB, SABMARK, OXBENCH and IRMBASE. Although CLUSTALW is still the most popular alignment tool to date, recent methods offer significantly better alignment quality and, in some cases, reduced computational cost.

Algorithms↗

A library-based bioinformatics services program.

Support for molecular biology researchers has been limited to traditional library resources and services in most academic health sciences libraries. The University of Washington Health Sciences Libraries have been providing specialized services to this user community since 1995. The library recruited a Ph.D. biologist to assess the molecular biological information needs of researchers and design strategies to enhance library resources and services. A survey of laboratory research groups identified areas of greatest need and led to the development of a three-pronged program: consultation, education, and resource development. Outcomes of this program include bioinformatics consultation services, library-based and graduate level courses, networking of sequence analysis tools, and a biological research Web site. Bioinformatics clients are drawn from diverse departments and include clinical researchers in need of tools that are not readily available outside of basic sciences laboratories. Evaluation and usage statistics indicate that researchers, regardless of departmental affiliation or position, require support to access molecular biology and genetics resources. Centralizing such services in the library is a natural synergy of interests and enhances the provision of traditional library resources. Successful implementation of a library-based bioinformatics program requires both subject-specific and library and information technology expertise.

Computational Biology↗

Does structural and chemical divergence play a role in precluding undesirable protein interactions?

To understand the evolutionary forces establishing, maintaining, breaking, or precluding protein-protein interactions, a comprehensive data set of protein complexes has been analyzed to examine the overlap between protein interfaces and the most conserved or divergent protein surface areas. The most divergent areas tend to be found predominantly away from protein interfaces, although when found at interfaces, they are associated with specific lack of cross-reactivity between close homologues, like in antibody-antigen complexes. Moreover, the amino acid composition of highly variable regions is significantly different from any other protein surfaces. The variable regions present higher structural plasticity as a result of insertions and deletions, and favor charged over hydrophobic residues, a known strategy to minimize aggregation. This suggests that (1) a rapid rate of mutations at these regions might be continuously altering their properties, making difficult the coadaptation, in shape and chemical complementarity, to potential interacting partners; and (2) the existence of some form of selective pressure for variable areas away from interfaces to accumulate charged residues, perhaps as an evolutionary mechanism to increase solubility and minimize undesirable interactions within the crowded cellular environment. Finally, these results are placed into the context of the aberrant oligomerization of sickle-cell anemia hemoglobin and prion proteins.

Amino Acid Sequence↗