PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

The global transcriptional regulatory network for metabolism in Escherichia coli exhibits few dominant functional states.

A principal aim of systems biology is to develop in silico models of whole cells or cellular processes that explain and predict observable cellular phenotypes. Here, we use a model of a genome-scale reconstruction of the integrated metabolic and transcriptional regulatory networks for Escherichia coli, composed of 1,010 gene products, to assess the properties of all functional states computed in 15,580 different growth environments. The set of all functional states of the integrated network exhibits a discernable structure that can be visualized in 3-dimensional space, showing that the transcriptional regulatory network governing metabolism in E. coli responds primarily to the available electron acceptor and the presence of glucose as the carbon source. This result is consistent with recently published experimental data. The observation that a complex network composed of 1,010 genes is organized to achieve few dominant modes demonstrates the utility of the systems approach for consolidating large amounts of genome-scale molecular information about a genome and its regulation to elucidate an organism's preferred environments and functional capabilities.

Bacterial Proteins↗

Fractals in pathology.

Many natural objects, including most objects studied in pathology, have complex structural characteristics and the complexity of their structures, for example the degree of branching of vessels or the irregularity of a tumour boundary, remains at a constant level over a wide range of magnifications. These structures also have patterns that repeat themselves at different magnifications, a property known as scaling self-similarity. This has important implications for measurement of parameters such as length and area, since Euclidean measurements of these may be invalid. The fractal system of geometry overcomes the limitations of the Euclidean geometry for such objects and measurement of the fractal dimension gives an index of their space-filling properties. The fractal dimension may be measured using image analysis systems and the box-counting, divider (perimeter-stepping) and pixel dilation methods have all been described in the published literature. Fractal analysis has found applications in the detection of coding of coding regions in DNA and measurement of the space-filling properties of tumours, blood vessels and neurones. Fractal concepts have also been usefully incorporated into models of biological processes, including epithelial cell growth, blood vessel growth, periodontal disease and viral infections.

Bone Diseases↗

Structure and biological activity of gonadotropin-releasing hormone isoforms isolated from rat and hamster brains.

Rat and hamster brain tissues were used to investigate the possible existence of a follicle stimulating hormone (FSH)-releasing factor with similar characteristics to the lamprey gonadotropin-releasing hormone III (lGnRH-III) form proposed in previous reports. The present studies involved isolation and purification of the molecule by high-performance liquid chromatography (HPLC), identification by radioimmunoassay, sequence analysis by automated Edman degradation, mass spectrometry and examination of biological activity. Hypothalamic extracts from both species contained an HPLC fraction that was immunoreactive to GnRH and coeluted with lGnRH-III and 9-hydroxyproline mGnRH ([Hyp(9)]GnRH). Determination of primary structure from purified total brain material demonstrated that the isolated molecule was [Hyp(9)]GnRH. This is the first report showing the presence of the posttranslationally modified form already known as [Hyp(9)]GnRH by primary sequence analysis. The biological activity of distinct GnRH peptides was also tested in vitro for gonadotropin release using rat pituitary primary cell cultures. The results showed that [Hyp(9)]GnRH stimulated both luteinizing hormone and FSH release, as already reported, whereas lGnRH-III had no action on the secretion of either gonadotropin.

Amino Acid Sequence↗

From 'omes to biology.

Technologies that have emerged from the genome project have dramatically increased our ability to generate data on the way in which organisms respond to their environments, how they execute their programmes of development and growth, and how these are altered in the development of disease states. However, our ability to analyse these large datasets has not kept pace with our ability to generate them and consequently new strategies must be developed to address the issues associated with their analysis. One approach that we have employed quite successfully is to look at data from microarrays (or proteomics or metabolomics experiments) not as independent datasets, but rather as elements of a much larger body of biological information across various scales that must be integrated with, and interpreted within, the context of such ancillary data. Here we outline the general approach and provide three examples from published studies of the way in which we have applied this strategy.

Animals↗

Hidden Markov Models, grammars, and biology: a tutorial.

Biological sequences and structures have been modelled using various machine learning techniques and abstract mathematical concepts. This article surveys methods using Hidden Markov Model and functional grammars for this purpose. We provide a formal introduction to Hidden Markov Model and grammars, stressing on a comprehensive mathematical description of the methods and their natural continuity. The basic algorithms and their application to analyzing biological sequences and modelling structures of bio-molecules like proteins and nucleic acids are discussed. A comparison of the different approaches is discussed, and possible areas of work and problems are highlighted. Related databases and softwares, available on the internet, are also mentioned.

Algorithms↗

Human and mouse alpha-synuclein genes: comparative genomic sequence analysis and identification of a novel gene regulatory element.

The human alpha-synuclein gene (SNCA) encodes a presynaptic nerve terminal protein that was originally identified as a precursor of the non-beta-amyloid component of Alzheimer's disease plaques. More recently, mutations in SNCA have been identified in some cases of familial Parkinson's disease, presenting numerous new areas of investigation for this important disease. Molecular studies would benefit from detailed information about the long-range sequence context of SNCA. To that end, we have established the complete genomic sequence of the chromosomal regions containing the human and mouse alpha-synuclein genes, with the objective of using the resulting sequence information to identify conserved regions of biological importance through comparative sequence analysis. These efforts have yielded approximately 146 and approximately 119 kb of high-accuracy human and mouse genomic sequence, respectively, revealing the precise genetic architecture of the alpha-synuclein gene in both species. A simple repeat element upstream of SNCA/Snca has been identified and shown to be necessary for normal expression in transient transfection assays using a luciferase reporter construct. Together, these studies provide valuable data that should facilitate more detailed analysis of this medically important gene.

Animals↗

White variants of Trichophyton violaceum isolated in Ethiopia.

Certain dermatophytes are geographically restricted and endemic in particular parts of the world, while other species may have a sporadic but worldwide distribution. Trichophyton violaceum is one of the most common dermatophytes causing tinea capitis, and is the predominant cause of tinea in Africa, South America and the Indian subcontinent. Among 1187 dermatophyte isolates collected from Ethiopian patients with various types of tinea, 32 isolates had uncharacteristic phenotypic features. Based on conventional methods complemented by sequence analysis of the rDNA ITS2 region, these isolates were identified as white variants of T. violaceum. This is the first time that white isolates of T. violaceum have been identified in Ethiopia.

Adolescent↗

[Isolation, identification and over- siderophores production of Pseudomonas fluorescens sp-f].

Strain sp-f was isolated, a siderophores over producing bacterium, using an improved universal Chrome Azurol S(CAS)-agar plate method from Donghu Lake. The result of the CAS solution siderophores quantitative determination showed the lowest As/Ar (OD680) ratio could be as low as 0.09 with Su (Siderophore Unit) of 90%. Some more experiments were made to make out the pertinence between its growth and siderophores production, indicating that its siderophores quantity reached maximum amount during the prophase of logarithmic growth. After then, siderophores concentration stopped accumulating and turned to be stable at stationary phase. Based on the characteristics of morphology, cultivation, physiology, (G + C) mol % content, 16S rDNA sequence and BIOLOG Station system analysis, it was identified as Pseudomonas fluorescens sp-f strain. RP-HPLC analysis showed there exist at least 3 kinds of catecholate siderophores, including fluorescent and non-fluorescent pyoverdins. But only fluorescent pyoverdin's excretion was completely repressed by the 200 micromol/L Fe2+ in the medium. And the non-pyoverdin siderophores excretion was induced at the same time, contrarily.

Chromatography, High Pressure Liquid↗

Genome sequence completed of Alcanivorax borkumensis, a hydrocarbon-degrading bacterium that plays a global role in oil removal from marine systems.

In this paper, we provide background to the genome sequencing project of Alcanivorax borkumensis, which is a marine bacterium that uses exclusively petroleum oil hydrocarbons as sources of carbon and energy (therefore designated "hydrocarbonoclastic"). It is found in low numbers in all oceans of the world and in high numbers in oil-contaminated waters. Its ubiquity and unusual physiology suggest it is globally important in the removal of hydrocarbons from polluted marine systems. A functional genomics analysis of Alcanivorax borkumensis strain SK2 was recently initiated, and its genome sequence has just been completed. Annotation of the genome, metabolome modelling, and functional genomics, will soon reveal important insights into the genomic basis of the properties and physiology of this fascinating and globally important bacterium.

Biodegradation, Environmental↗

Comparison of Pax1/9 locus reveals 500-Myr-old syntenic block and evolutionary conserved noncoding regions.

Identification of conserved genomic regions within and between different genomes is crucial when studying genome evolution. Here, we described regions of strong synteny conservation between vertebrate deuterostomes (tetrapods and teleosts) and invertebrate deuterostomes (amphioxus and sea urchin). The shared gene contents across phylogenetically distant species demonstrate that the conservation of the regions stemmed from an ancestral segment instead of a series of independent convergent events. Comparison of the syntenic regions allows us to postulate the primitive gene organization in the last common ancestor of deuterostomes and the evolutionary events that occurred to the 3 distinct lineages of sea urchin, amphioxus, and vertebrates after their separation. In addition, alignment of the syntenic regions led to the identification of 8 noncoding evolutionarily conserved regions shared between amphioxus and vertebrates. To our knowledge, this is the first report of conserved noncoding sequences shared by vertebrates and nonvertebrates. These noncoding sequences have high possibility of being elements that regulate neighboring genes. They are likely to be a factor in the maintenance of conserved synteny over long phylogenetic distance in different deuterostome lineages.

Amino Acid Sequence↗

[Alignment-free biomolecular sequence comparison method].

Biosequence analysis is the primary research field of bioinformatics. In this field, useful information can be extracted by comparison analysis methods. Among them, sequence alignment is the most common comparison method. However the sequence comparison by alignment, which assumes conservation of contiguity between homologous segments, is at odds with genetic recombination. Especially for the multisequence alignment, there exists the difficulty in the complexity of calculation. Therefore, alignment-free sequence comparison methods are required. In this paper, two main categories of alignment-free sequence comparison methods are reviewed. The first one is based on the word (oligomer) frequency and its distribution. The sequences are compared using the distances defined in a Cartesian space by the frequency vectors. In the second category, sequences are compared using Kolmogorov complexity and chaos theory.

Algorithms↗

ALGINATE LYASE: review of major sources and enzyme characteristics, structure-function analysis, biological roles, and applications.

Alginate lyases, characterized as either mannuronate (EC 4.2.2.3) or guluronate lyases (EC 4.2.2.11), catalyze the degradation of alginate, a complex copolymer of alpha-L-guluronate and its C5 epimer beta-D-mannuronate. Lyases have been isolated from a wide range of organisms, including algae, marine invertebrates, and marine and terrestrial microorganisms. This review catalogs the major characteristics of these lyases, the methods for analyzing these enzymes, as well as their biological roles. Analysis of primary sequence data identifies some markedly conserved motifs that should help elucidate functional domains. Information about the three-dimensional structure of a mannuronate lyase from Sphingomonas sp., combined with various mutagenesis studies, has identified residues that are important for catalytic activity in several lyases. Characterization of alginate lyases will enhance and expand the use of these enzymes to engineer novel alginate polymers for applications in various industrial, agricultural, and medical fields. In this review, we explore both past and present applications of this important enzyme and discuss its future prospects.

Alginates↗

Rubrivivax gelatinosus acsF (previously orf358) codes for a conserved, putative binuclear-iron-cluster-containing protein involved in aerobic oxidative cyclization of Mg-protoporphyrin IX monomethylester.

This study describes the characterization of orf358, an open reading frame of previously unidentified function, in the purple bacterium Rubrivivax gelatinosus. A strain in which orf358 was disrupted exhibited a phenotype similar to the wild type under photosynthesis or low-aeration respiratory growth conditions. In contrast, under highly aerated respiratory growth conditions, the wild type still produced bacteriochlorophyll a (Bchl a), while the disrupted strain accumulated a compound that had the same absorption and fluorescence emission spectra as Mg-protoporphyrin but was less polar, suggesting that it was Mg-protoporphyrin monomethylester (MgPMe). These data indicated a blockage in Bchl a synthesis at the oxidative cyclization stage and implied the coexistence of two different mechanisms for MgPMe cyclization in R. gelatinosus, an anaerobic mechanism active under photosynthesis or low oxygenation and an aerobic mechanism active under high-oxygenation growth conditions. Based on these results as well as on sequence analysis indicating the presence of conserved putative binuclear-iron-cluster binding motifs, the designation of orf358 as acsF (for aerobic cyclization system Fe-containing subunit) is proposed. Several homologs of AcsF were found in a wide range of photosynthetic organisms, including Chlamydonomas reinhardtii Crd1 and Pharbitis nil PNZIP, suggesting that this aerobic oxidative cyclization mechanism is conserved from bacteria to plants.

Aerobiosis↗

Whole metagenome sequencing: not deep enough for complete microbial function recovery.

BACKGROUND: Whole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data was annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. RESULTS: Three findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in, e.g., case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this study's sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. We further showed that despite biomes differences in functional diversity and host contamination levels (e.g., soil, fecal), incomplete functional recovery at commonly used sequencing depths was consistently observed. CONCLUSIONS: Together, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights. Video Abstract.

Humans↗

Defining a large set of full-length clones from a Xenopus tropicalis EST project.

Amphibian embryos from the genus Xenopus are among the best species for understanding early vertebrate development and for studying basic cell biological processes. Xenopus, and in particular the diploid Xenopus tropicalis, is also ideal for functional genomics. Understanding the behavior of genes in this accessible model system will have a significant and beneficial impact on the understanding of similar genes in other vertebrate systems. Here we describe the analysis of 219,270 X. tropicalis expressed sequence tags (ESTs) from four early developmental stages. From these, we have deduced a set of unique expressed sequences comprising approximately 20,000 clusters and 16,000 singletons. Furthermore, we developed a computational method to identify clones that contain the complete coding sequence and describe the creation for the first time of a set of approximately 7000 such clones, the full-length (FL) clone set. The entire EST set is cloned in a eukaryotic expression vector and is flanked by bacteriophage promoters for in vitro transcription, allowing functional experiments to be carried out without further subcloning. We have created a publicly available database containing the FL clone set and related clustering data (http://www.gurdon.cam.ac.uk/informatics/Xenopus.html) and we make the FL clone set publicly available as a resource to accelerate the process of gene discovery and function in this model organism. The creation of the unique set of expressed sequences and the FL clone set pave the way toward a large-scale systematic analysis of gene sequence, gene expression, and gene function in this vertebrate species.

Animals↗

GenePalette: a universal software tool for genome sequence visualization and analysis.

To make effective use of the growing host of complete genome sequences, biologists must have easy-to-use software tools that allow them to visualize, analyze, and modify genome data in an interactive and generalized manner. In an effort to bridge the gap between genome and researcher, we have created GenePalette (www.genepalette.org), a desktop application that can access any genome sequence and display the positions of various features [e.g., transcription factor binding sites (TFBSs)] relative to the introns and exons of annotated genes. Written in Java, GenePalette can run on all Java-supporting operating systems (Mac, PC, Unix, Linux). Annotated sequence encompassing the majority of public genome data is rapidly retrieved from GenBank or Ensembl. The software provides intuitive access to the selected genomic region through three interface components: a colorful graphical display showing a schematic of genes and features; an annotated sequence view in which features and genes are highlighted directly on the sequence; and the selectable raw sequence. The three interface components are fully integrated and presented on one page, permitting the user to move easily between representations at different levels of resolution, ranging from kilobases to individual nucleotides. GenePalette is a particularly powerful platform for analyzing the organization of cis-regulatory elements and designing wet-lab experiments to investigate them.

Computational Biology↗

Drosophila Knickkopf and Retroactive are needed for epithelial tube growth and cuticle differentiation through their specific requirement for chitin filament organization.

Precise epithelial tube diameters rely on coordinated cell shape changes and apical membrane enlargement during tube growth. Uniform tube expansion in the developing Drosophila trachea requires the assembly of a transient intraluminal chitin matrix, where chitin forms a broad cable that expands in accordance with lumen diameter growth. Like the chitinous procuticle, the tracheal luminal chitin cable displays a filamentous structure that presumably is important for matrix function. Here, we show that knickkopf (knk) and retroactive (rtv) are two new tube expansion mutants that fail to form filamentous chitin structures, both in the tracheal and cuticular chitin matrices. Mutations in knk and rtv are known to disrupt the embryonic cuticle, and our combined genetic analysis and chemical chitin inhibition experiments support the argument that Knk and Rtv specifically assist in chitin function. We show that Knk is an apical GPI-linked protein that acts at the plasma membrane. Subcellular mislocalization of Knk in previously identified tube expansion mutants that disrupt septate junction (SJ) proteins, further suggest that SJs promote chitinous matrix organization and uniform tube expansion by supporting polarized epithelial protein localization. We propose a model in which Knk and the predicted chitin-binding protein Rtv form membrane complexes essential for epithelial tubulogenesis and cuticle formation through their specific role in directing chitin filament assembly.

Animals↗

KOBAS server: a web-based platform for automated annotation and pathway identification.

There is an increasing need to automatically annotate a set of genes or proteins (from genome sequencing, DNA microarray analysis or protein 2D gel experiments) using controlled vocabularies and identify the pathways involved, especially the statistically enriched pathways. We have previously demonstrated the KEGG Orthology (KO) as an effective alternative controlled vocabulary and developed a standalone KO-Based Annotation System (KOBAS). Here we report a KOBAS server with a friendly web-based user interface and enhanced functionalities. The server can support input by nucleotide or amino acid sequences or by sequence identifiers in popular databases and can annotate the input with KO terms and KEGG pathways by BLAST sequence similarity or directly ID mapping to genes with known annotations. The server can then identify both frequent and statistically enriched pathways, offering the choices of four statistical tests and the option of multiple testing correction. The server also has a 'User Space' in which frequent users may store and manage their data and results online. We demonstrate the usability of the server by finding statistically enriched pathways in a set of upregulated genes in Alzheimer's Disease (AD) hippocampal cornu ammonis 1 (CA1). KOBAS server can be accessed at http://kobas.cbi.pku.edu.cn.

Alzheimer Disease↗