PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Altered proglucagon processing in an alpha-cell line derived from prohormone convertase 2 null mouse islets.

The endoproteolytic processing of proproteins in the secretory pathway depends on the expression of selected members of a family of subtilisin-like endoproteases known as the prohormone convertases (PCs). The main PC family members expressed in mammalian neuroendocrine cells are PC2 and PC1/3. The differential processing of proglucagon in pancreatic alpha-cells and intestinal L cells leads to production of distinct hormonal products with opposing physiological effects from the same precursor. Here we describe the establishment and characterization of a novel alpha-cell line (alphaTC-DeltaPC2) derived from PC2 homozygous null animals. The alphaTC-DeltaPC2 cells are shown to be similar to the well characterized alphaTC1-6 cell line in both morphology and overall gene expression. However, the absence of PC2 activity in alphaTC-DeltaPC2 leads to a complete block in the production of mature glucagon. Surprisingly, alphaTC-DeltaPC2 cells are able to efficiently cleave the interdomain site in proglucagon (KR 70-71). Further analysis reveals that alphaTC-DeltaPC2 cells, unlike alphaTC1-6 cells, express low levels of PC1/3 that lead to the generation of glicentin as well as low amounts of oxyntomodulin, GLP-1, truncated GLP-1, and N-terminally extended GLP-2. We conclude that alphaTC-DeltaPC2 cells provide additional evidence for PC2 as the major convertase in alpha-cells leading to mature glucagon production and provide a robust model for further analysis of the mechanisms of proprotein processing by the prohormone convertases.

Animals↗

WholePathwayScope: a comprehensive pathway-based analysis tool for high-throughput data.

BACKGROUND: Analysis of High Throughput (HTP) Data such as microarray and proteomics data has provided a powerful methodology to study patterns of gene regulation at genome scale. A major unresolved problem in the post-genomic era is to assemble the large amounts of data generated into a meaningful biological context. We have developed a comprehensive software tool, WholePathwayScope (WPS), for deriving biological insights from analysis of HTP data. RESULT: WPS extracts gene lists with shared biological themes through color cue templates. WPS statistically evaluates global functional category enrichment of gene lists and pathway-level pattern enrichment of data. WPS incorporates well-known biological pathways from KEGG (Kyoto Encyclopedia of Genes and Genomes) and Biocarta, GO (Gene Ontology) terms as well as user-defined pathways or relevant gene clusters or groups, and explores gene-term relationships within the derived gene-term association networks (GTANs). WPS simultaneously compares multiple datasets within biological contexts either as pathways or as association networks. WPS also integrates Genetic Association Database and Partial MedGene Database for disease-association information. We have used this program to analyze and compare microarray and proteomics datasets derived from a variety of biological systems. Application examples demonstrated the capacity of WPS to significantly facilitate the analysis of HTP data for integrative discovery. CONCLUSION: This tool represents a pathway-based platform for discovery integration to maximize analysis power. The tool is freely available at http://www.abcc.ncifcrf.gov/wps/wps_index.php.

Computer Graphics↗

Checking cell size in budding yeast: a systems biology approach.

The regulation of cell cycle progression via the attainment of a critical cell size is a conserved feature from simpler unicellular organisms to mammalian cells that is obtaining much attention recently. Genome wide analysis of Saccharomyces cerevisiae deletion strains, genetic epistasis, DNA microarray analysis have recently revealed an increasingly complex network of cell size modulation mechanisms. A systems biology-based approach, that is needed to structure the underlying complexity of cell cycle regulatory mechanisms, is described.

Cell Cycle↗

Text-mining approaches in molecular biology and biomedicine.

Biomedical articles provide functional descriptions of bioentities such as chemical compounds and proteins. To extract relevant information using automatic techniques, text-mining and information-extraction approaches have been developed. These technologies have a key role in integrating biomedical information through analysis of scientific literature. In this article, important applications such as the identification of biologically relevant entities in free text and the construction of literature-based networks of protein-protein interactions will be introduced. Also, the use of text mining to aid the interpretation of microarray data and the analysis of pathology reports will be discussed. Finally, we will consider the recent evolution of this field and the efforts for community-based evaluations.

Biomedical Research↗

Biological data warehousing system for identifying transcriptional regulatory sites from gene expressions of microarray data.

Identification of transcriptional regulatory sites plays an important role in the investigation of gene regulation. For this propose, we designed and implemented a data warehouse to integrate multiple heterogeneous biological data sources with data types such as text-file, XML, image, MySQL database model, and Oracle database model. The utility of the biological data warehouse in predicting transcriptional regulatory sites of coregulated genes was explored using a synexpression group derived from a microarray study. Both of the binding sites of known transcription factors and predicted over-represented (OR) oligonucleotides were demonstrated for the gene group. The potential biological roles of both known nucleotides and one OR nucleotide were demonstrated using bioassays. Therefore, the results from the wet-lab experiments reinforce the power and utility of the data warehouse as an approach to the genome-wide search for important transcription regulatory elements that are the key to many complex biological systems.

Algorithms↗

The UK National External Quality Assessment Scheme (UK NEQAS) for molecular genetic testing in haemophilia.

Molecular genetic analysis of families with haemophilia and other inherited bleeding disorders is now a common laboratory investigation. In contrast to phenotypic testing in which strict quality control is adhered to, in haemophilia molecular genetic testing there has been a lack of any external quality assurance schemes. In 1998 the UK National External Quality Assessment Scheme (UK NEQAS) established a pilot quality assurance scheme for molecular genetic testing in haemophilia. Results from three initial surveys highlighted problems with the quality of samples when used to screen for the intron 22 inversion within the F8 gene. The scheme was re-launched in 2003, and since that time there have been five exercises involving whole blood or immortalised cell line DNA. The results together with an overall summary of the exercise are subsequently returned to participants. Exercises to date have focused exclusively on haemophilia A and QA, material has included screening for the intron 1 and intron 22 inversions as well as sequence analysis. A paper exercise circulated in 2003 highlighted problems with the format of reports and, following feedback to participants, only a single error has been made in the subsequent four exercises. Participating laboratories now receive QA material every six months. Immortalised cell line material was introduced in 2005 and was shown to perform well. This will allow expansion of the scheme and a reduction in the dependence on blood donation.

Chromosome Inversion↗

rVISTA 2.0: evolutionary analysis of transcription factor binding sites.

Identifying and characterizing the transcription factor binding site (TFBS) patterns of cis-regulatory elements represents a challenge, but holds promise to reveal the regulatory language the genome uses to dictate transcriptional dynamics. Several studies have demonstrated that regulatory modules are under positive selection and, therefore, are often conserved between related species. Using this evolutionary principle, we have created a comparative tool, rVISTA, for analyzing the regulatory potential of noncoding sequences. Our ability to experimentally identify functional noncoding sequences is extremely limited, therefore, rVISTA attempts to fill this great gap in genomic analysis by offering a powerful approach for eliminating TFBSs least likely to be biologically relevant. The rVISTA tool combines TFBS predictions, sequence comparisons and cluster analysis to identify noncoding DNA regions that are evolutionarily conserved and present in a specific configuration within genomic sequences. Here, we present the newly developed version 2.0 of the rVISTA tool, which can process alignments generated by both the zPicture and blastz alignment programs or use pre-computed pairwise alignments of several vertebrate genomes available from the ECR Browser and GALA database. The rVISTA web server is closely interconnected with the TRANSFAC database, allowing users to either search for matrices present in the TRANSFAC library collection or search for user-defined consensus sequences. The rVISTA tool is publicly available at http://rvista.dcode.org/.

Algorithms↗

Challenges and prospects in the analysis of large-scale gene expression data.

Large heterogeneous expression data comprising a variety of cellular conditions hold the promise of a global view of transcriptional regulation. While standard analysis methods have been successfully applied to smaller data sets, large-scale data pose specific challenges that have prompted the development of new and more sophisticated approaches. This paper focuses on one such approach (the Signature Algorithm) and discusses the central challenges in the analysis of large data sets, and how they might be overcome. Biological questions that have been addressed using the Signature Algorithm are highlighted and a summary of other important methods from the literature is provided.

Algorithms↗

MitoRes: a resource of nuclear-encoded mitochondrial genes and their products in Metazoa.

BACKGROUND: Mitochondria are sub-cellular organelles that have a central role in energy production and in other metabolic pathways of all eukaryotic respiring cells. In the last few years, with more and more genomes being sequenced, a huge amount of data has been generated providing an unprecedented opportunity to use the comparative analysis approach in studies of evolution and functional genomics with the aim of shedding light on molecular mechanisms regulating mitochondrial biogenesis and metabolism. In this context, the problem of the optimal extraction of representative datasets of genomic and proteomic data assumes a crucial importance. Specialised resources for nuclear-encoded mitochondria-related proteins already exist; however, no mitochondrial database is currently available with the same features of MitoRes, which is an update of the MitoNuc database extensively modified in its structure, data sources and graphical interface. It contains data on nuclear-encoded mitochondria-related products for any metazoan species for which this type of data is available and also provides comprehensive sequence datasets (gene, transcript and protein) as well as useful tools for their extraction and export. DESCRIPTION: MitoRes http://www2.ba.itb.cnr.it/MitoRes/ consolidates information from publicly external sources and automatically annotates them into a relational database. Additionally, it also clusters proteins on the basis of their sequence similarity and interconnects them with genomic data. The search engine and sequence management tools allow the query/retrieval of the database content and the extraction and export of sequences (gene, transcript, protein) and related sub-sequences (intron, exon, UTR, CDS, signal peptide and gene flanking regions) ready to be used for in silico analysis. CONCLUSION: The tool we describe here has been developed to support lab scientists and bioinformaticians alike in the characterization of molecular features and evolution of mitochondrial targeting sequences. The way it provides for the retrieval and extraction of sequences allows the user to overcome the obstacles encountered in the integrative use of different bioinformatic resources and the completeness of the sequence collection allows intra- and interspecies comparison at different biological levels (gene, transcript and protein).

Animals↗

Microarray analysis of trophoblast differentiation: gene expression reprogramming in key gene function categories.

Placental development results from a highly dynamic differentiation program. We used DNA microarray analysis to characterize the process by which human cytotrophoblast cells differentiate into syncytiotrophoblast cells in a purified cell culture system. Of 6,918 genes analyzed, 141 genes were induced and 256 were downregulated by more than 2-fold. Dynamically regulated genes were divided by the K-means algorithm into 9 kinetic pattern groups, then by biologic classification into 6 overall functional categories: cell and tissue structural dynamics, cell cycle and apoptosis, intercellular communication, metabolism, regulation of gene expression, and expressed sequence tag (EST) and function unknown. Gene expression changes within key functional categories were tightly coupled to morphological changes. In several key gene function categories, such as cell and tissue structure, many gene members of the category were strongly activated while others were strongly repressed. These findings suggest that differentiation is augmented by "categorical reprogramming" in which the function of induced genes is enhanced by preventing the further synthesis of categorically related gene products.

Cell Differentiation↗

Augur--a computational pipeline for whole genome microbial surface protein prediction and classification.

UNLABELLED: The analysis of protein function is a challenge and a major bottleneck towards well-annotated and analysed microbial genomes. In particular, bacterial surface proteins present an opportunity for pharmacological intervention and vaccine development. We present Augur, an automatic prediction pipeline that integrates major surface prediction algorithms and enables comparative analysis, classification and visualization for gram-positive bacteria on a genomic scale. AVAILABILITY: http://bioinfo.mikrobio.med.uni-giessen.de/augur

Algorithms↗

Expression array technology in the diagnosis and treatment of breast cancer.

The most common group of cancers among American women involves malignancies of the breast. Breast cancer is a complex disease, involving several different types of tissues and specific cells with various functions, that is categorized into many distinct subtypes. Microarray analysis has recently revealed that different biological subtypes of breast cancer are accompanied by differences in their specific gene expression profile. Because breast tissue (and breast cancer) is heterogeneous, microarray analysis may provide clinicians with a better understanding of how to treat each specific case. Thus, microarray analysis may translate basic research data into more confident diagnoses, specifically designed treatment regimens geared to each patient's needs, and better clinical prognoses.

Breast Neoplasms↗

Delayed development and lifespan extension as features of metabolic lifestyle alteration in C. elegans under dietary restriction.

Studies of the model organism Caenorhabditis elegans have almost exclusively utilized growth on a bacterial diet. Such culturing presents a challenge to automation of experimentation and introduces bacterial metabolism as a secondary concern in drug and environmental toxicology studies. Axenic cultivation of C. elegans can avoid these problems, yet past work suggests that axenic growth is unhealthy for C. elegans. Here we employ a chemically defined liquid medium to culture C. elegans and find development slows, fecundity declines, lifespan increases, lipid and protein stores decrease, and gene expression changes relative to that on a bacterial diet. These changes do not appear to be random pathologies associated with malnutrition, as there are no developmental delays associated with starvation, such as L1 or dauer diapause. Additionally, development and reproductive period are fixed percentages of lifespan regardless of diet, suggesting that these alterations are adaptive. We propose that C. elegans can exist as a healthy animal with at least two distinct adult life histories. One life history maximizes the intrinsic rate of population increase, the other maximizes the efficiency of exploitation of the carrying capacity of the environment. Microarray analysis reveals increased transcript levels of daf-16 and downstream targets and past experiments demonstrate that DAF-16 (FOXO) acting on downstream targets can influence all of the phenotypes we see altered in maintenance medium. Thus, life history alteration in response to diet may be modulated by DAF-16. Our observations introduce a powerful system for automation of experimentation on healthy C. elegans and for systematic analysis of the profound impact of diet on animal physiology.

Animals↗

The human growth hormone locus: nucleotide sequence, biology, and evolution.

The human chromosomal growth hormone locus contained on cloned DNA and spanning approximately 66,500 bp was sequenced in its entirety to provide a framework for the analysis of its biology and evolution. This locus evolved by a series of duplications and contains in its present form five genes which display a remarkably high degree of sequence identity (approximately 95%) in all their domains. The DNA sequence of the locus reveals the presence of 48 middle repetitive sequence elements of the Alu type and one member of the KpnI family, all located in the intergenic regions. The expression of each gene was examined by screening pituitary and placental cDNA libraries by using gene-specific oligonucleotides. According to this analysis, the hGH-N gene is transcribed exclusively in the pituitary, whereas the other four genes (hCS-L, hCS-A, hGH-V, hCS-B) are expressed only in placental tissue, at levels characteristic for each gene. Particular DNA sequences found upstream of the individual promoter regions might account for the observed tissue specificity and different transcriptional activity of the genes. The hCS-L gene carries a G to A transition in a sequence used by the other four genes as an intronic 5' splice donor site. This mutation results in a different splicing pattern and, hence, in a novel sequence of the hCS-L gene mRNA and the deduced polypeptide.

Amino Acid Sequence↗

Purification and characterization of human recombinant interleukin-1 beta.

A human interleukin-1 (IL-1) beta cDNA was cloned, and the region coding for the mature protein was expressed in Escherichia coli. The 17-kDa biologically active product was purified in 40% yield to apparent homogeneity, without chaotropes, from the soluble fraction of sonicated cell lysates. The recombinant IL-1 beta was characterized by amino acid analysis, NH2- and COOH-terminal sequence analysis, sodium dodecyl sulfate-polyacrylamide gel electrophoresis, spectroscopy, and biological assay. Specific biological activity was 4.6 X 10(8) units/mg in a co-mitogenic IL-2 induction assay using cultured EL-4 T-lymphocytes. The molar extinction coefficient was determined to be 10,300 cm-1 M-1 at 280 nm. NH2-terminal sequence analysis revealed that 70% of the product begins with the Ala corresponding to the NH2 terminus of the natural protein, while 30% begins with the following Pro. No initiator Met was observed. Both of the sulfhydryl groups are reactive to Ellman's reagent and to iodoacetamide under nonreducing conditions, indicating that the Cys residues do not form disulfide bonds. S-Carboxamidomethyl-Cys-rIL-1 beta retained biological activity in the IL-2 induction assay. Circular dichroism suggested an extensive beta sheet structure for rIL-1 beta.

Amino Acid Sequence↗

Unsupervised multiscale clustering of single-cell transcriptomes to identify hierarchical structures of cell subtypes.

BACKGROUND: Cell clustering is an essential step in uncovering cellular architectures in single-cell RNA sequencing (scRNA-seq) data. However, the existing cell clustering approaches are not well designed to dissect complex structures of cellular landscapes at a finer resolution. RESULTS: Here, we develop a multiscale clustering (MSC) approach to construct a sparse cell-cell correlation network for unsupervised identification of de novo cell types and subtypes across multiple resolutions. Based upon simulated silver- and gold-standard data as well as real scRNA-seq data in diseases, MSC demonstrates significantly improved performance compared to established benchmark methods and reveals a biologically meaningful cell hierarchy to facilitate the discovery of novel disease-associated cell subtypes and mechanisms. CONCLUSIONS: We present MSC as a new single-cell multiscale clustering framework as a powerful tool for advancing discoveries in disease-associated cell populations using single-cell sequencing data.

Single-Cell Analysis↗

Legionella pneumophila - a human pathogen that co-evolved with fresh water protozoa.

The bacterial pathogen Legionella pneumophila is found ubiquitously in fresh water environments where it replicates within protozoan hosts. When inhaled by humans it can replicate within alveolar macrophages and cause a severe pneumonia, Legionnaires disease. Yet much needs to be learned regarding the mechanisms that allow Legionella to modulate host functions to its advantage and the regulatory network governing its intracellular life cycle. The establishment and publication of the complete genome sequences of three clinical L. pneumophila isolates paved the way for major breakthroughs in understanding the biology of L. pneumophila. Based on sequence analysis many new putative virulence factors have been identified foremost among them eukaryotic-like proteins that may be implicated in many different steps of the Legionella life cycle. This review summarizes what is currently known about regulation of the Legionella life cycle and gives insight in the Legionella-specific features as deduced from genome analysis.

Animals↗

Making mechanistic connections between cell signaling pathways and pathological endpoints.

Cell signaling is a term used to describe a complex interactive system of signals that act to regulate or mediate a cellular response. Therapies that target cell signaling pathways have the potential to effectively reverse molecular deregulation underlying disease. The inherent complexity of cell signaling presents a major challenge to designing such therapies however, because perturbation of pathways has the potential to produce dramatic adverse effects. Pathologists are in the primary position of detecting adverse responses in drug development and are essential members of teams whose goal is to determine the mechanisms underlying tissue responses. The pathologist therefore will be expected to integrate morphologic interpretation with data obtained from several laboratory-based methods and data derived from novel technologies. Approaches being used include several in silico tools that provide access to public databases and signal pathway visualization that can serve to focus on key mechanistic hypotheses. The main objective of this article is to discuss a basic mechanistic approach and methods that can be used to associate modulation of cell signaling pathways with pathologic endpoints. The approach suggested begins with diagnostic pathology and uses global gene expression analysis in conjunction with transcription factor profiling and confirmatory protein technologies, to elucidate pathways relevant to the biological mechanism. Another important objective is to highlight the use of in silico technologies to prioritize laboratory efforts and focus these efforts on key hypotheses.

Animals↗