PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

From patterns to pathways: gene expression data analysis comes of age.

Many different biological questions are routinely studied using transcriptional profiling on microarrays. A wide range of approaches are available for gleaning insights from the data obtained from such experiments. The appropriate choice of data-analysis technique depends both on the data and on the goals of the experiment. This review summarizes some of the common themes in microarray data analysis, including detection of differential expression, clustering, and predicting sample characteristics. Several approaches to each problem, and their relative merits, are discussed and key areas for additional research highlighted.

Cluster Analysis↗

Systematic gene function prediction from gene expression data by using a fuzzy nearest-cluster method.

BACKGROUND: Quantitative simultaneous monitoring of the expression levels of thousands of genes under various experimental conditions is now possible using microarray experiments. However, there are still gaps toward whole-genome functional annotation of genes using the gene expression data. RESULTS: In this paper, we propose a novel technique called Fuzzy Nearest Clusters for genome-wide functional annotation of unclassified genes. The technique consists of two steps: an initial hierarchical clustering step to detect homogeneous co-expressed gene subgroups or clusters in each possibly heterogeneous functional class; followed by a classification step to predict the functional roles of the unclassified genes based on their corresponding similarities to the detected functional clusters. CONCLUSION: Our experimental results with yeast gene expression data showed that the proposed method can accurately predict the genes' functions, even those with multiple functional roles, and the prediction performance is most independent of the underlying heterogeneity of the complex functional classes, as compared to the other conventional gene function prediction approaches.

Algorithms↗

Design of long oligonucleotide probes for functional gene detection in a microbial community.

MOTIVATION: Analysis of the functions of microorganisms and their dynamics in the environment is essential for understanding microbial ecology. For analysis of highly similar sequences of a functional gene family using microarrays, the previous long oligonucleotide probe design strategies have not been useful in generating probes. RESULTS: We developed a Hierarchical Probe Design (HPD) program that designs both sequence-specific probes and hierarchical cluster-specific probes from sequences of a conserved functional gene based on the clustering tree of the genes, specifically for analyses of functional gene diversity in environmental samples. HPD was tested on datasets for the nirS and pmoA genes. Our results showed that HPD generated more sequence-specific probes than several popular oligonucleotide design programs. With a combination of sequence-specific and cluster-specific probes, HPD generated a probe set covering all the sequences of each test set. AVAILABILITY: http://brcapp.kribb.re.kr/HPD/

Algorithms↗

Microarray analysis identifies Salmonella genes belonging to the low-shear modeled microgravity regulon.

The low-shear environment of optimized rotation suspension culture allows both eukaryotic and prokaryotic cells to assume physiologically relevant phenotypes that have led to significant advances in fundamental investigations of medical and biological importance. This culture environment has also been used to model microgravity for ground-based studies regarding the impact of space flight on eukaryotic and prokaryotic physiology. We have previously demonstrated that low-shear modeled microgravity (LSMMG) under optimized rotation suspension culture is a novel environmental signal that regulates the virulence, stress resistance, and protein expression levels of Salmonella enterica serovar Typhimurium. However, the mechanisms used by the cells of any species, including Salmonella, to sense and respond to LSMMG and identities of the genes involved are unknown. In this study, we used DNA microarrays to elucidate the global transcriptional response of Salmonella to LSMMG. When compared with identical growth conditions under normal gravity (1 x g), LSMMG differentially regulated the expression of 163 genes distributed throughout the chromosome, representing functionally diverse groups including transcriptional regulators, virulence factors, lipopolysaccharide biosynthetic enzymes, iron-utilization enzymes, and proteins of unknown function. Many of the LSMMG-regulated genes were organized in clusters or operons. The microarray results were further validated by RT-PCR and phenotypic analyses, and they indicate that the ferric uptake regulator is involved in the LSMMG response. The results provide important insight about the Salmonella LSMMG response and could provide clues for the functioning of known Salmonella virulence systems or the identification of uncharacterized bacterial virulence strategies.

Biomechanical Phenomena↗

A Cis-Regulatory Duplication in a Hox Hotspot Implicated in Mimetic Convergence in the Bumble Bee Bombus flavifrons.

Several species of North American bumble bees spanning the Pacific Coastal and Rocky Mountain regions converge onto distinct mimetic abdominal colour forms for each region by switching abdominal coloration from black to red. Previous genome-wide association studies (GWAS) of red and black transitions in two mimics (Bombus melanopygus and Bombus vancouverensis) revealed that black forms were generated by independently deleting a portion of the same cis-regulatory region near the Hox gene Abdominal-B (Abd-B). Here, we test the genetic basis of these mimetic colour forms in a third co-mimic, Bombus flavifrons, that has continuous variation in red and black that is shifted posteriorly one segment compared to its co-mimics. Using genome-wide association of red and black forms, we identified a structural variant <&#x2009;50&#x2009;bp away from the deletions in B. melanopygus and B. vancouverensis that was strongly associated with the colour phenotype. Sequencing across mimicry zones and closely related taxa revealed that all red forms of B. flavifrons and monomorphic red close relative Bombus centralis have a 319&#x2009;bp tandem duplication at this locus that has extensive modification to the duplicated copy. Black forms of B. flavifrons from the Cascades also have this duplication but without the modifications, while black forms in the western Rockies mostly lack this duplication, similar to ancestral black forms. This suggests independent mechanisms may regulate the black phenotypes in different populations and that ancestral sorting of variation and/or adaptive introgression generated these phenotypes. This study strengthens support for this Abd-B cis-regulatory region being a hotspot for regulating abdominal coloration in bumble bees, and features the role of regulatory region duplication in creating novel phenotypes.

Animals↗

Clustering biological annotations and gene expression data to identify putatively co-regulated biological processes.

MOTIVATION: Functional profiling is a key step of microarray gene expression data analysis. Identifying co-regulated biological processes could help for better understanding of underlying biological interactions within the studied biological frame. RESULTS: We present herein an original approach designed to search for putatively co-regulated biological processes sharing a significant number of co-expressed genes. An R language implementation named "FunCluster" was built and tested on two gene expression data sets. A discriminatory functional analysis of the first data set, related to experiments performed on separated adipocytes and stroma vascular fraction cells of human white adipose tissue, highlighted the prevalent role of nonadipose cells in the synthesis of inflammatory and immunity molecules in human adiposity. On the second data set, resulting from a model investigating insulin coordinated regulation of gene expression in human skeletal muscle, FunCluster analysis spotlighted novel functional classes of putatively co-regulated biological processes related to protein metabolism and the regulation of muscular contraction. AVAILABILITY: Supplementary information about the FunCluster tool is available on-line at http://corneliu.henegar.info/FunCluster.htm.

Algorithms↗

Coexpressionfinder: a new algorithm for finding groups of coexpressed genes.

RESULTS: A new algorithm is developed which is intended to find groups of genes whose expression values change in a concordant manner in a series of experiments with DNA arrays. This algorithm is named as CoexpressionFinder. It can find more complete and internally coordinated groups of gene expression vectors than hierarchical clustering. Also, it finds more genes having coordinated expression. The algorithm's design allows parallel execution. AVAILABILITY: The algorithm is implemented as a Java application which is freely available at: http://www.bioinformatics.ru/cf/index.jsp and http://bioinformatics.ru/cf/index.jsp.

Algorithms↗

Comparative transcriptional and functional profiling of clear cell and papillary renal cell carcinoma.

Renal cell carcinoma (RCC) is known to effectively prevent immune recognition. However, little is known about the mechanisms that underlie this phenomenon. Thus, the identification of immunogenic molecules associated with RCC and the elucidation of the corresponding signaling pathways are crucial to the development of effective treatments. We performed transcriptional and functional profiling with cDNA microarrays (1070 cDNA probes) on a total of 17 RCCs, 11 clear cell and 6 papillary, and on corresponding normal tissue. Samples were clustered based on their expression profiles. We found a total of 45 genes to be regulated equally by both tumor types compared to the normal tissue. A set of 13 differentially expressed genes was identified between the examined tumor subtypes. Functional analysis was performed for both gene sets and showed a significant enrichment of cell surface genes regulated in both tumor subtypes. Within these we found five surface marker genes to be upregulated (TNFRSF10B, CD70, TNFR1, PDGFRB, and BAFF) which are involved in immune responses via the regulation of lymphocytes and can also induce apoptosis. Their overexpression in both tumor subtypes suggests a possible involvement in the immune escape strategies of RCC. The combination of transcriptional and functional profiling revealed potential target molecules for novel therapy strategies that must be studied in more detail.

Carcinoma, Papillary↗

Purification and sequence analysis of bioactive atrial peptides (atriopeptins).

Mammalian cardiac atria have several biologically active peptides that exert profound effects on sodium excretion, urine volume, and smooth muscle tone. In the present study two such peptides of low molecular weight were purified and separated from each other on the basis of differences in charge, hydrophobicity, and biological profile. The first peptide, designated atriopeptin I, exhibits natriuretic and diuretic activity and selectivity relaxes intestinal smooth muscle but not vascular smooth muscle strips. The second peptide, atriopeptin II, is a potent natriuretic and diuretic that relaxes both intestinal and vascular strips. Sequence analysis of atriopeptin I indicates that it is composed of 21 amino acids, of which serine and glycine residues predominate. The amino terminal sequence of atriopeptin II up to residue 21 is the same as that of atriopeptin I, with the addition of the Phe-Arg extension at the carboxyl terminus. Both peptides appear to be derived from a common high molecular weight precursor (designated atriopeptigen); their biological selectivity and potency may be determined by the site of carboxyl terminal cleavage.

Amino Acid Sequence↗

SBEAMS-Microarray: database software supporting genomic expression analyses for systems biology.

BACKGROUND: The biological information in genomic expression data can be understood, and computationally extracted, in the context of systems of interacting molecules. The automation of this information extraction requires high throughput management and analysis of genomic expression data, and integration of these data with other data types. RESULTS: SBEAMS-Microarray, a module of the open-source Systems Biology Experiment Analysis Management System (SBEAMS), enables MIAME-compliant storage, management, analysis, and integration of high-throughput genomic expression data. It is interoperable with the Cytoscape network integration, visualization, analysis, and modeling software platform. CONCLUSION: SBEAMS-Microarray provides end-to-end support for genomic expression analyses for network-based systems biology research.

Chromosome Mapping↗

Single-molecule spectroscopy for nucleic acid analysis: a new approach for disease detection and genomic analysis.

Recently developed single-molecule spectroscopy (SMS) permits the analysis of fluorescent mixtures one molecule at a time. SMS methods provide the means to make rapid measurements on small, complex samples without the need for separations and target amplification enabling a new class of ultrasensitive nucleic acid assays. Here we give a brief overview of the current state of the art of SMS nucleic acid analysis and discuss ongoing work in our laboratory on two-color single-molecule fluorescence detection of specific nucleic acid sequences. In the future, two-color SMS nucleic acid assays will be used for a variety of applications including: gene expression analysis, disease detection and genomics.

DNA↗

Molecular analysis of deep-sea hydrothermal vent aerobic methanotrophs by targeting genes of 16S rRNA and particulate methane monooxygenase.

Molecular diversity of deep-sea hydrothermal vent aerobic methanotrophs was studied using both 16S ribosomalDNA and pmoA encoding the subunit A of particulate methane monooxygenase (pMOA). Hydrothermal vent plume and chimney samples were collected from back-arc vent at Mid-Okinawa Trough (MOT), Japan, and the Trans-Atlantic Geotraverse (TAG) site along Mid-Atlantic Ridge, respectively. The target genes were amplified by polymerase chain reaction from the bulk DNA using specific primers and cloned. Fifty clones from each clone library were directly sequenced. The 16S rDNA sequences were grouped into 3 operational taxonomic units (OTUs), 2 from MOT and 1 from TAG. Two OTUs (1 MOT and 1 TAG) were located within the branch of type I methanotrophic ?-Proteobacteria. Another MOT OTU formed a unique phylogenetic lineage related to type I methanotrophs. Direct sequencing of 50 clones each from the MOT and TAG samples yielded 17 and 4 operational pmoA units (OPUs), respectively. The phylogenetic tree based on the pMOA amino acid sequences deduced from OPUs formed diverse phylogenetic lineages within the branch of type I methanotrophs, except for the OPU MOT-pmoA-8 related to type X methanotrophs. The deduced pMOA topologies were similar to those of all known pMOA, which may suggest that the pmoA gene is conserved through evolution. Neither the 16S rDNA nor pmoA molecular analysis could detect type II methanotrophs, which suggests the absence of type II methanotrophs in the collected vent samples.

Animals↗

Protein-based analysis of alternative splicing in the human genome.

Understanding the functional significance of alternative splicing and other mechanisms that generate RNA transcript diversity is an important challenge facing modern-day molecular biology. Using homology-based, protein sequence analysis methods, it should be possible to investigate how transcript diversity impacts protein structure and function. To test this, a data mining technique ("DiffHit") was developed to identify and catalog genes producing protein isoforms which exhibit distinct profiles of conserved protein motifs. We found that out of a test set of over 1,300 alternatively spliced genes with solved genomic structure, over 30% exhibited a differential profile of conserved InterPro and/or Blocks protein motifs across distinct isoforms. These results suggest that motif databases such as Blocks and InterPro are potentially useful tools for investigating how alternative transcript structure affects gene function.

Algorithms↗

Toward efficient analysis of >70 kDa proteins with 100% sequence coverage.

For complete characterization of larger proteins, primary structural analysis by mass spectrometry must be made more efficient. A straightforward approach is illustrated here using two proteins of 159 and 199 kDa with five and nine Lys residues, respectively. These proteins were degraded by Lys-C to mixtures of peptides ranging in size from 5 to 48 kDa, whose multiply charged ions (from electrospray ionization) are far more amenable than the intact proteins to direct interrogation in a Fourier-transform mass spectrometer. For the 199 kDa PchF of approximately 60% purity, an unfractionated Lys-C digest gave 106 isotopic distributions from 71 components (most of which were below 6 kDa); 15% sequence coverage was obtained. For the > 90% pure PchE (159 kDa), complete sequence coverage was obtained from six Lys-C peptides of 5, 8, 26, 32, 40 and 48 kDa, with all but the largest of these measured at isotopic resolution on a 4.7 Tesla instrument. Practical strategies for implementing this characterization strategy on a proteomic scale are considered.

Amino Acid Sequence↗

The CAEV tat gene trans-activates the viral LTR and is necessary for efficient viral replication.

Caprine arthritis-encephalitis virus (CAEV) is a lentivirus which is closely related by nucleotide sequence and biological properties to visna virus. Sequence analysis of the CAEV genome revealed the presence of a small open reading frame (ORF) which shares amino acid identity with the visna virus tat gene. Using an infectious molecular clone of CAEV the role of the tat ORF in viral replication was examined. Mutations were made in the tat ORF that introduced two in frame stop codons six amino acids downstream of the tat AUG; in addition, a deletion mutant was made that removed most of the tat ORF. Both of these mutants had greatly reduced virus titers (> 1000-fold less than the wild type infectious clone). Co-transfection of a tat expressing plasmid with these viruses containing the tat ORF mutations resulted in higher levels of virus production demonstrating that the effects of both mutants are tat specific. These mutants provide data that the CAEV tat gene is necessary for efficient virus replication. Analysis of the RNA in these transfected cells showed that complementation of the tat gene was in trans and not the result of recombination. Analysis of the gag and rev proteins in the transfected cells demonstrated that these proteins were not detectable in cells transfected with the tat mutants but could be readily detected when the mutations were complemented in trans with a tat expression vector. To test for tat mediated trans-activation a plasmid expressing the CAEV tat ORF was co-transfected with plasmids containing either the CAEV or visna virus LTR driving transcription of the bacterial chloramphenicol acetyltransferase gene (CAT). These experiments indicate that one function of the CAEV tat protein is to trans-activate gene expression from the viral promoter. RNase protection analysis of CAT mRNA from co-transfected cells demonstrated that CAEV Tat trans-activates gene expression by increasing steady-state levels of mRNA.

Amino Acid Sequence↗

Molecular characterization of the segment 2 gene of epizootic hemorrhagic disease virus serotype 2: gene sequence and genetic diversity.

The complete nucleotide sequence of the gene encoding the major outer capsid protein VP2 from the Alberta isolate of epizootic hemorrhagic disease virus serotype 2 (EHDV-2) was determined. Complementary DNA (cDNA) corresponding to segment 2 was 3002 nucleotides in length with a single open reading frame that encoded a VP2 of 982 amino acids. Although the VP2 from EHDV-2 was only 34% homologous to the cognate protein from EHDV-1, their predicted hydropathic profiles were similar, suggesting that conservation of structure is important biologically to these capsid proteins. Sequence analysis of six North American EHDV-2 field isolates showed a high degree of comparative genetic identity (> 97%). Phylogenetic profiles constructed suggest that regionalization of the viruses within the North American continent has contributed to the genetic diversity.

Amino Acid Sequence↗

Comparative genomics of the HOG-signalling system in fungi.

Signal transduction pathways play crucial roles in cellular adaptation to environmental changes. In this study, we employed comparative genomics to analyse the high osmolarity glycerol pathway in fungi. This system contains several signalling modules that are used throughout eukaryotic evolution, such as a mitogen-activated protein kinase and a phosphorelay module. Here we describe the identification of pathway components in 20 fungal species. Although certain proteins proved difficult to identify due to low sequence conservation, a main limitation was incomplete, low coverage genomic sequences and fragmentary genome annotation. Still, the pathway was readily reconstructed in each species, and its architecture could be compared. The most striking difference concerned the Sho1 branch, which frequently does not appear to activate the Hog1 MAPK module, although its components are conserved in all but one species. In addition, two species lacked apparent orthologues for the Sln1 osmosensing histidine kinase. All information gathered has been compiled in an MS Excel sheet, which also contains interactive visualisation tools. In addition to primary sequence analysis, we employed analysis of protein size conservation. Protein size appears to be conserved largely independently from primary sequence and thus provides an additional tool for functional analysis and orthologue identification.

Cluster Analysis↗