PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The genome sequence and structure of rice chromosome 1.

The rice species Oryza sativa is considered to be a model plant because of its small genome size, extensive genetic map, relative ease of transformation and synteny with other cereal crops. Here we report the essentially complete sequence of chromosome 1, the longest chromosome in the rice genome. We summarize characteristics of the chromosome structure and the biological insight gained from the sequence. The analysis of 43.3 megabases (Mb) of non-overlapping sequence reveals 6,756 protein coding genes, of which 3,161 show homology to proteins of Arabidopsis thaliana, another model plant. About 30% (2,073) of the genes have been functionally categorized. Rice chromosome 1 is (G + C)-rich, especially in its coding regions, and is characterized by several gene families that are dispersed or arranged in tandem repeats. Comparison with a draft sequence indicates the importance of a high-quality finished sequence.

Arabidopsis↗

Applications of InterPro in protein annotation and genome analysis.

The applications of InterPro span a range of biologically important areas that includes automatic annotation of protein sequences and genome analysis. In automatic annotation of protein sequences InterPro has been utilised to provide reliable characterisation of sequences, identifying them as candidates for functional annotation. Rules based on the InterPro characterisation are stored and operated through a database called RuleBase. RuleBase is used as the main tool in the sequence database group at the EBI to apply automatic annotation to unknown sequences. The annotated sequences are stored and distributed in the TrEMBL protein sequence database. InterPro also provides a means to carry out statistical and comparative analyses of whole genomes. In the Proteome Analysis Database, InterPro analyses have been combined with other analyses based on CluSTr, the Gene Ontology (GO) and structural information on the proteins.

Amino Acid Sequence↗

Human endometrial receptivity: gene regulation.

Endometrial receptivity is a self-limited period in which the endometrial epithelium (EE) acquires a functional and transient ovarian steroid-dependent status that allows blastocyst adhesion. Termed as "the window of implantation", this specific period opens 4-5 days after progesterone production or administration and closes after 9-10 days. Scientific knowledge on the endometrial receptivity process is fundamental for the understanding of human reproduction, but so far none of the proposed biochemical markers for endometrial receptivity has been proven to be clinically useful. In this work, we present strategies of cDNA analysis technologies that aim to clarify the fragmented information in this field. Specifically, the objective is the differential identification, cloning and sequencing of genes linked to endometrial receptivity in humans, combining differential display PCR and cDNA microarray analysis of endometrial epithelial-derived cell lines and endometrial samples obtained in the same patient 2 and 7 days after the luteinizing hormone (LH) surge (day LH+2) and (day LH+7), respectively.

Cell Line↗

Medaka receptors for somatolactin and growth hormone: phylogenetic paradox among fish growth hormone receptors.

Somatolactin (SL) in fish belongs to the growth hormone/prolactin family. Its ortholog in tetrapods has not been identified and its function(s) remains largely unknown. The SL-deficient mutant of medaka (color interfere, ci) and an SL receptor (SLR) recently identified in salmon provide a fascinating field for investigating SL's function(s) in vivo. Here we isolated a medaka ortholog of the salmon SLR. The mRNA is transcribed in variable organs. Triglycerides and cholesterol contents in the ci are significantly higher than those in the wild type, providing the first evidence of SL's function in suppressing lipid accumulation to organs. Interestingly, phylogenetic comparisons between the medaka SLR and growth hormone receptor (GHR), which is also isolated in this study, in relation to GHRs of other fish, suggested that all GHRs reported from nonsalmonid species are, at least phylogenetically, SLRs. An extra intron inserted in medaka and pufferfish SLRs and flounder and sea bream GHRs also supports their orthologous relationship, but not with tetrapod GHRs. These results may indicate lineage-specific diversification of SLR and GHR functions among fish or just an inappropriate naming of these receptors. Further functional and comparative reassessments are necessary to address this question.

Amino Acid Sequence↗

Primer on medical genomics. Part III: Microarray experiments and data analysis.

Genomics has been defined as the comprehensive study of whole sets of genes, gene products, and their interactions as opposed to the study of single genes or proteins. Microarray technology is one of many novel tools that are allowing global and high-throughput analysis of genes and gene products. In addition to an introduction on underlying principles, the current review focuses on the use of both complementary DNA and oligodeoxynucleotide microarrays in gene expression analysis. Genome-wide experiments generate a massive amount of data points that require systematic methods of analysis to extract biologically useful information. Accordingly, the current educational communication discusses different methods of data analysis, including supervised and unsupervised clustering algorithms. Illustrative clinical examples show clinical applications, including (1) identification of candidate genes or pathological pathways (ie, elucidation of pathogenesis); (2) identification of "new" molecular classes of diseases that may be relevant in disease reclassification, prognostication, and treatment selection (ie, class discovery); and (3) use of expression profiles of known disease classes to predict diagnosis and classification of unknown samples (ie, class prediction). The current review should serve as an introduction to the subject for clinician investigators, physicians and medical scientists in training, practicing clinicians, and other students of medicine.

Breast Neoplasms↗

Mitochondrial protein phylogeny joins myriapods with chelicerates.

The animal phylum Arthropoda is very useful for the study of body plan evolution given its abundance of morphologically diverse species and our profound understanding of Drosophila development. However, there is a lack of consistently resolved phylogenetic relationships between the four extant arthropod subphyla, Hexapoda, Myriapoda, Chelicerata and Crustacea. Recent molecular studies have strongly supported a sister group relationship between Hexapoda and Crustacea, but have not resolved the phylogenetic position of Chelicerata and Myriapoda. Here we sequence the mitochondrial genome of the centipede species Lithobius forficatus and investigate its phylogenetic information content. Molecular phylogenetic analysis of conserved regions from the arthropod mitochondrial proteome yields highly resolved and congruent trees. We also find that a sister group relationship between Myriapoda and Chelicerata is strongly supported. We propose a model to explain the apparently parallel evolution of similar head morphologies in insects and myriapods.

Animals↗

Estimating time-dependent gene networks from time series microarray data by dynamic linear models with Markov switching.

In gene network estimation from time series microarray data, dynamic models such as differential equations and dynamic Bayesian networks assume that the network structure is stable through all time points, while the real network might changes its structure depending on time, affection of some shocks and so on. If the true network structure underlying the data changes at certain points, the fitting of the usual dynamic linear models fails to estimate the structure of gene network and we cannot obtain efficient information from data. To solve this problem, we propose a dynamic linear model with Markov switching for estimating time-dependent gene network structure from time series gene expression data. Using our proposed method, the network structure between genes and its change points are automatically estimated. We demonstrate the effectiveness of the proposed method through the analysis of Saccharomyces cerevisiae cell cycle time series data.

Algorithms↗

Immunome-derived vaccines.

Immune response to a subset of antigens and epitopes derived from an infectious pathogen may be sufficient for competent protection; immune recognition of every potential epitope derived from the pathogen's genome does not appear to be required. The pneumococcal and hepatitis vaccines, both of which are subunit vaccines, illustrate this premise. Similarly, 'immunome-derived vaccines' are based on the concept that response to the subset of antigens and epitopes that interface with the host immune system (the immunome) and not the whole organism (represented by the proteome or genome) can be sufficient for protection. Competent immune responses to cancer are also probably restricted to the neoplasm's 'immunome', although the set of antigens that drive successful immune response to cancer cells has proven more difficult to uncover. Researchers are now using bioinformatics sequence analysis tools, epitope mapping tools, microarrays and high-throughput immunology assays to discover the components of the immunome, which are then used to compose these new vaccines. At least one immunome-derived vaccine is in clinical trials and many others are in the vaccine pipeline. Due to the rapid improvement of immunoinformatics tools and immunological assays, the era of immunome-derived vaccines has begun.

Animals↗

MUSCLE: a multiple sequence alignment method with reduced time and space complexity.

BACKGROUND: In a previous paper, we introduced MUSCLE, a new program for creating multiple alignments of protein sequences, giving a brief summary of the algorithm and showing MUSCLE to achieve the highest scores reported to date on four alignment accuracy benchmarks. Here we present a more complete discussion of the algorithm, describing several previously unpublished techniques that improve biological accuracy and / or computational complexity. We introduce a new option, MUSCLE-fast, designed for high-throughput applications. We also describe a new protocol for evaluating objective functions that align two profiles. RESULTS: We compare the speed and accuracy of MUSCLE with CLUSTALW, Progressive POA and the MAFFT script FFTNS1, the fastest previously published program known to the author. Accuracy is measured using four benchmarks: BAliBASE, PREFAB, SABmark and SMART. We test three variants that offer highest accuracy (MUSCLE with default settings), highest speed (MUSCLE-fast), and a carefully chosen compromise between the two (MUSCLE-prog). We find MUSCLE-fast to be the fastest algorithm on all test sets, achieving average alignment accuracy similar to CLUSTALW in times that are typically two to three orders of magnitude less. MUSCLE-fast is able to align 1,000 sequences of average length 282 in 21 seconds on a current desktop computer. CONCLUSIONS: MUSCLE offers a range of options that provide improved speed and / or alignment accuracy compared with currently available programs. MUSCLE is freely available at http://www.drive5.com/muscle.

Cluster Analysis↗

Exploring tick saliva: from biochemistry to 'sialomes' and functional genomics.

Tick saliva, a fluid once believed to be only relevant for lubrication of mouthparts and water balance, is now well known to be a cocktail of potent anti-haemostatic, anti-inflammatory and immunomodulatory molecules that helps these arthropods obtain a blood meal from their vertebrate hosts. The repertoire of pharmacologically active components in this cocktail is impressive as well as the number of targets they specifically affect. These salivary components change the physiology of the host at the bite site and, consequently, some pathogens transmitted by ticks take advantage of this change and become more infective. Tick salivary proteins have therefore become an attractive target to control tick-borne diseases. Recent advances in molecular biology, protein chemistry and computational biology are accelerating the isolation, sequencing and analysis of a large number of transcripts and proteins from the saliva of different ticks. Many of these newly isolated genes code for proteins with homologies to known proteins allowing identification or prediction of their function. However, most of these genes code for proteins with unknown functions therefore opening the road to functional genomic approaches to identify their biological activities and roles in blood feeding and hence, vaccine development to control tick-borne diseases.

Animals↗

Gene expression arrays in cancer research: methods and applications.

During the last 5 years, the number of papers describing data obtained by microarray technology increased exponentially with about 3000 papers in 2003. Undoubtedly, cancer is by far the disease that received most of the attention as far as the amount of data generated. As array technology is rather new and highly dependent on bioinformatics, mathematics and statistics, a clear understanding of the knowledge and information derived from array-based experiments is not widely appreciated. We shall review herein some of the issues related to the construction of DNA arrays, quantities and heterogeneity of probes and targets, the consequences of the physical characteristics of the probes, data extraction and data analysis as well as the applications of array technology. Our goal is to bring to the general audience, some of the basics of array technology and its possible application in oncology. By discussing some of the basic aspects of the methodology, we hope to stimulate criticism concerning the conclusions proposed by authors, especially in the light of the very low degree of reproducibility already proven when commercially available platforms were compared . Regardless of its pitfalls, it is unquestionable that array technology will have a great impact in the management of cancer and its applications will range from the discovery of new drug targets, new molecular tools for diagnosis and prognosis as well as for a tailored treatment that will take into account the molecular determinants of a given tumor. Hence, we shall also highlight some of the already available and promising applications of array technology on the day-to-day practice of oncology.

Cluster Analysis↗

PCR primers based on different portions of insertion elements can assist genetic relatedness studies, strain fingerprinting and species identification in rhizobia.

Using the sequence of an insertion element originally found in Rhizobium sullae, the nitrogen-fixing bacterial symbiont of the legume Hedysarum coronarium, we devised three primer pairs (inbound, outbound and internal primers) for the following applications: (a) tracing genetic relatedness within rhizobia using a method independent of ribosomal inheritance, based on the presence and conservation of IS elements; (b) achieve sensitive and reproducible bacterial fingerprinting; (c) enable a fast and unambiguous detection of rhizobia at the species level. In terms of taxonomy, while in line with part of the 16S rRNA gene- and glutamine synthetase I-based clustering, the tools appeared nonetheless more coherent with the actual geographical ranges of origin of rhizobial species, strengthening the European-Mediterranean connections and discerning them from the asian and american taxa. The fingerprinting performance of the outward-pointing primers, designed upon the inverted repeats, was shown to be at least as sensitive as BOX PCR, and to be functional on a universal basis with all 13 bacterial species tested. The primers designed on the internal part of the transposase gene instead proved highly species-specific for R. sullae, enabling selective distinction from its most related species, and testing positive on every R. sullae strain examined, fulfilling the need of PCR-mediated species identification. A general use of other IS elements for a combined approach to rhizobial taxonomy and ecology is proposed.

Base Sequence↗

A dual selection based, targeted gene replacement tool for Magnaporthe grisea and Fusarium oxysporum.

Rapid progress in fungal genome sequencing presents many new opportunities for functional genomic analysis of fungal biology through the systematic mutagenesis of the genes identified through sequencing. However, the lack of efficient tools for targeted gene replacement is a limiting factor for fungal functional genomics, as it often necessitates the screening of a large number of transformants to identify the desired mutant. We developed an efficient method of gene replacement and evaluated factors affecting the efficiency of this method using two plant pathogenic fungi, Magnaporthe grisea and Fusarium oxysporum. This method is based on Agrobacterium tumefaciens-mediated transformation with a mutant allele of the target gene flanked by the herpes simplex virus thymidine kinase (HSVtk) gene as a conditional negative selection marker against ectopic transformants. The HSVtk gene product converts 5-fluoro-2'-deoxyuridine to a compound toxic to diverse fungi. Because ectopic transformants express HSVtk, while gene replacement mutants lack HSVtk, growing transformants on a medium amended with 5-fluoro-2'-deoxyuridine facilitates the identification of targeted mutants by counter-selecting against ectopic transformants. In addition to M. grisea and F. oxysporum, the method and associated vectors are likely to be applicable to manipulating genes in a broad spectrum of fungi, thus potentially serving as an efficient, universal functional genomic tool for harnessing the growing body of fungal genome sequence data to study fungal biology.

Agrobacterium tumefaciens↗

Bioinformatic methods for integrating whole-genome expression results into cellular networks.

Extracting a comprehensive overview from the huge amount of information arising from whole-genome analyses is a significant challenge. This review critically surveys the state of the art methods that are used to connect information from functional genomic studies to biological function. Cluster analysis methods for inferring the correlation between genes are discussed, as are the methods for integrating gene expression information with existing information on biological pathways and the methods that combine cluster analysis with biological information to reconstruct novel biological networks.

Cluster Analysis↗

Microarray data analysis: from disarray to consolidation and consensus.

In just a few years, microarrays have gone from obscurity to being almost ubiquitous in biological research. At the same time, the statistical methodology for microarray analysis has progressed from simple visual assessments of results to a weekly deluge of papers that describe purportedly novel algorithms for analysing changes in gene expression. Although the many procedures that are available might be bewildering to biologists who wish to apply them, statistical geneticists are recognizing commonalities among the different methods. Many are special cases of more general models, and points of consensus are emerging about the general approaches that warrant use and elaboration.

Algorithms↗

Functional genomic responses to cystic fibrosis transmembrane conductance regulator (CFTR) and CFTR(delta508) in the lung.

Cystic fibrosis (CF), a common lethal pulmonary disorder in Caucasians, is caused by mutations in the cystic fibrosis transmembrane conductance regulator gene (CFTR) that disturbs fluid homeostasis and host defense in target organs. The effects of CFTR and delta508-CFTR were assessed in transgenic mice that 1) lack CFTR expression (Cftr-/-); 2) express the human delta508 CFTR (CFTR(delta508)); 3) overexpress the normal human CFTR (CFTR(tg)) in respiratory epithelial cells. Genes were selected from Affymetrix Murine Gene-Chips analysis and subjected to functional classification, k-means clustering, promoter cis-elements/modules searching, literature mining, and pathway exploring. Genomic responses to Cftr-/- were not corrected by expression of CFTR(delta508). Genes regulating host defense, inflammation, fluid and electrolyte transport were similarly altered in Cftr-/- and CFTR(delta508) mice. CFTR(delta508) induced a primary disturbance in expression of genes regulating redox and antioxidant systems. Genomic responses to CFTR(tg) were modest and were not associated with lung pathology. CFTR(tg) and CFTR(delta508) induced genes encoding heat shock proteins and other chaperones but did not activate the endoplasmic reticulum-associated degradation pathway. RNAs encoding proteins that directly interact with CFTR were identified in each of the CFTR mouse models, supporting the hypothesis that CFTR functions within a multiprotein complex whose members interact at the level of protein-protein interactions and gene expression. Promoters of genes influenced by CFTR shared common regulatory elements, suggesting that their co-expression may be mediated by shared regulatory mechanisms. Genes and pathways involved in the response to CFTR may be of interest as modifiers of CF.

Animals↗

Study of coordinative gene expression at the biological process level.

MOTIVATION: Cellular processes are not isolated groups of events. Nevertheless, in most microarray analyses, they tend to be treated as standalone units. To shed light on how various parts of the interlocked biological processes are coordinated at the transcription level, there is a need to study the between-unit expressional relationship directly. RESULTS: We approach this issue by constructing an index of correlation function to convey the global pattern of coexpression between genes from one process and genes from the entire genome. Processes with similar signatures are then identified and projected to a process-to-process association graph. This top-down method allows for detailed gene-level analysis between linked processes to follow up. Using the cell-cycle gene-expression profiles for Saccharomyces cerevisiae, we report well-organized networks of biological processes that would be difficult to find otherwise. Using another dataset, we report a sharply different network structure featuring cellular responses under environmental stress. SUPPLEMENTARY INFORMATION: http://kiefer.stat.ucla.edu/lap2/download/KL_supplement.pdf.

Algorithms↗