PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Long branch attraction, taxon sampling, and the earliest angiosperms: Amborella or monocots?

BACKGROUND: Numerous studies, using in aggregate some 28 genes, have achieved a consensus in recognizing three groups of plants, including Amborella, as comprising the basal-most grade of all other angiosperms. A major exception is the recent study by Goremykin et al. (2003; Mol. Biol. Evol. 20:1499-1505), whose analyses of 61 genes from 13 sequenced chloroplast genomes of land plants nearly always found 100% support for monocots as the deepest angiosperms relative to Amborella, Calycanthus, and eudicots. We hypothesized that this conflict reflects a misrooting of angiosperms resulting from inadequate taxon sampling, inappropriate phylogenetic methodology, and rapid evolution in the grass lineage used to represent monocots. RESULTS: We used two main approaches to test this hypothesis. First, we sequenced a large number of chloroplast genes from the monocot Acorus and added these plus previously sequenced Acorus genes to the Goremykin et al. (2003) dataset in order to explore the effects of altered monocot sampling under the same analytical conditions used in their study. With Acorus alone representing monocots, strongly supported Amborella-sister trees were obtained in all maximum likelihood and parsimony analyses, and in some distance-based analyses. Trees with both Acorus and grasses gave either a well-supported Amborella-sister topology or else a highly unlikely topology with 100% support for grasses-sister and paraphyly of monocots (i.e., Acorus sister to "dicots" rather than to grasses). Second, we reanalyzed the Goremykin et al. (2003) dataset focusing on methods designed to account for rate heterogeneity. These analyses supported an Amborella-sister hypothesis, with bootstrap support values often conflicting strongly with cognate analyses performed without allowing for rate heterogeneity. In addition, we carried out a limited set of analyses that included the chloroplast genome of Nymphaea, whose position as a basal angiosperm was also, and very recently, challenged. CONCLUSIONS: These analyses show that Amborella (or Amborella plus Nymphaea), but not monocots, is the sister group of all other angiosperms among this limited set of taxa and that the grasses-sister topology is a long-branch-attraction artifact leading to incorrect rooting of angiosperms. These results highlight the danger of having lots of characters but too few and, especially, molecularly divergent taxa, a situation long recognized as potentially producing strongly misleading molecular trees. They also emphasize the importance in phylogenetic analysis of using appropriate evolutionary models.

Artifacts↗

Mitochondrial DNA and retroviral RNA analyses of archival oral polio vaccine (OPV CHAT) materials: evidence of macaque nuclear sequences confirms substrate identity.

Inoculation of live experimental oral poliovirus vaccines (OPV CHAT) during the 1950s in central Africa has been proposed to account for the introduction of HIV into human populations. For this to have occurred, it would have been necessary for chimpanzee rather than macaque kidney epithelial cells to have been included in the preparation of early OPV materials. Theoretically, this could have led to contamination with a progenitor of HIV-1 derived from a related simian immunodeficiency virus of chimpanzees (SIVCPZ). In this article we present further detailed analyses of two samples of OPV, CHAT 10A-11 and CHAT 6039/Yugo, which were used in early human trials of poliovirus vaccination. Recovery of poliovirus by culture techniques confirmed the biological viability of the vaccines and sequence analysis of poliovirus RNA specifically identified the presence of the CHAT strain. Independent nested sets of oligonucleotide primers specific for HIV-1/SIVCPZ and HIV-2/SIVMAC/SIVSM phylogenetic lineages, respectively, indicated no evidence of HIV/SIV RNA in either vaccine preparation, at a sensitivity of 100 RNA equivalents/ml. Analysis of cellular substrate by the amplification of two distinct regions of mitochondrial DNA (D-loop control region and 12S ribosomal sequences) revealed no evidence of chimpanzee cellular sequences. However, this approach positively identified rhesus and cynomolgus macaque DNA for the CHAT 10A-11 and CHAT 6039/Yugo vaccine preparations, respectively. Analysis of multiple clones of mtDNA 12S rDNA indicated a relatively high number of nuclear mitochondrial DNA sequences (numts) in the CHAT 10A-11 material, but confirmed the macaque origin of cellular substrate used in vaccine preparation. These data reinforce earlier findings on this topic providing no evidence to support the contention that poliovirus vaccination was responsible for the introduction of HIV into humans and sparking the AIDS pandemic.

Animals↗

Postgenomic bioinformatic analysis of yeast artificial chromosome sequence.

The free availability of multiple genomic sequences represents one of the greatest advances in biology of the new millennium, and promises to revolutionize our ability to determine and treat the causes of human disease. This chapter highlights a number of basic, freely available, and user-friendly bioinformatic techniques that can be used to predict the functional genetic contents of specific yeast artificial chromosome (YAC) clones. The content of this chapter is written for the level of graduate students, who may be relatively inexperienced with the use of computers for analyzing DNA sequences. The basic instructions that allow the identification of the genomic sequence of interest and to download this sequence onto a personal computer from an online database are presented. Simple instructions are also given on how to perform basic sequence manipulations, how to use online tools to design polymerase chain reaction primers, and how to map restriction sites. Also described are more complicated programs that rapidly and efficiently perform genome alignments that, in addition to predicting the location of protein coding sequences, allow the prediction of functional genomic sequences, such as cis regulatory elements and scaffold/matrix attachment sites. The availability of genomic sequences and the rapidly expanding numbers of predictive programs that allow the predictive analysis of these sequences promises to greatly facilitate the use of YAC clones in the search for the causes of disease.

Chromosomes, Artificial, Yeast↗

The Eukaryotic Promoter Database EPD: the impact of in silico primer extension.

The Eukaryotic Promoter Database (EPD) is an annotated non-redundant collection of eukaryotic POL II promoters, experimentally defined by a transcription start site (TSS). There may be multiple promoter entries for a single gene. The underlying experimental evidence comes from journal articles and, starting from release 73, from 5' ESTs of full-length cDNA clones used for so-called in silico primer extension. Access to promoter sequences is provided by pointers to TSS positions in nucleotide sequence entries. The annotation part of an EPD entry includes a description of the type and source of the initiation site mapping data, links to other biological databases and bibliographic references. EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets for comparative sequence analysis. Web-based interfaces have been developed that enable the user to view EPD entries in different formats, to select and extract promoter sequences according to a variety of criteria and to navigate to related databases exploiting different cross-references. Tools for analysing sequence motifs around TSSs defined in EPD are provided by the signal search analysis server. EPD can be accessed at http://www.epd. isb-sib.ch.

Animals↗

Analysis of Sendai virus mRNAs with cDNA clones of viral genes and sequences of biologically important regions of the fusion protein.

cDNA clones representing five of the genes of Sendai virus (P, HN, NP, F, and M) were isolated and used to identify the viral mRNAs by hybridization. Five mRNAs that were monocistronic transcripts of these genes were identified. A sixth transcript, which was identified on the basis of size and of hybridization to viral RNA but not to the cDNA of the other five genes, is thought to represent the message for the L protein. In addition, polycistronic transcripts of the NP and P genes and of the M and F genes were also found. The latter establishes the position of the F gene adjacent to the M gene; these results confirm and extend the previously reported partial gene order of the virus. Nucleotide sequences and derived amino acid sequences of two biologically important regions of the F protein--approximately 25% of F proximal to its COOH terminus and the region spanning the site of the proteolytic cleavage that activates the fusion activity of the protein--are presented. The F protein has an unusually large "cytoplasmic domain" of 42 amino acids beyond the hydrophobic region by which it is anchored in the viral membrane. A single possible trypsin cleavage site was found at the junction of the F1 and F2 polypeptides, and 26 hydrophobic amino acids extend from this cleavage site at the NH2 terminus of the F1 polypeptide.

Amino Acid Sequence↗

Unraveling Neuronal Identities Using SIMS: A Deep Learning Label Transfer Tool for Single-Cell RNA Sequencing Analysis.

Large single-cell RNA datasets have contributed to unprecedented biological insight. Often, these take the form of cell atlases and serve as a reference for automating cell labeling of newly sequenced samples. Yet, classification algorithms have lacked the capacity to accurately annotate cells, particularly in complex datasets. Here we present SIMS (Scalable, Interpretable Machine Learning for Single-Cell), an end-to-end data-efficient machine learning pipeline for discrete classification of single-cell data that can be applied to new datasets with minimal coding. We benchmarked SIMS against common single-cell label transfer tools and demonstrated that it performs as well or better than state of the art algorithms. We then use SIMS to classify cells in one of the most complex tissues: the brain. We show that SIMS classifies cells of the adult cerebral cortex and hippocampus at a remarkably high accuracy. This accuracy is maintained in trans-sample label transfers of the adult human cerebral cortex. We then apply SIMS to classify cells in the developing brain and demonstrate a high level of accuracy at predicting neuronal subtypes, even in periods of fate refinement, shedding light on genetic changes affecting specific cell types across development. Finally, we apply SIMS to single cell datasets of cortical organoids to predict cell identities and unveil genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. When cell types are obscured by stress signals, label transfer from primary tissue improves the accuracy of cortical organoid annotations, serving as a reliable ground truth. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Brain organoids↗

Targeted Next-Generation Sequencing in Rare Diseases.

Targeted next-generation sequencing (NGS) in rare disease focuses on genetic analysis of specific regions in genome that are linked to a rare disease. In addition to library preparation, sequencing, and data analysis, targeted NGS includes an additional step of target enrichment of selected genes and regions. It allows for more sensitive and profound sequencing, as it is a fast and cost-effective approach with less data burden and is therefore often a method of choice for identifying rare variants in known genes, especially in diagnostics of rare diseases. Several in silico tools address the pathogenicity predictions of rare variants of unknown significance (VUS) and can therefore facilitate clinical interpretation.

Rare Diseases↗

Reverse engineering of regulatory networks in human B cells.

Cellular phenotypes are determined by the differential activity of networks linking coregulated genes. Available methods for the reverse engineering of such networks from genome-wide expression profiles have been successful only in the analysis of lower eukaryotes with simple genomes. Using a new method called ARACNe (algorithm for the reconstruction of accurate cellular networks), we report the reconstruction of regulatory networks from expression profiles of human B cells. The results are suggestive a hierarchical, scale-free network, where a few highly interconnected genes (hubs) account for most of the interactions. Validation of the network against available data led to the identification of MYC as a major hub, which controls a network comprising known target genes as well as new ones, which were biochemically validated. The newly identified MYC targets include some major hubs. This approach can be generally useful for the analysis of normal and pathologic networks in mammalian cells.

Algorithms↗

[Advances in molecular biology of dermatophytes].

During the 44th meeting of The Japanese Society for Medical Mycology in Nagasaki, 2000, a forum was held entitled Advances in Molecular Biology of Dermatophytes. Based on the subject, target molecules and kind of approach, we selected seven presentations from over 100 of the poster abstracts. Six of them concerned identification and one concerned viability. Summaries of the 7 presentations are given in this article. Of presentations on the identification methods, 5 demonstrated their usefulness: 1) A sequence analysis of ITS 1 region in ribosomal DNA of several Microsporum species showed ITS 1 genospecies Arthroderma otae to be composed of A. otae, M. canis, M. equinum and M. audouinii. 2) RAPD may be useful for identifying isolates which are not clearly identifiable by conventional biological techniques. 3) Sequence analysis of CHS 1 was shown to be a rapid tool for species level identification of M. gypseum. 4) PCR-SSCP analysis was also useful for discrimination of dermatophytes with high reproducibility and sensitivity. 5) Strain identification of A. benhamiae isolates may be possible using RFLP analysis of NTS regions in ribosomal DNA. The other presentation concerning identification pointed out some important problems: RFLP of mitochondrial DNA and ITS sequencing of A. benhamiae showed that the results are sometimes in conflict with those obtained from biological techniques, or in some cases, between other molecular techniques. This implies that our concept of fungal species needs to be re-examined and perhaps amended. The presentation on viability introduced quantitative analysis of mRNA of ACT gene, a new application of a molecular technique. Since the mRNA expresses only in living cells, the method is highly useful as an indicator of fungal viability.

Actins↗

Qualitative analysis of the relation between DNA microarray data and behavioral models of regulation networks.

We introduce a mathematical framework that allows to test the compatibility between differential data and knowledge on genetic and metabolic interactions. Within this framework, a behavioral model is represented by a labeled oriented interaction graph; its predictions can be compared to experimental data. The comparison is qualitative and relies on a system of linear qualitative equations derived from the interaction graph. We show how to partially solve the qualitative system, how to identify incompatibilities between the model and the data, and how to detect competitions in the biological processes that are modeled. This approach can be used for the analysis of transcriptomic, metabolic or proteomic data.

Fatty Acids↗

A Web interface generator for molecular biology programs in Unix.

MOTIVATION: Almost all users encounter problems using sequence analysis programs. Not only are they difficult to learn because of the parameters, syntax and semantic, but many are different. That is why we have developed a Web interface generator for more than 150 molecular biology command-line driven programs, including: phylogeny, gene prediction, alignment, RNA, DNA and protein analysis, motif discovery, structure analysis and database searching programs. The generator uses XML as a high-level description language of the legacy software parameters. Its aim is to provide users with the equivalent of a basic Unix environment, with program combination, customization and basic scripting through macro registration. RESULTS: The program has been used for three years by about 15000 users throughout the world; it has recently been installed on other sites and evaluated as a standard user interface for EMBOSS programs.

Computational Biology↗

Temporal evolution of mouse striatal gene expression following MPTP injury.

The gradual loss of striatal dopamine and dopaminergic neurons residing in the substantia nigra (SN) causes parkinsonism characterized by slow, halting movements, rigidity, and resting tremor when neuronal loss exceeds a threshold of approximately 80%. It is estimated that there is extensive compensation for several years prior to symptom onset, during which vulnerable neurons asynchronously die. Recent evidence would argue that much of the compensatory response of the nigrostriatal system is multimodal including both pre-synaptic and striatal mechanisms. Although parkinsonism may have multiple causes, the classic syndrome, Parkinson's disease (PD), is frequently modeled in small animals by repeated administration of the selective neurotoxin 1-methyl-4-phenyl-1,2,3,6-tetrahydropyridine (MPTP). Because the MPTP model of PD recapitulates many of the known behavioral and pathological features of human PD, we asked whether the striatal cells of mice treated with MPTP in a semi-chronic paradigm enact a transcriptional program that would help elucidate the response to dopamine denervation. Our findings reveal a time-dependent dysregulation in the striatum of a set of genes whose products may impact both the viability and ability to communicate of dopamine neurons in the SN.

1-Methyl-4-phenyl-1,2,3,6-tetrahydropyridine↗

Amplification and cloning of near full-length HIV-2 genomes.

The genomes of human immunodeficiency virus type 2 (HIV-2), like those of HIV-1, are not only extremely variable but are also highly recombinogenic. Determination of subtypes based on partial genomes cannot predict the subtype classification of other regions of the genome owing to the frequent occurrence of recombinant genomes among subtypes. To fully understand the genetic variation and evolution of HIV-2s, full-length viral genomes need to be obtained for genetic analysis. Full-length HIV-2 genomes can also be used as infectious clones to study viral biological characteristics and as reference sequences for phylogenetic analysis. More important, all genes in the obtained genomes can be cloned into expression vectors to produce proteins that can be used as antigens for diagnostic reagents or as immunogens for vaccine development. The long-range polymerase chain reaction technique has recently become a more preferred method than the lambda phage cloning method to obtain near-full-length HIV-2 genomes.

Cloning, Molecular↗

Effects of site and plant species on rhizosphere community structure as revealed by molecular analysis of microbial guilds.

The bacterial and fungal rhizosphere communities of strawberry (Fragaria ananassa Duch.) and oilseed rape (Brassica napus L.) were analysed using molecular fingerprints. We aimed to determine to what extent the structure of different microbial groups in the rhizosphere is influenced by plant species and sampling site. Total community DNA was extracted from bulk and rhizosphere soil taken from three sites in Germany in two consecutive years. Bacterial, fungal and group-specific (Alphaproteobacteria, Betaproteobacteria and Actinobacteria) primers were used to PCR-amplify 16S rRNA and 18S rRNA gene fragments from community DNA prior to denaturing gradient gel electrophoresis (DGGE) analysis. Bacterial fingerprints of soil DNA revealed a high number of equally abundant faint bands, while rhizosphere fingerprints displayed a higher proportion of dominant bands and reduced richness, suggesting selection of bacterial populations in this environment. Plant specificity was detected in the rhizosphere by bacterial and group-specific DGGE profiles. Different bulk soil community fingerprints were revealed for each sampling site. The plant species was a determinant factor in shaping similar actinobacterial communities in the strawberry rhizosphere from different sites in both years. Higher heterogeneity of DGGE profiles within soil and rhizosphere replicates was observed for the fungi. Plant-specific composition of fungal communities in the rhizosphere could also be detected, but not in all cases. Cloning and sequencing of 16S rRNA gene fragments obtained from dominant DGGE bands detected in the bacterial profiles of the Rostock site revealed that Streptomyces sp. and Rhizobium sp. were among the dominant ribotypes in the strawberry rhizosphere, while sequences from Arthrobacter sp. corresponded to dominant bands from oilseed rape bacterial fingerprints.

Actinobacteria↗

Detecting correlation between sequence and expression divergences in a comparative analysis of human serpin genes.

Physiological functions and characteristic structures of the serpin gene superfamily have been studied extensively, yet the evolution of the serpin genes remains unclear. Gene duplication in this superfamily may shed light on this issue. Two models are used to predict the preservation of duplicated genes: the classical model and the duplication-degeneration-complementation (DDC) model. In this study, we analyzed the phylogenetic relationships of 33 human serpin genes and the expression data of some members of the serpin superfamily from a DNA microarray of human leukemia U937 cells with stably inducible expression of the leukemia-related AML1-ETO gene. We then determined the utility of the DDC model by mapping serpin superfamily expression data to the phylogenetic tree. The correlation between sequence and expression divergences as measured by the Pearson correlation coefficient indicated that human serpin genes evolved under the DDC model. Our study provides a new strategy for comparative analysis of gene sequences and microarray data.

Cell Line, Tumor↗