PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

A biological question and a balanced (orthogonal) design: the ingredients to efficiently analyze two-color microarrays with Confirmatory Factor Analysis.

BACKGROUND: Factor analysis (FA) has been widely applied in microarray studies as a data-reduction-tool without any a-priori assumption regarding associations between observed data and latent structure (Exploratory Factor Analysis).A disadvantage is that the representation of data in a reduced set of dimensions can be difficult to interpret, as biological contrasts do not necessarily coincide with single dimensions. However, FA can also be applied as an instrument to confirm what is expected on the basis of pre-established hypotheses (Confirmatory Factor Analysis, CFA). We show that with a hypothesis incorporated in a balanced (orthogonal) design, including 'SelfSelf' hybridizations, dye swaps and independent replications, FA can be used to identify the latent factors underlying the correlation structure among the observed two-color microarray data. An orthogonal design will reflect the principal components associated with each experimental factor. We applied CFA to a microarray study performed to investigate cisplatin resistance in four ovarian cancer cell lines, which only differ in their degree of cisplatin resistance. RESULTS: Two latent factors, coinciding with principal components, representing the differences in cisplatin resistance between the four ovarian cancer cell lines were easily identified. From these two factors 315 genes associated with cisplatin resistance were selected, 199 genes from the first factor (False Discovery Rate (FDR): 19%) and 152 (FDR: 24%) from the second factor, while both gene sets shared 36. The differential expression of 16 genes was validated with reverse transcription-polymerase chain reaction. CONCLUSION: Our results show that FA is an efficient method to analyze two-color microarray data provided that there is a pre-defined hypothesis reflected in an orthogonal design.

Antineoplastic Agents, Alkylating↗

Microfluidic devices for DNA sequencing: sample preparation and electrophoretic analysis.

Modern DNA sequencing 'factories' have revolutionized biology by completing the human genome sequence, but in the race to completion we are left with inefficient, cumbersome, and costly macroscale processes and supporting facilities. During the same period, microfabricated DNA sequencing, sample processing and analysis devices have advanced rapidly toward the goal of a 'sequencing lab-on-a-chip'. Integrated microfluidic processing dramatically reduces analysis time and reagent consumption, and eliminates costly and unreliable macroscale robotics and laboratory apparatus. A microfabricated device for high-throughput DNA sequencing that couples clone isolation, template amplification, Sanger extension, purification, and electrophoretic analysis in a single microfluidic circuit is now attainable.

Base Sequence↗

Reproductive biology and genetic diversity of a cryptoviviparous mangrove aegiceras corniculatum (Myrsinaceae) using allozyme and intersimple sequence repeat (ISSR) analysis

Mangroves consist of a group of taxonomically diverse species representing about 20 families of angiosperms. However, little is known about their reproductive biology, genetic structure, and the ecological and genetic factors affecting this structure. Comparative studies of various mangrove species are needed to fill such gaps in our knowledge. The pollination biology, outcrossing rate, and genetic diversity of Aegiceras corniculatum were investigated in this study. Pollination experiments suggested that the species is predominantly pollinator-dependent in fruit setting. A quantitative analysis of the mating system was performed using progeny arrays assayed for intersimple sequence repeat (ISSR) markers. The multilocus outcrossing rate (tm) was estimated to be 0.653 in a wild population. Both allozyme and ISSR were used to investigate genetic variation within and among populations. The combined effects of founder events and enhanced local gene flow through seedling dispersal by ocean currents apparently played an important role in shaping the population genetic structure in this mangrove species. Both allozyme variation (P = 4.76%, A = 1.05, HE = 0.024) and ISSR diversity (P = 16.18%, A = 1.061, HE = 0.039) were very low at the species level, in comparison with other woody plants with mixed-mating or outcrossing systems. Gene differentiation among populations was also low: allozyme GST = 0.106 and ISSR GST = 0.178. The unusually high genetic identities (0.997 for allozyme and 0.992 for ISSR loci), however, suggest that these populations are probably all descended from a common ancestral population with low polymorphism.

Journal Article↗

GraBCas: a bioinformatics tool for score-based prediction of Caspase- and Granzyme B-cleavage sites in protein sequences.

Caspases and granzyme B are proteases that share the primary specificity to cleave at the carboxyl terminal of aspartate residues in their substrates. Both, caspases and granzyme B are enzymes that are involved in fundamental cellular processes and play a central role in apoptotic cell death. Although various targets are described, many substrates still await identification and many cleavage sites of known substrates are not identified or experimentally verified. A more comprehensive knowledge of caspase and granzyme B substrates is essential to understand the biological roles of these enzymes in more detail. The relatively high variability in cleavage site recognition sequence often complicates the identification of cleavage sites. As of yet there is no software available that allows identification of caspase and/or granzyme with cleavage sites differing from the consensus sequence. Here, we present a bioinformatics tool 'GraBCas' that provides score-based prediction of potential cleavage sites for the caspases 1-9 and granzyme B including an estimation of the fragment size. We tested GraBCas on already known substrates and showed its usefulness for protein sequence analysis. GraBCas is available at http://wwwalt.med-rz.uniklinik-saarland.de/med_fak/humangenetik/software/index.html.

Caspases↗

Adaptive radiation within New Zealand endemic species of the cockroach genus Celatoblatta Johns (Blattidae): a response to Plio-Pleistocene mountain building and climate change.

The South Island of New Zealand offers unique opportunities to study insect evolution due to long-term physical isolation, recent alpine habitats and high levels of biotic endemism. Using DNA sequence data from cytochrome oxidase subunit 1, we investigated the phylogeographical pattern among 10 endemic cockroach species within the genus Celatoblatta Johns (Blattidae). We tested the hypothesis that an ancestral cockroach species underwent rapid speciation in response to major climatic differentiation induced by mountain building. Results suggest that speciation was a twofold process, with an interspecific radiation of Pliocene/Pleistocene age followed by intraspecific diversification during the mid Pleistocene. Average genetic distance (maximum likelihood GTR + I + Gamma) was 9.17%, with a maximum of 14.5%. Data revealed eight deep well-supported branches, each with terminal clades. Six clades were differentiated according to morphological species, while the seventh was composed of three sympatric species. We consider the latter to be a phylogenetic species, possibly as a result of hybridization within a defined geographical area. This finding seriously challenges species distinctions for these three cockroach species. Correlation between genetic distances and a Climate Similarity Index (CSI) was negative, suggesting that species found in similar habitats are also genetically closely related. A Mantel test on within-clade genetic distances vs. linear geographical distance was positive, suggesting allopatric isolation for those haplotypes. We present a model of speciation for South Island Celatoblatta.

Adaptation, Biological↗

Separation of nearly identical repeats in shotgun assemblies using defined nucleotide positions, DNPs.

An increasingly important problem in genome sequencing is the failure of the commonly used shotgun assembly programs to correctly assemble repetitive sequences. The assembly of non-repetitive regions or regions containing repeats considerably shorter than the average read length is in practice easy to solve, while longer repeats have been a difficult problem. We here present a statistical method to separate arbitrarily long, almost identical repeats, which makes it possible to correctly assemble complex repetitive sequence regions. The differences between repeat units may be as low as 1% and the sequencing error may be up to ten times higher. The method is based on the realization that a comparison of only a part of all overlapping sequences at a time in a data set does not generate enough information for a conclusive analysis. Our method uses optimal multi-alignments consisting of all the overlaps of each read. This makes it possible to determine defined nucleotide positions, DNPs, which constitute the differences between the repeat units. Differences between repeats are distinguished from sequencing errors using statistical methods, where the probabilities of obtaining certain combinations of candidate DNPs are calculated using the information from the multi-alignments. The use of DNPs and combinations of DNPs will allow for optimal and rapid assemblies of repeated regions. This method can solve repeats that differ in only two positions in a read length, which is the theoretical limit for repeat separation. We predict that this method will be highly useful in shotgun sequencing in the future.

Algorithms↗

Assignment of disulfide bond location in prothoracicotropic hormone of the silkworm, Bombyx mori: a homodimeric peptide.

The disulfide bond location of a homodimeric peptide, prothoracicotropic hormone (PTTH) of the silkworm, Bombyx mori, was determined by a combination of partial reduction and sequence analysis of peptide fragments generated through a partial reduction of PTTH followed by alkylation and enzyme digestion. The partial reduction and S-alkylation broke the interchain disulfide bond but did not affect the intrachain disulfide bonds, generating monomeric PTTH whose intrachain disulfide bonds were kept intact. This monomeric PTTH has about one-half the biological activity of intact PTTH. Sequence analysis of the fragments generated by lysyl endopeptidase digestion of this monomeric PTTH after complete reduction and S-alkylation by another S-alkylating reagent showed that only the Cys15 residue was reduced and S-alkylated by the foregoing partial reduction, indicating that this residue formed the interchain disulfide bond. The other disulfide bonds which formed intrachain bridgings were determined by sequence and mass analyses of the fragments generated by two successive enzyme digestions of the monomeric PTTH. In conclusion, the disulfide bond location of PTTH was assigned to Cys15-Cys15' as an interchain disulfide linkage and Cys17-Cys54, Cys40-Cys96, and Cys48-Cys98 as intrachain disulfide linkages.

Amino Acid Sequence↗

The Lusitano horse maternal lineage based on mitochondrial D-loop sequence variation.

The analysis of mitochondrial D-loop sequences (408 bp) from 145 Lusitano founder mares yielded a total of 27 different haplotypes. The distribution of these mtDNA sequences was quite unequal, with the three most frequent ones representing 56.5% of all the Lusitano founder mares and 14 haplotypes (51.9%) being rare variants found only once in the sampling. Four main haplotype clusters were present in the Lusitano breed. The comparison of these sequences with other equine haplotypes shows that they fall in groups shared with other horse breeds. These data support the hypothesis of multiple domestication events in many distinct geographic areas over a broad time span. However, the analysis of 145 Lusitano, 55 Pura Raza Espanola and 18 Sorraia sequences indicates that half of the samples (50.9%) fall in one specific-cluster (A), which has previously been described as characteristic of the Iberian and Northern African horse breeds. The presence of a phylogeographic structure in cluster A associated with its star-like structure was interpreted as suggestive of a centre of horse domestication in the Iberia Peninsula.

Animals↗

A novel method for the rapid detection of specific nucleotide sequences in crude biological samples without blotting or radioactivity; application to the analysis of hepatitis B virus in human serum.

The detection of a little as 0.2 pg (60,000 molecules) of hepatitis B viral (HBV) DNA in human serum samples in 4 h has been demonstrated using a solution-hybridization and bead-capture method. An amplification method based on chemically crosslinked oligodeoxyribonucleotides was coupled with a horseradish peroxidase-labeling scheme for the ultimate detection of the analyte. Two sets of HBV complementary synthetic oligodeoxyribonucleotide probes containing one of two types of single-stranded (ss) overhangs were employed. These ss overhangs were used to capture the probe-analyte complex onto a bead and subsequently to label it. Detection was achieved with either a chemiluminescent or colorimetric output substrate for the enzyme. Only in the presence of the virus was label specifically bound to the support. The assay was relatively unaffected by either sample composition or by the presence of heterologous nucleic acids.

Base Sequence↗

Site-directed mutations in the Sindbis virus E2 glycoprotein identify palmitoylation sites and affect virus budding.

The assembly and budding of Sindbis virus, a prototypic member of the alphavirus subgroup in the family Togaviridae, requires a specific interaction between the nucleocapsid core and the membrane-embedded glycoproteins E1 and E2. These glycoproteins are modified posttranslationally by the addition of palmitic acid, and inhibitors of acylation interfere with this budding process (M.J. Schlesinger and C. Malfer, J. Biol. Chem. 257:9887-9890, 1982). This report describes the use of site-directed mutagenesis to identify two of the acylation sites in the E2 glycoprotein as the cysteines near the carboxyl terminus of the protein which is oriented to the cytoplasmic domain of this type 1 transmembrane protein. Additional mutations were made at two prolines within a hydrophobic sequence of E2 that is highly conserved among several alphaviruses, and the mutant viruses were aberrant in assembly and particle formation. These data support earlier studies indicating that the native structure of the cytoplasmic domain of E2 is essential for proper assembly of this enveloped virus.

Amino Acid Sequence↗

Identification of cyanobacterial non-coding RNAs by comparative genome analysis.

BACKGROUND: Whole genome sequencing of marine cyanobacteria has revealed an unprecedented degree of genomic variation and streamlining. With a size of 1.66 megabase-pairs, Prochlorococcus sp. MED4 has the most compact of these genomes and it is enigmatic how the few identified regulatory proteins efficiently sustain the lifestyle of an ecologically successful marine microorganism. Small non-coding RNAs (ncRNAs) control a plethora of processes in eukaryotes as well as in bacteria; however, systematic searches for ncRNAs are still lacking for most eubacterial phyla outside the enterobacteria. RESULTS: Based on a computational prediction we show the presence of several ncRNAs (cyanobacterial functional RNA or Yfr) in several different cyanobacteria of the Prochlorococcus-Synechococcus lineage. Some ncRNA genes are present only in two or three of the four strains investigated, whereas the RNAs Yfr2 through Yfr5 are structurally highly related and are encoded by a rapidly evolving gene family as their genes exist in different copy numbers and at different sites in the four investigated genomes. One ncRNA, Yfr7, is present in at least seven other cyanobacteria. In addition, control elements for several ribosomal operons were predicted as well as riboswitches for thiamine pyrophosphate and cobalamin. CONCLUSION: This is the first genome-wide and systematic screen for ncRNAs in cyanobacteria. Several ncRNAs were both computationally predicted and their presence was biochemically verified. These RNAs may have regulatory functions and each shows a distinct phylogenetic distribution. Our approach can be applied to any group of microorganisms for which more than one total genome sequence is available for comparative analysis.

Base Sequence↗

Current bioinformatics tools in genomic biomedical research (Review).

On the advent of a completely assembled human genome, modern biology and molecular medicine stepped into an era of increasingly rich sequence database information and high-throughput genomic analysis. However, as sequence entries in the major genomic databases currently rise exponentially, the gap between available, deposited sequence data and analysis by means of conventional molecular biology is rapidly widening, making new approaches of high-throughput genomic analysis necessary. At present, the only effective way to keep abreast of the dramatic increase in sequence and related information is to apply biocomputational approaches. Thus, over recent years, the field of bioinformatics has rapidly developed into an essential aid for genomic data analysis and powerful bioinformatics tools have been developed, many of them publicly available through the World Wide Web. In this review, we summarize and describe the basic bioinformatics tools for genomic research such as: genomic databases, genome browsers, tools for sequence alignment, single nucleotide polymorphism (SNP) databases, tools for ab initio gene prediction, expression databases, and algorithms for promoter prediction.

Computational Biology↗

XenDB: full length cDNA prediction and cross species mapping in Xenopus laevis.

BACKGROUND: Research using the model system Xenopus laevis has provided critical insights into the mechanisms of early vertebrate development and cell biology. Large scale sequencing efforts have provided an increasingly important resource for researchers. To provide full advantage of the available sequence, we have analyzed 350,468 Xenopus laevis Expressed Sequence Tags (ESTs) both to identify full length protein encoding sequences and to develop a unique database system to support comparative approaches between X. laevis and other model systems. DESCRIPTION: Using a suffix array based clustering approach, we have identified 25,971 clusters and 40,877 singleton sequences. Generation of a consensus sequence for each cluster resulted in 31,353 tentative contig and 4,801 singleton sequences. Using both BLASTX and FASTY comparison to five model organisms and the NR protein database, more than 15,000 sequences are predicted to encode full length proteins and these have been matched to publicly available IMAGE clones when available. Each sequence has been compared to the KOG database and approximately 67% of the sequences have been assigned a putative functional category. Based on sequence homology to mouse and human, putative GO annotations have been determined. CONCLUSION: The results of the analysis have been stored in a publicly available database XenDB http://bibiserv.techfak.uni-bielefeld.de/xendb/. A unique capability of the database is the ability to batch upload cross species queries to identify potential Xenopus homologues and their associated full length clones. Examples are provided including mapping of microarray results and application of 'in silico' analysis. The ability to quickly translate the results of various species into 'Xenopus-centric' information should greatly enhance comparative embryological approaches.

Animals↗

Genome sequences and structures of two biologically distinct strains of Grapevine leafroll-associated virus 2 and sequence analysis.

Grapevine leafroll-associated virus 2 (GLRaV-2), a member of the genus Closterovirus within Closteroviridae, is implicated in several important diseases of grapevines including "leafroll", "graft-incompatibility", and "quick decline" worldwide. Several GLRaV-2 isolates have been detected from different grapevine genotypes. However, the genomes of these isolates were not sequenced or only partially sequenced. Consequently, the relationship of these viral isolates at the molecular level has not been determined. Here, we group the various GLRaV-2 isolates into four strains based on their coat protein gene sequences. We show that isolates "PN" (originated from Vitis vinifera cv. "Pinot noir"), "Sem" (from V. vinifera cv. "Semillon") and "94/970" (from V. vinifera cv. "Muscat of Alexandria") belong to the same strain, "93/955" (from hybrid "LN-33") and "H4" (from V. rupestris "St. George") each represents a distinct strain, while Grapevine rootstock stem lesion-associated virus.

Amino Acid Sequence↗

The early dark-response in Arabidopsis thaliana revealed by cDNA microarray analysis.

Despite intense research on light responses in plants, the consequences of a simple shift from light to darkness remain poorly characterized. We have examined the transcriptome of Arabidopsis thaliana seedling leaves upon a shift from constant light to darkness for between 1 and 8 h, while excluding most effects associated with circadian oscillation. Expression clustering and gene ontology analyses identified about 790 responsive genes implicated in diverse cellular processes. Compared to the better-studied long-term dark adaptation response, the early response to darkness is partially overlapping yet clearly distinct, encompassing early transient, early sustained, and late response clusters. The repressor of photomorphogenesis, COP1 (constitutive photomorphogenic 1), is not a chief regulator of the early response to darkness, in contrast to its well-established role during long-term dark adaptation and etiolation. Only part of the early dark response can be understood as the opposite of the response following a dark-to-light transition and as a response to sugar deprivation. Bioinformatic comparisons with published microarray datasets further suggest that abscisic acid (ABA) signaling plays a prominent role in the early response to darkness, although this effect is not mediated by an increase in the ABA level. The potential basis for the co-regulation by darkness and ABA is discussed in light of sugar and redox signaling.

Abscisic Acid↗