PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Visualizing associations between genome sequences and gene expression data using genome-mean expression profiles.

The combination of genome-wide expression patterns and full genome sequences offers a great opportunity to further our understanding of the mechanisms and logic of transcriptional regulation. Many methods have been described that identify sequence motifs enriched in transcription control regions of genes that share similar gene expression patterns. Here we present an alternative approach that evaluates the transcriptional information contained by specific sequence motifs by computing for each motif the mean expression profile of all genes that contain the motif in their transcription control regions. These genome-mean expression profiles (GMEP's) are valuable for visualizing the relationship between genome sequences and gene expression data, and for characterizing the transcriptional importance of specific sequence motifs. Analysis of GMEP's calculated from a dataset of 519 whole-genome microarray experiments in Saccharomyces cerevisiae show a significant correlation between GMEP's of motifs that are reverse complements, a result that supports the relationship between GMEP's and transcriptional regulation. Hierarchical clustering of GMEP's identifies clusters of motifs that correspond to binding sites of well-characterized transcription factors. The GMEP's of these clustered motifs have patterns of variation across conditions that reflect the known activities of these transcription factors. Software that computed GMEP's from sequence and gene expression data is available under the terms of the Gnu Public License from http://rana.lbl.gov/.

Algorithms↗

Long-range correlation properties of coding and noncoding DNA sequences: GenBank analysis.

An open question in computational molecular biology is whether long-range correlations are present in both coding and noncoding DNA or only in the latter. To answer this question, we consider all 33301 coding and all 29453 noncoding eukaryotic sequences--each of length larger than 512 base pairs (bp)--in the present release of the GenBank to dtermine whether there is any statistically significant distinction in their long-range correlation properties. Standard fast Fourier transform (FFT) analysis indicates that coding sequences have practically no correlations in the range from 10 bp to 100 bp (spectral exponent beta=0.00 +/- 0.04, where the uncertainty is two standard deviations). In contrast, for noncoding sequences, the average value of the spectral exponent beta is positive (0.16 +/- 0.05) which unambiguously shows the presence of long-range correlations. We also separately analyze the 874 coding and the 1157 noncoding sequences that have more than 4096 bp and find a larger region of power-law behavior. We calculate the probability that these two data sets (coding and noncoding) were drawn from the same distribution and we find that it is less than 10(-10). We obtain independent confirmation of these findings using the method of detrended fluctuation analysis (DFA), which is designed to treat sequences with statistical heterogeneity, such as DNA's known mosaic structure ("patchiness") arising from the nonstationarity of nucleotide concentration. The near-perfect agreement between the two independent analysis methods, FFT and DFA, increases the confidence in the reliability of our conclusion.

Animals↗

Interior-branch and bootstrap tests of phylogenetic trees.

We have compared statistical properties of the interior-branch and bootstrap tests of phylogenetic trees when the neighbor-joining tree-building method is used. For each interior branch of a predetermined topology, the interior-branch and bootstrap tests provide the confidence values, PC and PB, respectively, that indicate the extent of statistical support of the sequence cluster generated by the branch. In phylogenetic analysis these two values are often interpreted in the same way, and if PC and PB are high (say, > or = 0.95), the sequence cluster is regarded as reliable. We have shown that PC is in fact the complement of the P-value used in the standard statistical test, but PB is not. Actually, the bootstrap test usually underestimates the extent of statistical support of species clusters. The relationship between the confidence values obtained by the two tests varies with both the topology and expected branch lengths of the true (model) tree. The most conspicuous difference between PC and PB is observed when the true tree is starlike, and there is a tendency for the difference to increase as the number of sequences in the tree increases. The reason for this is that the bootstrap test tends to become progressively more conservative as the number of sequences in the tree increases. Unlike the bootstrap, the interior-branch test has the same statistical properties irrespective of the number of sequences used when a predetermined tree is considered. Therefore, the interior-branch test appears to be preferable to the bootstrap test as long as unbiased estimators of evolutionary distances are used. However, when the interior-branch is applied to a tree estimated from a given data set, PC may give an overestimate of statistical confidence. For this case, we developed a method for computing a modified version (P'C) of the PC value and showed that this P'C tends to give a conservative estimate of statistical confidence, though it is not as conservative as PB. In this paper we have introduced a model in which evolutionary distances between sequences follow a multivariate normal distribution. This model allowed us to study the relationships between the two tests analytically.

Computer Simulation↗

Molecular parentage analysis in experimental newt populations: the response of mating system measures to variation in the operational sex ratio.

Molecular studies of parentage have been extremely influential in the study of sexual selection in the last decade, but a consensus statistical method for the characterization of genetic mating systems has not yet emerged. Here we study the utility of alternative mating system measures by experimentally altering the intensity of sexual selection in laboratory-based breeding populations of the rough-skinned newt. Our experiment involved skewed sex ratio (high sexual selection) and even sex ratio (low sexual selection) treatments, and we assessed the mating system by assigning parentage with microsatellite markers. Our results show that mating system measures based on Bateman's principles accurately reflect the intensity of sexual selection. One key component of this way of quantifying mating systems is the Bateman gradient, which is currently underutilized in the study of genetic mating systems. We also compare inferences based on Bateman's principles with those obtained using two other mating system measures that have been advocated recently (Morisita's index and the index of resource monopolization), and our results produce no justification for the use of these alternative measures. Overall, our results show that Bateman's principles provide the best available method for the statistical characterization of mating systems in nature.

Animals↗

The optimal discovery procedure for large-scale significance testing, with applications to comparative microarray experiments.

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the optimal discovery procedure (ODP), which has recently been introduced and theoretically shown to optimally perform multiple significance tests. Whereas existing procedures essentially use data from only one feature at a time, the ODP approach uses the relevant information from the entire data set when testing each feature. In particular, we propose a generally applicable estimate of the ODP for identifying differentially expressed genes in microarray experiments. This microarray method consistently shows favorable performance over five highly used existing methods. For example, in testing for differential expression between two breast cancer tumor types, the ODP provides increases from 72% to 185% in the number of genes called significant at a false discovery rate of 3%. Our proposed microarray method is freely available to academic users in the open-source, point-and-click EDGE software package.

Apoptosis Regulatory Proteins↗

An alignment-free strategy for circulating tumor DNA detection and tumor fraction estimation from whole-genome sequencing data.

Circulating tumor DNA (ctDNA) is emerging as a promising biomarker for postoperative monitoring of cancer patients. Precise estimation of circulating tumor fraction is crucial for evaluating treatment effects and timely detection of disease recurrence. All current ctDNA detection methods that utilize whole-genome sequencing (WGS) data rely on the reference genome alignment of sequencing reads and often apply separate tools for detecting different variant types. However, various bioinformatic analysis confounders and the application of external variant calling tools could be avoided by analyzing k-mers from unaligned sequencing reads. While k-mer-based methods have successfully been applied for somatic variant validation and detection, the potential of k-mer-based ctDNA detection is unexplored. We have developed a tumor-informed alignment-free ctDNA detection tool called ctDNAmer that detects tumor-specific somatic variation directly from unaligned sequencing data by identifying k-mers unique to the tumor DNA. ctDNAmer detects variant information across the genome by comparing the primary tumor and germline WGS data and accounts for sample-specific germline variability and technical noise in the same framework. We tested the utility of ctDNAmer for tumor fraction estimation on postoperative plasma cfDNA WGS data (mean sequencing depth ~ 28x) from 90 stage III colorectal cancer patients with three years of follow-up. The tumor fraction (TF) estimates agreed with the available clinical information and ctDNA was detected in 77% (17/22) of recurring patients with a median lead time of 8 months compared to radiological imaging. We further validated ctDNAmer's tumor fraction estimates based on a comparison with the mean cfDNA allele frequencies of somatic clonal SNVs identified from aligned primary tumor sequencing data. The TF estimates showed a strong Pearson correlation of 0.897 with the mean allele frequencies and improved ctDNA detection results across samples with an AUC of 0.79 compared to 0.75 if the mean allele frequency of clonal mutations is used.

Circulating Tumor DNA↗

Genetic analysis of pigment biosynthesis in Xanthobacter autotrophicus Py2 using a new, highly efficient transposon mutagenesis system that is functional in a wide variety of bacteria.

A highly efficient method of transposon mutagenesis was developed for genetic analysis of Xanthobacter autotrophicus Py2. The method makes use of a transposon delivery vector that encodes a hyperactive Tn 5 transposase that is 1,000-fold more active than the wild-type transposase. In this construct, the transposase is expressed from the promoter of the tetA gene of plasmid RP4, which is functional in a wide variety of organisms. The transposon itself contains a kanamycin resistance gene as a selectable marker and the origin of replication from plasmid R6K to facilitate subsequent cloning of the resulting insertion site. To test the effectiveness of this method, mutants unable to produce the characteristic yellow pigment (zeaxanthin dirhamnoside) of X. autotrophicus Py2 were isolated and analyzed. Transposon insertions were obtained at high frequency: approximately 1 x 10(-3) per recipient cell. Among these, pigment mutants were observed at a frequency of approximately 10(-3). Such mutants were found to have transposon insertions in genes homologous to known carotenoid biosynthetic genes previously characterized in other pigmented bacteria. Mutants were also isolated in Pseudomonas stutzeri and in an Alcaligenes faecalis, demonstrating the effectiveness of the method in diverse Proteobacteria. Preliminary results from other laboratories have confirmed the effectiveness of this method in additional phylogenetically diverse species.

Cloning, Molecular↗

Accurate anchoring alignment of divergent sequences.

MOTIVATION: Obtaining high quality alignments of divergent homologous sequences for cross-species sequence comparison remains a challenge. RESULTS: We propose a novel pairwise sequence alignment algorithm, ACANA (ACcurate ANchoring Alignment), for aligning biological sequences at both local and global levels. Like many fast heuristic methods, ACANA uses an anchoring strategy. However, unlike others, ACANA uses a Smith-Waterman-like dynamic programming algorithm to recursively identify near-optimal regions as anchors for a global alignment. Performance evaluations using a simulated benchmark dataset and real promoter sequences suggest that ACANA is accurate and consistent, especially for divergent sequences. Specifically, we use a simulated benchmark dataset to show that ACANA has the highest sensitivity to align constrained functional sites compared to BLASTZ, CHAOS and DIALIGN for local alignment and compared to AVID, ClustalW, DIALIGN and LAGAN for global alignment. Applied to 6007 pairs of human-mouse orthologous promoter sequences, ACANA identified the largest number of conserved regions (defined as over 70% identity over 100 bp) compared to AVID, ClustalW, DIALIGN and LAGAN. In addition, the average length of conserved region identified by ACANA was the longest. Thus, we suggest that ACANA is a useful tool for identifying functional elements in cross-species sequence analysis, such as predicting transcription factor binding sites in non-coding DNA. AVAILABILITY: ACANA software and test sequence data are publicly available at http://BioMedEmpire.org/

Algorithms↗

Activation of calpain in lens: a review and proposed mechanism.

The purpose of these experiments was to develop a hypothesis to explain activation of m-calpain in cataractogenesis observed in rodents. The in vitro model used to study m-calpain activation was to correlate breakdown of the 'reporter' protein alpha-crystallin with the appearance of activated m-calpain using protein sequencing and casein zymography. Incubation of alpha-crystallins with m-calpain and Ca2+ caused proteolysis of alpha-crystallins and accumulation of new polypeptides. E64 and calpain inhibitor I each inhibited proteolysis of alpha-crystallins. The N-terminus of the 80 kDa subunit of m-calpain was blocked at time 0 (pro calpain). After incubation with Ca2+, the remaining 80 kDa subunit of m-calpain gave a N-terminal sequence of KDREAAEGLG, indicating loss of nine amino acid from the N-terminus (autolysed calpain). The new 43 kDa m-calpain fragment also gave a N-terminal sequence of KDREAAEGLG, indicating the same loss of the first nine amino acids on the N-terminus as well as a major loss of the C-terminal half of the subunit (degraded calpain). In contrast, the N-terminus of the 80 kDa subunit of m-calpain remained blocked when E64 was present (unautolysed form). Moreover, the Ca2+ concentration required for proteolysis decreased when calpain was pre-incubated with Ca2+, although proteolysis of alpha-crystallin required a higher Ca2+ concentration than proteolysis of casein. These data suggested that the sequence of events for m-calpain activation were unautolysed, autolysed and finally degraded calpain. Unautolysed and/or autolysed calpains may be proteolytically active against alpha-crystallin.

Amino Acid Sequence↗

The use of Bayesian analysis of PCR-derived genomic DNA sequences to enable a biologically relevant interpretation of immunoglobulin variable region gene expression.

We analyzed by means of polymerase chain reactions (PCRs) and DNA sequencing techniques the immunoglobulin heavy chain variable region genes of bone marrow B lineage cells. We first formulated an explanatory model to guide understanding of the biological mechanisms determining both the size of the total available pool of relevant genes and clonal expansion following heavy chain gene rearrangement. We then followed Box's paradigm of criticism and estimation to interpret our experimental findings.

Base Sequence↗

Molecular phylogenetics of Calamus (Palmae) and related rattan genera based on 5S nrDNA spacer sequence data.

Phylogenetic relationships among the rattan palm genera Calamus, Daemonorops, Ceratolobus, Calospatha, Pogonotium, and Retispatha were investigated using DNA sequences from the nontranscribed spacer of 5S nrDNA. Moderate levels of intragenome polymorphism were identified, indicating that concerted evolution is not completely homogenizing the multiple copies of the 5S nrDNA repeat present in the nuclear genome. The existence of intragenome polymorphism did not excessively interfere with phylogeny reconstruction because, in the majority of cases, multiple clones obtained from individual species were resolved as monophyletic groups. The highly speciose genus Calamus was found to be nonmonophyletic with all five remaining genera being embedded within it. A number of major lineages within Calamus were resolved, one of which included the monotypic genus Calospatha, another included the monotypic genus Retispatha, and a third included a monophyletic group comprising Daemonorops, Ceratolobus, and Pogonotium. While the findings indicate that generic circumscriptions require revision, a nomenclatural solution was not sought at this stage because inadequate sampling and lack of support at basal nodes suggested that the topologies obtained might not be entirely reliable. Under these circumstances, name changes to such an important group would be both unhelpful and irresponsible.

Africa↗

OsPPR1, a pentatricopeptide repeat protein of rice is essential for the chloroplast biogenesis.

In this paper, we report a novel pentatricopeptide repeat (PPR) protein gene in rice. PPR, a characteristic repeat motif consisted of tandem 35 amino acids, has been found in various biological systems including plant. Sequence analysis revealed that the gene designated OsPPR1 consisted of an open reading frame of 2433 nucleotides encoding 810 amino acids that include 11 PPR motifs. Blast search result indicated that the gene did not align with any of the characterized PPR genes in plant. The OsPPR1 gene was found to contain a putative chloroplast transit peptide in the N-terminal region, suggesting that the gene product targets to the chloroplast. Southern blot hybridization indicated that the OsPPR1 is the member of a gene family within the rice genome. Expression analysis and immunoblot analysis suggested that the OsPPR1 was accumulated mainly in rice leaf. Antisense transgenic strategy was used to suppress the expression of OsPPR1 and the resulted transgenic rice showed the typical phenotypes of chlorophyll-deficient mutants; albinism and lethality. Cytological observation using microscopy revealed that the antisense transgenic plant contained a significant defect in the chloroplast development. Taken together, the results suggest that the OsPPR1 is a nuclear gene of rice, encoding the PPR protein that might play a role in the chloroplast biogenesis. This is the first report on the PPR protein required for the chloroplast biogenesis in rice.

Amino Acid Sequence↗

Iron metabolism in insect disease vectors: mining the Anopheles gambiae translated protein database.

All animals require iron for survival. This requirement reflects the role of this mineral as a cofactor of numerous proteins. However, under physiological conditions, Fe(2+) oxidizes to Fe(3+) encouraging the formation of toxic free radicals. In mammals, the potential for oxidative damage from iron is minimized by binding iron to proteins. Mammalian iron metabolism is complex and numerous proteins are involved in iron absorption, transport, uptake and utilization. We have analyzed the Anopheles gambiae translated protein database for candidates that show identity to proteins involved in mammalian iron metabolism (Holt et al., 2002. The genome sequence of the malaria mosquito Anopheles gambiae. Science 298, 129-149). Our results indicate that proteins involved in iron absorption and intracellular iron utilization are, for the most part, conserved in A. gambiae. In contrast, proteins involved in the pathways of iron export from the gut, transport in hemolymph and uptake at peripheral tissues in mosquitos differ from those for mammals.

Animals↗

Phylogeny of the genus Aphis Linnaeus, 1758 (Homoptera: Aphididae) inferred from mitochondrial DNA sequences.

Aphis is the largest aphid genus in the world and contains several of the most injurious aphid pests. It is also the most reluctant aphid genus to any comprehensive taxonomic treatment: while most species are easily classified into "species groups" that form well defined entities, numerous species within these groups are difficult to tell apart morphologically and identification keys remain ambiguous and mostly rely on host plant affiliation. In this paper, we used partial sequences of COI/COII and CytB genes to reconstruct the first phylogeny of Aphis and discuss the present systematics. The monophyly of the subgenus Bursaphis and of the tree major species groups, Black aphid, Black backed aphid and frangulae-like species was recovered by all phylogenetic analyses. However our data suggested that the nominal subgenus was not monophyletic. Relationships between major species groups were often ambiguous but "Black" and "Black backed" species groups appeared as sister clades. The most striking result of this study was that our molecular data met the same limits as the morphological characters used in classifications: mitochondrial DNA did not allow the differentiation of species that are difficult to identify. Further, interspecies relationships within groups of species for which taxonomic treatment is difficult stayed unresolved. This suggests that species delineation in the genus Aphis is often ambiguous and that diversification might have been a rapid process.

Animals↗

Molecular biology for the pediatric surgeon.

Molecular biology is leading a revolution in our understanding, diagnosis, and treatment of disease and will continue to do so. Medicine in the future will require a greater understanding of this field and its methods by medical practitioners. This report reviews the basic aspects of the field including recombinant DNA methods. Of particular importance is how molecular biology will impact pediatric surgeons. Accordingly, the final section of this report briefly reviews the molecular biology of three diseases commonly treated by pediatric surgeons.

DNA, Recombinant↗