PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

gff2ps: visualizing genomic annotations.

gff2psis a program for visualizing annotations of genomic sequences. The program takes the annotated features on a genomic sequence in GFF format as input, and produces a visual output in PostScript. While it can be used in a very simple way, it also allows for a great degree of customization through a number of options and/or customization files.

Computational Biology↗

wEMBOSS: a web interface for EMBOSS.

UNLABELLED: wEMBOSS provides a web environment from which the user can access EMBOSS in a user-friendly way. wEMBOSS supplies each user with space and tools to organize and review his or her work. AVAILABILITY: wEMBOSS can be downloaded at http://www.wemboss.org CONTACT: msarachu@biol.unlp.edu.ar.

Computer Graphics↗

Prediction of subcellular localization using sequence-biased recurrent networks.

MOTIVATION: Targeting peptides direct nascent proteins to their specific subcellular compartment. Knowledge of targeting signals enables informed drug design and reliable annotation of gene products. However, due to the low similarity of such sequences and the dynamical nature of the sorting process, the computational prediction of subcellular localization of proteins is challenging. RESULTS: We contrast the use of feed forward models as employed by the popular TargetP/SignalP predictors with a sequence-biased recurrent network model. The models are evaluated in terms of performance at the residue level and at the sequence level, and demonstrate that recurrent networks improve the overall prediction performance. Compared to the original results reported for TargetP, an ensemble of the tested models increases the accuracy by 6 and 5% on non-plant and plant data, respectively. AVAILABILITY: The Protein Prowler incorporating the recurrent network predictor described in this paper is available online at http://pprowler.imb.uq.edu.au/

Algorithms↗

Molecular evolution of PAS domain-containing proteins of filamentous cyanobacteria through domain shuffling and domain duplication.

When the entire genome of a filamentous heterocyst-forming N2-fixing cyanobacterium, Anabaena sp. PCC 7120 (Anabaena) was determined in 2001, a large number of PAS domains were detected in signal-transducing proteins. The draft genome sequence is also available for the cyanobacterium, Nostoc punctiforme strain ATCC 29133 (Nostoc), that is closely related to Anabaena. In this study, we extracted all PAS domains from the Nostoc genome sequence and analyzed them together with those of Anabaena. Clustering analysis of all the PAS domains gave many specific pairings, indicative of evolutionary conservations. Ortholog analysis of PAS-containing proteins showed composite multidomain architecture in some cases of conserved domains and domains of disagreement between the two species. Further inspection of the domains of disagreement allowed us to trace them back in evolution. Thus, multidomain proteins could have been generated by duplication or shuffling in these cyanobacteria. The conserved PAS domains in the orthologous proteins were analyzed by structural fitting to the known PAS domains. We detected several subclasses with unique sequence features, which will be the target of experimental analysis.

Amino Acid Sequence↗

BIONJ: an improved version of the NJ algorithm based on a simple model of sequence data.

We propose an improved version of the neighbor-joining (NJ) algorithm of Saitou and Nei. This new algorithm, BIONJ, follows the same agglomerative scheme as NJ, which consists of iteratively picking a pair of taxa, creating a new mode which represents the cluster of these taxa, and reducing the distance matrix by replacing both taxa by this node. Moreover, BIONJ uses a simple first-order model of the variances and covariances of evolutionary distance estimates. This model is well adapted when these estimates are obtained from aligned sequences. At each step it permits the selection, from the class of admissible reductions, of the reduction which minimizes the variance of the new distance matrix. In this way, we obtain better estimates to choose the pair of taxa to be agglomerated during the next steps. Moreover, in comparison with NJ's estimates, these estimates become better and better as the algorithm proceeds. BIONJ retains the good properties of NJ--especially its low run time. Computer simulations have been performed with 12-taxon model trees to determine BIONJ's efficiency. When the substitution rates are low (maximum pairwise divergence approximately 0.1 substitutions per site) or when they are constant among lineages, BIONJ is only slightly better than NJ. When the substitution rates are higher and vary among lineages,BIONJ clearly has better topological accuracy. In the latter case, for the model trees and the conditions of evolution tested, the topological error reduction is on the average around 20%. With highly-varying-rate trees and with high substitution rates (maximum pairwise divergence approximately 1.0 substitutions per site), the error reduction may even rise above 50%, while the probability of finding the correct tree may be augmented by as much as 15%.

Algorithms↗

Inferring population history from molecular phylogenies.

Variable molecular sequences sampled from a population can be used to infer its dynamic history. Graphical methods are developed and applied to real data, illustrating ways of navigating through hypothesis space with two landmarks for reference: constant population size and exponentially growing population size.

Genetics, Population↗

Animal evolution and the molecular signature of radiations compressed in time.

The phylogenetic relationships among most metazoan phyla remain uncertain. We obtained large numbers of gene sequences from metazoans, including key understudied taxa. Despite the amount of data and breadth of taxa analyzed, relationships among most metazoan phyla remained unresolved. In contrast, the same genes robustly resolved phylogenetic relationships within a major clade of Fungi of approximately the same age as the Metazoa. The differences in resolution within the two kingdoms suggest that the early history of metazoans was a radiation compressed in time, a finding that is in agreement with paleontological inferences. Furthermore, simulation analyses as well as studies of other radiations in deep time indicate that, given adequate sequence data, the lack of resolution in phylogenetic trees is a signature of closely spaced series of cladogenetic events.

Animals↗

Rapid isolation and sequencing of purified plasmid DNA from Bacillus subtilis.

We report two methods for isolation of plasmid DNA from the gram-positive bacterium Bacillus subtilis. The protoplast alkaline lysis procedure was developed for general use, and the protoplast alkaline lysis magic procedure was developed for isolation of DNA for sequencing. Both procedures yielded large amounts of high-quality DNA in less than 1 h, while current protocols require 4 to 7 h to perform and give lower yields and quality. Plasmid DNA was obtained from strains containing either high- or low-copy-number plasmids. In addition, the procedures were easily adapted to yield large amounts of plasmid DNA suitable for sequencing from another gram-positive organism, Staphylococcus aureus. Further, we demonstrated that neither chloramphenicol, used for plasmid selection, nor the mutation recE4 reduced plasmid DNA yield from the strains we examined.

Bacillus subtilis↗

Structural and biological analysis of integrated polyoma virus DNA and its adjacent host sequences cloned from transformed rat cells.

EcoRI fragments containing integrated viral and adjacent host sequences were cloned from two polyoma virus-transformed cell lines (7axT and 7axB) which each contain a single insert of polyoma virus DNA. Cloned DNA fragments which contained a complete coding capacity for the polyoma virus middle and small T-antigens were capable of transforming rat cells in vitro. Analysis of the flanking sequences indicated that rat DNA had been reorganized or deleted at the sites of polyoma virus integration, but none of the hallmarks of retroviral integration, such as the duplication of host DNA, were apparent. There was no obvious similarity of DNA sequences in the four virus-host joins. In one case the virus-host junction sequence predicted the virus-host fusion protein which was detected in the transformed cell line. DNA homologous to the flanking sequences of three out of four of the joins was present in single copy in untransformed cells. One copy of the flanking host sequences existed in an unaltered form in the two transformed cell lines, indicating that a haploid copy of the viral transforming sequences is sufficient to maintain transformation. The flanking sequences from one cell line were further used as a probe to isolate a target site (unoccupied site) for polyoma virus integration from uninfected cellular DNA. The restriction map of this DNA was in agreement with that of the flanking sequences, but the sequence of the unoccupied site indicated that viral integration did not involve a simple recombination event between viral and cellular sequences. Instead, sequence rearrangements or alterations occurred immediately adjacent to the viral insert, possibly as a consequence of the integration of viral DNA.

Animals↗

Construction and biological analysis of deletion mutants of Fujinami sarcoma virus: 5'-fps sequence has a role in the transforming activity.

Fujinami sarcoma virus (FSV) genome codes for the gag-fps fusion protein FSV-P130. The amino acid sequence of the 3' one-third portion in v-fps is partially homologous to the 3' half of pp60src, or the kinase domain, but the sequence of the 5' portion is unique to v-fps. To identify a possible domain structure in the v-fps sequence responsible for cell transformation, we constructed various deletion mutants of FSV with molecularly cloned viral DNA. Their transforming activities were assayed by measuring focus formation on chicken embryo fibroblasts and rat 3Y1 cells and tumor formation in chickens. The mutants carrying a deletion at the 3' portion in v-fps, the kinase domain, lost transforming activity. The mutants carrying an approximately 1-kilobase deletion within the 5' portion of the v-fps sequence retained focus-forming activity and tumorigenicity in the chicken system, but the efficiency of focus formation was about 10 times lower than that of the wild type. The morphology of these transformed cells was distinct from that observed in cells infected with wild-type FSV. Furthermore, these mutants could not transform rat 3Y1 cells, although wild-type FSV DNA transformed rat 3Y1 cells at a high frequency. The mutants carrying a larger deletion in the 5' portion of fps completely lacked the transforming activity. These results suggest that the 3' portion of the v-fps sequence is necessary but not sufficient for cell transformation and that the 5' portion of v-fps has a role in the transforming activity.

Animals↗

Characterization of an extracellular protease and its cDNA from the nematode-trapping fungus Monacrosporium microscaphoides.

To better exploit the biocontrol potential of nematophagous fungi, it is important to fully understand the molecular background of the infection process. In this paper, several nematode-trapping fungi were surveyed for nematocidal activity. From the culture filtrate of Monacrosporium microscaphoides, a neutral serine protease (designated Mlx) was purified by chromatography. This protease could immobilize the nematode Penagrellus redivivus in vitro and degrade its purified cuticle, suggesting that Mlx could serve as a virulence factor during infection. Characterization of the purified protease revealed a molecular mass of approximately 39 kDa, an isoelectric point of 6.8, and optimum activity at pH 9 at 65 degrees C. Mlx has broad substrate specificity, and it hydrolyzes protein substrates, including casein, skimmed milk, collagen, and bovine serum albumin. The gene encoding Mlx was also cloned and the nucleotide sequence was determined. The deduced amino acid sequence contained the conserved catalytic triad of aspartic acid--histidine--serine and showed high similarity with two cuticle-degrading proteases (PII and Aoz1), which were purified from the nematode-trapping fungus Arthrobotrys oligospora. Research on infection mechanisms of nematode-trapping fungi has thus far only focused on A. oligospora. However, little is known about other nematode-trapping fungi. Our report is among the first to describe the purification and cloning of an infectious protease from a different nematode-trapping fungus.

Animals↗

Making sense of EST sequences by CLOBBing them.

BACKGROUND: Expressed sequence tags (ESTs) are single pass reads from randomly selected cDNA clones. They provide a highly cost-effective method to access and identify expressed genes. However, they are often prone to sequencing errors and typically define incomplete transcripts. To increase the amount of information obtainable from ESTs and reduce sequencing errors, it is necessary to cluster ESTs into groups sharing significant sequence similarity. RESULTS: As part of our ongoing EST programs investigating 'orphan' genomes, we have developed a clustering algorithm, CLOBB (Cluster on the basis of BLAST similarity) to identify and cluster ESTs. CLOBB may be used incrementally, preserving original cluster designations. It tracks cluster-specific events such as merging, identifies 'superclusters' of related clusters and avoids the expansion of chimeric clusters. Based on the Perl scripting language, CLOBB is highly portable relying only on a local installation of NCBI's freely available BLAST executable and can be usefully applied to > 95 % of the current EST datasets. Analysis of the Danio rerio EST dataset demonstrates that CLOBB compares favourably with two less portable systems, UniGene and TIGR Gene Indices. CONCLUSIONS: CLOBB provides a highly portable EST clustering solution and is freely downloaded from: http://www.nematodes.org/CLOBB

Algorithms↗

The interaction of boar sperm proacrosin with its natural substrate, the zona pellucida, and with polysulfated polysaccharides.

Boar sperm acrosin is an acrosomal protease with trypsin-like specificity, and it functions in fertilization by assisting sperm passage through the zona pellucida by limited hydrolysis of this extracellular matrix. In addition to a proteolytic active site domain, acrosin binds the zona pellucida at a separate binding domain that is lost during proacrosin autolysis. In this study, we quantitate the binding of proacrosin to the physiological substrate for acrosin, the zona pellucida, and to a non-substrate, the polysulfated polysaccharide fucoidan. Binding was analogous to sea urchin sperm bindin that binds egg jelly fucan and the vitelline envelope of sea urchin eggs. Proacrosin was found to bind to fucoidan and to the zona pellucida with binding affinities similar to bindin interaction with egg jelly fucan. These interactions were competitively inhibited by similar relative molecular mass polysulfated polymers. Since bindin and proacrosin have distinctly different amino acid sequences, their interaction with acidic sulfate esters demonstrates an example of convergent evolution wherein different macromolecules localized in analogous sperm compartments have the same biological function. From cDNA sequence analysis of proacrosin, this binding may be mediated through a consensus sequence for binding sulfated glycoconjugates. Proacrosin binding to the zona pellucida may serve as both a recognition or primary sperm receptor, as well as maintaining the sperm on the zona pellucida once the acrosome reaction has occurred.

Acrosin↗