PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

The maize genome as a model for efficient sequence analysis of large plant genomes.

The genomes of flowering plants vary in size from about 0.1 to over 100 gigabase pairs (Gbp), mostly because of polyploidy and variation in the abundance of repetitive elements in intergenic regions. High-quality sequences of the relatively small genomes of Arabidopsis (0.14 Gbp) and rice (0.4 Gbp) have now been largely completed. The sequencing of plant genomes that have a more representative size (the mean for flowering plant genomes is 5.6 Gbp) has been seen as a daunting task, partly because of their size and partly because of the numerous highly conserved repeats. Nevertheless, creative strategies and powerful new tools have been generated recently in the plant genetics community, so that sequencing large plant genomes is now a realistic possibility. Maize (2.4-2.7 Gbp) will be the first gigabase-size plant genome to be sequenced using these novel approaches. Pilot studies on maize indicate that the new gene-enrichment, gene-finishing and gene-orientation technologies are efficient, robust and comprehensive. These strategies will succeed in sequencing the gene-space of large genome plants, and in locating all of these genes and adjacent sequences on the genetic and physical maps.

DNA, Plant↗

Structure-function inferences based on molecular modeling, sequence-based methods and biological data analysis of snake venom lectins.

Lectins are a structurally and functionally diverse group of proteins from different sources, capable to recognize and bind specifically carbohydrates. Several snake venoms contain calcium-dependent true lectins (SVLs) that recognize galactose. Herein, in order to enlighten some of the structure-function relationships of snake venom lectins (SVLs), we constructed theoretical models for 10 SVLs based on the Crotalus atrox lectin (CaL), the only SVL crystal structure available, and compared with other animal and plant lectins, and C-type lectin-like proteins (CLPs) that do not bind carbohydrates. Although these are theoretical structures, we could identify some SVL features, including: (i) a singular intrachain disulfide bond (Cys(38)-Cys(133)) that is not present in CLPs; (ii) a significant reorientation (39-41A) of the 80's loop position that folds back to the globular domain, assists the carbohydrate recognition domain (CRD), and orients the dimer formation, even in BfL-1 and BfL-2, which did not present the Cys(86) interchain; (iii) a CRD presenting a negative and concave surface that allows the interaction with the specific saccharide hydroxyl groups and calcium ion; (iv) the role of water molecules in some interchain interactions, similar to other animal and plant lectins; and (v) the inability of forming oligomers in contrast to CaL and some CLPs, such as convulxin.

Amino Acid Sequence↗

Isolation and characterization of TGF-beta 2 and TGF-beta 5 from medium conditioned by Xenopus XTC cells.

TGF-beta 2 and -beta 5 have been purified from medium conditioned by Xenopus cultured cells (XTC) and identified based on their N-terminal amino acid sequence analysis and biological activity. When applied in high concentrations, Xenopus TGF-beta 2, like porcine TGF-beta 2, induces expression of mesodermal markers from cultured Xenopus ectodermal explants, whereas TGF-beta 5 is inactive in this assay. However, the TGF-beta 's could be separated from the major mesoderm-inducing activity present in XTC medium. Xenopus TGF-beta 2 and -beta 5 are approximately equivalent to TGF-beta 1 in their abilities to inhibit the growth of mink lung CCL-64 cells, induce anchorage-independent growth of rat NRK cells, inhibit the proliferation and antibody secretion of human B-lymphocytes, and stimulate chemotaxis of human monocytes. These data establish the functional activity of TGF-beta 5 and suggest that more complex multicellular systems, in contrast to most isolated cells, discriminate between the different TGF-beta s.

Amino Acid Sequence↗

Probabilistic and statistical properties of words: an overview.

In the following, an overview is given on statistical and probabilistic properties of words, as occurring in the analysis of biological sequences. Counts of occurrence, counts of clumps, and renewal counts are distinguished, and exact distributions as well as normal approximations, Poisson process approximations, and compound Poisson approximations are derived. Here, a sequence is modelled as a stationary ergodic Markov chain; a test for determining the appropriate order of the Markov chain is described. The convergence results take the error made by estimating the Markovian transition probabilities into account. The main tools involved are moment generating functions, martingales, Stein's method, and the Chen-Stein method. Similar results are given for occurrences of multiple patterns, and, as an example, the problem of unique recoverability of a sequence from SBH chip data is discussed. Special emphasis lies on disentangling the complicated dependence structure between word occurrences, due to self-overlap as well as due to overlap between words. The results can be used to derive approximate, and conservative, confidence intervals for tests.

Base Sequence↗

SAV, an archaebacterial gene with extensive homology to a family of highly conserved eukaryotic ATPases.

Nucleotide sequencing of a region of the hyperthermophilic archaebacterium Sulfolobus acidocaldarius allowed us to identify an open reading frame of 780 amino acids strikingly similar to a family of eukaryotic ATPases, involved in a variety of biological functions. Sequence analysis of the predicted polypeptide revealed 63 to 66% similarity with S. cerevisiae CDC 48p and its related genes in amphibians (p97ATPase) and mammals (Valosin Containing Protein, VCP), all possibly involved in the regulation of the cell cycle. The finding of an archaebacterial equivalent of these proteins with a high degree of similarity suggests that it represents the same gene in these various species. The new archaebacterial ORF, called SAV (S. acidocaldarius VCP-like) exhibited the usual signature of all members of the family, a highly conserved domain of about 200 amino acids, which is duplicated. Thus, apart from the VCP-like proteins, SAV also appeared similar, although less clearly, to other ATPases, members of the family, involved in vesicle-mediated transport (NSF, Sec18p), peroxysome assembly (PAS1p), and gene expression in yeast (SUG1p) and in human immunodeficiency virus (TBP-1). Finally, the discovery of the archaebacterial gene could enlighten not only the evolutionary relationships between the members of this complex ATPase family, but also the cellular function of these proteins, that is presently obscure.

ATPases Associated with Diverse Cellular Activitie↗

Molecular and insecticidal characterization of a Cry1I protein toxic to insects of the families Noctuidae, Tortricidae, Plutellidae, and Chrysomelidae.

The most notable characteristic of Bacillus thuringiensis is its ability to produce insecticidal proteins. More than 300 different proteins have been described with specific activity against insect species. We report the molecular and insecticidal characterization of a novel cry gene encoding a protein of the Cry1I group with toxic activity towards insects of the families Noctuidae, Tortricidae, Plutellidae, and Chrysomelidae. PCR analysis detected a DNA sequence with an open reading frame of 2.2 kb which encodes a protein with a molecular mass of 80.9 kDa. Trypsin digestion of this protein resulted in a fragment of ca. 60 kDa, typical of activated Cry1 proteins. The deduced sequence of the protein has homologies of 96.1% with Cry1Ia1, 92.8% with Cry1Ib1, and 89.6% with Cry1Ic1. According to the Cry protein classification criteria, this protein was named Cry1Ia7. The expression of the gene in Escherichia coli resulted in a protein that was water soluble and toxic to several insect species. The 50% lethal concentrations for larvae of Earias insulana, Lobesia botrana, Plutella xylostella, and Leptinotarsa decemlineata were 21.1, 8.6, 12.3, and 10.0 microg/ml, respectively. Binding assays with biotinylated toxins to E. insulana and L. botrana midgut membrane vesicles revealed that Cry1Ia7 does not share binding sites with Cry1Ab or Cry1Ac proteins, which are commonly present in B. thuringiensis-treated crops and commercial B. thuringiensis-based bioinsecticides. We discuss the potential of Cry1Ia7 as an active ingredient which can be used in combination with Cry1Ab or Cry1Ac in pest control and the management of resistance to B. thuringiensis toxins.

Amino Acid Sequence↗

Sequence analysis of insecticidal genes from Xenorhabdus nematophilus PMFI296.

Three strains of Xenorhabdus nematophilus showed insecticidal activity when fed to Pieris brassicae (cabbage white butterfly) larvae. From one of these strains (X. nematophilus PMFI296) a cosmid genome library was prepared in Escherichia coli and screened for oral insecticidal activity. Two overlapping cosmid clones were shown to encode insecticidal proteins, which had activity when expressed in E. coli (50% lethal concentration [LC(50)] of 2 to 6 microg of total protein/g of diet). The complete sequence of one cosmid (cHRIM1) was obtained. On cHRIM1, five genes (xptA1, -A2, -B1, -C1, and -D1) showed homology with up to 49% identity to insecticidal toxins identified in Photorhabdus luminescens, and also a smaller gene (chi) showed homology to a putative chitinase gene (38% identity). Transposon mutagenesis of the cosmid insert indicated that the genes xptA2, xptD1, and chi were not important for the expression of insecticidal activity toward P. brassicae. One gene (xptA1) was found to be central for the expression of activity, and the genes xptB1 and xptC1 were needed for full activity. The location of these genes together on the chromosome and therefore present on a single cosmid insert probably accounted for the detection of insecticidal activity in this E. coli clone. Although multiple genes may be needed for full activity, E. coli cells expressing the xptA1 gene from the bacteriophage lambda P(L) promoter were shown to have insecticidal activity (LC(50) of 112 microg of total protein/g of diet). This is contrary to the toxin genes identified in P. luminescens, which were not insecticidal when expressed individually in E. coli. High-level gene expression and the use of a sensitive insect may have aided in the detection of insecticidal activity in the E. coli clone expressing xptA1. The location of these toxin genes and the chitinase gene and the presence of mobile elements (insertion sequence) and tRNA genes on cHRIM1 indicates that this region of DNA represents a pathogenicity island on the genome of X. nematophilus PMFI296.

Animals↗

Molecular cloning of a new crystal protein gene cry1Af1 of Bacillus thuringiensis NT0423 from Korean sericultural farms.

A new cry1Ab-type gene encoding the 130 kDa protein of Bacillus thuringiensis NT0423 bipyramidal crystals was cloned, sequenced, and expressed in a crystal-negative B. thuringiensis host. Hybridization experiments revealed that the crystal protein gene is located on a 44 MDa plasmid of B. thuringiensis NT0423. A strong positive signal detected on the 6.6 kb HindIII fragment from B. thuringiensis NT0423 plasmid DNA was cloned and sequenced. The cry1Ab-type gene, designated cry1Af1, consisted of open reading frame of 3453 bp, encoding a protein of 1151 amino acid residues. The polypeptide has the deduced amino acid sequences predicting molecular masses of 130,215 Da. With both Bt I and Br II promoter sequences were found, the B. thuringiensis NT0423 crystal protein gene promoter closely aligned with those of cry1A-type crystal protein gene. When compared with known sequences of other Cry and Cyt proteins, the Cry1Af1 protein showed maximum 93% sequence identity to Cry1Ab protein of B. thuringiensis subsp. kurstaki. The expressed Cry1Af1 protein in a crystal-negative B. thuringiensis host appears to have strong insecticidal activity against lepidopteran larvae (Plutella xylostella). Crystals containing Cry1Af1 were about six times more toxic than the wild-type crystals of B. thuringiensis NT0423.

Agriculture↗

Detection of length-dependent effects of tandem repeat alleles by 3-D geometric decomposition of craniofacial variation.

Topologically conservative morphological transformations typify the succession of species in the fossil record and also typify more subtle morphological variation within species. Isolation and quantification of morphological variation along its various intermingled modes becomes increasingly difficult as the structures under consideration increase in complexity. Here, we describe a comparative morphometric and genomic study in dogs in which complex three-dimensional craniofacial variation is mathematically distilled into simpler geometric components to test the hypothesis that incremental mutations at developmental loci result in simple geometric deformations of morphology. Combinations of candidate transforms are computationally evaluated for their ability to accurately transform a reference three-dimensional skull model into those of distinct breeds. A set of five simple basis functions are found to be sufficient to describe most craniofacial variation among dogs. Allele lengths of amino acid repeat length variants in developmental regulator genes, which frequently have quantitative effects on phenotype, were compared to geometric terms using Pearson correlation and regression. The coordinated quantitative representation of both phenotype and genotype improves the statistical power for the detection of causative genotype-phenotype relationships and enabled the characterization of the influence of Runx-2 coding repeat length on craniofacial variation among domestic dogs.

Animals↗

Characterization of a vaccinia virus strain used to produce smallpox vaccine in Argentina between 1937 and 1970.

Due to recent political developments, smallpox has re-emerged as a serious threat. We recovered and characterized an old batch of smallpox vaccine, Malbrán strain, produced between 1945 and 1949. The virus was re-isolated and characterized by sequence analysis and biological activity in animals. Phylogenetic analysis using the hemagglutinin and A45R genes showed that the Malbrán strain was closely related to the Lister strain of vaccinia virus. In animals, the Malbrán strain exhibited low pathogenicity, confirming historical records. Mice immunized with the Malbrán strain survived a lethal challenge with cowpox virus. Thus, this strain of vaccinia virus remains a viable candidate as a smallpox vaccine.

Argentina↗

T cell epitope identification for bovine vaccines: an epitope mapping method for BoLA A-11.

T cell responses play an important role in immunity to parasites and other microbial agents of infectious diseases, therefore a number of T cell-directed vaccines are in development. Computer-driven algorithms that facilitate the discovery of T cell epitopes from protein and genome sequences are now being used to accelerate preclinical studies of human vaccines. Similar tools are not yet available for predicting T cell epitopes for animal vaccines, but there may be sufficient data available to begin the process of compiling the algorithms. We describe the construction of a novel mathematical 'matrix' that describes the properties of bovine major histocompatibility complex (BoLA) system antigen (BoLA) A-11 peptide ligands, developed for use with EpiMatrix, an existing T cell epitope-mapping algorithm. An alternative means of developing BoLA matrices, using the pocket profile method, is also discussed. Matrices such as the one described here may be used to develop T cell epitope-mapping tools for cattle and other ruminants. Epitope-mapping algorithms offer a significant advantage over other methods of epitope selection, such as the screening of synthetic overlapping peptides, because high throughput screening can be performed in silico, followed by ex vivo confirmatory studies. Furthermore, using epitope-mapping algorithms, putative T cell epitopes can be derived directly from genomic sequences, allowing researchers to circumvent labor-intensive cloning steps in the genome-to-vaccine discovery pathway.

Algorithms↗

Manuscript evolution.

Frequently, letters, words and sentences are used in undergraduate textbooks and the popular press as an analogy for the coding, transfer and corruption of information in DNA. We discuss here how the converse can be exploited, by using programs designed for biological analysis of sequence evolution to uncover the relationships between different manuscript versions of a text. We point out similarities between the evolution of DNA and the evolution of texts.

Evolution, Molecular↗

Manuscript evolution.

Frequently, letters, words and sentences are used in undergraduate textbooks and the popular press as an analogy for the coding, transfer and corruption of information in DNA. We discuss here how the converse can be exploited, by using programs designed for biological analysis of sequence evolution to uncover the relationships between different manuscript versions of a text. We point out similarities between the evolution of DNA and the evolution of texts.

DNA↗

Cloning and characterization of a Bacillus thuringiensis serovar higo gene encoding a novel class of the delta-endotoxin protein, Cry27A, specifically active on the Anopheles mosquito.

A novel gene encoding a 98-kDa mosquitocidal delta-endotoxin protein, designated Cry27A, was cloned from a Bacillus thuringiensis serovar higo strain. The Cry27A protein contained the five sequence blocks of amino acids commonly conserved in most B. thuringiensis Cry proteins. Relatively high homologies, ranging from 43.0% to 84.4%, existed between the Cry27A protein and several established classes of mosquitocidal Cry proteins (Cry4A, Cry10A, Cry19A, Cry19B, and Cry20A) in the sequence of 51 N-terminal amino acids. The complete sequence of this protein, however, showed low levels (<40%) of amino acid identity to those of the known Cry proteins. Although the expression level of the cry27A gene was low in the transformants under the control of its own promoter, the use of the cyt1A promoter resulted in high-level expression of the gene, leading to the formation of inclusions. The expressed Cry27A protein showed larvicidal activity highly specific for Anopheles stephensi, but lacked the toxicity against Culex pipiens molestus and Aedes aegypti. The results suggest that the Cry27A protein is responsible for the Anopheles-preferential toxicity of the B. thuringiensis serovar higo strain.

Amino Acid Sequence↗

Whole genomes: the foundation of new biology and medicine.

Our genomic DNA sequence provides a unique glimpse of the provenance and evolution of our species, the migration of peoples, and the causation of disease. Understanding the genome may help resolve previously unanswerable questions, including perhaps which human characteristics are innate or acquired. Such an understanding will make it possible to study how genomic DNA sequence varies among populations and among individuals, including the role of such variation in the pathogenesis of important illnesses and responses to pharmaceuticals. The study of the genome and the associated proteomics of free-living organisms will eventually make it possible to localize and annotate every human gene, as well as the regulatory elements that control the timing, organ-site specificity, extent of gene expression, protein levels, and post-translational modifications. For any given physiological process, we will have a new paradigm for addressing its evolution, development, function, and mechanism.

Animals↗

Bassiacridin, a protein toxic for locusts secreted by the entomopathogenic fungus Beauveria bassiana.

A toxic protein, bassiacridin, was purified from a strain of the entomopathogenic fungus Beauveria bassiana isolated from a locust, using chromatographic methods. The final toxic fraction contained between 0.1 and 0.3% of the proteins of the crude extract. Bassiacridin showed no affinity for ion exchangers, was characterised as a monomer with a mol. wt of 60 kDa and an isoelectric point of 9.5, and exhibited beta-glucosidase, beta-galactosidase and N-acetylglucosaminidase activities. Injection of fourth instar nymphs of Locusta migratoria with the pure protein at relatively low dosage (3.3 microg toxin g body wt(-1)) caused a rate of mortality near to 50%. The effects of the crude and pure fractions were characterized at tissular and cellular levels. The formation of melanised spots on tracheae and air sacs and of melanised nodules in contact with the fat body was observed in injected locusts. Alterations of the fine structure of epithelial cells of tracheae, air bags, and integument were also revealed. The insecticidal protein showed a specific activity against locusts. Bassiacridin is different from the other macromolecular toxins of entomopathogenic fungi already described. Microsequencing of peptides generated by trypsic digestion of bassiacridin confirmed that it is a novel molecule and showed that it exhibits a probably limited similarity with a chitin binding protein from yeast.

Animals↗

Detection of new DNA polymerase genes of known and potentially novel herpesviruses by PCR with degenerate and deoxyinosine-substituted primers.

A consensus primer PCR approach was used to (i) investigate the presence of herpesviruses in wild and zoo equids (zebra, wild ass, tapir) and to (ii) study the genetic relationship of the herpesvirus of pigeons (columbid herpesvirus 1) to other herpesvirus species. The PCR assay, based on degenerate primers targeting highly conserved regions of the DNA polymerase gene of herpesviruses, was modified by using a mixture of degenerate and deoxyinosine-substituted primers. The applicability of the modification was validated by amplification of published DNA polymerase genes of 16 herpesvirus species and of the previously uncharacterized DNA polymerase genes of equine herpesvirus 3 (EHV-3) and equine herpesvirus 5 (EHV-5). The modified assay was then used for partial amplification of the polymerase of columbid herpesvirus 1 which is presently classified as a beta-herpesvirus based on biological criteria. Sequence analysis of amplicons obtained from four different viral strains revealed a close relationship of columbid herpesvirus 1 to members of the subfamily Alphaherpesvirinae, especially to Marek's disease herpesvirus. This was confirmed by characterization of additional 1.6kb of the columbid herpesvirus 1 polymerase. Consensus PCR analysis of blood samples from zebras, a wild ass and a tapir revealed amplicons showing high percentages ( > 50%) of sequence identity to DNA polymerases of gamma-herpesviruses. In particular, the zebra and the wild ass sequence were closely related to each other and to the polymerases of the equine gamma-herpesviruses EHV-2 and EHV-5 with sequence identities of > 80%. This is a first indication that novel gamma-herpesviruses are present in wild and zoo equids.

Amino Acid Sequence↗