PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

[Bioinformatic analysis of Vibrio parahaemolyticus thermolabile hemolysin gene].

OBJECTIVE: To carry out bioinformatic analysis of Vibrio parahaemolyticus thermolabile hemolysin gene (tlh) obtained by PCR amplification. METHODS: The tlh gene amplified by PCR was cloned into the vector pET32a(+) and sequenced, followed by analysis of the biological information by with presenting the sequences to the websites of bioinformatics on the Internet. RESULTS AND CONCLUSIONS: The sequenced tlh gene (named tlh14-90) was entered into GenBank with the accession number of AY289609. Tlh14-90 has a length of 1 257 bp with both start and stop codons, having 99% homology with the tlh gene of WP1. Tlh14-90 is predicted to encode a protein containing 418 amino acids (named TLH14-90, with the molecular formula of C(2131)H(3184)N(548)O(649)S(16), molecular weight of 47 392.9, and the theoretical PI of 4.92). This protein consists of 4 Cys and the contents of Ala, Leu and Asn are 11.0%, 7.4% and 7.2%, respectively, having a hydrophobic parameter of -34 calculated using kappa-D method. Predicted as a alpha/protein for its secondary structure, TLH14-90 has not been identified for its tertiary structure.

Bacterial Proteins↗

SIRW: A web server for the Simple Indexing and Retrieval System that combines sequence motif searches with keyword searches.

SIRW (http://sirw.embl.de/) is a World Wide Web interface to the Simple Indexing and Retrieval System (SIR) that is capable of parsing and indexing various flat file databases. In addition it provides a framework for doing sequence analysis (e.g. motif pattern searches) for selected biological sequences through keyword search. SIRW is an ideal tool for the bioinformatics community for searching as well as analyzing biological sequences of interest.

Abstracting and Indexing↗

Clustering of Schistosoma mansoni mRNA sequences and analysis of the most transcribed genes: implications in metabolism and biology of different developmental stages.

The study of the Schistosoma mansoni genome, one of the etiologic agents of human schistosomiasis, is essential for a better understanding of the biology and development of this parasite. In order to get an overview of all S. mansoni catalogued gene sequences, we performed a clustering analysis of the parasite mRNA sequences available in public databases. This was made using softwares PHRAP and CAP3. The consensus sequences, generated after the alignment of cluster constituent sequences, allowed the identification by database homology searches of the most expressed genes in the worm. We analyzed these genes and looked for a correlation between their high expression and parasite metabolism and biology. We observed that the majority of these genes is related to the maintenance of basic cell functions, encoding genes whose products are related to the cytoskeleton, intracellular transport and energy metabolism. Evidences are presented here that genes for aerobic energy metabolism are expressed in all the developmental stages analyzed. Some of the most expressed genes could not be identified by homology searches and may have some specific functions in the parasite.

Amino Acid Sequence↗

Improved method for preparation of lipopolysaccharide-binding protein from human serum by electrophoretic and chromatographic separation techniques.

Recent work has established the importance of serum proteins which interact with endotoxin (lipopolysaccharide, LPS) from Gram-negative bacteria. Thus human monocytes are activated after binding LPS complexed with a serum protein. LPS-binding protein (LBP) is a protein present in both normal and acute phase sera which binds LPS with high affinity. We describe the purification of LBP from human acute phase serum. The purification procedures combine preparative isoelectric focusing (IEF) and either preparative polyacrylamide gel electrophoresis (PAGE) or alternatively an anion-exchange chromatographic step using a Mono Q HR 5/5 column. This allows the isolation of biologically active LBP. LBP was characterized by N-terminal sequence analysis and by measuring the biological activity using flow cytometry (fluorescence-activated cell sorter, FACS) and a luminol enhanced chemiluminescence (LECL) assay.

Acute-Phase Proteins↗

Sequence analysis of heparan sulphate and heparin oligosaccharides.

The biological activity of heparan sulphate (HS) and heparin largely depends on internal oligosaccharide sequences that provide specific binding sites for an extensive range of proteins. Identification of such structures is crucial for the complete understanding of glycosaminoglycan (GAG)-protein interactions. We describe here a simple method of sequence analysis relying on the specific tagging of the sugar reducing end by 3H radiolabelling, the combination of chemical scission and specific enzymic digestion to generate intermediate fragments, and the analysis of the generated products by strong-anion-exchange HPLC. We present full sequence data on microgram quantities of four unknown oligosaccharides (three HS-derived hexasaccharides and one heparin-derived octasaccharide) which illustrate the utility and relative simplicity of the technique. The results clearly show that it is also possible to read sequences of inhomogeneous preparations. Application of this technique to biologically active oligosaccharides should accelerate progress in the understanding of HS and heparin structure-function relationships and provide new insights into the primary structure of these polysaccharides.

Animals↗

Abundantly and rarely expressed Lhc protein genes exhibit distinct regulation patterns in plants.

We have analyzed gene regulation of the Lhc supergene family in poplar (Populus spp.) and Arabidopsis (Arabidopsis thaliana) using digital expression profiling. Multivariate analysis of the tissue-specific, environmental, and developmental Lhc expression patterns in Arabidopsis and poplar was employed to characterize four rarely expressed Lhc genes, Lhca5, Lhca6, Lhcb7, and Lhcb4.3. Those genes have high expression levels under different conditions and in different tissues than the abundantly expressed Lhca1 to 4 and Lhcb1 to 6 genes that code for the 10 major types of higher plant light-harvesting proteins. However, in some of the datasets analyzed, the Lhcb4 and Lhcb6 genes as well as an Arabidopsis gene not present in poplar (Lhcb2.3) exhibited minor differences to the main cooperative Lhc gene expression pattern. The pattern of the rarely expressed Lhc genes was always found to be more similar to that of PsbS and the various light-harvesting-like genes, which might indicate distinct physiological functions for the rarely and abundantly expressed Lhc proteins. The previously undetected Lhcb7 gene encodes a novel plant Lhcb-type protein that possibly contains an additional, fourth, transmembrane N-terminal helix with a highly conserved motif. As the Lhcb4.3 gene seems to be present only in Eurosid species and as its regulation pattern varies significantly from that of Lhcb4.1 and Lhcb4.2, we conclude it to encode a distinct Lhc protein type, Lhcb8.

Amino Acid Sequence↗

Discovery and profiling of bovine microRNAs from immune-related and embryonic tissues.

MicroRNAs are small approximately 22 nucleotide-long noncoding RNAs capable of controlling gene expression by inhibiting translation. Alignment of human microRNA stem-loop sequences (mir) against a recent draft sequence assembly of the bovine genome resulted in identification of 334 predicted bovine mir. We sequenced five tissue-specific cDNA libraries derived from the small RNA fractions of bovine embryo, thymus, small intestine, and lymph node to validate these predictions and identify new mir. This strategy combined with comparative sequence analysis identified 129 sequences that corresponded to mature microRNAs (miR). A total of 107 sequences aligned to known human mir, and 100 of these matched expressed miR. The other seven sequences represented novel miR expressed from the complementary strand of previously characterized human mir. The 22 sequences without matches displayed characteristic mir secondary structures when folded in silico, and 10 of these retained sequence conservation with other vertebrate species. Expression analysis based on sequence identity counts revealed that some miR were preferentially expressed in certain tissues, while bta-miR-26a and bta-miR-103 were prevalent in all tissues examined. These results support the premise that species differences in regulation of gene expression by miR occur primarily at the level of expression and processing.

Animals↗

Protocol to predict gene expression from transcriptomic data using PREDICT.

Linking DNA sequence variation to context-specific transcriptional programs is a critical challenge in regulatory genomics, especially for non-model organisms. Here, we present PREDICT, a modular Python package for discovering cis-regulatory elements and transcription factor binding motifs. We describe steps to identify enriched k-mers from differentially expressed genes, map them to known motifs, quantify their impact on gene expression, and visualize motif co-occurrences. PREDICT provides a robust, k-mer-based approach to uncover regulatory logic in diverse genomic systems. For complete details on the use and execution of this protocol, please refer to Yen et al. and Liu et al.1,2.

Gene Expression Profiling↗

Learning interpretable SVMs for biological sequence classification.

BACKGROUND: Support Vector Machines (SVMs)--using a variety of string kernels--have been successfully applied to biological sequence classification problems. While SVMs achieve high classification accuracy they lack interpretability. In many applications, it does not suffice that an algorithm just detects a biological signal in the sequence, but it should also provide means to interpret its solution in order to gain biological insight. RESULTS: We propose novel and efficient algorithms for solving the so-called Support Vector Multiple Kernel Learning problem. The developed techniques can be used to understand the obtained support vector decision function in order to extract biologically relevant knowledge about the sequence analysis problem at hand. We apply the proposed methods to the task of acceptor splice site prediction and to the problem of recognizing alternatively spliced exons. Our algorithms compute sparse weightings of substring locations, highlighting which parts of the sequence are important for discrimination. CONCLUSION: The proposed method is able to deal with thousands of examples while combining hundreds of kernels within reasonable time, and reliably identifies a few statistically significant positions.

Algorithms↗

Nonrandom sequence of slope-intercept estimates in longitudinal gompertzian analysis suggests biological relevance.

The significance of intersections in age-specific mortality rate distributions could be attributed to a fundamental statistical relationship between estimates of slope and intercept. A strong negative correlation between estimates of slope and intercept is often observed in simple linear regression problems. The net result is that the family of lines generated by repetitive estimates of slope and intercept in a static experimental situation will tend to intersect at a common point. This statistical relationship between slope and intercept, however, should be random with respect to the time-ordered sequence of slope and intercept estimates. Annual paired slope and intercept estimates derived using the method of longitudinal Gompertzian analysis of age-specific mortality rates for men and women in the United States from 1900 to 1988 were analyzed to determine if they varied randomly. The probability that the observed sequence in these annual paired slope-intercept estimates was random is less than 10(-50) for both men and women. This finding essentially excludes the possibility that intersections in age-specific mortality rate distributions reflect a fundamental statistical relationship between slope and intercept and further suggests biological relevance for the method of longitudinal Gompertzian analysis.

Female↗

Nucleotide sequence of the type A staphylococcal enterotoxin gene.

We determined the nucleotide sequence of the gene encoding staphylococcal enterotoxin A (entA). The gene, composed of 771 base pairs, encodes an enterotoxin A precursor of 257 amino acid residues. A 24-residue N-terminal hydrophobic leader sequence is apparently processed, yielding the mature form of staphylococcal enterotoxin A (Mr, 27,100). Mature enterotoxin A has 82, 72, 74, and 34 amino acid residues in common with staphylococcal enterotoxins B and C1, type A streptococcal exotoxin, and toxic shock syndrome toxin 1, respectively. This level of homology was determined to be significant based on the results of computer analysis and biological considerations. DNA sequence homology between the entA gene and genes encoding other types of staphylococcal enterotoxins was examined by DNA-DNA hybridization analysis with probes derived from the entA gene. A 624-base-pair DNA probe that represented an internal fragment of the entA gene hybridized well to DNA isolated from EntE+ strains and some EntA+ strains. In contrast, a 17-base oligonucleotide probe that encoded a peptide conserved among staphylococcal enterotoxins A, B, and C1 hybridized well to DNA isolated from EntA+, EntB+, EntC1+, and EntD+ strains. These hybridization results indicate that considerable sequence divergence has occurred within this family of exotoxins.

Amino Acid Sequence↗

Two new SINE elements, p-SINE2 and p-SINE3, from rice.

p-SINE1 was the first plant SINE element identified in the Waxy gene in Oryza sativa, and since then a large number of p-SINE1-family members have been identified from rice species with the AA or non-AA genome. In this paper, we report two new rice SINE elements, designated p-SINE2 and p-SINE3, which form distinct families from that of p-SINE1. Each of the two new elements is significantly homologous to p-SINE1 in their 5'-end regions with that of the polymerase III promoter (A box and B box), but not significantly homologous in the 3'-end regions, although they all have a T-rich tail at the 3' terminus. Despite the three elements sharing minimal homology in their 3'-end regions, the deduced RNA secondary structures of p-SINE1, p-SINE2 and p-SINE3 were found to be similar to one another, such that a stem-loop structure seen in the 3'-end region of each element is well conserved, suggesting that the structure has an important role on the p-SINE retroposition. These findings suggest that the three p-SINE elements originated from a common ancestor. Similar to members of the p-SINE1 family, the members of p-SINE2 or p-SINE3 are almost randomly dispersed in each of the 12 rice chromosomes, but appear to be preferentially inserted into gene-rich regions. The p-SINE2 members were present at respective loci not only in the strains of the species with the AA genome in the O. sativa complex, but also in those of other species with the BB, CC, DD, or EE genome in the O. officinalis complex. The p-SINE3 members were, however, only present in strains of species in the O. sativa complex. These findings suggest that p-SINE2 originated in an ancestral species with the AA, BB, CC, DD and EE genomes, like p-SINE1, whereas p-SINE3 originated in an ancestral strain of the species with the AA genome. The nucleotide sequences of p-SINE1 members are more divergent than those of p-SINE2 or p-SINE3, indicating that p-SINE1 is likely to be older than p-SINE2 and p-SINE3. This suggests that p-SINE2 and p-SINE3 have been derived from p-SINE1.

Base Sequence↗

A comprehensive set of sequence analysis programs for the VAX.

The University of Wisconsin Genetics Computer Group (UWGCG) has been organized to develop computational tools for the analysis and publication of biological sequence data. A group of programs that will interact with each other has been developed for the Digital Equipment Corporation VAX computer using the VMS operating system. The programs available and the conditions for transfer are described.

Base Sequence↗

Characterization and partial genome sequence analysis of Clostera anachoreta granulovirus.

The morphological and biological properties as well as partial genomic sequencing of a granulovirus isolated from Clostera anachoreta (Lepidoptera: Notodontidae), C. anachoreta granulovirus (ClanGV), were carried out. The ovoidal occlusion bodies were 337 nm x 170 nm in size, and each granule contained one single rod-shape virion, with a mean size of 250 nm x 46 nm. Granulin had a molecular weight of approximately 30 kDa. ClanGV genome size was estimated as 104.34 kb based on the restriction fragments. The restriction pattern of the ClanGV genome was different from other GVs. A restriction fragment genomic library of ClanGV genome was constructed. The library consisted of nine SalI fragments, seven HindIII fragments and seven EcoRI fragments. One 4.8 kb fragment of the genome, digested by SalI, was sequenced and analyzed. This region was composed of eight unknown ORFs, two baculoviruses homologous gene (vp1054 and lef10) and partial sequence of lef-8. The unknown ORFs included three unique to ClanGV, the other five ORFs were related to baculoviruses. The ORFs, located within this restriction fragment, were compared to homologues in other GVs. The results indicated that ClanGV, CpGV, ClGV, AoGV and PoGV had similar arrangement and orientation of the homologous ORFs. Phylogenetic analysis of VP1054 proteins from 20 baculoviruses indicated that ClanGV was more closely related to CpGV, ClGV, AoGV and PoGV than to other baculoviruses.

Amino Acid Sequence↗

Parallel molecular genetic analysis.

We describe recent progress in parallel molecular genetic analyses using DNA microarrays, gel-based systems, and capillary electrophoresis and utilization of these approaches in a variety of molecular biology assays. These applications include use of polymorphic markers for mapping of genes and disease-associated loci and carrier detection for genetic diseases. Application of these technologies in molecular diagnostics as well as fluorescent technologies in DNA analysis using immobilized oligonucleotide arrays on silicon or glass microchips are discussed. The array-based assays include sequencing by hybridization, cDNA expression profiling, comparative genome hybridization and genetic linkage analysis. Developments in non microarray-based, parallel analyses of mutations and gene expression profiles are reviewed. The promise of and recent progress in capillary array electrophoresis for parallel DNA sequence analysis and genotyping is summarized. Finally, a framework for decision making in selecting available technology options for specific molecular genetic analyses is presented.

DNA↗

Human and murine interleukin 1 possess sequence and structural similarities.

The molecular cloning and sequence analysis for human and murine interleukin 1 precursor have recently been described. Comparison of the amino acid sequences resulting from these data can be used to aid in the identification of conserved regions essential to biological activity. We report results which confirm the relationship between these two molecules and suggest that specific regions may be essential for activity. Amino terminal sequence analysis of a 19,000 Mr biologically active IL-1 isolated from stimulated human monocytes reveals a sequence which is in good agreement with that inferred from the human cDNA and, furthermore, locates the processed amino terminus at a site similar to that described for the murine sequence.

Amino Acid Sequence↗

Genome and proteome analysis of Chlamydia.

It has been difficult to study the molecular biology of the obligate intracellular bacterium Chlamydia due to lack of genetic transformation systems. Therefore, genome sequencing has greatly expanded the information concerning the biology of these pathogens. Comparing the genomes of seven sequenced Chlamydia genomes has provided information of the common gene content and gene variation. In addition, the genome sequences have enabled global investigation of both transcript and protein content during the developmental cycle of chlamydiae. During this cycle Chlamydia alternates between an infectious extracellular form and an intracellular dividing form surrounded by a phagosome membrane termed the chlamydial inclusion. Proteins secreted from the chlamydial inclusion into the host cell may interact with host cell proteins and modify the host cell's response to infection. However, identification of such proteins has been difficult because the host cell cytoplasm of Chlamydia infected cells cannot be purified. This problem has been circumvented by comparative proteomics.

Bacterial Physiological Phenomena↗