PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Comparison of various algorithms for recognizing short coding sequences of human genes.

MOTIVATION: Since the early 1980s of the twentieth century, there has been great progress in the development of computational gene-finding algorithms. Some problems, however, have not yet been solved currently. Recognizing short genes in prokaryotes and short exons in eukaryotes is one of such problems. The paper is devoted to assessing various algorithms, including those currently available and the new ones proposed here, in order to find the best algorithm to solve the issue. RESULTS: The databases consisting of phase-specific coding and non-coding sequences of human genes with length of 192, 162, 129, 108, 87, 63 and 42 bp, respectively, have been established. Based on the databases and a standard benchmark, 19 algorithms were evaluated, which include the methods of Markov models with orders of 1 through 5, codon usage, hexamer usage, codon preference, amino acid usage, codon prototype, Fourier transform and 8 Z curve methods with various numbers of parameters. Consequently, the Z curve methods with 69 and 189 parameters are the best ones among them, based on the databases constructed here. In addition to the highest recognition accuracy confirmed by 10-fold cross-validation tests, the Z curve methods are much simpler computationally than the second best one, the fifth-order Markov chain model, in which 12 288 parameters are used. We hope that the Z curve methods presented in this paper would be beneficial to the further development of gene-finding algorithms. AVAILABILITY: The programs of various Z curve methods are available on request.

Algorithms↗

BaGGLS: a Bayesian shrinkage framework for interpretable modeling of interactions in high-dimensional biological data.

MOTIVATION: Biological data is often high dimensional, noisy, and governed by complex interactions among sparse signals. This poses major challenges for interpretability and reliable feature selection. Tasks such as identifying motif interactions in genomics exemplify these difficulties, as only a small subset of biologically relevant features (e.g. motifs) are typically active, and their effects are often non-linear and context-dependent. While statistical approaches often result in more interpretable models, deep learning models have proven effective in modeling complex interactions and prediction accuracy, yet their black-box nature limits interpretability. RESULTS: We introduce BaGGLS, a flexible and interpretable probabilistic binary regression model designed for high-dimensional biological inference involving feature interactions. BaGGLS incorporates a Bayesian group global-local shrinkage prior, aligned with the group structure introduced by interaction terms. This prior encourages sparsity while retaining interpretability, helping to isolate meaningful signals and suppress noise. To enable scalable inference, we employ a partially factorized variational approximation that captures posterior skewness and supports efficient learning even in large feature spaces. In extensive simulations, we compare BaGGLS to frequentist probit regressions (unconstrained and with L1-penalty) as well as a probit model with Markov Chain Monte Carlo (MCMC) sampling under a horseshoe prior. We can show that BaGGLS outperforms the other methods with regard to interaction detection and is many times faster than MCMC sampling under the horseshoe prior. We also demonstrate the usefulness of BaGGLS in the context of interaction discovery from motif scanner outputs (e.g. Find Individual Motif Occurrences (FIMO)) and noisy attribution scores from deep learning models. This shows that BaGGLS is a promising approach for uncovering biologically relevant interaction patterns, with potential applicability across a range of high-dimensional tasks in computational biology. AVAILABILITY: Code is available at gitlab.com/dacs-hpi/baggls.

Bayes Theorem↗

Heterologous overexpression of glucose dehydrogenase from the halophilic archaeon Haloferax mediterranei, an enzyme of the medium chain dehydrogenase/reductase family.

The first gene encoding a glucose dehydrogenase (GDH) from a halophilic organism has been sequenced. Amino acid sequence alignments of GDH from Haloferax mediterranei show a high degree of homology with the thermoacidophilic GDHs and with other enzymes from the medium chain dehydrogenase/reductase family. Heterologous overexpression using the mesophilic organism Escherichia coli as the host has been performed and the expression product was obtained as inclusion bodies. To obtain the halophilic enzyme in its native form refolding and reactivation in a saline environment were required. A pure and highly concentrated sample of the enzyme was obtained using a purification procedure based on the protein's halophilicity. This method may be useful as a general procedure for purifying other halophilic proteins from mesophilic hosts.

Amino Acid Sequence↗

Rapid identification of fungal pathogens in BacT/ALERT, BACTEC, and BBL MGIT media using polymerase chain reaction and DNA sequencing of the internal transcribed spacer regions.

We report a direct polymerase chain reaction/sequence (d-PCRS)-based method for the rapid identification of clinically significant fungi from 5 different types of commercial broth enrichment media inoculated with clinical specimens. Media including BacT/ALERT FA (BioMérieux, Marcy l'Etoile, France) (n = 87), BACTEC Plus Aerobic/F (Becton Dickinson, Microbiology Systems, Sparks, MD) (n = 16), BACTEC Peds Plus/F (Becton Dickinson) (n = 15), BACTEC Lytic/10 Anaerobic/F (Becton Dickinson) (n = 11) bottles, and BBL MGIT (Becton Dickinson) (n = 11) were inoculated with specimens from 138 patients. A universal DNA extraction method was used combining a novel pretreatment step to remove PCR inhibitors with a column-based DNA extraction kit. Target sequences in the noncoding internal transcribed spacer regions of the rRNA gene were amplified by PCR and sequenced using a rapid (24 h) automated capillary electrophoresis system. Using sequence alignment software, fungi were identified by sequence similarity with sequences derived from isolates identified by upper-level reference laboratories or isolates defined as ex-type strains. We identified Candida albicans (n = 14), Candida parapsilosis (n = 8), Candida glabrata (n = 7), Candida krusei (n = 2), Scedosporium prolificans (n = 4), and 1 each of Candida orthopsilosis, Candida dubliniensis, Candida kefyr, Candida tropicalis, Candida guilliermondii, Saccharomyces cerevisiae, Cryptococcus neoformans, Aspergillus fumigatus, Histoplasma capsulatum, and Malassezia pachydermatis by d-PCRS analysis. All d-PCRS identifications from positive broths were in agreement with the final species identification of the isolates grown from subculture. Earlier identification of fungi using d-PCRS may facilitate prompt and more appropriate antifungal therapy.

Costs and Cost Analysis↗

Quantification of single nucleotide polymorphisms by automated DNA sequencing.

Single nucleotide polymorphisms (SNPs) are linked to phenotypes associated with diseases and drug responses. Many techniques are now available to identify and quantify such SNPs in DNA or RNA pools, although the information on the latter is limited. The majority of these methodologies require prior knowledge of target sequences, normally obtained through DNA sequencing. Direct quantitation of SNPs from DNA sequencing raw data will save time and money for large amount sample analysis. A high throughput DNA sequencing assay, in combination with a SNP quantitative algorithm, was developed for the quantitation of a SNP present in HCV RNA sequences. For a side-by-side comparison, a Pyrosequencing assay was also developed. Quantitation performance was evaluated for both methods. The direct DNA sequencing quantitation method was shown to be more linear, accurate, sensitive, and reproducible than the Pyrosequencing method for the quantitation of the SNP present in HCV RNA molecules.

Algorithms↗

Detection of hepatitis delta virus recombinants in cultured cells co-transfected with cloned genotypes I and IIb DNA sequences.

It was reported previously that hepatitis delta virus (HDV), the only animal virus in which replication is performed by cellular RNA polymerase(s), undergoes RNA recombination. However, the previous RNA transfection system was somewhat limited in terms of practical application. Cultured cells were transfected with plasmids expressing replication-competent genotypes I and IIb HDV genomic RNAs to develop a better system for studying the fundamental aspects of HDV RNA recombination and HDV-related RNA species were examined using restriction fragment length polymorphisms and sequence analysis of cloned RT-PCR products. This novel experimental system generated efficiently recombinants between the two parental HDV sequences, but not between replication-defective HDV constructs. The genome organization of the HDV recombinants produced in this system resembled that observed previously in cultured cells co-transfected with genome I and IIb RNAs. These data indicate that replication-dependent HDV RNA recombination can be catalyzed by host RNA polymerases in cultured cells co-transfected with two cloned HDV sequences. This new DNA-based system is simpler than the previous RNA-based method of study, and generates a higher recombination frequency, facilitating study of HDV RNA recombination.

Animals↗

Identification of DRB alleles in rhesus monkeys using polymerase chain reaction-sequence-specific primers (PCR-SSP) amplification.

Major histocompatibility complex (MHC) class In molecules play a vital role in the regulation of T-cell functions in the mammalian immune system. Two key features characterize the polymorphism of MHC haplotypes in humans and non-human primates: the existence of a large number of alleles, and the high degree of genetic diversity between those alleles. Rhesus monkeys and Chimpanzees have been extensively used as relevant models for human diseases and transplantation We have investigated DRB genes in 19 macaques, members of 3 families, using polymerase chain reaction with sequence-specific primers (PCR-SSP) and denaturing gradient gel electrophoresis (DGGE). After amplification PCR products were purified and subjected direct sequencing. Seven animals (Madison #1) were typed by DDGE also. We report that the DRB haplotypes defined by PCR-SSP exhibit a high degree of concordance with the data obtained by DGGE and direct sequening. Our data show prominent variability in the number of DRB1 alleles ranging from 1-4 per genotype within these families. This analysis demonstrated that most of the amplicons were identical to Mamu-DRB alleles that our PCR primers were to amplify. However, 98-99% similarity was noticed in the case of Mamu-DRB1*0303, Mamu-DRB6*0103 and Mamu-DRB*W201 alleles. The observed mismatches were located in non-polymorphic regions. Thus, family studies in rhesus macaques performed by molecular methods confirmed the multiplicity of Mamu-DRB1 alleles per haplotype and the existence of allelic associations published earlier. In addition, we propose 3 more DRB allele associations (haplotypes): Mamu-DRB1*04-DRB5*03; Mamu-DRB1*04-*DRB*W5; Mamu-DRB1*04*W2. The proposed medium-resolution PCR-SSP technique appears to be a highly reproducible and discriminatory typing method for detecting polymorphisms of DRB genes in rhesus monkeys.

Alleles↗

Dual-genome primer design for construction of DNA microarrays.

MOTIVATION: Microarray experiments using probes covering a whole transcriptome are expensive to initiate, and a major part of the costs derives from synthesizing gene-specific PCR primers or hybridization probes. The high costs may force researchers to limit their studies to a single organism, although comparing gene expression in different species would yield valuable information. RESULTS: We have developed a method, implemented in the software DualPrime, that reduces the number of primers required to amplify the genes of two different genomes. The software identifies regions of high sequence similarity, and from these regions selects PCR primers shared between the genomes, such that either one or, preferentially, both primers in a given PCR can be used for amplification from both genomes. To assure high microarray probe specificity, the software selects primer pairs that generate products of low sequence similarity to other genes within the same genome. We used the software to design PCR primers for 2182 and 1960 genes from the hyperthermophilic archaea Sulfolobus solfataricus and Sulfolobus acidocaldarius, respectively. Primer pairs were shared among 705 pairs of genes, and single primers were shared among 1184 pairs of genes, resulting in a saving of 31% compared to using only unique primers. We also present an alternative primer design method, in which each gene shares primers with two different genes of the other genome, enabling further savings. 3. AVAILABILITY: The software is freely available at http://www.biotech.kth.se/molbio/microarray/.

Algorithms↗

Rapid differentiation of current infectious bronchitis virus vaccine strains and field isolates in Australia.

OBJECTIVE: Rapid differentiation of vaccine strains of infectious bronchitis virus (IBV) from wild type strains would enhance investigations of disease outbreaks. This study aimed to develop a reverse transcription-polymerase chain reaction (RT-PCR) assay to differentiate between Australian vaccine strains of IBV and field isolates. PROCEDURE: A fragment of 6.5 kilobases that contains the S, M and N genes was amplified by RT-PCR from ten different IBV strains, including vaccine strains and field isolates, and then sequenced. RESULTS: Comparison of the sequences of these strains revealed a deletion of 58 bases in the 3' untranslated region (UTR) of IBV vaccine strains but not in the field isolates. Two primers were designed to amplify a fragment of the 3' UTR that differed in size between the vaccine strains and field isolates. RT-PCR was performed using these two primers to screen 20 IBV strains, including field isolates and the vaccine strains. All strains were correctly identified as either vaccine strains or field isolates. CONCLUSION: This procedure is a rapid, sensitive and inexpensive method for discrimination between most current Australian vaccine strains and field isolates of IBV.

Animals↗

Detection of Blastocystis hominis in unpreserved stool specimens by using polymerase chain reaction.

Blastocystis hominis is a common enteric parasite of worldwide distribution. Its pathogenetic potential has not yet been established, although numerous case reports suggest that B. hominis may cause the development of various gastrointestinal symptoms and disorders. The detection of the parasite in stool specimens is conventionally done by microscopy of direct smears, fecal concentrates, or permanently stained smears; however, morphology-based diagnosis is problematic. The aim of this study was to develop and evaluate a polymerase chain reaction (PCR) technique for the direct detection of B. hominis in human stool samples. Primers were based on small subunit ribosomal DNA and able to detect > or =32 parasites/200 mg stool artificially spiked with cultured B. hominis. In the evaluation of 43 clinical specimens, the PCR was tested against the formol ethyl acetate concentration technique (FECT) and a culture technique, proving 100% test specificity and a significantly higher sensitivity than the FECT. The PCR method is recommended for screening clinical specimens for B. hominis infection and for use in prevalence studies.

Animals↗

Genomic organization of the mouse T-cell receptor beta-chain gene family.

We have combined three different methods, deletion mapping of T-cell lines, field-inversion gel electrophoresis, and the restriction mapping of a cosmid clone, to construct a physical map of the murine T-cell receptor beta-chain gene family. We have mapped 19 variable (V beta) gene segments and the two clusters of diversity (D beta) and joining (J beta) gene segments and constant (C beta) genes. These members of the beta-chain gene family span approximately equal to 450 kilobases of DNA, excluding one potential gap in the DNA fragment alignments.

Animals↗

Prediction of protein secondary structure using the 3D-1D compatibility algorithm.

A new method for the prediction of protein secondary structure is proposed, which relies totally on the global aspect of a protein. The prediction scheme is as follows. A structural library is first scanned with a query sequence by the 3D-1D compatibility method developed before. All the structures examined are sorted with the compatibility score and the top 50 in the list are picked out. Then, all the known secondary structures of the 50 proteins are globally aligned against the query sequence, according to the 3D-1D alignments. Prediction of either alpha helix, beta strand or coil is made by taking the majority among the observations at each residue site. Besides 325 proteins in the structural library, 77 proteins were selected from the latest release of the Brookhaven Protein Data Bank, and they were divided into three data sets. Data set 1 was used as a training set for which several adjustable parameters in the method were optimized. Then, the final form of the method was applied to a testing set (data set 2) which contained proteins of chain length < or = 400 residues. The average prediction accuracy was as high as 69% in the three-state assessment of alpha, beta and coil. On the other hand, data set 3 contains only those proteins of length > 400 residues, for which the present method would not work properly because of the size effect inherent in the 3D-1D compatibility method. The proteins in data set 3 were, therefore, subdivided into constituent domains (data set 4) before being fed into the prediction program. The prediction accuracy for data set 4 was 66% on average, a few percent lower than that for data set 2. Possible causes for this discrepancy are discussed.

Algorithms↗

Kabat Database and its applications: future directions.

The Kabat Database was initially started in 1970 to determine the combining site of antibodies based on the available amino acid sequences. The precise delineation of complementarity determining regions (CDR) of both light and heavy chains provides the first example of how properly aligned sequences can be used to derive structural and functional information of biological macromolecules. This knowledge has subsequently been applied to the construction of artificial antibodies with prescribed specificities, and to many other studies. The Kabat database now includes nucleotide sequences, sequences of T cell receptors for antigens (TCR), major histocompatibility complex (MHC) class I and II molecules, and other proteins of immunological interest. While new sequences are continually added into this database, we have undertaken the task of developing more analytical methods to study the information content of this collection of aligned sequences. New examples of analysis will be illustrated on a yearly basis. The Kabat Database and its applications are freely available at http://immuno.bme.nwu.edu.

Animals↗

Sequence similarity as a predictor of the transmembrane topology of membrane-intrinsic subunits of bacterial respiratory chain enzymes.

Integral membrane proteins usually have a predominantly alpha-helical secondary structure in which transmembrane segments are connected by membrane-extrinsic loops. Although a number of membrane protein structures have been reported in recent years, in most cases transmembrane topologies are initially predicted using a variety of theoretical techniques, including hydropathy analyses and the "positive inside" rule. We have explored the use of plots of the distribution of sequence similarity within families of membrane proteins comprising homeomorphic domains as a new method for the prediction/verification of the orientation of transmembrane topology models within certain families of multimeric respiratory chain enzymes. Within such proteins, analyses of sequence similarity can: i) identify heme and/or quinol binding sites; ii) identify potential electron-transfer conduits to/from prosthetic groups; and iii) locate regions defining potential subunit-subunit interactions. We mined emerging bioinformatic data for sequences of 11 families of membrane-intrinsic proteins that are part of multimeric respiratory chain complexes that also have membrane-extrinsic subunits. The sequences of each family were then aligned and the resultant alignments converted into a graphical format recording an empirical measure of the sequence similarity plotted versus residue position. In each case, this plot was compared to the predicted transmembrane topology. With one exception, there is a strong correlation between the existence

Amino Acid Sequence↗

Improving functional annotation of non-synonomous SNPs with information theory.

Automated functional annotation of nsSNPs requires that amino-acid residue changes are represented by a set of descriptive features, such as evolutionary conservation, side-chain volume change, effect on ligand-binding, and residue structural rigidity. Identifying the most informative combinations of features is critical to the success of a computational prediction method. We rank 32 features according to their mutual information with functional effects of amino-acid substitutions, as measured by in vivo assays. In addition, we use a greedy algorithm to identify a subset of highly informative features. The method is simple to implement and provides a quantitative measure for selecting the best predictive features given a set of features that a human expert believes to be informative. We demonstrate the usefulness of the selected highly informative features by cross-validated tests of a computational classifier, a support vector machine (SVM). The SVM's classification accuracy is highly correlated with the ranking of the input features by their mutual information. Two features describing the solvent accessibility of "wild-type" and "mutant" amino-acid residues and one evolutionary feature based on superfamily-level multiple alignments produce comparable overall accuracy and 6% fewer false positives than a 32-feature set that considers physiochemical properties of amino acids, protein electrostatics, amino-acid residue flexibility, and binding interactions.

Analysis of Variance↗

Molecular epidemiology of 'Norwalk-like viruses' associated with gastroenteritis outbreaks in New Zealand.

Outbreaks of gastroenteritis are a major public health problem in New Zealand. The introduction of molecular detection methods has now shown that the 'Norwalk-like viruses' (NLVs) are the major cause of food and waterborne nonbacterial gastroenteritis. Reverse transcription and polymerase chain reaction (RT-PCR) were used to determine the presence of NLVs in faecal specimens from 83 nonbacterial gastroenteritis outbreaks occurring in New Zealand between August 1995 and July 1999. Further characterisation of the NLVs for epidemiological purposes was carried out by dot blot DNA hybridisation and DNA sequencing of representative outbreak strains. The majority of NLV strains occurring in New Zealand since August 1995 are similar to those occurring overseas. The predominant New Zealand strain is genetically similar to the Bristol/Lordsdale virus group. Several New Zealand outbreaks were attributed to Auckland virus, a Mexico-like NLV strain identified as the most likely cause of gastroenteritis after consumption of contaminated oysters in 1994. A new strain, designated Napier virus, has been identified in six outbreaks since 1996. A number of strains closely resembling internationally recognised strains, including Southampton virus, Saratoga virus; Desert Shield virus and Melksham virus have been associated with gastroenteritis outbreaks across New Zealand. Application of these typing methods has provided information on disease transmission for epidemiological investigations of public health significance.

Amino Acid Sequence↗

Detection and identification of mycoplasmas by amplification of rDNA.

Alignment of published 16S rRNA sequences allowed the definition of a pair of oligonucleotides suitable for polymerase chain reaction (PCR). Using this pair of PCR primers, several mycoplasmas including the four human parasites Mycoplasma genitalium, M. hominis, M. salivarium and M. orale were detected. This DNA amplification was restricted to species of the genus Mycoplasma while no cross-reaction was observed with DNA from other bacteria and eukaryotic cells. Subsequent analysis of amplified products by either specific oligonucleotide hybridization or dideoxy sequencing specified the identity of the detected mycoplasmas. This method offers a highly discriminating and sensitive assay for the direct detection and identification of these microorganisms without the need for prior cultivation.

Base Sequence↗

PKCnu, a new member of the protein kinase C family, composes a fourth subfamily with PKCmu.

Members of the protein kinase C (PKC) family of serine/threonine kinases are thought to play critical roles in the regulation of cellular differentiation and proliferation in many cell types. An additional member of the PKC family was identified through human expressed sequence tag (EST) database search and its full length cDNA was isolated. Sequence analysis revealed that the predicted translation product was composed of 890 amino acid residues and that the protein has 77.3% similarity to human PKC mu (PKCmu) and 77. 4% similarity to mouse PKD (the mouse homolog of PKCmu). We designated the new member as protein kinase C nu (PKCnu). The PKCnu messenger RNA was ubiquitously expressed in various tissues when analyzed by Northern blots and reverse transcriptase-coupled polymerase chain reaction (PCR) analyses. The chromosomal location of the gene was determined between markers WI-9798 and D2S177 on chromosome 2p21 region by PCR-based methods with both a human/rodent monochromosomal hybrid cell panel and a radiation hybrid mapping panel.

Amino Acid Sequence↗