PubMed Health⌕ Search

Biomedical subjects

Alex Pothen

Publications and source records attributed to Alex Pothen.

3 recordsLinked to original sources

Genome prediction of putative genome-linked viral protein (VPg) of astroviruses.

Positive-sense single-stranded RNA (+ssRNA) viruses replicate by uncoating the RNA genome for translation to provide viral proteins essential for genome replication and the production of new viral particles. The viral proteins are synthesized from a polyprotein precursor, which is cleaved nascently. The synthesized proteins include viral RNA-dependent RNA polymerase (RdRP), viral genome-linked protein (VPg), and a helicase. VPg is covalently attached to the genomic form of +ssRNA viruses. Helicases and NTPase unwind the RNA before replication. VPg and helicases have been identified in +ssRNA families, however, the presence of VPg and helicase in the Astroviridae, another +ssRNA family, has not been fully elucidated. Computational tools were utilized to provide sequence analysis evidence for the presence and genomic location of astrovirus VPg and helicase. HMMER program v2.1.1 was used to build Hidden Markov Model (HMM) profile for calicivirus VPg to search for conserved motifs in the astrovirus genome. We performed phylogenetic analysis of two genomic regions of astroviruses and caliciviruses (encoding the RdRP and VPg). We identified a putative VPg coding region in astrovirus. This region was located in open reading frame 1a (ORF1 a) and included sites with high sequence similarity to the VPg coding regions of Caliciviridae, Piconaviridae, and Potyviridae. A region encoding a putative astrovirus helicase identified conserved motifs only with pestivirus helicase sequences. Sequence analysis and comparison to other +ssRNA viruses supports the presence of VPg in the Astroviridae. Further laboratory analysis will be necessary to confirm these findings.

Amino Acid Sequence↗

Computational protein biomarker prediction: a case study for prostate cancer.

BACKGROUND: Recent technological advances in mass spectrometry pose challenges in computational mathematics and statistics to process the mass spectral data into predictive models with clinical and biological significance. We discuss several classification-based approaches to finding protein biomarker candidates using protein profiles obtained via mass spectrometry, and we assess their statistical significance. Our overall goal is to implicate peaks that have a high likelihood of being biologically linked to a given disease state, and thus to narrow the search for biomarker candidates. RESULTS: Thorough cross-validation studies and randomization tests are performed on a prostate cancer dataset with over 300 patients, obtained at the Eastern Virginia Medical School using SELDI-TOF mass spectrometry. We obtain average classification accuracies of 87% on a four-group classification problem using a two-stage linear SVM-based procedure and just 13 peaks, with other methods performing comparably. CONCLUSIONS: Modern feature selection and classification methods are powerful techniques for both the identification of biomarker candidates and the related problem of building predictive models from protein mass spectrometric profiles. Cross-validation and randomization are essential tools that must be performed carefully in order not to bias the results unfairly. However, only a biological validation and identification of the underlying proteins will ultimately confirm the actual value and power of any computational predictions.

Biomarkers, Tumor↗

Protocols for disease classification from mass spectrometry data.

We report our results in classifying protein matrix-assisted laser desorption/ionization-time of flight mass spectra obtained from serum samples into diseased and healthy groups. We discuss in detail five of the steps in preprocessing the mass spectral data for biomarker discovery, as well as our criterion for choosing a small set of peaks for classifying the samples. Cross-validation studies with four selected proteins yielded misclassification rates in the 10-15% range for all the classification methods. Three of these proteins or protein fragments are down-regulated and one up-regulated in lung cancer, the disease under consideration in this data set. When cross-validation studies are performed, care must be taken to ensure that the test set does not influence the choice of the peaks used in the classification. Misclassification rates are lower when both the training and test sets are used to select the peaks used in classification versus when only the training set is used. This expectation was validated for various statistical discrimination methods when thirteen peaks were used in cross-validation studies. One particular classification method, a linear support vector machine, exhibited especially robust performance when the number of peaks was varied from four to thirteen, and when the peaks were selected from the training set alone. Experiments with the samples randomly assigned to the two classes confirmed that misclassification rates were significantly higher in such cases than those observed with the true data. This indicates that our findings are indeed significant. We found closely matching masses in a database for protein expression in lung cancer for three of the four proteins we used to classify lung cancer. Data from additional samples, increased experience with the performance of various preprocessing techniques, and affirmation of the biological roles of the proteins that help in classification, will strengthen our conclusions in the future.

Biomarkers↗