PubMed Health⌕ Search

Biomedical subjects

Duncan A E Cochran

Publications and source records attributed to Duncan A E Cochran.

4 recordsLinked to original sources

Algorithms for sequence analysis via mutagenesis.

MOTIVATION: Despite many successes of conventional DNA sequencing methods, some DNAs remain difficult or impossible to sequence. Unsequenceable regions occur in the genomes of many biologically important organisms, including the human genome. Such regions range in length from tens to millions of bases, and may contain valuable information such as the sequences of important genes. The authors have recently developed a technique that renders a wide range of problematic DNAs amenable to sequencing. The technique is known as sequence analysis via mutagenesis (SAM). This paper presents a number of algorithms for analysing and interpreting data generated by this technique. RESULTS: The essential idea of SAM is to infer the target sequence using the sequences of mutants derived from the target. We describe three algorithms used in this process. The first algorithm predicts the number of mutants that will be required to infer the target sequence with a desired level of accuracy. The second algorithm infers the target sequence itself, using the mutant sequences. The third algorithm assigns quality values to each inferred base. The algorithms are illustrated using mutant sequences generated in the laboratory.

Algorithms↗

Unlocking hidden genomic sequence.

Despite the success of conventional Sanger sequencing, significant regions of many genomes still present major obstacles to sequencing. Here we propose a novel approach with the potential to alleviate a wide range of sequencing difficulties. The technique involves extracting target DNA sequence from variants generated by introduction of random mutations. The introduction of mutations does not destroy original sequence information, but distributes it amongst multiple variants. Some of these variants lack problematic features of the target and are more amenable to conventional sequencing. The technique has been successfully demonstrated with mutation levels up to an average 18% base substitution and has been used to read previously intractable poly(A), AT-rich and GC-rich motifs.

AT Rich Sequence↗

Proteomic analysis of chronic lymphocytic leukemia subtypes with mutated or unmutated Ig V(H) genes.

Chronic lymphocytic leukemia (CLL) is a common hematopoietic malignant disease with variable outcome. CLL has been divided into distinct groups based on whether somatic hypermutation has occurred in the variable region of the immunoglobulin heavy-chain locus or alternatively if the cells express higher levels of the CD38 protein. We have analyzed the proteome of 12 cases of CLL (six mutated (M-CLL) and six unmutated (UM-CLL) immunoglobulin heavy-chain loci; seven CD38-negative and five CD38-positive) using two-dimensional electrophoresis and mass spectrometry. Statistical evaluation using principal component analysis indicated significant differences in patterns of protein expression between the cases with and without somatic mutation. Specific proteins indicated by principal component analysis as varying between the prognostic groups were characterized using mass spectrometry. The levels of F-actin-capping protein beta subunit, 14-3-3 beta protein, and laminin-binding protein precursor were significantly increased in M-CLL relative to UM-CLL. In addition, primary sequence data from tandem mass spectrometry showed that nucleophosmin was present as several protein spots in M-CLL but was not detected in UM-CLL samples, suggesting that several post-translationally modified forms of nucleophosmin vary between these two sample groups. No specific differences were found between CD38-positive and -negative patient samples using the same approach. The results presented show that proteomic analysis can complement other approaches in identifying proteins that may have potential value in the biological and diagnostic distinction between important clinical subtypes of CLL.

14-3-3 Proteins↗

A simulated annealing algorithm for finding consensus sequences.

MOTIVATION: A consensus sequence for a family of related sequences is, as the name suggests, a sequence that captures the features common to most members of the family. Consensus sequences are important in various DNA sequencing applications and are a convenient way to characterize a family of molecules. RESULTS: This paper describes a new algorithm for finding a consensus sequence, using the popular optimization method known as simulated annealing. Unlike the conventional approach of finding a consensus sequence by first forming a multiple sequence alignment, this algorithm searches for a sequence that minimises the sum of pairwise distances to each of the input sequences. The resulting consensus sequence can then be used to induce a multiple sequence alignment. The time required by the algorithm scales linearly with the number of input sequences and quadratically with the length of the consensus sequence. We present results demonstrating the high quality of the consensus sequences and alignments produced by the new algorithm. For comparison, we also present similar results obtained using ClustalW. The new algorithm outperforms ClustalW in many cases.

Algorithms↗