PubMed Health⌕ Search

Biomedical subjects

Russell Schwartz

Publications and source records attributed to Russell Schwartz.

12 recordsLinked to original sources

Frequencies of hydrophobic and hydrophilic runs and alternations in proteins of known structure.

Patterns of alternation of hydrophobic and polar residues are a profound aspect of amino acid sequences, but a feature not easily interpreted for soluble proteins. Here we report statistics of hydrophobicity patterns in proteins of known structure in a current protein database as compared with results from earlier, more limited structure sets. Previous studies indicated that long hydrophobic runs, common in membrane proteins, are underrepresented in soluble proteins. Long runs of hydrophobic residues remain significantly underrepresented in soluble proteins, with none longer than 16 residues observed. These long runs most commonly occur as buried alpha helices, with extended hydrophobic strands less common. Avoiding aggregation of partially folded intermediates during intracellular folding remains a viable explanation for the rarity of long hydrophobic runs in soluble proteins. Comparison between database editions reveals robustness of statistics on aqueous proteins despite an approximately twofold increase in nonredundant sequences. The expanded database does now allow us to explain several deviations of hydrophobicity statistics from models of random sequence in terms of requirements of specific secondary structure elements. Comparison to prior membrane-bound protein sequences, however, shows significant qualitative changes, with the average hydrophobicity and frequency of long runs of hydrophobic residues noticeably increasing between the database editions. These results suggest that the aqueous proteins of solved structure may represent an essentially complete sample of the universe of aqueous sequences, while the membrane proteins of known structure are not yet representative of the universe of membrane-associated proteins, even by relatively simple measures of hydrophobic patterns.

Computational Biology↗

Evaluating spatial constraints in cellular assembly processes using a monte carlo approach.

Biomolecular behavior commonly involves complex sets of interacting components that are challenging to understand through solution-based chemical theories. Molecular assembly is especially intriguing in the cellular environment because of its links to cell structure in processes such as chemotaxis. We use a coarse-grained Monte Carlo simulation to elucidate the importance of spatial constraints in molecular assembly. We have performed a study of actin filament polymerization through this space-aware probabilistic lattice-based model. Quantitative results are compared with nonspatial models and show convergence over a wide parameter space, but marked divergence over realistic levels corresponding to macromolecular crowding inside cells and localized actin concentrations found at the leading edge during cell motility. These conclusions have direct implications for cell shape and structure, as well as tumor cell migration.

Actin Cytoskeleton↗

Relaxing haplotype block models for association testing.

The arrival of publicly available genome-wide variation data is creating new opportunities for reconciling model-based methods for associating genotypes and phenotypes with the complexities of real genome data. Such data is particularly valuable for testing the utility of models of conserved haplotype structure to association studies. While there is much interest in "haplotype block" models that assume population-wide regions of low diversity, there is also evidence that such models eliminate correlations potentially useful to association studies. We investigate the value of relaxing the rigidity of block models by developing an association testing method using the previously developed "haplotype motif" model, which retains the notion of representing haploid sequences as concatenations of conserved haplotypes but abandons the assumption of population-wide block boundaries. We compare the effectiveness of motif, block, and single-variant models at finding association with simulated phenotypes using real and simulated data. We conclude that the benefits of haplotype models in any form are modest, but that haplotype models in general and block-free models in particular are useful in picking up correlations near the boundaries of the detectable level.

Chromosomes, Human, Pair 22↗

Simulation study of the contribution of oligomer/oligomer binding to capsid assembly kinetics.

The process by which hundreds of identical capsid proteins self-assemble into icosahedral structures is complex and poorly understood. Establishing constraints on the assembly pathways is crucial to building reliable theoretical models. For example, it is currently an open question to what degree overall assembly kinetics are dominated by one or a few most efficient pathways versus the enormous number theoretically possible. The importance of this question, however, is often overlooked due to the difficulties of addressing it in either theoretical or experimental practice. We apply a computer model based on a discrete-event simulation method to evaluate the contributions of nondominant pathways to overall assembly kinetics. This is accomplished by comparing two possible assembly models: one allowing growth to proceed only by the accretion of individual assembly subunits and the other allowing the binding of sterically compatible assembly intermediates any sizes. Simulations show that the two models perform almost identically under low binding rate conditions, where growth is strongly nucleation-limited, but sharply diverge under conditions of higher association rates or coat protein concentrations. The results suggest the importance of identifying the actual binding pattern if one is to build reliable models of capsid assembly or other complex self-assembly processes.

Biophysics↗

Optimal haplotype block-free selection of tagging SNPs for genome-wide association studies.

It is widely hoped that the study of sequence variation in the human genome will provide a means of elucidating the genetic component of complex diseases and variable drug responses. A major stumbling block to the successful design and execution of genome-wide disease association studies using single-nucleotide polymorphisms (SNPs) and linkage disequilibrium is the enormous number of SNPs in the human genome. This results in unacceptably high costs for exhaustive genotyping and presents a challenging problem of statistical inference. Here, we present a new method for optimally selecting minimum informative subsets of SNPs, also known as "tagging" SNPs, that is efficient for genome-wide selection. We contrast this method to published methods including haplotype block tagging, that is, grouping SNPs into segments of low haplotype diversity and typing a subset of the SNPs that can discriminate all common haplotypes within the blocks. Because our method does not rely on a predefined haplotype block structure and makes use of the weaker correlations that occur across neighboring blocks, it can be effectively applied across chromosomal regions with both high and low local linkage disequilibrium. We show that the number of tagging SNPs selected is substantially smaller than previously reported using block-based approaches and that selecting tagging SNPs optimally can result in a two- to threefold savings over selecting random SNPs.

Algorithms↗

Haplotype parsing: methods for extracting information from human genetic variations.

While the shared consensus genetic sequence of our species contains a great deal of information about our common biology, there is also much to be learned from the subtle genetic variations across our species. These variations are believed to be generally of little or no direct functional significance and predominantly reflect the chance accumulation of small genetic changes since our emergence as a species. Therefore, they carry little useful information when observed in a single individual. When tallied across a whole population though, these chance mutations can teach us a great deal about our evolutionary history and the patterns of inheritance in particular individuals. In particular, frequently observed patterns of single nucleotide polymorphisms (SNPs) in a population can identify segments of chromosome that have been passed down largely intact through long stretches of our evolution. Finding these frequently conserved chromosomal segments, or haplotypes, and developing methods to identify haplotype patterns in particular individuals, will in turn help us to identify those particular segments that carry genetic factors influencing risk for many common human diseases. To make the best use of this data, we will need to develop new models for the encoding of information in genome variations--the "language of genetic variation"--and new algorithms for fitting datasets to those models. This article surveys past work by the author and colleagues on this problem, utilising computational methods for locating frequent patterns in haploid sequence data, and "parsing" sequences so as to optimally explain them given the knowledge of the general population structure. The author's recent work in this area has been compiled into a set of computational tools available at http://www-2.cs.cmu.edu/~russells/software/hapmotif.html.

Algorithms↗

Algorithms for association study design using a generalized model of haplotype conservation.

There is considerable interest in computational methods to assist in the use of genetic polymorphism data for locating disease-related genes. Haplotypes, contiguous sets of correlated variants, may provide a means of reducing the difficulty of the data analysis problems involved. The field to date has been dominated by methods based on the "haplotype block" hypothesis, which assumes discrete population-wide boundaries between conserved genetic segments, but there is strong reason to believe that haplotype blocks do not fully capture true haplotype conservation patterns. In this paper, we address the computational challenges of using a more flexible, block-free representation of haplotype structure called the "haplotype motif" model for downstream analysis problems. We develop algorithms for htSNP selection and missing data inference using this more generalized model of sequence conservation. Application to a dataset from the literature demonstrates the practical value of these block-free methods.

Algorithms↗

Understanding actin organization in cell structure through lattice based Monte Carlo simulations.

Understanding the connection between mechanics and cell structure requires the exploration of the key molecular constituents responsible for cell shape and motility. One of these molecular bridges is the cytoskeleton, which is involved with intracellular organization and mechanotransduction. In order to examine the structure in cells, we have developed a computational technique that is able to probe the self-assembly of actin filaments through a lattice based Monte Carlo method. We have modeled the polymerization of these filaments based upon the interactions of globular actin through a probabilistic model encompassing both inert and active proteins. The results show similar response to classic ordinary differential equations at low molecular concentrations, but a bi-phasic divergence at realistic concentrations for living mammalian cells. Further, by introducing localized mobility parameters, we are able to simulate molecular gradients that are observed in nonhomogeneous protein distributions in vivo. The method and results have potential applications in cell and molecular biology as well as self-assembly for organic and inorganic systems.

Actins↗

Robustness of inference of haplotype block structure.

In this report, we examine the validity of the haplotype block concept by comparing block decompositions derived from public data sets by variants of several leading methods of block detection. We first develop a statistical method for assessing the concordance of two block decompositions. We then assess the robustness of inferred haplotype blocks to the specific detection method chosen, to arbitrary choices made in the block-detection algorithms, and to the sample analyzed. Although the block decompositions show levels of concordance that are very unlikely by chance, the absolute magnitude of the concordance may be low enough to limit the utility of the inference. For purposes of SNP selection, it seems likely that methods that do not arbitrarily impose block boundaries among correlated SNPs might perform better than block-based methods.

Algorithms↗

Haplotype motifs: an algorithmic approach to locating evolutionarily conserved patterns in haploid sequences.

The promise of plentiful data on common human genetic variations has given hope that we will be able to uncover genetic factors behind common diseases that have proven difficult to locate by prior methods. Much recent interest in this problem has focused on using haplotypes (contiguous regions of correlated genetic variations), instead of the isolated variations, in order to reduce the size of the statistical analysis problem. In order to most effectively use such variation data, we will need a better understanding of haplotype structure, including both the general principles underlying haplotype structure in the human population and the specific structures found in particular genetic regions or sub-populations. This paper presents a probabilistic model for analyzing haplotype structure in a population using conserved motifs found in statistically significant sub-populations. It describes the model and computational methods for deriving the predicted motif set and haplotype structure for a population. It further presents results on simulated data, in order to validate the method, and on two real datasets from the literature, in order to illustrate its practical application.

Algorithms↗

Epitope prediction algorithms for peptide-based vaccine design.

Peptide-based vaccines, in which small peptides derived from target proteins (eptiopes) are used to provoke an immune reaction, have attracted considerable attention recently as a potential means both of treating infectious diseases and promoting the destruction of cancerous cells by a patient's own immune system. With the availability of large sequence databases and computers fast enough for rapid processing of large numbers of peptides, computer aided design of peptide-based vaccines has emerged as a promising approach to screening among billions of possible immune-active peptides to find those likely to provoke an immune response to a particular cell type. In this paper, we describe the development of three novel classes of methods for the prediction problem. We present a quadratic programming approach that can be trained on quantitative as well as qualitative data. The second method uses linear programming to counteract the fact that our training data contains mostly positive examples. The third class of methods uses sequence profiles obtained by clustering known epitopes to score candidate peptides. By integrating these methods, using a simple voting heuristic, we achieve improved accuracy over the state of the art.

Algorithms↗

Algorithmic strategies for the single nucleotide polymorphism haplotype assembly problem.

With the consensus human genome sequenced and many other sequencing projects at varying stages of completion, greater attention is being paid to the genetic differences among individuals and the abilities of those differences to predict phenotypes. A significant obstacle to such work is the difficulty and expense of determining haplotypes--sets of variants genetically linked because of their proximity on the genome--for large numbers of individuals for use in association studies. This paper presents some algorithmic considerations in a new approach for haplotype determination: inferring haplotypes from localised polymorphism data gathered from short genome 'fragments.' Formalised models of the biological system under consideration are examined, given a variety of assumptions about the goal of the problem and the character of optimal solutions. Some theoretical results and algorithms for handling haplotype assembly given the different models are then sketched. The primary conclusion is that some important simplified variants of the problem yield tractable problems while more general variants tend to be intractable in the worst case.

Algorithms↗