PubMed HealthSearch

Biomedical subjects

M S Boguski

Publications and source records attributed to M S Boguski.

At least 19 recordsLinked to original sources

Positionally cloned human disease genes: patterns of evolutionary conservation and functional motifs.

Positional cloning has already produced the sequences of more than 70 human genes associated with specific diseases. In addition to their medical importance, these genes are of interest as a set of human genes isolated solely on the basis of the phenotypic effect of the respective mutations. We analyzed the protein sequences encoded by the positionally cloned disease genes using an iterative strategy combining several sensitive computer methods. Comparisons to complete sequence databases and to separate databases of nematode, yeast, and bacterial proteins showed that for most of the disease gene products, statistically significant sequence similarities are detectable in each of the model organisms. Only the nematode genome encodes apparent orthologs with conserved domain architecture for the majority of the disease genes. In yeast and bacterial homologs, domain organization is typically not conserved, and sequence similarity is limited to individual domains. Generally, human genes complement mutations only in orthologous yeast genes. Most of the positionally cloned genes encode large proteins with several globular and nonglobular domains, the functions of some or all of which are not known. We detected conserved domains and motifs not described previously in a number of proteins encoded by disease genes and predicted functions for some of them. These predictions include an ATP-binding domain in the product of hereditary nonpolyposis colon cancer gene (a MutL homolog), which is conserved in the HS90 family of chaperone proteins, type II DNA topoisomerases, and histidine kinases, and a nuclease domain homologous to bacterial RNase D and the 3'-5' exonuclease domain of DNA polymerase I in the Werner syndrome gene product.

Amino Acid Sequence

Positional cloning of the gene for multiple endocrine neoplasia-type 1.

Multiple endocrine neoplasia-type 1 (MEN1) is an autosomal dominant familial cancer syndrome characterized by tumors in parathyroids, enteropancreatic endocrine tissues, and the anterior pituitary. DNA sequencing from a previously identified minimal interval on chromosome 11q13 identified several candidate genes, one of which contained 12 different frameshift, nonsense, missense, and in-frame deletion mutations in 14 probands from 15 families. The MEN1 gene contains 10 exons and encodes a ubiquitously expressed 2.8-kilobase transcript. The predicted 610-amino acid protein product, termed menin, exhibits no apparent similarities to any previously known proteins. The identification of MEN1 will enable improved understanding of the mechanism of endocrine tumorigenesis and should facilitate early diagnosis.

Amino Acid Sequence

GenBank.

The GenBank sequence database incorporates DNA sequences from all available public sources, primarily through the direct submission of sequence data from authors and from large-scale sequencing projects. Data exchange with the EMBL Data Library and the DNA Data Bank of Japan helps ensure comprehensive coverage. GenBank continues to focus on quality control and annotation while expanding data coverage and retrieval services. An integrated retrieval system, known asEntrez, incorporates data from the major DNA and protein sequence databases, along with genome maps and protein structure information. MEDLINE abstracts from published articles describing the sequences are also included as an additional source of biological annotation. Sequence similarity searching is offered through the BLAST family of programs. All of NCBI's services are offered through the World Wide Web. In addition, there are specialized server/client versions as well as FTP and e-mail server access.

Amino Acid Sequence

Genome cross-referencing and XREFdb: implications for the identification and analysis of genes mutated in human disease.

Comparative genomics approaches and multi-organismal biology are valuable tools for genetic analysis. Cross-species connections between genes mutated in human disease states and homologues in model organisms can be particularly powerful, as model-organism gene function data and experimental approaches can shed light on the molecular mechanisms defective in the disease. We describe a project that is systematically identifying novel expressed sequence tag (EST) sequences that are highly related to genes in model organisms and mapping them to positions on the mouse and human maps. This process effectively cross-references model organism genes with mapped mammalian phenotypes, facilitating the identification of genes mutated in human disease states via the positional candidate approach. A public database, XREFdb (http:@www.ncbi.nlm.nih.gov/XREFdb/), disseminates similarity search, mapping and mammalian phenotype information and increases the rate at which these cross-species connections are established.

Animals

A gene map of the human genome.

The human genome is thought to harbor 50,000 to 100,000 genes, of which about half have been sampled to date in the form of expressed sequence tags. An international consortium was organized to develop and map gene-based sequence tagged site markers on a set of two radiation hybrid panels and a yeast artificial chromosome library. More than 16,000 human genes have been mapped relative to a framework map that contains about 1000 polymorphic genetic markers. The gene map unifies the existing genetic and physical maps with the nucleotide and protein sequence databases in a fashion that should speed the discovery of genes underlying inherited human disease. The integrated resource is available through a site on the World Wide Web at http://www.ncbi.nlm.nih.gov/SCIENCE96/.

Amino Acid Sequence

Comparative analysis of 1196 orthologous mouse and human full-length mRNA and protein sequences.

A large set of mRNA and encoded protein sequences, from orthologous murine and human genes, was compiled to analyze statistical, biological, and evolutionary properties of coding and noncoding transcribed sequences. Protein sequence conservation varied between 36% and 100% identity, with an average value of 85%. The average degree of nucleotide sequence identity for the corresponding coding sequences was also approximately 85%, whereas 5' and 3' untranslated regions (UTRs) were less conserved, with aligned identities of 67% and 69%, respectively. For some mouse and human genes, nucleotide sequences are more highly conserved than the encoded protein sequences. A subset of 32 sequences, consisting of only mouse/human protein pairs for which the human sequence represents a positionally cloned disease gene, had properties very similar to the larger data set, suggesting that our data are representative of the genome as a whole. With respect to sequence conservation, two interesting outliers are the breast cancer (BRCAI) gene product and the testis-determining factor (SRY), both of which display among the lowest degrees of sequence identity. The occurrence of both introns and repetitive elements (e.g., Alu, Bl) in 5' and 3' UTRs was also studied. These results provide one benchmark for the "comparative genomics" of mice and humans, with practical implications for the cross-referencing of transcript maps. Also, they should prove useful in estimating the additional sampling diversity provided by mouse EST sequencing projects designed to complement the existing human cDNA collection.

Amino Acid Sequence

Threading analysis suggests that the obese gene product may be a helical cytokine.

The ob gene encodes a protein that, in mutant form, is associated with obesity and type II diabetes in mice. Sequence analysis has revealed no similarities to other proteins, however, and no clues as to possible functions. The possibility nonetheless remains that ob is functionally or ancestrally related to other proteins, whose sequences are divergent to the point that only a comparison of three-dimensional structures might detect relationship. To explore this possibility, we conduct a 'threading' search of a 3-dimensional structure database, to determine whether the ob protein might adopt a fold similar to any known structure. This search reveals that the ob sequence is compatible, at a significance level of P < 0.05, with structures from the family of helical cytokines that includes interleukin-2 and growth hormone. A structural model of ob based upon these results is physically and biologically plausible and leads to testable predictions, including the prediction that ob may activate the JAK-STAT pathway, via binding to a receptor resembling those of the cytokine family.

Amino Acid Sequence

A novel RING finger protein interacts with the cytoplasmic domain of CD40.

CD40 is a member of the tumor necrosis factor receptor family and, like other members, it appears to possess no intrinsic signaling capacity (e.g. kinase activity), suggesting that signal transduction is likely mediated by associating molecules. To identify such molecules, we have utilized the yeast two-hybrid system to clone cDNAs encoding proteins that bind the CD40 cytoplasmic domain. One such interacting protein, designated CD40-binding protein, has a N-terminal RING finger motif that is found in a number of DNA-binding proteins, including the V(D)J recombination activating gene RAG1. In addition, it contains a prominent central coiled-coil segment that may allow homo- or hetero-oligomerization. The C terminus possesses substantial homology to the tumor necrosis factor receptor-associated factor (TRAF) domain that is found in two proteins (TRAF1 and TRAF2) that associate with the cytoplasmic domain of the related 75-kDa tumor necrosis factor receptor. This is the first identification of a molecule that interacts with CD40 and whose sequence suggests a potential role in signaling.

Amino Acid Sequence

Bioinformatics.

Computer databases, networks and software tools are essential materials and methods for biomedical research and are involved in almost every aspect of disease gene mapping and positional cloning. Public databases of DNA and protein sequences and genetic and physical map information are increasing rapidly in size and complexity and are also improving in quality, comprehensiveness, interoperability and access. A new generation of software tools for navigating through the biomedical literature has become available. Programs for sequence homology searching and genetic map construction have become more sophisticated, yet easier to use. Global computer networks are bringing state-of-the-art capabilities to all.

Chromosome Mapping

Issues in searching molecular sequence databases.

Sequence similarity search programs are versatile tools for the molecular biologist, frequently able to identify possible DNA coding regions and to provide clues to gene and protein structure and function. While much attention had been paid to the precise algorithms these programs employ and to their relative speeds, there is a constellation of associated issues that are equally important to realize the full potential of these methods. Here, we consider a number of these issues, including the choice of scoring systems, the statistical significance of alignments, the masking of uninformative or potentially confounding sequence regions, the nature and extent of sequence redundancy in the databases and network access to similarity search services.

Algorithms

Genes conserved in yeast and humans.

Evolutionary conservation of homologous gene products from distantly related organisms provides an information resource of great value for elucidating protein structure and function. Sequence similarities also serve as molecular cross-references between diverse organisms that offer different, or complementary, experimental approaches for analyzing gene expression and biochemistry in normal and abnormal states. There are now countless examples of information about a protein from one species contributing to the understanding of biological phenomena or disease in another species. Such connections are often unanticipated and surprising, but there is an opportunity to make them more systematically as concerted genome sequencing projects progress. In the present review we focus on connections between yeast and human proteins and their functional implications. We present several 'case studies' as well as survey results derived from comprehensive sequence comparisons among all yeast and human proteins currently present in the public databases.

Conserved Sequence