PubMed Health⌕ Search

Biomedical subjects

Sarah K Kummerfeld

Publications and source records attributed to Sarah K Kummerfeld.

7 recordsLinked to original sources

Discrimination of non-protein-coding transcripts from protein-coding mRNA.

Several recent studies indicate that mammals and other organisms produce large numbers of RNA transcripts that do not correspond to known genes. It has been suggested that these transcripts do not encode proteins, but may instead function as RNAs. However, discrimination of coding and non-coding transcripts is not straightforward, and different laboratories have used different methods, whose ability to perform this discrimination is unclear. In this study, we examine ten bioinformatic methods that assess protein-coding potential and compare their ability and congruency in the discrimination of non-coding from coding sequences, based on four underlying principles: open reading frame size, sequence similarity to known proteins or protein domains, statistical models of protein-coding sequence, and synonymous versus non-synonymous substitution rates. Despite these different approaches, the methods show broad concordance, suggesting that coding and non-coding transcripts can, in general, be reliably discriminated, and that many of the recently discovered extra-genic transcripts are indeed non-coding. Comparison of the methods indicates reasons for unreliable predictions, and approaches to increase confidence further. Conversely and surprisingly, our analyses also provide evidence that as much as approximately 10% of entries in the manually curated protein database Swiss-Prot are erroneous translations of actually non-coding transcripts.

Algorithms↗

DBD: a transcription factor prediction database.

Regulation of gene expression influences almost all biological processes in an organism; sequence-specific DNA-binding transcription factors are critical to this control. For most genomes, the repertoire of transcription factors is only partially known. Hitherto transcription factor identification has been largely based on genome annotation pipelines that use pairwise sequence comparisons, which detect only those factors similar to known genes, or on functional classification schemes that amalgamate many types of proteins into the category of 'transcription factor'. Using a novel transcription factor identification method, the DBD transcription factor database fills this void, providing genome-wide transcription factor predictions for organisms from across the tree of life. The prediction method behind DBD identifies sequence-specific DNA-binding transcription factors through homology using profile hidden Markov models (HMMs) of domains. Thus, it is limited to factors that are homologus to those HMMs. The collection of HMMs is taken from two existing databases (Pfam and SUPERFAMILY), and is limited to models that exclusively detect transcription factors that specifically recognize DNA sequences. It does not include basal transcription factors or chromatin-associated proteins, for instance. Based on comparison with experimentally verified annotation, the prediction procedure is between 95% and 99% accurate. Between one quarter and one-half of our genome-wide predicted transcription factors represent previously uncharacterized proteins. The DBD (www.transcriptionfactor.org) consists of predicted transcription factor repertoires for 150 completely sequenced genomes, their domain assignments and the hand curated list of DNA-binding domain HMMs. Users can browse, search or download the predictions by genome, domain family or sequence identifier, view families of transcription factors based on domain architecture and receive predictions for a protein sequence.

Animals↗

Comparative genomics of trypanosomatid parasitic protozoa.

A comparison of gene content and genome architecture of Trypanosoma brucei, Trypanosoma cruzi, and Leishmania major, three related pathogens with different life cycles and disease pathology, revealed a conserved core proteome of about 6200 genes in large syntenic polycistronic gene clusters. Many species-specific genes, especially large surface antigen families, occur at nonsyntenic chromosome-internal and subtelomeric regions. Retroelements, structural RNAs, and gene family expansion are often associated with syntenic discontinuities that-along with gene divergence, acquisition and loss, and rearrangement within the syntenic regions-have shaped the genomes of each parasite. Contrary to recent reports, our analyses reveal no evidence that these species are descended from an ancestor that contained a photosynthetic endosymbiont.

Animals↗

Relative rates of gene fusion and fission in multi-domain proteins.

During evolution genes can produce more complex proteins by gene fusion or less complex proteins by gene fission. Considering proteins from 131 completely sequenced genomes from all three kingdoms of life, we identified 2869 groups of multi-domain proteins as a single protein in certain organisms and as two or more smaller proteins with equivalent domain architectures in other organisms. We found that fusion events are approximately four times more common than fission events, and we established that, in most cases, any particular fusion or fission event only occurred once during the course of evolution.

Artificial Gene Fusion↗

The SUPERFAMILY database in 2004: additions and improvements.

The SUPERFAMILY database provides structural assignments to protein sequences and a framework for analysis of the results. At the core of the database is a library of profile Hidden Markov Models that represent all proteins of known structure. The library is based on the SCOP classification of proteins: each model corresponds to a SCOP domain and aims to represent an entire superfamily. We have applied the library to predicted proteins from all completely sequenced genomes (currently 154), the Swiss-Prot and TrEMBL databases and other sequence collections. Close to 60% of all proteins have at least one match, and one half of all residues are covered by assignments. All models and full results are available for download and online browsing at http://supfam.org. Users can study the distribution of their superfamily of interest across all completely sequenced genomes, investigate with which other superfamilies it combines and retrieve proteins in which it occurs. Alternatively, concentrating on a particular genome as a whole, it is possible first, to find out its superfamily composition, and secondly, to compare it with that of other genomes to detect superfamilies that are over- or under-represented. In addition, the webserver provides the following standard services: sequence search; keyword search for genomes, superfamilies and sequence identifiers; and multiple alignment of genomic, PDB and custom sequences.

Animals↗

Tyramide signal amplification enhances the detectable distribution of connexin-43 positive gap junctions across the ventricular wall of the rabbit heart.

Previous mapping studies examinig the distribution and pattern of staining for connexin-43 expression (the major ventricular gap junction protein) across the ventricular wall have yielded variable findings. The aim of this study was to determine if variations in the distribution of connexin-43 were due to histochemical detection problems, i.e. cross-linking of antigenic sites as a consequence of aldehyde fixation and/or due to low levels of protein expression within the epicardial or endocardial regions of the heart. Immunoperoxidase staining of connexin-43 using the ABC method was carried out in crosssections of rabbit hearts at the level of the papillary muscle. The following treatments were examined: the antibody (Ab) only, Ab with 1/2 Tyramide Signal Amplification (TSA) or full TSA; antibody with microwave antigen retrieval (AR); Ab + 1/2 TSA + AR and finally Ab + TSA + AR. Under light microscopy and using computerized image analysis the percentages of ventricular cross-sectional transmural staining for the different treatment groups were calculated: Ab amounted to only 55%; Ab + 1/2 TSA 63%; Ab + TSA 78%; Ab + AR 72%; Ab + AR + 1/2 TSA 72% and Ab + AR + TSA 88%. The percentages of transumural connexin-43 staining in both TSA + Ab and Ab + TSA + AR groups when compared to Ab only were significantly greater p < 0.01. The antigenic cross-linking due to aldehyde fixation and low levels expression of connexin-43 are contributing factors that influence the immunohistochemical detection of connexin-43 in the mammalian heart. Methodological enhancement for the detection of connexin-43 in this study was derived primarily from amplification of low background levels of connexin-43 being expressed using the TSA protocol. This is supported by the significant differences encountered when TSA was utilized in the protocol and compared with antibody treatment only.

Animals↗

AMID: autonomous modeler of intragenic duplication.

Intragenic duplication is an evolutionary process where segments of a gene become duplicated. While there has been much research into whole-gene or domain duplication, there have been very few studies of non-tandem intragenic duplication. The identification of intragenically replicated sequences may provide insight into the evolution of proteins, helping to link sequence data with structure and function. This paper describes a tool for autonomously modelling intragenic duplication. AMID provides: identification of modularly repetitive genes; an algorithm for identifying repeated modules; and a scoring system for evaluating the modules' similarity. An evaluation of the algorithms and use cases are presented.

Algorithms↗