PubMed HealthSearch

Biomedical subjects

T Etzold

Publications and source records attributed to T Etzold.

11 recordsLinked to original sources

Using views for retrieving data from extremely heterogeneous databanks.

Information in molecular biology or biology in general is contained in a multitude of different databanks each specializing in a certain area. This specialization is useful since it improves the maintainability of the data and the flexibility of the overall structure which until now manages to exist without almost any standards. However, information gathering from these extremely heterogeneous sources is difficult since the desired data may be scattered over many databanks. This paper presents the concept of views implemented in the retrieval system SRS that uses links between databanks and a sophisticated parsing engine for information extraction to provide a flexible way of searching and obtaining data across databank boundaries. The implementation provides homogeneous access to all databanks using two types of views: the list-view which displays selected data-fields in their original format and the table-view available for HTML and plain ASCII text which tries to represent the data in a format independent from that of the source databank.

Amino Acid Sequence

A Tcl-based SRS v. 4 interface.

A new SRS (Sequence Retrieval System) user interface has been developed for SRS v.4. Key features are the support of simple character-oriented (ASCII, VT100) terminals by coding in Tcl augmented by some dedicated Curses calls, support of graphics terminals in an X-Windows version by using the Tk extension to Tcl, and support of a client/server environment by using the TDP extension to Tcl. The Sequence Retrieval System (SRS) is a powerful tool for the fast extraction of information from flat file libraries (Etzold and Argos, 1993) and has rapidly established itself as a major research instrument for the bio-informatics community. Internally the system employs a query language, which is user accessible through either a command-line user interface, 'getz', or a more user friendly, character-oriented window interface. For SRS versions up to release v. 3, this window interface supported VT100-compatible terminals. Because of major changes in the underlying SRS libraries, the v. 3 interface became fully incompatible with the most recent version of SRS (v. 4.x). Thus the many users with only a simple terminal/terminal emulator connection were either deprived of access to SRS, or were forced to use the ASCII WWW client LYNX. This prompted us to develop a character-oriented SRS v. 4 window interface with the look and feel of its SRS v. 3.1 predecessor and coded to be as library independent as possible to maintain compatibility with future SRS releases. In addition, some 'extensions' were coded to widen the applicability to graphics terminals and to a client/server environment. At the time of preparation of this paper, the SRS interface described had been implemented in one form or another on most EM Bnet nodes and on all the platforms given in Table II. The code has been stored at the EMBL in Heidelberg, where it will be available, with installation instructions and scripts, as part of the SRS distribution.

Computer Graphics

An assessment of amino acid exchange matrices in aligning protein sequences: the twilight zone revisited.

The sensitivity of most protein sequence alignment methods depends strongly on the quality of the comparison matrices used. These matrices, which assign weights or similarity scores to every possible amino acid substitution pair, are utilized to differentiate amongst the various possible alignments of two or more sequences. There are many ways to generate these exchange weights and new matrices are constantly published. There has been no overall assessment of these various matrices when applied in different alignment techniques and over many protein folds and families, both close and distant and with the use of several gap penalty values. In this work, a set of amino acid sequences matched by superposition of known protein tertiary topologies is used to test the alignment accuracy of the different method/matrix/penalty combinations. The comparisons show relatively similar results for the top scoring matrices, a preference for the global alignment method of Needleman and Wunsch, and the importance of matrix modification and optimized gap penalties. The relationship between the percentage identity in a resulting alignment and the level of correctness to be expected are given for the top-performing matrix, resulting in a better definition of the so-called "twilight zone". Estimates are made for the probability that two sequences, aligned at a certain level of residue percentage identity, are in fact unrelated.

Amino Acid Sequence

VIP36, a novel component of glycolipid rafts and exocytic carrier vesicles in epithelial cells.

In simple epithelial cells, apical and basolateral proteins and lipids in transit to the cell surface are sorted in the trans-Golgi network. We have recently isolated detergent-insoluble complexes from Madin-Darby canine kidney cells that are enriched in glycosphingolipids, apical cargo and a subset of the proteins of the exocytic carrier vesicles. The vesicular proteins are thought to be involved in protein sorting and include VIP21-caveolin. The vesicular protein VIP36 (36 kDa vesicular integral membrane protein) has been purified from a CHAPS-insoluble residue and a cDNA encoding VIP36 has been isolated. The N-terminal 31 kDa luminal/exoplasmic domain of the encoded protein shows homology to leguminous plant lectins. The transiently expressed protein is localized to the Golgi apparatus, endosomal and vesicular structures and the plasma membrane, as predicted for a protein involved in transport between the Golgi and the cell surface. It is diffusely localized on the plasma membrane but can be redistributed by antibody modulation into caveolae and clathrin-coated pits. We speculate that VIP36 binds to sugar residues of glycosphingolipids and/or glycosylphosphatidyl-inositol anchors and might provide a link between the extracellular/luminal face of glycolipid rafts and the cytoplasmic protein segregation machinery.

Amino Acid Sequence

SRS--an indexing and retrieval tool for flat file data libraries.

SRS (Sequence Retrieval System) is an information indexing and retrieval system designed for libraries with a flat file format such as the EMBL nucleotide sequence databank, the SwissProt protein sequence databank or the Prosite library of protein subsequence consensus patterns. SRS supports the data structure of these libraries by providing special indices for implementing lists of subentities (e.g. feature tables) or hierarchically structured data-fields (e.g. taxonomic classification). A language (ODD) has been designed for the convenient specification of library format and organization, representation of individual data-fields within the system (design of indices) and structuring other data needed during retrieval. This ensures flexibility required for coping with different library formats, which are subject to continuous change. Queries and inspection of retrieved entries can be performed from a user interface with pull-down menus and windows. SRS supports various input and output formats but is particularly well adapted to the GCG programs.

Abstracting and Indexing

Transforming a set of biological flat file libraries to a fast access network.

SRS (Sequence Retrieval System), an indexing system for flat file libraries, provides fast access to individual library entries via retrieval by keywords from various data fields. SRS is now also able to build indices using cross-references that most libraries provide. Fifteen libraries of DNA and protein sequences and structures have been selected. These libraries interact with at least one other by means of cross-references. Indexing these cross-references allows a complete network of libraries to be built. In the network an entry from one library can be linked in principle to every other library. If two libraries are not directly cross-referenced, the linkage can be made with a succession of single links between neighbouring, cross-referenced libraries. A new operator has been added to the query language of SRS for convenient specification of links amongst complete libraries or entry sets generated by previous queries on particular libraries. All the information in the network can now be used to retrieve an entry in a specific library, e.g. the full information given in amino acid sequence entries from SwissProt can now be used to retrieve related tertiary structure entries from PDB. Furthermore, a search in a single library can be extended to a search in the complete library network, e.g. all entries in all databases pertaining to elastase can be found.

Algorithms

Cloning of a growth arrest-specific and transforming growth factor beta-regulated gene, TI 1, from an epithelial cell line.

By cDNA cloning and differential screening, five genes that are regulated by transforming growth factor beta (TGF beta) in mink lung epithelial cells were identified. A novel membrane protein gene, TI 1, was identified which was downregulated by TGF beta and serum in quiescent cells. In actively growing cells, the TI 1 gene is rapidly and transiently induced by TGF beta, and it is overexpressed in the presence of protein synthesis inhibitors. It appears to be related to a family of transmembrane glycoproteins that are expressed on lymphocytes and tumor cells. The four other genes were all induced by TGF beta and correspond to the genes of collagen alpha type I, fibronectin, plasminogen activator inhibitor 1, and the monocyte chemotactic cell-activating factor (JE gene) previously shown to be TGF beta regulated.

Amino Acid Sequence

Point mutations in the 23 S rRNA genes of four lincomycin resistant Nicotiana plumbaginifolia mutants could provide new selectable markers for chloroplast transformation.

Experiments designed to establish stable chloroplast transformation require selectable marker genes encoded by the chloroplast genome. The antibiotic lincomycin is a specific inhibitor of chloroplast ribosomal activity and is known to bind to the large ribosomal subunit. We have investigated a defined region of the chloroplast 23 S rRNA genes from four lincomycin resistant Nicotiana plumbaginifolia mutants and from wild-type N. plumbaginifolia. The mutants LR415, LR421 and LR446 have A to G transitions at positions equivalent to the nucleotides 2058 and 2059 in the Escherichia coli 23 S rRNA. The mutant, LR400, possesses a G to A transition at a position corresponding to nucleotide 2032 of the E. coli 23 S rRNA.

Base Sequence