PubMed HealthSearch

Biomedical subjects

S Suhai

Publications and source records attributed to S Suhai.

15 recordsLinked to original sources

X-HUSAR, an X-based graphical interface for the analysis of genomic sequences.

Management and analysis of nucleotide and protein sequence and structure data constitute a traditional area of bioinformatics. Since the analytical programs are frequently developed by researchers, rather than software engineers, they tend to suffer from idiosyncratic and non-ergonomic man-machine interfaces. We report on HUSAR, our 140+ collection of third-party, as well as in-house developed or adapted, sequence manipulation and analysis tools, well integrated into the UNIX operating system environment and accessible via consistent menu-aware interface. Most of the HUSAR programs can be completely specified by UNIX command-line options; they can thus be run in batches or combined into pipes. Adding such a program into the HUSAR environment is almost a 'plug-and-play' exercise. HUSAR has been recently complemented with a graphical client interface, X-HUSAR, to support users on UNIX platforms with X11 windowing systems. The whole X-HUSAR interface is based on a single generic program, COMLIGEN, and a number of specific configuration files. COMLIGEN interprets those files and renders appropriate windows, menus, and other interactive elements, which help the end user in selecting application programs and specifying their options. Efforts of extending both HUSAR and X-HUSAR are roughly linear to the size of the collection.

Animals

A parallel neural network simulator on the connection machine CM-5.

We here present a parallel implementation of artificial neural networks on the connection machine CM-5 and compare it with other parallel implementations on SIMD and MIMD architectures. This parallel implementation was developed with the goal of efficiently training large neural networks with huge training pattern sets for applications in molecular biology, in particular the prediction of coding regions in DNA sequences. The implementation uses training pattern parallelism and makes use of the parallel I/O facilities of the CM-5 and its efficient reduction operations available within the control network to achieve a high scalability. The parallel simulator obtains a maximum speed of 149.25 MCUPS for training feedforward networks with backpropagation on a 512 processor CM-5 system without using the CM-5 vector facility. The implementation poses no restriction on the type of network topology and works with different batch training algorithms like BP. Quickprop and Rprop.

Algorithms

Prediction of hypervariable CDR-H3 loop structures in antibodies.

The structure of the most variable antibody hypervariable loop, CDR-H3, has been predicted from amino acid sequence alone. In contrast to other approaches predictions are made for loop lengths up to 17 residues. The predictions have been achieved using artificial neural networks which are trained on a large set of loops from the Brookhaven Protein Databank which have structures similar to CDR-H3. The loop structures are described by the two backbone dihedral angles phi and psi for each residue. For 21 CDR-H3 loops unique to the neural network, the prediction of dihedral angles leads to an average root mean square deviation in the Cartesian coordinates of 2.65 A. The present method, when combined with existing modelling protocols, provides an important addition to the structural prediction of the complementarity determining regions of antibodies.

Amino Acid Sequence

Computer simulation of antisense DNA containing enantio-deoxynucleotides in the double helix.

Computer modeling of DNA double helices containing L-oligodeoxynucleotides and their complementary D-beta-strands in different orientations and conformations was performed using empirical force-field and semiempirical methods. In particular, the parallel and antiparallel orientations of L-alpha and L-beta-configured strands have been extensively simulated. The parallel oriented double helix with L-alpha-2'-deoxynucleotides has been found to be the enantio-deoxynucleotides. The energy difference between this enantio-DNA and native DNA is approximately 6kcal/mol per base pair, favoring the latter. The theoretical results compared well with preliminary experimental data on the hybridization of recently synthesized 11-mer L-oligodeoxynucleotides, L-d-5' (TpCpGpCpTpGpCpTpTpCpT)3' and L-d-5' (TpCpTpTpCpGpTpCpGpCpT)3' containing the complementary D-beta-oligodeoxynucleotides. As predicted by the calculations, experimental results can be interpreted by assuming that hybridization occurs only in the parallel L-alpha-2'-oligodeoxynucleotides.

Base Sequence

Prototype implementation of the integrated genomic database.

We aim to develop an open software system to handle human genome data. The system, called Integrated Genomic Database (IGD), will integrate information from many genomic databases and experimental resources into a comprehensive target-end database (IGD TED). Users will access front-end client systems (IGD FRED) to download data of interest to their computers and merge them with their own local data. FREDs will provide persistent storage of, and instant access to, retrieved data; a friendly graphical interface; tools for querying, browsing, analyzing, and editing local data; interface to external analysis; and tools for communicating with the outside world. The TED will be accessible over the network (online and offline) as a read-only resource for multiple clients. It collects data from major databases for nucleotide and protein sequences and structures, genome maps, experimental reagents, phenotypes, and bibliographic data, and sets of raw data produced at genome centers and laboratories. Beside character-based access via Gopher, WAIS, FTP, and several query language interfaces to the TED, we will develop a specialized front-end client, IGD FRED, with its own database manager, based on the ACEDB program. The FRED will support graphical display methods for sequence feature maps, chromosomal genetic and physical maps, and experimental objects like clone grids, etc. FRED will also provide an interface to important analysis software packages and tools for submitting data to external databases in their own format.

Computer Communication Networks

Secondary structure of the Arg-Gly-Asp recognition site in proteins involved in cell-surface adhesion. Evidence for the occurrence of nested beta-bends in the model hexapeptide GRGDSP.

The primary sequence Arg-Gly-Asp has been found in a number of proteins which bind to cell surface receptors. Studies with synthetic peptides have shown that the presence of charged side chains alone is not sufficient to confer binding activity. Application of folding algorithms to proteins and peptides having similar sequences indicates that binding activity is strongly correlated with the presence of two or more closely spaced residues that each have a high probability of initiating a beta-bend. Circular dichroic studies on the hexapeptide GRGDSP, whose sequence is contained in fibronectin and which also shows binding activity, demonstrate that it adopts an unusual conformation in aqueous solution. 1H-NMR spectra of the peptide in aqueous solution show that the two amide hydrogens of Asp4 and Ser5 exchange very slowly. Computer-assisted modeling using restrained molecular dynamics and energy minimization results in conformations that include two beta-bends of type III-III or III-I (hydrogen bonds 4----1 and 5----2), fully consistent with constraints imposed by 1H- and 13C-NMR data. It is suggested that this unusual secondary structure provides an additional specificity determinant.

Amino Acid Sequence

Human papillomavirus type 16 DNA sequence.

The complete nucleotide sequence of HPV16 DNA (7904 bp) cloned from an invasive cervical carcinoma was determined. Homology comparisons allowed us to align the major open reading frames with the other published papilloma virus DNA sequences. The general organization of the open reading frames is similar to that of the other four papillomavirus (BPV1, HPV1a, HPV6b, CRPV) already sequenced. The sequence reveals an interruption of the reading frame coding for a suspected E1 protein.

Base Sequence

The nucleotide sequence of the early region of the Tupaia adenovirus DNA corresponding to the oncogenic region E1b of human adenovirus 7.

The nucleotide sequence of the early region E1b of the tree shrew (Tupaia) adenovirus (TAV) DNA has been determined. The sequenced region includes the genes for polypeptides of Mr 15 000, 44 000 and 13 400, which are analogous to the small and large E1b proteins and protein IX, respectively, of the three human adenovirus serotypes 5, 7, and 12. The hexanucleotide consensus signal AATAAA occurs only at the 3' terminus of the gene for protein IX suggesting that the E1 region of TAV encompasses one transcription unit. The amino acid sequences of the TAV polypeptides have a higher degree of homology to those of Ad7 and Ad5 than to those of Ad12.

Adenoviridae

DNA sequence and genome organization of genital human papillomavirus type 6b.

The complete nucleotide sequence of the circular double-stranded DNA of the genital human papillomavirus type 6b (HPV6b) comprising 7902 bp was determined and compared with the DNA sequences of human papillomavirus type 1a (HPV1a) and bovine papillomavirus type 1 (BPV1). All major open reading frames are located on one DNA strand only. Their arrangement reveals that the genomic organization of HPV6b is similar to that of HPV1a and BPV1. The putative early region includes two large open reading frames E1 and E2 with marked amino acid sequence homologies to HPV1a and BPV1 which are flanked by several smaller frames. The internal part of E2 completely overlaps with another open reading frame E4. The putative late region contains two large open reading frames L1 and L2. The L1 amino acid sequences are highly conserved among analyzed papillomavirus types. By sequence comparison, potential promoter, splicing and polyadenylation signals can be localized in HPV6b DNA suggesting possible mechanisms of genital papillomavirus gene expression.

Amino Acid Sequence

Energy bands and charge transfer in proteins.

The effects of salts on protein--the causing of a shift in isoelectric point and the altering of the melting temperature--are proposed to be the result of binding to the protein peptide chain, which is considered as a one-dimensional solid. The interaction of methylglyoxal with protein and polylysine to give charge-transfer complexes and allow electrical conductivity are viewed as further support for the band structure of proteins. Calculations on protein chains resembling real proteins show that conductivity should be much less than expected for homopolypeptides.

Electric Conductivity