PubMed Health⌕ Search

Biomedical subjects

E Barillot

Publications and source records attributed to E Barillot.

At least 19 recordsLinked to original sources

XML, bioinformatics and data integration.

MOTIVATION: The eXtensible Markup Language (XML) is an emerging standard for structuring documents, notably for the World Wide Web. In this paper, the authors present XML and examine its use as a data language for bioinformatics. In particular, XML is compared to other languages, and some of the potential uses of XML in bioinformatics applications are presented. The authors propose to adopt XML for data interchange between databases and other sources of data. Finally the discussion is illustrated by a test case of a pedigree data model in XML. CONTACT: Emmanuel.Barillot@infobiogen.fr

Computational Biology↗

The art of pedigree drawing: algorithmic aspects.

MOTIVATION: Giving a meaningful representation of a pedigree is not obvious when it includes consanguinity loops, individuals with multiple mates or several related families. RESULTS: We show that finding a perfectly meaningful representation of a pedigree is equivalent to the interval graph sandwich problem and we propose an algorithm for drawing pedigrees.

Algorithms↗

DBcat: a catalog of 500 biological databases.

The DBcat (http://www.infobiogen.fr/services/dbcat ) is a comprehensive catalog of biological databases, maintained and curated at Infobiogen. It contains 500 databases classified by application domains. The DBcat is a structured flat-file library, that can be searched by means of an SRS server or a dedicated Web interface. The files are available for download from Infobiogen anonymous ftp server.

Biology↗

XML: a lingua franca for science?

XML is a new language designed to solve one of the biggest problems of the World Wide Web: its main language, HTML, is not extensible. In this article, the authors discuss the current successes and limitations of the World Wide Web, briefly explain the basics of XML and present the benefits of using XML as a data-exchange language. Finally, they discuss real-life applications that have been developed using XML, with a focus on biology.

Internet↗

MappetShow: non-linear visualization for genome data.

The genome mapping projects now produce very dense maps with up to several thousands of markers per chromosome. Besides synteny plays a increasing role in mapping: enrichment of poor maps from the maps of close genomes (in terms of evolution) is a high-reward task. We propose a map viewer adapted to this situation: MappetShow gives a clear view of very dense maps and compares efficiently several maps. MappetShow is based on non-linear viewing and is written in Java. A map description language isolates the software from the data sources. This software was easily used on data coming from as different sources as an Object Request Broker, an Object-Oriented Database, or a flat data stream. MappetShow can be browsed at the URL http:¿www.infobiogen.fr/services/Mappet. More generally we discuss how to use the non-linear viewing concept in molecular biology data visualization.

Chromosome Mapping↗

DBcat: a catalog of biological databases.

The DBcat (http://www.infobiogen.fr/services/dbcat) is a comprehensive catalog of biological databases, maintained and curated on a daily basis at GIS Infobiogen. It contains more than 400 databases classified by application domains. The DBcat is a structured flat file library, that can be searched by means of an SRS server or a dedicated Web interface. The files are available for downloading from Infobiogen anonymous ftp server.

Biology↗

Virgil database for rich links (1999 update).

With so many databases available for research in the Human Genome Project, it is crucial to efficiently relate information from different resources. For that purpose, we maintain Virgil, a database of rich links for data browsing, data analysis and database interconnection. Virgil current version contains more than 40 000 rich links from five major databases: SWISS-PROT, GenBank, PDB, GDB and OMIM. Materials described in this paper are available from http://www.infobiogen.fr/services/virgil/

Animals↗

The HuGeMap Database: interconnection and visualization of human genome maps.

The HuGeMap database stores the major genetic and physical maps of the human genome. HuGeMap is accessible on the Web at http://www. infobiogen.fr/services/Hugemap and through a CORBA server. A standard genome map data format for the interconnection of genome map databases was defined in collaboration with the EBI. The HuGeMap CORBA server provides this interconnection using the interface definition language IDL. Two graphical user interfaces were developed for the visualization of the HuGeMap data: ZoomMap (http://www.infobiogen.fr/services/zomit/Zoom Map.html) for navigation by zooming and data transformation via magic lenses, and MappetShow (http://www.infobiogen.fr/services/Mappet) for visualizing and comparing maps.

Animals↗

Strategies for detecting susceptibility genes in a complex disease.

One of the current issues in genetic epidemiology is detecting susceptibility genes on the genome. It is common now to undertake systematic screening of the genome using approaches based on a measure of the haplotype sharing in sib pairs. Here, we compare the efficiency of two statistics, the maximum likelihood score (MLS) and the nonparametric linkage score (NPLa) on the simulated data provided for GAW11. A question often raised is whether it is better to perform a single-step or a two-step strategy. For the simulated model, and whatever the strategy used, we show here that the answer is not unequivocal. In both cases, the power to detect susceptibility genes in a single replicate with MLS or NPL is extremely low. With two replicates, only one of the four simulated loci could be detected with reasonable power. When gametic disequilibrium is suspected, methods testing for both linkage and association might be more powerful.

Genetic Linkage↗

A proposal for a standard CORBA interface for genome maps.

MOTIVATION: The scientific community urgently needs to standardize the exchange of biological data. This is helped by the use of a common protocol and the definition of shared data structures. We have based our standardization work on CORBA, a technology that has become a standard in the past years and allows interoperability between distributed objects. RESULTS: We have defined an IDL specification for genome maps and present it to the scientific community. We have implemented CORBA servers based on this IDL to distribute RHdb and HuGeMap maps. The IDL will co-evolve with the needs of the mapping community. AVAILABILITY: The standard IDL for genome maps is available at http:// corba.ebi.ac.uk/RHdb/EUCORBA/MapIDL.htm l. The IORs to browse maps from Infobiogen and EBI are at http://www.infobiogen.fr/services/Hugemap/IOR and http://corba.ebi.ac.uk/RHdb/EUCORBA/IOR CONTACT: manu@infobiogen.fr, tome@ebi.ac.uk

Animals↗

CoPE: a collaborative pedigree drawing environment.

SUMMARY: We developed a collaborative pedigree environment called CoPE. This environment includes a Java program for drawing pedigrees and a standardized system for pedigree storage. Unlike other existing pedigree programs, this software is particularly intended for epidemiologists in the sense that it allows customized automatic drawing of large numbers of pedigrees and remote and distributed consultation of pedigrees. AVAILABILITY: At http://www.infobiogen.fr/services/CoPE

Pedigree↗

Virgil: a database of rich links between GDB and GenBank.

Database interconnection requires the development of links between related objects from different databases. We built a database of links, called Virgil, to manage and distribute rich (documented) links between GDB genes and GenBank human sequences. Virgil contains 18 667 unique links. In addition to a simple Web form for ad-hoc queries, we propose a generic Web interface and a prototype CORBA server for link distribution. Materials described in this paper are available from http://www.infobiogen.fr/services/virgil/home. html

Computer Communication Networks↗

HuGeMap: a distributed and integrated Human Genome Map database.

The HuGeMap database stores the major genetic and physical maps of the human genome. It is also interconnected with the gene radiation hybrid mapping database RHdb. HuGeMap is accessible through a Web server for interactive browsing at URL http://www.infobiogen. fr/services/Hugemap , as well as through a CORBA server for effective programming. HuGeMap is intended as an attempt to build open, interconnected databases, that is databases that distribute their objects worldwide in compliance with a recognized standard of distribution. Maps can be displayed and compared with a java applet (http://babbage.infobiogen.fr:15000/Mappet/Show. html ) that queries the HuGeMap ORB server as well as the RHdb ORB server at the EBI.

Chromosome Mapping↗

The new Virgil database: a service of rich links.

MOTIVATION: Links between biological objects are frequently used by researchers in biology. However, many of the links found in public databases are insufficiently documented and difficult to retrieve. Virgil introduces the idea of a rich link, i.e. the link itself and the related pieces of information. Virgil was developed to collect, manage and distribute such links. RESULTS: At the moment, Virgil is a prototype database that contains rich links between GDB genes and Genbank sequences. The Virgil data model is rich enough to describe comprehensively a link between two biological objects. Two different means to access the information were developed: a schema-driven Web interface and a CORBA server. AVAILABILITY: http://www.infobiogen. fr/services/virgil/home.html CONTACT: Frederic.Achard@infobiogen.fr

Computer Communication Networks↗

Zomit: biological data visualization and browsing.

MOTIVATION: The problems caused by the difficulty in visualizing and browsing biological databases have become crucial. Scientists can no longer interact directly with the huge amount of available data. However, future breakthroughs in biology depend on this interaction. We propose a new metaphor for biological data visualization and browsing that allows navigation in very large databases in an intuitive way. The concepts underlying our approach are based on navigation and visualization with zooming, semantic zooming and portals; and on data transformation via magic lenses. We think that these new visualization and navigation techniques should be applied globally to a federation of biological databases. RESULTS: We have implemented a generic tool, called Zomit, that provides an application programming interface for developing servers for such navigation and visualization, and a generic architecture-independent client (Javatrade mark applet) that queries such servers. As an illustration of the capabilities of our approach, we have developed ZoomMap, a prototype browser for the HuGeMap human genome map database. AVAILABILITY: Zomit and ZoomMap are available at the URL http://www.infobiogen.fr/services/zomit.

Biological Science Disciplines↗

Ubiquitous distributed objects with CORBA.

Database interoperation is becoming a bottleneck for the research community in biology. In this paper, we first discuss the question of interoperability and give a brief overview of CORBA. Then, an example is explained in some detail: a simple but realistic data bank of STSs is implemented. The Object Request Broker is the media for communication between an object server (the data bank) and a client (possibly a genome center). Since CORBA enables easy development of networked applications, we meant this paper to provide an incentive for the bioinformatics community to develop distributed objects.

Base Sequence↗

Isolation of chromosome 21-specific yeast artificial chromosomes from a total human genome library.

A new approach for the isolation of chromosome-specific subsets from a human genomic yeast artificial chromosome (YAC) library is described. It is based on the hybridization with an Alu polymerase chain reaction (PCR) probe. We screened a 1.5 genome equivalent YAC library of megabase insert size with Alu PCR products amplified from hybrid cell lines containing human chromosome 21, and identified a subset of 63 clones representative of this chromosome. The majority of clones were assigned to chromosome 21 by the presence of specific STSs and in situ hybridization. Twenty-nine of 36 STSs that we tested were detected in the subset, and a contig spanning 20 centimorgans in the genetic map and containing 8 STSs in 4 YACs was identified. The proposed approach can greatly speed efforts to construct physical maps of the human genome.

Base Sequence↗

Theoretical analysis of library screening using a N-dimensional pooling strategy.

A solution to the problem of library screening is analysed. We examine how to retrieve those clones that are positive for a single copy landmark from a whole library while performing only a minimum number of laboratory tests: the clones are arranged on a matrix (i.e in 2 dimensions) and pooled according to the rows and columns. A fingerprint is determined for each pool and an analysis allows selection of a list containing all the positive clones, plus a few false positives. These false positives are eliminated by using another (or several other) matrix which has to be reconfigured in a way as different as possible from the previous one. We examine the use of cubes (3 dimensions) or hypercubes of any dimension instead of matrices and analyse how to reconfigure them in order to eliminate the false positives as efficiently as possible. The advantage of the method proposed is the low number of tests required and the low number of pools that require to be prepared [only 258 pools and 282 tests (258 + 24 verifications) are needed to screen the 72,000 clones of the CEPH YAC library (1) with a sequence-tagged site]. Furthermore, this method allows easy and systematic screenings and can be applied to a large physical mapping project, which will lead to an interesting map with a low, precisely known, rate of error: when fingerprinting a 150 Mb chromosome with the CEPH YAC library and 1750 sequence-tagged sites, 903,000 tests would be necessary to obtain about 20 contigs of an average length of 6.7 Mb, while only about one false positive would be expected in the resultant map. Finally, STSs can be ordered by dividing a clone library into sublibraries (corresponding to groups of microplates for example) and testing each STS on pooled clones from each sublibrary. This allows to dedicate to each STSs a fingerprint that consists in the list of the positive pools. In many cases these fingerprints will be enough to order the STSs. Indeed if large YACs (greater than 1 Mb) can be obtained, the combined screening of DNA families and YAC DNA pools would allow an integrated construction of both genetic and physical maps of the human genome, that will also reduce the optimal number of meioses needed for a 1 centimorgan linkage map.

Chromosome Mapping↗