PubMed Health⌕ Search

Biomedical subjects

Jaak Vilo

Publications and source records attributed to Jaak Vilo.

6 recordsLinked to original sources

ArrayExpress--a public repository for microarray gene expression data at the EBI.

ArrayExpress is a new public database of microarray gene expression data at the EBI, which is a generic gene expression database designed to hold data from all microarray platforms. ArrayExpress uses the annotation standard Minimum Information About a Microarray Experiment (MIAME) and the associated XML data exchange format Microarray Gene Expression Markup Language (MAGE-ML) and it is designed to store well annotated data in a structured way. The ArrayExpress infrastructure consists of the database itself, data submissions in MAGE-ML format or via an online submission tool MIAMExpress, online database query interface, and the Expression Profiler online analysis tool. ArrayExpress accepts three types of submission, arrays, experiments and protocols, each of these is assigned an accession number. Help on data submission and annotation is provided by the curation team. The database can be queried on parameters such as author, laboratory, organism, experiment or array types. With an increasing number of organisations adopting MAGE-ML standard, the volume of submissions to ArrayExpress is increasing rapidly. The database can be accessed at http://www.ebi.ac.uk/arrayexpress.

Animals↗

A first-generation linkage disequilibrium map of human chromosome 22.

DNA sequence variants in specific genes or regions of the human genome are responsible for a variety of phenotypes such as disease risk or variable drug response. These variants can be investigated directly, or through their non-random associations with neighbouring markers (called linkage disequilibrium (LD)). Here we report measurement of LD along the complete sequence of human chromosome 22. Duplicate genotyping and analysis of 1,504 markers in Centre d'Etude du Polymorphisme Humain (CEPH) reference families at a median spacing of 15 kilobases (kb) reveals a highly variable pattern of LD along the chromosome, in which extensive regions of nearly complete LD up to 804 kb in length are interspersed with regions of little or no detectable LD. The LD patterns are replicated in a panel of unrelated UK Caucasians. There is a strong correlation between high LD and low recombination frequency in the extant genetic map, suggesting that historical and contemporary recombination rates are similar. This study demonstrates the feasibility of developing genome-wide maps of LD.

Chromosome Mapping↗

From genomes to vaccines: Leishmania as a model.

The 35 Mb genome of Leishmania should be sequenced by late 2002. It contains approximately 8500 genes that will probably translate into more than 10 000 proteins. In the laboratory we have been piloting strategies to try to harness the power of the genome-proteome for rapid screening of new vaccine candidate. To this end, microarray analysis of 1094 unique genes identified using an EST analysis of 2091 cDNA clones from spliced leader libraries prepared from different developmental stages of Leishmania has been employed. The plan was to identify amastigote-expressed genes that could be used in high-throughput DNA-vaccine screens to identify potential new vaccine candidates. Despite the lack of transcriptional regulation that polycistronic transcription in Leishmania dictates, the data provide evidence for a high level of post-transcriptional regulation of RNA abundance during the developmental cycle of promastigotes in culture and in lesion-derived amastigotes of Leishmania major. This has provided 147 candidates from the 1094 unique genes that are specifically upregulated in amastigotes and are being used in vaccine studies. Using DNA vaccination, it was demonstrated that pooling strategies can work to identify protective vaccines, but it was found that some potentially protective antigens are masked by other disease-exacerbatory antigens in the pool. A total of 100 new vaccine candidates are currently being tested separately and in pools to extend this analysis, and to facilitate retrospective bioinformatic analysis to develop predictive algorithms for sequences that constitute potentially protective antigens. We are also working with other members of the Leishmania Genome Network to determine whether RNA expression determined by microarray analyses parallels expression at the protein level. We believe we are making good progress in developing strategies that will allow rapid translation of the sequence of Leishmania into potential interventions for disease control in humans.

Animals↗

Microarray data representation, annotation and storage.

Management and analysis of the huge amounts of data produced by microarray experiments is becoming one of the major bottlenecks in the utilization of this high-throughput technology. We describe the basic design of a microarray gene expression database to help microarray users and their informatics teams to set up their information services. We describe two data models--a simpler one called ArrayExpressB and the complete model ArrayExpressC, and discuss some implementation issues. For latest developments see http: wwwebi.ac.uk/arrayexpress

Algorithms↗

Protein interaction verification and functional annotation by integrated analysis of genome-scale data.

Assays capable of determining the properties of thousands of genes in parallel present challenges with regard to accurate data processing and functional annotation. Collections of microarray expression data are applied here to assess the quality of different high-throughput protein interaction data sets. Significant differences are found. Confidence in 973 out of 5342 putative two-hybrid interactions from S. cerevisiae is increased. Besides verification, integration of expression and interaction data is employed to provide functional annotation for over 300 previously uncharacterized genes. The robustness of these approaches is demonstrated by experiments that test the in silico predictions made. This study shows how integration improves the utility of different types of functional genomic data and how well this contributes to functional annotation.

Animals↗

Correlating gene promoters and expression in gene disruption experiments.

MOTIVATION: Finding putative transcription factor binding sites in the upstream sequences of similarly expressed genes has recently become a subject of intensive studies. In this paper we investigate how much gene expression regulation can be attributed to the presence of various binding sites in the gene promoters by correlating the binding sites and the changes in gene expression resulting from gene disruptions (e.g. knockouts). RESULTS: We have developed a data analysis method for comparing mRNA measurements of gene disruption experiments with information about gene promoters. The method was applied to a well-known dataset to uncover correlations between known transcription factor binding site motifs in the upstream regions of all S. cerevisiae genes and the gene expression changes in various gene disruption experiments. The possible explanations of the correlations were categorized and analyzed using e.g. expression cascades. Several correlations turned out to be consistent with existing biological knowledge while some new ones suggest themselves for further study. AVAILABILITY: The resulting tables are available at http://www.cs.helsinki.fi/u/kpalin/CorrDisrupt/.

Algorithms↗