PubMed Health⌕ Search

Biomedical subjects

Xavier Messeguer

Publications and source records attributed to Xavier Messeguer.

10 recordsLinked to original sources

M-GCAT: interactively and efficiently constructing large-scale multiple genome comparison frameworks in closely related species.

BACKGROUND: Due to recent advances in whole genome shotgun sequencing and assembly technologies, the financial cost of decoding an organism's DNA has been drastically reduced, resulting in a recent explosion of genomic sequencing projects. This increase in related genomic data will allow for in depth studies of evolution in closely related species through multiple whole genome comparisons. RESULTS: To facilitate such comparisons, we present an interactive multiple genome comparison and alignment tool, M-GCAT, that can efficiently construct multiple genome comparison frameworks in closely related species. M-GCAT is able to compare and identify highly conserved regions in up to 20 closely related bacterial species in minutes on a standard computer, and as many as 90 (containing 75 cloned genomes from a set of 15 published enterobacterial genomes) in an hour. M-GCAT also incorporates a novel comparative genomics data visualization interface allowing the user to globally and locally examine and inspect the conserved regions and gene annotations. CONCLUSION: M-GCAT is an interactive comparative genomics tool well suited for quickly generating multiple genome comparisons frameworks and alignments among closely related species. M-GCAT is freely available for download for academic and non-commercial use at: http://alggen.lsi.upc.es/recerca/align/mgcat/intro-mgcat.html.

Algorithms↗

Transcription factor map alignment of promoter regions.

We address the problem of comparing and characterizing the promoter regions of genes with similar expression patterns. This remains a challenging problem in sequence analysis, because often the promoter regions of co-expressed genes do not show discernible sequence conservation. In our approach, thus, we have not directly compared the nucleotide sequence of promoters. Instead, we have obtained predictions of transcription factor binding sites, annotated the predicted sites with the labels of the corresponding binding factors, and aligned the resulting sequences of labels--to which we refer here as transcription factor maps (TF-maps). To obtain the global pairwise alignment of two TF-maps, we have adapted an algorithm initially developed to align restriction enzyme maps. We have optimized the parameters of the algorithm in a small, but well-curated, collection of human-mouse orthologous gene pairs. Results in this dataset, as well as in an independent much larger dataset from the CISRED database, indicate that TF-map alignments are able to uncover conserved regulatory elements, which cannot be detected by the typical sequence alignments.

Algorithms↗

ABS: a database of Annotated regulatory Binding Sites from orthologous promoters.

Information about the genomic coordinates and the sequence of experimentally identified transcription factor binding sites is found scattered under a variety of diverse formats. The availability of standard collections of such high-quality data is important to design, evaluate and improve novel computational approaches to identify binding motifs on promoter sequences from related genes. ABS (http://genome.imim.es/datasets/abs2005/index.html) is a public database of known binding sites identified in promoters of orthologous vertebrate genes that have been manually curated from bibliography. We have annotated 650 experimental binding sites from 68 transcription factors and 100 orthologous target genes in human, mouse, rat or chicken genome sequences. Computational predictions and promoter alignment information are also provided for each entry. A simple and easy-to-use web interface facilitates data retrieval allowing different views of the information. In addition, the release 1.0 of ABS includes a customizable generator of artificial datasets based on the known sites contained in the collection and an evaluation tool to aid during the training and the assessment of motif-finding programs.

Animals↗

FibreHelix, a program for calculating the X-ray diffraction pattern of macromolecules with helical symmetry: application to DNA coiled coils.

A program has been developed to determine the diffraction pattern given by partially ordered fibres formed by macromolecules with helical symmetry. It is particularly useful for visualizing the splitting of layer lines typical of coiled coils. The program produces as output the diffraction diagram calculated for helices that are oriented along their axis but are randomly oriented in other directions. The results can be numerically analyzed and also visualized on-screen. The program has been applied to the diffraction patterns given by DNA and protein coiled coils.

Crystallography, X-Ray↗

MREPATT: detection and analysis of exact consecutive repeats in genomic sequences.

UNLABELLED: We have developed a program to determine the number, length and position of exact consecutive repeats of short sequences in DNA fragments or whole genomes. The program also gives the statistical significance of results by comparing them with those expected for a random sequence generated according to a Markovian model. AVAILABILITY: MREPATT can be accessed on line at http://www.lsi.upc.es/~alggen under the RESEARCH and SEARCH links.

Algorithms↗

DnaSP, DNA polymorphism analyses by the coalescent and other methods.

SUMMARY: DnaSP is a software package for the analysis of DNA polymorphism data. Present version introduces several new modules and features which, among other options allow: (1) handling big data sets (approximately 5 Mb per sequence); (2) conducting a large number of coalescent-based tests by Monte Carlo computer simulations; (3) extensive analyses of the genetic differentiation and gene flow among populations; (4) analysing the evolutionary pattern of preferred and unpreferred codons; (5) generating graphical outputs for an easy visualization of results. AVAILABILITY: The software package, including complete documentation and examples, is freely available to academic users from: http://www.ub.es/dnasp

Algorithms↗

Identification of patterns in biological sequences at the ALGGEN server: PROMO and MALGEN.

In this paper we present several web-based tools to identify conserved patterns in sequences. In particular we present details on the functionality of PROMO version 2.0, a program for the prediction of transcription factor binding site in a single sequence or in a group of related sequences and, of MALGEN, a tool to visualize sequence correspondences among long DNA sequences. The web tools and associated documentation can be accessed at http://www.lsi.upc.es/~alggen (RESEARCH link).

Animals↗

Genome-wide analysis of the Emigrant family of MITEs of Arabidopsis thaliana.

Miniature inverted-repeat transposable elements (MITEs) are structurally similar to defective class II elements, but their high copy number and the size and sequence conservation of most MITE families suggest that they can be amplified by a replicative mechanism. Here we present a genome-wide analysis of the Emigrant family of MITEs from Arabidopsis thaliana. In order to be able to detect divergent ancient copies, and low copy number subfamilies with a different internal sequence we have developed a computer program to look for Emigrant elements based solely on the terminal inverted-repeat sequence. We have detected 151 Emigrant elements of different subfamilies. Our results show that different bursts of amplification, probably of few active, or master, elements, have occurred at different times during Arabidopsis evolution. The analysis of the insertion sites of the Emigrant elements shows that recently inserted Emigrant elements tend to be located far from open reading frames, whereas more ancient Emigrant subfamilies are preferentially found associated to genes.

Amino Acid Sequence↗

[Automatization of a hospital-based tumor registry].

INTRODUCTION: To increase data reliability and reduce the costs associated with the HTR, the Catalan Institute of Oncology programmed the manual procedures of data collection from databases by means of a computer application (ASEDAT). MATERIAL AND METHOD: ASEDAT detects the incident tumors of the registry from the databases of the pathology records (PR) and discharge records (DR) and selects the basic information from both databases. Data from the HTR data was collected for the period 1999-2000 by means of 2 procedures: manual and automatized collection and the results obtained were compared. RESULTS: 10,498 cancer patients were detected. Manual resolution detected 8,309 incident tumors and 2,374 prevalent tumors. ASEDAT automatically detected 8,901 patients (84.8%), in whom 8,367 incident tumors were detected (58 more tumors than the manual procedure). Validation of agreement was performed in the incident tumors detected by both methods (7,063 tumors). In 6,185 tumors (87.6%) the information agreed in all the variables. Of the discordant tumors, 692 (9.8%) were obtained by the RHT staff using manual resolution, and the remainder (186; 2.6%) were obtained by the application (automatic resolution). CONCLUSIONS: Cancer registry automatization is feasible when PR and DR databases are available, coded and automatized.

Hospital Records↗