PubMed Health⌕ Search

Biomedical subjects

Matthew Bellgard

Publications and source records attributed to Matthew Bellgard.

9 recordsLinked to original sources

Statistical evaluation and comparison of a pairwise alignment algorithm that a priori assigns the number of gaps rather than employing gap penalties.

MOTIVATION: Although pairwise sequence alignment is essential in comparative genomic sequence analysis, it has proven difficult to precisely determine the gap penalties for a given pair of sequences. A common practice is to employ default penalty values. However, there are a number of problems associated with using gap penalties. First, alignment results can vary depending on the gap penalties, making it difficult to explore appropriate parameters. Second, the statistical significance of an alignment score is typically based on a theoretical model of non-gapped alignments, which may be misleading. Finally, there is no way to control the number of gaps for a given pair of sequences, even if the number of gaps is known in advance. RESULTS: In this paper, we develop and evaluate the performance of an alignment technique that allows the researcher to assign a priori set of the number of allowable gaps, rather than using gap penalties. We compare this approach with the Smith-Waterman and Needleman-Wunsch techniques on a set of structurally aligned protein sequences. We demonstrate that this approach outperforms the other techniques, especially for short sequences (56-133 residues) with low similarity (<25%). Further, by employing a statistical measure, we show that it can be used to assess the quality of the alignment in relation to the true alignment with the associated optimal number of gaps. AVAILABILITY: The implementation of the described methods SANK_AL is available at http://cbbc.murdoch.edu.au/ CONTACT: matthew@cbbc.murdoch.edu.au.

Algorithms↗

Comparative organization of wheat homoeologous group 3S and 7L using wheat-rice synteny and identification of potential markers for genes controlling xanthophyll content in wheat.

EST and genomic DNA sequencing efforts for rice and wheat have provided the basis for interpreting genome organization and evolution. In this study we have used EST and genomic sequencing information and a bioinformatic approach in a two-step strategy to align portions of the wheat and rice genomes. In the first step, wheat ESTs were used to identify rice orthologs and it was shown that wheat 3S and rice 1 contain syntenic units with intrachromosomal rearrangements. Further analysis using anchored rice contiguous sequences and TBLASTX alignments in a second alignment step showed interruptions by orthologous genes that map elsewhere in the wheat genome. This indicates that gene content and order is not as conserved as large chromosomal blocks as previously predicted. Similarly, chromosome 7L contains syntenic units with rice 6 and 8 but is interrupted by combinations of intrachromosomal and interchromosomal rearrangements involving syntenic units and single gene orthologs from other rice chromosome groups. We have used the rice sequence annotations to identify genes that can be used to develop markers linked to biosynthetic pathways on 3BS controlling xanthophyll production in wheat and thus involved in determining flour colour.

Amino Acid Sequence↗

Integrative annotation of 21,037 human genes validated by full-length cDNA clones.

The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.

Alternative Splicing↗

Genes controlling seed dormancy and pre-harvest sprouting in a rice-wheat-barley comparison.

Pre-harvest sprouting results in significant economic loss for the grain industry around the world. Lack of adequate seed dormancy is the major reason for pre-harvest sprouting in the field under wet weather conditions. Although this trait is governed by multiple genes it is also highly heritable. A major QTL controlling both pre-harvest sprouting and seed dormancy has been identified on the long arm of barley chromosome 5H, and it explains over 70% of the phenotypic variation. Comparative genomics approaches among barley, wheat and rice were used to identify candidate gene(s) controlling seed dormancy and hence one aspect of pre-harvest sprouting. The barley seed dormancy/pre-harvest sprouting QTL was located in a region that showed good synteny with the terminal end of the long arm of rice chromosome 3. The rice DNA sequences were annotated and a gene encoding GA20-oxidase was identified as a candidate gene controlling the seed dormancy/pre-harvest sprouting QTL on 5HL. This chromosomal region also shared synteny with the telomere region of wheat chromosome 4AL, but was located outside of the QTL reported for seed dormancy in wheat. The wheat chromosome 4AL QTL region for seed dormancy was syntenic to both rice chromosome 3 and 11. In both cases, corresponding QTLs for seed dormancy have been mapped in rice.

Amino Acid Sequence↗

MASV--Multiple (BLAST) Annotation System Viewer.

UNLABELLED: Multiple (BLAST) Annotation System Viewer (MASV) is a tool designed to aid in the annotation of genomic sequences. MASV enables the researcher to compare and analyse differences in annotation and analysis, resulting from changes in databases, analysis program parameters and results. This provides a unique capability for the user to conduct further bioinformatics analysis from the information obtained. AVAILABILITY: http://cbbc.murdoch.edu.au/projects/masv/

Abstracting and Indexing↗

Genomic and phylogenetic analysis of the S100A7 (Psoriasin) gene duplications within the region of the S100 gene cluster on human chromosome 1q21.

The human S100 gene family encodes the EF-hand superfamily of calcium-binding proteins, with at least 14 family members clustered relatively closely together on chromosome 1q21. We have analyzed the most recently available genomic sequence of the human S100 gene cluster for evidence of tandem gene duplications during primate evolutionary history. The sequences obtained from both GenBank and GoldenPath were analyzed in detail using various comparative sequence analysis tools. We found that of the S100A genes clustered relatively closely together within a genomic region of 260 kb, only the S100A7 (psoriasin) gene region showed evidence of recent duplications. The S100A7 gene duplicated region is composed of three distinct genomic regions, 33, 11, and 31 kb, respectively, that together harbor at least five identifiable S100A7-like genes. Regions 1 and 3 are in opposite orientation to each other, but each region carries two S100A7-like genes separated by an 11-kb intergenic region (region 2) that has only one S100A7-like gene, providing limited sequence resemblance to regions 1 and 3. The duplicated genomic regions 1 and 3 share a number of different retroelements including five Alu subfamily members that serve as molecular clocks. The shared (paralogous) Alu S insertions suggest that regions 1 and 3 were probably duplicated during or after the phase of AluS amplification some 30-40 mya. We used PCR to amplify an indel within intron 1 of the S100A7a and S100A7c genes that gave the same two expected product sizes using 40 human DNA samples and 1 chimpanzee sample, therefore confirming the presence of the region 1 and 3 duplication in these species. Comparative genomic analysis of the other S100 gene members shows no similarity between intergenic regions, suggesting that they diverged long before the emergence of the primates. This view was supported by the phylogenetic analysis of different human S100 proteins including the human S100A7 protein members. The S100A7 protein, also known as psoriasin, has important functions as a mediator and regulator in skin differentiation and disease (psoriasis), in breast cancer, and as a chemotactic factor for inflammatory cells. This is the first report of five copies of the S100A7 gene in the human genome, which may impact on our understanding of the possible dose effects of these genes in inflammation and normal skin development and pathogenesis.

Amino Acid Sequence↗

FBSA: feature-based sequence alignment technique for very large sequences.

The ability to align pairs of very large molecular sequences is essential for a range of comparative genomic studies. However, given the complexity of genomic sequences, it has been difficult to devise a systematic method that can align - even within the same species - pairs of large sequences. Most existing approaches typically attempt to align nucleotide sequences while ignoring valuable features contained within them, eg they filter out low-complexity regions and retroelements before aligning the sequences. However, features are then added post-alignment for visualisation and analysis purposes. We argue that repetitive elements and other features (such as genes, exons and regulatory elements) should be part of the alignment process. A hierarchical approach that aligns the biologically relevant features before aligning the detailed nucleotide sequences has a number of interesting characteristics: (1) features define 'alignment anchor points' that can guide meaningful nucleotide alignment; (2) features can be weighted; (3) a hierarchical approach would identify only meaningful regions to be aligned; (4) nucleotide sequences can be described as sequences of features and non-features, providing a natural mechanism to divide the sequences for processing; and (5) computational speed is significantly faster than other approaches. In this paper, we describe and discuss a feature-based approach to aligning large genome sequences. We refer to this as 'feature-based sequence alignment'.

Algorithms↗

Gap mapping: a paradigm for aligning two sequences.

Pairwise sequence alignment is one of the most essential tools in comparative genomic sequence analysis. It is used to compare the sequences of genes and proteins with the aim of inferring structural, functional and evolutionary relationships. However, current 'mainstream' alignment algorithms have optimisation criteria based primarily on computational efficiency using parameters such as gap penalties, which are not biologically motivated. In addition, current alignment algorithms such as the Smith and Waterman technique provide a single alignment that could be sensitive to rather arbitrary choices in parameters such as gap penalties. This paper explores the range of properties resulting from posing the alignment problem more as a 'mapping gaps in sequences' exercise. We argue that this approach is intuitive and provides greater control over the number of gaps placed within an alignment. This type of approach was proposed by Sankoff (1972), but unfortunately has not received much attention. We report and discuss our findings by comparing this approach to other techniques using structurally confirmed aligned sequences from a benchmark alignment database. Interestingly, this approach consistently provides optimal and near optimal alignments and is thus a viable approach to sequence alignment.

Algorithms↗

Microarray analysis using bioinformatics analysis audit trails (BAATs).

Bioinformatics analysis plays an integrative role in genomics and functional genomics. The ability to conduct quality managed, hypothesis-driven bioinformatics analysis with the plethora of data available is mandatory. Biological interpretation of this data is dependent on versions of databases, programs and the parameters used. Thus, tracking and auditing the analyses process is important. This paper outlines what we term Bioinformatics Analysis Audit Trails (BAATs) and describes YABI, a bioinformatics environment that implements BAATs. YABI can incorporate most bioinformatics tools within the same environment, making it a valuable resource.

Computational Biology↗