PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Aligning experimental design with bioinformatics analysis to meet discovery research objectives.

The utility of genomic technology and bioinformatic analytical support to provide new and needed insight into the molecular basis of disease, development, and diversity continues to grow as more research model systems and populations are investigated. Yet deriving results that meet a specific set of research objectives requires aligning or coordinating the design of the experiment, the laboratory techniques, and the data analysis. The following paragraphs describe several important interdependent factors that need to be considered to generate high quality data from the microarray platform. These factors include aligning oligonucleotide probe design with the sample labeling strategy if oligonucleotide probes are employed, recognizing that compromises are inherent in different sample procurement methods, normalizing 2-color microarray raw data, and distinguishing the difference between gene clustering and sample clustering. These factors do not represent an exhaustive list of technical variables in microarray-based research, but this list highlights those variables that span both experimental execution and data analysis.

Computational Biology↗

Bioinformatic analysis of exon repetition, exon scrambling and trans-splicing in humans.

MOTIVATION: Using bioinformatic approaches we aimed to characterize poorly understood abnormalities in splicing known as exon scrambling, exon repetition and trans-splicing. RESULTS: We developed a software package that allows large-scale comparison of all human expressed sequence tags (EST) sequences to the entire set of human gene sequences. Among 5,992,495 EST sequences, 401 cases of exon repetition and 416 cases of exon scrambling were found. The vast majority of identified ESTs contain fragments rather than full-length repeated or scrambled exons. Their structures suggest that the scrambled or repeated exon fragments may have arisen in the process of cDNA cloning and not from splicing abnormalities. Nevertheless, we found 11 cases of full-length exon repetition showing that this phenomenon is real yet very rare. In searching for examples of trans-splicing, we looked only at reproducible events where at least two independent ESTs represent the same putative trans-splicing event. We found 15 ESTs representing five types of putative trans-splicing. However, all 15 cases were derived from human malignant tissues and could have resulted from genomic rearrangements. Our results provide support for a very rare but physiological occurrence of exon repetition, but suggest that apparent exon scrambling and trans-splicing result, respectively, from in vitro artifact and gene-level abnormalities. AVAILABILITY: Exon-Intron Database (EID) is available at http://www.meduohio.edu/bioinfo/eid. Programs are available at http://www.meduohio.edu/bioinfo/software.html. The Laboratory website is available at http://www.meduohio.edu/medicine/fedorov SUPPLEMENTARY INFORMATION: Supplementary file is available at http://www.meduohio.edu/bioinfo/software.html.

Algorithms↗

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation↗

A high-throughput approach for subcellular proteome: identification of rat liver proteins using subcellular fractionation coupled with two-dimensional liquid chromatography tandem mass spectrometry and bioinformatic analysis.

Four fractions from rat liver (a crude mitochondria (CM) and cytosol (C) fraction obtained with differential centrifugation, a purified mitochondrial (PM) fraction obtained with nycodenz density gradient centrifugation, and a total liver (TL) fraction) were analyzed with two-dimensional liquid chromatography tandem mass spectrometry analysis. A total of 564 rat proteins were identified and were bioinformatically annotated according to their physicochemical characteristics and functions. While most extreme alkaline ribosomal proteins were identified in the TL fraction, the C fraction mainly included neutral enzymes and the PM fraction enriched alkaline proteins and proteins with electron transfer activity or oxygen binding activity. Such characteristics were more apparent in proteins identified only in the TL, C, or PM fraction. The Swiss-Prot annotation and the bioinformatic prediction results proved that the C and PM fractions had enriched cytoplasmic or mitochondrial proteins, respectively. Combination usage of subcellular fractionation with two-dimensional liquid chromatography tandem mass spectrometry was proved to be a high-throughput, sensitive, and effective analytical approach for subcellular proteomics research. Using such a strategy, we have constructed the largest proteome database to date for rat liver (564 rat proteins) and its cytosol (222 rat proteins) and mitochondrial fractions (227 rat proteins). Moreover, the 352 proteins with Swiss-Prot subcellular location annotation in the 564 identified proteins were used as an actual subcellular proteome dataset to evaluate the widely used bioinformatics tools such as PSORT, TargetP, TMHMM, and GRAVY.

Animals↗

Identification of potential biomarkers and mechanisms for keloid disorder based on comprehensive bioinformatics analysis and machine learning algorithms.

BACKGROUND: Keloid disorder (KD) encompasses a spectrum of fibroproliferative dermal conditions, the pathogenesis remains complex and incompletely understood. This study sought to identify biomarkers and potential therapeutic targets for KD through an integrative bioinformatics approach and machine learning analysis of RNA sequencing data. METHODS: RNA sequencing was performed on skin tissue samples from 13 patients with KD and 14 healthy controls. Using weighted gene co-expression network analysis and differential expression analysis revealed differentially expressed key module genes, and the CytoHubba plugin identified candidate genes. Subsequently analyzed using least absolute shrinkage and selection operator (LASSO) and support vector machine recursive feature elimination (SVM-RFE) methods to pinpoint feature genes associated with KD. Following this, biomarkers were determined through expression level validation, enrichment analysis, and immune infiltration analysis. RESULTS: A total of 420 differentially expressed key module genes were identified, and the top 10 genes with DMNC values were selected as candidate genes. Five feature genes were selected through LASSO and SVM-RFE, with NID2, MFAP2, COL8A1, and P4HA3 showing significant expression differences between KD and control samples, along with consistent expression patterns across datasets, identified as potential biomarkers. These four biomarkers were proved to possess high diagnostic potential, and they were found to exhibit significant positive correlations with one another. Functional enrichment analysis indicated that the primary KEGG pathways associated with these biomarkers included "steroid hormone biosynthesis" and "cytokine-cytokine receptor interaction." Moreover, immune infiltration analysis revealed that the four biomarkers were negatively correlated with type 17 T helper cells and positively correlated with 15 immune cell types, including activated B cells and central memory CD4 T cells. CONCLUSION: In conclusion, NID2, MFAP2, COL8A1, and P4HA3 were identified as key biomarkers for KD, offering new avenues for more targeted and effective diagnostic and therapeutic strategies for managing this condition.

Humans↗

Bioinformatics analysis of mycoplasma metabolism: important enzymes, metabolic similarities, and redundancy.

In this work we apply a bioinformatics approach to determine the most important enzymes of the metabolic network of mycoplasmas. The genomes of several mycoplasmas shared predicted important enzymes. Our method allows us to determine both enzymes that are isolated from the metabolic network of the organism and those that are redundant. We also compare the similarities of the mycoplasmas metabolic networks with the phylogenetic relationships predicted from their 16s rRNA sequences.

Computational Biology↗

The cost and cost trajectory of genome sequencing and bioinformatics analysis for Indigenous children with suspected rare diseases.

PURPOSE: Indigenous peoples are underrepresented in reference genome libraries. Consequently, rare disease diagnosis may require bespoke bioinformatics analyses of genome sequences. Establishing diagnostic cost is crucial to support policy development for equitable diagnosis of rare diseases. We estimated the cost and cost trajectory of diagnostic genome sequencing and bioinformatics for Indigenous participants with suspected rare diseases. METHODS: We conducted a microcosting study of Indigenous children and their families receiving genome sequencing through Canada's Silent Genomes Project. Invoice data informed the costs of genome sequencing. We conducted a time-and-motion study for bioinformatics analyses, including labor, computing, and data storage costs. RESULTS: With standard bioinformatics, costs ranged from C$3645 (SD: 455) for singletons to C$7402 (SD: 566) for trios. With advanced, bespoke bioinformatics, costs ranged from C$5344 (SD: 634) for singletons to C$9760 (SD: 822) for trios. Genome sequencing was a primary cost driver; however, sequencing costs decreased by 61% over 4 years. Bioinformatics costs ranged from 21.3% to 58.3% of the total costs. The time required for bioinformatics ranged from 71 hours to 215 hours for standard and advanced analyses, respectively. CONCLUSION: Genome sequencing costs decreased over time. Bioinformatics is a significant cost driver, particularly for bespoke analyses arising from nonrepresentative reference libraries.

Humans↗

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, “DESeq” and “Limma” R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology↗

Bioinformatic analysis reveals the potential association of ESRP1 with the splicing of cytoskeleton-associated genes in doxorubicin-resistant MCF7 breast cancer cells.

BACKGROUND: Breast cancer remains one of the most prevalent malignancies among women, with doxorubicin resistance posing a significant challenge that undermines treatment success and survival outcomes. Aberrant alternative splicing (AS), driven by dysregulation or mutations in splicing factors (SFs), is implicated in cancer initiation, progression, and drug resistance. This study aims to investigate the association of the epithelial cell-specific splicing factor ESRP1 with doxorubicin resistance in breast cancer, focusing on how ESRP1 deficiency correlates with AS changes that promote chemoresistance. METHODS: We analyzed RNA-sequencing (RNA-seq) data from doxorubicin-resistant (MCF7-DR) and parental (MCF7) breast cancer cell lines to identify enhanced alternative splicing events (ASEs) and changes in ESRP1 expression; we further leveraged The Cancer Genome Atlas (TCGA)-BRCA cohort to construct an SF-RASE correlation network for screening core SFs (including ESRP1). An integrative analysis combining crosslinking immunoprecipitation (CLIP-seq) data and The Cancer Genome Atlas (TCGA) database was performed to validate ESRP1 binding targets and assess the association between ESRP1-related splicing and cytoskeleton organization. RESULTS: We observed extensive AS changes and significantly downregulated ESRP1 expression in MCF7-DR cells. Integrative analysis identified 61 high-confidence ASEs that correlate with ESRP1 expression. Further bioinformatic integration suggests that ESRP1 expression is associated with the splicing patterns of SPTBN1, MAP2K7, FGFR3, and CYB561A3-four genes involved in cytoskeleton organization-though direct experimental verification to confirm a causal regulatory relationship between ESRP1 and the splicing of these genes is still pending. CONCLUSIONS: Our findings suggest that ESRP1 expression is closely associated with doxorubicin resistance in breast cancer cells, with concomitant alterations in key ASEs linked to cytoskeletal remodeling that correlate with ESRP1. Exploring the ESRP1-related splicing network may offer new strategies to overcome chemoresistance and improve patient outcomes. However, the small cell line sample size (n = 2 per group) constrains the robustness of ASE and SF-ASE correlation findings, and these results should be interpreted with caution and require further validation with larger sample cohorts.

Alternative splicing↗

Genetic and bioinformatic analysis of 41C and the 2R heterochromatin of Drosophila melanogaster: a window on the heterochromatin-euchromatin junction.

Genomic sequences provide powerful new tools in genetic analysis, making it possible to combine classical genetics with genomics to characterize the genes in a particular chromosome region. These approaches have been applied successfully to the euchromatin, but analysis of the heterochromatin has lagged somewhat behind. We describe a combined genetic and bioinformatics approach to the base of the right arm of the Drosophila melanogaster second chromosome, at the boundary between pericentric heterochromatin and euchromatin. We used resources provided by the genome project to derive a physical map of the region, examine gene density, and estimate the number of potential genes. We also carried out a large-scale genetic screen for lethal mutations in the region. We identified new alleles of the known essential genes and also identified mutations in 21 novel loci. Fourteen complementation groups map proximal to the assembled sequence. We used PCR to map the endpoints of several deficiencies and used the same set of deficiencies to order the essential genes, correlating the genetic and physical map. This allowed us to assign two of the complementation groups to particular "computed/curated genes" (CGs), one of which is Nipped-A, which our evidence suggests encodes Drosophila Tra1/TRRAP.

Animals↗

[Bioinformatic analysis of adenoma-normal mucosa SSH library of colon].

We established a colonic adenoma-normal mucosa suppressive subtraction hybridization (SSH) library in 1999. In this study, we wanted to explore the expression profile of all candidate genes in this library. We developed an EST pipeline which contained two in-house software packages, nucleic acid analytical software and GetUni. The nucleic acid analytical software, an integrator of the universal bioinformatics tools including phred, phd2fasta, cross_match, repeatmasker and blast2.0, can blast sequences of differential clones with the downloaded non-redundant nucleotide (NR) database. GetUni can cluster these NR sequences into Unigene via matching with the downloaded Homo Sapiens UniGene database. Sixty-two candidate genes in A-N library were obtained via the high throughput automatic gene expression bioinformatics pipeline. Gene Ontology online analysis revealed that ribosome genes and immunity-regulating genes were the two most common categories in the KEGG or Biocarta Pathway. We also detected the expression of 2 genes with highest hits, Reg4 and FAM46A, by semi-quantitative RT-PCR. Both genes were up-regulated in 10 or 9 out of 10 adenomas in comparison with the paired normal mucosa, respectively. The candidate genes in A-N library would be of great significance in disclosing the molecular mechanism underlying in colonic adenoma initiation and progression.

Adenoma↗

Gene expression of human T lymphocytes cell cycle: experimental and bioinformatic analysis.

Human lymphocytes gene expression is monitored before and after PHA stimulation over 72 h, using DNA microarray technology. Results are then compared with our previous bioinformatics predictions, which identified six leader genes of highest importance in human T lymphocytes cell cycle. Experimental data are strikingly compatible with bioinformatic predictions of the specific role and interaction of PCNA, CDC2, and CCNA2 at all phases of the cell cycle and of CHEK1 in regulating DNA repair and preservation. It does not escape our notice that the conception and use of ad hoc arrays, based on a bioinformatics prediction which identifies the most important genes involved in a particular biological process, can really be an added value in cell biology and cancer research alternative to massive frequently misleading molecular genomics.

CDC2 Protein Kinase↗

The current excitement in bioinformatics-analysis of whole-genome expression data: how does it relate to protein structure and function?

Whole-genome expression profiles provide a rich new data-trove for bioinformatics. Initial analyses of the profiles have included clustering and cross-referencing to 'external' information on protein structure and function. Expression profile clusters do relate to protein function, but the correlation is not perfect, with the discrepancies partially resulting from the difficulty in consistently defining function. Other attributes of proteins can also be related to expression-in particular, structure and localization-and sometimes show a clearer relationship than function.

Computational Biology↗

Bioinformatic analysis of autism positional candidate genes using biological databases and computational gene network prediction.

Common genetic disorders are believed to arise from the combined effects of multiple inherited genetic variants acting in concert with environmental factors, such that any given DNA sequence variant may have only a marginal effect on disease outcome. As a consequence, the correlation between disease status and any given DNA marker allele in a genomewide linkage study tends to be relatively weak and the implicated regions typically encompass hundreds of positional candidate genes. Therefore, new strategies are needed to parse relatively large sets of 'positional' candidate genes in search of actual disease-related gene variants. Here we use biological databases to identify 383 positional candidate genes predicted by genomewide genetic linkage analysis of a large set of families, each with two or more members diagnosed with autism, or autism spectrum disorder (ASD). Next, we seek to identify a subset of biologically meaningful, high priority candidates. The strategy is to select autism candidate genes based on prior genetic evidence from the allelic association literature to query the known transcripts within the 1-LOD (logarithm of the odds) support interval for each region. We use recently developed bioinformatic programs that automatically search the biological literature to predict pathways of interacting genes (PATHWAYASSIST and GENEWAYS). To identify gene regulatory networks, we search for coexpression between candidate genes and positional candidates. The studies are intended both to inform studies of autism, and to illustrate and explore the increasing potential of bioinformatic approaches as a compliment to linkage analysis.

Autistic Disorder↗

Identification of novel highly expressed genes in pancreatic ductal adenocarcinomas through a bioinformatics analysis of expressed sequence tags.

In most microarray experiments, a significant fraction of the differentially expressed mRNAs identified correspond to expressed sequence tags (ESTs) and are generally discarded from further analyses. We used careful bioinformatics analyses to characterize those ESTs that were found to be highly overexpressed in a series of pancreatic adenocarcinomas. cDNA was prepared from 60 non-neoplastic samples (normal pancreas [n = 20], normal colon [n = 10], or normal duodenal mucosal [n = 30]) and from 64 pancreatic cancers (resected cancers [n = 50] or cancer cell lines [n = 14]) and hybridized to the complete Affymetrix Human Genome U133 GeneChip(R) set (arrays U133A and B) for simultaneous analysis of 45,000 fragments corresponding to 33,000 known genes and 6,000 ESTs. The GeneExpress(R) software system Fold Change Analysis Tool was used and 60 ESTs were identified that were expressed at levels at least 3-fold greater in the pancreatic cancers as compared to normal tissues. Searches against the human genomic sequence and comparative genomic analysis of human and mouse genomes was carried out using basic local alignment search tools (BLAST), BLASTN, and BLASTX, for identifying protein coding genes corresponding to the ESTs. Subsequently, in order to pick the most relevant candidate genes for a more detailed analysis, we looked for domains/motifs in the open reading frames using SMART and Pfam programs. We were able to definitively map 43 of the 60 ESTs to known or novel genes, and 15 of the ESTs could be localized in close proximity to a gene in the human genome although we were unable to establish that the EST was indeed derived from those genes. The differential expression of a subset of genes was confirmed at the protein level by immunohistochemical labeling of tissue microarrays (inhibin beta A [INHBA] and CD29) and/or at the transcript level by RT-PCR (INHBA, AKAP12, ELK3, FOXQ1, EIF5A2, and EFNA5). We conclude that bioinformatics tools can be used to characterize differentially overexpressed ESTs, and that some of these ESTs may represent diagnostically and therapeutically useful targets that might be missed using data solely from currently annotated databases.

Adenocarcinoma↗

Identification of a new Schistosoma mansoni membrane-bound protein through bioinformatic analysis.

Progress in schistosome genome research has enabled investigators to move rapidly from genome sequences to vaccine development. Proteins bound to the surface of parasites are potential vaccine candidates, or they can be used for diagnosis. We analyzed 4342 proteins deduced from the Schistosoma mansoni transcriptome with bioinformatic computer programs. Thirty-four proteins had membrane-bound motifs. Within this group, we selected the Sm29 protein to be further characterized by in silico analysis. Sm29 was found to have a signal peptide made up of 26 amino acids, with a cleavage site between Ser26 and Val27. The glycosylation site search revealed three threonines (39, 132 and 133) with high probability of O-glycosylation and two asparagines (58 and 115) with high probability of N-glycosylation. Only one transmembrane helix was found in the C-terminal region of the protein from Leu169 to Lis191. The search for similarities and conserved motifs show that Sm29 is a protein with high identity to proteins present in S. japonicum (53, 52, 49, and 37% of identity) and it possesses disulfide-rich conserved domains. Apparently, Sm29 is a membrane bound protein, and it may be an important molecule in host-parasite interactions.

Amino Acid Sequence↗

Comparative bioinformatic analysis of complete proteomes and protein parameters for cross-species identification in proteomics.

Peptide mass fingerprinting (PMF) remains the most amenable technique for protein identification in proteomics, using mass spectrometry as the primary analytical technique coupled with bioinformatics. This relies on the presence of the amino acid sequence of the protein in the current databanks. Despite this, it is desirable to be able to use the technique for organisms whose genomes are not yet fully sequenced and apply cross-species protein identification. In this study, we have re-examined the feasibility of such approaches by considering the extent of protein similarity between genome sequences using a data set of 29 complete bacterial and two eukaryotic genomes. A range of protein and peptide features are considered, including protein isoelectric focussing point, protein mass, and amino acid conservation. The effectiveness of PMF approaches has then been tested with a series of computer simulations with varying peptide number and mass accuracy for several cross-species tests. The results show that PMF alone is unsuitable in general for divergent species jumps, or when protein similarity is less than 70% identity. Despite this, there exists a considerable enrichment above random of tryptic peptide conservation and PMF promises to remain useful when combined with other data than just peptide masses for cross-species protein identification.

Algorithms↗

Structural features of normal and mutant human lysosomal glycoside hydrolases deduced from bioinformatics analysis.

Lysosomal storage diseases are due to inherited deficiencies in various enzymes involved in basic metabolic processes. As with other genetic diseases, accurate structure data for these enzymatic proteins should help in better understanding the molecular effects of mutations identified in patients with the corresponding lysosomal diseases; however, no such three-dimensional (3D) structure data are available for many lysosomal enzymes. Thus, we herein intend to illustrate for an audience of molecular geneticists how structure information can nonetheless be obtained via a bioinformatics approach in the case of five human lysosomal glycoside hydrolases. Indeed, using the two-dimensional hydrophobic cluster analysis method to decipher the sequence information available in data banks for the large group of glycoside hydrolases (clan GH-A) to which these human lysosomal enzymes belong, we could deduce structure predictions for their catalytic domains and propose explanations for the molecular effects of mutations described in patients. In addition, in the case of human beta-glucuronidase for which experimental 3D data have been reported, we also show here that bioinformatics methods relying on the available 3D structure information can be used to obtain further insights into the effects of various mutations described in patients with Sly disease. In a broader perspective, our work stresses that, in the context of a rapid increase in protein sequence information through genome sequencing, bioinformatics approaches might be highly useful for generating structure-function predictions based on sequence-structure interrelationships.

Crystallography, X-Ray↗