PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Structural relatedness of plant food allergens with specific reference to cross-reactive allergens: an in silico analysis.

BACKGROUND: The body of sequence and structural information on allergens and the sequence analysis of whole plant genomes are facilitating the application of bioinformatic approaches to identifying and defining plant allergens. OBJECTIVE: An in silico approach was used to quantify the distribution of plant food allergen sequences across protein families and to develop and apply a novel means of assessing conserved surface features important for IgE cross-reactivity. METHODS: Plant food allergen sequences were classified into Pfam families on the basis of sequence homology. Contact surface areas of selected proteins were calculated with MOLMOL by using a 1.4-A probe, corrected by removing contributions from IgE inaccessible main chains and side chains forming the ligand binding sites. RESULTS: A set of 129 food allergen sequences were classified into only 20 of 3849 possible Pfam families, with 4 families accounting for more than 65% of food allergens. Structural bioinformatic analysis of conserved exterior main chains and amino acid side chains in cross-reactive homologues of Bet v 1 and nonspecific lipid transfer proteins showed higher levels of similarity than shown by simple sequence comparisons. Thus, 75% of the Mal d 1 surface is likely to bind anti-Bet v 1 antibodies, compared with a sequence identity of approximately 56%. CONCLUSION: Most plant food allergens belong to only 4 structural families, indicating that conserved structures and biological activities may play a role in determining or promoting allergenic properties. Structural bioinformatic analysis shows that conservation of 3-dimensional structure should be included in any assessment of potential IgE cross-reactivity in, for example, novel proteins.

Allergens↗

Multiomics Analysis Reveals Therapeutic Targets for Chronic Kidney Disease With Sarcopenia.

BACKGROUND: The presence of sarcopenia in patients with chronic kidney disease (CKD) is associated with poor prognosis. The mechanism underlying CKD-induced muscle wasting has not yet been fully explored. This study investigates the influence of renal secretions on muscles using multiomics sequencing. METHODS: The kidney transcriptome analysis by RNA-seq and protein profiling by tandem mass tag (TMT), serum TMT and muscle TMT were performed in CKD established using 0.2% adenine and control mice. Spp1 recombinant protein was used to study its effect on myotube atrophy in&#xa0;vitro. In animal experiments on CKD, pharmacological inhibition of Spp1 was used to explore the role of Spp1 in skeletal muscle wasting. Transcriptome analysis was performed to identify differentially expressed genes (DEGs) in the gastrocnemius muscle following Spp1 pharmacological inhibition. RESULTS: In the renal transcriptome and TMT, 503 and 377 proteins/genes respectively were co-upregulated and co-downregulated. In the serum TMT of CKD and normal control (NC) mice, 22 upregulated and 7 downregulated differentially expressed proteins (DEPs) showed the same expression patterns as those in the kidney transcriptome and TMT analysis. Based on bioinformatics analysis and reported studies, we selected Spp1 for further validation. Spp1 recombinant protein was added to C2C12 myotubes in&#xa0;vitro, and the results indicated that Spp1 significantly increased the protein levels of the muscle atrophy marker (Murf-1) and promoted the smaller myotubes (all p&#x2009;<&#x2009;0.05). Compared with NC mice, Spp1 mRNA and protein levels were significantly upregulated in the kidneys of CKD mice, and the serum concentration of Spp1 was also markedly increased (all p&#x2009;<&#x2009;0.05). In animal experiments, pharmacological inhibition of Spp1 increased the weights of gastrocnemius and tibialis anterior muscles (p&#x2009;<&#x2009;0.05) and improved muscle atrophy phenotype. Transcriptome analysis showed that DEGs in the gastrocnemius muscle following Spp1 pharmacological inhibition were enriched in protein digestion and absorption, glucagon signalling pathway, apelin signalling pathway and ECM-receptor interaction pathway. CONCLUSIONS: Our study is the first to establish a regulatory network of kidney-muscle crosstalk to explore the potential mechanism of CKD-related sarcopenia. Employing multiomics analysis, cellular assessment and animal experiments, we have identified that Spp1 could potentialy serve as a promising therapeutic target for CKD patients with sarcopenia.

Sarcopenia↗

Comparison of transcripts in Phalaenopsis bellina and Phalaenopsis equestris (Orchidaceae) flowers to deduce monoterpene biosynthesis pathway.

BACKGROUND: Floral scent is one of the important strategies for ensuring fertilization and for determining seed or fruit set. Research on plant scents has hampered mainly by the invisibility of this character, its dynamic nature, and complex mixtures of components that are present in very small quantities. Most progress in scent research, as in other areas of plant biology, has come from the use of molecular and biochemical techniques. Although volatile components have been identified in several orchid species, the biosynthetic pathways of orchid flower fragrance are far from understood. We investigated how flower fragrance was generated in certain Phalaenopsis orchids by determining the chemical components of the floral scent, identifying floral expressed-sequence-tags (ESTs), and deducing the pathways of floral scent biosynthesis in Phalaneopsis bellina by bioinformatics analysis. RESULTS: The main chemical components in the P. bellina flower were shown by gas chromatography-mass spectrometry to be monoterpenoids, benzenoids and phenylpropanoids. The set of floral scent producing enzymes in the biosynthetic pathway from glyceraldehyde-3-phosphate (G3P) to geraniol and linalool were recognized through data mining of the P. bellina floral EST database (dbEST). Transcripts preferentially expressed in P. bellina were distinguished by comparing the scent floral dbEST to that of a scentless species, P. equestris, and included those encoding lipoxygenase, epimerase, diacylglycerol kinase and geranyl diphosphate synthase. In addition, EST filtering results showed that transcripts encoding signal transduction and Myb transcription factors and methyltransferase, in addition to those for scent biosynthesis, were detected by in silico hybridization of the P. bellina unigene database against those of the scentless species, rice and Arabidopsis. Altogether, we pinpointed 66% of the biosynthetic steps from G3P to geraniol, linalool and their derivatives. CONCLUSION: This systems biology program combined chemical analysis, genomics and bioinformatics to elucidate the scent biosynthesis pathway and identify the relevant genes. It integrates the forward and reverse genetic approaches to knowledge discovery by which researchers can study non-model plants.

Acyclic Monoterpenes↗

[In silico data mining of the human programmed cell death 5 (PDCD5) sequences].

OBJECTIVE: To lay foundation for the functional studies of programmed cell death 5 (PDCD5) and develop new technical pathway for bioinformatics analysis of human functional genes. METHODS: Using PDCD5 as the target molecule, intensive bioinformatics analysis of the nucleic acid and protein sequences were conducted. Data mining and comprehensive analysis by sequence against database similarity searching, ortholog structure comparison, expression profile analysis and gene "neighbor" listing were performed. RESULTS: Two human putative pseudogenes on chromosomes 12 and 5, and one mouse putative pseudogene on chromosome 1 were identified. The methanobacterium thermoautotrophicum ortholog was classified as the same fold as ubiquitin and ribosomal protein S13. The C. elegans ortholog, ubiquitin and IAP (inhibitor of apoptosis proteins) belonged to the same expression profile cluster. This cluster was related to biosynthesis and protein synthesis. PDCD5 orthologs in various genomes were adjacent to various ribosomal proteins on the chromosome. CONCLUSION: The human genome contains at least two processed pseudogenes of PDCD5. Besides the relationship with cell apoptosis, PDCD5 is predicted to have functional relationship with ubiquitin and participate in the translation regulation.

Amino Acid Sequence↗

Identification of novel genes preferentially expressed in the retina using a custom human retina cDNA microarray.

PURPOSE: To construct a custom cDNA microarray for comprehensive human retinal gene expression profiling and apply it to the identification of genes that are preferentially expressed in the retina. METHODS: A cDNA microarray was constructed based on the predicted human retina gene expression profile according to expressed sequence tag (EST) databases. Gene expression profiles were obtained from five human retinas, two livers, and the cerebral cortical regions of two brains. Each sample was studied in duplicate, using a reference sample experimental design. Retina-enriched genes were identified by using the significance analysis for microarray (SAM) algorithm. Quantitative real time PCR was used to confirm microarray results. Bioinformatic analysis was performed to compare the array results with expression data available from public databases. RESULTS: The cDNA microarray contains 10,034 sequences: 67% represent known genes and 33% represent ESTs. Differential hybridization with the array identified, in addition to known retinal genes, 186 retina-enriched genes that do not have known retinal function. Of these, 96 represent novel genes. Quantitative real-time PCR of 11 of the identified genes and ESTs confirmed their retina-enriched expression pattern. Bioinformatic analysis of EST databases suggests that of the 186 genes, approximately 40% are predominantly expressed in the retina, whereas the remainder show significant expression in other tissues. Comparison of this study's microarray-based retina-enriched gene set with three published similar sets identified using complementary high-throughput approaches demonstrated only limited overlap of the identified genes. CONCLUSIONS: Because previous studies have demonstrated that many retina-enriched genes are crucial for maintaining normal retinal function, the genes identified here are likely to include ones that have important roles in the retina and ones that when mutated can cause or modulate retinal disease. In addition, the retina custom array should provide a useful resource for comparing expression profiles between normal and diseased human retinas.

Adult↗

Esophageal cancer in Chinese population: no polymorphism in codon 149 of P21(Waf1/Cip1) cyclin dependent kinase gene.

It has recently been suggested that people of the Indian population who carried the codon 149 polymorphism (GAT-->GGT) of P21(Waf1/Cip1) gene were more susceptible to esophageal cancer and oral cancer than the individuals without that polymorphism. Since esophageal cancer is a high incident neoplasm in China, we analysed the same codon of P21(Waf1/Cip1) in the Chinese population. Blood samples from 80 esophageal cancer patients and 80 normal blood donors were collected for DNA extraction. Methods of Polymerase Chain Reaction (PCR) and direct sequencing were used for detection of the polymorphism in codon 149 of P21(Waf1/Cip1). Bioinformatics analysis was also thoroughly performed for this gene. No polymorphism was found in all samples tested. Bioinformatics analysis revealed that the so-called polymorphism of codon 149 reported previously was a wrong one. In conclusion, no polymorphism exists in codon 149 of P21(Waf1/Cip1). It is not appropriate to use it as a susceptible site of the gene in cancer study.

Asian People↗

[Advances in the applications of DNA chips and proteomics approaches for neurosciences].

DNA chips and proteomics are two of the recently developed high-throughput technologies that allow us to simultaneously analyze the expression levels of multiple genes and the interactions of their products in the brain. Their applications in neuroscience provide us with the unprecedented opportunities for understanding of the brain. A typical gene chip experiment contains a series of steps, including preparations (or purchase) of DNA chips, target DNA and probes, hybridization of chips with target DNA, chip scanning, and imaging analysis. The technologies in proteomics are more complicated, which comprise of three main aspects of technologies: protein isolation, identification and bioinformatic analysis of data. If protein isolation is based on 2-D, it includes the preparation of protein sample, 2-D, staining gels, spot excision, proteolysis, mass spectrometry of purified protein and finally undergoing bioinformatic analysis. Here we review the two technologies with emphasis on their applications in neurosciences. The challenges, advantages and disadvantages of the two techniques, and perspectives for their developments are discussed.

Animals↗

[Tumor relevance analysis of a highly conserved gene by using gene microarray hybridization].

Preliminary function research of a highly conserved human gene,which was cloned from human fetal cDNA library during large-scale cDNA sequencing,is illustrated in this article. Bioinformatics analysis indicates that this gene is highly conserved in human, mouse, fruit fly, thaliana and fission yeast. Other bioinformatics analysis implies its relevance with tumors. RT-PCR analysis shows its wide-ranging expression patterns. Its expression in 16 cancer cases (including 7 liver cancer cases, 5 pancreas cancer cases, 2 larynx cancer cases and 2 lung cancer cases) is studied by using gene microarray analysis. The result shows its relevance with tumors and implies it may have different status in different classification of tumors.

English Abstract↗

hCLCA1 and mCLCA3 are secreted non-integral membrane proteins and therefore are not ion channels.

Proteins of the CLCA gene family have been proposed to mediate calcium-activated chloride currents. In this study, we used detailed bioinformatics analysis and found that no transmembrane domains are predicted in hCLCA1 or mCLCA3 (Gob-5). Further analysis suggested that they are globular proteins containing domains that are likely to be involved in protein-protein interactions. In support of the bioinformatics analysis, biochemical studies showed that hCLCA1 and mCLCA3, when expressed in HEK293 cells, could be removed from the cell surface and could be detected in the extracellular medium, even after short incubation times. The accumulation in the medium was shown to be brefeldin A-sensitive, demonstrating that hCLCA1 is constitutively secreted. The N-terminal cleavage products of hCLCA1 and mCLCA3 could be detected in bronchoalveolar lavage fluid taken from asthmatic subjects and ovalbumin-challenged mice, demonstrating release from cells in a physiological setting. We conclude that hCLCA1 and mCLCA3 are non-integral membrane proteins and therefore cannot be chloride channels in their own right.

Animals↗

Genome-wide expression studies of atherosclerosis: critical issues in methodology, analysis, interpretation of transcriptomics data.

During the past 6 years, gene expression profiling of atherosclerosis has been used to identify genes and pathways relevant in vascular (patho)physiology. This review discusses some critical issues in the methodology, analysis, and interpretation of the data of gene expression studies that have made use of vascular specimens from animal models and humans. Analysis of gene expression studies has evolved toward the genome-wide expression profiling of large series of individual samples of well-characterized donors. Despite the advances in statistical and bioinformatical analysis of expression data sets, studies have not yet fully exploited the potential of gene expression data sets to obtain novel insights into the molecular mechanisms underlying atherosclerosis. To assess the potential of published expression data, we compared the data of a CC chemokine gene cluster between 18 murine and human gene expression profiling articles. Our analysis revealed that an adequate comparison is mainly hindered by the incompleteness of available data sets. The challenge for future vascular genomic profiling studies will be to further improve the experimental design, statistical, and bioinformatical analysis and to make data sets freely accessible.

Animals↗

KEGG as a glycome informatics resource.

Bioinformatics approaches to carbohydrate research have recently begun using large amounts of protein and carbohydrate data. In this field called glycome informatics, the foremost necessity is a comprehensive resource for genome-scale bioinformatics analysis of glycan data. Although the accumulation of experimental data may be useful as a reference of biological and biochemical information on carbohydrates, this is insufficient for bioinformatics analysis. Thus, we have developed a glycome informatics resource (http://www.genome.jp/kegg/glycan/) in KEGG (Kyoto Encyclopedia of Genes and Genomes), an integrated knowledge base of protein networks, genomic information, and chemical information. This review describes three noteworthy features: (1) GLYCAN, a database of carbohydrate structures; (2) glycan-related pathways; and (3) Composite Structure Map (CSM), a map illustrating all possible variations of carbohydrate structures within organisms. GLYCAN includes two useful tools: an intuitive drawing tool called KegDraw, and an efficient glycan search and alignment tool called KEGG Carbohydrate Matcher (KCaM). KEGG's glycan biosynthesis and metabolism pathways, integrating carbohydrate structures, proteins, and reactions, are also a pivotal resource. CSM is constructed as a bridge between carbohydrate functions and structures. CSM is able to display, for example, expression data of glycosyltransferases in a compact manner. In all the KEGG resources, various objects including KEGG pathways, chemical compounds, as well as carbohydrate structures are commonly represented as graphs, which are widely studied and utilized in the computer science field.

Carbohydrates↗

ChimerDB--a knowledgebase for fusion sequences.

Chromosome translocation and gene fusion are frequent events in the human genome and are often the cause of many types of tumor. ChimerDB is the database of fusion sequences encompassing bioinformatics analysis of mRNA and expressed sequence tag (EST) sequences in the GenBank, manual collection of literature data and integration with other known database such as OMIM. Our bioinformatics analysis identifies the fusion transcripts that have non-overlapping alignments at multiple genomic loci. Fusion events at exon-exon borders are selected to filter out the cloning artifacts in cDNA library preparation. The result is classified into two groups--genuine chromosome translocation and fusion between neighboring genes owing to intergenic splicing. We also integrated manually collected literature and OMIM data for chromosome translocation as an aid to assess the validity of each fusion event. The database is available at http://genome.ewha.ac.kr/ChimerDB/ for human, mouse and rat genomes.

Animals↗

Application of in silico positional cloning and bioinformatic mutation analysis to the study of eye diseases.

A vast amount of DNA and protein sequence is now available and a plethora of programs have been developed to analyse the data. The bewildering variety of analyses that can be performed via the World-Wide Web can deter researchers from applying bioinformatics to augment their traditional genetic research. Focusing on the inherited eye diseases, this paper provides a guide to the appropriate software required for identification of candidate genes through to the detection and analysis of mutations.

Chromosome Mapping↗

Unveiling novel antimicrobial peptides from the ruminant gastrointestinal microbiomes: A deep learning-driven approach yields an anti-MRSA candidate.

INTRODUCTION: Antimicrobial peptides (AMPs) present a promising avenue to combat the growing threat of antibiotic resistance. The ruminant gastrointestinal microbiome serves as a unique ecosystem that offers untapped potential for AMP discovery. OBJECTIVES: The aims of this study are to develop an effective methodology for the identification of novel AMPs from ruminant gastrointestinal microbiomes, followed by evaluating their antimicrobial efficacy and elucidating the mechanisms underlying their activity. METHODS: We developed a deep learning-based model to identify AMP candidates from a dataset comprising 120 metagenomes and 10,373 metagenome-assembled genomes derived from the ruminant gastrointestinal tract. Both in vivo and in vitro experiments were performed to examine and validate the antimicrobial activities of the AMP candidates that were selected through bioinformatic analysis and subsequently synthesized chemically. Additionally, molecular dynamics simulations were conducted to explore the action mechanism of the most potent AMP candidate. RESULTS: The deep learning model identified 27,192 potential secretory AMP candidates. Following bioinformatic analysis, 39 candidates were synthesized and tested. Remarkably, all synthesized peptides demonstrated antimicrobial activity against Staphylococcus aureus, with 79.5% showing effectiveness against multiple pathogens. Notably, Peptide 4, which exhibited the highest antimicrobial activity against methicillin-resistant Staphylococcus aureus (MRSA), confirmed this effect in a mouse model with wound infection, exhibiting a low propensity for resistance development and minimal cytotoxicity and hemolysis towards mammalian cells. Molecular dynamics simulations provided insights into the mechanism of Peptide 4, primarily its ability to disrupt bacterial cell membranes, leading to cell death. CONCLUSION: This study highlights the power of combining deep learning with microbiome research to uncover novel therapeutic candidates, paving the way for the development of next-generation antimicrobials like Peptide 4 to combat the growing threat of MRSA would infections. It also underscores the value of utilizing ruminant microbial resources.

Animals↗

Elevated circulating IL-8 correlates with poor prognosis in urological cancers: a meta-analysis and bioinformatic validation.

BACKGROUND: Interleukin-8 (IL-8) is a key cytokine that has been implicated in multiple aspects of cancer progression and therapeutic resistance. Elevated levels of circulating IL-8&#xa0;(cIL-8) have been implicated in adverse clinical outcomes among patients with urological cancers. However, definitive evidence consolidating these observations remains lacking. The present study aims to synthesize the existing research findings to provide a comprehensive, evidence-based reference for clinical practice. METHODS: A systematic literature search was conducted to identify relevant studies that reported on the prognostic impact of cIL-8 levels in urological cancer patients. Hazard ratios (HRs) for overall survival (OS) and progression-free survival (PFS) were extracted and pooled to estimate the overall effect. Furthermore, Kaplan-Meier's survival analyses were conducted using RNA-seq data from The Cancer Genome Atlas (TCGA) through the Gene Expression Profiling Interactive Analysis 2 (GEPIA 2) online tool to validate the observed associations. RESULTS: A total of 19 cohorts encompassing 2740 patients from 12 studies were included in the meta-analysis. The findings revealed that elevated cIL-8 levels were significantly associated with inferior OS (HR: 1.86; 95% confidence intervals (CI): 1.72-2.02) and PFS (HR: 1.59; 95%CI: 1.25-2.03) in patients with urological cancers. The consistency and validity of these results were further supported by survival analyses performed using the GEPIA 2 tool. CONCLUSIONS: This study, which is the first meta-analysis to systematically examine the prognostic significance of cIL-8 in urological cancers, supported by bioinformatics validation, confirms that elevated cIL-8 levels serve as a potential biomarker for predicting adverse outcomes. Our findings underscore the importance of targeting IL-8 as a therapeutic strategy to overcome treatment resistance and improve outcomes for urological cancer patients. Further research into IL-8-targeted therapies and their integration into clinical practice is urgently needed to enhance the treatment landscape for urological cancers.

Humans↗

RNF115 aggravates tumor progression through regulation of CDK10 degradation in thyroid carcinoma.

BACKGROUND: RING Finger Protein 115 (RNF115), a notable E3 ligase, is known to modulate tumorigenesis and metastasis. In our investigation, we endeavor to unravel the putative function and inherent mechanism through which RNF115 influences the evolution of thyroid carcinoma (THCA). METHODS: We analyzed RNF115 expression in THCA using the Cancer Genome Atlas (TCGA) database. The influence of RNF115 on the progression of THCA was evaluated using both in vitro and in vivo experimental approaches. The protein regulated by RNF115 was identified through bioinformatics analysis, and its biological significance was further explored. RESULTS: In both THCA tissues and cells, RNF115 showed elevated expression levels. Enhanced expression of RNF115 fostered cell proliferation, tumor growth, and the exacerbation of epithelial-mesenchymal transition (EMT) in THCA, while also promoting tumor lung metastasis. Bioinformatics analysis identified cyclin-dependent kinase 10 (CDK10) as a downstream target of RNF115, which was found to be ubiquitinated and degraded by RNF115 in THCA cells. Functionally, overexpression of CDK10 was found to counteract the promotion of malignant phenotype in THCA induced by RNF115. From a mechanistic perspective, RNF115 activated the Raf-1 pathway and enhanced cancer cell cycle progression by degrading CDK10 in THCA cells. CONCLUSION: RNF115 triggers cell proliferation, EMT, and tumor metastasis by ubiquitinating and degrading CDK10. The regulation of the Raf-1 pathway and cell cycle progression in THCA may be profoundly influenced by this process.

Humans↗

Imprint of evolutionary conservation and protein structure variation on the binding function of protein tyrosine kinases.

MOTIVATION: According to the models of divergent molecular evolution, the evolvability of new protein function may depend on the induction of new phenotypic traits by a small number of mutations of the binding site residues. Evolutionary relationships between protein kinases are often employed to infer inhibitor binding profiles from sequence analysis. However, protein kinases binding profiles may display inhibitor selectivity within a given kinase subfamily, while exhibiting cross-activity between kinases that are phylogenetically remote from the prime target. The emerging insights into kinase function and evolution combined with a rapidly growing number of publically available crystal structures of protein kinases complexes have motivated structural bioinformatics analysis of sequence-structure relationships in determining the binding function of protein tyrosine kinases. RESULTS: In silico profiling of Imatinib mesylate and PD-173955 kinase inhibitors with protein tyrosine kinases is conducted on kinome scale by using evolutionary analysis and fingerprinting inhibitor-protein interactions with the panel of all publically available protein tyrosine kinases crystal structures. We have found that sequence plasticity of the binding site residues alone may not be sufficient to enable protein tyrosine kinases to readily evolve novel binding activities with inhibitors. While evolutionary signal derived solely from the tyrosine kinase sequence conservation can not be readily translated into the ligand binding phenotype, the proposed structural bioinformatics analysis can discriminate a functionally relevant kinase binding signal from a simple phylogenetic relationship. The results of this work reveal that protein conformational diversity is intimately linked with sequence plasticity of the binding site residues in achieving functional adaptability of protein kinases towards specific drug binding. This study offers a plausible molecular rationale to the experimental binding profiles of the studied kinase inhibitors and provides a theoretical basis for constructing functionally relevant kinase binding trees.

Amino Acid Sequence↗

Putative lipoproteins of Streptococcus agalactiae identified by bioinformatic genome analysis.

Streptococcus agalactiae is a significant pathogen causing invasive disease in neonates and thus an understanding of the molecular basis of the pathogenicity of this organism is of importance. N-terminal lipidation is a major mechanism by which bacteria can tether proteins to membranes. Lipidation is directed by the presence of a cysteine-containing 'lipobox' within specific signal peptides and this feature has greatly facilitated the bioinformatic identification of putative lipoproteins. We have designed previously a taxon-specific pattern (G+LPP) for the identification of Gram-positive bacterial lipoproteins, based on the signal peptides of experimentally verified lipoproteins (Sutcliffe I.C. and Harrington D.J. Microbiology 148: 2065-2077). Patterns searches with this pattern and other bioinformatic methods have been used to identify putative lipoproteins in the recently published genomes of S. agalactiae strains 2603/V and NEM316. A core of 39 common putative lipoproteins was identified, along with 5 putative lipoproteins unique to strain 2603/V and 2 putative lipoproteins unique to strain NEM316. Thus putative lipoproteins represent ca. 2% of the S. agalactiae proteome. As in other Gram-positive bacteria, the largest functional category of S. agalactiae lipoproteins is that predicted to comprise of substrate binding proteins of ABC transport systems. Other roles include lipoproteins that appear to participate in adhesion (including the previously characterised Lmb protein), protein export and folding, enzymes and several species-specific proteins of unknown function. These data suggest lipoproteins may have significant roles that influence the virulence of this important pathogen.

Amino Acid Sequence↗