PubMed Health⌕ Search

Biomedical subjects

Denis C Shields

Publications and source records attributed to Denis C Shields.

18 recordsLinked to original sources

Open and sustainable AI: challenges, opportunities and the road ahead in the life sciences.

Artificial intelligence (AI) has seen transformative breakthroughs in the life sciences, expanding possibilities to interpret biological information at an unprecedented capacity. To maximize return on growing investments and accelerate progress, it is urgent to address long-standing research challenges arising from the rapid adoption of AI methods. We review the erosion of trust in AI outputs driven by poor reusability and reproducibility, and highlight their impact on environmental sustainability. Furthermore, we discuss the fragmented components of the AI ecosystem and lack of guiding pathways to support open and sustainable AI model development. In response, this Perspective introduces practical open and sustainable AI recommendations mapped to over 300 ecosystem components and provides guiding implementation pathways. Our work connects researchers with relevant AI resources, facilitating the implementation of sustainable, reusable and reproducible AI. Built upon community consensus and aligned to existing efforts, these outputs will aid future policy development and structured pathways for guiding AI implementation.

Artificial Intelligence↗

The impact of genetic variation in the region of the GPIIIa gene, on Pl expression bias and GPIIb/IIIa receptor density in platelets.

Some studies have suggested that genetic variability in the glycoprotein (GP) IIIa gene modulates expression of platelet GPIIb/IIIa (alpha(2b)beta(3)). We sought to determine as to whether combinations of genetic variants within the GPIIIa gene (haplotypes) influenced the expression of GPIIIa RNA and protein levels in human platelets. Three promoter polymorphisms, Pl(A1/A2) genotype and platelet receptor densities were determined in 207 acute coronary syndrome (ACS) patients. Allele-specific quantitative reverse transcription-polymerase chain reaction of platelet RNA from Pl(A1/A2) heterozygotes identified a greater expression of Pl(A2) bearing transcripts among heterozygotes. Among the patients studied, the ratio of Pl(A1)/Pl(A2) RNA expression was significantly influenced by promoter haplotype (P < 0.01). However, this effect reflected carriership of rare not common haplotypes (P = 0.2). There was a threefold variation between subjects in the number of GPIIb/IIIa receptors expressed per platelet, although no association between receptor density and the Pl(A2) (P = 0.93) or promoter polymorphisms was demonstrated (-468A, P = 0.52; -425C, P = 0.59; -400A, P = 0.52). Among common haplotypes, Pl(A1)/Pl(A2) RNA expression was negatively correlated with adjusted GPIIb/IIIa receptor density (P = 0.04). The overall trend towards higher expression of Pl(A2) bearing message in Pl(A1/A2) heterozygotes, and the existence of rare haplotypes with more pronounced changes indicate the existence of cis-acting genetic factors that remain to be identified.

Analysis of Variance↗

Polymorphisms of the Flavin containing monooxygenase 3 (FMO3) gene do not predispose to essential hypertension in Caucasians.

BACKGROUND: The recessive disorder trimethylaminuria is caused by defects in the FMO3 gene, and may be associated with hypertension. We investigated whether common polymorphisms of the FMO3 gene confer an increased risk for elevated blood pressure and/or essential hypertension. METHODS: FMO3 genotypes (E158K, V257M, E308G) were determined in 387 healthy subjects with ambulatory systolic and diastolic blood pressure measurements, and in a cardiovascular disease population of 1649 individuals, 691(41.9%) of whom had a history of hypertension requiring drug treatment. Haplotypes were determined and their distribution noted. RESULTS: There was no statistically significant association found between any of the 4 common haplotypes and daytime systolic blood pressure in the healthy population (p = 0.65). Neither was a statistically significant association found between the 4 common haplotypes and hypertension status among the cardiovascular disease patients (p = 0.80). CONCLUSION: These results suggest that the variants in the FMO3 gene do not predispose to essential hypertension in this population.

Adult↗

BADASP: predicting functional specificity in protein families using ancestral sequences.

SUMMARY: Burst After Duplication with Ancestral Sequence Predictions (BADASP) is a software package for identifying sites that may confer subfamily-specific biological functions in protein families following functional divergence of duplicated proteins. A given protein phylogeny is grouped into subfamilies based on orthology/paralogy relationships and/or user definitions. Ancestral sequences are then predicted from the sequence alignment and the functional specificity is calculated using variants of the Burst After Duplication method, which tests for radical amino acid substitutions following gene duplications that are subsequently conserved. Statistics are output along with subfamily groupings and ancestral sequences for an easy analysis with other packages. AVAILABILITY: BADASP is freely available from http://www.bioinformatics.rcsi.ie/~redwards/badasp/

Algorithms↗

Tandem repeat copy-number variation in protein-coding regions of human genes.

BACKGROUND: Tandem repeat variation in protein-coding regions will alter protein length and may introduce frameshifts. Tandem repeat variants are associated with variation in pathogenicity in bacteria and with human disease. We characterized tandem repeat polymorphism in human proteins, using the UniGene database, and tested whether these were associated with host defense roles. RESULTS: Protein-coding tandem repeat copy-number polymorphisms were detected in 249 tandem repeats found in 218 UniGene clusters; observed length differences ranged from 2 to 144 nucleotides, with unit copy lengths ranging from 2 to 57. This corresponded to 1.59% (218/13,749) of proteins investigated carrying detectable polymorphisms in the copy-number of protein-coding tandem repeats. We found no evidence that tandem repeat copy-number polymorphism was significantly elevated in defense-response proteins (p = 0.882). An association with the Gene Ontology term 'protein-binding' remained significant after covariate adjustment and correction for multiple testing. Combining this analysis with previous experimental evaluations of tandem repeat polymorphism, we estimate the approximate mean frequency of tandem repeat polymorphisms in human proteins to be 6%. Because 13.9% of the polymorphisms were not a multiple of three nucleotides, up to 1% of proteins may contain frameshifting tandem repeat polymorphisms. CONCLUSION: Around 1 in 20 human proteins are likely to contain tandem repeat copy-number polymorphisms within coding regions. Such polymorphisms are not more frequent among defense-response proteins; their prevalence among protein-binding proteins may reflect lower selective constraints on their structural modification. The impact of frameshifting and longer copy-number variants on protein function and disease merits further investigation.

Frameshift Mutation↗

A sequence sub-sampling algorithm increases the power to detect distant homologues.

Searching databases for distant homologues using alignments instead of individual sequences increases the power of detection. However, most methods assume that protein evolution proceeds in a regular fashion, with the inferred tree of sequences providing a good estimation of the evolutionary process. We investigated the combined HMMER search results from random alignment subsets (with three sequences each) drawn from the parent alignment (Rand-shuffle algorithm), using the SCOP structural classification to determine true similarities. At false-positive rates of 5%, the Rand-shuffle algorithm improved HMMER's sensitivity, with a 37.5% greater sensitivity compared with HMMER alone, when easily identified similarities (identifiable by BLAST) were excluded from consideration. An extension of the Rand-shuffle algorithm (Ali-shuffle) weighted towards more informative sequence subsets. This approach improved the performance over HMMER alone and PSI-BLAST, particularly at higher false-positive rates. The improvements in performance of these sequence sub-sampling methods may reflect lower sensitivity to alignment error and irregular evolutionary patterns. The Ali-shuffle and Rand-shuffle sequence homology search programs are available by request from the authors.

Algorithms↗

Overdispersion of allele frequency differences between populations: implications for meta-analyses of genotypic disease associations.

Methods correcting case-control studies of genetic polymorphisms for unmeasured genetic population substructure by modelling the variation at a number of variant loci provide no standard and easily implemented approach to meta-analysis, which is a key to understanding the effects of minor genotypic risks on complex diseases. A correction of the odds ratio estimate and its confidence interval is shown to be easy to implement using a mixed effects logistic regression. The method is shown to substantially reduce bias and to give accurate coverage even when there is substantial overdispersion of allele frequency differences between populations. Major sequence classes of single-nucleotide polymorphism (SNP) are likely to act as valid controls for each other, since CpG SNPs did not differ in the extent of population structure from other SNPs. Agreement among investigators and journals to provide these straightforward statistics in publications of polymorphism studies will enhance the ability of future investigators to perform meta-analyses of weak genetic effects across accumulated studies that allow for population structure.

Case-Control Studies↗

Uroplakin III is not a major candidate gene for primary vesicoureteral reflux.

Vesicoureteral reflux (VUR) is the retrograde flow of urine from the bladder into the ureter and towards the kidneys. VUR is the most common cause of end stage renal failure in both children and adults and it is a major cause of severe hypertension in children. VUR is seen in approximately 1-2% of newborn Caucasians. Substantial evidence exists that VUR is a genetic disorder. Uroplakins are integral membrane proteins found in the bladder wall. Knockout studies in mice have suggested uroplakin III (UPK3) as a candidate gene for VUR. We have used parametric and nonparametric linkage analysis and tests for association, to investigate this possibility in a cohort of 126 sibling pairs affected with primary VUR. None of the analyses showed any substantial evidence for linkage or association of markers at the UPK3 locus to VUR. Our results do not support a role for UPK3 in primary VUR.

Child↗

Genetic stratification of pathogen-response-related and other variants within a homogeneous Caucasian Irish population.

Selection pressures from pathogens impact on the worldwide geographic distribution of polymorphisms in certain pathogen-response-associated genes. Such gene-specific effects could lead to confounding by geographic disease associations. We wished to determine if such constraints impinge on the genetic structure of a population of Irish patients and whether variants associated with responses to pathogens showed greater stratification. The counties of origin of each subject's grandparents were used as the geographic variable. F(st), proportional to the extent of population structure, was low (mean F(st)=0.004 across 25 SNPs, range 0.001-0.008) and it was not significantly higher for pathogen response SNPs (F(st)=0.004) than for other SNPs (F(st)=0.003, P=0.21). Correspondence analysis revealed weak trends primarily in approximately northeast to southwest and secondarily in northwest to southeast directions. One-dimensional spatial autocorrelation analysis revealed a weak (Moran's I autocorrelation of -0.10) tendency for SNP frequencies to diverge with greater distance. Two-dimensional autocorrelation indicated a northeast to southwest gradient that was similar for both the pathogen response and other SNPs. The southeastern county, Wexford, showed a distinctive pattern, perhaps consistent with Anglo-Norman settlements. In conclusion, these results indicate that pathogen response SNPs do not exhibit significantly more population structure than other SNPs within this Caucasian population. This suggests that the specific population structure of particular genes may not typically be a cause of strong confounding in genetic studies where population structure is controlled.

Arylsulfotransferase↗

GASP: Gapped Ancestral Sequence Prediction for proteins.

BACKGROUND: The prediction of ancestral protein sequences from multiple sequence alignments is useful for many bioinformatics analyses. Predicting ancestral sequences is not a simple procedure and relies on accurate alignments and phylogenies. Several algorithms exist based on Maximum Parsimony or Maximum Likelihood methods but many current implementations are unable to process residues with gaps, which may represent insertion/deletion (indel) events or sequence fragments. RESULTS: Here we present a new algorithm, GASP (Gapped Ancestral Sequence Prediction), for predicting ancestral sequences from phylogenetic trees and the corresponding multiple sequence alignments. Alignments may be of any size and contain gaps. GASP first assigns the positions of gaps in the phylogeny before using a likelihood-based approach centred on amino acid substitution matrices to assign ancestral amino acids. Important outgroup information is used by first working down from the tips of the tree to the root, using descendant data only to assign probabilities, and then working back up from the root to the tips using descendant and outgroup data to make predictions. GASP was tested on a number of simulated datasets based on real phylogenies. Prediction accuracy for ungapped data was similar to three alternative algorithms tested, with GASP performing better in some cases and worse in others. Adding simple insertions and deletions to the simulated data did not have a detrimental effect on GASP accuracy. CONCLUSIONS: GASP (Gapped Ancestral Sequence Prediction) will predict ancestral sequences from multiple protein alignments of any size. Although not as accurate in all cases as some of the more sophisticated maximum likelihood approaches, it can process a wide range of input phylogenies and will predict ancestral sequences for gapped and ungapped residues alike.

Amino Acid Sequence↗

Elevated white cell count in acute coronary syndromes: relationship to variants in inflammatory and thrombotic genes.

BACKGROUND: Elevated white blood cell counts (WBC) in acute coronary syndromes (ACS) increase the risk of recurrent events, but it is not known if this is exacerbated by pro-inflammatory factors. We sought to identify whether pro-inflammatory genetic variants contributed to alterations in WBC and C-reactive protein (CRP) in an ACS population. METHODS: WBC and genotype of interleukin 6 (IL-6 G-174C) and of interleukin-1 receptor antagonist (IL1RN intronic repeat polymorphism) were investigated in 732 Caucasian patients with ACS in the OPUS-TIMI-16 trial. Samples for measurement of WBC and inflammatory factors were taken at baseline, i.e. Within 72 hours of an acute myocardial infarction or an unstable angina event. RESULTS: An increased white blood cell count (WBC) was associated with an increased C-reactive protein (r = 0.23, p < 0.001) and there was also a positive correlation between levels of beta-fibrinogen and C-reactive protein (r = 0.42, p < 0.0001). IL1RN and IL6 genotypes had no significant impact upon WBC. The difference in median WBC between the two homozygote IL6 genotypes was 0.21/mm3 (95% CI = -0.41, 0.77), and -0.03/mm3 (95% CI = -0.55, 0.86) for IL1RN. Moreover, the composite endpoint was not significantly affected by an interaction between WBC and the IL1 (p = 0.61) or IL6 (p = 0.48) genotype. CONCLUSIONS: Cytokine pro-inflammatory genetic variants do not influence the increased inflammatory profile of ACS patients.

Acute Disease↗

Significantly different patterns of amino acid replacement after gene duplication as compared to after speciation.

We have performed a large-scale analysis of amino acid sequence evolution after gene duplication by comparing evolution after gene duplication with evolution after speciation in over 1,800 phylogenetic trees constructed from manually curated alignments of protein domains downloaded from the PFAM database. The site-specific rate of evolution is significantly altered by gene duplication. A significant increase in the proportion of amino acid substitutions at constrained (slowly evolving) sites after duplication was observed. An increase in the proportion of replacements at normally constrained amino acid sites could result from relaxation of purifying selective pressure. However, the proportion of amino acid replacements involving radical changes in amino acid properties after duplication does not appear to be significantly increased by relaxed selective pressure. The increased proportion of replacements at constrained sites was observed over a relatively large range of protein change (up to 25% amino acid replacements per site). These findings have implications for our understanding of the nature of evolution after duplication and may help to shed light on the evolution of novel protein functions through gene duplication.

Algorithms↗

Wrapping up BLAST and other applications for use on Unix clusters.

UNLABELLED: We have developed two programs that speed up common bioinformatic applications by spreading them across a UNIX cluster.(1) BLAST.pm, a new module for the 'MOLLUSC' package. (2) WRAPID, a simple tool for parallelizing large numbers of small instances of programs such as BLAST, FASTA and CLUSTALW. AVAILABILITY: The packages were developed in Perl on a 20-node Linux cluster and are provided together with a configuration script and documentation. They can be freely downloaded from http://wolfe.gen.tcd.ie/wrapper.

Computer Communication Networks↗

Genetic variability in the extracellular matrix as a determinant of cardiovascular risk: association of type III collagen COL3A1 polymorphisms with coronary artery disease.

Although common genetic variants in platelet collagen receptors influence platelet activation and thrombosis, the impact of polymorphisms in collagen genes on cardiovascular disease is unknown. To evaluate this, we genotyped a highly polymorphic intronic tandem repeat of the COL3A1 gene, encoding collagen type III, alpha 1. This revealed 4 common alleles (COL3A1-1, -2, -3, and -4). The 2 populations studied were as follows: (1) a cross-sectional study of 703 acute coronary syndrome (ACS) patients with myocardial infarction (MI) and unstable angina, and (2) a prospective study of 924 Caucasian patients from the OPUS (Orbofiban in Patients with Unstable coronary Syndromes)-TIMI-16 trial of the oral GPIIb/IIIa antagonist orbofiban. In addition, we studied 306 control subjects and 224 patients with stable angina. In the case-control population, COL3A1-4 carriers were protected against ACS (odds ratio [OR] = 0.57, 95% CI = 0.35-0.91, P =.02) and stable angina (OR = 0.35, 95% CI = 0.16-0.74, P =.006). In the OPUS population, allele 4 again appeared protective against composite end points (death, MI, stroke, recurrent ischemia, and urgent rehospitalization) (relative risk [RR] = 0.41, 95% CI = 0.17-1.00). There were significant interactions between COL3A1-1 and -3 variants and treatment. Allele COL3A1-3 was associated with an increased risk of the composite end point (RR = 1.65, 95% CI = 1.07-2.55) in patients randomized to orbofiban, but appeared protective in placebo patients (RR = 0.53, 95% CI = 0.28-0.98). We conclude that variants in the COL3A1 gene, the product of which is a vessel-wall protein and platelet ligand, modulate the risk of coronary artery disease and could also modulate the response to antithrombotic therapy. This is the first reported association between polymorphisms of extracellular matrix components and cardiovascular risk.

Alanine↗

Platelet glycoprotein Ib alpha receptor polymorphisms and recurrent ischaemic events in acute coronary syndrome patients.

AIMS: To examine the relationship between polymorphisms in the platelet receptor glycoprotein (GP) Ib(alpha) and recurrent ischaemic events, and assess their impact on response to anti-platelet treatment. METHODS AND RESULTS: 1014 patients presenting with unstable coronary syndrome were recruited from the OPUS-TIMI 16 clinical trial of the platelet GPIIb/IIIa antagonist, orbofiban. The subjects were genotyped for two polymorphisms in the gene for GPIb(alpha). These were a T-5C polymorphism in the 5' untranslated Kozak region of the GPIb(alpha) gene, and the variable number of tandem repeats (VNTR) in the macroglycopeptide region.165 patients had events (recurrent ischaemia, urgent revascularisation, myocardial infarction (MI), stroke and death). There was no effect of the number of -5C alleles on composite endpoint frequency among Caucasian subjects (test for trend, p = 0.47). However, MI risk increased with the number of -5C alleles carried, with MI occurring in 2.3% of patients with the -5T/-5T genotype, 5.0% of -5T/-5C, and 16.7% of -5C/-5C (p < 0.01). The effect of treatment on MI outcome was not significantly modified by genotype (test for interaction, p = 0.10). The overall risk of bleeding was not strongly influenced by either the -5C or the VNTR polymorphisms. CONCLUSION: In an unstable coronary syndrome population the T-5C polymorphism in GPIb(alpha) influences risk of subsequent MI.

Acute Disease↗

Improved database searches for orthologous sequences by conditioning on outgroup sequences.

MOTIVATION: Searches of biological sequence databases are usually focussed on distinguishing significant from random matches. However, the increasing abundance of related sequences on databases present a second challenge: to distinguish the evolutionarily most closely related sequences (often orthologues) from more distantly related homologues. This is particularly important when searching a database of partial sequences, where short orthologous sequences from a non-conserved region will score much more poorly than non-orthologous (outgroup) sequences from a conserved region. RESULTS: Such inferences are shown to be improved by conditioning the search results on the scores of an outgroup sequence. The log-odds score for each target sequence identified on the database has the log-odds score of the outgroup sequence subtracted from it. A test group of Caenorhabditis elegans kinase sequences and their identified C.elegans outgroups were searched against a test database of human Expressed Sequence Tag (EST) sequences, where the sets of true target sequences were known in advance. The outgroup conditioned method was shown to identify 58% more true positives ahead of the first false positive, compared to the straightforward search without an outgroup. A test dataset of 151 proteins drawn from the C.elegans genome, where the putative 'outgroup' was assigned automatically, similarly found 50% more true positives using outgroup conditioning. Thus, outgroup conditioning provides a means to improve the results of database searching with little increase in the search computation time.

Algorithms↗

Human tissue profiling with multidimensional protein identification technology.

Profiling of tissues and cell types through systematic characterization of expressed genes or proteins shows promise as a basic research tool, and has potential applications in disease diagnosis and classification. We used multidimensional protein identification protein identification technology (MudPIT) to analyze proteomes for enriched nuclear extracts of eight human tissues: brain, heart, liver, lung, muscle, pancreas, spleen, and testis. We show that the method is approximately 80% reproducible. We address issues of relative abundance, tissue-specificity, and selectivity, and the significance of proteins whose expression does not correlate with that of the corresponding mRNA. Surprisingly, most proteins are detected in a single tissue. These proteins tend to fulfill specialist (and potentially tissue-specific) functions compared to proteins expressed in two or more tissues.

Algorithms↗