PubMed Health⌕ Search

Biomedical subjects

Praveen F Cherukuri

Publications and source records attributed to Praveen F Cherukuri.

3 recordsLinked to original sources

Tandem splice acceptor sites: Profiling their relevance to human disease.

PURPOSE: Interpretation of variation, particularly the creation or disruption of tandem splice acceptor sites (NAGNnAG variants), challenges genomic medicine practice. METHODS: We analyzed the creation and disruption of dinucleotide AG sites within ±30 bases of natural splice-acceptor sites in the GRCh37 human reference genome. These results were compared with variant data from the ClinVar and gnomAD databases, as well as with data from 779 National Institutes of Health Undiagnosed Diseases Program study participants. Using RNA sequencing, we assessed the splicing at NAGNnAG variants for 107 of the Undiagnosed Diseases Program participants and compared the empirical data with SpliceAI predictions. RESULTS: Creation or disruption of NAGNnAG sites within 30 bases of the natural splice acceptor are enriched in ClinVar compared with gnomAD; however, such variants in the 2 databases are rarely differentiated by SpliceAI scores. Empirical evaluation via RNA sequencing analysis supported novel acceptor site usage from -21 to +30; splice-altering variants did not predominate in a specific region or have SpliceAI scores invariantly, suggesting increased spliceogenicity. CONCLUSION: NAGNnAG variants within 30 bp of the natural splice acceptor have a high probability of clinical relevance and are poorly contextualized for clinical utility. Their interpretation benefits from empirical evaluation via RNA analysis.

Humans↗

Co-evolutionary analysis of domains in interacting proteins reveals insights into domain-domain interactions mediating protein-protein interactions.

Recent advances in functional genomics have helped generate large-scale high-throughput protein interaction data. Such networks, though extremely valuable towards molecular level understanding of cells, do not provide any direct information about the regions (domains) in the proteins that mediate the interaction. Here, we performed co-evolutionary analysis of domains in interacting proteins in order to understand the degree of co-evolution of interacting and non-interacting domains. Using a combination of sequence and structural analysis, we analyzed protein-protein interactions in F1-ATPase, Sec23p/Sec24p, DNA-directed RNA polymerase and nuclear pore complexes, and found that interacting domain pair(s) for a given interaction exhibits higher level of co-evolution than the non-interacting domain pairs. Motivated by this finding, we developed a computational method to test the generality of the observed trend, and to predict large-scale domain-domain interactions. Given a protein-protein interaction, the proposed method predicts the domain pair(s) that is most likely to mediate the protein interaction. We applied this method on the yeast interactome to predict domain-domain interactions, and used known domain-domain interactions found in PDB crystal structures to validate our predictions. Our results show that the prediction accuracy of the proposed method is statistically significant. Comparison of our prediction results with those from two other methods reveals that only a fraction of predictions are shared by all the three methods, indicating that the proposed method can detect known interactions missed by other methods. We believe that the proposed method can be used with other methods to help identify previously unrecognized domain-domain interactions on a genome scale, and could potentially help reduce the search space for identifying interaction sites.

Amino Acid Sequence↗

CDD: a Conserved Domain Database for protein classification.

The Conserved Domain Database (CDD) is the protein classification component of NCBI's Entrez query and retrieval system. CDD is linked to other Entrez databases such as Proteins, Taxonomy and PubMed, and can be accessed at http://www.ncbi.nlm.nih.gov/entrez/query.fcgi?db=cdd. CD-Search, which is available at http://www.ncbi.nlm.nih.gov/Structure/cdd/wrpsb.cgi, is a fast, interactive tool to identify conserved domains in new protein sequences. CD-Search results for protein sequences in Entrez are pre-computed to provide links between proteins and domain models, and computational annotation visible upon request. Protein-protein queries submitted to NCBI's BLAST search service at http://www.ncbi.nlm.nih.gov/BLAST are scanned for the presence of conserved domains by default. While CDD started out as essentially a mirror of publicly available domain alignment collections, such as SMART, Pfam and COG, we have continued an effort to update, and in some cases replace these models with domain hierarchies curated at the NCBI. Here, we report on the progress of the curation effort and associated improvements in the functionality of the CDD information retrieval system.

Amino Acid Sequence↗