PubMed HealthSearch

SEARCH · PubMed Health

Results for “Open data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Privacy-preserving framework for genomic computations via multi-key homomorphic encryption.

MOTIVATION: The affordability of genome sequencing and the widespread availability of genomic data have opened up new medical possibilities. Nevertheless, they also raise significant concerns regarding privacy due to the sensitive information they encompass. These privacy implications act as barriers to medical research and data availability. Researchers have proposed privacy-preserving techniques to address this, with cryptography-based methods showing the most promise. However, existing cryptography-based designs lack (i) interoperability, (ii) scalability, (iii) a high degree of privacy (i.e. compromise one to have the other), or (iv) multiparty analyses support (as most existing schemes process genomic information of each party individually). Overcoming these limitations is essential to unlocking the full potential of genomic data while ensuring privacy and data utility. Further research and development are needed to advance privacy-preserving techniques in genomics, focusing on achieving interoperability and scalability, preserving data utility, and enabling secure multiparty computation. RESULTS: This study aims to overcome the limitations of current cryptography-based techniques by employing a multi-key homomorphic encryption scheme. By utilizing this scheme, we have developed a comprehensive protocol capable of conducting diverse genomic analyses. Our protocol facilitates interoperability among individual genome processing and enables multiparty tests, analyses of genomic databases, and operations involving multiple databases. Consequently, our approach represents an innovative advancement in secure genomic data processing, offering enhanced protection and privacy measures. AVAILABILITY AND IMPLEMENTATION: All associated code and documentation are available at https://github.com/farahpoor/smkhe.

Computer Security

Dihydrolipoamide dehydrogenase from Haloferax volcanii: gene cloning, complete primary structure, and comparison to other dihydrolipoamide dehydrogenases.

We used the N-terminal amino acid sequence of dihydrolipoamide dehydrogenase from Haloferax volcanii, to design and synthesize two oligonucleotide probes that were used to identify and clone a 4.3 kilobase pair (kbp) fragment from MboI restriction endonuclease digestion of Hf. volcanii genomic DNA. The nucleotide sequence of a 1.5-kbp region of this clone was determined and this revealed an open reading frame that translated into a protein with good homology to dihydrolipoamide dehydrogenase from other sources. The first 48 amino acids were identical with the N-terminal sequence data obtained from the purified protein. The complete primary structure of the halophilic dihydrolipoamide dehydrogenase was analyzed in terms of its homologies to dihydrolipoamide dehydrogenases from other sources and its molecular adaptations to high intracellular ionic strength.

Amino Acid Sequence

The nucleotide sequence between genes 31 and 30 of bacteriophage T4.

The nucleotide sequence of the 2994 bp T4 phage DNA fragment between genes 31 and 30 is presented. The fragment contains 7 complete open reading frames in the direction of early transcription and two early promoters, PE128.6 and PE128.2, which we show to cause difficulties in cloning DNA from this genomic region. Our data complete the nucleotide sequence and the organization of genes in the genomic region between T4 genes 31 and 30.

Amino Acid Sequence

Molecular cloning and characterisation of the 22-kilodalton adult Schistosoma mansoni antigen recognised by antibodies from mice protectively vaccinated with isolated tegumental surface membranes.

A cDNA clone from an adult Schistosoma mansoni lambda gt11 expression library (A12) encoding an antigenic polypeptide of 22 kDa is described. A12 is 797 bp long and has one open reading frame encoding a protein of 190 amino acids which does not contain a signal sequence or membrane anchor motif and has no homologies with any sequences on the currently available data bases. Its product (sm22.6) is recognised by antibodies from mice protectively vaccinated with purified adult S. mansoni tegumental membranes and by serum from S. mansoni-infected Brazilians. It is present in all post-snail life cycle stages except the egg, is not sex-specific, and is found in 9 species of Schistosoma, but not in a range of other helminths. Data are presented which suggest that sm22.6 is a soluble, peripheral membrane protein.

Amidohydrolases

Analysis of the UL36 open reading frame encoding the large tegument protein (ICP1/2) of herpes simplex virus type 1.

Using peptide antisera specific for regions within the N terminus and C terminus of the predicted UL36 gene product, immunoblotting experiments were performed to demonstrate definitively that ICP1/2 is encoded by the UL36 gene. These data also suggest that both the cell- and the virion-associated forms of ICP1/2 are colinear with the complete predicted amino acid sequence of the UL36 gene. Computer-assisted analyses of the predicted amino acid sequence of the UL36 gene revealed the presence of two putative leucine zipper-type motifs and a potential ATP-binding domain. The possible functions of these consensus domains will also be discussed.

Amino Acid Sequence

Genome organization and nucleotide sequence of human papillomavirus type 39.

The 7833-bp nucleotide sequence of human papillomavirus type 39 (HPV39), which is associated with genital intraepithelial neoplasias and invasive carcinomas, has been determined. The genome organization deduced from the sequence shares characteristic features with other genital papillomaviruses. According to sequence comparisons, HPV39 most closely resembles HPV18 and may be a member of a subgroup of genital papillomaviruses distinct from the HPV16/31/33 group. As a novel feature, we report a 1.3-kb open reading frame on the DNA strand which lacks major open reading frames in the other sequenced HPV genomes.

Base Sequence

DNA sequence and analysis of a cryptic 4.2-kb plasmid from the filamentous cyanobacterium, Plectonema sp. strain PCC 6402.

The 4194-bp plasmid, pRF1, from Plectonema sp. Strain PCC 6402 was completely sequenced and analyzed. Seven potential open reading frames were identified. The predicted amino acid sequence of open reading frame C (ORF C) had identities of 34, 29, and 25% with Rep B from the Staphylococcus aureus plasmid, pUB110; Rep from the Bacillus amyloliquefaciens plasmid, pFTB14; and protein A from the S. aureus plasmid, pC194, respectively. A 75-amino-acid region conserved in these proteins (Rep B, Rep, and protein A) also was highly conserved in ORF C with identities of 45, 37, and 40%, respectively. Significantly, 16 of the 21 amino acids conserved in Rep B, Rep, and protein A were found at the same positions in ORF C. This ORF may encode a replication protein that includes a region conserved in some eubacteria. Additional structural features include a 425-bp region that contains palindromes, tandem repeats, and short direct repeats which may correspond to the origin of replication. An 18-bp inverted repeat was located between two open reading frames, A and G.

Amino Acid Sequence

HINN: Hierarchical Input Neural Network identifies multi-omics biomarker for cognitive decline.

Understanding complex diseases requires models that can integrate diverse layers of biological data while yielding insights that are biologically interpretable. Although multi-omics integration with machine learning (ML) has advanced disease prediction and biomarker discovery, most existing approaches overlook the hierarchical and regulatory relationships that connect these molecular layers. Here, we present the Hierarchical Input Neural Network (HINN), a deep learning framework that incorporates known cross-omics relationships directly into its architecture, capturing the flow of information from genomics to epigenomics, transcriptomics, and downstream biological processes. By embedding these relationships, HINN improves both predictive performance and biological interpretability. We applied HINN to blood-derived multi-omics data from individuals with Alzheimer's disease or mild cognitive impairment to predict cognitive scores from standardized assessments. HINN outperformed both baseline and state-of-the-art models and pinpointed multi-omics biomarkers-including SNPs and promoter-region CpG sites in ATP6V1C1 and RCHY1 -that were significantly correlated with plasma p-Tau181 levels. These features map to biologically relevant processes with potential implications for cognitive decline. Our findings demonstrate how combining deep learning with biological knowledge can uncover interpretable, blood-based biomarkers for cognitive decline due to complex diseases such as Alzheimer's. All code and data are openly available at https://github.com/bozdaglab/HINN.

Alzheimer’s disease

Identification and sequencing of the Choristoneura biennis entomopoxvirus DNA polymerase gene.

A degenerate oligonucleotide probe corresponding to a highly conserved amino acid sequence in several DNA polymerases was used to locate the DNA polymerase gene in the Choristoneura biennis entomopoxvirus. Southern blot analysis of the entomopoxvirus genome using the degenerate oligonucleotide probe showed specific interaction between the probe and an eight kilobasepair EcoRI fragment from the entomopoxvirus genome. Sequencing this EcoRI fragment revealed an open reading frame 2892 nucleotides in length, capable of encoding a protein about 115 kilodaltons. Homology search of this open reading frame against other proteins indicated a high degree of homology in four distinct regions with DNA polymerases from other organisms. The highest degree of homology (24.9% at the amino acid level) was found between the vaccinia DNA polymerase gene and the entomopoxvirus open reading frame.

Amino Acid Sequence

Expression of the gene encoded by a family of macronuclear chromosomes generated by alternative DNA processing in Oxytricha fallax.

Hypotrichous ciliated protozoa, such as Oxytricha fallax, produce tiny chromosomes during generation of the transcriptionally active macronucleus. The 81-MAC family of macronuclear chromosomes is produced by alternative DNA processing, such that the chromosomes share a common region of 1.6 kbp. Transcription of a 1.3 kb mRNA from the common region has been analyzed. Transcription starts very near the telomere (34 bp), in a 23 bp region of pure A + T DNA. Polyadenylation sites are very near the other telomere (26 bp), also in a region of nearly pure A + T DNA. Three introns are clustered in the first third of the gene. Intron removal can follow polyadenylation, and the order of removal is not fixed. All three known sequence versions of the 81-MAC chromosomes are represented in the mRNA pool, with no evidence of any further versions. The A + T sequences surrounding the transcription starts and polyadenylation sites are conserved among versions. Introns have conserved 5' and 3' ends and a putative branch-point sequence (YYRAT), but otherwise are highly diverged and are AT-rich. A single long open reading frame, interrupted by the three introns, encodes a homolog of known mitochondrial solute carriers, and contains the codon TAA, which does not encode 'stop,' but a conserved glutamine; TAG appears also to encode glutamine. The results significantly enlarge the small data set of transcription start and polyadenylation sites, of intron features, and of translation signals for hypotrichs.

Animals

The complete sequence of a 6146 bp fragment of Saccharomyces cerevisiae chromosome III contains two new open reading frames.

As part of the EEC project to sequence the entire chromosome III of Saccharomyces cerevisiae we have sequenced a total of 11,040 bp from near the right end of the chromosome. A new protein kinase gene was found at one extremity of the sequenced region (Wilson et al., 1992), while the previously sequenced actin binding protein gene, ABP1, (Drubin et al., 1990) was found at the other extremity. We present here the sequence of the region between these two genes which has the potential to code for two new open reading frames (ORFs).

Amino Acid Sequence

The TSM1 gene of Saccharomyces cerevisiae overlaps the MAT locus.

We have cloned the region from MAT to THR4 on chromosome III of Saccharomyces cerevisiae. Although the region is only 15 kb, the two loci are genetically separated by 22 cM. This is in sharp contrast to the very low level of recombination (2 cM in 22 kb) that is observed in the adjacent CRY1-MAT interval, and suggests that there may be a "hot spot" for recombination in the MAT-THR4 region. The DNA sequence of the first 4.4 kb distal to MAT reveals an open reading frame that we have identified as the essential gene, TSM1. Surprisingly, the TSM1 open reading frame of 1,410 amino acids extends into the MAT locus, such that the 3'-end of the MAT alpha 1 transcript ends 15 bp from the 3'-end of the TSM1 open reading frame.

Alleles

Sequence of the CDC10 region at chromosome III of Saccharomyces cerevisiae.

A 4.74 kb DNA fragment from the right arm of chromosome III of Saccharomyces cerevisiae, adjacent to the centromere region was sequenced. Four open reading frames with an ATG initiation codon and larger than 200 bp were found in this fragment. The largest open reading frame of 966 bp was identified as the CDC10 gene.

Base Sequence

Cloning and analysis of YMR26, the nuclear gene for a mitochondrial ribosomal protein in Saccharomyces cerevisiae.

The nuclear gene for a mitochondrial ribosomal protein, termed YMR26, of Saccharomyces cerevisiae strain DC-5 was cloned by hybridization with synthetic oligonucleotide mixtures corresponding to the N-terminal amino acid sequence of this protein. The gene was found to occur in a single copy on either chromosome VII or chromosome XV. The nucleotide sequence of the cloned segment containing this gene showed the presence of an open reading frame capable of encoding a basic protein of 18.5 kDa with 158 amino acid residues. The deduced amino acid sequence showed no significant similarity to any known ribosomal proteins of prokaryotic or eukaryotic origin or to any other proteins in the NBRF protein data bank. When the gene was disrupted by insertion of a 2.9 kb restriction fragment containing LEU2, cells became PET- indicating that the gene is essential for yeast mitochondria. Northern blot analysis indicated that the size of the transcript from the YMR26 gene was approximately 530 nucleotides long. The expression level of the YMR26 gene was monitored upon catabolite repression, in strains with various mitochondrial genetic backgrounds and in strains harboring an increased dosage of the YMR26 gene. In rho+ cells, the transcription of the YMR26 gene was more repressed in a medium with glucose than in the presence of either galactose or nonfermentable carbon sources. However, in rho o cells, its transcription appeared not to be repressed even by high concentrations of glucose. The amount of the YMR26 mRNA was increased 10-fold when cells carried the YMR26 gene on a high-copy number plasmid.

Amino Acid Sequence

Eukaryotic initiation factor (eIF)-4F. Implications for a role in internal initiation of translation.

In order to study the eukaryotic translation initiation mechanisms of "internal initiation," "re-initiation," and/or "coupled internal initiation," a series of model mRNAs have been constructed which contain two non-overlapping open reading frames (ORFs) that encode different lengths of rabbit alpha globin. These mRNAs, along with the bicistronic constructs TK/CAT and TK/P2CAT developed by Pelletier and Sonenberg (Pelletier, J., and Sonenberg, N. (1988) Nature 334, 320-325, 1988), were used to program an in vitro rabbit reticulocyte lysate translation system. Cap-dependent and cap-independent translation were distinguished by monitoring translation in the presence or absence of exogenously added cap analog (m7GTP). Messenger RNAs which translate both ORF1 and ORF2 by a cap-dependent mechanism, as well as mRNAs that translate ORF2 by a cap-independent mechanism while still translating ORF1 in a cap-dependent fashion have been obtained. These same alpha globin mRNAs differ by no more than 45 nucleotides in intercistronic length. Initiation factor addition studies were performed in this same in vitro translation system. Both eukaryotic initiation factor (eIF)-4F and, to a lesser extent, eIF-4B can stimulate translation of an internally located ORF independent of upstream ORF translation and in a manner not dependent on mRNA cap recognition. This indicates that the cap-recognition initiation factor, eIF-4F, and eIF-4B facilitate cap-independent and internal initiation of an open reading frame.

Animals

The DNA sequence and structural organization of the GC2 plasmid from the red alga Gracilaria chilensis.

The complete DNA sequence of the circular GC2 plasmid from the red alga Gracilaria chilensis was obtained. It contains 3827 bp and has a base composition of 75% A + T nucleotides. The sequence revealed that GC2 has two inverted repeats, each of 290 nucleotides, that border four long direct tandem repeats of 216 nucleotides. five short, direct, tandem repeats of 21-22 nucleotides were also found in the plasmid. The plasmid sequence has one major open reading frame that could encode a 411 amino acid polypeptide. Finally, the GC2 plasmid is transcriptionally active.

Amino Acid Sequence