PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Genetic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

HCSeeker: A classification tool for human genetic variant hot and cold spots designed for PM1 and benign criteria in the ACMG-AMP guideline.

PURPOSE: The PM1 criterion, which states that a variant is located in a mutational hot spot and/or critical and well-established functional domain without benign variation (such as the active site of an enzyme), is considered moderate evidence for assessing its pathogenicity. Although guidelines from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology are widely adopted, the PM1 criterion remains limited from lacking a reliable database of variant hot spots. Compared with hot spots, cold spots are neglected by the guidelines. To improve variant classification, we suggest including cold spots for supporting benign classifications. Consequently, we have developed the HCSeeker to provide data support for PM1 and the "Benign" criteria. METHODS: HCSeeker uses the Kernel Density Estimation and the Expectation-Maximization algorithm to identify hot- and cold-spot regions. RESULTS: Through HCSeeker, we identified 988 hot spots and 682 cold spots across 889 genes and provided a public database (http://www.genemed.tech/hcseeker/) for researchers and clinicians to query variant locations, facilitating the application of American College of Medical Genetics and Genomics and the Association for Molecular Pathology PM1 or "Benign" criteria. CONCLUSION: We developed the HCSeeker tool, which can effectively identify variant hot and cold spots within genes to enhance the interpretability of gene variants.

Humans

Distribution and complementarity of hydropathy in multisubunit proteins.

A survey of 40 multisubunit proteins and 2 protein-protein complexes was performed to assay quantitatively the distribution of hydropathy among the exterior surface, interior, contact surface, and noncontact exterior surface of the isolated subunits. We suggest a useful way to present this distribution by using a "hydropathy level diagram." Additionally, we have devised a function called "hydropathy complementarity" to quantitate the degree to which interacting surfaces have matching hydropathy distributions. Our survey revealed the following patterns: (1) The difference in hydropathy between the interior and exterior of subunits is a fairly invariant quantity. (2) On average, the hydropathy of the contact surface is higher than that of the exterior surface, but is not greater than that of the protein as a whole. There was variation, however, among the proteins. In some instances, the contact surface was more hydrophilic than the noncontact exterior, and in a few cases the contact surface was as hydrophobic as the protein interior. (3) The average interface manifests significant hydropathy complementarity, signifying that proteins interact by placing hydrophobic centers of one surface against hydrophobic centers of the other surface, and by similarly matching hydrophilic centers. As a measure of recognition and specificity, hydropathy complementarity could be a useful tool for predicting correct docking of interacting proteins. We suggest that high hydropathy complementarity is associated with static inflexible interactions. (4) We have found that some subunits that bind predominantly through hydrophilic forces, such as hydrogen bonds, ionic pairs, and water and metal bridges, are involved in dynamic quaternary organization and allostery.

Animals

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence

Large scale database experiments to assess the significance of matching DNA profiles.

Over 5,700 three-probe VNTR DNA profiles collected by several United Kingdom (UK) laboratories have been compared to examine the probability of randomly matching 2 samples from different individuals. In over 16 million comparisons, using a matching rule corresponding to the matching guideline employed by the UK Forensic Science Service, no profiles were found to match at the 3 loci D1S7 (MS1), D7S21(MS31) and D12S11 (MS43a). The frequency of occurrence of a set of Caucasian profiles have been estimated with 6 reference databases. The results show that there were greater differences in the frequency estimates when using a database of Afro-Caribbean or Asian profiles, rather than a different Caucasian database. The results further demonstrate the power and robustness of the VNTR DNA profiling technique for forensic casework.

Bayes Theorem

The branching order of mammals: phylogenetic trees inferred from nuclear and mitochondrial molecular data.

In order to clarify some controversial phylogenies such as those regarding the triplet of human, rodent, and cow and the evolutionary position of Lagomorpha with respect to other mammals, we have analyzed both nuclear and mitochondrial genes using the stationary Markov model developed in our laboratory. We found that the two sets of genes give different results. In particular the mitochondrial tree showed rabbit linked first to rodents and the rabbit-rodents branch linked to artiodactyls with human as the outgroup. The most favorite nuclear tree showed human linked first to artiodactyls and the human-artiodactyls branch linked to rabbit with rodents as the outgroup. The obvious questions, (1) which tree is the correct one, or (2) both trees can be incorrect, and (3) how can we explain such an evolutionary pattern, are discussed on the basis of our limited knowledge of factors that influence the clocklike behavior of biological macromolecules.

Animals

Importance of purine and pyrimidine content of local nucleotide sequences (six bases long) for evolution of the human immunodeficiency virus type 1.

Human immunodeficiency virus type 1 evolves rapidly, and random base change is thought to act as a major factor in this evolution. However, segments of the viral genome differ in their variability: there is the highly variable env gene, particularly hypervariable regions located within env, and, in contrast, the conservative gag and pol genes. Computer analysis of the nucleotide sequences of human immunodeficiency virus type 1 isolates reveals that base substitution in this virus is nonrandom and affected by local nucleotide sequences. Certain local sequences 6 base pairs long are excessively frequent in the hypervariable regions. These sequences exhibit base-substitution hotspots at specific positions in their 6 bases. The hotspots tend to be nonsilent letters of codons in the hypervariable regions--thus leading to marked amino acid substitutions there. Conversely, in the conservative gag and pol genes the hotspots tend to be silent letters because of a difference in codon frame from the hypervariable regions. Furthermore, base substitutions in the local sequences that frequently appear in the conservative genes occurred at a low level, even within the variable env. Thus, despite the high variability of this virus, the conservative genes and their products could be conserved. These may be some of the strategies evolved in human immunodeficiency virus type 1 to allow for positive-selection pressures, such as the host immune system, and negative-selection pressures on the conservative gene products.

Acquired Immunodeficiency Syndrome

Corruption of genomic databases with anomalous sequence.

We describe evidence that DNA sequences from vectors used for cloning and sequencing have been incorporated accidentally into eukaryotic entries in the GenBank database. These incorporations were not restricted to one type of vector or to a single mechanism. Many minor instances may have been the result of simple editing errors, but some entries contained large blocks of vector sequence that had been incorporated by contamination or other accidents during cloning. Some cases involved unusual rearrangements and areas of vector distant from the normal insertion sites. Matches to vector were found in 0.23% of 20,000 sequences analyzed in GenBank Release 63. Although the possibility of anomalous sequence incorporation has been recognized since the inception of GenBank and should be easy to avoid, recent evidence suggests that this problem is increasing more quickly than the database itself. The presence of anomalous sequence may have serious consequences for the interpretation and use of database entries, and will have an impact on issues of database management. The incorporated vector fragments described here may also be useful for a crude estimate of the fidelity of sequence information in the database. In alignments with well-defined ends, the matching sequences showed 96.8% identity to vector; when poorer matches with arbitrary limits were included, the aggregate identity to vector sequence was 94.8%.

Base Sequence

The Oxford Grid.

The 'Oxford Grid' is a term used to denote a method of displaying homologous loci from two species. It may also be used as a basis for historical inferences relating to co-ancestry. These uses are discussed with special reference to man and mouse. It is inferred that as few as 30 reciprocal translocations are sufficient to explain the differences defined by the present grid and that the telocentric karyotype, through which the mouse differs from both man and many closely related rodents, including the rat, must have evolved mainly through the formation and suppression of centromeres rather than through pericentric inversions. Mouse and man share the widest distribution of any mammal and their success appears to be related to being omnivorous with behavioural modifications allowing a wide range of habitat.

Animals

An estimate of the sequencing error frequency in the DNA sequence databases.

We have examined vector sequences fortuitously present in the EMBL sequence database as contaminating parts of submitted sequences, and found a sequencing error frequency of 3.55% in this subset of release 27 of the database. We discuss the possibility that this value may be representative for corresponding errors in the database as a whole.

Base Sequence

A variant database.

Variant biomacromolecules, either natural or artificially created, are proving to be useful in determining structure-function relationships. Research in this field would be facilitated by the compilation of data of variants into a central and widely available repository. The Variant Database acts as a repository for data concerning variant molecules. It complements the Biological Activity and the Physicochemical Property Databases as well as provides variant sequences.

Amino Acid Sequence

[DNA Data Bank].

Explore the source record for details and available documents.

Animals

The Saccharomyces Genome Database-a history of ideas and accomplishments, 1994-2026.

The Saccharomyces Genome Database (SGD) is one of the longest-running and most consequential biological databases in the world. Founded in the early 1990s at Stanford University under the visionary leadership of David Botstein and developed under the long-term technical direction of J. Michael Cherry, SGD has served for more than three decades not only as the authoritative knowledge center for the budding yeast Saccharomyces cerevisiae, but also as the source for much of the fundamentals of eukaryotic biology. This history traces the arc of a remarkable intellectual and scientific project: beginning with the challenge of building the very first integrated eukaryotic genome database and evolving across 30 years into a global knowledge hub for genetics, functional genomics, and human disease research. The history is organized chronologically, with each section highlighting the central ideas, technical developments, and concrete accomplishments of that period.

Databases, Genetic

PAHG: the database of human multi-gene families.

BACKGROUND: In the early vertebrate history, gene duplications, including single-gene, segmental-gene (SSD), and whole-genome duplication (WGD), formed multigene families. Despite efforts to classify metazoan multigene families hierarchically for evolutionary insight, a gap exists in accessible, curated resources for human/vertebrate multigene families. RESULTS: Addressing this, we present the Phylogenomic Analysis of Human Genome (PAHG) database. It focuses on curated multigene families in the human genome, particularly within four paralogons: HOX-bearing (Hsa:2/7/12/17), FGFR-bearing (Hsa:4/5/8/10), MHC-bearing (Hsa:1/6/9/19), and chromosomes 1/2/8/20. CONCLUSION: The current PAHG version details the phylogenetic history of 221 human multigene families (1247 gene members) with 15,231 protein sequences from diverse metazoans. It provides insights into gene duplication timings, co-duplication events, and their relationships with human genome syntenic organization. The PAHG database addresses the lack of accessible resources, offering valuable information on human/vertebrate multigene family evolution. Access the PAHG database at: https://www.pahgncb.com/ and http://pahg.qau.edu.pk/ . This resource enriches our understanding of vertebrate genetic evolution.

Humans

Integration of gene maps: updating chromosome 1.

The first integrated map of chromosome 1 was published in 1992. We present an updated summary map of 371 loci constructed from a location database that includes physical and genetic data. The summary map subsumes a composite physical location, sex-specific genetic location, cytogenetic assignment, mouse homology, rank and references to physical maps. The genetic length is 208 cM for the male map, in close agreement with the chiasma map, and 371 cM for the female map. There is evidence for a high level of interference on chromosome 1. The location database comprising both data and analytical software is discussed in relation to alternative approaches and possible enhancements.

Algorithms