PubMed HealthSearch

SEARCH · PubMed Health

Results for “cDNA sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Genetic analysis of four cases of Poirier Bienvenu neurodevelopmental syndrome associated with CSNK2B variant.

BACKGROUND: CSNK2B deficiency underlies the pathogenesis of Poirier-Bienvenu neurodevelopmental syndrome (POBINDS). In this study, we present four cases of pediatric seizures caused by de novo variants in CSNK2B, with the aim to reinforce the clinical and variant data pertaining to early genetic factors associated with epilepsy. METHODS: Trio whole exome sequencing were used to detect variants in the proband and her family members, and bioinformatics annotation was performed for the variant. Sanger sequencing and CSNK2B cDNA sequencing were employed to ascertain the carrier status of additional family members and evaluate the potential impact of variants on splicing. RESULTS: All four cases presented with epilepsy as the initial manifestation, accompanied by global developmental delay, particularly in language and motor developmental delay. Cases 1, 3 and 4 exhibited full-scale tonic-clonic seizures, while case 2 displayed myoclonic and typical absence seizures. Furthermore, case 2 demonstrated delayed growth and development compared to age-matched peers. No abnormality was detected in the head magnetic resonance imaging (MRI). Genetic analysis revealed novel heterozygous variants in the CSNK2B gene in all four cases, including c.175 + 1G > A, c.73-2A > G, c.291 + 1G > A and c.481delA. In case 2, reverse transcription analysis of CSNK2B mRNA revealed the retention of the 3' end sequence of Intron 2 and deletion of the 5' end sequence of Exon 3. In treatment, four case received a combination of one to three types of antiseizure medication and rehabilitation training individually. Case 1 continued to experience seizures to varying degrees, while cases 2-4 demonstrated effective seizure control. Overall motor and intellectual development improved in all four cases, however, there was slow recovery in language function. CONCLUSION: This study elucidates the molecular etiology of epilepsy in four cases with POBINDS and expands the mutational spectrum of pathogenic variants in the CSNK2B, highlighting their impact on splicing. The highly genetic heterogeneous phenotype of POBINDS relies on the detection of pathogenic variants in CSNK2B. Conventional antiseizure medication effectively control seizures, while rehabilitation treatment can significantly improve intelligence and motor function to varying degrees; however, language recovery tends to be relatively slow.

Humans

Clinical and functional characterization of a novel homozygous non-canonical splice mutation (c.1910-15_1910-11delinsTTACA) in CEP290 causing Joubert syndrome.

BACKGROUND: Joubert syndrome (JS) is a rare, predominantly autosomal recessive neurodevelopmental disorder characterized by hypotonia, motor delay, intellectual disability, oculomotor apraxia, and the hallmark "molar tooth sign" on axial view of MRI. JS is genetically heterogeneous, with pathogenic variants identified in more than 40 genes involved in primary cilia function. Among these, CEP290 is one of the most frequently mutated genes. RESULTS: In this study, we investigated two children-an 11-year-old boy (the proband) and his 5-year-old sister-both presenting with a similar phenotype consistent with JS. The parents, who self-identified as Chechen, reported distant consanguinity. The family also included a healthy 13-year-old daughter. The proband had previously been evaluated by a neurologist and underwent whole-genome sequencing (WGS); however, no causative variants were identified initially. After phenotype reassessment by a clinical geneticist, we performed a reanalysis of the raw WGS data and identified a novel homozygous intronic variant of uncertain significance (VUS), c.1910-15_1910-11delinsTTACA in CEP290 (NM_025114.4). Sanger sequencing confirmed that both the proband and his affected sister were homozygous for this variant, which they inherited from their heterozygous parents. Their healthy sister did not carry the variant. mRNA-sequencing and targeted cDNA sequencing (read depth ~ 100,000x) demonstrated that this intronic variant causes completely aberrant splicing of CEP290 pre-mRNA. Predominantly this variant causes the skipping of exon 20 in the main CEP290 transcript. Alternatively, the variant results in partial inclusion of intron 19 into the mRNA, elongation of exon 20 by 58 nucleotides, and a homozygous substitution chr12:88114573 (ACTGTGTA> TTACAGTA). No canonical mRNA isoform was detected when the variant was homozygous. Both the predicted severe truncation and the likely degradation of aberrant transcripts through nonsense-mediated decay (NMD) would correspond to complete loss of CEP290 function. Following the reclassification of this VUS to likely pathogenic, the family was able to pursue in vitro fertilization (IVF) with preimplantation genetic testing for monogenic disorders (PGT-M). CONCLUSION: Our study highlights the critical importance of proper phenotyping prior to referral for WES/WGS as well as of combining NGS with functional mRNA studies to achieve a molecular diagnosis for patients with predicted splice-site mutations in JS-associated genes. It also emphasizes the need for functional reassessment of VUS when genomic data are expected to guide reproductive decision-making within affected families.

Humans

AnoEST: toward A. gambiae functional genomics.

Here, we present an analysis of 215,634 EST and cDNA sequences of a major vector of human malaria Anopheles gambiae structured into the AnoEST database. The expressed sequences are grouped into clusters using genomic sequence as template and associated with inferred functional annotation, including the following: corresponding Ensembl gene prediction, putative orthologous genes in other species, homology to known proteins, protein domains, associated Gene Ontology terms, and corresponding classification into broad GO-slim functional groups. AnoEST is a vital resource for interpretation of expression profiles derived using recently developed A. gambiae cDNA microarrays. Using these cDNA microarrays, we have experimentally confirmed the expression of 7961 clusters during mosquito development. Of these, 3100 are not associated with currently predicted genes. Moreover, we found that clusters with confirmed expression are nonbiased with respect to the current gene annotation or homology to known proteins. Consequently, we expect that many as yet unconfirmed clusters are likely to be actual A. gambiae genes. [AnoEST is publicly available at http://komar.embl.de, and is also accessible as a Distributed Annotation Service (DAS).].

Animals

[Genetic and functional characterization of a novel KIT splicing variant in a Chinese three-generation pedigree with piebaldism].

OBJECTIVES: To investigate the genetic etiology of a three-generation pedigree affected with piebaldism. METHODS: Next-generation sequencing and Sanger sequencing were employed to detect and verify gene variants. Bioinformatics tools were used to predict the effects of candidate variants on splicing and protein function. RT-PCR and Sanger sequencing were further performed to validate the impact of the variant on RNA splicing, and homology modeling was applied to predict its effect on the three-dimensional structure of the KIT protein. The pathogenicity of the variant was then classified according to the guidelines of the American College of Medical Genetics and Genomics (ACMG) and the UK Association for Clinical Genomic Science (ACGS). RESULTS: A heterozygous insertion variant near the splice site, c.1990+8_1990+9insTGCACCATTGGAGGTAAA, was identified in the KIT gene in the proband and was found to co-segregate with the phenotype within the family. RT-PCR and cDNA sequencing revealed that this variant led to aberrant splicing during transcription, resulting in a 21 bp in-frame insertion in the mRNA, which encodes an extra 7 amino acids within the tyrosine kinase domain and may thus affect protein function. In silico predictions, together with the experimental findings, supported classification of this variant as likely pathogenic according to relevant variant interpretation guidelines. CONCLUSIONS: The heterozygous splice-site insertion variant KIT:c.1990+8_1990+9insTGCACCATTGGAGGTAAA is the genetic cause of piebaldism in this pedigree.

Genetics diagnosis

Assessment of different promoters in lentiviral vectors for expression of the N-acetyl-galactosamine-6-sulfate sulfatase gene.

Mucopolysaccharidosis IVA (MPS IVA) is caused by pathogenic variants in the GALNS gene encoding N-acetylgalactosamine-6-sulfate sulfatase (GALNS) enzyme, leading to glycosaminoglycan (GAG) accumulation in multiple tissues, resulting in progressive skeletal dysplasia and poor quality of life. There is currently no effective treatment for this skeletal disease. This study proposes a novel lentiviral vector (LV)-based gene therapy that produces and secretes the active GALNS enzyme at supraphysiologic levels within the cells. LVs carrying the native GALNS encoding sequence (cDNA) were made under three different promoters: CBh, COL2A1, and CD11b. Moreover, we designed LVs carrying the native GALNS cDNA tagged with D8 octapeptide under the CD11b promoter and a human codon-optimized GALNS cDNA under the CBh promoter, respectively. Transduced HEK293 cells, HepG2 cells, and MPS IVA fibroblasts and chondrocytes were cultured for 8 and 30 days, and the media were collected every three days. The enzyme activity, GAG levels, and vector copy numbers (VCNs) in these cells and media were analyzed. LV with the COL2A1 promoter produced the highest enzyme activity in HEK293, HepG2, MPS IVA fibroblasts, and chondrocytes, followed by LV with the CBh promoter. VCNs were higher in MPS IVA fibroblasts treated with LV-CBh-hGALNS and in HepG2 cells treated with LV-CD11b-hGALNS than in HEK293 cells. Accumulated GAGs were normalized to wild-type levels by the LV gene therapy, especially with CBh and COL2A1 promoters. These findings, if further validated, could significantly impact the treatment of MPS IVA, offering a more effective and feasible treatment option.

Humans

Cloning and characterization of H4 (D10S170), a gene involved in RET rearrangements in vivo.

H4(D10S170) is a gene which we isolated because of its frequent rearrangement with the RET proto-oncogene in vivo. Its fusion to RET generates the RET/PTC1 oncogene, which has been detected in about 20% of human thyroid papillary carcinomas. We have cloned and sequenced the cDNA corresponding to the H4(D10S170) gene from a human normal thyroid cDNA library. The nucleotide sequence of the H4(D10S170) 3 kb transcript shows no significant homology to known genes and contains an open reading frame (ORF) of 585 amino acids. H4(D10S170) predicted protein has no transmembrane domain and shows extensive regions in the alpha helical conformation, which are 30% homologous to the alpha-helical domains of several proteins including tropomyosin, vimentin, keratin and the tail region of myosin heavy chain. A putative SH3 binding site is present at the carboxy terminus, which suggests that H4(D10S170) might be a cytoskeletal protein.

Amino Acid Sequence

Hybridization capture increases on-target nanopore sequencing of plant RNA tobamovirus- derived cDNA libraries.

High-throughput sequencing (HTS) can support plant virus surveillance, but host nucleic acids often reduce on-target read recovery. We evaluated a targeted hybridization-capture workflow in which barcoded double-stranded cDNA (ds-cDNA) libraries generated from plant RNA extracts spiked with lyophilized tobamovirus-positive controls were enriched before Oxford Nanopore sequencing. Biotinylated probes targeted conserved regions of cucumber green mottle mosaic virus (CGMMV), species Tobamovirus viridimaculae; pepper mild mottle virus (PMMoV), species Tobamovirus capsici; and tobacco mosaic virus (TMV), species Tobamovirus tabaci. Across four pairs per virus, relative target-read abundance increased after capture from 0.76 ± 0.33% to 37.62 ± 15.72% for CGMMV, 8.16 ± 3.86% to 24.68 ± 12.34% for PMMoV, and 15.62 ± 10.40% to 36.83 ± 30.33% for TMV. Exact two-sided Wilcoxon signed-rank tests yielded P = 0.125 for each virus; with four nonzero differences in a common direction, this was the minimum attainable two-sided P value. Genome-coverage breadth was maintained. Retrospective duplex qPCR supported an increased virus-to-18S ratio for CGMMV, showed a variable PMMoV response, and showed a decreased virus-to-18S ratio for TMV because the 18S signal shifted earlier by as much as or more than the TMV signal. The findings provide proof-of-concept evidence for target-dependent library enrichment but do not establish analytical sensitivity, diagnostic performance, or field validity. Validation with naturally infected, low-titer, and mixed-infection samples and comparison with simpler targeted workflows are required.

biosecurity

Molecular characterization of RET/PTC3; a novel rearranged version of the RETproto-oncogene in a human thyroid papillary carcinoma.

The RET proto-oncogene encodes a transmembrane receptor of the tyrosine kinase family and has frequently been found activated in human thyroid carcinomas of the papillary subtype. In most cases the activation consisted of the fusion of its tyrosine-kinase domain with the 5'-terminal region of a gene designated H4 or D10S170. We have named the resulting H4/RET chimeric oncogene RET/PTC. Another activated form of the RET oncogene has subsequently been found in a thyroid carcinoma and is now referred to as RET/PTC2. Here we report the identification and cloning of a novel rearranged version of the RET oncogene in a human thyroid papillary carcinoma. In this case the tyrosine-kinase domain of RET was fused to a sequence 790 bp long belonging to a new gene that we have named RFG (RET Fused Gene). This novel chimeric oncogene has been designated RET/PTC3. In order to have more insights into the function of RFG we have completely cloned and sequenced its cDNA. RFG predicted amino-acid sequence does not have any significant homology to any already known genes and is ubiquitously expressed in human and mouse tissues. Finally we provide evidence indicating that the rearrangement leading to the generation of RET/PTC3 occurred in vivo in the original tumor DNA.

Amino Acid Sequence

RNA Sequencing Protocols for Short-Read Sequencing.

RNA sequencing (RNA-seq) methodologies allow the discovery of novel variants and transcripts. These comprise three general steps: (1) capture of RNA species of interest, (2) conversion of RNA to complementary DNA (cDNA), and (3) modification of cDNA to fit the sequencing platform. Here we describe four different library preparation protocols for short-read sequencing: cDNA synthesis with poly(A) selection, library preparation with ribosomal depletion, and cDNA synthesis with SMART® (Switching Mechanism at 5' end of RNA Template) technology for low and Pico inputs.

Gene Library

Long-read transcriptomics corrects Trichomonas vaginalis intron annotations and refines transcript-end features.

BACKGROUND: Trichomonas vaginalis causes the most prevalent non-viral sexually transmitted infection worldwide. Despite its large genome (181.5 Mb; 36,310 predicted protein-coding genes in NYU_TvagG3_2), intron annotations remain limited and inconsistently validated. A recent short-read RNA-seq study reported 63 putative active introns, but short reads can misassign splice boundaries and cannot resolve complete transcript structures. METHODS: We integrated Oxford Nanopore direct RNA sequencing (DRS), ONT cDNA long-read sequencing, and Illumina RNA-seq to refine intron annotations, transcript-end features, and UTR boundaries in T. vaginalis. Candidate introns were validated by targeted PCR and Sanger sequencing, and representative splicing events were further assessed using public SRA datasets. RESULTS: Starting from 31 historically annotated introns, motif-guided long-read screening and orthogonal validation identified 17 additional validated introns, increasing the curated set to 48 confirmed introns. Among these 17 events, three were previously unrecognized in the current NYU_TvagG3_2 reference annotation. We also corrected five reported loci, including two false-positive introns, two splice-coordinate misannotations, and one gene-sequence error. DRS further supported transcript termination site mapping, UAAA polyadenylation-signal profiling relative to poly(A) addition sites, and single-molecule poly(A)-tail estimation. StringTie mixed-mode assemblies provided updated UTR boundaries for intron-bearing transcripts and transcripts without curated introns. CONCLUSIONS: This study provides a rigorously validated, long-read-refined resource of intron annotations, UTR boundaries, and UAAA-guided transcript-end features for T. vaginalis, together with a reproducible workflow for non-model protists. These refinements improve the current reference annotation and support future studies of functional genomics, parasite biology, pathogenesis, and diagnostic development.

Trichomonas vaginalis

Transcriptomic and RNAi analyses reveal chloride channel 3-associated osmoregulation in Litopenaeus vannamei under low-salinity stress.

Chloride channels and transporters are important for cellular volume regulation and salinity adaptation in euryhaline crustaceans, yet the intestinal transcriptional relationship between plasma-membrane and intracellular chloride pathways remains unclear in Litopenaeus vannamei. In this study, RNA interference of anoctamin 1 (ANO1) was combined with intestinal transcriptome sequencing under the production-relevant low-salinity condition of salinity 3. ANO1 silencing produced a focused transcriptional response, with 16 differentially expressed genes (DEGs) identified (11 upregulated and 5 downregulated). Functional enrichment indicated that these genes were associated with transporter activity, cytoskeletal organization, extracellular matrix-receptor interaction, membrane lipid metabolism, and vesicular processes. Notably, a transcript encoding chloride channel protein 3 (CLC-3) was significantly upregulated following ANO1 knockdown, suggesting a potential transcriptional relationship between ANO1 and CLC-3 in chloride homeostasis. Based on this finding, CLC-3 was selected for full-length cDNA cloning, sequence characterization, salinity-gradient expression analysis, and RNAi-based functional assessment. The cloned CLC-3 cDNA was 2883 bp in length and encoded an 850 amino acid protein containing a conserved voltage-gated chloride channel (Voltage-CLC) domain and two cystathionine β-synthase domains. Phylogenetic analysis placed LvCLC-3 within the intracellular CLC-c clade, and tissue distribution analysis showed the highest CLC-3 expression in the intestine. Intestinal CLC-3 expression responded nonlinearly to salinity variation, peaking at salinity 20. Under salinity 3, CLC-3 knockdown reduced ANO1, Na+/K+-ATPase alpha subunit, and Na+-K+-2Cl- cotransporter transcript levels, whereas glutamate-gated chloride channel expression increased. Mild hepatopancreatic structural alterations were also observed after CLC-3 knockdown. These findings suggest that CLC-3 is a salinity-responsive intracellular chloride-transporter candidate associated with intestinal ion-transport-related transcriptional responses after ANO1 suppression in L. vannamei, although the underlying physiological mechanism requires further validation.

Animals

Somatic and germinal mosaicism of a canonical splicing variant causing limb-girdle muscular dystrophy type 1B.

Limb-girdle muscular dystrophy type 1B is one of several muscular dystrophies caused by pathogenic variants in the LMNA gene. In this study, we investigated the clinical, pathological, and genetic findings of an LGMD1B family. Genetic sequencing identified the proband and her younger brother both carried the canonical splicing c.513 + 1G > A variant in the LMNA gene. The variant was absent in the proband's mother, and a certain percentage of the LMNA variant was identified in the venous blood, urine, and semen sample of the proband's father by pyrophosphate sequencing. Further cDNA analysis demonstrated that the canonical splicing c.513 + 1G > A variant in intron 2 induced retention of the first 45 bp of intron 2, resulting in an in-frame insertion of 15 amino acids. Our study directly confirmed the presence of somatic and germinal mosaicism in the LGMD1B family and the pathogenicity of the canonical splicing variant in the LMNA gene.

Humans

IMAGE cDNA clones, UniGene clustering, and ACeDB: an integrated resource for expressed sequence information.

In this study we describe a new information resource that provides integrated access to information on IMAGE (integrated molecular analysis of genomes and their expression) cDNA library clones and derived expressed sequence tags (ESTs). We have developed an automated procedure that collates data from various public sources into a single ACeDB database. This database is a valuable tool for electronic cloning experiments and gene expression studies. It allows researchers to find information about cDNA libraries, plate addresses, insert sizes, and sequence data for IMAGE clones, the assignment of ESTs to UniGene clusters, and the chromosomal location of those genes in an efficient, graphically oriented manner.

Cloning, Molecular

Deciphering mixed infections by plant RNA virus and reconstructing complete genomes simultaneously present within-host.

Local co-circulation of multiple phylogenetic lineages is particularly likely for rapidly evolving pathogens in the current context of globalisation. When different phylogenetic lineages co-occur in the same fields, they may be simultaneously present in the same host plant (i.e. mixed infection), with potentially important consequences for disease outcome. This is the case in Burkina Faso for the rice yellow mottle virus (RYMV), which is endemic to Africa and a major constraint on rice production. We aimed to decipher the distinct RYMV isolates that simultaneously infect a single rice plant and to sequence their genomes. To this end, we tested different sequencing strategies, and we finally combined direct cDNA ONT (Oxford Nanopore Technology) sequencing with the bioinformatics tool RVhaplo. This method was validated by the successful reconstruction of two viral genomes that were less than a hundred nucleotides apart (out of a genome of 4450nt length, i.e. 2-3%), and present in artificial mixes at a ratio of up to a 99/1. We then used this method to subsequently analyze mixed infections from field samples, revealing up to three RYMV isolates within one single rice plant sample from Burkina Faso. In most cases, the complete genome sequences were obtained, which is particularly important for a better estimation of viral diversity and the detection of recombination events. The method described thus allows to identify various haplotypes of RYMV simultaneously infecting a single rice plant, obtaining their full-length sequences, as well as a rough estimate of relative frequencies within the sample. It is efficient, cost-effective, as well as portable, so that it could further be implemented where RYMV is endemic. Prospects include unravelling mixed infections with other RNA viruses that threaten crop production worldwide.

Genome, Viral

Human pegivirus 1 in healthy blood donors and in acute febrile patients from Mato Grosso, Midwestern Brazil.

The association between Human Pegivirus 1 (HPgV-1, Pegivirus hominis, family Hepaciviridae) with disease etiology remains unclear, yet this virus is widely distributed among people exposed to contaminated blood. This study aimed to identify the prevalence and characterize HPgV-1 genomes obtained from acute febrile patients (n = 92 from 2019) and from blood donors (n = 633 from 2022) sampled in Mato Grosso, Brazil. The serum samples were tested by RT-qPCR for a conserved 5' UTR region of HPgV-1. Nucleic acid of 28 positive samples was converted to ds-cDNA, purified and sequenced with an Illumina NextSeq platform. The prevalence of HPgV-1 among acute febrile patients was 4.35% (4/92), who were mostly female (3/4; 75.00%) mean 28.86 (12-67) years-old; whereas among blood donors it was 3.79% (24/633), which were predominantly male (16/24; 66.67%), aged 30-45 years-old (13/24; 54.17%). Eight HPgV-1 genomic sequences (7,708-9,248) were recovered; two belong to subgenotype 2a and clustered with strains from United States, Brazil, France and Japan, while the remaining six genomes from subgenotype 2b grouped with strains from Pará, Brazil and abroad. HPgV-1 detection in acute febrile patients and in healthy blood donors highlights the necessity to investigate the epidemiological aspects of this infection, as well as the possible impact to public health in Midwestern Brazil.

Cross-Sectional Studies

Nanopore sequencing to detect A-to-I editing sites.

Adenosine-to-inosine (A-to-I) RNA editing, mediated by the ADAR family of enzymes, is pervasive in metazoans and functions as an important mechanism to diversify the proteome and control gene expression. Over the years, there have been multiple efforts to comprehensively map the editing landscape in different organisms and in different disease states. As inosine (I) is recognized largely as guanosine (G) by cellular machineries including the reverse transcriptase, editing sites can be detected as A-to-G changes during sequencing of complementary DNA (cDNA). However, such an approach is indirect and can be confounded by genomic single nucleotide polymorphisms (SNPs) and DNA mutations. Moreover, past studies rely primarily on the Illumina platform, which generates short sequencing reads that can be challenging to map. Recently, nanopore direct RNA sequencing has emerged as a powerful technology to address the issues. Here, we describe the use of the technology together with deep learning models that we have developed, named Dinopore (Detection of inosine with nanopore sequencing), to interrogate the A-to-I editome of any organism.

Inosine

Obstacles in quantifying A-to-I RNA editing by Sanger sequencing.

Adenosine-to-Inosine (A-to-I) RNA editing is the most prevalent type of RNA editing, in which adenosine within a completely or largely double-stranded RNA (dsRNA) is converted to inosine by deamination. RNA editing was shown to be involved in many neurological diseases and cancer; therefore, detection of A-to-I RNA editing and quantitation of editing levels are necessary for both basic and clinical biomedical research. While high-throughput sequencing (HTS) is widely used for global detection of editing events, Sanger sequencing is the method of choice for precise characterization of editing site clusters (hyper-editing) and for comparing levels of editing at a particular site under different environmental conditions, developmental stages, genetic backgrounds, or disease states. To detect A-to-I editing events and quantify them using Sanger sequencing, RNA samples are reverse transcribed, cDNA is amplified using gene-specific primers, and then sequenced. The chromatogram outputs are then compared to the genomic DNA sequence. As editing occurs in the context of dsRNA, the reverse transcription step is performed at a temperature as high as 65 °C, using thermostable reverse transcriptase to open double-stranded structures. However, this measure alone is insufficient for transcripts possessing long stems comprised of hundreds of nucleotide pairs. Consequently, the editing levels detected by Sanger sequencing are significantly lower than those obtained by HTS, and the amplification yield is low. We suggest that the reverse transcription is biased towards unedited transcripts, and the severity of the bias is dependent on the transcript's secondary structure. Here, we show how this bias can be significantly reduced to allow reliable detection of editing levels and sufficient product yield.

RNA Editing

EnsMart: a generic system for fast and flexible access to biological data.

The EnsMart system (www.ensembl.org/EnsMart) provides a generic data warehousing solution for fast and flexible querying of large biological data sets and integration with third-party data and tools. The system consists of a query-optimized database and interactive, user-friendly interfaces. EnsMart has been applied to Ensembl, where it extends its genomic browser capabilities, facilitating rapid retrieval of customized data sets. A wide variety of complex queries, on various types of annotations, for numerous species are supported. These can be applied to many research problems, ranging from SNP selection for candidate gene screening, through cross-species evolutionary comparisons, to microarray annotation. Users can group and refine biological data according to many criteria, including cross-species analyses, disease links, sequence variations, and expression patterns. Both tabulated list data and biological sequence output can be generated dynamically, in HTML, text, Microsoft Excel, and compressed formats. A wide range of sequence types, such as cDNA, peptides, coding regions, UTRs, and exons, with additional upstream and downstream regions, can be retrieved. The EnsMart database can be accessed via a public Web site, or through a Java application suite. Both implementations and the database are freely available for local installation, and can be extended or adapted to 'non-Ensembl' data sets.

Animals