Time to defend what we have won.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to D States.
Explore the source record for details and available documents.
Untranslated regions (UTR) play important roles in the posttranscriptional regulation of mRNA processing. There is a wealth of UTR-related information to be mined from the rapidly accumulating EST collections. A computational tool, UTR-extender, has been developed to infer UTR sequences from genomically aligned ESTs. It can completely and accurately reconstruct 72% of the 3' UTRs and 15% of the 5' UTRs when tested using 908 functionally cloned transcripts. In addition, it predicts extensions for 11% of the 5' UTRs and 28% of the 3' UTRs. These extension regions are validated by examining splicing frequencies and conservation levels. We also developed a method called polyadenylation site scan (PASS) to precisely map polyadenylation sites in human genomic sequences. A PASS analysis of 908 genic regions estimates that 40-50% of human genes undergo alternative polyadenylation. Using EST redundancy to assess expression levels, we also find that genes with short 3' UTRs tend to be highly expressed.
A YAC/STS map of the X chromosome has reached an inter-STS resolution of 75 kb. The map density is sufficient to provide YACs or other large-insert clones that are cross-validated as sequencing substrates across the chromosome. Marker density also permits estimates of regional gene content and a detailed comparison of genetic and physical map distances. Five regions are detected with relatively high G + C, correlated with gene richness; and a 17-Mb region with very low recombination is revealed between the Xq13.3 [XIST] and Xq21.3 XY homology loci.
Sets of new gene sequences from human, nematode, and yeast were compared with each other and with a set of Escherichia coli genes in order to detect ancient evolutionarily conserved regions (ACRs) in the encoded proteins. Nearly all of the ACRs so identified were found to be homologous to sequences in the protein databases. This suggests that currently known proteins may already include representatives of most ACRs and that new sequences not similar to any database sequence are unlikely to contain ACRs. Preliminary analyses indicate that moderately expressed genes may be more likely to contain ACRs than rarely expressed genes. It is estimated that there are fewer than 900 ACRs in all.