<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Peter+Ralph</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Peter+Ralph"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Peter_Ralph"/>
	<updated>2026-09-24T22:18:15Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SNP_Call_Set_Properties&amp;diff=5182</id>
		<title>SNP Call Set Properties</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SNP_Call_Set_Properties&amp;diff=5182"/>
		<updated>2012-09-10T23:08:11Z</updated>

		<summary type="html">&lt;p&gt;Peter Ralph: fixing an oops: see http://mbe.oxfordjournals.org/content/17/1/32.full for reference&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;There are typically a number of properties that we check in SNP call sets. This page gives some useful pointers on what these quantities are and what to look for.&lt;br /&gt;
&lt;br /&gt;
== Proportion of dbSNPs ==&lt;br /&gt;
&lt;br /&gt;
Most of the genetic variants in any one individual have been previously observed in other individuals. Thus, it is usually a good diagnostic to investigate what fraction of variants in an individual genome have been previously described in [http://www.ncbi.nlm.nih.gov/projects/SNP/ dbSNP]. &lt;br /&gt;
&lt;br /&gt;
The expected proportion of previously discovered SNPs (those already catalogued in dbSNP) and novel SNPs (those that haven&#039;t been previously discovered) will change overtime. The dbSNP database is being constantly updated so that currently (mid-2010) we&#039;d expect &amp;gt;90% of the variants in an individual genome to have been previously discovered. &lt;br /&gt;
&lt;br /&gt;
If many individuals are sequenced, the vast majority of common variants (those shared among many individuals) are expected to be in dbSNP, whereas a smaller fraction of newly discovered variants should be in dbSNP.&lt;br /&gt;
&lt;br /&gt;
== Transition to Transversion Ratio ==&lt;br /&gt;
&lt;br /&gt;
Human mutations don&#039;t occur randomly. In fact, transitions (changes from A &amp;lt;-&amp;gt; G and C &amp;lt;-&amp;gt; T) are expected to occur twice as frequently as transversions (changes from A &amp;lt;-&amp;gt; C, A &amp;lt;-&amp;gt; T, G &amp;lt;-&amp;gt; C or G &amp;lt;-&amp;gt; T). Thus, another useful diagnostic is the ratio of transitions to transversions in a particular set of SNP calls. This ratio is often evaluated separately for previously discovered and novel SNPs.&lt;br /&gt;
&lt;br /&gt;
Across the entire genome the ratio of transitions to transversions is typically around 2. In protein coding regions, this ratio is typically higher, often a little above 3. The higher ratio occurs because, especially when they occur in the third base of a codon, transversions are much more likely to change the encoded amino acid. A refinement to this analysis, in protein coding regions, is to examine the transition to transversion ratio separately for non-degenerate, two-fold degenerate, three-fold degenerate and four-fold degenerate sites.&lt;br /&gt;
&lt;br /&gt;
== Why Are Reciprocal Changes Not Equally Frequent? ==&lt;br /&gt;
&lt;br /&gt;
One of the most surprising features of many variant lists in humans is that C-&amp;gt;T changes (C reference, T variant) are more frequent than T-&amp;gt;C changes. Likewise, G-&amp;gt;A changes are more frequent than A-&amp;gt;G changes.&lt;br /&gt;
&lt;br /&gt;
At first, this might seem a bit puzzling. For example, perhaps we might expect that the two counts should be extremely similar. However, the reference makes perfect biological sense -- and the explanation below is due to [[Tom Blackwell]]. &lt;br /&gt;
&lt;br /&gt;
The major mechanism for new mutations (in warm-blooded animals) is deamination of 5&#039;-methyl C to uracil (equivalently T) producing (C -&amp;gt; T) or, on the complementary strand, (G -&amp;gt; A).  This was first studied for CpG dinucleotide sites, but it also occurs at lower rates throughout the genome at any C whether followed by G or not.&lt;br /&gt;
&lt;br /&gt;
More often than not, we expect that the reference genome will include the most common allele, which is also likely to be the ancestral allele. Thus, if C-&amp;gt;T mutations are more common than T-&amp;gt;C mutations, we expect to see an imbalance of C-&amp;gt;T versus T-&amp;gt;C changes. Further, when comparing rare and common variants, we expect the imbalance to be stronger for lower frequency variants.&lt;/div&gt;</summary>
		<author><name>Peter Ralph</name></author>
	</entry>
</feed>