<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Youna</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Youna"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Youna"/>
	<updated>2026-09-24T12:12:04Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Main_Page&amp;diff=8936</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Main_Page&amp;diff=8936"/>
		<updated>2013-11-04T07:10:10Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Sequence Analysis Tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--        BANNER ACROSS TOP OF PAGE        --&amp;gt; &lt;br /&gt;
&lt;br /&gt;
{| style=&amp;quot;width:100%; background:#fcfcfc; margin-top:1.2em; border:1px solid #ccc;&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
| style=&amp;quot;width:100%; text-align:center; white-space:nowrap; color:#000;&amp;quot; | &amp;lt;div style=&amp;quot;font-size:162%; border:none; margin:0; padding:.1em; color:#000;&amp;quot;&amp;gt;Abecasis Group Wiki&amp;lt;/div&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:Abecasis_group_photo_cropped.jpg|900px|center|Group Photo 2013]]&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Below is old image from 2009 Retreat --&amp;gt;&lt;br /&gt;
&amp;lt;!-- &amp;lt;br&amp;gt; [[Image:2009.08 Group Retreat Photo.jpg|center|400px|Group Photo]]--&amp;gt;&lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
== Welcome!  ==&lt;br /&gt;
&lt;br /&gt;
Welcome to our brand new wiki! &lt;br /&gt;
&lt;br /&gt;
If you would like to contribute, [[Special:UserLogin|log-in]] or [[Special:RequestAccount|request an account]]. We recommend using your e-mail address or Michigan uniqname as your user id. &lt;br /&gt;
&lt;br /&gt;
For basic instructions, see [http://en.wikipedia.org/wiki/Wikipedia:Tutorial the Wikipedia Tutorial]. &lt;br /&gt;
&lt;br /&gt;
== Sequence Analysis Tools  ==&lt;br /&gt;
&lt;br /&gt;
We are developing [[Software|software tools]] for the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
These tools include: &lt;br /&gt;
&lt;br /&gt;
#Variant Calling with [[GlfSingle]] and [[GlfMultiples]] &lt;br /&gt;
#Variant Calling and De Novo Mutation Detection in Families with [[Polymutt]] &lt;br /&gt;
#Variant Annotations using [[VcfCodingSnps]] &lt;br /&gt;
#Rare Variant Analysis using [[RvTests]] &lt;br /&gt;
#Rare Variant Association Analysis in family samples [[FamRvTest]]&lt;br /&gt;
#Quality control using [[C++ Executable: fastQValidator|FastQValidator]], [[VerifyBamID]], and [[BamValidator]] &lt;br /&gt;
#C++ APIs for sequence analsysis using [[C++ Library: libStatGen]] &lt;br /&gt;
#Meta-analysis of single variant or gene-level associations [[RAREMETAL-SOFTWARE]]&lt;br /&gt;
#Sequencing study design helper [[Rarefy]]&lt;br /&gt;
#Local ancestry inference (ancestry painting) using off-targeted sequence data [[SEQMIX]]&lt;br /&gt;
&lt;br /&gt;
These tools and additional tools can be found on the [[Software]] page. &lt;br /&gt;
&lt;br /&gt;
We are developing Genome/Sequencing Processing Pipelines for anyone to use: [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
== High Level Tutorials  ==&lt;br /&gt;
&lt;br /&gt;
Some high-level tutorials on the analysis of next generation sequence data: &lt;br /&gt;
&lt;br /&gt;
#[[Evaluating a Read Mapper on Simulated Data]] &lt;br /&gt;
#[[SNP Call Set Properties]] &lt;br /&gt;
#[[Generic Exome Analysis Plan]]&lt;br /&gt;
&lt;br /&gt;
== Projects  ==&lt;br /&gt;
&lt;br /&gt;
[[SardiNIA]] - The SardiNIA longitudinal study of aging. &lt;br /&gt;
&lt;br /&gt;
[[1000 Genomes Project Pilot 1 SNP Calling]]&lt;br /&gt;
&lt;br /&gt;
[[EMADS|Exome Meta-analysis of Drinking and Smoking (EMADS)]]&lt;br /&gt;
&lt;br /&gt;
== Learn Genetics  ==&lt;br /&gt;
&lt;br /&gt;
Faculty in the group teach in a variety of formal and informal settings. [[Class Notes|Class notes]] and relevant discussion are archived here. &lt;br /&gt;
&lt;br /&gt;
== General Resources  ==&lt;br /&gt;
&lt;br /&gt;
*[[Computer How-Tos]]&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8935</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8935"/>
		<updated>2013-11-04T07:07:35Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] for SEQMIX (version 0.1) which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for Europeans&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Maintainer == &lt;br /&gt;
&lt;br /&gt;
Please contact Youna Hu (&#039;&#039;youna@umich.edu&#039;&#039;) if you have any questions or suggestions for SEQMIX.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8934</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8934"/>
		<updated>2013-11-04T07:07:10Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for Europeans&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Maintainer == &lt;br /&gt;
&lt;br /&gt;
Please contact Youna Hu (&#039;&#039;youna@umich.edu&#039;&#039;) if you have any questions or suggestions for SEQMIX.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8933</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8933"/>
		<updated>2013-11-04T07:06:56Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Questions and Suggestions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Maintainer == &lt;br /&gt;
&lt;br /&gt;
Please contact Youna Hu (&#039;&#039;youna@umich.edu&#039;&#039;) if you have any questions or suggestions for SEQMIX.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8932</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8932"/>
		<updated>2013-11-04T07:06:25Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Questions and Suggestions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Questions and Suggestions == &lt;br /&gt;
&lt;br /&gt;
Please contact Youna Hu (&#039;&#039;youna@umich.edu&#039;&#039;) if you have any questions or suggestions for SEQMIX.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8931</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8931"/>
		<updated>2013-11-04T07:06:12Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Questions and Suggestions == &lt;br /&gt;
&lt;br /&gt;
Please contact Youna Hu &#039;&#039;youna@umich.edu&#039;&#039; if you have any questions or suggestions for SEQMIX. &lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8930</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8930"/>
		<updated>2013-11-04T07:05:06Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command file as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8929</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8929"/>
		<updated>2013-11-04T07:04:41Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Overview */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference. The paper is currently accepted by AJHG and will appear at the November issue (link coming soon).&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8928</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8928"/>
		<updated>2013-11-04T07:03:40Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to have them. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8927</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8927"/>
		<updated>2013-11-04T07:03:24Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the &#039;&#039;&#039;example.sh&#039;&#039;&#039; file), please refer to the &#039;&#039;&#039;Readme.txt&#039;&#039;&#039; file for detailed explanations of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (&#039;&#039;youna@umich.edu&#039;&#039;) if you would like to use it. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8926</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8926"/>
		<updated>2013-11-04T07:00:40Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo Abecasis&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the example.sh file), please refer to the Readme.txt file for detailed explanation of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (youna@umich.edu) if you would like to use it. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8925</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8925"/>
		<updated>2013-11-04T07:00:16Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
Before you run SEQMIX, or even while you are running SEQMIX examples (summarized by the example.sh file), please refer to the Readme.txt file for detailed explanation of the two steps for running SEQMIX. &lt;br /&gt;
&lt;br /&gt;
Note that SEQMIX requires these specified files&lt;br /&gt;
&lt;br /&gt;
* allele frequency for Africans&lt;br /&gt;
* allele frequency for European&lt;br /&gt;
* genetic distance file &lt;br /&gt;
* input vcf &lt;br /&gt;
&lt;br /&gt;
The input vcf file is generated from sequencing experiment and the downstream data processing steps. The allele frequency and genetic distance files should be prepared by users and sometimes are tedious to do. The good news is that I have used these files for the whole genome level. Please contact me (youna@umich.edu) if you would like to use it. I will point you to the path if you are internal user and will figure out a way to share them (These files are fairly big) if you are a external users.&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8924</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8924"/>
		<updated>2013-11-04T06:53:51Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
Here is a tar ball [[Media:SEQMIX_0.1.tar]] which compresses the following three folders&lt;br /&gt;
&lt;br /&gt;
* libsrc: a folder contains source code from code written by Goncalo&lt;br /&gt;
* src: source code for SEQMIX&lt;br /&gt;
* Release: example files and command as well as the Readme.txt file that explains how to run SEQMIX&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:SEQMIX_0.1.tar&amp;diff=8923</id>
		<title>File:SEQMIX 0.1.tar</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:SEQMIX_0.1.tar&amp;diff=8923"/>
		<updated>2013-11-04T06:51:01Z</updated>

		<summary type="html">&lt;p&gt;Youna: This file has the main code, the goncalo library code and released example for using SEQMIX. I&amp;#039;ve provided precompiled executable file (localAncestry_release) as well as source code. If the precompiled version doesn&amp;#039;t run due to a different computer ar...&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This file has the main code, the goncalo library code and released example for using SEQMIX. I&#039;ve provided precompiled executable file (localAncestry_release) as well as source code. If the precompiled version doesn&#039;t run due to a different computer architecture, please go to the src folder, to run make to compile your own version of SEQMIX.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8922</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8922"/>
		<updated>2013-11-04T06:47:45Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
[[Media:SEQMIX_0.1.tar]]&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8921</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=8921"/>
		<updated>2013-11-04T06:46:04Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
[[SEQMIX_0.1.tar]]&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6784</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6784"/>
		<updated>2013-03-13T18:14:43Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Method */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to LD prune your data so that pairs of sites in high LD (r^2 &amp;gt; 0.1) are identified and only the one with a higher sequence depth are included into the model. As the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6783</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6783"/>
		<updated>2013-03-13T18:13:07Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6782</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6782"/>
		<updated>2013-03-13T18:12:56Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6781</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6781"/>
		<updated>2013-03-13T18:12:09Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are not user friendly and are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6780</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6780"/>
		<updated>2013-03-13T18:11:35Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are not user friendly and are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [http://genome.sph.umich.edu/wiki/LASER/ LASER].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6779</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6779"/>
		<updated>2013-03-13T18:11:20Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/\~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are not user friendly and are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [http://genome.sph.umich.edu/wiki/LASER/ LASER].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6778</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6778"/>
		<updated>2013-03-13T18:10:19Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are not user friendly and are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [http://genome.sph.umich.edu/wiki/LASER/ LASER].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6777</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6777"/>
		<updated>2013-03-13T18:09:54Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP] (Warning: These programs are not user friendly and are very difficult to run.).&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6776</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6776"/>
		<updated>2013-03-13T18:09:19Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [http://lamp.icsi.berkeley.edu/lamp/ LAMP], [http://genepath.med.harvard.edu/~reich/Software.htm/ ANCESTRYMAP]. (Warning: These programs are not user friendly and are very difficult to get it running)&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6775</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6775"/>
		<updated>2013-03-13T18:07:37Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [http://www.stats.ox.ac.uk/~myers/software.html/ HAPMIX], [[LAMP]], [[ANCESTRYMAP]].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=HAPMIX&amp;diff=6774</id>
		<title>HAPMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=HAPMIX&amp;diff=6774"/>
		<updated>2013-03-13T18:05:20Z</updated>

		<summary type="html">&lt;p&gt;Youna: Created page with &amp;#039;http://www.stats.ox.ac.uk/~myers/software.html&amp;#039;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;http://www.stats.ox.ac.uk/~myers/software.html&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6773</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6773"/>
		<updated>2013-03-13T18:04:36Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [[HAPMIX]], [[LAMP]], [[ANCESTRYMAP]].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[LASER]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6772</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6772"/>
		<updated>2013-03-13T18:04:11Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [[HAPMIX]], [[[Lamp]]], [[[ANCESTRYMAP]]].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[[LASER]]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6771</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6771"/>
		<updated>2013-03-13T18:03:56Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Related Programs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
Local ancestry inference with high density genotype array data can be done with existing software [[[HAPMIX]]], [[[Lamp]]], [[[ANCESTRYMAP]]].&lt;br /&gt;
&lt;br /&gt;
Whole genome ancestry inference with ultra low coverage sequence data can be analyzed with [[[LASER]]].&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6770</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6770"/>
		<updated>2013-03-13T18:00:41Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Method */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is important to pre-process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth into the model. Since the sequence depth distribution is sample dependent, it is necessary to prune the sequence data for each individual.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6769</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6769"/>
		<updated>2013-03-13T17:59:25Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Method */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method ==&lt;br /&gt;
&lt;br /&gt;
Before running SEQMIX, it is useful to process your data with a LD pruning step, which identify sites that are in high LD (r^2 &amp;gt; 0.1) and keep the sites with a higher sequence depth.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6768</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6768"/>
		<updated>2013-03-13T17:53:28Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Download */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method == &lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
(Coming soon)&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6767</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6767"/>
		<updated>2013-03-13T17:53:04Z</updated>

		<summary type="html">&lt;p&gt;Youna: /* Introduction */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Overview ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;SEQMIX&#039;&#039;&#039; is a C++ program that takes advantage of off-targeted sequence reads from exome/targeted sequencing experiments for accurate local ancestry inference.&lt;br /&gt;
&lt;br /&gt;
== Method == &lt;br /&gt;
&lt;br /&gt;
== Download == &lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6764</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6764"/>
		<updated>2013-03-13T17:49:10Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Introduction == &lt;br /&gt;
&lt;br /&gt;
== Method == &lt;br /&gt;
&lt;br /&gt;
== Download == &lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6763</id>
		<title>SEQMIX</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SEQMIX&amp;diff=6763"/>
		<updated>2013-03-13T17:48:35Z</updated>

		<summary type="html">&lt;p&gt;Youna: Created page with &amp;#039;== Introduction ==   == Download ==   == Workflow ==   == Related Programs ==&amp;#039;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Introduction == &lt;br /&gt;
&lt;br /&gt;
== Download == &lt;br /&gt;
&lt;br /&gt;
== Workflow == &lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=LibStatGen:_VCF&amp;diff=6248</id>
		<title>LibStatGen: VCF</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=LibStatGen:_VCF&amp;diff=6248"/>
		<updated>2013-01-23T20:09:56Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:C++]]&lt;br /&gt;
[[Category:libStatGen]]&lt;br /&gt;
[[Category:libStatGen VCF]]&lt;br /&gt;
&lt;br /&gt;
= Variant Call Format (VCF) =&lt;br /&gt;
&lt;br /&gt;
The documentation on the Variant Call Format (VCF) can be found at: http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41&lt;br /&gt;
&lt;br /&gt;
&amp;lt;span style=&amp;quot;color:#D2691E&amp;quot;&amp;gt;&#039;&#039;&#039;Our APIs are currently in the initial test phase, and are not yet available for public use.&#039;&#039;&#039;&amp;lt;/span&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== API for Reading VCF Files ==&lt;br /&gt;
&lt;br /&gt;
You will use both an &amp;lt;code&amp;gt;VcfFileReader&amp;lt;/code&amp;gt; and an &amp;lt;code&amp;gt;VCfRecord&amp;lt;/code&amp;gt; for reading VCF files.&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;code&amp;gt;VcfFileReader&amp;lt;/code&amp;gt; ===&lt;br /&gt;
An instance of the VcfFileReader class is used to read VCF files.&lt;br /&gt;
&lt;br /&gt;
==== include file ====&lt;br /&gt;
&amp;lt;code&amp;gt;VcfFileReader&amp;lt;/code&amp;gt; is declared in &amp;lt;code&amp;gt;VcfFileReader.h&amp;lt;/code&amp;gt;, so be sure to include that file.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
#include &amp;quot;VcfFileReader.h&amp;quot;&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Opening the VCF File ====&lt;br /&gt;
&amp;lt;code&amp;gt;open&amp;lt;/code&amp;gt; opens the specified file and by default throws an exception if it was not successfully opened.&lt;br /&gt;
&lt;br /&gt;
There are two methods for opening a file for reading.  One is standard while the other allows the specification of a subset of samples to keep.&lt;br /&gt;
&lt;br /&gt;
Both methods take the filename to be read as well as a reference to the VcfHeader to be populated.&lt;br /&gt;
&lt;br /&gt;
The standard method is:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Reading a Subset of Samples ====&lt;br /&gt;
To select only a subset of samples to keep, when opening the file also specify the name of the file containing the names of the samples to keep and the delimiter separating the sample names (default is a new line, &#039;\n&#039;).&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    // Subset 1 is delimited by new lines, &#039;\n&#039;.&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header, &amp;quot;subsetFile1.txt&amp;quot;, NULL, NULL);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    // Subset 2 is delimited by &#039;;&#039;&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header, &amp;quot;subsetFile2.txt&amp;quot;, NULL, NULL, &#039;;&#039;);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
To select a subset of samples to remove/ignore, when opening the file also specify the name of the file containing the names of the samples to remove/ignore and the delimiter separating the sample names (default is a new line, &#039;\n&#039;).&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    // Subset 1 is delimited by new lines, &#039;\n&#039;.&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header, NULL, NULL, &amp;quot;subsetFile1.txt&amp;quot;);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    // Subset 2 is delimited by &#039;;&#039;&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header, NULL, NULL, &amp;quot;subsetFile2.txt&amp;quot;, &#039;;&#039;);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If you just have 1 sample to be excluded, you can specify it directly in the open line.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    // Subset 1 is delimited by new lines, &#039;\n&#039;.&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header, NULL, &amp;quot;BAD_SAMPLE&amp;quot;, NULL);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Reading a VCF Record ====&lt;br /&gt;
Once the file is open, get the next record using &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
It returns true if a record was successfully found and false on EOF or an error.  Typically on a reading error an exception is thrown.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; takes a reference to a &amp;lt;code&amp;gt;VcfRecord&amp;lt;/code&amp;gt; object as a parameter.  When true is returned, the &amp;lt;code&amp;gt;VcfRecord&amp;lt;/code&amp;gt; is updated with the next record.&lt;br /&gt;
&lt;br /&gt;
If you only process one record at a time (are done with a record before reading the next one), you can loop until &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; return false, reusing the same record for each call.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
        VcfRecord record;&lt;br /&gt;
        while(reader.readRecord(record))&lt;br /&gt;
        {&lt;br /&gt;
            // Your record specific processing here.&lt;br /&gt;
        }&lt;br /&gt;
        // Done reading the file.&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If any subsetting was specified when the file was opened, the record that is returned will have only the specified subset of samples in it.&lt;br /&gt;
&lt;br /&gt;
When reading a record, you can also specify discard rules.  Internal processing of &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; will apply the discard rules on the record that was read and will continue reading records and not return from the method until a record is found that should not be discarded (or until the end of the file).  True is returned if a record that should be kept was found, false if the end of the file was hit without finding a record to keep.&lt;br /&gt;
&lt;br /&gt;
==== Specifying Discard Rules ====&lt;br /&gt;
===== Basic Rules =====&lt;br /&gt;
When specifying the discard rules, you should use the constants found at the top of VcfFileReader.h&lt;br /&gt;
&lt;br /&gt;
For Example (see the file for the current set of Discard Rules and the associated values):&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    static const int DISCARD_NON_PHASED = 0x1;&lt;br /&gt;
    static const int DISCARD_MISSING_GT = 0x2;&lt;br /&gt;
    static const int DISCARD_FILTERED = 0x4;&lt;br /&gt;
    static const int DISCARD_MULTIPLE_ALTS = 0x8;&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
You can set the Discard using &amp;lt;code&amp;gt;setDiscardRules(uint32_t)&amp;lt;/code&amp;gt;.  You can add additional rules to what has already been set by using &amp;lt;code&amp;gt;addDiscardRules(uint32_t)&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
        reader.setDiscardRules(DISCARD_FILTERED);&lt;br /&gt;
        reader.addDiscardRules(DISCARD_NON_PHASED);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
is the same as&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
        reader.setDiscardRules(DISCARD_FILTERED | DISCARD_NON_PHASED);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
and discards reads that do not have &amp;lt;code&amp;gt;PASS&amp;lt;/code&amp;gt; in the &amp;lt;code&amp;gt;FILTER&amp;lt;/code&amp;gt; field and reads that have a genotype that is not phased or have no &amp;lt;code&amp;gt;GT&amp;lt;/code&amp;gt; in the &amp;lt;code&amp;gt;FORMAT&amp;lt;/code&amp;gt; fields.&lt;br /&gt;
&lt;br /&gt;
===== Minimum Alternate Allele Count =====&lt;br /&gt;
To Discard any records without a minimum number of alternate alleles, use:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
VcfFileReader::addDiscardMinAltAlleleCount(int32_t minAltAlleleCount, VcfSubsetSamples* subset)&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;minAltAlleleCount&amp;lt;/code&amp;gt; parameter is the minimum number of alternate alleles found in the specified subset (if specified) in order for the record to be kept.&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;VcfSubsetSamples* subset&amp;lt;/code&amp;gt; parameter is a pointer to the subset of samples that you want to include when counting the number of alternate alleles.  If all samples that are read/kept are to be included, NULL should be passed in.  &lt;br /&gt;
&lt;br /&gt;
See [[#Handling a Subset of Samples|Handling a Subset of Samples]] for how to use &amp;lt;code&amp;gt;VcfSubsetSamples&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Use the following method to remove the DiscardMinAltAlleleCount rule:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
VcfFileReader::rmDiscardMinAltAlleleCount()&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Example:  Minimum Alternate Allele Count = 4&lt;br /&gt;
 Sample1  Sample2  Sample3  Keep/Discard&lt;br /&gt;
  0|0       1|1      2|2    Keep&lt;br /&gt;
  0|0       0|1      2|2    Discard, only 3 Alternates (1 Allele1 &amp;amp; 2 Allele 2)&lt;br /&gt;
  0|0       1|1      1|2    Keep&lt;br /&gt;
  0|2       1|1      2|2    Keep&lt;br /&gt;
  2|1       0|1      2|0    Keep&lt;br /&gt;
&lt;br /&gt;
Example:  Minimum Alternate Allele Count = 3 &amp;amp; Exclude Sample2 (without the exclusion, all would be kept)&lt;br /&gt;
 Sample1  Sample2  Sample3  Keep/Discard&lt;br /&gt;
  0|0       1|1      2|2    Discard, only 2 Alternates (0 Allele1 &amp;amp; 2 Allele 2)&lt;br /&gt;
  0|0       0|1      2|2    Discard, only 2 Alternates (0 Allele1 &amp;amp; 2 Allele 2)&lt;br /&gt;
  0|0       1|1      1|2    Discard, only 2 Alternates (1 Allele1 &amp;amp; 1 Allele 2)&lt;br /&gt;
  0|2       1|1      2|2    Keep&lt;br /&gt;
  2|1       0|1      2|0    Keep&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===== Minimum Minor Allele Count =====&lt;br /&gt;
To Discard any records without a minimum number of minor alleles, use:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
VcfFileReader::addDiscardMinMinorAlleleCount(int32_t minMinorAlleleCount, VcfSubsetSamples* subset)&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;minMinorAlleleCount&amp;lt;/code&amp;gt; parameter is the minimum number of minor alleles found in the specified subset (if specified) in order for the record to be kept.&lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;VcfSubsetSamples* subset&amp;lt;/code&amp;gt; parameter is a pointer to the subset of samples that you want to include when counting the number of alleles.  If all samples that are read/kept are to be included, NULL should be passed in.  &lt;br /&gt;
&lt;br /&gt;
See [[#Handling a Subset of Samples|Handling a Subset of Samples]] for how to use &amp;lt;code&amp;gt;VcfSubsetSamples&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Use the following method to remove the DiscardMinMinorAlleleCount rule:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
VcfFileReader::rmDiscardMinMinorAlleleCount()&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Example:  Minimum Minor Allele Count = 2&lt;br /&gt;
 Sample1  Sample2  Sample3  Keep/Discard&lt;br /&gt;
  0|0       1|1      2|2    Keep&lt;br /&gt;
  0|0       0|1      2|2    Discard, only 1 Allele1&lt;br /&gt;
  0|0       1|1      1|2    Discard, only 1 Allele2&lt;br /&gt;
  0|2       1|1      2|2    Discard, only 1 Allele0&lt;br /&gt;
  2|1       0|1      2|0    Keep&lt;br /&gt;
&lt;br /&gt;
Example:  Minimum Minor Allele Count = 1 &amp;amp; Exclude Sample2 (without the exclusion, all would be kept)&lt;br /&gt;
 Sample1  Sample2  Sample3  Keep/Discard&lt;br /&gt;
  0|0       1|1      2|2    Discard, 0 Allele1&lt;br /&gt;
  0|0       0|1      2|2    Discard, 0 Allele1&lt;br /&gt;
  0|0       1|1      1|2    Keep&lt;br /&gt;
  0|2       1|1      2|2    Discard, 0 Allele1&lt;br /&gt;
  2|1       0|1      2|0    Keep&lt;br /&gt;
&lt;br /&gt;
==== Read only Certain Sections of the File / Using a VCF Index (TABIX) File ====&lt;br /&gt;
The VCF library has the capability of using a tabix file to allow random access in the VCF file.&lt;br /&gt;
&lt;br /&gt;
To use the INDEX File to read sections of the bam file:&lt;br /&gt;
# Open the VCF file using: &amp;lt;code&amp;gt;VcfFileReader::open&amp;lt;/code&amp;gt;&lt;br /&gt;
# Open the VCF Index file using: &amp;lt;code&amp;gt;VcfFileReader::readVcfIndex&amp;lt;/code&amp;gt;&lt;br /&gt;
# Set the Section to be read: &amp;lt;code&amp;gt;VcfFileReader::setReadSection&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;VcfFileReader::set1BasedReadSection&amp;lt;/code&amp;gt;&lt;br /&gt;
# Read records from the section: &amp;lt;code&amp;gt;VcfFileReader::readRecord&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;VcfFileReader::readVcfIndex&amp;lt;/code&amp;gt; returns true if the index file was successfully read and false if not.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    bool readVcfIndex(const char * filename);&lt;br /&gt;
    bool readVcfIndex();&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
If you specify a filename, it looks for the index file with that path/name.  If you do not specify the index file name, it will attempt to find it.  First it searches for a file with your VCF FileName + a &amp;quot;.tbi&amp;quot; extension.  If that isn&#039;t found, it removes the &amp;quot;.vcf&amp;quot; if found in your VCF FileName and looks for a file with that name and a &amp;quot;.tbi&amp;quot; extension.&lt;br /&gt;
&lt;br /&gt;
After opening both the VCF File and the Index file, set the section that should be read next using &amp;lt;code&amp;gt;setReadSection&amp;lt;/code&amp;gt; or &amp;lt;code&amp;gt;set1BasedReadSection&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    bool setReadSection(const char* chromName);&lt;br /&gt;
    bool set1BasedReadSection(const char* chromName, &lt;br /&gt;
                              int32_t start, int32_t end,&lt;br /&gt;
                              bool overlap = false);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
Both methods currently always return true, since in the current implementation this can&#039;t fail.  It currently allows the VcfIndex file to be read after this call.&lt;br /&gt;
&lt;br /&gt;
The read section is reset when a new file is opened, so do not set this prior to opening the file.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;setReadSection&amp;lt;/code&amp;gt; will set the code to read the entire specified chromosome.  On the first &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; call it will jump to the beginning of the chromosome.  It will continue to read the chromosome for each &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; call made until it has read the entire chromosome.  Once the whole chromosome has been read, &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; will return false until a new read section is specified.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;set1BasedReadSection&amp;lt;/code&amp;gt; will set the code to read the specified chromosome starting at the specified 1-based position up to, but not including, the specified 1-based end position.  On the first &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; call it will jump the file to the beginning of this section.  It will continue to read the section for each &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; call made until it has read the entire section.  Once the entire section has been read, &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; will return false until a new read section is specified. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;set1BasedReadSection&amp;lt;/code&amp;gt; takes an optional parameter, &amp;lt;code&amp;gt;overlap&amp;lt;/code&amp;gt;.  It is defaulted to false which means that only records that start within the specified region will be returned by &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt;.  If &amp;lt;code&amp;gt;overlap&amp;lt;/code&amp;gt; is set to true, &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; will also return records that start prior to the specified region, but whose deletions occur in the specified region. &lt;br /&gt;
&lt;br /&gt;
Implementation NOTE: internally it may read extra records prior to the section, but &amp;lt;code&amp;gt;readRecord&amp;lt;/code&amp;gt; will keep reading and will not return until it finds a record in the section or until the section has been passed (no records in the section).&lt;br /&gt;
&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Open the vcf file for reading.&lt;br /&gt;
    VcfFileReader reader;&lt;br /&gt;
    VcfHeader header;&lt;br /&gt;
    VcfRecord record&lt;br /&gt;
    reader.open(&amp;quot;vcfFileName.vcf&amp;quot;, header);&lt;br /&gt;
&lt;br /&gt;
    // Index File name is &amp;quot;vcfFileName.vcf.tbi&amp;quot; or &amp;quot;vcfFileName.tbi&amp;quot;&lt;br /&gt;
    reader.readVcfIndex();&lt;br /&gt;
&lt;br /&gt;
    reader.set1BasedReadSection(&amp;quot;1&amp;quot;, 16384, 32767);&lt;br /&gt;
    while(reader.readRecord(record))&lt;br /&gt;
    {&lt;br /&gt;
        // Process a record from this section.&lt;br /&gt;
    }&lt;br /&gt;
&lt;br /&gt;
    // Set another section, chromosome &amp;quot;X&amp;quot;&lt;br /&gt;
    reader.setReadSection(&amp;quot;X&amp;quot;);&lt;br /&gt;
    while(reader.readRecord(record))&lt;br /&gt;
    {&lt;br /&gt;
        // Process record on chromosome X&lt;br /&gt;
    }&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Alternatively, the index file name can be specified.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    // Index File name is specified&lt;br /&gt;
    reader.readVcfIndex(&amp;quot;myIndex.tbi&amp;quot;);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Check if EOF was reached ====&lt;br /&gt;
To see if the End Of the File has been reached, use &amp;lt;code&amp;gt;isEOF()&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
        reader.isEOF();&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
True is returned if the End of the File has been reached, false if not.&lt;br /&gt;
&lt;br /&gt;
==== See How Many Records Were Read ====&lt;br /&gt;
To see the total number of records in the file that were read, including any that were discarded, use &amp;lt;code&amp;gt;getNumRecords&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
To see only the records in the file that were kept (not discarded) use &amp;lt;code&amp;gt;getNumKeptRecords&amp;lt;/code&amp;gt;.&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
        std::cout &amp;lt;&amp;lt; reader.getNumRecords() &amp;lt;&amp;lt; &amp;quot; were found in the file\n.&amp;quot;&lt;br /&gt;
                  &amp;lt;&amp;lt; &amp;quot;but only &amp;quot; &amp;lt;&amp;lt; reader.getNumKeptRecords() &amp;lt;&amp;lt; &amp;quot; were processed\n.&amp;quot;;&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
==== Closing the VCF File ====&lt;br /&gt;
&amp;lt;code&amp;gt;close&amp;lt;/code&amp;gt; closes the opened VCF file.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
    reader.close();&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== &amp;lt;code&amp;gt;VcfRecord&amp;lt;/code&amp;gt; ===&lt;br /&gt;
VCF records are read from VCF files using &amp;lt;code&amp;gt;VcfFileReader::readRecord&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
Once you have the record, use methods from the &amp;lt;code&amp;gt;VcfRecord&amp;lt;/code&amp;gt; to extract the desired information.&lt;br /&gt;
&lt;br /&gt;
==== Extracting Data from the Record ====&lt;br /&gt;
Do not delete any of the returned pointers.  They are references to strings within the record and will not be valid when the record changes or is deleted.&lt;br /&gt;
&lt;br /&gt;
* Position Information&lt;br /&gt;
** &amp;lt;code&amp;gt;getChromStr()&amp;lt;/code&amp;gt; returns const char* to this record&#039;s chromosome string&lt;br /&gt;
** &amp;lt;code&amp;gt;get1BasedPosition()&amp;lt;/code&amp;gt; returns an integer containing the 1-based position of the record&lt;br /&gt;
* Identifiers&lt;br /&gt;
** &amp;lt;code&amp;gt;getIDStr()&amp;lt;/code&amp;gt; returns const char* to this record&#039;s ID.  It includes the entire list of unique identifiers delimited by &#039;;&#039;&lt;br /&gt;
* Information on Bases&lt;br /&gt;
** &amp;lt;code&amp;gt;getRefStr()&amp;lt;/code&amp;gt; returns const char* to the entire reference string&lt;br /&gt;
** &amp;lt;code&amp;gt;getAltStr()&amp;lt;/code&amp;gt; returns const char* to the entire alternate string (alternates seperated by &#039;,&#039;)&lt;br /&gt;
** &amp;lt;code&amp;gt;getNumAlts()&amp;lt;/code&amp;gt; returns an integer containing the number of alternates in the alt string.&lt;br /&gt;
* Quality Information&lt;br /&gt;
** &amp;lt;code&amp;gt;getQual()&amp;lt;/code&amp;gt; returns a float containing the phred quality score&lt;br /&gt;
** &amp;lt;code&amp;gt;getQualStr()&amp;lt;/code&amp;gt; returns a string containing the phred quality score&lt;br /&gt;
* Filter Information - see [[#VcfRecordFilter|VcfRecordFilter]] for more information on extracting information&lt;br /&gt;
** &amp;lt;code&amp;gt;getFilter()&amp;lt;/code&amp;gt; returns a reference to the Info field information, a VcfRecordInfo object reference&lt;br /&gt;
** &amp;lt;code&amp;gt;passedAllFilters()&amp;lt;/code&amp;gt; returns true if the FILTER field indicates &amp;lt;code&amp;gt;PASS&amp;lt;/code&amp;gt;, false if not.&lt;br /&gt;
* Aditional Information&lt;br /&gt;
** &amp;lt;code&amp;gt;getInfo()&amp;lt;/code&amp;gt; returns a reference to the filter information, a VcfRecordFilter object reference&lt;br /&gt;
* Genotype Information&lt;br /&gt;
** &amp;lt;code&amp;gt;getGenotypeInfo()&amp;lt;/code&amp;gt; returns a reference to the information from the genotype fields, a VcfRecordGenotype object reference&lt;br /&gt;
** &amp;lt;code&amp;gt;allPhased()&amp;lt;/code&amp;gt; returns true if all the samples are phased and none unphased and false if any are not phased&lt;br /&gt;
** &amp;lt;code&amp;gt;allUnphased()&amp;lt;/code&amp;gt; returns true if all the samples are unphased and none phased and false if any are not unphased&lt;br /&gt;
** &amp;lt;code&amp;gt;hasAllGenotypeAlleles()&amp;lt;/code&amp;gt; returns true if all the samples have all the genotype alleles specified and false if any are missing or the GT field is missing&lt;br /&gt;
** &amp;lt;code&amp;gt;getNumSamples()&amp;lt;/code&amp;gt; returns the number of samples (that are kept)&lt;br /&gt;
** &amp;lt;code&amp;gt;getNumGTs(int index)&amp;lt;/code&amp;gt; returns the number of GTs for the specified sample index (starts at 0)&lt;br /&gt;
** &amp;lt;code&amp;gt;getGT(int sampleNum, unsigned int gtIndex)&amp;lt;/code&amp;gt; returns the integer GT value for the specified sampleNum and GT index (both start at 0).  A GT of VcfGenotypeSample::INVALID_GT is returned if the sampleNum or GT index is out of range.  A GT of VcfGenotypeSample::MISSING_GT if the GT is &#039;.&#039;  (This method is also found in VcfRecordGenotype.&lt;br /&gt;
&lt;br /&gt;
=== VcfRecordFilter ===&lt;br /&gt;
&amp;lt;code&amp;gt;VcfRecords&amp;lt;/code&amp;gt; contain the data from the &amp;lt;code&amp;gt;INFO&amp;lt;/code&amp;gt; field in a &amp;lt;code&amp;gt;VcfRecordFilter&amp;lt;/code&amp;gt; object.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Handling a Subset of Samples ==&lt;br /&gt;
&lt;br /&gt;
When reading a file if you only want to process/keep a subset of samples, use [[#Reading a Subset of Samples|Reading a Subset of Samples]].  When that method is used, only the specified samples are stored.  Any further processing will only be on those samples.&lt;br /&gt;
&lt;br /&gt;
Some methods allow the user to specify a subset of samples to operate on.  The subset specified when reading the VCF file, if any, is automatically applied since only those samples were stored.  If a different/additional subset needs to be applied for other processing, you can use the &amp;lt;code&amp;gt;VcfSubsetSamples&amp;lt;/code&amp;gt; class.&lt;br /&gt;
&lt;br /&gt;
To setup a VcfSubsetSamples object, pass the already set VCF header to:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
void VcfSubsetSamples::init(const VcfHeader&amp;amp; header, bool include)&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
Set the &amp;lt;code&amp;gt;include&amp;lt;/code&amp;gt; parameter to:&lt;br /&gt;
* true if all samples should be included except any that are specified as excluded. &lt;br /&gt;
* false if all samples should be excluded except any that are specified as included.&lt;br /&gt;
&lt;br /&gt;
NOTE: the header is not modified to add/remove any samples.&lt;br /&gt;
&lt;br /&gt;
To mark a specific sample as excluded use:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
bool VcfSubsetSamples::addExcludeSample(const char* sampleName);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;br /&gt;
To mark a specific sample as included use:&lt;br /&gt;
&amp;lt;source lang=&amp;quot;cpp&amp;quot;&amp;gt;&lt;br /&gt;
bool VcfSubsetSamples::addIncludeSample(const char* sampleName);&lt;br /&gt;
&amp;lt;/source&amp;gt;&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2853</id>
		<title>File:VcfReader.v1.tar</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2853"/>
		<updated>2011-02-04T15:50:34Z</updated>

		<summary type="html">&lt;p&gt;Youna: uploaded a new version of &amp;quot;File:VcfReader.v1.tar&amp;quot;:&amp;amp;#32;1. fixed the bug of IDfile.
2. included the functionality to subset variants of a given positions.&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2852</id>
		<title>File:VcfReader.v1.tar</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2852"/>
		<updated>2011-02-04T15:48:20Z</updated>

		<summary type="html">&lt;p&gt;Youna: uploaded a new version of &amp;quot;File:VcfReader.v1.tar&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2851</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2851"/>
		<updated>2011-02-04T15:47:45Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++. Please contact Youna Hu (youna@umich.edu) for comments, suggestions or questions.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder and type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2849</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2849"/>
		<updated>2011-02-03T22:57:05Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++. Please contact Youna Hu (youna@umich.edu) for comments, suggestions or questions.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder and type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2842</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2842"/>
		<updated>2011-02-01T22:10:53Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++. Please contact Youna Hu (youna@umich.edu) for comments, suggestions or questions.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder and type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2841</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2841"/>
		<updated>2011-02-01T19:29:43Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++. Please contact Youna Hu (youna@umich.edu) for comments, suggestions or questions.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2840</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2840"/>
		<updated>2011-02-01T19:29:29Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++. Please contact Youna Hu (youna@umich.edu) for comments, suggestion or questions.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2835</id>
		<title>File:VcfReader.v1.tar</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:VcfReader.v1.tar&amp;diff=2835"/>
		<updated>2011-01-31T17:18:02Z</updated>

		<summary type="html">&lt;p&gt;Youna: uploaded a new version of &amp;quot;File:VcfReader.v1.tar&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:RV3Tests.v1.tar&amp;diff=2834</id>
		<title>File:RV3Tests.v1.tar</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:RV3Tests.v1.tar&amp;diff=2834"/>
		<updated>2011-01-31T14:33:51Z</updated>

		<summary type="html">&lt;p&gt;Youna: uploaded a new version of &amp;quot;File:RV3Tests.v1.tar&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2833</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2833"/>
		<updated>2011-01-30T20:03:21Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD keep the log file from the annotation, which will be used to create the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2832</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2832"/>
		<updated>2011-01-30T20:02:44Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code,&lt;br /&gt;
            you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first &lt;br /&gt;
 to output a annotated vcf file. You SHOULD record the log file, which will be later used to define the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2831</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2831"/>
		<updated>2011-01-30T20:02:01Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code, you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: You should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on your vcf file first to output a annotated vcf file. You SHOULD record the log file, which will be later used to define the gene list.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2830</id>
		<title>RvTests</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=RvTests&amp;diff=2830"/>
		<updated>2011-01-30T20:00:42Z</updated>

		<summary type="html">&lt;p&gt;Youna: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
= Overview =&lt;br /&gt;
A few rare variants tests (Li-Leal&#039;s CMC and Madsen-Browning&#039;s weighted method) are implemented in the logisitc regression framework using C++.&lt;br /&gt;
&lt;br /&gt;
The source code is located at [[File:RV3Tests.v1.tar]]&lt;br /&gt;
&lt;br /&gt;
You can just download the tar file, extract it and go to the RV3Test.v1 folder to type make all to compile the code, the binary file will then be in the exectuables folder. &lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
See a detailed [[example]] here.&lt;br /&gt;
&lt;br /&gt;
= Syntax =&lt;br /&gt;
&lt;br /&gt;
This software uses command line interface as follows&lt;br /&gt;
&lt;br /&gt;
RARE VARIANT ANALYSIS OPTIONS:&lt;br /&gt;
                 GENOTYPE : --genofile [pos.012],&lt;br /&gt;
                            --geneList [outGeneSorted.txt], --cutoff [0.010],&lt;br /&gt;
                            --collapseChoice [or]&lt;br /&gt;
                PHENOTYPE : --phenofile [LDL.y.ID]&lt;br /&gt;
               COVARIATES : --covConsider, --covfile [covFile.ID.2.txt]&lt;br /&gt;
              PERMUTATION : --nPermute [10], --PermutationSeed [1]&lt;br /&gt;
   GENE LEVEL TEST RESULT : --geneGlobalTestOut [globalPermuteSummary.txt],&lt;br /&gt;
                            --geneTestpvalueFile [geneTestPvalues.txt]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
;GENOTYPE&lt;br /&gt;
&lt;br /&gt;
;--genofile: A genotype 012 matrix (.012 is the file) This file can be prepared by using the prepare012s &lt;br /&gt;
  source code [[File:vcfReader.v1.tar]]&lt;br /&gt;
  Again, extract the tar file and then go into the directory to type make all to compile the code, you then will find binary file in the executables folder. &lt;br /&gt;
&lt;br /&gt;
 Note: If you going to analyze nonsynonymous and stop annotated variants, &lt;br /&gt;
  you should use Yanming&#039;s vcf annotation [http://genome.sph.umich.edu/wiki/VcfCodingSnps] on the vcf file.&lt;br /&gt;
&lt;br /&gt;
Data File PREPARATION&lt;br /&gt;
          Input files : --vcf [LDL.test.vcf], --log [], --IDfile []&lt;br /&gt;
   Subsetting choices : --All&lt;br /&gt;
         Output files : --outputPrefix [subsetGeno],&lt;br /&gt;
                        --outputGeneList [LDL.geneList.txt]&lt;br /&gt;
   --vcf: Input vcf file &lt;br /&gt;
   --log: This is the log file from Yanming&#039;s annotation output, we use this log to obtain the gene list&lt;br /&gt;
   --IDfile specifies a file with one column of subject IDs to subject from the vcf file. &lt;br /&gt;
    If it is not specified, then all subjects are included for the format conversion.&lt;br /&gt;
   --All:  specifies 1 to include all variants and 0 to include only nonsyn and stop annotated variants.&lt;br /&gt;
   -- outputPrefix: Specify the prefix for the four output files which will be used in rvTests&lt;br /&gt;
   *.012: A genotype matrix with subjects as rows and variant sites as columns.&lt;br /&gt;
   *.012.pos: Chromosome and position numbers. &lt;br /&gt;
   *.012.indv: Subject IDs.&lt;br /&gt;
   *.012.frq: The frequency of the included variants.&lt;br /&gt;
  --outputGeneList:  Specify a file to store the gene list which will be used in rvTest.&lt;br /&gt;
  The list file looks like this &lt;br /&gt;
  1	OR4F5	69090	70008&lt;br /&gt;
  1	SAMD11	860529	871276&lt;br /&gt;
  1	NOC2L	879583	893918&lt;br /&gt;
  1	KLHL17	895966	901095&lt;br /&gt;
  1	PLEKHN1	901876	910482&lt;br /&gt;
  1	C1orf170	910578	912021&lt;br /&gt;
&lt;br /&gt;
;--geneList: This file is an output from prepare012s using the option --outputGeneList  with columns as chromosome number, gene Name, start position, end position. There should be no header for this file. &lt;br /&gt;
&lt;br /&gt;
THE CHROMOSOME NUMBERS SHOULD BE NUMERICS!!!! 1 - chromosome 1, DO NOT USE chr1.&lt;br /&gt;
&lt;br /&gt;
;--cutoff: This is the minor allele frequency, you can specify it as 0.01, 0.05 or etc.&lt;br /&gt;
;--collapseChoice: Specify one of {or,sum,wt}. or: Li-Leal&#039;s CMC test, sum: Use the number of rare variants for each subject as the score, wt: Madeson-Browning&#039;s weighted rare variant score.&lt;br /&gt;
&lt;br /&gt;
;PHENOTYPE&lt;br /&gt;
;--phenofile: A file where the first column is subject ID and the second column is phenotype (0 or 1).&lt;br /&gt;
&lt;br /&gt;
;COVARIATES&lt;br /&gt;
;--covConsider: Default = 0, no covariate is considered. 1. covariate is considered.&lt;br /&gt;
;--covfile: Covariate file with the first column as subject ID and the other columns are covariates needed to be considered in the model.&lt;br /&gt;
&lt;br /&gt;
;PERMUTATION&lt;br /&gt;
;--nPermute: Number of permutation for the evaluation of p values.&lt;br /&gt;
;-- PermutationSeed: Default = 1. Can be changed to other numbers too.&lt;br /&gt;
&lt;br /&gt;
;GENE LEVEL TEST RESULT:&lt;br /&gt;
;--geneGlobalTestOut: This file stores the 5% and 95% quantiles of the p values for all the genes at each permutation&lt;br /&gt;
;--geneTestPvalueFile: This file gives you the gene name, number of rare variants, count of variants in case/control and p values from the RV test specified by collapseChoice.&lt;/div&gt;</summary>
		<author><name>Youna</name></author>
	</entry>
</feed>