<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Ben+Lerch</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Ben+Lerch"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Ben_Lerch"/>
	<updated>2026-09-26T04:24:14Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Biostatistics_830:_Main_Page&amp;diff=8525</id>
		<title>Biostatistics 830: Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Biostatistics_830:_Main_Page&amp;diff=8525"/>
		<updated>2013-09-07T03:27:28Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Target Audience */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Objective ==&lt;br /&gt;
&lt;br /&gt;
Gene mapping studies study the relationship between genetic variation and susceptibility to human disease. These studies are changing rapidly with the availability of techniques for very large scale genetic analysis, whether based on sequencing or on genotyping. Biostatistics 830 is a Ph.D. level course that dissects some recently developed methods and the principles behind their implementation. It is meant to provide students with a toolkit to facilitate development and implementation of new statistical methods.&lt;br /&gt;
&lt;br /&gt;
For additional information, see also [[Biostatistics 830: Core Competencies|Core Competencies in Biostatistics Program covered by this course]].&lt;br /&gt;
&lt;br /&gt;
== Target Audience ==&lt;br /&gt;
&lt;br /&gt;
It is highly recommended that students registering for Biostatistics 830 should have previously completed [[Biostatistics 666]] and [[Biostatistics 615/815]], which are courses introducing methods for genetic analysis and programming principles, respectively.&lt;br /&gt;
&lt;br /&gt;
== Scheduling ==&lt;br /&gt;
&lt;br /&gt;
For Fall 2013, classes are scheduled for Mondays and Wednesdays, 3:00 - 4:30 pm. &lt;br /&gt;
&lt;br /&gt;
The final grade will take into account your performance in problem sets and worksheets as well as your participation in class.&lt;br /&gt;
&lt;br /&gt;
== Standards of Academic Conduct ==&lt;br /&gt;
&lt;br /&gt;
The following is an extract from the School of Public Health&#039;s Student Code of Conduct [http://www.sph.umich.edu/academics/policies/conduct.html]:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Student academic misconduct includes behavior involving plagiarism, cheating, fabrication, falsification of records or official documents, intentional misuse of equipment or materials, and aiding and abetting the perpetration of such acts. The preparation of reports, papers, and examinations, assigned on an individual basis, must represent each student’s own effort. Reference sources should be indicated clearly. The use of assistance from other students or aids of any kind during a written examination, except when the use of books or notes has been approved by an instructor, is a violation of the standard of academic conduct.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
In the context of this course, any work you hand-in should be your own and any material that is a transcript (or interpreted transcript) of work by others must be clearly labeled as such.&lt;br /&gt;
&lt;br /&gt;
== Required Reading ==&lt;br /&gt;
&lt;br /&gt;
* Albers CA, Lunter G, MacArthur DG, McVean G, Ouwehand WH, Durbin R (2010) Dindel: accurate indel calls from short-read data. Genome Res. 21:961-73&lt;br /&gt;
&lt;br /&gt;
* Browning SR, Browning BL (2007) Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering. &#039;&#039;Am J Hum Genet.&#039;&#039; &#039;&#039;&#039;81&#039;&#039;&#039;:1084-97. PMID: 17924348&lt;br /&gt;
&lt;br /&gt;
* Coventry A, Bull-Otterson LM, Liu X, Clark AG, Maxwell TJ, Crosby J, Hixson JE, Rea TJ, Muzny DM, Lewis LR, Wheeler DA, Sabo A, Lusk C, Weiss KG, Akbar H, Cree A, Hawes AC, Newsham I, Varghese RT, Villasana D, Gross S, Joshi V, Santibanez J, Morgan M, Chang K, Iv WH, Templeton AR, Boerwinkle E, Gibbs R, Sing CF (2010) Deep resequencing reveals excess rare recent variants consistent with explosive population growth. &#039;&#039;Nat Commun.&#039;&#039; &#039;&#039;&#039;1&#039;&#039;&#039;:131. PMID: 21119644&lt;br /&gt;
&lt;br /&gt;
* Delaneau O, Zagury JF, Marchini J (2013) Improved whole-chromosome phasing for disease and population genetic studies. &#039;&#039;Nat Methods.&#039;&#039; &#039;&#039;&#039;10&#039;&#039;&#039;:5-6. PMID: 23269371&lt;br /&gt;
&lt;br /&gt;
* Howie B, Fuchsberger C, Stephens M, Marchini J, Abecasis GR (2012) Fast and accurate genotype imputation in genome-wide association studies through pre-phasing. &#039;&#039;Nat Genet.&#039;&#039; &#039;&#039;&#039;44&#039;&#039;&#039;:955-9. PMID: 22820512&lt;br /&gt;
&lt;br /&gt;
* Iqbal Z, Caccamo M, Turner I, Flicek P, McVean G (2012) De novo assembly and genotyping of variants using colored de Bruijn graphs. &#039;&#039;Nat Genet.&#039;&#039; &#039;&#039;&#039;44&#039;&#039;&#039;:226-32. PMID: 22231483&lt;br /&gt;
&lt;br /&gt;
* Jun G, Flickinger M, Hetrick KN, Romm JM, Doheny KF, Abecasis GR, Boehnke M, Kang HM (2012) Detecting and estimating contamination of human DNA samples in sequencing and array-based genotype data. &#039;&#039;Am J Hum Genet.&#039;&#039; &#039;&#039;&#039;91&#039;&#039;&#039;:839-48. PMID: 23103226&lt;br /&gt;
&lt;br /&gt;
* Li H, Ruan J, Durbin R (2008) Mapping short DNA sequencing reads and calling variants using mapping quality scores. &#039;&#039;Genome Res.&#039;&#039; &#039;&#039;&#039;18&#039;&#039;&#039;:1851-8. PMID: 18714091&lt;br /&gt;
&lt;br /&gt;
* Li H, Durbin R (2009) Fast and accurate short read alignment with Burrows-Wheeler transform. &#039;&#039;Bioinformatics.&#039;&#039; &#039;&#039;&#039;25&#039;&#039;&#039;:1754-60. PMID: 19451168&lt;br /&gt;
&lt;br /&gt;
* Li H, Durbin R (2011) Inference of human population history from individual whole-genome sequences. &#039;&#039;Nature.&#039;&#039; &#039;&#039;&#039;475&#039;&#039;&#039;:493-6. PMID: 21753753&lt;br /&gt;
&lt;br /&gt;
* Li Y, Willer CJ, Ding J, Scheet P, Abecasis GR (2010) MaCH: using sequence and genotype data to estimate haplotypes and unobserved genotypes. &#039;&#039;Genet Epidemiol.&#039;&#039; &#039;&#039;&#039;34&#039;&#039;&#039;:816-34. PMID: 21058334&lt;br /&gt;
&lt;br /&gt;
* Lin DY, Zeng D (2010) Meta-analysis of genome-wide association studies: no efficiency gain in using individual participant data. &#039;&#039;Genet Epidemiol.&#039;&#039; &#039;&#039;&#039;34&#039;&#039;&#039;:60-6. PMID: 19847795&lt;br /&gt;
&lt;br /&gt;
* Liu et al (2013) http://arxiv.org/abs/1305.1318&lt;br /&gt;
&lt;br /&gt;
* Wen X, Stephens M (2010) Using linear predictors to impute allele frequencies from summary or pooled genotype data. &#039;&#039;Ann Appl Stat.&#039;&#039; &#039;&#039;&#039;4&#039;&#039;&#039;:1158-1182. PMID: 21479081 &lt;br /&gt;
&lt;br /&gt;
* Wu MC, Lee S, Cai T, Li Y, Boehnke M, Lin X (2011) Rare-variant association testing for sequencing data with the sequence kernel association test. Am J Hum Genet. 89:82-93&lt;br /&gt;
&lt;br /&gt;
* Zerbino DR, Birney E (2008) Velvet: algorithms for de novo short read assembly using de Bruijn graphs. &#039;&#039;Genome Res.&#039;&#039; &#039;&#039;&#039;18&#039;&#039;&#039;:821-9. PMID: 18349386&lt;br /&gt;
&lt;br /&gt;
== Course History ==&lt;br /&gt;
&lt;br /&gt;
This course is an ad-hoc course, first taught by Goncalo Abecasis in the Fall of 2013.&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7751</id>
		<title>QPLOT</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7751"/>
		<updated>2013-08-06T16:53:23Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Parameters */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Introduction =&lt;br /&gt;
&lt;br /&gt;
The qplot program calculates various summary statistics some of which are plotted in a PDF file. These statistics can be used to assess the sequencing quality of sequence reads mapped to the reference genome. The main statistics are empirical Phred scores which are calculated based on the background mismatch rate. Background mismatch rate is the rate that sequenced bases are different from the reference genome, EXCLUDING dbSNP positions. Other statistics include GC biases, insert size distribution, depth distribution, genome coverage, empirical Q20 count, and so on. &lt;br /&gt;
&lt;br /&gt;
In the following sections, we will guide you through: [[#Where to Find It |how to obtain qplot]], [[#Usage |how to use qplot]], [[#Built-in example |example outputs]], [[#anchorOfInteractiveQplot |interactive diagnostic plots]], and [[#Diagnose sequencing quality |real applications]] in which qplot has helped identify sequencing problems.&lt;br /&gt;
&lt;br /&gt;
= Where to Find It =&lt;br /&gt;
&lt;br /&gt;
You can obtain qplot in two ways: &lt;br /&gt;
&lt;br /&gt;
(1) Download the pre-compiled binary along with the source code as described in [[#Binary Download|Binary Download]]. &lt;br /&gt;
&lt;br /&gt;
(2) Download source code only and compile it on your own machine. Please follow the instruction in [[#Source Code Distribution|Source Code Distribution]] on fetching source code and building instructions.&lt;br /&gt;
&lt;br /&gt;
== Binary Download ==&lt;br /&gt;
&lt;br /&gt;
We have prepared a pre-compiled (under Ubuntu) qplot along with source code . You can download it from: [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot.20130627.tar.gz qplot.20130627.tar.gz (File Size: 1.7G)] &lt;br /&gt;
&lt;br /&gt;
The executable file is under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
In addition, we provided the necessary input files under qplot/data/ (NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100).&lt;br /&gt;
&lt;br /&gt;
You can also find an example BAM input file under qplot/example/chrom20.9M.10M.bam. It is taken from the 1000 Genome Project with sequencing reads aligned to chromosome 20 positions 8M to 9M.&lt;br /&gt;
&lt;br /&gt;
== Source Code Distribution ==&lt;br /&gt;
&lt;br /&gt;
We provide a source code only download in [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-source.20130627.tar.gz qplot-source.20130627.tar.gz]. Optionally, you can download example file and/or data file:&lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-example.tar.gz  example]: example input file, and expected outputs if you following the [[#Built-in example | direction]]. &lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-data.tar.gz resources data]: necessary input files for qplot, including NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100.&lt;br /&gt;
&lt;br /&gt;
You can put above file(s) in the same folder and follow these steps:&lt;br /&gt;
&lt;br /&gt;
* 1. Unarchive downloaded file&lt;br /&gt;
 tar zvxf qplot-source.20130627.tar.gz&lt;br /&gt;
&lt;br /&gt;
A new folder &#039;&#039;qplot&#039;&#039; will be created.&lt;br /&gt;
&lt;br /&gt;
* 2. Build libStatGen&lt;br /&gt;
 cd qplot&lt;br /&gt;
 (cd ../libStatGen; make cloneLib)&lt;br /&gt;
&lt;br /&gt;
This step will download a necessary software library [http://genome.sph.umich.edu/wiki/C%2B%2B_Library:_libStatGen libStatGen] and compile source code into a binary code library.&lt;br /&gt;
&lt;br /&gt;
* 3. Build qplot&lt;br /&gt;
 make &lt;br /&gt;
&lt;br /&gt;
This step will then build qplot. Upon success, the executable qplot can be found under qplot/bin/.&lt;br /&gt;
&lt;br /&gt;
* 4. (Optional) unarchive example and/or data&lt;br /&gt;
 tar zvxf qplot-example.tar.gz&lt;br /&gt;
&lt;br /&gt;
An example file, &#039;&#039;chrom20.9M.10M.bam&#039;&#039;, will be extracted to qplot/example/. It contains ~1.1 million aligned Illumina sequencing reads of NA12878 from 1000 Genome Project. Example command line, &#039;&#039;cmd.sh&#039;&#039;, example outputs, &#039;&#039;qplot.pdf&#039;&#039;, &#039;&#039;qplot.stats&#039;&#039;, and &#039;&#039;qplot.R&#039;&#039; are also provided and will be extracted qplot/example/ as well. &lt;br /&gt;
&lt;br /&gt;
 tar zvxf qplot-data.tar.gz&lt;br /&gt;
&lt;br /&gt;
Three files will be extracted to qplot/data/: &#039;&#039;human.g1k.v37-bs.umfa&#039;&#039; is binary NCBI reference genome build 37; &#039;&#039;dbSNP130.UCSC.coordinates.tbl&#039;&#039; is dbSNP version 130; and &#039;&#039;human.g1k.w100.gc&#039;&#039; is pre-calculated GC content with windows size 100.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Please download source code from [[]], the building &lt;br /&gt;
{{ToolGitRepo|repoName=qplot|noDownload=}}&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Usage =&lt;br /&gt;
&lt;br /&gt;
== Command line ==&lt;br /&gt;
&lt;br /&gt;
After you obtain the qplot executable (either by compiling the source code or by downloading the pre-compiled binary file), you will find the executable file under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
Here is the qplot help page by invoking qplot without any command line arguments:&lt;br /&gt;
&lt;br /&gt;
  some_linux_host &amp;gt; qplot/bin/qplot&lt;br /&gt;
    The following parameters are available.  Ones with &amp;quot;[]&amp;quot; are in effect:&lt;br /&gt;
    &lt;br /&gt;
    &lt;br /&gt;
    &lt;br /&gt;
                    References : --reference [/net/fantasia/home/zhanxw/software/qplot/data/human.g1k.v37.fa],&lt;br /&gt;
                                 --dbsnp [/net/fantasia/home/zhanxw/software/qplot/data/dbSNP130.UCSC.coordinates.tbl]&lt;br /&gt;
       GC content file options : --winsize [100]&lt;br /&gt;
                   Region list : --regions [], --invertRegion&lt;br /&gt;
                  Flag filters : --read1_skip, --read2_skip, --paired_skip,&lt;br /&gt;
                                 --unpaired_skip&lt;br /&gt;
                Dup and QCFail : --dup_keep, --qcfail_keep&lt;br /&gt;
               Mapping filters : --minMapQuality [0.00]&lt;br /&gt;
            Records to process : --first_n_record [-1]&lt;br /&gt;
              Lanes to process : --lanes []&lt;br /&gt;
         Read group to process : --readGroup []&lt;br /&gt;
            Input file options : --noeof&lt;br /&gt;
                  Output files : --plot [], --stats [], --Rcode [], --xml []&lt;br /&gt;
                   Plot labels : --label [], --bamLabel []&lt;br /&gt;
        Obsoleted (DO NOT USE) : --gccontent [], --create_gc&lt;br /&gt;
&lt;br /&gt;
== Input files ==&lt;br /&gt;
&lt;br /&gt;
qplot runs on the input BAM/SAM file(s) specified on the command-line after all other parameters.&lt;br /&gt;
&lt;br /&gt;
Additionally, three (3) precomputed files are required. &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--reference&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The reference genome is the same as karma reference genome. If the index files do not exist, qplot will create the index files &#039;&#039;&#039;automatically&#039;&#039;&#039; using the input reference fasta file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--dbsnp&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This file has two columns. First column is the chromosome name which must be consistent with the reference created above. Second column is 1-based SNP position. If you want to create your own dbSNP data from downloaded UCSC dbSNP file, one way to do it is: &amp;lt;code&amp;gt;cat dbsnp_129_b36.rod|grep &amp;quot;single&amp;quot; | awk &#039;$4-$3==1&#039; |cut -f2,4 &amp;gt; dbSNP_129_b36.tbl&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt; **OBSOLETED** --gccontent, --create_gc &amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Although GC content can be calculated on the fly each time, it is much more efficient to load a precomputed GC content from a file. &lt;br /&gt;
GC content file name is automatically determined in this format: &amp;lt;reference_genome_base_file_name&amp;gt;.winsize&amp;lt;gc_content_window_size&amp;gt;.gc.&lt;br /&gt;
For example, if your reference genome is human.g1k.v37.fa and the window size is 100, then the GC content file name is: human.g1k.v37.winsize100.gc .&lt;br /&gt;
&lt;br /&gt;
As it said, there is no need to use --gccontent to specify GC content file in each run.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt; input files &amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
QPLOT take SAM/BAM files.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Note&#039;&#039;: Before running qplot, it is critical to check how the chromosome names are coded. Some BAM/SAM files use just numbers, others use chr + numbers. &#039;&#039;&#039;You need to make sure that the chromosome names from the reference and dbSNP are consistent with the BAM/SAM files.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Parameters ==&lt;br /&gt;
&lt;br /&gt;
Some of the command line parameters are described here, but most are self explanatory.&lt;br /&gt;
&lt;br /&gt;
*Flag filter&lt;br /&gt;
&lt;br /&gt;
By default all reads are processed. If it is desired to check only the first read of a pair, use &amp;lt;code&amp;gt;--read2_skip&amp;lt;/code&amp;gt; to ignore the second read. And so on.&lt;br /&gt;
&lt;br /&gt;
*Duplication and QCFail&lt;br /&gt;
&lt;br /&gt;
By default reads marked as duplication and QCFail are ignored but can be retained by &lt;br /&gt;
 --dup_keep &lt;br /&gt;
or &lt;br /&gt;
 --qcfail_keep&lt;br /&gt;
&lt;br /&gt;
*Records to process &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--first_n_record&amp;lt;/code&amp;gt; option followed by a number, &#039;&#039;&#039;n&#039;&#039;&#039;, will enable qplot to read the first &#039;&#039;&#039;n&#039;&#039;&#039; reads to test the bam files and verify it works.&lt;br /&gt;
&lt;br /&gt;
* Lanes to process (only works for Illumina sequences)&lt;br /&gt;
&lt;br /&gt;
If the input bam files have more than one lane and only some of them need to be checked, use something like &amp;lt;code&amp;gt;--lanes 1,3,5&amp;lt;/code&amp;gt; to specify that only lanes 1, 3, and 5 need to be checked.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTE&#039;&#039;&#039; In order for this to work, the lane info has to be encoded in the read name such that the lane number is the second field with the delimiter &amp;quot;:&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
* Read group to process : &lt;br /&gt;
&lt;br /&gt;
Read group option can restrict qplot to process a subset of reads. For example, if BAM contain the following @RG tags:&lt;br /&gt;
&lt;br /&gt;
 @RG	ID:UM0348_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
&lt;br /&gt;
If specify nothing or not using &amp;quot;--readGroup&amp;quot;, QPLOT by default will process all reads; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348&amp;quot;, then only read group UM0348_1, UM_0348_2, UM_0348_3, UM_0348_4 will be processed; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348_1&amp;quot;, then only one read group UM0348_1 will be processed.&lt;br /&gt;
&lt;br /&gt;
* Input file options :&lt;br /&gt;
&lt;br /&gt;
BAM files are compress by BGZF algorithm and it should contain EOF by default. QPLOT will by default stop working when it does not found a valid EOF tag inside BAM files. &lt;br /&gt;
However, you can force QPLOT to continue process using --noeof. But you should be aware the input files may be corrupted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Mapping filters&lt;br /&gt;
&lt;br /&gt;
Qplot will exclude reads with lower mapping qualities than the user specified parameter, &amp;lt;code&amp;gt;--minMapQuality&amp;lt;/code&amp;gt;. By default, mapped reads with all mapping quality will be included in the analysis.&lt;br /&gt;
&lt;br /&gt;
*Region list&lt;br /&gt;
&lt;br /&gt;
If the interest of qplot is a list of regions, e.g. exons, this can be achieved by providing a list of regions. The regions should be in the form of &amp;quot;chr start end label&amp;quot; each line in the file (NOTE: &#039;&#039;start&#039;&#039; and &#039;&#039;end&#039;&#039; position are inclusive and they follow the convention of [http://genome.ucsc.edu/FAQ/FAQformat#format1 BED file]). &lt;br /&gt;
In order for this option to work, within each chromosome (contig) the regions have to be sorted by starting position, and also the input bam files have to be sorted. &lt;br /&gt;
For example, you can create a text file, region.txt like following:&lt;br /&gt;
&lt;br /&gt;
 1 100 500 region_A&lt;br /&gt;
 1 600 800 region_B&lt;br /&gt;
 2 100 300 region_C&lt;br /&gt;
 &lt;br /&gt;
Then specifying &amp;lt;code&amp;gt; --regions region.txt&amp;lt;/code&amp;gt; enables qplot to calculate various statistics out of sequenced bases only within the above 3 regions.&lt;br /&gt;
&lt;br /&gt;
Qplot also provides the &amp;lt;code&amp;gt;--invertRegion&amp;lt;/code&amp;gt; option. Enabling this option tells qplot to operate on those sequence bases that are outside the given region.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Plot labels&lt;br /&gt;
&lt;br /&gt;
Two kinds of labels are enabled. &amp;lt;code&amp;gt;--label&amp;lt;/code&amp;gt; is the label for the plot (default is empty) which is appended to the title of each subplot. &amp;lt;code&amp;gt;--bamLabels&amp;lt;/code&amp;gt; followed by a column separated list of labels provides the labels for each input SAM/BAM file, e.g. sample ID (default is numbers 1, 2, ... until the number of input bam files). For example:&lt;br /&gt;
 --label Run100 --bamLabels s1,s2,s3,s4,s5,s6,s7,s8&lt;br /&gt;
&lt;br /&gt;
== Output files ==&lt;br /&gt;
&lt;br /&gt;
There are three (optional) output files.&lt;br /&gt;
* &amp;lt;code&amp;gt;--plot &#039;&#039;qa.pdf&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a PDF file named &#039;&#039;qa.pdf&#039;&#039; containing 2 pages each with 4 figures. The plot is generated using Rscript.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--stats &#039;&#039;qa.stats&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a text file named &#039;&#039;qa.stats&#039;&#039; containing various summary statistics for each input BAM/SAM file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--Rcode &#039;&#039;qa.R&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate &#039;&#039;qa.R&#039;&#039; which is the R code used for plotting the figures in the &#039;&#039;qa.pdf&#039;&#039; file. If Rscript is not installed in the system, you can use the qa.R to generate the figures on other machines, or extract plotting data from each run and combine multiple runs together to generate more comprehensive plots (See [[#Example | Example]]).&lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
Qplot can generate diagnostic graphs, related R code, and summary statistics for each SAM/BAM file.&lt;br /&gt;
&lt;br /&gt;
== Built-in example ==&lt;br /&gt;
&lt;br /&gt;
In the pre-compiled binary download, you will find a subdirectory named examples. We provide a sample file from the 1000 Genome project, it contains aligned reads on chromosome 20 from position 8 Mbp to 9Mbp. You can invoke qplot using the following command line:&lt;br /&gt;
&lt;br /&gt;
 ../bin/qplot --reference ../data/human.g1k.v37.umfa --dbsnp ../data/dbSNP130.UCSC.coordinates.tbl --gccontent ../data/human.g1k.w100.gc --plot qplot.pdf --stats qplot.stats --Rcode qplot.R --label &amp;quot;chr20:9M-10M&amp;quot; chrom20.9M.10M.bam&lt;br /&gt;
&lt;br /&gt;
Sample outputs are listed below:&lt;br /&gt;
&lt;br /&gt;
1) Figure: [[Media:qplot.pdf | qplot.pdf]]&lt;br /&gt;
&lt;br /&gt;
2) Summary statistics:&lt;br /&gt;
 Stats\BAM       chrom20.9M.10M.bam&lt;br /&gt;
 TotalReads(e6)  1.11&lt;br /&gt;
 MappingRate(%)  97.24&lt;br /&gt;
 MapRate_MQpass(%)       97.24&lt;br /&gt;
 TargetMapping(%)        0.00&lt;br /&gt;
 ZeroMapQual(%)  2.39&lt;br /&gt;
 MapQual&amp;lt;10(%)   2.86&lt;br /&gt;
 PairedReads(%)  83.76&lt;br /&gt;
 ProperPaired(%) 71.34&lt;br /&gt;
 MappedBases(e9) 0.04&lt;br /&gt;
 Q20Bases(e9)    0.04&lt;br /&gt;
 Q20BasesPct(%)  88.63&lt;br /&gt;
 MeanDepth       42.22&lt;br /&gt;
 GenomeCover(%)  0.03&lt;br /&gt;
 EPS_MSE 1.81&lt;br /&gt;
 EPS_Cycle_Mean  18.71&lt;br /&gt;
 GCBiasMSE       0.01&lt;br /&gt;
 ISize_mode      137&lt;br /&gt;
 ISize_medium    184&lt;br /&gt;
 DupRate(%)      5.90&lt;br /&gt;
 QCFailRate(%)   0.00&lt;br /&gt;
 BaseComp_A(%)   29.9&lt;br /&gt;
 BaseComp_C(%)   20.1&lt;br /&gt;
 BaseComp_G(%)   20.2&lt;br /&gt;
 BaseComp_T(%)   29.8&lt;br /&gt;
 BaseComp_O(%)   0.1&lt;br /&gt;
&lt;br /&gt;
== Gallery of examples ==&lt;br /&gt;
&lt;br /&gt;
Here we show qplot can be applied in various sequencing scenarios. Also users can customize statistics generated by qplot to their needs.&lt;br /&gt;
&lt;br /&gt;
* Whole genome sequencing with 24-multiplexing&lt;br /&gt;
&lt;br /&gt;
With a customized script, we aggregated 24 bar-coded samples in the same graph.&lt;br /&gt;
The graph will help compare sequencing quality between samples. &lt;br /&gt;
&lt;br /&gt;
[[Media: qplot.Pool.9847.pdf | QPlot of 24 samples(PDF) ]]&lt;br /&gt;
&lt;br /&gt;
* Interactive qplot &lt;br /&gt;
&lt;br /&gt;
&amp;lt;span id=&amp;quot;anchorOfInteractiveQplot&amp;quot;&amp;gt;&amp;lt;/span&amp;gt;&lt;br /&gt;
Qplot can be interactive. In the following example, you can use mouse scroll to zoom in and zoom out on each graph and pan to a certain part of the graph.&lt;br /&gt;
By presenting qplot data on a web page, users can easily identify problematic sequencing samples. Users of qplot can customize its outputs into web page format greatly easing the data exploring process.&lt;br /&gt;
&lt;br /&gt;
[http://www-personal.umich.edu/~zhanxw/qplot.Pool.9847.html  QPlot of 24 samples(HTML) ]&lt;br /&gt;
&lt;br /&gt;
== Diagnose sequencing quality ==&lt;br /&gt;
&lt;br /&gt;
Qplot is designed and implemented for the need of checking sequencing quality. &lt;br /&gt;
Besides the example of analyzing RNA-seq data as shown in our manuscript, &lt;br /&gt;
here we demonstrate two additional scenarios in which qplot can help identify problems after obtaining sequencing data. &lt;br /&gt;
&lt;br /&gt;
* Base quality distributed abnormally&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBaseQual.pdf | Example of qplot helping to identify wrong phred base quality]]&lt;br /&gt;
&lt;br /&gt;
By checking the first graph &amp;quot;Empirical vs reported Phred score&amp;quot;, we found reported base qualities are shifted to the right.&lt;br /&gt;
In this particular example, &#039;33&#039; was incorrectly added to all base qualities. &lt;br /&gt;
When such data used in variant calling, we may increase false positive SNP variants.&lt;br /&gt;
&lt;br /&gt;
* Bar-coded samples&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBarCoding.pdf | Example of qplot identifying the effect of ignoring bar-coding]]&lt;br /&gt;
&lt;br /&gt;
By checking &amp;quot;Empirical phred score by cycle&amp;quot; (top right graph on the first page), we noticed the empirical qualities in the first several cycles are abnormally low. This phenomenon leads us to hypothesize that the first several bases have different properties. Further investigation confirmed that this sequencing was done using bar-coded DNA samples, but the analysis did not properly de-multiplex each sample.&lt;br /&gt;
&lt;br /&gt;
= Contact =&lt;br /&gt;
&lt;br /&gt;
Questions and requests should be sent to Bingshan Li ([mailto:bingshan@umich.edu bingshan@umich.edu]) or Xiaowei Zhan ([mailto:zhanxw@umich.edu zhanxw@umich.edu]) or Goncalo Abecasis ([mailto:goncalo@umich.edu goncalo@umich.edu])&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=EPACTS&amp;diff=7721</id>
		<title>EPACTS</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=EPACTS&amp;diff=7721"/>
		<updated>2013-08-02T17:12:13Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Creating marker group file */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;EPACTS&#039;&#039;&#039; (Efficient and Parallelizable Association Container Toolbox) is a versatile software pipeline to perform various statistical tests for identifying genome-wide association from sequence data through a user-friendly interface, both to scientific analysts and to method developers.&lt;br /&gt;
&lt;br /&gt;
== Join in EPACTS mailing list ==&lt;br /&gt;
&lt;br /&gt;
Please join in the [http://groups.google.com/group/epacts EPACTS Google Group] to ask / discuss / comment about EPACTS.&lt;br /&gt;
&lt;br /&gt;
== Lastest ChangeLog ==&lt;br /&gt;
* March 25th, 2013 : EPACTS v3.2.3 release&lt;br /&gt;
** Relaxed the checking of low-rank matrix in SKAT tests (to avoid unncessary skipping of genes)&lt;br /&gt;
* March 13th, 2013 : EPACTS v3.2.2 release&lt;br /&gt;
** Fixed an error which occasionally report mismatches in the number of samples&lt;br /&gt;
* March 9th, 2013 : EPACTS v3.2.1 release&lt;br /&gt;
**Fixed errors in loading the dynamic library&lt;br /&gt;
** Fixed errors in SKAT-O (thanks to Anubha Mahajan and Jason Flannick)&lt;br /&gt;
** Fixed bugs in emmax-CMC&lt;br /&gt;
** Added emmax-SKAT (contributed by Seunngeun Lee)&lt;br /&gt;
** And additional minor bug fixes&lt;br /&gt;
See [[#Full ChangeLog]] for full details&lt;br /&gt;
&lt;br /&gt;
== Key Features ==&lt;br /&gt;
&lt;br /&gt;
EPACTS currently provides the following set of key features&lt;br /&gt;
* Robust support for widely used format of sequence-based genotypes (VCF) and phenotypes with pedigree (PED)&lt;br /&gt;
** Efficient library for accessing VCF file to reduce computational burden to analyze large-scale sequencing data&lt;br /&gt;
** Support selecting markers by arbitrary combination of substring matching. &lt;br /&gt;
** Support for using genotype dosages instead of hard genotype calls&lt;br /&gt;
** Utilize PED format to perform test across multiple traits.&lt;br /&gt;
* Supports a large number of widely used statistical tests for single variant association and burden tests.&lt;br /&gt;
** See the &amp;quot;Currently Supported Statistical Tests&amp;quot; section below for more information&lt;br /&gt;
* Easy to Highly Parallelize Jobs&lt;br /&gt;
** Makefile-based partition into and ligation of multiple subtasks&lt;br /&gt;
** Parallel run of job is simply adding one parameter when running EPACTS &lt;br /&gt;
* Integrative and versatile framework that allows easy addition of additional statistical test&lt;br /&gt;
** Core input/output routines are implemented in C++&lt;br /&gt;
** Most statistical tests (except for EMMAX) are implemented in R&lt;br /&gt;
** Adding a simple R function to implement additional statistical test (See [[#Implementing Additional Statistical Tests]] for details)&lt;br /&gt;
* Useful utilities for post-association-analysis tasks&lt;br /&gt;
** Automatic functional annotation of associated variants&lt;br /&gt;
** Automatic generation of QQ and Manhattan Plot&lt;br /&gt;
** (TBA) Zoom plot for the significant associations&lt;br /&gt;
&lt;br /&gt;
== Obtaining EPACTS ==&lt;br /&gt;
&lt;br /&gt;
* The official release of EPACTS software is available at http://www.sph.umich.edu/csg/kang/epacts/ . &lt;br /&gt;
** From the CSG cluster, it is available at /net/fantasia/home/bin/epacts/&lt;br /&gt;
* Note that R (version 2.10 or higher) and gnuplot (version 4.2 or higher) must be installed in order to run EPACTS correctly.&lt;br /&gt;
&lt;br /&gt;
== Currently Supported Statistical Tests ==&lt;br /&gt;
&lt;br /&gt;
EPACTS supports the following sets of widely used statistical tests for single variant tests and burden tests&lt;br /&gt;
&lt;br /&gt;
=== Single Variant Tests ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;noinclude&amp;gt;&lt;br /&gt;
{|&amp;lt;/noinclude&amp;gt; border=&amp;quot;1&amp;quot; cellpadding=&amp;quot;4&amp;quot; cellspacing=&amp;quot;0&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse; font-size: 95%; clear: center;&amp;quot;&amp;lt;noinclude&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
! Test Name&lt;br /&gt;
! Phenotypes&lt;br /&gt;
! Covariates&lt;br /&gt;
! Computational Time&lt;br /&gt;
! Description&lt;br /&gt;
| Implemented by&lt;br /&gt;
|- &lt;br /&gt;
| b.wald &lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint)&lt;br /&gt;
| Slow&lt;br /&gt;
| Logisitic Wald Test &lt;br /&gt;
| Hyun Min Kang &amp;lt;br&amp;gt; (simply used glm in R)&lt;br /&gt;
|-&lt;br /&gt;
| b.score&lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out)&lt;br /&gt;
| Fast&lt;br /&gt;
| Logistic Score Test &amp;lt;br&amp;gt; (from Lin DY and Tang ZZ, AJHG 2011 89:354-67)&lt;br /&gt;
| Clement Ma &amp;amp; Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| b.firth&lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint)&lt;br /&gt;
| Slow&lt;br /&gt;
| Firth Bias-Corrected Logistic Likelihood Ratio Test &lt;br /&gt;
| Clement Ma&lt;br /&gt;
|-&lt;br /&gt;
| b.lrt&lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint)&lt;br /&gt;
| Slow&lt;br /&gt;
| Likelihood Ratio Test &lt;br /&gt;
| Clement Ma&lt;br /&gt;
|-&lt;br /&gt;
| b.glrt&lt;br /&gt;
| Binary&lt;br /&gt;
| NO&lt;br /&gt;
| Fast&lt;br /&gt;
| Genotype Likelihood Ratio Test &amp;lt;br&amp;gt; (use GL or PL field in VCF to perform case-control test)&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.lm&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint)&lt;br /&gt;
| Slow&lt;br /&gt;
| Linear Wald Test &lt;br /&gt;
| Hyun Min Kang &amp;lt;br&amp;gt; (as implemented in lm in R)&lt;br /&gt;
|-&lt;br /&gt;
| q.score&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out)&lt;br /&gt;
| Fast&lt;br /&gt;
| Quantitative Score Test &amp;lt;br&amp;gt; (from Lin DY and Tang ZZ, AJHG 2011 89:354-67)&lt;br /&gt;
| Clement Ma&lt;br /&gt;
|-&lt;br /&gt;
| q.linear&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out)&lt;br /&gt;
| Fast&lt;br /&gt;
| Linear Wald Test&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.reverse&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint)&lt;br /&gt;
| Slow&lt;br /&gt;
| Reverse regression &amp;lt;br&amp;gt; of phenotypes on binary genotypes (dominant model)&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.wilcox&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| Nonparametric Reverse regression &amp;lt;br&amp;gt; of phenotypes on binary genotypes (dominant model)&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.emmax&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| EMMAX &amp;lt;br&amp;gt; ( Kang et al (2010) Nat Genet 42:348-54 )&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
=== Gene-wise or group-wise tests ===&lt;br /&gt;
&lt;br /&gt;
&amp;lt;noinclude&amp;gt;&lt;br /&gt;
{|&amp;lt;/noinclude&amp;gt; border=&amp;quot;1&amp;quot; cellpadding=&amp;quot;4&amp;quot; cellspacing=&amp;quot;0&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse; font-size: 95%; clear: center;&amp;quot;&amp;lt;noinclude&amp;gt;&lt;br /&gt;
|-&lt;br /&gt;
! Test Name&lt;br /&gt;
! Phenotypes&lt;br /&gt;
! Covariates&lt;br /&gt;
! Computational Time&lt;br /&gt;
! Description&lt;br /&gt;
| Implemented by&lt;br /&gt;
|- &lt;br /&gt;
| b.collapse&lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint Estimation)&lt;br /&gt;
| Slow&lt;br /&gt;
| Logistic Wald Test between binary phenotypes and 0/1 collapsed variables&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| b.madsen&lt;br /&gt;
| Binary&lt;br /&gt;
| NO&lt;br /&gt;
| Slow&lt;br /&gt;
| Wilcoxon Rank Sum Test between binary phenotypes and weighted rare variant scores (slightly different version from the published method - it uses pooled allele frequency across cases and controls for weighting each variant)&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| b.wcnt&lt;br /&gt;
| Binary&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint Estimation)&lt;br /&gt;
| Slow&lt;br /&gt;
| Logistic Wald Test between binary phenotypes and weighted rare variant scores&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.reverse&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint Estimation)&lt;br /&gt;
| Slow&lt;br /&gt;
| Reverse regression of phenotypes on binary collapsed variables&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| q.wilcox&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| Nonparametric Reverse regression of phenotypes on collapsed variables&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| skat&lt;br /&gt;
| Binary/Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Joint Estimation)&lt;br /&gt;
| Slow&lt;br /&gt;
| SKAT-O Test by Lee et al, Biostatistics (2012)&lt;br /&gt;
| Seunggeun Lee &amp;lt;br&amp;gt; (adaptive by Xueling Sim and Hyun Min Kang)&lt;br /&gt;
|-&lt;br /&gt;
| VT&lt;br /&gt;
| Binary/Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed out first)&lt;br /&gt;
| Slow&lt;br /&gt;
| Variable Threshold Test &amp;lt;br&amp;gt; with adaptive permutation &amp;lt;br&amp;gt; Price et al, AJHG (2010) 86:832-8&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| emmaxCMC&lt;br /&gt;
| Binary/Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| Collapsing burden test using EMMAX&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| emmaxVT&lt;br /&gt;
| Binary/Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| Variable-threshold burden test using EMMAX&lt;br /&gt;
| Hyun Min Kang&lt;br /&gt;
|-&lt;br /&gt;
| mmskat&lt;br /&gt;
| Quantitative&lt;br /&gt;
| YES &amp;lt;br&amp;gt; (Regressed Out First)&lt;br /&gt;
| Slow&lt;br /&gt;
| SKAT test using EMMAX&lt;br /&gt;
| Seunggeun Lee &amp;amp; Hyun Min Kang&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
== Installation Details  ==&lt;br /&gt;
&lt;br /&gt;
If you want to use EPACTS in an Ubuntu platform, following the step below &lt;br /&gt;
&lt;br /&gt;
*Download EPACTS source distribution at http://www.sph.umich.edu/csg/kang/epacts/download/EPACTS-3.2.3.tar.gz (100MB) &lt;br /&gt;
*Uncompress EPACTS package, and install the package using the following set of commands&lt;br /&gt;
&lt;br /&gt;
  tar xzvf EPACTS-3.2.3.tar.gz&lt;br /&gt;
  cd EPACTS-3.2.3&lt;br /&gt;
  ./configure --prefix=/path/to/install&lt;br /&gt;
  make&lt;br /&gt;
  make install&lt;br /&gt;
&lt;br /&gt;
(Important Note: &#039;&#039;&#039;make sure to specify --prefix=/path/to/install&#039;&#039;&#039; to avoid installing to the default path /usr/local/, which you may not have the permission. /home/your_userid/epacts might be a good one, if you are not sure where to install)&lt;br /&gt;
  &lt;br /&gt;
* Now ${EPACTS_DIR} represents the &#039;/path/to/install&#039; directory&lt;br /&gt;
&lt;br /&gt;
* Download the reference FASTA files from 1000 Genomes FTP automatically by running the following commands&lt;br /&gt;
&lt;br /&gt;
  ${EPACTS_DIR}/bin/epacts download&lt;br /&gt;
&lt;br /&gt;
 (For advanced users, to save time for downloading the FASTA files (~900MB), you may copy a local copy of GRCh37 FASTA file and the index file to ${EPACTS_DIR}/share/EPACTS/)&lt;br /&gt;
&lt;br /&gt;
*Perform a test run by running the following command&lt;br /&gt;
&lt;br /&gt;
  ${EPACTS_DIR}/bin/test_run_epacts.sh&lt;br /&gt;
&lt;br /&gt;
In order to use EPACTS in the CSG cluster, you do not need to install them. You can directly use or make a copy of the in-house release version at &lt;br /&gt;
&lt;br /&gt;
 /net/fantasia/home/hmkang/bin/epacts/&lt;br /&gt;
&lt;br /&gt;
== Getting Started With Examples ==&lt;br /&gt;
If you are using EPACTS from the CSG cluster, please set the following environment variable&lt;br /&gt;
 EPACTS_DIR=/net/fantasia/home/hmkang/bin/epacts (in bash)&lt;br /&gt;
 setenv EPACTS_DIR /net/fantasia/home/hmkang/bin/epacts (in csh)&lt;br /&gt;
&lt;br /&gt;
If you downloaded EPACTS binary and please set EPACTS_DIR to the full path of the downloaded and uncompressed directory.&lt;br /&gt;
&lt;br /&gt;
=== All-in-one example ===&lt;br /&gt;
&lt;br /&gt;
To get started with EPACTS, run the following command will perform an example run&lt;br /&gt;
 ${EPACTS_DIR}/bin/test_run_epacts.sh&lt;br /&gt;
 &lt;br /&gt;
You will find a series of lines in test_run_epacts.sh script commented out for each possible test. &lt;br /&gt;
&lt;br /&gt;
The example phenotype (PED format) and genotype (VCF format) can be found at&lt;br /&gt;
 ${EPACTS_DIR}/share/EPACTS/&lt;br /&gt;
&lt;br /&gt;
=== Single Variant Test ===&lt;br /&gt;
&lt;br /&gt;
Or You can run EPACTS command yourself by running&lt;br /&gt;
 ${EPACTS_DIR}/epacts single \&lt;br /&gt;
   --vcf  ${EPACTS_DIR}/data/1000G_exome_chr20_example_softFiltered.calls.vcf.gz \&lt;br /&gt;
   --ped  ${EPACTS_DIR}/data/1000G_dummy_pheno.ped  \&lt;br /&gt;
   --min-maf 0.001 --chr 20 --pheno DISEASE --cov AGE --cov SEX --test b.score --anno \ &lt;br /&gt;
   --out out/test --run 2&lt;br /&gt;
&lt;br /&gt;
The command above will perform single variant association test using a dummy case-control phenotype file and a subset of 1000 genomes exome VCF file (chr20) using score test statistic for all variants over 1% of higher MAF using 2 parallel runs.&lt;br /&gt;
&lt;br /&gt;
You will see the 4 output files as the main outcome of the analysis&lt;br /&gt;
&lt;br /&gt;
==== Output Text of All Test Statistics ====&lt;br /&gt;
&lt;br /&gt;
The filename is out/test.single.b.score.epacts.gz and the contents will look like&lt;br /&gt;
 $ zcat out/test.single.b.score.epacts.gz | head&lt;br /&gt;
 #CHROM	BEGIN	END	MARKER_ID	NS	AC	CALLRATE	MAF	PVALUE	SCORE	N.CASE	N.CTRL	AF.CASE	AF.CTRL&lt;br /&gt;
 20	68303	68303	20:68303_A/G_Upstream:DEFB125	266	1	1	0.0018797	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	68319	68319	20:68319_C/A_Upstream:DEFB125	266	1.4467e-36	1	0	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	68396	68396	20:68396_C/T_Nonsynonymous:DEFB125	266	1	1	0.0018797	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76635	76635	20:76635_A/T_Intron:DEFB125	266	1.534e-37	1	0	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76689	76689	20:76689_T/C_Synonymous:DEFB125	266	0	1	0	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76690	76690	20:76690_T/C_Nonsynonymous:DEFB125	266	1	1	0.0018797	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76700	76700	20:76700_G/A_Nonsynonymous:DEFB125	266	0	1	0	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76726	76726	20:76726_C/G_Nonsynonymous:DEFB125	266	0	1	0	NA	NA	NA	NA	NA	NA&lt;br /&gt;
 20	76771	76771	20:76771_C/T_Nonsynonymous:DEFB125	266	3	1	0.0056391	0.68484	0.40587	145	121	0.013793	0.0082645&lt;br /&gt;
&lt;br /&gt;
==== Output Text of Top Associations ====&lt;br /&gt;
&lt;br /&gt;
Same type of file but containing top 5,000 association will be stored at out/test.epacts.top5000&lt;br /&gt;
&lt;br /&gt;
 $ head out/test.single.b.score.epacts.top5000 &lt;br /&gt;
 #CHROM	BEGIN	END	MARKER_ID	NS	AC	CALLRATE	MAF	PVALUE	SCORE	N.CASE	N.CTRL	AF.CASE	AF.CTRL&lt;br /&gt;
 20	1610894	1610894	20:1610894_G/A_Synonymous:SIRPG	266	138.64	1	0.26061	6.9939e-05	3.9765	145	121	0.65177	0.36476&lt;br /&gt;
 20	4162411	4162411	20:4162411_T/C_Intron:SMOX	266	204	1	0.38346	0.00055583	-3.4523	145	121	0.62759	0.93388&lt;br /&gt;
 20	34061918	34061918	20:34061918_T/C_Intron:CEP250	266	41.815	1	0.0786	0.00095471	3.3035	145	121	0.22543	0.075436&lt;br /&gt;
 20	4155948	4155948	20:4155948_G/A_Intron:SMOX	266	215	1	0.40414	0.0020792	-3.0787	145	121	0.68276	0.95868&lt;br /&gt;
 20	4680251	4680251	20:4680251_A/G_Nonsynonymous:PRNP	266	186	1	0.34962	0.0025962	3.0119	145	121	0.8069	0.57025&lt;br /&gt;
 20	36668874	36668874	20:36668874_G/A_Synonymous:RPRD1B	266	96	1	0.18045	0.003031	2.9646	145	121	0.44828	0.2562&lt;br /&gt;
 20	36641871	36641871	20:36641871_G/A_Synonymous:TTI1	266	10	1	0.018797	0.004308	-2.8547	145	121	0.0068966	0.07438&lt;br /&gt;
 20	1616892	1616892	20:1616892_A/G_Synonymous:SIRPG	266	144	1	0.27068	0.0051239	2.7991	145	121	0.63449	0.42975&lt;br /&gt;
 20	25038372	25038372	20:25038372_G/A_Intron:ACSS1	266	103.3	1	0.19418	0.005748	2.7618	145	121	0.47201	0.28813&lt;br /&gt;
&lt;br /&gt;
The key columns represents:&lt;br /&gt;
* &#039;&#039;&#039;NS&#039;&#039;&#039; : Number of phenotyped samples with non-missing genotypes &lt;br /&gt;
* &#039;&#039;&#039;AC&#039;&#039;&#039; : Total Non-reference Allele Count&lt;br /&gt;
* &#039;&#039;&#039;CALLRATE&#039;&#039;&#039; : Fraction of non-missing genotypes.&lt;br /&gt;
* &#039;&#039;&#039;MAF&#039;&#039;&#039; : Minor allele frequencies&lt;br /&gt;
* &#039;&#039;&#039;PVALUE&#039;&#039;&#039; : P-value of single variant test&lt;br /&gt;
* &#039;&#039;&#039;AF.CASE&#039;&#039;&#039; : Non-reference allele frequencies for cases&lt;br /&gt;
* &#039;&#039;&#039;AF.CTRL&#039;&#039;&#039; : Non-reference allele frequencies for controls&lt;br /&gt;
&lt;br /&gt;
==== Q-Q plot of test statistics (stratified by MAF) ====&lt;br /&gt;
&lt;br /&gt;
The file out/test.b.score.epacts.qq.pdf will be generated as shown below&lt;br /&gt;
&lt;br /&gt;
[[File:test_b_score_epacts_qq.png]]&lt;br /&gt;
&lt;br /&gt;
==== Manhattan Plot of Test Statistics ====&lt;br /&gt;
&lt;br /&gt;
The file out/test.b.score.epacts.mh.pdf will be generated for chr20 only. &lt;br /&gt;
&lt;br /&gt;
[[File:test_b_score_epacts_mh.png]]&lt;br /&gt;
&lt;br /&gt;
An example Genome-wide manhattan plot (from a genome-wide run) will look like below&lt;br /&gt;
&lt;br /&gt;
[[File:tes_b_score_epacts_mh_gw.png]]&lt;br /&gt;
&lt;br /&gt;
=== Gene-wise or group-wise burden test ===&lt;br /&gt;
&lt;br /&gt;
Gene-wise or group-wise burden test requires two steps. First, &#039;group&#039; file containing the list of &lt;br /&gt;
markers per group needs to be generated. Second, group-wise burden test needs to be run&lt;br /&gt;
&lt;br /&gt;
==== Creating marker group file ====&lt;br /&gt;
&lt;br /&gt;
The marker group file has the following format&lt;br /&gt;
&lt;br /&gt;
 [GROUP_ID]  [MARKER_ID_1]   [MARKER_ID_2]  .... [MARKER_ID_N]&lt;br /&gt;
&lt;br /&gt;
where &lt;br /&gt;
* [GROUP_ID] is a string representing the group (e.g. gene name)&lt;br /&gt;
* [MARKER_ID_K] is a marker key as a format of [CHROM]:[POS]_[REF]/[ALT] (NOTE THAT THIS IS DIFFERENT FROM TYPICAL VCF MARKER ID field)&lt;br /&gt;
&lt;br /&gt;
Note that [MARKER_ID_K] has to be sorted by increasing order of genomic coordinate&lt;br /&gt;
&lt;br /&gt;
In order to create gene-level group file from typically formatted VCF file, one may use the following utility &lt;br /&gt;
&lt;br /&gt;
 ${EPACTS_DIR}/epacts make-group --vcf [input-vcf] --out [output-group-file] --format [epacts, annovar, chaos or gatk] --nonsyn&lt;br /&gt;
&lt;br /&gt;
The above command create a file [output-group-file] containing a list of missense and nonsense variants per each gene. To incorporate different types of functional annotations, use --type option as follows&lt;br /&gt;
&lt;br /&gt;
 ${EPACTS_DIR}/epacts make-group --vcf [input-vcf] --out [output-group-file] --format [epacts, annovar, chaos or gatk] --type [function_type_1] --type [function_type_2] ...&lt;br /&gt;
&lt;br /&gt;
Type &#039;epacts makegroup -man&#039; for the detailed documentation&lt;br /&gt;
&lt;br /&gt;
==== Annotating VCF file using EPACTS ====&lt;br /&gt;
&lt;br /&gt;
If the VCF is not annotated, &#039;epacts makegroup&#039; cannot be used. In order to annotate VCF, one can use the example VCF using ANNOVAR as follows:&lt;br /&gt;
&lt;br /&gt;
 ${EPACTS_DIR}/epacts anno \&lt;br /&gt;
    --in ${EPACTS_DIR}/data/1000G_exome_chr20_example_softFiltered.calls.vcf.gz \&lt;br /&gt;
    --out ${EPACTS_DIR}/data/1000G_exome_chr20_example_softFiltered.calls.anno.vcf.gz&lt;br /&gt;
&lt;br /&gt;
The epacts anno script will add &amp;quot;ANNO=[function]:[genename]&amp;quot; entry into the INFO field based on gencodeV7 (default) or refGene database.&lt;br /&gt;
&lt;br /&gt;
It is important to check whether the VCF file is already annotated or not in order to avoid no or redundant annotation.&lt;br /&gt;
&lt;br /&gt;
==== Running Groupwise Test ====&lt;br /&gt;
&lt;br /&gt;
To perform a groupwise burden test on the example VCF (annotated as above), run the following command&lt;br /&gt;
&lt;br /&gt;
 ${EPACTS_DIR}/epacts group --vcf ${EPACTS_DIR}/data/1000G_exome_chr20_example_softFiltered.calls.anno.vcf.gz \&lt;br /&gt;
   --groupf ${EPACTS_DIR}/data/1000G_exome_chr20_example_softFiltered.calls.anno.grp --out out/test.gene.skat \&lt;br /&gt;
   --ped ${EPACTS_DIR}/data/1000G_dummy_pheno.ped --maxAF 0.05 \&lt;br /&gt;
   --chr 20 --pheno QT --cov AGE --cov SEX --test skat --skat-o --run 2&lt;br /&gt;
&lt;br /&gt;
==== Example Output ====&lt;br /&gt;
 $ head out/test.gene.skat.epacts.top5000&lt;br /&gt;
 #CHROM BEGIN   END     MARKER_ID       NS      FRAC_WITH_RARE     NUM_ALL_VARS    NUM_PASS_VARS   NUM_SING_VARS   PVALUE  STATRHO&lt;br /&gt;
 20     62607037        62608720        20:62607037-62608720_SAMD10     266     0.14662 9       5       1       0.0020064       1&lt;br /&gt;
 20     2816211 2820493 20:2816211-2820493_FAM113A      266     0.011278        12      2       1       0.0032542       0&lt;br /&gt;
 20     47245987        47361692        20:47245987-47361692_PREX1      266     0.1391  54      9       6       0.0054849       1&lt;br /&gt;
 20     34761734        34810279        20:34761734-34810279_EPB41L1    266     0.071429        14      7       5       0.0068492       0.2&lt;br /&gt;
 20     61340671        61391602        20:61340671-61391602_NTSR1      266     0.11278 24      9       3       0.011063        1&lt;br /&gt;
 20     48561952        48568644        20:48561952-48568644_RNF114     266     0.011278        4       2       1       0.015175        0.2&lt;br /&gt;
 20     60962895        60963559        20:60962895-60963559_RPS21      266     0.06015 6       3       2       0.016409        0&lt;br /&gt;
 20     55904961        55917801        20:55904961-55917801_SPO11      266     0.011278        11      3       3       0.018031        0&lt;br /&gt;
&lt;br /&gt;
The key columns represents:&lt;br /&gt;
* &#039;&#039;&#039;NS&#039;&#039;&#039; : Number of phenotyped samples with non-missing genotypes &lt;br /&gt;
* &#039;&#039;&#039;FRAC_WITH_RARE&#039;&#039;&#039; : Fraction of individual carrying rare variants below --max-maf (default : 0.05) threshold.&lt;br /&gt;
* &#039;&#039;&#039;NUM_ALL_VARS&#039;&#039;&#039; : Number of all variants defining the group.&lt;br /&gt;
* &#039;&#039;&#039;NUM_PASS_VARS&#039;&#039;&#039; : Number of variants passing the --min-maf, --min-mac, --max-maf, --min-callrate thresholds&lt;br /&gt;
* &#039;&#039;&#039;NUM_SING_VARS&#039;&#039;&#039; : Number of singletons among variants in NUM_PASS_VARS&lt;br /&gt;
* &#039;&#039;&#039;PVALUE&#039;&#039;&#039; : P-value of burden tests&lt;br /&gt;
* Other columns are test specific auxiliary columns. For example, in the VT test, the optimal MAF threshold is recorded as an auxiliary output column.&lt;br /&gt;
&lt;br /&gt;
=== Specialized Instruction for EMMAX tests ===&lt;br /&gt;
&lt;br /&gt;
EMMAX (Efficient Mixed Model Association eXpedited - Kang et al (2010) Nat Genet 42:348-54) is an efficient implementation of mixed model association accounting for sample structure including population structure and hidden relatedness. Currently EPACTS supports EMMAX association mapping in single variant test and CMC-like burden tests. &lt;br /&gt;
&lt;br /&gt;
Because EMMAX is based on linear model, the method fits better to quantiative traits than binary traits. However, p-values for binary traits are expected to be valid in the spirit of Armitage trend test, although the estimated effect size may not be precise.&lt;br /&gt;
&lt;br /&gt;
In order to run EMMAX analysis from sequence-based genotypes. We recommend running EPACTS multiple times using the following procedure.&lt;br /&gt;
&lt;br /&gt;
==== Single Variant EMMAX Association Analysis ====&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Creating Kinship Matrix&#039;&#039;&#039; : From VCF, we recommend to set a MAF (e.g. 0.01) and call rate (e.g. 0.95) threshold to select high-quality markers to generate kinship matrix as follows.&lt;br /&gt;
 ${EPACTS_DIR}/epacts make-kin \&lt;br /&gt;
  --vcf [input.vcf.gz] --ped  [input.ped (Optional)] --min-maf 0.01 --minCallRate 0.95 \&lt;br /&gt;
  --sepchr (if VCF is separated by chromosome) --out [outprefix.kinf] --run [# of parallel jobs]&lt;br /&gt;
&lt;br /&gt;
If you provide [input.ped] file, then it will calculate the subset the individuals contained in the PED file. &lt;br /&gt;
&lt;br /&gt;
The procedure above will create a file [outprefix.kinf] after splitting and merging the genomes into multiple pieces. If only a certain subset of SNPs needs to be considered due to target regions, LD-pruning, or any other reasons, a VCF containing the subset of markers must be created beforehand and should be used as input VCF file.&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Perform Single Variant Association&#039;&#039;&#039; : From VCF and PED, we recommend to use less stringent MAF threshold (e.g. 0.001) and call rate (e.g. 0.50) to perform single variant association&lt;br /&gt;
 ${EPACTS_DIR}/epacts single \&lt;br /&gt;
  --vcf [input.vcf.gz] --ped  [input.ped] --min-maf 0.001 --kin [outputprefix.kinf] \&lt;br /&gt;
  --sepchr --pheno [PHENO_NAME] --cov [COV1] --cov [COV2] --test q.emmax \&lt;br /&gt;
  --out [outprefix] --run [# of parallel jobs]&lt;br /&gt;
&lt;br /&gt;
The procedure above will perform single variant association analysis compatible to other types of single variant association analyses implemented in EPACTS&lt;br /&gt;
&lt;br /&gt;
==== Burden-style EMMAX Association Analysis ====&lt;br /&gt;
&lt;br /&gt;
In order to run EMMAX analysis from sequence-based genotypes. We recommend running EPACTS multiple times using the following procedure.&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Creating Kinship Matrix&#039;&#039;&#039; : See &#039;Creating Kinship Matrix&#039; section in [[#Single Variant EMMAX Association Analysis]]&lt;br /&gt;
* &#039;&#039;&#039;Create Marker Group&#039;&#039;&#039;&lt;br /&gt;
** By annotating the VCF and extracting missense and nonsense variants&lt;br /&gt;
*** [[#Annotating VCF file using ANNOVAR]] - This step will be required to create marker group file&lt;br /&gt;
*** [[#Creating marker group file]] - Assume that [group.grp] file is produced&lt;br /&gt;
** Or, by creating your own marker group information&lt;br /&gt;
*** See [[#Creating marker group file]] for details&lt;br /&gt;
* Run CMC-style burden test by&lt;br /&gt;
 ${EPACTS_DIR}/epacts group --groupf [group.grp] \&lt;br /&gt;
  --vcf [input.vcf.gz] --ped  [input.ped] --max-maf [max-MAF-for-rare-variants] \&lt;br /&gt;
  --kin [outputprefix.kinf] --sepchr --pheno [PHENO_NAME] --cov [COV1] --cov [COV2] \&lt;br /&gt;
  --test emmaxCMC --out [outprefix] &lt;br /&gt;
* Run Variable Threshold burden test by&lt;br /&gt;
 ${EPACTS_DIR}/epacts group --groupf [group.grp] \&lt;br /&gt;
  --vcf [input.vcf.gz] --ped  [input.ped] --max-maf [max-MAF-for-rare-variants] \&lt;br /&gt;
  --kin [outputprefix.kinf] --sepchr --pheno [PHENO_NAME] --cov [COV1] --cov [COV2] \&lt;br /&gt;
  --test emmaxVT --out [outprefix]&lt;br /&gt;
&lt;br /&gt;
== Preparing Your Own Input Data ==&lt;br /&gt;
&lt;br /&gt;
=== VCF file for Genotypes ===&lt;br /&gt;
&lt;br /&gt;
EPACTS support VCF files as input for association with the following requirement&lt;br /&gt;
* Input VCF file must be bgzipped and tabixed before running association to allow efficient random access of the file. Below is an example command to conver plain VCF into bgzipped and tabixed VCF&lt;br /&gt;
  bgzip input.vcf     ## this command will produce input.vcf.gz&lt;br /&gt;
  tabix -pvcf -f input.vcf.gz  ## this command will produce input.vcf.gz.tbi&lt;br /&gt;
* If the VCF file is separated by chromosome, the VCF file must contain the string &amp;quot;chr1&amp;quot; in the chromosome 1 file, and corresponding chromosome name for other chromosomes.&lt;br /&gt;
* Sample IDs in the VCF file must be consistent to those from PED file&lt;br /&gt;
* Currently EPACTS only support bi-allelic variants, but it handles SNPs, INDELs, snd SVs.&lt;br /&gt;
* Currently, EPACTS only support VCF aligned with NCBI build 37 coordinates&lt;br /&gt;
* An example VCF file from 1000 genome project is below. &lt;br /&gt;
 $ zcat example/1000G_integrated_phase1_chr20.vcf.gz | cut -f 1-10 | head -50 &lt;br /&gt;
 ##fileformat=VCFv4.1&lt;br /&gt;
 ##INFO=&amp;lt;ID=LCSNP,Number=0,Type=Flag,Description=&amp;quot;Genotype likelihood in Low coverage VCF in data integration&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=EXSNP,Number=0,Type=Flag,Description=&amp;quot;Genotype likelihood in Exome VCF in data integration&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=INDEL,Number=0,Type=Flag,Description=&amp;quot;Genotype likelihood in INDEL VCF in data integration&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=SV,Number=0,Type=Flag,Description=&amp;quot;Genotype likelihood in SV VCF in data integration&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=BAVGPOST,Number=1,Type=Float,Description=&amp;quot;Average posterior probability from beagle&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=BRSQ,Number=1,Type=Float,Description=&amp;quot;Genotype imputation quality estimate from beagle&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=LDAF,Number=1,Type=Float,Description=&amp;quot;MLE Allele Frequency Accounting for LD&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=AVGPOST,Number=1,Type=Float,Description=&amp;quot;Average posterior probability from MaCH/Thunder&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=RSQ,Number=1,Type=Float,Description=&amp;quot;Genotype imputation quality from MaCH/Thunder&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=ERATE,Number=1,Type=Float,Description=&amp;quot;Per-marker Mutation rate from MaCH/Thunder&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=THETA,Number=1,Type=Float,Description=&amp;quot;Per-marker Transition rate from MaCH/Thunder&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=CIEND,Number=2,Type=Integer,Description=&amp;quot;Confidence interval around END for imprecise variants&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=CIPOS,Number=2,Type=Integer,Description=&amp;quot;Confidence interval around POS for imprecise variants&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=END,Number=1,Type=Integer,Description=&amp;quot;End position of the variant described in this record&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=HOMLEN,Number=.,Type=Integer,Description=&amp;quot;Length of base pair identical micro-homology at event breakpoints&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=HOMSEQ,Number=.,Type=String,Description=&amp;quot;Sequence of base pair identical micro-homology at event breakpoints&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=SOURCE,Number=.,Type=String,Description=&amp;quot;Source of deletion call&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=SVLEN,Number=1,Type=Integer,Description=&amp;quot;Difference in length between REF and ALT alleles&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=SVTYPE,Number=1,Type=String,Description=&amp;quot;Type of structural variant&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=AC,Number=.,Type=Integer,Description=&amp;quot;Alternate Allele Count&amp;quot;&amp;gt;&lt;br /&gt;
 ##INFO=&amp;lt;ID=AN,Number=1,Type=Integer,Description=&amp;quot;Total Allele Count&amp;quot;&amp;gt;&lt;br /&gt;
 ##ALT=&amp;lt;ID=DEL,Description=&amp;quot;Deletion&amp;quot;&amp;gt;&lt;br /&gt;
 ##FORMAT=&amp;lt;ID=GT,Number=1,Type=String,Description=&amp;quot;Genotype&amp;quot;&amp;gt;&lt;br /&gt;
 ##FORMAT=&amp;lt;ID=DS,Number=1,Type=Float,Description=&amp;quot;Genotype dosage from MaCH/Thunder&amp;quot;&amp;gt;&lt;br /&gt;
 ##FORMAT=&amp;lt;ID=GL,Number=.,Type=Float,Description=&amp;quot;Genotype Likelihoods&amp;quot;&amp;gt;&lt;br /&gt;
 ##FORMAT=&amp;lt;ID=BD,Number=1,Type=Float,Description=&amp;quot;Genotype dosage from beagle&amp;quot;&amp;gt;&lt;br /&gt;
 #CHROM POS ID  REF ALT QUAL    FILTER  INFO    FORMAT  HG00096&lt;br /&gt;
 20 60479   .   C   T   100 PASS    LCSNP;EXSNP;BAVGPOST=1.000;BRSQ=0.894;LDAF=0.0020;AVGPOST=0.9995;RSQ=0.8779;ERATE=0.0005;THETA=0.0008;AC=4;AN=2184  GT:DS:GL:BD 0|0:0.000:-0.19,-0.46,-2.68:0.0022&lt;br /&gt;
 20 60522   .   T   TC  1588    PASS    INDEL;BAVGPOST=1.000;BRSQ=0.994;LDAF=0.0116;AVGPOST=0.9980;RSQ=0.9327;ERATE=0.0004;THETA=0.0167;AC=24;AN=2184   GT:DS:GL:BD 0|0:0.000:0.00,-0.90,-9.20:0&lt;br /&gt;
 20 60571   .   C   A   100 PASS    LCSNP;EXSNP;BAVGPOST=0.999;BRSQ=0.813;LDAF=0.0029;AVGPOST=0.9986;RSQ=0.8085;ERATE=0.0014;THETA=0.0014;AC=5;AN=2184  GT:DS:GL:BD 0|0:0.000:-0.05,-0.96,-5.00:0.0008&lt;br /&gt;
 20 60795   .   G   C   100 PASS    LCSNP;EXSNP;BAVGPOST=1.000;BRSQ=0.930;LDAF=0.0006;AVGPOST=0.9996;RSQ=0.7205;ERATE=0.0003;THETA=0.0041;AC=1;AN=2184  GT:DS:GL:BD 0|0:0.000:-0.03,-1.21,-5.00:0.0001&lt;br /&gt;
 20 60810   .   G   GA  127 PASS    INDEL;BAVGPOST=1.000;BRSQ=0.862;LDAF=0.0013;AVGPOST=0.9987;RSQ=0.5684;ERATE=0.0004;THETA=0.0061;AC=2;AN=2184    GT:DS:GL:BD 0|0:0.000:0.00,-1.80,-18.80:0&lt;br /&gt;
&lt;br /&gt;
=== PED file for Phenotypes and Covariates ===&lt;br /&gt;
&lt;br /&gt;
EPACTS accepts a PED format supported by MERLIN or PLINK software to represent phenotypes. For example, the example.ped file and example.dat file can represent the phenotypes and corresponding column name (from 6th column and after). &lt;br /&gt;
&lt;br /&gt;
 $ head example.ped&lt;br /&gt;
 13281  NA12344 NA12347 NA12348 1   1   94.17   66.1&lt;br /&gt;
 13281  NA12347 0   0   1   1   109.54  44.0&lt;br /&gt;
 13281  NA12348 0   0   2   2   119.40  46.6&lt;br /&gt;
 1328   NA06984 0   0   1   2   87.72   39.3&lt;br /&gt;
 1328   NA06989 0   0   2   1   100.60  41.7&lt;br /&gt;
 1328   NA12329 NA06984 NA06989 2   1   100.85  46.4&lt;br /&gt;
 13291  NA06986 0   0   1   2   91.94   61.9&lt;br /&gt;
 13291  NA06995 NA07435 NA07037 1   2   104.36  57.4&lt;br /&gt;
 13291  NA06997 NA06986 NA07045 2   2   107.53  53.1&lt;br /&gt;
&lt;br /&gt;
 $ cat example.dat&lt;br /&gt;
 A DISEASE&lt;br /&gt;
 T QT&lt;br /&gt;
 T AGE&lt;br /&gt;
&lt;br /&gt;
EPACTS also accept a PED format with header information. The above file can be combined into one file as follows&lt;br /&gt;
&lt;br /&gt;
 $ head data/1000G_dummy_pheno.ped&lt;br /&gt;
 #FAM_ID    IND_ID  FAT_ID  MOT_ID  SEX DISEASE QT  AGE&lt;br /&gt;
 13281  NA12344 NA12347 NA12348 1   1   94.17   66.1&lt;br /&gt;
 13281  NA12347 0   0   1   1   109.54  44.0&lt;br /&gt;
 13281  NA12348 0   0   2   2   119.40  46.6&lt;br /&gt;
 1328   NA06984 0   0   1   2   87.72   39.3&lt;br /&gt;
 1328   NA06989 0   0   2   1   100.60  41.7&lt;br /&gt;
 1328   NA12329 NA06984 NA06989 2   1   100.85  46.4&lt;br /&gt;
 13291  NA06986 0   0   1   2   91.94   61.9&lt;br /&gt;
 13291  NA06995 NA07435 NA07037 1   2   104.36  57.4&lt;br /&gt;
 13291  NA06997 NA06986 NA07045 2   2   107.53  53.1&lt;br /&gt;
&lt;br /&gt;
The column names can be used to identify the names of phenotypes and covariates in the analysis.&lt;br /&gt;
&lt;br /&gt;
== Frequently Asked Questions ==&lt;br /&gt;
=== Installation ===&lt;br /&gt;
# How should I install EPACTS? &lt;br /&gt;
#* See [[EPACTS#Installation_Details | Installation Details]]&lt;br /&gt;
# I am having the following error message &#039;&#039;&#039;configure: error: libR.{so,a} was not found. Please install it at http://www.r-project.org/ first&#039;&#039;&#039;. What do I have to do?&lt;br /&gt;
#* First, you need to find out where R was installed. Try to type &amp;quot;locate libR.so&amp;quot; and see if it returns anything&lt;br /&gt;
#* If &amp;quot;locate libR.so&amp;quot; returns you something, as explained [[EPACTS#Installation_Details | Installation Details]], try to add &amp;quot;LDFLAGS=-L/path/to/R/library&amp;quot; and rerun &#039;&#039;&#039;configure&#039;&#039;&#039; and &#039;&#039;&#039;make&#039;&#039;&#039;&lt;br /&gt;
#* If you cannot find libR.so, you make have to recompile R with --enable-R-shlib option as described in http://cran.r-project.org/doc/manuals/R-admin.html#Installation&lt;br /&gt;
&lt;br /&gt;
=== Input Files ===&lt;br /&gt;
# What is VCF?&lt;br /&gt;
#* VCF refers to Variant Call Format&lt;br /&gt;
#* See [[http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41 1000 Genomes wiki page]] for the detailed description of VCF format&lt;br /&gt;
# Should input VCF be compressed into certain format?&lt;br /&gt;
#* Correct. EPACTS assumes that VCF file is bgzipped and tabixed already.&lt;br /&gt;
#* See [[#VCF file for Genotypes]] for details.&lt;br /&gt;
# What are the additional requirements for input VCF file?&lt;br /&gt;
#* Input VCF file used for association mapping must contain individual genotype information at 10-th or higher order columns.&lt;br /&gt;
#* GT field must be encoded as haploid or diploid&lt;br /&gt;
#* Bi-allelic SNPs only : Currently EPACTS may not handle multi-allelic SNPs correctly.&lt;br /&gt;
#* If non-GT field is used, the field is considered as dosage and should be a single numeric value.&lt;br /&gt;
# What are the acceptable input format to encode phenotypes and covariates?&lt;br /&gt;
#* See [[#PED file for Phenotypes and Covariates]] for the detailed information&lt;br /&gt;
# How should I encode binary phenotypes?&lt;br /&gt;
#* If you encode your phenotypes into two different numeric values (e.g. 0/1 or 1/2), EPACTS will automatically recognize them as binary phenotypes and encode them into 1/2 values. Higher value will be considered as cases for case-control association&lt;br /&gt;
# How should I encode missing genotypes?&lt;br /&gt;
#* The default code missing phenotypes in EPACTS are &#039;NA&#039;&lt;br /&gt;
#* One may use --missing option to specify different types of missing values&lt;br /&gt;
#* The encoding of missing genotypes follows the VCF specificiation&lt;br /&gt;
# How do I match the relationship between VCF and PED input files?&lt;br /&gt;
#* EPACTS will assume that the individual IDs in each VCF and PED file are unique, and they follow the saming convention. Thus, the individual IDs overlapping between VCF and PED files will be considered in the associations&lt;br /&gt;
# How the individuals with missing phenotypes are handled?&lt;br /&gt;
#* Currently, EPACTS will automatically remove the individuals without phenotypes or covariates. If one wants to use imputed covariates to increase sample size, the PED file must contain the imputed covariate values.&lt;br /&gt;
#* Markers with missing genotypes won&#039;t be discarded automatically. It can be explicitly discarded by --minCallRate option when performing association&lt;br /&gt;
&lt;br /&gt;
=== Output Files ===&lt;br /&gt;
# Which output files should I be looking at?&lt;br /&gt;
#* [[#Output Text of Top Associations]] is the key file to look at the individual top associations&lt;br /&gt;
#* [[#Q-Q plot of test statistics (stratified by MAF)]] will be important to see the global distribution of test statistics and examine if there are apparent inflation of test statistics&lt;br /&gt;
#* [[#Manhattan Plot of Test Statistics]] will inform us the genome-wide distribution of association signals&lt;br /&gt;
#* [[#Output Text of All Test Statistics]] will contain the full information of test results across all units tested&lt;br /&gt;
# The Q-Q and Manhattan plots cannot be found. Why?&lt;br /&gt;
#* It is probably because gnuplot 4.2 or higher is not installed in your system, or they are included but cannot be found in your ${PATH}. Please visit [[http://gnuplot.info/ GNUPLOT web page]] for installation.&lt;br /&gt;
&lt;br /&gt;
=== More questions ===&lt;br /&gt;
# If you have more questions, please contact [[mailto:hmkang@umich.edu Hyun Min Kang]].&lt;br /&gt;
&lt;br /&gt;
== Detailed Options ==&lt;br /&gt;
&lt;br /&gt;
The detailed options can viewed by running the following commands&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts -man           (for overall structure) &lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts single -man    (for single variant test)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts group -man     (for groupwise test)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts anno -man      (for annotation)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts plot -man      (for QQ and Manhattan plot)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts zoom -man      (for zoom plot)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts meta -man      (for meta-analysis)&lt;br /&gt;
 ${EPACTS_DIR}/bin/epacts makegroup -man (for creating gene group)&lt;br /&gt;
&lt;br /&gt;
== Implementing Additional Statistical Tests ==&lt;br /&gt;
&lt;br /&gt;
In order to add additional statistical test to EPACTS, the following procedure are recommended&lt;br /&gt;
&lt;br /&gt;
# Create a file named &#039;single.[testname].R&#039; for single variant test or &#039;gene.[testname].R&#039; for gene-level test under ${EPACTS_DIR}/share/EPACTS/&lt;br /&gt;
# Test your implementation using --test [testname] option to perform sanity check and debugging&lt;br /&gt;
# If you want to add your test in the official in-house version, please send your code to Hyun&lt;br /&gt;
&lt;br /&gt;
Below is an example of a single variant test implementation ( single.q.lm.R )&lt;br /&gt;
 ## Core functions of EPACTS to perform association&lt;br /&gt;
 &lt;br /&gt;
 ##################################################################&lt;br /&gt;
 ## SINGLE VARIANT TEST&lt;br /&gt;
 ## INPUT VARIABLES:&lt;br /&gt;
 ##   n        : total # of individuals&lt;br /&gt;
 ##   NS       : number of called samples&lt;br /&gt;
 ##   AC       : allele count&lt;br /&gt;
 ##   MAF      : minor allele frequency&lt;br /&gt;
 ##   vids     : indices from 1:nrow(NS) after AF/AC threshold&lt;br /&gt;
 ##   genos    : genotype matrix (after AF/AC threshold)&lt;br /&gt;
 ## EXPECTED OUTPUT : list(p, addcols, addnames) for each genos row&lt;br /&gt;
 ##   p        : p-value&lt;br /&gt;
 ##   add      : additional columns to add&lt;br /&gt;
 ##   cname    : column names for additional columns&lt;br /&gt;
 ##################################################################  &lt;br /&gt;
 &lt;br /&gt;
 ## single.lm() : Use built-in lm() function to perform association&lt;br /&gt;
 ## KEY FEATURES : SIMPLE, BUT MAY BE SLOW&lt;br /&gt;
 ##                GOOD SNIPPLET TO START A NEW FUNCTION&lt;br /&gt;
 ## TRAITS  : QUANTITATIVE&lt;br /&gt;
 ## RETURNS : PVALUE, BETA, SEBETA, TSTAT&lt;br /&gt;
 ## MISSING VALUES : IGNORED&lt;br /&gt;
 single.q.lm &amp;lt;- function() {&lt;br /&gt;
   cname &amp;lt;- c(&amp;quot;BETA&amp;quot;,&amp;quot;SEBETA&amp;quot;,&amp;quot;TSTAT&amp;quot;) # column names for additional variables in the EPACTS output&lt;br /&gt;
   m &amp;lt;- nrow(genos)&lt;br /&gt;
   p &amp;lt;- rep(NA,m)&lt;br /&gt;
   add &amp;lt;- matrix(NA,m,3) ## BETA, SEBETA, TSTAT&lt;br /&gt;
   if ( m &amp;gt; 0 ) {&lt;br /&gt;
    for(i in 1:m) {&lt;br /&gt;
      r &amp;lt;- summary(lm(pheno~genos[i,]+cov-1))$coefficients[1,]  # run simple linear regression&lt;br /&gt;
      p[i] &amp;lt;- r[4]   # store p-value to p[i]&lt;br /&gt;
      add[i,] &amp;lt;- r[1:3] # store additional variables to add[i,]&lt;br /&gt;
    }&lt;br /&gt;
  }&lt;br /&gt;
  return(list(p=p,add=add,cname=cname))&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
As described in the comment, you may assume that the following variables are available for use for testing association across m markers&lt;br /&gt;
* n (scalar) : total number of individuals&lt;br /&gt;
* NS (M * 1 vector) : Number of called samples for each marker&lt;br /&gt;
* AC (M * 1 vector) : Non-reference allele count for each marker&lt;br /&gt;
* MAF (M * 1 vector) : Minor allele frequency&lt;br /&gt;
* vids (m * 1 vector) : indices of markers passing the inclusion criteria (e.g. MAF threshold) among 1:M &lt;br /&gt;
* genos (m * n matrix) : genotype matrix as a input for association test&lt;br /&gt;
&lt;br /&gt;
The output variables to generate is as follows&lt;br /&gt;
* p (m * 1 vector) : p-value matrix as output&lt;br /&gt;
* add (m * c matrix) : additional columns as output of test (such as SCORE, BETA, etc)&lt;br /&gt;
* cname (c * 1 vector) : column names of add&lt;br /&gt;
&lt;br /&gt;
In the output files, the following columns will be displayed&lt;br /&gt;
# MARKER : Marker ID&lt;br /&gt;
# NS : Number of called samples&lt;br /&gt;
# AC : Non-ref allele count&lt;br /&gt;
# CALLRATE : Call rate = NS/n&lt;br /&gt;
# MAF : Minor allele frequency&lt;br /&gt;
# PVALUE : P-values&lt;br /&gt;
# Additional columns specified by return values &#039;add&#039;&lt;br /&gt;
&lt;br /&gt;
Below is an example of a gene-lvel variant test implementation ( single.q.lm.R )&lt;br /&gt;
&lt;br /&gt;
 ##################################################################&lt;br /&gt;
 ## GENE-LEVEL BURDEN TEST&lt;br /&gt;
 ## INPUT VARIABLES: &lt;br /&gt;
 ##   n        : total # of individuals&lt;br /&gt;
 ##   genos    : genotype matrix for each gene&lt;br /&gt;
 ##   NS       : number of called samples for each marker&lt;br /&gt;
 ##   AC       : allele count for each marker&lt;br /&gt;
 ##   MAC      : minor allele count for each marker&lt;br /&gt;
 ##   MAF      : minor allele frequency&lt;br /&gt;
 ##   vids     : indices from 1:n after AF/AC threshold&lt;br /&gt;
 ## EXPECTED OUTPUT : list(p, addcols, addnames) for each genos row&lt;br /&gt;
 ##   p        : p-value&lt;br /&gt;
 ##   add      : additional column values&lt;br /&gt;
 ##   cname    : additional column names&lt;br /&gt;
 ##################################################################      &lt;br /&gt;
 &lt;br /&gt;
 ## gene.q.reverse() : Reverse logistic regression&lt;br /&gt;
 ## KEY FEATURES : 0/1 collapsing variable ~ rare variants&lt;br /&gt;
 ## TRAITS  : QUANTITATIVE (GAUSSIAN)&lt;br /&gt;
 ## RETURNS : PVALUE, BETA, SEBETA, ZSTAT&lt;br /&gt;
 ## MISSING VALUE : IMPUTED AS MAJOR ALLELES&lt;br /&gt;
 gene.q.reverse &amp;lt;- function() {&lt;br /&gt;
   cname &amp;lt;- c(&amp;quot;BETA&amp;quot;,&amp;quot;SEBETA&amp;quot;,&amp;quot;ZSTAT&amp;quot;)&lt;br /&gt;
   m &amp;lt;- nrow(genos)&lt;br /&gt;
   if ( m &amp;gt; 0 ) {&lt;br /&gt;
     g &amp;lt;- as.double(colSums(genos,na.rm=T) &amp;gt; 0)&lt;br /&gt;
     sg &amp;lt;- sum(g)&lt;br /&gt;
     if ( ( sg &amp;gt; 0 ) &amp;amp;&amp;amp; ( sg &amp;lt; n ) ) {&lt;br /&gt;
       r &amp;lt;- glm(g~pheno+cov-1,family=binomial)&lt;br /&gt;
        if ( ( r$converged ) &amp;amp;&amp;amp; ( ! r$boundary ) ) {&lt;br /&gt;
         return(list(p=summary(r)$coefficients[1,4],&lt;br /&gt;
                     add=summary(r)$coefficients[1,1:3],&lt;br /&gt;
                     cname=cname))&lt;br /&gt;
       }&lt;br /&gt;
     }&lt;br /&gt;
   }&lt;br /&gt;
   return(list(p=NA,add=rep(NA,3),cname=cname))&lt;br /&gt;
 }&lt;br /&gt;
&lt;br /&gt;
Similar to gene-level test, you may assume the following variables exist for testing A SINGLE GENE. Note that M is the number of markers spanning the gene region&lt;br /&gt;
&lt;br /&gt;
* n (scalar) : total number of individuals&lt;br /&gt;
* NS (M * 1 vector) : Number of called samples for each marker &lt;br /&gt;
* AC (M * 1 vector) : Non-reference allele count for each marker&lt;br /&gt;
* MAC (M * 1 vector) : Minor allele count&lt;br /&gt;
* MAF (M * 1 vector) : Minor allele frequency&lt;br /&gt;
* vids (m * 1 vector) : indices of markers passing the inclusion criteria (e.g. MAF threshold) among 1:M &lt;br /&gt;
* genos (m * n matrix) : genotype matrix as a input for association test&lt;br /&gt;
&lt;br /&gt;
The output variables to generate is as follows&lt;br /&gt;
* p (scalar) : p-value matrix as output&lt;br /&gt;
* add (c * 1 vector) : additional columns as output of test (such as SCORE, BETA, etc)&lt;br /&gt;
* cname (c * 1 vector) : column names of add&lt;br /&gt;
&lt;br /&gt;
In the output files, the following columns will be displayed&lt;br /&gt;
# MARKER : Marker ID&lt;br /&gt;
# NS : Number of called samples&lt;br /&gt;
# MAF_BURDEN : MAF of 0/1 collapsing variables (existence of rare variants)&lt;br /&gt;
# NUM_ALL_VARS : Number of all variants within the gene&lt;br /&gt;
# NUM_RARE_VARS : Number of rare variants below the max-MAF threshold&lt;br /&gt;
# NUM_SING_VARS : Number of singleton variants&lt;br /&gt;
# PVALUE : P-value from the test&lt;br /&gt;
# Additional columns specified by return values &#039;add&#039;&lt;br /&gt;
&lt;br /&gt;
== Full ChangeLog ==&lt;br /&gt;
* March 25th, 2013 : EPACTS v3.2.3 release&lt;br /&gt;
** Relaxed the checking of low-rank matrix in SKAT tests (to avoid unncessary skipping of genes)&lt;br /&gt;
* March 13th, 2013 : EPACTS v3.2.2 release&lt;br /&gt;
** Fixed an error which occasionally report mismatches in the number of samples&lt;br /&gt;
* March 9th, 2013 : EPACTS v3.2.1 release&lt;br /&gt;
**Fixed errors in loading the dynamic library&lt;br /&gt;
** Fixed errors in SKAT-O (thanks to Anubha Mahajan and Jason Flannick)&lt;br /&gt;
** Fixed bugs in emmax-CMC&lt;br /&gt;
** Added emmax-SKAT (contributed by Seunngeun Lee)&lt;br /&gt;
** And additional minor bug fixes&lt;br /&gt;
* February 28th, 2013 : EPACTS v3.2.0 release&lt;br /&gt;
** R package installation bug (for some users) was fixed&lt;br /&gt;
** A bug in the MAF error for high frequency variants (AF&amp;gt;0.25) was now fixed&lt;br /&gt;
** SKAT version is updated to 0.81&lt;br /&gt;
** --bprange option is added to allow testing for small region size&lt;br /&gt;
** Additional minor bug fixes&lt;br /&gt;
* December 4th, 2012 : EPACTS v3.1.0 release&lt;br /&gt;
** Removed dependency on libR.so&lt;br /&gt;
** Additional minor bug fixes&lt;br /&gt;
** --bprange option is added to allow testing for small region size&lt;br /&gt;
** November 25th, 2012 : EPACTS v3.0.0 release&lt;br /&gt;
** Restructured with source code release (with autoconf / automake / libtools)&lt;br /&gt;
** Added zoom plot feature&lt;br /&gt;
** FRAC_BURDEN keyword was replace to FRAC_WITH_RARE for groupwise testing&lt;br /&gt;
* October 26th, 2012 : EPACTS v2.2.0-beta is released with the following updates&lt;br /&gt;
** Added --max-mac option&lt;br /&gt;
** Fixed Firth&#039;s bias-corrected test (by Clement Ma)&lt;br /&gt;
** Added more informative warning messages when index files do not exist&lt;br /&gt;
** Fixed the bug in the epacts-plot in plotting ties&lt;br /&gt;
** Fixed errors in the MAF estimates per case and control&lt;br /&gt;
** Fixed bug in --minRSQ option&lt;br /&gt;
* September 28, 2012 : EPACTS v2.11-beta is released with the following updates&lt;br /&gt;
** Counts and allele frequencies for case/control added for binary tests&lt;br /&gt;
** --max-maf parameter is added&lt;br /&gt;
** Fixed EMMAX error in MAF in the output&lt;br /&gt;
** More informative error messages &lt;br /&gt;
* September 27, 2012 : EPACTS v2.1-beta is released with the following updates&lt;br /&gt;
** EMMAX interface is changed. --kinOnly option is related with a new command &#039;&#039;&#039;make-kin&#039;&#039;&#039; &lt;br /&gt;
** SKAT-O is upgraded to version 0.77 with additional configurable parameter settings&lt;br /&gt;
** Some parameter names are renamed (e.g. --min-maf, --min-mac)&lt;br /&gt;
** Many minor bugs are fixed&lt;br /&gt;
* Jul 6, 2012 : EPACTS v2.01-beta is released with the following updates&lt;br /&gt;
** SKAT-O is upgraded to version 0.76&lt;br /&gt;
** Fixed minor bugs in option names (Thanks to Xueling Sim)&lt;br /&gt;
* Jul 3, 2012 : EPACTS v2.0-beta is released with the following updates&lt;br /&gt;
** Major restructuring of the software&lt;br /&gt;
** Annotation software is switched with built-in application&lt;br /&gt;
** Addition of SKAT-O and EMMAX burden test&lt;br /&gt;
** Minor bug fixes&lt;br /&gt;
* Apr 8, 2012 : EPACTS v1.2-alpha is released with the following updates, in addition to the following updates&lt;br /&gt;
** EMMAX bug in handling covariates was fixed&lt;br /&gt;
** Variable Threshold Test is added&lt;br /&gt;
** Variable Threshold Test with genomic score (e.g. GERP or PhyloP) is added.&lt;br /&gt;
* Apr 4, 2012 : EPACTS v1.1-alpha is released with the following updates, in addition to minor updates&lt;br /&gt;
** EMMAX burden test (Hyun Min Kang)&lt;br /&gt;
** Likelihood ratio test (Clement Ma)&lt;br /&gt;
** Updated version of Firth bias-corrected likelihood ratio test (Clement Ma)&lt;br /&gt;
** Updated version of EMMAX single variant test (Hyun Min Kang) &lt;br /&gt;
* Mar 29, 2012 : EPACTS v1.0-alpha is released&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=VerifyBamID&amp;diff=7717</id>
		<title>VerifyBamID</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=VerifyBamID&amp;diff=7717"/>
		<updated>2013-08-02T14:40:03Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Command Line Options */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software|VerifyBamID]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;verifyBamID&#039;&#039;&#039; is a software that verifies whether the reads in particular file match previously known genotypes for an individual (or group of individuals), and checks whether the reads are contaminated as a mixture of two samples. &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; can detect sample contamination and swaps when external genotypes are available. When external genotypes are not available, &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; still robustly detects sample swaps.&lt;br /&gt;
&lt;br /&gt;
== Download verifyBamID  ==&lt;br /&gt;
&lt;br /&gt;
To get a copy go to the [http://www.sph.umich.edu/csg/kang/verifyBamID/download VerifyBamID Download] download page.&lt;br /&gt;
&lt;br /&gt;
== Join in verifyBamID mailing list ==&lt;br /&gt;
&lt;br /&gt;
Please join in the [http://groups.google.com/group/verifybamid VerifyBamID Google Group] to ask / discuss / comment about verifyBamID.&lt;br /&gt;
&lt;br /&gt;
== What&#039;s new ==&lt;br /&gt;
&lt;br /&gt;
(2012/06/20) &lt;br /&gt;
* Fixed a bug of incorrect estimate of contamination when --chip-full option was used (Thanks to Richard Smith)&lt;br /&gt;
* Fixed a bug of incorrect per-readgroup output in --chip-* parameter&lt;br /&gt;
&lt;br /&gt;
(2012/05/24) &lt;br /&gt;
* Fixed a bug of incorrect per-readgroup output (Thanks to Matthew Flickinger)&lt;br /&gt;
* &#039;&#039;&#039;(IMPORTANT)&#039;&#039;&#039; Add an option to remove either side of overlapping fragment. This option is turned on by default, and can be turned off usig --ignoreOverlapPair. If your sequence data has very short insert size, this update may increase the sensitivity of estimated contamination.&lt;br /&gt;
* Changes in the directory structure and Makefile&lt;br /&gt;
&lt;br /&gt;
(2012/05/18) The new release of verifyBamID have undergone major change since the last version (as of 2011 April). Here are the highlights&lt;br /&gt;
* The genotype / allele frequency file is now based on VCF format rather than PLINK format.&lt;br /&gt;
* The reference sequence information is no longer required&lt;br /&gt;
* Uses Brent&#039;s method for precise estimation of contamination parameters&lt;br /&gt;
* Generate the depth distribution statistics.&lt;br /&gt;
* Estimated reference-bias parameters (useful mostly for ABI SOLiD sequence data)&lt;br /&gt;
&lt;br /&gt;
== Build verifyBamID  ==&lt;br /&gt;
&lt;br /&gt;
The binary download of verifyBamID is available. You may use the version in Ubuntu 64-bit platform. To build verifyBamID, download the statgen library and run the following series of commands&lt;br /&gt;
 tar xzvf verifyBamID.20120620.tar.gz&lt;br /&gt;
 cd verifyBamID&lt;br /&gt;
 make cloneLib (in the case ../libStatGen does not exist)&lt;br /&gt;
 make&lt;br /&gt;
 ./bin/verifyBamID&lt;br /&gt;
&lt;br /&gt;
Note that &#039;&#039;&#039;git clone&#039;&#039;&#039; command will create a directory ./libStatGen under your working directory, and &#039;&#039;&#039;make&#039;&#039;&#039; will create binary of verifyBamID under verifyBamID/bin/&lt;br /&gt;
&lt;br /&gt;
verifyBamID is designed to be reasonably portable. &lt;br /&gt;
&lt;br /&gt;
However, since development occurs only on Ubuntu 9.10 x86 and x64 platforms, and later, there are likely other portability issues. &lt;br /&gt;
&lt;br /&gt;
Currently we support verifyBamID only on Ubuntu 9.10 and later on 64-bit processors.&lt;br /&gt;
&lt;br /&gt;
== Basic Usage ==&lt;br /&gt;
&lt;br /&gt;
A key step in any genetic analysis is to verify whether data being generated matches expectations. &#039;&#039;verifyBamID&#039;&#039; checks whether reads in a BAM file match previous genotypes for a specific sample. In addition, it detects possible sample mixture from population allele frequency only, which can be particularly useful when the genotype data is not available.&lt;br /&gt;
&lt;br /&gt;
Using a mathematical model that relates observed sequence reads to an hypothetical true genotype, &#039;&#039;verifyBamID&#039;&#039; tries to decide whether sequence reads match a particular individual or are more likely to be contaminated (including a small proportion of foreign DNA), derived from a closely related individual, or derived from a completely different individual.&lt;br /&gt;
&lt;br /&gt;
== Basic Usage Example ==&lt;br /&gt;
&lt;br /&gt;
Here is a typical command line:&lt;br /&gt;
&lt;br /&gt;
 verifyBamID --vcf [input.vcf] --bam [input.bam] --out [output.prefix] --verbose --ignoreRG&lt;br /&gt;
 &lt;br /&gt;
 where&lt;br /&gt;
 [input.bam] is a BAM (Binary Alignment Map) file of a sequence reads&lt;br /&gt;
 [input.vcf] is input VCF file containing individual genotypes or AF or AC/AN fields in the INFO field. gzipped VCF is also allowed.&lt;br /&gt;
 [outPrefix] is output prefix of output files - [outPrefix].{selfRG,selfSM,bestRG,bestSM,depthRG,depthSM} will be created.&lt;br /&gt;
&lt;br /&gt;
More detailed description of command line input is below&lt;br /&gt;
&lt;br /&gt;
== Preparing input files ==&lt;br /&gt;
&lt;br /&gt;
verifyBamID requires two input files - VCF file containing external genotypes or allele frequency information, and the BAM file.&lt;br /&gt;
&lt;br /&gt;
=== VCF input genotype file ===&lt;br /&gt;
&lt;br /&gt;
The input VCF file contains (1) external genotype information and/or (2) allele frequency information as AF entry or AC/AN entries in the INFO field. (See [http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41 | VCF specification] for further details). If neither information is provided, verifyBamID will not work properly.&lt;br /&gt;
&lt;br /&gt;
If external genotype information is provided, sequence+array method will identify contamination and sample swaps by comparing the concordance between the external genotypes and the sequence reads. Additionally, sequence-only method will provide additional contamination estimates by modeling the sequence reads as mixture of two unknown samples based on the allele frequency information in the VCF file.&lt;br /&gt;
&lt;br /&gt;
Input VCF file needs to meet several additional contraints need to meet in order to properly run verifyBamID.&lt;br /&gt;
* The VCF is assumed to be well-formed. For example, verifyBamID does not check whether REF allele actually matches with reference sequence.  &lt;br /&gt;
* The VCF should only contain SNPs. Current version of verifyBamID does not accept INDELs, MNPs, Structural Variations, or other complex variants.&lt;br /&gt;
* The individual IDs in the VCF file, must be identical with the individual identifier in the BAM file. Otherwise, --smID option can override the sample ID information of the BAM file to the ID that matches to the individual IDs in the VCF file.&lt;br /&gt;
* IMPORTANT : For targeted sequencing data, it is important to subselect the markers to only include on-target markers in the genotype file. Off-target markers are not likely to have multiple non-duplicated reads at the marker position, and it may create artifacts in the analysis due to overlapping fragments.&lt;br /&gt;
* Currently, verifyBamID takes only autosomal chromosomes as input VCF.&lt;br /&gt;
&lt;br /&gt;
An example input VCF file (without external genotype) is provided below. Note that AC and AC entries exists in the INFO field for the allele frequency information.&lt;br /&gt;
&lt;br /&gt;
 #CHROM	POS	ID	REF	ALT	QUAL	FILTER	INFO&lt;br /&gt;
 20	61651	SNP20-9651	C	A	.	PASS	CR=99.86851;GentrainScore=0.7055;HW=0.077647716;AN=2180;AC=11&lt;br /&gt;
 20	63231	SNP20-11231	T	G	.	PASS	CR=99.93036;GentrainScore=0.7837;HW=0.035481825;AN=2182;AC=275&lt;br /&gt;
 20	63244	rs6139074	A	C	.	PASS	CR=98.893394;GentrainScore=0.8001;HW=7.327299E-7;AN=2162;AC=501&lt;br /&gt;
 20	63799	rs1418258	C	T	.	PASS	CR=99.75217;GentrainScore=0.8170;HW=0.6653377;AN=2182;AC=881&lt;br /&gt;
&lt;br /&gt;
=== Input BAM file ===&lt;br /&gt;
&lt;br /&gt;
verifyBamID requires a sorted, indexed, base quality recalibrated, and duplication-marked BAM file. It also requires to contain &amp;quot;@RG&amp;quot; header lines to annotation different readGroups (sequencing runs and lanes). The SM tag in the &amp;quot;@RG&amp;quot; header should match with one of the genotyped sample. Otherwise, verifyBamID may not be able to test whether the sequenced sample matches with genotyped sample, but will try to detect sample mixture from allele frequency, and will try to detect the best-matching sample among the genotyped sample.&lt;br /&gt;
&lt;br /&gt;
== What the default option does ==&lt;br /&gt;
&lt;br /&gt;
The default option of &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; is the recommended setting for the most sequencing studies to provide a rapid and informative response. The default option provides the following features:&lt;br /&gt;
* --free-mix is turned on for estimating contamination using sequence-only method&lt;br /&gt;
* --chip-mix is turned on for estimating contamination or swap using sequence+array method, if the external genotype file is provided in the VCF&lt;br /&gt;
* --self is turnd on : The default option does not try to compare the sequence reads to identify the best matching individual (which is possible with --best option). It only compares with the external genotypes from the same individual to the sequenced individual.&lt;br /&gt;
* --maxDepth 20 is used without --precise option : The default option is intended for whole genome low coverage sequencing. For the targeted exome sequencing, --maxDepth 1000 and --precise is recommended.&lt;br /&gt;
* --ignoreRG is not a default option, but a recommended option, when you want to check the contamination for the entire BAM rather than examining each read group separately. This option will increase the computational efficiency especially in the case whether the sequence reads are multiplexed across many sequencing runs.&lt;br /&gt;
&lt;br /&gt;
== Interpreting output files ==&lt;br /&gt;
&lt;br /&gt;
See also [[Understanding VerifyBamID output]].&lt;br /&gt;
&lt;br /&gt;
=== Output files ===&lt;br /&gt;
When verifyBamID runs successfully, the following sets of files may be generated.&lt;br /&gt;
* [outPrefix].selfSM - Per-sample statistics describing how well the sample matches to the annotated sample.&lt;br /&gt;
* [outPrefix].depthSM - The depth distribution of the sequence reads per sample&lt;br /&gt;
* [outPrefix].selfRG - Per-readGroup statistics describing how well each lane matches to the annotated sample. (available only without --ignoreRG option)&lt;br /&gt;
* [outPrefix].depthRG - The depth distribution of the sequence reads per readGroup. (available only without --ignoreRG option)&lt;br /&gt;
* [outPrefix].bestSM - Per-sample best-match statistics with best-matching sample among the genotyped sample (available only with --best option)&lt;br /&gt;
* [outPrefix].bestRG - Per-readgroup best-match statistics with best-matching sample among the genotyped sample (available only with --best and without --ignoreRG option)&lt;br /&gt;
&lt;br /&gt;
=== Column information in the output files ===&lt;br /&gt;
The .selfSM/.selfRG/.bestSM/.bestRG files have the following 19 columns per sample, or per readgroup (lane). &lt;br /&gt;
&lt;br /&gt;
# SEQ_SM : Sample ID of the sequenced sample. Obtained from @RG header / SM tag in the BAM file&lt;br /&gt;
# RG : ReadGroup ID of sequenced lane. For [outPrefix].selfSM and [outPrefix].bestSM, these values are &amp;quot;ALL&amp;quot;&lt;br /&gt;
# CHIP_ID : Sample ID compared to in the genotype file. For [outPrefix].selfRG and [outPrefix].selfSM, these values should be identical to [SEQ_SM] or &amp;quot;NA&amp;quot; if the genotype of sequenced samples are unavailable. For [outPrefix].bestRG and [outPrefix].bestSM, these values should be the ID of best-matching sample among the genotype files compared to.&lt;br /&gt;
# # SNPs : # of SNPs passing the criteria from the VCF file&lt;br /&gt;
# # READS : Total # of reads loaded from the BAM file&lt;br /&gt;
# # AVG_DP : Average sequencing depth at the sites in the VCF file&lt;br /&gt;
# FREEMIX : Sequence-only estimate of contamination (0-1 scale)&lt;br /&gt;
# FREELK1 : Maximum log-likelihood of the sequence reads given estimated contamination under sequence-only method&lt;br /&gt;
# FREELK0 : Log-ikelihood of the sequence reads given no contamination under sequence-only method&lt;br /&gt;
# FREE_RH : Estimated reference bias parameter Pr(refBase|HET) (when --free-refBias or --free-full is used)&lt;br /&gt;
# FREE_RA : Estimated reference bias parameter Pr(refBase|HOMALT) (when --free-refBias or --free-full is used)&lt;br /&gt;
# CHIPMIX : Sequence+array estimate of contamination (NA if the external genotype is unavailable) (0-1 scale)&lt;br /&gt;
# CHIPLK1 : Maximum log-likelihood of the sequence reads given estimated contamination under sequence+array method (NA if the external genotypes are unavailable)&lt;br /&gt;
# CHIPLK0 : Log-likelihood of the sequence reads given no contamination under sequence+array method (NA if the external genotypes are unavailable)&lt;br /&gt;
# CHIP_RH : Estimated reference bias parameter Pr(refBase|HET) (when --chip-refBias or --chip-full is used)&lt;br /&gt;
# CHIP_RA : Estimated reference bias parameter Pr(refBase|HOMALT) (when --chip-refBias or --chip-full is used)&lt;br /&gt;
# DPREF : Depth (Coverage) of HomRef site (based on the genotypes of (SELF_SM/BEST_SM), passing mapQ, baseQual, maxDepth thresholds.&lt;br /&gt;
# RDPHET : DPHET/DPREF, Relative depth at Heterozygous site.&lt;br /&gt;
# RDPALT : DPHET/DPREF, Relative depth at HomAlt site.&lt;br /&gt;
&lt;br /&gt;
=== A guideline to interpret output files ===&lt;br /&gt;
&lt;br /&gt;
verifyBamID provides a series of information that is informative to determine whether the sample is possibly contaminated or swapped, but there is no single criteria that works for every circumstances. There are a few unmodeled factor in the estimation of [SELF-IBD]/[BEST-IBD] and [%MIX], so please note that the MLE estimation may not always exactly match to the true amount of contamination. Here we provide a guideline to flag potentially contaminated/swapped samples &lt;br /&gt;
&lt;br /&gt;
*  Each sample or lane can be checked in this way. When [CHIPMIX] &amp;gt;&amp;gt; 0.02 and/or [FREEMIX] &amp;gt;&amp;gt; 0.02, meaning 2% or more of non-reference bases are observed in reference sites, we recommend to examine the data more carefully for the possibility of contamination.&lt;br /&gt;
* We recommend to check each lane for the possibility of sample swaps. When [CHIPMIX] ~ 1 AND [FREEMIX] ~ 0, then it is possible that the sample is swapped with another sample. When [CHIPMIX] ~ 0 in .bestSM file, [CHIP_ID] might be actually the swapped sample. Otherwise, the swapped sample may not exist in the genotype data you have compared. &lt;br /&gt;
* When genotype data is not available but allele-frequency-based estimates of [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, then it is possible that the sample is contaminated with other sample. We recommend to use per-sample data rather than per-lane data for checking this for low coverage data, because the inference will be more confident when there are large number of bases with depth 2 or higher.&lt;br /&gt;
&lt;br /&gt;
== Command Line Options ==&lt;br /&gt;
&lt;br /&gt;
 The following parameters are available.  Ones with &amp;quot;[]&amp;quot; are in effect:&lt;br /&gt;
 &lt;br /&gt;
 Available Options&lt;br /&gt;
                             Input Files : --vcf [], --bam [], --subset [],&lt;br /&gt;
                                           --smID []&lt;br /&gt;
                    VCF analysis options : --genoError [1.0e-03],&lt;br /&gt;
                                           --minAF [0.01],&lt;br /&gt;
                                           --minCallRate [0.50]&lt;br /&gt;
   Individuals to compare with chip data : --site, --self, --best&lt;br /&gt;
          Chip-free optimization options : --free-none, --free-mix [ON],&lt;br /&gt;
                                           --free-refBias, --free-full&lt;br /&gt;
          With-chip optimization options : --chip-none, --chip-mix [ON],&lt;br /&gt;
                                           --chip-refBias, --chip-full&lt;br /&gt;
                    BAM analysis options : --ignoreRG, --ignoreOverlapPair,&lt;br /&gt;
                                           --noEOF, --precise, --minMapQ [10],&lt;br /&gt;
                                           --maxDepth [20], --minQ [13],&lt;br /&gt;
                                           --maxQ [40], --grid [0.05]&lt;br /&gt;
                 Modeling Reference Bias : --refRef [1.00], --refHet [0.50],&lt;br /&gt;
                                           --refAlt [0.00]&lt;br /&gt;
                          Output options : --out [], --verbose&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Each option provides the following features:&lt;br /&gt;
* --vcf : specify required VCF file&lt;br /&gt;
* --bam : specify required BAM file (indexed with .bam.bai or .bai file)&lt;br /&gt;
* --subset : list of individual IDs to calculate the allele frequency. All individuals are used if unspecified&lt;br /&gt;
* --smID : If the individual ID in the BAM file and VCF file does not match, substitute the BAM file&#039;s ID into the specified argument&lt;br /&gt;
* --genoError : error rate of the external genotype file&lt;br /&gt;
* --minAF : minimum allele frequency of the markers to include&lt;br /&gt;
* --minAF : minimum call rate of the markers to include&lt;br /&gt;
* --site : If set, use only site information in the VCF and do not compare with the actual genotypes&lt;br /&gt;
* --self : Only compare the ID-matching individuals between the VCF and BAM file&lt;br /&gt;
* --best : Find the best matching individuals (.bestSM and .bestRG files will be produced). This option is substantially longer than the default option&lt;br /&gt;
* --free-none : Do not perform sequence-only method to estimate parameters&lt;br /&gt;
* --free-mix : (default) Estimate contamination using sequence-only method with Brent&#039;s single dimensional optimization.&lt;br /&gt;
* --free-refBias : Estimate the reference bias parameters using sequence-only method with Simplex method&lt;br /&gt;
* --free-full : Estimate both reference bias parameters and the contamination parameters using sequence-only method&lt;br /&gt;
* --chip-none : Do not perform sequence+array method to estimate parameters&lt;br /&gt;
* --free-mix : (default) Estimate contamination using sequence+array method with Brent&#039;s single dimensional optimization.&lt;br /&gt;
* --free-refBias : Estimate the refernece bias parameters using sequence+array method with Simplex method&lt;br /&gt;
* --free-full : Estimate both reference bias parameters and the contamination parameters using sequence+array method&lt;br /&gt;
* --ignoreRG : ignore the read grouup level comparison and compare samples only (recommended for an expedited run)&lt;br /&gt;
* --ignoreOverlapPair : ignore overlapping pair end fragment covering the same base. Disabling this option may decrease the sensitivity of the method when the insert size is short (with slight gain in the computational speed)&lt;br /&gt;
* --noEOF : do not check the EOF marker of the BAM file (for earlier version of BAM)&lt;br /&gt;
* --precise : calculate the likelihood in log-scale for high-depth data (recommended when --maxDepth is greater than 20. Can be a little bit slower)&lt;br /&gt;
* --minMapQ : minimum mapping quality of the sequence reads to compare&lt;br /&gt;
* --minQ : minimum base quality to include&lt;br /&gt;
* --maxQ : maximum base quality to cap&lt;br /&gt;
* --grid : the grid interval to search the optimum before running Brent&#039;s algorithm.&lt;br /&gt;
* --refRef : Initial Pr(refBase|HOMREFGeno) parameter&lt;br /&gt;
* --refHet : Initial Pr(refBase|HETGeno) parameter&lt;br /&gt;
* --refAlt : Initial Pr(refBase|HOMALTGeno) parameter&lt;br /&gt;
* --out : output file prefix (required)&lt;br /&gt;
* --verbose : print the progress of the method on the screeen&lt;br /&gt;
&lt;br /&gt;
== Principle of Operation ==&lt;br /&gt;
&lt;br /&gt;
Each read group in a BAM file is evaluated independently. This means that in file with multiple read groups, problems will be flagged at the read group level (a plus). However, it also means that it might be hard to discern the correct assignment of read groups with very little data.&lt;br /&gt;
&lt;br /&gt;
For each aligned base that overlaps a known genotype, we calculate the probability the probability that it was derived from a particular known genotype. This comparison considers only bases that overlap previously known genotypes and that meet the base quality and mapping quality thresholds.&lt;br /&gt;
&lt;br /&gt;
Each individual in a pedigree has a different combination of genotypes, and bamGenotypeCheck will systematically search for the individual whose genotypes best match the observed read data.&lt;br /&gt;
&lt;br /&gt;
For more about the technical details, see the page [[Verifying Sample Identities - Implementation]]&lt;br /&gt;
&lt;br /&gt;
== Reference ==&lt;br /&gt;
&lt;br /&gt;
Please cite the following paper:&lt;br /&gt;
&lt;br /&gt;
G. Jun, M. Flickinger, K. N. Hetrick, Kurt, J. M. Romm, K. F. Doheny, G. Abecasis, M. Boehnke,and H. M. Kang, &#039;&#039;Detecting and Estimating Contamination of Human DNA Samples in Sequencing and Array-Based Genotype Data&#039;&#039;, American journal of human genetics doi:10.1016/j.ajhg.2012.09.004 (volume 91 issue 5 pp.839 - 848) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Contamination in Array Data ==&lt;br /&gt;
&lt;br /&gt;
[[VerifyIDintensity]] or [[BAFRegress]] can estimate sample contamination from Illumina genotype array data.&lt;br /&gt;
&lt;br /&gt;
== Acknowledgements ==&lt;br /&gt;
&lt;br /&gt;
VerifyBamID is a result from collaborative effort by Hyun Min Kang, Goo Jun, Matthew Flickinger, Mary Kate Wing, and Goncalo Abecasis. Please email to Hyun Min Kang [[mailto:hmkang@umich.edu| hmkang@umich.edu ]] for any questions.&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Variant_Calling_Pipeline&amp;diff=7573</id>
		<title>GotCloud: Variant Calling Pipeline</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Variant_Calling_Pipeline&amp;diff=7573"/>
		<updated>2013-07-01T14:41:03Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Results */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
The Variant Calling Pipeline (previously called &#039;UMAKE&#039;) makes genotype calls from recalibrated BAM files. These genotype calls are output into [http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41 VCF (Variant Call Format) files].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Running the GotCloud Variant Calling Pipeline =&lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline (umake) is run using &amp;lt;code&amp;gt;gotcloud snpcall&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;gotcloud ldrefine&amp;lt;/code&amp;gt;.  &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; is found under &amp;lt;code&amp;gt;gotcloud/&amp;lt;/code&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
==Running the Automatic Test==&lt;br /&gt;
&lt;br /&gt;
The automatic test runs the variant calling pipeline on a small testset and checks the results against expected results validating that GotCloud is installed correctly.&lt;br /&gt;
&lt;br /&gt;
*Run variant calling pipeline test:&lt;br /&gt;
 gotcloud snpcall --test OUTPUT_DIR&lt;br /&gt;
where OUTPUT_DIR is the directory where you want to store the test results&lt;br /&gt;
&lt;br /&gt;
If you see &amp;quot;Successfully ran the test case, congratulations!&amp;quot;, then you are ready to align samples.&lt;br /&gt;
&lt;br /&gt;
= Overview of Variant Calling Pipeline Steps =&lt;br /&gt;
Here is an overview of the Variant Calling Pipeline:&lt;br /&gt;
&lt;br /&gt;
[[File: umakeSteps.png]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Input Data=&lt;br /&gt;
*Aligned/Processed/Recalibrated BAM files&lt;br /&gt;
*Index file containing Sample IDs &amp;amp; BAM file names&lt;br /&gt;
*Reference files&lt;br /&gt;
*(Optional) Configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
== BAM files ==&lt;br /&gt;
The BAM files need to be duplicate-marked and base-quality recalibrated in order to obtain high quality SNP calls. Generating these BAM files from original FASTQs is documented elsewhere as part of the [[Alignment Pipeline]] of gotCloud.&lt;br /&gt;
&lt;br /&gt;
== Index File ==&lt;br /&gt;
Each line of the index file represents each individual under the following format. Note that multiple BAMs per individual may be provided. Note that if all samples are from the same population, just specify &amp;quot;ALL&amp;quot; for the population label for each sample.&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if applicable)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N&lt;br /&gt;
&lt;br /&gt;
== Reference Files ==&lt;br /&gt;
The variant calling pipeline requires multiple reference files in order to work correctly. &lt;br /&gt;
&lt;br /&gt;
* Reference Sequence in fasta format.&lt;br /&gt;
** Configuration File Setting:  &amp;lt;code&amp;gt;REF = path/file.fa&amp;lt;/code&amp;gt;&lt;br /&gt;
* Indel VCF File Prefix&lt;br /&gt;
** Configuration File Setting:  &amp;lt;code&amp;gt;INDEL_PREFIX = path/indels.sites.hg19&amp;lt;/code&amp;gt;&lt;br /&gt;
** &amp;lt;code&amp;gt;path/&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;indels.sites.hg19.chr20.vcf&amp;lt;/code&amp;gt; for each chromosome being processed&lt;br /&gt;
* DBSNP File vcf.gz file (must be indexed with tabix)&lt;br /&gt;
** Configuration File Setting:  &amp;lt;code&amp;gt;DBSNP_VCF = path/dbsnp_135.b37.vcf.gz&amp;lt;/code&amp;gt;&lt;br /&gt;
** &amp;lt;code&amp;gt;path/&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;dbsnp_135_b37.rod.chr20.map&amp;lt;/code&amp;gt; for each chromosome being processed&lt;br /&gt;
* HapMap3 polymorphic site vcf.gz file (must be indexed with tabix)&lt;br /&gt;
** Configuration File Setting:  &amp;lt;code&amp;gt;HM3_VCF = path/hapmap_3.3.b37.sites.vcf.gz&amp;lt;code&amp;gt;&lt;br /&gt;
** &amp;lt;code&amp;gt;path/&amp;lt;/code&amp;gt; contains &amp;lt;code&amp;gt;hapmap3.qc.poly.chr20.bim&amp;lt;/code&amp;gt; &amp;amp; &amp;lt;code&amp;gt;hapmap3.qc.poly.chr20.frq&amp;lt;/code&amp;gt; for each chromosome being processed&lt;br /&gt;
&lt;br /&gt;
A set of reference files can be downloaded from: [[ftp://share.sph.umich.edu/1000genomes/umake-resources/ | FTP Download of Full Resource Files]]&lt;br /&gt;
&lt;br /&gt;
Configuration File Example Reference Settings:&lt;br /&gt;
 REF = path/file.fa&lt;br /&gt;
 INDEL_PREFIX = path/indels.sites.hg19&lt;br /&gt;
 DBSNP_VCF = path/dbsnp_135_b37.vcf.gz&lt;br /&gt;
 HM3_VCF = path/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Configuration File ==&lt;br /&gt;
Configuration file contains the run-time options including the software binaries and command line arguments.  A default configuration file is automatically loaded.  Users must specify their own configuration file specifying just the values different than the defaults.&lt;br /&gt;
&lt;br /&gt;
Comments begin with a &amp;lt;code&amp;gt;#&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Format: KEY = value&lt;br /&gt;
&lt;br /&gt;
Where KEY is the item being set and value is its new value&lt;br /&gt;
&lt;br /&gt;
===Required User Config Files Settings===&lt;br /&gt;
The following Config File Settings must be specified by the user:&lt;br /&gt;
* CHRS = space separated list of chromosomes you want&lt;br /&gt;
* BAM_INDEX = path to the Index File of BAMs&lt;br /&gt;
&lt;br /&gt;
===Required on Command-Line or in Config File===&lt;br /&gt;
The following Command-Line or Config File Settings must be specified by the user:&lt;br /&gt;
* --outdir/OUT_DIR= path to desired output directory&lt;br /&gt;
&lt;br /&gt;
===Targeted/Exome Sequencing Settings===&lt;br /&gt;
If you are running Targeted/Exome Sequencing, the user should specify:&lt;br /&gt;
* Write loci file when performing pileup&lt;br /&gt;
** WRITE_TARGET_LOCI = TRUE&lt;br /&gt;
* Specify the output sub-directory to store target information, for example: targetDir&lt;br /&gt;
** Should not be a full path as this will co under the OUT_DIR directory.&lt;br /&gt;
** TARGET_DIR = targetDir&lt;br /&gt;
&lt;br /&gt;
If all individuals have the same target:&lt;br /&gt;
* Specify the single bed file, for example: target.bed&lt;br /&gt;
** UNIFORM_TARGET_BED = target.bed&lt;br /&gt;
&lt;br /&gt;
If not all individuals have the same target:&lt;br /&gt;
* Specify the file containing the sample id -&amp;gt; bed map, for example: targetMap.txt&lt;br /&gt;
** MULTIPLE_TARGET_MAP = targetMap.txt&lt;br /&gt;
*** Each line of the file contains [SM_ID] [TARGET_BED]&lt;br /&gt;
&lt;br /&gt;
Optional Settings:&lt;br /&gt;
* Extend the target region by a given number of bases, for example: 50&lt;br /&gt;
** OFFSET_OFF_TARGET = 50&lt;br /&gt;
&lt;br /&gt;
=== Configure Reference Files ===&lt;br /&gt;
See [[#Reference Files| Reference Files]] for information on how to specify the reference files.&lt;br /&gt;
&lt;br /&gt;
=== Chromosome X Calling ===&lt;br /&gt;
Making calls on the X chromosome requires the user to specifty a PED file with sex information.&lt;br /&gt;
* PED_INDEX = pedfile.ped&lt;br /&gt;
&lt;br /&gt;
== Example Configuration File ==&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam index file in path/freeze5&lt;br /&gt;
 CHRS = 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22&lt;br /&gt;
 BAM_INDEX = /path/freeze5/freeze5.bam.index  ### The BAM index file described above&lt;br /&gt;
 OUT_DIR = /path/freeze5/output               ### Directory in which to put all gotcloud output&lt;br /&gt;
 REF = /path/reference/hs37d5.fa              ### Reference sequence&lt;br /&gt;
 INDEL_PREFIX = /path/reference/1kg.pilot_release.merged.indels.sites.hg19   ### Known indel sites&lt;br /&gt;
 HM3_VCF = /path/reference/hapmap3_r3_b37.sites.vcf.gz    ### HapMap variants (requires tabix index file in same directory)&lt;br /&gt;
 DBSNP_VCF = /path/reference/dbsnp_135.b37.sites.vcf.gz   ### dbSNP variants (requires tabix index file in same directory)&lt;br /&gt;
&lt;br /&gt;
= Running =&lt;br /&gt;
&lt;br /&gt;
Running variant calling is straightforward:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;gotcloud snpcall --conf vc.conf --numjobs 2&lt;br /&gt;
 &#039;&#039;&#039;gotcloud ldrefine --conf vc.conf --numjobs 2&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Replace vc.conf with the approprate path/name of the user&#039;s configuration file.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;code&amp;gt;OUT_DIR&amp;lt;/code&amp;gt; is not defined in the configuration file, add &amp;lt;code&amp;gt;--outdir&amp;lt;/code&amp;gt; followed by the path to the user&#039;s desired output directory.&lt;br /&gt;
&lt;br /&gt;
Update the value following &amp;lt;code&amp;gt;--numjobs&amp;lt;/code&amp;gt; to the appropriate number of jobs to be run in parallel.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Running on a Cluster ==&lt;br /&gt;
To run on the Cluster, the following settings need to be added to the configuration file:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;--- Following may need revision ---&amp;gt;&lt;br /&gt;
TODO: COMING SOON&lt;br /&gt;
 SLEEP_MULT =     20&lt;br /&gt;
 REMOTE_PREFIX =  # REMOTE_PREFIX : Set if cluster node see the directory differently (e.g. /net/mymachine/[original-dir])&lt;br /&gt;
&amp;lt;--- End: Following may need revision ---&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Here&#039;s the same configuration file we used above but now made to run on a cluster computer with MOSIX.&lt;br /&gt;
 == Example Configuration File ==&lt;br /&gt;
 CHRS = 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22&lt;br /&gt;
 BAM_INDEX = /path/freeze5/freeze5.bam.index &lt;br /&gt;
 OUT_DIR = /path/freeze5/output              &lt;br /&gt;
 REF = /path/reference/hs37d5.fa             &lt;br /&gt;
 INDEL_PREFIX = /path/reference/1kg.pilot_release.merged.indels.sites.hg19&lt;br /&gt;
 HM3_VCF = /path/reference/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
 DBSNP_VCF = /path/reference/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 BATCH_TYPE = mosix             ### Specify MOSIX as the batch system&lt;br /&gt;
 BATCH_OPTS = -j10,11,12,13     ### Specify available MOSIX compute nodes&lt;br /&gt;
&lt;br /&gt;
= Results =&lt;br /&gt;
&lt;br /&gt;
If there is a failure, you should see a message like: &lt;br /&gt;
 make: *** [...] Error 1&lt;br /&gt;
Where ... is filled in with other text indicating what step failed.&lt;br /&gt;
&lt;br /&gt;
On SNP Call success, you should see the following output sub-directories under your output directory:&lt;br /&gt;
* glfs with a bams &amp;amp; samples subdirectory&lt;br /&gt;
* pvcfs with a subdirectory per chromosome and then per region&lt;br /&gt;
* split with a subdirectory per chromosome&lt;br /&gt;
* vcfs with a subdirectory per chromosome&lt;br /&gt;
* (optionally your target directory)&lt;br /&gt;
&lt;br /&gt;
Under the vcf/chrXX directory, there should be:&lt;br /&gt;
* chrXX.filtered.sites.vcf&lt;br /&gt;
* chrXX.filtered.sites.vcf.norm.log&lt;br /&gt;
* chrXX.filtered.sites.vcf.summary&lt;br /&gt;
* chrXX.filtered.vcf.gz&lt;br /&gt;
* chrXX.filtered.vcf.gz.OK&lt;br /&gt;
* chrXX.filtered.vcf.gz.tbi&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf.log&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf.summary&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz.OK&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz.tbi&lt;br /&gt;
* chrXX.merged.sites.vcf&lt;br /&gt;
* chrXX.merged.stats.vcf&lt;br /&gt;
* chrXX.merged.vcf&lt;br /&gt;
* chrXX.merged.vcf.OK&lt;br /&gt;
&lt;br /&gt;
The .merged.vcf is the merged together versions of the separate regions in the same chromosome.&lt;br /&gt;
&lt;br /&gt;
The filtered is the merged.vcf after it has been run through filters and is marked with PASS/FAIL.&lt;br /&gt;
&lt;br /&gt;
Under the split/chrXX directory, there should be:&lt;br /&gt;
* chrXX.filtered.PASS.split.[N].vcf.gz&lt;br /&gt;
* chrXX.filtered.PASS.split.err&lt;br /&gt;
* chrXX.filtered.PASS.split.vcflist&lt;br /&gt;
* chrXX.filtered.PASS.gz&lt;br /&gt;
* subset.OK&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7549</id>
		<title>Tutorial: GotCloud</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7549"/>
		<updated>2013-06-26T13:52:11Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* STEP 2 : Run GotCloud Alignment Pipeline */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= GotCloud Tutorial =&lt;br /&gt;
In this tutorial, we illustrate some of the essential steps in the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
For a background on GotCloud and Sequence Analysis Pipelines, see [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
While GotCloud can run on a cluster of machines or instances, this tutorial is just a small test that just runs on the machine the commands are run on.&lt;br /&gt;
&lt;br /&gt;
GotCloud and this basic tutorial were presented at the [http://ibg.colorado.edu/dokuwiki/doku.php?id=workshop:2013:announcement 2013 IBG Workshop].  It was presented in two sessions.  On Wednesday an overview was presented with steps for running the tutorial data: [[Media:IBG2013GotCloud.pdf|IBG2013GotCloud.pdf]].  On Friday more detail on the input files and what goes into generating the input files was presented: [[Media:GotCloudIBGWorkshop2013Friday.pdf|GotCloudIBGWorkshop2013Friday.pdf]].&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;This tutorial is in the process of being updated for gotcloud version 1.06 (April 17. 2013).&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== STEP 1 : Setup GotCloud ==&lt;br /&gt;
&lt;br /&gt;
[[GotCloud]] has been developed and tested on Linux Ubuntu 12.10 and 12.04.2 LTS but has not been tested on other Linux operating systems. It is not available for Windows. If you do not have your own set of machines to run on, GotCloud is also available for Ubuntu running on the Amazon Elastic Compute Cloud, see [[Amazon_Snapshot]] for more information.&lt;br /&gt;
&lt;br /&gt;
We will use 3 different directories for this tutorial:&lt;br /&gt;
# path to the directory where gotcloud is installed, default ~/gotcloud/&lt;br /&gt;
# path to the directory where the example data is installed, default ~/gotcloudExample&lt;br /&gt;
# path to your output directory, default ~/gotcloudTutorialOut/&lt;br /&gt;
&lt;br /&gt;
If the directories specified above do not reflect the directories you would like to use, replace their occurrances in the instructions below with the appropriate paths.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Step 1a: Install GotCloud ===&lt;br /&gt;
In order to run this tutorial, you need to make sure you have GotCloud installed on your system.  &lt;br /&gt;
&lt;br /&gt;
If you have root and would like to install gotcloud on your system, follow: [[GotCloud#Install_GotCloud_Software| root access installation instructions]]&lt;br /&gt;
&lt;br /&gt;
Otherwise, you can install it in your own directory:&lt;br /&gt;
# Change to the directory where you want gotcloud/ installed&lt;br /&gt;
# Download the gotcloud tar from the ftp site.&lt;br /&gt;
# Extract the tar&lt;br /&gt;
# Build (compile) the source&lt;br /&gt;
#* Note: as the source builds, many messages will scroll through your terminal.  You may even see some warnings.  These messages are normal and expected.  As long as the build does not end with an error, you have successfully built the source.&lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloud_latest.tgz  # Download&lt;br /&gt;
 tar xf gotcloud_latest.tgz     # Extracts into gotcloud/&lt;br /&gt;
 cd ~/gotcloud/src; make         # Build source&lt;br /&gt;
 &lt;br /&gt;
GotCloud requires the following tools to be installed.&lt;br /&gt;
You can run ~/gotcloud/scripts/check_requirements.sh&lt;br /&gt;
...TBD – put in required programs/tools.&lt;br /&gt;
* java (java-common default-jre on ubuntu)&lt;br /&gt;
* make (make on ubuntu)&lt;br /&gt;
* libssl (libssl0.9.8 on ubuntu)&lt;br /&gt;
* gcc 4.4 or newer&lt;br /&gt;
&lt;br /&gt;
=== Step 1b: Install Example Dataset ===&lt;br /&gt;
Our dataset consists of 60 individuals from Great Britain (GBR) sequenced by the 1000 Genomes Project. These individuals have been sequenced to an average depth of about 4x.&lt;br /&gt;
&lt;br /&gt;
To conserve time and disk-space, our analysis will focus on a small region on chromosome 20, 42900000 - 43200000. &lt;br /&gt;
&lt;br /&gt;
The tutorial will run the alignment pipeline on 2 of the individuals (HG00096, HG00100).  The fastqs used for this step are reduced to reads that fall into our target region.&lt;br /&gt;
&lt;br /&gt;
The tutorial will then used previously aligned/mapped reads for the full 60 individuals to generate a list of polymorphic sites and estimate accurate genotypes at each of these sites. &lt;br /&gt;
&lt;br /&gt;
The example dataset we&#039;ll be using is available at: ftp://share.sph.umich.edu/gotcloud/gotcloudExample.tgz &lt;br /&gt;
&lt;br /&gt;
# Change directory to where you want to install the Tutorial data &lt;br /&gt;
# Download the dataset tar from the ftp site &lt;br /&gt;
# Extract the tar &lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloudExample_latest.tgz  # Download &lt;br /&gt;
 tar xvf gotcloudExample_latest.tgz    # Extracts into gotcloudExample/&lt;br /&gt;
&lt;br /&gt;
== STEP 2 : Run GotCloud Alignment Pipeline == &lt;br /&gt;
The first step in processing next generation sequence data is mapping the reads to the reference genome, generating per sample BAM files. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Align the fastqs to the reference genome &lt;br /&gt;
#* handles both single &amp;amp; paired end &lt;br /&gt;
# Merge the results from multiple fastqs into 1 file per sample &lt;br /&gt;
# Mark Duplicate Reads are marked &lt;br /&gt;
# Recalibrate Base Qualities &lt;br /&gt;
&lt;br /&gt;
This processing results in 1 BAM file per sample. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline also includes Quality Control (QC) steps: &lt;br /&gt;
# Visualization of various quality measures (QPLOT) &lt;br /&gt;
# Screen for sample contamination &amp;amp; swap (VerifyBamID) &lt;br /&gt;
&lt;br /&gt;
Run the alignment pipeline (the example aligns 2 samples) : &lt;br /&gt;
 ~/gotcloud/gotcloud align --conf ~/gotcloudExample/[[#Alignment Configuration File|GBR2align.conf]] --outdir [[#Alignment Output Directory|~/gotcloudTutorialOut]] --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the alignment pipeline (about 1-3 minutes), you will see the following message: &lt;br /&gt;
 Processing finished in n secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The final BAM files produced by the alignment pipeline are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/bams&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* BAM (.bam) files - 1 per sample&lt;br /&gt;
** HG00096.recal.bam &lt;br /&gt;
** HG00100.recal.bam &lt;br /&gt;
* BAM index files (.bai) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.bai &lt;br /&gt;
** HG00100.recal.bam.bai &lt;br /&gt;
* Indicator files that the step completed successfully:&lt;br /&gt;
** HG00096.recal.bam.done &lt;br /&gt;
** HG00100.recal.bam.done &lt;br /&gt;
&lt;br /&gt;
The Quality Control (QC) files are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/QCFiles&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* VerifyBamID output files:&lt;br /&gt;
** HG00096.genoCheck.depthRG &lt;br /&gt;
** HG00096.genoCheck.depthSM &lt;br /&gt;
** HG00096.genoCheck.selfRG &lt;br /&gt;
** HG00096.genoCheck.selfSM &lt;br /&gt;
** HG00100.genoCheck.depthRG &lt;br /&gt;
** HG00100.genoCheck.depthSM &lt;br /&gt;
** HG00100.genoCheck.selfRG &lt;br /&gt;
** HG00100.genoCheck.selfSM &lt;br /&gt;
&lt;br /&gt;
* VerifyBamID step completion files – 1 per sample&lt;br /&gt;
** HG00096.genoCheck.done &lt;br /&gt;
** HG00100.genoCheck.done &lt;br /&gt;
&lt;br /&gt;
* QPLOT output files&lt;br /&gt;
** HG00096.qplot.R &lt;br /&gt;
** HG00096.qplot.stats &lt;br /&gt;
** HG00100.qplot.R &lt;br /&gt;
** HG00100.qplot.stats &lt;br /&gt;
&lt;br /&gt;
* QPLOT step completion files – 1 per sample&lt;br /&gt;
** HG00096.qplot.done &lt;br /&gt;
** HG00100.qplot.done &lt;br /&gt;
&lt;br /&gt;
For information on the VerifyBamID output, see: [[Understanding VerifyBamID output]] &lt;br /&gt;
&lt;br /&gt;
For information on the QPLOT output, see: [[Understanding QPLOT output]]&lt;br /&gt;
&lt;br /&gt;
== STEP 3 : Run GotCloud Variant Calling Pipeline == &lt;br /&gt;
The next step is to analyze BAM files by calling SNPs and generating a VCF file containing the variant calls. &lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Filter out reads with low mapping quality &lt;br /&gt;
# Per Base Alignment Quality Adjustment (BAQ) &lt;br /&gt;
# Resolve overlapping paired end reads &lt;br /&gt;
# Generate genotype likelihood files &lt;br /&gt;
# Perform variant calling &lt;br /&gt;
# Extract features from variant sites &lt;br /&gt;
# Perform variant filtering &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
To speed variant calling, each chromosome is broken up into smaller regions which are processed separately.  While initially split by sample, the per sample data gets merged and is processed together for each region.  These regions are later merged to result in a single Variant Call File (VCF) per chromosome.  For the tutorial all of the data falls within a single region.&lt;br /&gt;
&lt;br /&gt;
Run the variant calling pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud snpcall --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --region 20:42900000-43200000 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the variant calling pipeline (about 3-4 minutes), you will see the following message: &lt;br /&gt;
  Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
On SNP Call success, the VCF files of interest are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered*&lt;br /&gt;
&lt;br /&gt;
This gives you the following files:&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.vcf.gz &#039;&#039;&#039; - vcf for whole chromosome after it has been run through hardfilters and SVM filters and marked with PASS/FAIL including per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf.norm.log - log file&lt;br /&gt;
* chr20.filtered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
* chr20.filtered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
* chr20.filtered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
&lt;br /&gt;
Also in the ~/gotcloudTutorialOut/vcfs/chr20 directory are intermediate files:&lt;br /&gt;
* the whole chromosome variant calls prior to any filtering: &lt;br /&gt;
** chr20.merged.sites.vcf - without per sample genotypes&lt;br /&gt;
** chr20.merged.stats.vcf &lt;br /&gt;
** chr20.merged.vcf - including per sample genotypes&lt;br /&gt;
** chr20.merged.vcf.OK - indicator that the step completed successfully&lt;br /&gt;
* the hardfiltered (pre-svm filtered) variant calls:&lt;br /&gt;
** chr20.filtered.vcf.gz - vcf for whole chromosome after it has been run through hard filters&lt;br /&gt;
** chr20.hardfiltered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.log - log file&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
* 40000001.45000000 subdirectory contains the data for just that region.&lt;br /&gt;
&lt;br /&gt;
The ~/gotcloudTutorialOut/split/chr20 folder contains a VCF with just the sites that pass the filters.&lt;br /&gt;
 ls ~/gotcloudTutorialOut/split/chr20/&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.PASS.vcf.gz &#039;&#039;&#039; – vcf of just sites that pass all filters&lt;br /&gt;
* chr20.filtered.PASS.split.1.vcf.gz - intermediate file&lt;br /&gt;
* chr20.filtered.PASS.split.err - log file&lt;br /&gt;
* chr20.filtered.PASS.split.vcflist - list of intermediate files&lt;br /&gt;
* subset.OK &lt;br /&gt;
&lt;br /&gt;
In addition to the vcfs subdirectory, there are additional intermediate files/directories:&lt;br /&gt;
* glfs – holds genotype likelihood format [[GLF]] files split by chromosome, region, and sample&lt;br /&gt;
* pvcfs – holds intermediate vcf files split by chromosome and region&lt;br /&gt;
&lt;br /&gt;
Note: the tutorial does not produce a target directory, but if you run with targeted data, you may see that.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 4 : Run GotCloud Genotype Refinement Pipeline == &lt;br /&gt;
The next step is to perform genotype refinement using linkage disequilibrium information using [http://faculty.washington.edu/browning/beagle/beagle.html Beagle] &amp;amp; [[ThunderVCF]]. &lt;br /&gt;
&lt;br /&gt;
Run the LD-aware genotype refinement pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud ldrefine --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of this pipeline (about 10 minutes), you will see the following message: &lt;br /&gt;
 Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The output from the beagle step of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
The output from the thunderVcf (final step) of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 5 : Run GotCloud Association Analysis Pipeline (EPACTS) == &lt;br /&gt;
&lt;br /&gt;
We will assume that the EPACTS are installed in the following directory&lt;br /&gt;
 setenv EPACTS /path/to/epacts&lt;br /&gt;
(If you need to install EPACTS, please refer to the documentation at [[EPACTS#Installation_Details]])&lt;br /&gt;
&lt;br /&gt;
 $EPACTS/epacts single --vcf ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered.vcf.gz --ped ~/gotcloudExample/test.GBR60.ped \\&lt;br /&gt;
    --out ~/gotcloudTutorialOut/epacts --test q.linear --run 1 --top 1 --chr 20&lt;br /&gt;
&lt;br /&gt;
Upon successful run, you will see files starting with ~/gotcloudTutorialOut/epacts&lt;br /&gt;
 ls ~/gotcloudTutorialOut/epacts*&lt;br /&gt;
&lt;br /&gt;
To see the top associated variants, you can run&lt;br /&gt;
 less ~/gotcloudTutorialOut/epacts.epacts.top5000&lt;br /&gt;
&lt;br /&gt;
To see the locus-zoom like plot, you can type the following command (assuming GNU gnuplot 4.2 or higher version was installed)&lt;br /&gt;
 xpdf ~/gotcloudTutorialOut/epacs.zoom.20.42987877.pdf&lt;br /&gt;
&lt;br /&gt;
Click [[Media:EPACTS TEST.zoom.20.42987877.pdf | Exampe LocusZoom PDF]] to see the expected output pdf&lt;br /&gt;
&lt;br /&gt;
= Frequently Asked Questions (FAQs) =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;I ran the tutorial example successfully, how can I run it with my real sequence data?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Congratulations for your successful run of your [[GotCloud]] Tutorial. Please see [[#Tutorial Inputs]] section to prepare your own input files for your sequence data. You will need to specify the FASTQ files associated with its sample names as explained. In addition, you will need to download the full reference and resource file across whole genome (the Tutorial contains only chr20 portion to make it compact) See [[#Alignment Configuration File]] section for the detailed information. Also, please refer to the original documentation of [[GotCloud]] for more detailed guide on installation beyond the scope of the tutorial.&lt;br /&gt;
&lt;br /&gt;
= Input Files for GotCloud Tutorial = &lt;br /&gt;
&lt;br /&gt;
This section describes the input files needed for the GotCloud tutorial. You don&#039;t need to know this detail to run the tutorial, but if you&#039;re interested in understanding the structure of GotCloud pipeline and run with your own sample, this would be a good starting point&lt;br /&gt;
&lt;br /&gt;
== Alignment Pipeline == &lt;br /&gt;
=== List of Input Files needed for Alignment ===&lt;br /&gt;
The command-line inputs to the tutorial alignment pipeline are: &lt;br /&gt;
# [[#Alignment Configuration File|Configuration File (--conf)]] &lt;br /&gt;
#* Specifies the configuration file to use when running&lt;br /&gt;
# [[#Alignment Output Directory|Output Directory (--outdir)]] &lt;br /&gt;
#* Directory where the output should be placed.&lt;br /&gt;
&lt;br /&gt;
Additional information required to run the alignment pipeline:&lt;br /&gt;
# [[#Alignment FASTQ Index File|Index file of FASTQs]]&lt;br /&gt;
# [[#Alignment Reference Files|Alignment Reference Files]]&lt;br /&gt;
For the tutorial, these values are specified in the configuration file.&lt;br /&gt;
&lt;br /&gt;
=== Alignment Configuration File === &lt;br /&gt;
The configuration file contains KEY = VALUE settings that override defaults and set specific values for the given run. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt; &lt;br /&gt;
INDEX_FILE = GBR60fastq.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_DIR = chr20Ref &lt;br /&gt;
AS = NCBI37 &lt;br /&gt;
FA_REF = $(REF_DIR)/human_g1k_v37_chr20.fa &lt;br /&gt;
DBSNP_VCF =  $(REF_DIR)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF = $(REF_DIR)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&amp;lt;/pre&amp;gt; &lt;br /&gt;
&lt;br /&gt;
This configuration file sets: &lt;br /&gt;
* [[#Alignment FASTQ Index File|INDEX_FILE]] - file containing the fastqs to be processed as well as the read group information for these fastqs.&lt;br /&gt;
* Reference Information: see [[#Alignment Reference Files|Alignment Reference Files]] for more information&lt;br /&gt;
** AS - assembly value to put in the BAM &lt;br /&gt;
&lt;br /&gt;
The index file and chromosome 20 references used in this tutorial are included with the example data under the ~/gotcloudExample directory.  The tutorial uses chromosome 20 only references in order to speed the processing time.  &lt;br /&gt;
&lt;br /&gt;
When running with your own data, you will need to update the:&lt;br /&gt;
* Index File to contain the information for your own fastq files&lt;br /&gt;
** See [[#Alignment FASTQ Index File|Alignment FASTQ Index File]] for more information on the contents of the index file.&lt;br /&gt;
* The reference files to be whole genome references&lt;br /&gt;
** If you are just running chromosome 20, you can use the tutorial references&lt;br /&gt;
** Whole genome reference files can be downloaded from [[GotCloudReference]].&lt;br /&gt;
** See [[#Alignment Reference Files|Alignment Reference Files]] for more information.&lt;br /&gt;
&lt;br /&gt;
Note: It is recommended that you use absolute paths (full path names, like “/home/mktrost/gotcloudReference” rather than just “gotcloudReference”).  This example does not use absolute paths in order to be flexible to where the data is installed, but using relative paths requires it to be run from the correct directory.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
&lt;br /&gt;
Reference files are required for running both the alignment and variant calling pipelines.  &lt;br /&gt;
&lt;br /&gt;
The configuration keys for setting these are:&lt;br /&gt;
* FA_REF - Genome sequence reference files (needed for both pipelines)&lt;br /&gt;
* DBSNP_VCF – DBSNP site VCF file (needed for both pipelines)&lt;br /&gt;
* HM3_VCF - HAPMAP site VCF file (needed for both pipelines)&lt;br /&gt;
* INDEL_PREFIX - Indel sites file (need for variant calling pipeline)&lt;br /&gt;
&lt;br /&gt;
The tutorial configuration file is setup to point to the required chromosome 20 reference files which are included with the tutorial example data in ~/gotcloudExample/chr20Ref/. &lt;br /&gt;
&lt;br /&gt;
If you are running more than just chromosome 20, you will need whole genome reference files which can be downloaded from [[GotCloudReference]].&lt;br /&gt;
&lt;br /&gt;
The configuration settings for these files are setup in the default configuration so do not need to be specified.  You just need to set REF_DIR in your configuration file to the path where you installed your reference files.&lt;br /&gt;
&lt;br /&gt;
To learn more about the reference files that are required, see [[GotCloud: Reference Files]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Alignment Output Directory === &lt;br /&gt;
This setting tells the alignment pipeline where to write the output and intermediate files. &lt;br /&gt;
&lt;br /&gt;
The output directory will be created if it doesn&#039;t already exist and will contain the following Directories/files: &lt;br /&gt;
* bams - directory containing the final bams/bai files &lt;br /&gt;
** HG00096.OK - indicates that this sample completed alignment processing &lt;br /&gt;
** HG00100.OK - indicates that this sample completed alignment processing &lt;br /&gt;
* failLogs - directory containing logs from steps that failed &lt;br /&gt;
** this directory is only created if an error is detected&lt;br /&gt;
* Makefiles - directory containing the makefiles with commands for processing each sample &lt;br /&gt;
** biopipe_HG00096.Makefile – commands for processing sample HG00096&lt;br /&gt;
** biopipe_HG00100.Makefile – commands for processing sample HG00100&lt;br /&gt;
** biopipe_HG00096.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
** biopipe_HG00100.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
* QCFiles - directory containing the QC Results &lt;br /&gt;
** &lt;br /&gt;
* tmp - directory containing temporary alignment files &lt;br /&gt;
** bwa.sai.t – contains temporary files for the 1st step of bwa that generates sai files from fastq files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
**** HG0096&lt;br /&gt;
** alignment.bwa – contains temporary files for the 2nd step of bwa that generates BAM files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
** alignment.pol - contains temporary files for the polish bam step that cleans up BAM files&lt;br /&gt;
** alignment.dedup – contains temporary files for the deduping step&lt;br /&gt;
*** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
*** *.metrics – metrics files for the deduping step&lt;br /&gt;
** alignment.recal – contains temporary files for the deduping step&lt;br /&gt;
*** *.log – contains information about the recalibration step&lt;br /&gt;
*** *.qemp – contains the recalibration tables used for recalibrating each BAM file&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Alignment FASTQ Index File=== &lt;br /&gt;
&lt;br /&gt;
There are four fastq files in {ROOT_DIR}/test/align/fastq/Sample_1 and four fastq files in {ROOT_DIR}/test/align/fastq/Sample_2, both in paired-end format.  Normally, we would need to build an index file for these files. Conveniently, an index file (indexFile.txt) already exists for the automatic test samples.  It can be found in {ROOT_DIR}/test/align/, and contains the following information in tab-delimited format: &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME FASTQ1                           FASTQ2                           RGID   SAMPLE    LIBRARY CENTER PLATFORM &lt;br /&gt;
 Sample1    fastq/Sample_1/File1_R1.fastq.gz fastq/Sample_1/File1_R2.fastq.gz RGID1  SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample1    fastq/Sample_1/File2_R1.fastq.gz fastq/Sample_1/File2_R2.fastq.gz RGID1a SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File1_R1.fastq.gz fastq/Sample_2/File1_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File2_R1.fastq.gz fastq/Sample_2/File2_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
&lt;br /&gt;
If you are in the {ROOT_DIR}/test/align directory, you can use this file as-is.  If you prefer, you can create a new index file and change the MERGE_NAME, RGID, SAMPLE, LIBRARY, CENTER, or PLATFORM values. It is recommended that you do not modify existing files in {ROOT_DIR}/test/align. &lt;br /&gt;
&lt;br /&gt;
If you want to run this example from a different directory, make sure the FASTQ1 and FASTQ2 paths are correct.  That is, each of the FASTQ1 and FASTQ2 entry in the index file should look like the following: &lt;br /&gt;
&lt;br /&gt;
 {ROOT_DIR}/test/align/fastq/Sample_1/File1_R1.fastq.gz &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to run this example from a different directory, but do not want to edit the index file, you can create a relative path to the test fastq files so their path agrees with that listed in the index file: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/align/fastq fastq &lt;br /&gt;
&lt;br /&gt;
This will create a symbolic link to the test fastq directory from your current directory. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Mapping_Pipeline#Sequence_Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Analyzing a Sample== &lt;br /&gt;
&lt;br /&gt;
Using UMAKE, you can analyze BAM files by calling SNPs, and generate a VCF file containing the results.  Once again, we can analyze BAM files used in the automatic test.  For this example, we have 60 BAM files, which can be found in {ROOT_DIR}/test/umake/bams.  These contain sequence information for a targeted region in chromosome 20. &lt;br /&gt;
&lt;br /&gt;
In addition to the BAM files, you will need three files to run UMAKE: an index file, a configuration file, and a bed file (needed to analyze BAM files from targeted/exome sequencing). &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
===Index file=== &lt;br /&gt;
&lt;br /&gt;
First, you need a list of all the BAM files to be analyzed. Conveniently, the a test index file (umake_test.index) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
 NA12272 ALL     bams/NA12272.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 NA12004 ALL     bams/NA12004.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 ... &lt;br /&gt;
 NA12874 ALL     bams/NA12874.mapped.LS454.ssaha2.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
&lt;br /&gt;
You can use this file directly if you change your current directory to {ROOT_DIR}/test/umake/. &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to copy and use this index file to a different directory, you can create a symbolic link to the bams folder as follows: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/umake/bams bams &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===BED file=== &lt;br /&gt;
&lt;br /&gt;
This file contains a single line: &lt;br /&gt;
&lt;br /&gt;
 chr20   20000050        20300000 &lt;br /&gt;
&lt;br /&gt;
You can copy this to the current directory and use it as-is. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Targeted.2FExome_Sequencing_Settings|targeted/exome sequencing settings]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Configuration file=== &lt;br /&gt;
&lt;br /&gt;
A configuration file (umake_test.conf) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
CHRS = 20 &lt;br /&gt;
BAM_INDEX = GBR60bam.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_ROOT = chr20Ref &lt;br /&gt;
# &lt;br /&gt;
REF = $(REF_ROOT)/human_g1k_v37_chr20.fa &lt;br /&gt;
INDEL_PREFIX = $(REF_ROOT)/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
DBSNP_VCF =  $(REF_ROOT)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF =  $(REF_ROOT)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 CHRS = 20 &lt;br /&gt;
 TEST_ROOT = $(UMAKE_ROOT)/test/umake &lt;br /&gt;
 BAM_INDEX = $(TEST_ROOT)/umake_test.index &lt;br /&gt;
 OUT_PREFIX = umake_test &lt;br /&gt;
 REF_ROOT = $(TEST_ROOT)/ref &lt;br /&gt;
 # &lt;br /&gt;
 REF = $(REF_ROOT)/karma.ref/human.g1k.v37.chr20.fa &lt;br /&gt;
 INDEL_PREFIX = $(REF_ROOT)/indels/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
 DBSNP_PREFIX =  $(REF_ROOT)/dbSNP/dbsnp_135_b37.rod &lt;br /&gt;
 HM3_PREFIX =  $(REF_ROOT)/HapMap3/hapmap3_r3_b37_fwd.consensus.qc.poly &lt;br /&gt;
 # &lt;br /&gt;
 RUN_INDEX = TRUE        # create BAM index file &lt;br /&gt;
 RUN_PILEUP = TRUE       # create GLF file from BAM &lt;br /&gt;
 RUN_GLFMULTIPLES = TRUE # create unfiltered SNP calls &lt;br /&gt;
 RUN_VCFPILEUP = TRUE    # create PVCF files using vcfPileup and run infoCollector &lt;br /&gt;
 RUN_FILTER = TRUE       # filter SNPs using vcfCooker &lt;br /&gt;
 RUN_SPLIT = TRUE        # split SNPs into chunks for genotype refinement &lt;br /&gt;
 RUN_BEAGLE = FALSE  # BEAGLE - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 RUN_SUBSET = FALSE  # SUBSET FOR THUNDER - MAY BE SET WITH BEAGLE STEP TOGETHER &lt;br /&gt;
 RUN_THUNDER = FALSE # THUNDER - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 ############################################################################### &lt;br /&gt;
 WRITE_TARGET_LOCI = TRUE  # FOR TARGETED SEQUENCING ONLY -- Write loci file when performing pileup &lt;br /&gt;
 UNIFORM_TARGET_BED = $(TEST_ROOT)/umake_test.bed # Targeted sequencing : When all individuals has the same target. Otherwise, comment it out &lt;br /&gt;
 OFFSET_OFF_TARGET = 50 # Extend target by given # of bases &lt;br /&gt;
 MULTIPLE_TARGET_MAP =  # Target per individual : Each line contains [SM_ID] [TARGET_BED] &lt;br /&gt;
 TARGET_DIR = target    # Directory to store target information &lt;br /&gt;
 SAMTOOLS_VIEW_TARGET_ONLY = TRUE # When performing samtools view, exclude off-target regions (may make command line too long) &lt;br /&gt;
&lt;br /&gt;
If you are running this from a different directory, you will want to change some of the lines as follows: &lt;br /&gt;
&lt;br /&gt;
 BAM_INDEX = {CURRENT_DIR}/umake_test.index &lt;br /&gt;
 UNIFORM_TARGET_BED = {CURRENT_DIR}/umake_test.bed &lt;br /&gt;
&lt;br /&gt;
where {CURRENT_DIR} is the absolute path to the directory that contains the index and bed files.  &lt;br /&gt;
&lt;br /&gt;
An additional option can be added in the configuration file: &lt;br /&gt;
&lt;br /&gt;
 OUT_DIR = {OUT_DIR} &lt;br /&gt;
&lt;br /&gt;
where {OUT_DIR} is the name of directory in which you want the output to be stored.  If you do not specify this in the configuration file, you will need to add an extra parameter when you run UMAKE in the next step. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Configuration_File|the configuration file]], [[Variant_Calling_Pipeline_(UMAKE)#Reference_Files|reference files]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Further Information== &lt;br /&gt;
&lt;br /&gt;
[[Mapping_Pipeline|Mapping (Alignment) Pipeline]] &lt;br /&gt;
&lt;br /&gt;
[[Variant_Calling_Pipeline_(UMAKE)|Variant Calling Pipeline (UMAKE)]]&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7548</id>
		<title>Tutorial: GotCloud</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7548"/>
		<updated>2013-06-26T12:09:34Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* STEP 2 : Run GotCloud Alignment Pipeline */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= GotCloud Tutorial =&lt;br /&gt;
In this tutorial, we illustrate some of the essential steps in the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
For a background on GotCloud and Sequence Analysis Pipelines, see [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
While GotCloud can run on a cluster of machines or instances, this tutorial is just a small test that just runs on the machine the commands are run on.&lt;br /&gt;
&lt;br /&gt;
GotCloud and this basic tutorial were presented at the [http://ibg.colorado.edu/dokuwiki/doku.php?id=workshop:2013:announcement 2013 IBG Workshop].  It was presented in two sessions.  On Wednesday an overview was presented with steps for running the tutorial data: [[Media:IBG2013GotCloud.pdf|IBG2013GotCloud.pdf]].  On Friday more detail on the input files and what goes into generating the input files was presented: [[Media:GotCloudIBGWorkshop2013Friday.pdf|GotCloudIBGWorkshop2013Friday.pdf]].&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;This tutorial is in the process of being updated for gotcloud version 1.06 (April 17. 2013).&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== STEP 1 : Setup GotCloud ==&lt;br /&gt;
&lt;br /&gt;
[[GotCloud]] has been developed and tested on Linux Ubuntu 12.10 and 12.04.2 LTS but has not been tested on other Linux operating systems. It is not available for Windows. If you do not have your own set of machines to run on, GotCloud is also available for Ubuntu running on the Amazon Elastic Compute Cloud, see [[Amazon_Snapshot]] for more information.&lt;br /&gt;
&lt;br /&gt;
We will use 3 different directories for this tutorial:&lt;br /&gt;
# path to the directory where gotcloud is installed, default ~/gotcloud/&lt;br /&gt;
# path to the directory where the example data is installed, default ~/gotcloudExample&lt;br /&gt;
# path to your output directory, default ~/gotcloudTutorialOut/&lt;br /&gt;
&lt;br /&gt;
If the directories specified above do not reflect the directories you would like to use, replace their occurrances in the instructions below with the appropriate paths.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Step 1a: Install GotCloud ===&lt;br /&gt;
In order to run this tutorial, you need to make sure you have GotCloud installed on your system.  &lt;br /&gt;
&lt;br /&gt;
If you have root and would like to install gotcloud on your system, follow: [[GotCloud#Install_GotCloud_Software| root access installation instructions]]&lt;br /&gt;
&lt;br /&gt;
Otherwise, you can install it in your own directory:&lt;br /&gt;
# Change to the directory where you want gotcloud/ installed&lt;br /&gt;
# Download the gotcloud tar from the ftp site.&lt;br /&gt;
# Extract the tar&lt;br /&gt;
# Build (compile) the source&lt;br /&gt;
#* Note: as the source builds, many messages will scroll through your terminal.  You may even see some warnings.  These messages are normal and expected.  As long as the build does not end with an error, you have successfully built the source.&lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloud_latest.tgz  # Download&lt;br /&gt;
 tar xf gotcloud_latest.tgz     # Extracts into gotcloud/&lt;br /&gt;
 cd ~/gotcloud/src; make         # Build source&lt;br /&gt;
 &lt;br /&gt;
GotCloud requires the following tools to be installed.&lt;br /&gt;
You can run ~/gotcloud/scripts/check_requirements.sh&lt;br /&gt;
...TBD – put in required programs/tools.&lt;br /&gt;
* java (java-common default-jre on ubuntu)&lt;br /&gt;
* make (make on ubuntu)&lt;br /&gt;
* libssl (libssl0.9.8 on ubuntu)&lt;br /&gt;
* gcc 4.4 or newer&lt;br /&gt;
&lt;br /&gt;
=== Step 1b: Install Example Dataset ===&lt;br /&gt;
Our dataset consists of 60 individuals from Great Britain (GBR) sequenced by the 1000 Genomes Project. These individuals have been sequenced to an average depth of about 4x.&lt;br /&gt;
&lt;br /&gt;
To conserve time and disk-space, our analysis will focus on a small region on chromosome 20, 42900000 - 43200000. &lt;br /&gt;
&lt;br /&gt;
The tutorial will run the alignment pipeline on 2 of the individuals (HG00096, HG00100).  The fastqs used for this step are reduced to reads that fall into our target region.&lt;br /&gt;
&lt;br /&gt;
The tutorial will then used previously aligned/mapped reads for the full 60 individuals to generate a list of polymorphic sites and estimate accurate genotypes at each of these sites. &lt;br /&gt;
&lt;br /&gt;
The example dataset we&#039;ll be using is available at: ftp://share.sph.umich.edu/gotcloud/gotcloudExample.tgz &lt;br /&gt;
&lt;br /&gt;
# Change directory to where you want to install the Tutorial data &lt;br /&gt;
# Download the dataset tar from the ftp site &lt;br /&gt;
# Extract the tar &lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloudExample_latest.tgz  # Download &lt;br /&gt;
 tar xvf gotcloudExample_latest.tgz    # Extracts into gotcloudExample/&lt;br /&gt;
&lt;br /&gt;
== STEP 2 : Run GotCloud Alignment Pipeline == &lt;br /&gt;
The first step in processing next generation sequence data is mapping the reads to the reference genome, generating per sample BAM files. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Align the fastqs to the reference genome &lt;br /&gt;
#* handles both single &amp;amp; paired end &lt;br /&gt;
# Merge the results from multiple fastqs into 1 file per sample &lt;br /&gt;
# Mark Duplicate Reads are marked &lt;br /&gt;
# Recalibrate Base Qualities &lt;br /&gt;
&lt;br /&gt;
This processing results in 1 BAM file per sample. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline also includes Quality Control (QC) steps: &lt;br /&gt;
# Visualization of various quality measures (QPLOT) &lt;br /&gt;
# Screen for sample contamination &amp;amp; swap (VerifyBamID) &lt;br /&gt;
&lt;br /&gt;
Run the alignment pipeline (the example aligns 2 samples) : &lt;br /&gt;
 ~/gotcloud/gotcloud align --conf ~/gotcloudExample/[[#Alignment Configuration File|GBR2align.conf]] --outdir [[#Alignment Output Directory|~/gotcloudTutorialOut]] --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the alignment pipeline (about 1-3 minutes), you will see the following message: &lt;br /&gt;
 Processing finished in n secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The final BAM files produced by the alignment pipeline are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/bams&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* BAM (.bam) files - 1 per sample&lt;br /&gt;
** HG00096.recal.bam &lt;br /&gt;
** HG00100.recal.bam &lt;br /&gt;
* BAM index files (.bai) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.bai &lt;br /&gt;
** HG00100.recal.bam.bai &lt;br /&gt;
* BAM checksum files (.md5) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.md5 &lt;br /&gt;
** HG00100.recal.bam.md5 &lt;br /&gt;
* Indicator files that the step completed successfully:&lt;br /&gt;
** HG00096.recal.bam.done &lt;br /&gt;
** HG00100.recal.bam.done &lt;br /&gt;
&lt;br /&gt;
The Quality Control (QC) files are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/QCFiles&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* VerifyBamID output files:&lt;br /&gt;
** HG00096.genoCheck.depthRG &lt;br /&gt;
** HG00096.genoCheck.depthSM &lt;br /&gt;
** HG00096.genoCheck.selfRG &lt;br /&gt;
** HG00096.genoCheck.selfSM &lt;br /&gt;
** HG00100.genoCheck.depthRG &lt;br /&gt;
** HG00100.genoCheck.depthSM &lt;br /&gt;
** HG00100.genoCheck.selfRG &lt;br /&gt;
** HG00100.genoCheck.selfSM &lt;br /&gt;
&lt;br /&gt;
* VerifyBamID step completion files – 1 per sample&lt;br /&gt;
** HG00096.genoCheck.done &lt;br /&gt;
** HG00100.genoCheck.done &lt;br /&gt;
&lt;br /&gt;
* QPLOT output files&lt;br /&gt;
** HG00096.qplot.R &lt;br /&gt;
** HG00096.qplot.stats &lt;br /&gt;
** HG00100.qplot.R &lt;br /&gt;
** HG00100.qplot.stats &lt;br /&gt;
&lt;br /&gt;
* QPLOT step completion files – 1 per sample&lt;br /&gt;
** HG00096.qplot.done &lt;br /&gt;
** HG00100.qplot.done &lt;br /&gt;
&lt;br /&gt;
For information on the VerifyBamID output, see: [[Understanding VerifyBamID output]] &lt;br /&gt;
&lt;br /&gt;
For information on the QPLOT output, see: [[Understanding QPLOT output]]&lt;br /&gt;
&lt;br /&gt;
== STEP 3 : Run GotCloud Variant Calling Pipeline == &lt;br /&gt;
The next step is to analyze BAM files by calling SNPs and generating a VCF file containing the variant calls. &lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Filter out reads with low mapping quality &lt;br /&gt;
# Per Base Alignment Quality Adjustment (BAQ) &lt;br /&gt;
# Resolve overlapping paired end reads &lt;br /&gt;
# Generate genotype likelihood files &lt;br /&gt;
# Perform variant calling &lt;br /&gt;
# Extract features from variant sites &lt;br /&gt;
# Perform variant filtering &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
To speed variant calling, each chromosome is broken up into smaller regions which are processed separately.  While initially split by sample, the per sample data gets merged and is processed together for each region.  These regions are later merged to result in a single Variant Call File (VCF) per chromosome.  For the tutorial all of the data falls within a single region.&lt;br /&gt;
&lt;br /&gt;
Run the variant calling pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud snpcall --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --region 20:42900000-43200000 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the variant calling pipeline (about 3-4 minutes), you will see the following message: &lt;br /&gt;
  Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
On SNP Call success, the VCF files of interest are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered*&lt;br /&gt;
&lt;br /&gt;
This gives you the following files:&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.vcf.gz &#039;&#039;&#039; - vcf for whole chromosome after it has been run through hardfilters and SVM filters and marked with PASS/FAIL including per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf.norm.log - log file&lt;br /&gt;
* chr20.filtered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
* chr20.filtered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
* chr20.filtered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
&lt;br /&gt;
Also in the ~/gotcloudTutorialOut/vcfs/chr20 directory are intermediate files:&lt;br /&gt;
* the whole chromosome variant calls prior to any filtering: &lt;br /&gt;
** chr20.merged.sites.vcf - without per sample genotypes&lt;br /&gt;
** chr20.merged.stats.vcf &lt;br /&gt;
** chr20.merged.vcf - including per sample genotypes&lt;br /&gt;
** chr20.merged.vcf.OK - indicator that the step completed successfully&lt;br /&gt;
* the hardfiltered (pre-svm filtered) variant calls:&lt;br /&gt;
** chr20.filtered.vcf.gz - vcf for whole chromosome after it has been run through hard filters&lt;br /&gt;
** chr20.hardfiltered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.log - log file&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
* 40000001.45000000 subdirectory contains the data for just that region.&lt;br /&gt;
&lt;br /&gt;
The ~/gotcloudTutorialOut/split/chr20 folder contains a VCF with just the sites that pass the filters.&lt;br /&gt;
 ls ~/gotcloudTutorialOut/split/chr20/&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.PASS.vcf.gz &#039;&#039;&#039; – vcf of just sites that pass all filters&lt;br /&gt;
* chr20.filtered.PASS.split.1.vcf.gz - intermediate file&lt;br /&gt;
* chr20.filtered.PASS.split.err - log file&lt;br /&gt;
* chr20.filtered.PASS.split.vcflist - list of intermediate files&lt;br /&gt;
* subset.OK &lt;br /&gt;
&lt;br /&gt;
In addition to the vcfs subdirectory, there are additional intermediate files/directories:&lt;br /&gt;
* glfs – holds genotype likelihood format [[GLF]] files split by chromosome, region, and sample&lt;br /&gt;
* pvcfs – holds intermediate vcf files split by chromosome and region&lt;br /&gt;
&lt;br /&gt;
Note: the tutorial does not produce a target directory, but if you run with targeted data, you may see that.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 4 : Run GotCloud Genotype Refinement Pipeline == &lt;br /&gt;
The next step is to perform genotype refinement using linkage disequilibrium information using [http://faculty.washington.edu/browning/beagle/beagle.html Beagle] &amp;amp; [[ThunderVCF]]. &lt;br /&gt;
&lt;br /&gt;
Run the LD-aware genotype refinement pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud ldrefine --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of this pipeline (about 10 minutes), you will see the following message: &lt;br /&gt;
 Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The output from the beagle step of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
The output from the thunderVcf (final step) of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 5 : Run GotCloud Association Analysis Pipeline (EPACTS) == &lt;br /&gt;
&lt;br /&gt;
We will assume that the EPACTS are installed in the following directory&lt;br /&gt;
 setenv EPACTS /path/to/epacts&lt;br /&gt;
(If you need to install EPACTS, please refer to the documentation at [[EPACTS#Installation_Details]])&lt;br /&gt;
&lt;br /&gt;
 $EPACTS/epacts single --vcf ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered.vcf.gz --ped ~/gotcloudExample/test.GBR60.ped \\&lt;br /&gt;
    --out ~/gotcloudTutorialOut/epacts --test q.linear --run 1 --top 1 --chr 20&lt;br /&gt;
&lt;br /&gt;
Upon successful run, you will see files starting with ~/gotcloudTutorialOut/epacts&lt;br /&gt;
 ls ~/gotcloudTutorialOut/epacts*&lt;br /&gt;
&lt;br /&gt;
To see the top associated variants, you can run&lt;br /&gt;
 less ~/gotcloudTutorialOut/epacts.epacts.top5000&lt;br /&gt;
&lt;br /&gt;
To see the locus-zoom like plot, you can type the following command (assuming GNU gnuplot 4.2 or higher version was installed)&lt;br /&gt;
 xpdf ~/gotcloudTutorialOut/epacs.zoom.20.42987877.pdf&lt;br /&gt;
&lt;br /&gt;
Click [[Media:EPACTS TEST.zoom.20.42987877.pdf | Exampe LocusZoom PDF]] to see the expected output pdf&lt;br /&gt;
&lt;br /&gt;
= Frequently Asked Questions (FAQs) =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;I ran the tutorial example successfully, how can I run it with my real sequence data?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Congratulations for your successful run of your [[GotCloud]] Tutorial. Please see [[#Tutorial Inputs]] section to prepare your own input files for your sequence data. You will need to specify the FASTQ files associated with its sample names as explained. In addition, you will need to download the full reference and resource file across whole genome (the Tutorial contains only chr20 portion to make it compact) See [[#Alignment Configuration File]] section for the detailed information. Also, please refer to the original documentation of [[GotCloud]] for more detailed guide on installation beyond the scope of the tutorial.&lt;br /&gt;
&lt;br /&gt;
= Input Files for GotCloud Tutorial = &lt;br /&gt;
&lt;br /&gt;
This section describes the input files needed for the GotCloud tutorial. You don&#039;t need to know this detail to run the tutorial, but if you&#039;re interested in understanding the structure of GotCloud pipeline and run with your own sample, this would be a good starting point&lt;br /&gt;
&lt;br /&gt;
== Alignment Pipeline == &lt;br /&gt;
=== List of Input Files needed for Alignment ===&lt;br /&gt;
The command-line inputs to the tutorial alignment pipeline are: &lt;br /&gt;
# [[#Alignment Configuration File|Configuration File (--conf)]] &lt;br /&gt;
#* Specifies the configuration file to use when running&lt;br /&gt;
# [[#Alignment Output Directory|Output Directory (--outdir)]] &lt;br /&gt;
#* Directory where the output should be placed.&lt;br /&gt;
&lt;br /&gt;
Additional information required to run the alignment pipeline:&lt;br /&gt;
# [[#Alignment FASTQ Index File|Index file of FASTQs]]&lt;br /&gt;
# [[#Alignment Reference Files|Alignment Reference Files]]&lt;br /&gt;
For the tutorial, these values are specified in the configuration file.&lt;br /&gt;
&lt;br /&gt;
=== Alignment Configuration File === &lt;br /&gt;
The configuration file contains KEY = VALUE settings that override defaults and set specific values for the given run. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt; &lt;br /&gt;
INDEX_FILE = GBR60fastq.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_DIR = chr20Ref &lt;br /&gt;
AS = NCBI37 &lt;br /&gt;
FA_REF = $(REF_DIR)/human_g1k_v37_chr20.fa &lt;br /&gt;
DBSNP_VCF =  $(REF_DIR)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF = $(REF_DIR)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&amp;lt;/pre&amp;gt; &lt;br /&gt;
&lt;br /&gt;
This configuration file sets: &lt;br /&gt;
* [[#Alignment FASTQ Index File|INDEX_FILE]] - file containing the fastqs to be processed as well as the read group information for these fastqs.&lt;br /&gt;
* Reference Information: see [[#Alignment Reference Files|Alignment Reference Files]] for more information&lt;br /&gt;
** AS - assembly value to put in the BAM &lt;br /&gt;
&lt;br /&gt;
The index file and chromosome 20 references used in this tutorial are included with the example data under the ~/gotcloudExample directory.  The tutorial uses chromosome 20 only references in order to speed the processing time.  &lt;br /&gt;
&lt;br /&gt;
When running with your own data, you will need to update the:&lt;br /&gt;
* Index File to contain the information for your own fastq files&lt;br /&gt;
** See [[#Alignment FASTQ Index File|Alignment FASTQ Index File]] for more information on the contents of the index file.&lt;br /&gt;
* The reference files to be whole genome references&lt;br /&gt;
** If you are just running chromosome 20, you can use the tutorial references&lt;br /&gt;
** Whole genome reference files can be downloaded from [[GotCloudReference]].&lt;br /&gt;
** See [[#Alignment Reference Files|Alignment Reference Files]] for more information.&lt;br /&gt;
&lt;br /&gt;
Note: It is recommended that you use absolute paths (full path names, like “/home/mktrost/gotcloudReference” rather than just “gotcloudReference”).  This example does not use absolute paths in order to be flexible to where the data is installed, but using relative paths requires it to be run from the correct directory.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
&lt;br /&gt;
Reference files are required for running both the alignment and variant calling pipelines.  &lt;br /&gt;
&lt;br /&gt;
The configuration keys for setting these are:&lt;br /&gt;
* FA_REF - Genome sequence reference files (needed for both pipelines)&lt;br /&gt;
* DBSNP_VCF – DBSNP site VCF file (needed for both pipelines)&lt;br /&gt;
* HM3_VCF - HAPMAP site VCF file (needed for both pipelines)&lt;br /&gt;
* INDEL_PREFIX - Indel sites file (need for variant calling pipeline)&lt;br /&gt;
&lt;br /&gt;
The tutorial configuration file is setup to point to the required chromosome 20 reference files which are included with the tutorial example data in ~/gotcloudExample/chr20Ref/. &lt;br /&gt;
&lt;br /&gt;
If you are running more than just chromosome 20, you will need whole genome reference files which can be downloaded from [[GotCloudReference]].&lt;br /&gt;
&lt;br /&gt;
The configuration settings for these files are setup in the default configuration so do not need to be specified.  You just need to set REF_DIR in your configuration file to the path where you installed your reference files.&lt;br /&gt;
&lt;br /&gt;
To learn more about the reference files that are required, see [[GotCloud: Reference Files]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Alignment Output Directory === &lt;br /&gt;
This setting tells the alignment pipeline where to write the output and intermediate files. &lt;br /&gt;
&lt;br /&gt;
The output directory will be created if it doesn&#039;t already exist and will contain the following Directories/files: &lt;br /&gt;
* bams - directory containing the final bams/bai files &lt;br /&gt;
** HG00096.OK - indicates that this sample completed alignment processing &lt;br /&gt;
** HG00100.OK - indicates that this sample completed alignment processing &lt;br /&gt;
* failLogs - directory containing logs from steps that failed &lt;br /&gt;
** this directory is only created if an error is detected&lt;br /&gt;
* Makefiles - directory containing the makefiles with commands for processing each sample &lt;br /&gt;
** biopipe_HG00096.Makefile – commands for processing sample HG00096&lt;br /&gt;
** biopipe_HG00100.Makefile – commands for processing sample HG00100&lt;br /&gt;
** biopipe_HG00096.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
** biopipe_HG00100.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
* QCFiles - directory containing the QC Results &lt;br /&gt;
** &lt;br /&gt;
* tmp - directory containing temporary alignment files &lt;br /&gt;
** bwa.sai.t – contains temporary files for the 1st step of bwa that generates sai files from fastq files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
**** HG0096&lt;br /&gt;
** alignment.bwa – contains temporary files for the 2nd step of bwa that generates BAM files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
** alignment.pol - contains temporary files for the polish bam step that cleans up BAM files&lt;br /&gt;
** alignment.dedup – contains temporary files for the deduping step&lt;br /&gt;
*** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
*** *.metrics – metrics files for the deduping step&lt;br /&gt;
** alignment.recal – contains temporary files for the deduping step&lt;br /&gt;
*** *.log – contains information about the recalibration step&lt;br /&gt;
*** *.qemp – contains the recalibration tables used for recalibrating each BAM file&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Alignment FASTQ Index File=== &lt;br /&gt;
&lt;br /&gt;
There are four fastq files in {ROOT_DIR}/test/align/fastq/Sample_1 and four fastq files in {ROOT_DIR}/test/align/fastq/Sample_2, both in paired-end format.  Normally, we would need to build an index file for these files. Conveniently, an index file (indexFile.txt) already exists for the automatic test samples.  It can be found in {ROOT_DIR}/test/align/, and contains the following information in tab-delimited format: &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME FASTQ1                           FASTQ2                           RGID   SAMPLE    LIBRARY CENTER PLATFORM &lt;br /&gt;
 Sample1    fastq/Sample_1/File1_R1.fastq.gz fastq/Sample_1/File1_R2.fastq.gz RGID1  SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample1    fastq/Sample_1/File2_R1.fastq.gz fastq/Sample_1/File2_R2.fastq.gz RGID1a SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File1_R1.fastq.gz fastq/Sample_2/File1_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File2_R1.fastq.gz fastq/Sample_2/File2_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
&lt;br /&gt;
If you are in the {ROOT_DIR}/test/align directory, you can use this file as-is.  If you prefer, you can create a new index file and change the MERGE_NAME, RGID, SAMPLE, LIBRARY, CENTER, or PLATFORM values. It is recommended that you do not modify existing files in {ROOT_DIR}/test/align. &lt;br /&gt;
&lt;br /&gt;
If you want to run this example from a different directory, make sure the FASTQ1 and FASTQ2 paths are correct.  That is, each of the FASTQ1 and FASTQ2 entry in the index file should look like the following: &lt;br /&gt;
&lt;br /&gt;
 {ROOT_DIR}/test/align/fastq/Sample_1/File1_R1.fastq.gz &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to run this example from a different directory, but do not want to edit the index file, you can create a relative path to the test fastq files so their path agrees with that listed in the index file: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/align/fastq fastq &lt;br /&gt;
&lt;br /&gt;
This will create a symbolic link to the test fastq directory from your current directory. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Mapping_Pipeline#Sequence_Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Analyzing a Sample== &lt;br /&gt;
&lt;br /&gt;
Using UMAKE, you can analyze BAM files by calling SNPs, and generate a VCF file containing the results.  Once again, we can analyze BAM files used in the automatic test.  For this example, we have 60 BAM files, which can be found in {ROOT_DIR}/test/umake/bams.  These contain sequence information for a targeted region in chromosome 20. &lt;br /&gt;
&lt;br /&gt;
In addition to the BAM files, you will need three files to run UMAKE: an index file, a configuration file, and a bed file (needed to analyze BAM files from targeted/exome sequencing). &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
===Index file=== &lt;br /&gt;
&lt;br /&gt;
First, you need a list of all the BAM files to be analyzed. Conveniently, the a test index file (umake_test.index) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
 NA12272 ALL     bams/NA12272.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 NA12004 ALL     bams/NA12004.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 ... &lt;br /&gt;
 NA12874 ALL     bams/NA12874.mapped.LS454.ssaha2.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
&lt;br /&gt;
You can use this file directly if you change your current directory to {ROOT_DIR}/test/umake/. &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to copy and use this index file to a different directory, you can create a symbolic link to the bams folder as follows: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/umake/bams bams &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===BED file=== &lt;br /&gt;
&lt;br /&gt;
This file contains a single line: &lt;br /&gt;
&lt;br /&gt;
 chr20   20000050        20300000 &lt;br /&gt;
&lt;br /&gt;
You can copy this to the current directory and use it as-is. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Targeted.2FExome_Sequencing_Settings|targeted/exome sequencing settings]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Configuration file=== &lt;br /&gt;
&lt;br /&gt;
A configuration file (umake_test.conf) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
CHRS = 20 &lt;br /&gt;
BAM_INDEX = GBR60bam.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_ROOT = chr20Ref &lt;br /&gt;
# &lt;br /&gt;
REF = $(REF_ROOT)/human_g1k_v37_chr20.fa &lt;br /&gt;
INDEL_PREFIX = $(REF_ROOT)/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
DBSNP_VCF =  $(REF_ROOT)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF =  $(REF_ROOT)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 CHRS = 20 &lt;br /&gt;
 TEST_ROOT = $(UMAKE_ROOT)/test/umake &lt;br /&gt;
 BAM_INDEX = $(TEST_ROOT)/umake_test.index &lt;br /&gt;
 OUT_PREFIX = umake_test &lt;br /&gt;
 REF_ROOT = $(TEST_ROOT)/ref &lt;br /&gt;
 # &lt;br /&gt;
 REF = $(REF_ROOT)/karma.ref/human.g1k.v37.chr20.fa &lt;br /&gt;
 INDEL_PREFIX = $(REF_ROOT)/indels/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
 DBSNP_PREFIX =  $(REF_ROOT)/dbSNP/dbsnp_135_b37.rod &lt;br /&gt;
 HM3_PREFIX =  $(REF_ROOT)/HapMap3/hapmap3_r3_b37_fwd.consensus.qc.poly &lt;br /&gt;
 # &lt;br /&gt;
 RUN_INDEX = TRUE        # create BAM index file &lt;br /&gt;
 RUN_PILEUP = TRUE       # create GLF file from BAM &lt;br /&gt;
 RUN_GLFMULTIPLES = TRUE # create unfiltered SNP calls &lt;br /&gt;
 RUN_VCFPILEUP = TRUE    # create PVCF files using vcfPileup and run infoCollector &lt;br /&gt;
 RUN_FILTER = TRUE       # filter SNPs using vcfCooker &lt;br /&gt;
 RUN_SPLIT = TRUE        # split SNPs into chunks for genotype refinement &lt;br /&gt;
 RUN_BEAGLE = FALSE  # BEAGLE - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 RUN_SUBSET = FALSE  # SUBSET FOR THUNDER - MAY BE SET WITH BEAGLE STEP TOGETHER &lt;br /&gt;
 RUN_THUNDER = FALSE # THUNDER - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 ############################################################################### &lt;br /&gt;
 WRITE_TARGET_LOCI = TRUE  # FOR TARGETED SEQUENCING ONLY -- Write loci file when performing pileup &lt;br /&gt;
 UNIFORM_TARGET_BED = $(TEST_ROOT)/umake_test.bed # Targeted sequencing : When all individuals has the same target. Otherwise, comment it out &lt;br /&gt;
 OFFSET_OFF_TARGET = 50 # Extend target by given # of bases &lt;br /&gt;
 MULTIPLE_TARGET_MAP =  # Target per individual : Each line contains [SM_ID] [TARGET_BED] &lt;br /&gt;
 TARGET_DIR = target    # Directory to store target information &lt;br /&gt;
 SAMTOOLS_VIEW_TARGET_ONLY = TRUE # When performing samtools view, exclude off-target regions (may make command line too long) &lt;br /&gt;
&lt;br /&gt;
If you are running this from a different directory, you will want to change some of the lines as follows: &lt;br /&gt;
&lt;br /&gt;
 BAM_INDEX = {CURRENT_DIR}/umake_test.index &lt;br /&gt;
 UNIFORM_TARGET_BED = {CURRENT_DIR}/umake_test.bed &lt;br /&gt;
&lt;br /&gt;
where {CURRENT_DIR} is the absolute path to the directory that contains the index and bed files.  &lt;br /&gt;
&lt;br /&gt;
An additional option can be added in the configuration file: &lt;br /&gt;
&lt;br /&gt;
 OUT_DIR = {OUT_DIR} &lt;br /&gt;
&lt;br /&gt;
where {OUT_DIR} is the name of directory in which you want the output to be stored.  If you do not specify this in the configuration file, you will need to add an extra parameter when you run UMAKE in the next step. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Configuration_File|the configuration file]], [[Variant_Calling_Pipeline_(UMAKE)#Reference_Files|reference files]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Further Information== &lt;br /&gt;
&lt;br /&gt;
[[Mapping_Pipeline|Mapping (Alignment) Pipeline]] &lt;br /&gt;
&lt;br /&gt;
[[Variant_Calling_Pipeline_(UMAKE)|Variant Calling Pipeline (UMAKE)]]&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7546</id>
		<title>GotCloud: Alignment Pipeline</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7546"/>
		<updated>2013-06-21T20:37:09Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Optional Configurable Settings */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Alignment Pipeline = &lt;br /&gt;
&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
The Alignment/Mapping Pipeline takes FASTQ files and generates recalibrated BAM files from them. &lt;br /&gt;
&lt;br /&gt;
== Running the GotCloud Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline is run using the &amp;lt;code&amp;gt;align&amp;lt;/code&amp;gt; option of the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; script.  This option calls &amp;lt;code&amp;gt;align.pl&amp;lt;/code&amp;gt; found in the &amp;lt;code&amp;gt;bin/&amp;lt;/code&amp;gt; directory under the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; installation. &lt;br /&gt;
&lt;br /&gt;
===Running the Automated Test=== &lt;br /&gt;
&lt;br /&gt;
The automated test runs the alignment pipeline on a small set of test data and checks that the results against expected results validating that GotCloud is installed correctly. &lt;br /&gt;
&lt;br /&gt;
*Run alignment pipeline test: &lt;br /&gt;
 gotcloud align --test OUTPUT_DIR &lt;br /&gt;
where OUTPUT_DIR is the directory where you want to store the test results &lt;br /&gt;
&lt;br /&gt;
If you see &amp;quot;Successfully ran the test case, congratulations!&amp;quot;, then you are ready to align samples. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Overview of Alignment Pipeline Steps == &lt;br /&gt;
Here is an overview of the Alignment Pipeline: &lt;br /&gt;
&lt;br /&gt;
[[File:MappingSteps.png]] &lt;br /&gt;
&lt;br /&gt;
== Input Data:== &lt;br /&gt;
*Raw Sequence (FASTQ) files &lt;br /&gt;
*Sequence Index file containing fastqs &amp;amp; RG info &lt;br /&gt;
*Reference files &lt;br /&gt;
*(Optional) Configuration file to override default options &lt;br /&gt;
&lt;br /&gt;
=== Raw Sequence (FASTQ) files === &lt;br /&gt;
&lt;br /&gt;
These are the FASTQ files that need to be mapped to BAM files. &lt;br /&gt;
&lt;br /&gt;
These files are specified in the [[#Sequence Index File|Sequence Index File]]. &lt;br /&gt;
&lt;br /&gt;
=== Sequence Index File === &lt;br /&gt;
This file specifies the FASTQ files that need to be processed and the Read Group information for them. &lt;br /&gt;
&lt;br /&gt;
This file is specified either via the command line parameter &amp;lt;code&amp;gt;--index_file&amp;lt;/code&amp;gt; or via the configuration file setting &amp;lt;code&amp;gt;INDEX_FILE&amp;lt;/code&amp;gt;.  &lt;br /&gt;
&lt;br /&gt;
The command-line setting takes precedence over the configuration file setting. &lt;br /&gt;
&lt;br /&gt;
The Sequence Index is a tab delimited file that starts with a header line.  The columns may be in any order. &lt;br /&gt;
&lt;br /&gt;
Following the header line, there is one line per single-end read and one line per paired-end read (only 1 line per pair). &lt;br /&gt;
&lt;br /&gt;
Required Column Names: &lt;br /&gt;
* MERGE_NAME - base name for the resulting BAM file for the sample (used to group multiple fastqs or fastq pairs into a single BAM) &lt;br /&gt;
* FASTQ1 - name of the fastq or the first in the pair if paired-end.  (Only 1 line per pair) &lt;br /&gt;
&lt;br /&gt;
Optional Column Names: &lt;br /&gt;
* FASTQ2 - name of the 2nd fastq in paired-end reads.  Specify &#039;.&#039; if the column exists, but this line is single-ended. &lt;br /&gt;
* RGID - Read Group ID for this entry &lt;br /&gt;
* SAMPLE - Sample Name for this entry &lt;br /&gt;
* LIBRARY - Library for this entry &lt;br /&gt;
* CENTER - Center Name for this entry &lt;br /&gt;
* PLATFORM - Platform for this entry &lt;br /&gt;
&lt;br /&gt;
The RGID, SAMPLE, LIBRARY, CENTER, and PLATFORM are used to populate the Read Group information for this entry.  These fields are optional.  Either leave the column header out of the file or specify &#039;.&#039; if the column header exists, but the data is N/A.  As long as the RGID field is specified non-N/A fields are added to the BAM file. &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME	FASTQ1	FASTQ2	RGID	SAMPLE	LIBRARY	CENTER	PLATFORM &lt;br /&gt;
 Sample1	fastq/S1/F1_R1.fastq.gz	fastq/S1/F1_R2.fastq.gz	RGID1	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample1	fastq/S1/F2_R1.fastq.gz	fastq/S1/F2_R2.fastq.gz	RGID1a	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F1_R1.fastq.gz	fastq/S2/F1_R2.fastq.gz	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F2.fastq.gz	.	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--fastq&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;FASTQ&amp;lt;/code&amp;gt; setting can be used to specify a prefix to the FASTQ1/FASTQ2 file paths that should be applied before using the files. &lt;br /&gt;
&lt;br /&gt;
=== Reference Files === &lt;br /&gt;
&lt;br /&gt;
The following Reference Files are required: &lt;br /&gt;
* Reference File fasta files &lt;br /&gt;
** Files required: .fa, -bs.umfa, .GCContent, .amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa &lt;br /&gt;
*** If you don&#039;t have the -bs.umfa file, the software will try to create it in the same directory as the reference fasta. &lt;br /&gt;
*** .GCContent can be generated using qplot, see: [[QPLOT#Input_files| QPLOT: Input Files: --gccontent]] and name the resulting file as &amp;lt;code&amp;gt;.fa.GCcontent&amp;lt;/code&amp;gt; &lt;br /&gt;
*** Use &amp;lt;code&amp;gt;bin/bwa index ref.fa&amp;lt;/code&amp;gt; if you need to generate the bwa reference files (.amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa) &lt;br /&gt;
** Configuration Name: FA_REF - specify the ref.fa/ref.fa.gz name &lt;br /&gt;
* DBSNP File &lt;br /&gt;
** tab delimited file/VCF, can be compressed &lt;br /&gt;
*** 1st column -&amp;gt; chromosome &lt;br /&gt;
*** 2nd column -&amp;gt; 1-based position &lt;br /&gt;
** Configuration Name: DBSNP_VCF &lt;br /&gt;
* PLINK-compatible binary genotype files &lt;br /&gt;
** Files required: .bed, .bin, .fam &lt;br /&gt;
** Configuration Name: PLINK &lt;br /&gt;
&lt;br /&gt;
=== Configuration File === &lt;br /&gt;
Configuration file contains the run-time options including the software binaries and command line arguments.  A default configuration file is automatically loaded.  Users may specify their own configuration file specifying just the values different than the defaults.  The configuration file is not required if there are no values to override. &lt;br /&gt;
&lt;br /&gt;
Comments begin with a &amp;lt;code&amp;gt;#&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Format: KEY = value &lt;br /&gt;
&lt;br /&gt;
Where KEY is the item being set and value is its new value &lt;br /&gt;
&lt;br /&gt;
See [[#Command-Line Options|Command-Line Options]] for values that can be set either via command line or via configuration. &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings &lt;br /&gt;
&lt;br /&gt;
==== Required Settings ==== &lt;br /&gt;
See [[#Reference Files|Reference Files]] for the required reference file settings. &lt;br /&gt;
&lt;br /&gt;
See [[#Sequence Index File|Sequence Index File]] for how to set the index file either via command line options or via configuration. &lt;br /&gt;
&lt;br /&gt;
==== Turning Off Optional Steps==== &lt;br /&gt;
Quality Control steps can be disabled. &lt;br /&gt;
&lt;br /&gt;
To Disable QPLOT, set: &lt;br /&gt;
 RUN_QPLOT = 0 &lt;br /&gt;
&lt;br /&gt;
To Disable VerifyBamID, set: &lt;br /&gt;
 RUN_VERIFY_BAM_ID = 0 &lt;br /&gt;
&lt;br /&gt;
==== Optional Configurable Settings ==== &lt;br /&gt;
You may want to adjust the amount of memory/threads that are used: &lt;br /&gt;
&lt;br /&gt;
There are additional configurable settings, but these are the ones most likely to be adjusted. &lt;br /&gt;
&lt;br /&gt;
* BWA_THREADS = -t N &lt;br /&gt;
** Fill in the N with the number of threads you want BWA to run with, default is 1 &lt;br /&gt;
* BWA_MAX_MEM = 2000000000 &lt;br /&gt;
** Maximum amount of memory used by samtools sort after running bwa &lt;br /&gt;
* JAVA_MEM = -Xmx4g &lt;br /&gt;
** Set the maximum size of the java memory allocation pool.  Default is 4g, adjust that as necessary.&lt;br /&gt;
*BATCH_TYPE = mosix&lt;br /&gt;
** Tells the cluster gateway to use mosix to send jobs to the client nodes.&lt;br /&gt;
*BATCH_OPTS = -j36,37,38,39,40,41,45,46,47,48,49&lt;br /&gt;
** Specifies which client nodes mosix should send jobs to.&lt;br /&gt;
&lt;br /&gt;
== Running the Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
=== Command-Line Options === &lt;br /&gt;
* help - print usage &lt;br /&gt;
* test OUTPUT_DIR - run the test example placing the output in a user specified OUTPUT_DIR.  No other options are required. &lt;br /&gt;
* out_dir OUTPUT_DIR - directory for the output &lt;br /&gt;
** May also be specified via OUT_DIR in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* conf CONFIG_FILE - configuration file &lt;br /&gt;
* index_file INDEX_FILE_NAME  - name of the index file &lt;br /&gt;
** May also be specified via INDEX_FILE in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* ref_dir REFERENCE_DIR - value to set config key REF_DIR to, overriding other values, REF_DIR can then be used inside config files. &lt;br /&gt;
** May also be specified via REF_DIR in the configuration file &lt;br /&gt;
* fastq FASTQ_PATH - prefix path to the fastq files specified in the INDEX_FILE &lt;br /&gt;
** May also be specified via FASTQ in the configuration file &lt;br /&gt;
* keepTmp - Do not remove the temporary files (removed by default) &lt;br /&gt;
** May also be specified via KEEP_TMP in the configuration file &lt;br /&gt;
* numcs N - Replace N with the number of samples that should be processed in parallel&lt;br /&gt;
* numjobs N - Replace N with the number of targets in each makefile that should be run in parallel &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings&lt;br /&gt;
&lt;br /&gt;
===Running the Alignment Pipeline=== &lt;br /&gt;
Run &amp;lt;code&amp;gt;gotcloud align&amp;lt;/code&amp;gt; with the appropriate command-line parameters. &lt;br /&gt;
&lt;br /&gt;
Example: &lt;br /&gt;
 gotcloud align --conf config.txt --outdir output &lt;br /&gt;
&lt;br /&gt;
This step generates 1 Makefile per sample in the output/Makefiles/ directory and then automatically runs them.  The Makefiles contain all of the information to run each sample. &lt;br /&gt;
&lt;br /&gt;
If you only want to generate the makefiles and not run them, use the &amp;lt;code&amp;gt;--dryrun&amp;lt;/code&amp;gt; option.  It will generate the Makefiles and print instructions for running the Makefiles. &lt;br /&gt;
&lt;br /&gt;
Each Makefile is independent and can be run in parallel and across a cloud. &lt;br /&gt;
&lt;br /&gt;
On success, you will see:&lt;br /&gt;
 Processing finished in nn secs with no errors reported &lt;br /&gt;
and should see the following subdirectories under the user specified output directory: &lt;br /&gt;
* bams/ &lt;br /&gt;
* Makefiles/ &lt;br /&gt;
* QCFiles/ (if all quality control is not disabled) &lt;br /&gt;
* tmp/ &lt;br /&gt;
&lt;br /&gt;
You should see a &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; for each Sample in the index file. &lt;br /&gt;
&lt;br /&gt;
If you do not see these &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; files, then your Alignment Pipeline failed. &lt;br /&gt;
&lt;br /&gt;
On success, the bams/ directory contains the final BAMs and bais.&lt;br /&gt;
&lt;br /&gt;
If processing fails part way through, you can pick up where you left off by rerunning gotcloud or the make command.&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7545</id>
		<title>GotCloud: Alignment Pipeline</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7545"/>
		<updated>2013-06-21T20:31:11Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Command-Line Options */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Alignment Pipeline = &lt;br /&gt;
&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
The Alignment/Mapping Pipeline takes FASTQ files and generates recalibrated BAM files from them. &lt;br /&gt;
&lt;br /&gt;
== Running the GotCloud Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline is run using the &amp;lt;code&amp;gt;align&amp;lt;/code&amp;gt; option of the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; script.  This option calls &amp;lt;code&amp;gt;align.pl&amp;lt;/code&amp;gt; found in the &amp;lt;code&amp;gt;bin/&amp;lt;/code&amp;gt; directory under the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; installation. &lt;br /&gt;
&lt;br /&gt;
===Running the Automated Test=== &lt;br /&gt;
&lt;br /&gt;
The automated test runs the alignment pipeline on a small set of test data and checks that the results against expected results validating that GotCloud is installed correctly. &lt;br /&gt;
&lt;br /&gt;
*Run alignment pipeline test: &lt;br /&gt;
 gotcloud align --test OUTPUT_DIR &lt;br /&gt;
where OUTPUT_DIR is the directory where you want to store the test results &lt;br /&gt;
&lt;br /&gt;
If you see &amp;quot;Successfully ran the test case, congratulations!&amp;quot;, then you are ready to align samples. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Overview of Alignment Pipeline Steps == &lt;br /&gt;
Here is an overview of the Alignment Pipeline: &lt;br /&gt;
&lt;br /&gt;
[[File:MappingSteps.png]] &lt;br /&gt;
&lt;br /&gt;
== Input Data:== &lt;br /&gt;
*Raw Sequence (FASTQ) files &lt;br /&gt;
*Sequence Index file containing fastqs &amp;amp; RG info &lt;br /&gt;
*Reference files &lt;br /&gt;
*(Optional) Configuration file to override default options &lt;br /&gt;
&lt;br /&gt;
=== Raw Sequence (FASTQ) files === &lt;br /&gt;
&lt;br /&gt;
These are the FASTQ files that need to be mapped to BAM files. &lt;br /&gt;
&lt;br /&gt;
These files are specified in the [[#Sequence Index File|Sequence Index File]]. &lt;br /&gt;
&lt;br /&gt;
=== Sequence Index File === &lt;br /&gt;
This file specifies the FASTQ files that need to be processed and the Read Group information for them. &lt;br /&gt;
&lt;br /&gt;
This file is specified either via the command line parameter &amp;lt;code&amp;gt;--index_file&amp;lt;/code&amp;gt; or via the configuration file setting &amp;lt;code&amp;gt;INDEX_FILE&amp;lt;/code&amp;gt;.  &lt;br /&gt;
&lt;br /&gt;
The command-line setting takes precedence over the configuration file setting. &lt;br /&gt;
&lt;br /&gt;
The Sequence Index is a tab delimited file that starts with a header line.  The columns may be in any order. &lt;br /&gt;
&lt;br /&gt;
Following the header line, there is one line per single-end read and one line per paired-end read (only 1 line per pair). &lt;br /&gt;
&lt;br /&gt;
Required Column Names: &lt;br /&gt;
* MERGE_NAME - base name for the resulting BAM file for the sample (used to group multiple fastqs or fastq pairs into a single BAM) &lt;br /&gt;
* FASTQ1 - name of the fastq or the first in the pair if paired-end.  (Only 1 line per pair) &lt;br /&gt;
&lt;br /&gt;
Optional Column Names: &lt;br /&gt;
* FASTQ2 - name of the 2nd fastq in paired-end reads.  Specify &#039;.&#039; if the column exists, but this line is single-ended. &lt;br /&gt;
* RGID - Read Group ID for this entry &lt;br /&gt;
* SAMPLE - Sample Name for this entry &lt;br /&gt;
* LIBRARY - Library for this entry &lt;br /&gt;
* CENTER - Center Name for this entry &lt;br /&gt;
* PLATFORM - Platform for this entry &lt;br /&gt;
&lt;br /&gt;
The RGID, SAMPLE, LIBRARY, CENTER, and PLATFORM are used to populate the Read Group information for this entry.  These fields are optional.  Either leave the column header out of the file or specify &#039;.&#039; if the column header exists, but the data is N/A.  As long as the RGID field is specified non-N/A fields are added to the BAM file. &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME	FASTQ1	FASTQ2	RGID	SAMPLE	LIBRARY	CENTER	PLATFORM &lt;br /&gt;
 Sample1	fastq/S1/F1_R1.fastq.gz	fastq/S1/F1_R2.fastq.gz	RGID1	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample1	fastq/S1/F2_R1.fastq.gz	fastq/S1/F2_R2.fastq.gz	RGID1a	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F1_R1.fastq.gz	fastq/S2/F1_R2.fastq.gz	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F2.fastq.gz	.	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--fastq&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;FASTQ&amp;lt;/code&amp;gt; setting can be used to specify a prefix to the FASTQ1/FASTQ2 file paths that should be applied before using the files. &lt;br /&gt;
&lt;br /&gt;
=== Reference Files === &lt;br /&gt;
&lt;br /&gt;
The following Reference Files are required: &lt;br /&gt;
* Reference File fasta files &lt;br /&gt;
** Files required: .fa, -bs.umfa, .GCContent, .amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa &lt;br /&gt;
*** If you don&#039;t have the -bs.umfa file, the software will try to create it in the same directory as the reference fasta. &lt;br /&gt;
*** .GCContent can be generated using qplot, see: [[QPLOT#Input_files| QPLOT: Input Files: --gccontent]] and name the resulting file as &amp;lt;code&amp;gt;.fa.GCcontent&amp;lt;/code&amp;gt; &lt;br /&gt;
*** Use &amp;lt;code&amp;gt;bin/bwa index ref.fa&amp;lt;/code&amp;gt; if you need to generate the bwa reference files (.amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa) &lt;br /&gt;
** Configuration Name: FA_REF - specify the ref.fa/ref.fa.gz name &lt;br /&gt;
* DBSNP File &lt;br /&gt;
** tab delimited file/VCF, can be compressed &lt;br /&gt;
*** 1st column -&amp;gt; chromosome &lt;br /&gt;
*** 2nd column -&amp;gt; 1-based position &lt;br /&gt;
** Configuration Name: DBSNP_VCF &lt;br /&gt;
* PLINK-compatible binary genotype files &lt;br /&gt;
** Files required: .bed, .bin, .fam &lt;br /&gt;
** Configuration Name: PLINK &lt;br /&gt;
&lt;br /&gt;
=== Configuration File === &lt;br /&gt;
Configuration file contains the run-time options including the software binaries and command line arguments.  A default configuration file is automatically loaded.  Users may specify their own configuration file specifying just the values different than the defaults.  The configuration file is not required if there are no values to override. &lt;br /&gt;
&lt;br /&gt;
Comments begin with a &amp;lt;code&amp;gt;#&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Format: KEY = value &lt;br /&gt;
&lt;br /&gt;
Where KEY is the item being set and value is its new value &lt;br /&gt;
&lt;br /&gt;
See [[#Command-Line Options|Command-Line Options]] for values that can be set either via command line or via configuration. &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings &lt;br /&gt;
&lt;br /&gt;
==== Required Settings ==== &lt;br /&gt;
See [[#Reference Files|Reference Files]] for the required reference file settings. &lt;br /&gt;
&lt;br /&gt;
See [[#Sequence Index File|Sequence Index File]] for how to set the index file either via command line options or via configuration. &lt;br /&gt;
&lt;br /&gt;
==== Turning Off Optional Steps==== &lt;br /&gt;
Quality Control steps can be disabled. &lt;br /&gt;
&lt;br /&gt;
To Disable QPLOT, set: &lt;br /&gt;
 RUN_QPLOT = 0 &lt;br /&gt;
&lt;br /&gt;
To Disable VerifyBamID, set: &lt;br /&gt;
 RUN_VERIFY_BAM_ID = 0 &lt;br /&gt;
&lt;br /&gt;
==== Optional Configurable Settings ==== &lt;br /&gt;
You may want to adjust the amount of memory/threads that are used: &lt;br /&gt;
&lt;br /&gt;
There are additional configurable settings, but these are the ones most likely to be adjusted. &lt;br /&gt;
&lt;br /&gt;
* BWA_THREADS = -t N &lt;br /&gt;
** Fill in the N with the number of threads you want BWA to run with, default is 1 &lt;br /&gt;
* BWA_MAX_MEM = 2000000000 &lt;br /&gt;
** Maximum amount of memory used by samtools sort after running bwa &lt;br /&gt;
* JAVA_MEM = -Xmx4g &lt;br /&gt;
** Set the maximum size of the java memory allocation pool.  Default is 4g, adjust that as necessary. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Running the Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
=== Command-Line Options === &lt;br /&gt;
* help - print usage &lt;br /&gt;
* test OUTPUT_DIR - run the test example placing the output in a user specified OUTPUT_DIR.  No other options are required. &lt;br /&gt;
* out_dir OUTPUT_DIR - directory for the output &lt;br /&gt;
** May also be specified via OUT_DIR in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* conf CONFIG_FILE - configuration file &lt;br /&gt;
* index_file INDEX_FILE_NAME  - name of the index file &lt;br /&gt;
** May also be specified via INDEX_FILE in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* ref_dir REFERENCE_DIR - value to set config key REF_DIR to, overriding other values, REF_DIR can then be used inside config files. &lt;br /&gt;
** May also be specified via REF_DIR in the configuration file &lt;br /&gt;
* fastq FASTQ_PATH - prefix path to the fastq files specified in the INDEX_FILE &lt;br /&gt;
** May also be specified via FASTQ in the configuration file &lt;br /&gt;
* keepTmp - Do not remove the temporary files (removed by default) &lt;br /&gt;
** May also be specified via KEEP_TMP in the configuration file &lt;br /&gt;
* numcs N - Replace N with the number of samples that should be processed in parallel&lt;br /&gt;
* numjobs N - Replace N with the number of targets in each makefile that should be run in parallel &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings&lt;br /&gt;
&lt;br /&gt;
===Running the Alignment Pipeline=== &lt;br /&gt;
Run &amp;lt;code&amp;gt;gotcloud align&amp;lt;/code&amp;gt; with the appropriate command-line parameters. &lt;br /&gt;
&lt;br /&gt;
Example: &lt;br /&gt;
 gotcloud align --conf config.txt --outdir output &lt;br /&gt;
&lt;br /&gt;
This step generates 1 Makefile per sample in the output/Makefiles/ directory and then automatically runs them.  The Makefiles contain all of the information to run each sample. &lt;br /&gt;
&lt;br /&gt;
If you only want to generate the makefiles and not run them, use the &amp;lt;code&amp;gt;--dryrun&amp;lt;/code&amp;gt; option.  It will generate the Makefiles and print instructions for running the Makefiles. &lt;br /&gt;
&lt;br /&gt;
Each Makefile is independent and can be run in parallel and across a cloud. &lt;br /&gt;
&lt;br /&gt;
On success, you will see:&lt;br /&gt;
 Processing finished in nn secs with no errors reported &lt;br /&gt;
and should see the following subdirectories under the user specified output directory: &lt;br /&gt;
* bams/ &lt;br /&gt;
* Makefiles/ &lt;br /&gt;
* QCFiles/ (if all quality control is not disabled) &lt;br /&gt;
* tmp/ &lt;br /&gt;
&lt;br /&gt;
You should see a &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; for each Sample in the index file. &lt;br /&gt;
&lt;br /&gt;
If you do not see these &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; files, then your Alignment Pipeline failed. &lt;br /&gt;
&lt;br /&gt;
On success, the bams/ directory contains the final BAMs and bais.&lt;br /&gt;
&lt;br /&gt;
If processing fails part way through, you can pick up where you left off by rerunning gotcloud or the make command.&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7544</id>
		<title>GotCloud: Alignment Pipeline</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Pipeline&amp;diff=7544"/>
		<updated>2013-06-21T20:29:52Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Command-Line Options */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Alignment Pipeline = &lt;br /&gt;
&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
The Alignment/Mapping Pipeline takes FASTQ files and generates recalibrated BAM files from them. &lt;br /&gt;
&lt;br /&gt;
== Running the GotCloud Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline is run using the &amp;lt;code&amp;gt;align&amp;lt;/code&amp;gt; option of the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; script.  This option calls &amp;lt;code&amp;gt;align.pl&amp;lt;/code&amp;gt; found in the &amp;lt;code&amp;gt;bin/&amp;lt;/code&amp;gt; directory under the &amp;lt;code&amp;gt;gotcloud&amp;lt;/code&amp;gt; installation. &lt;br /&gt;
&lt;br /&gt;
===Running the Automated Test=== &lt;br /&gt;
&lt;br /&gt;
The automated test runs the alignment pipeline on a small set of test data and checks that the results against expected results validating that GotCloud is installed correctly. &lt;br /&gt;
&lt;br /&gt;
*Run alignment pipeline test: &lt;br /&gt;
 gotcloud align --test OUTPUT_DIR &lt;br /&gt;
where OUTPUT_DIR is the directory where you want to store the test results &lt;br /&gt;
&lt;br /&gt;
If you see &amp;quot;Successfully ran the test case, congratulations!&amp;quot;, then you are ready to align samples. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Overview of Alignment Pipeline Steps == &lt;br /&gt;
Here is an overview of the Alignment Pipeline: &lt;br /&gt;
&lt;br /&gt;
[[File:MappingSteps.png]] &lt;br /&gt;
&lt;br /&gt;
== Input Data:== &lt;br /&gt;
*Raw Sequence (FASTQ) files &lt;br /&gt;
*Sequence Index file containing fastqs &amp;amp; RG info &lt;br /&gt;
*Reference files &lt;br /&gt;
*(Optional) Configuration file to override default options &lt;br /&gt;
&lt;br /&gt;
=== Raw Sequence (FASTQ) files === &lt;br /&gt;
&lt;br /&gt;
These are the FASTQ files that need to be mapped to BAM files. &lt;br /&gt;
&lt;br /&gt;
These files are specified in the [[#Sequence Index File|Sequence Index File]]. &lt;br /&gt;
&lt;br /&gt;
=== Sequence Index File === &lt;br /&gt;
This file specifies the FASTQ files that need to be processed and the Read Group information for them. &lt;br /&gt;
&lt;br /&gt;
This file is specified either via the command line parameter &amp;lt;code&amp;gt;--index_file&amp;lt;/code&amp;gt; or via the configuration file setting &amp;lt;code&amp;gt;INDEX_FILE&amp;lt;/code&amp;gt;.  &lt;br /&gt;
&lt;br /&gt;
The command-line setting takes precedence over the configuration file setting. &lt;br /&gt;
&lt;br /&gt;
The Sequence Index is a tab delimited file that starts with a header line.  The columns may be in any order. &lt;br /&gt;
&lt;br /&gt;
Following the header line, there is one line per single-end read and one line per paired-end read (only 1 line per pair). &lt;br /&gt;
&lt;br /&gt;
Required Column Names: &lt;br /&gt;
* MERGE_NAME - base name for the resulting BAM file for the sample (used to group multiple fastqs or fastq pairs into a single BAM) &lt;br /&gt;
* FASTQ1 - name of the fastq or the first in the pair if paired-end.  (Only 1 line per pair) &lt;br /&gt;
&lt;br /&gt;
Optional Column Names: &lt;br /&gt;
* FASTQ2 - name of the 2nd fastq in paired-end reads.  Specify &#039;.&#039; if the column exists, but this line is single-ended. &lt;br /&gt;
* RGID - Read Group ID for this entry &lt;br /&gt;
* SAMPLE - Sample Name for this entry &lt;br /&gt;
* LIBRARY - Library for this entry &lt;br /&gt;
* CENTER - Center Name for this entry &lt;br /&gt;
* PLATFORM - Platform for this entry &lt;br /&gt;
&lt;br /&gt;
The RGID, SAMPLE, LIBRARY, CENTER, and PLATFORM are used to populate the Read Group information for this entry.  These fields are optional.  Either leave the column header out of the file or specify &#039;.&#039; if the column header exists, but the data is N/A.  As long as the RGID field is specified non-N/A fields are added to the BAM file. &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME	FASTQ1	FASTQ2	RGID	SAMPLE	LIBRARY	CENTER	PLATFORM &lt;br /&gt;
 Sample1	fastq/S1/F1_R1.fastq.gz	fastq/S1/F1_R2.fastq.gz	RGID1	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample1	fastq/S1/F2_R1.fastq.gz	fastq/S1/F2_R2.fastq.gz	RGID1a	SampleID1	Lib1	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F1_R1.fastq.gz	fastq/S2/F1_R2.fastq.gz	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
 Sample2	fastq/S2/F2.fastq.gz	.	RGID2	SampleID2	Lib2	UM	ILLUMINA &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--fastq&amp;lt;/code&amp;gt;/&amp;lt;code&amp;gt;FASTQ&amp;lt;/code&amp;gt; setting can be used to specify a prefix to the FASTQ1/FASTQ2 file paths that should be applied before using the files. &lt;br /&gt;
&lt;br /&gt;
=== Reference Files === &lt;br /&gt;
&lt;br /&gt;
The following Reference Files are required: &lt;br /&gt;
* Reference File fasta files &lt;br /&gt;
** Files required: .fa, -bs.umfa, .GCContent, .amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa &lt;br /&gt;
*** If you don&#039;t have the -bs.umfa file, the software will try to create it in the same directory as the reference fasta. &lt;br /&gt;
*** .GCContent can be generated using qplot, see: [[QPLOT#Input_files| QPLOT: Input Files: --gccontent]] and name the resulting file as &amp;lt;code&amp;gt;.fa.GCcontent&amp;lt;/code&amp;gt; &lt;br /&gt;
*** Use &amp;lt;code&amp;gt;bin/bwa index ref.fa&amp;lt;/code&amp;gt; if you need to generate the bwa reference files (.amb, .ann, .bwt, .pac, .rbwt, .rpac, .rsa, .sa) &lt;br /&gt;
** Configuration Name: FA_REF - specify the ref.fa/ref.fa.gz name &lt;br /&gt;
* DBSNP File &lt;br /&gt;
** tab delimited file/VCF, can be compressed &lt;br /&gt;
*** 1st column -&amp;gt; chromosome &lt;br /&gt;
*** 2nd column -&amp;gt; 1-based position &lt;br /&gt;
** Configuration Name: DBSNP_VCF &lt;br /&gt;
* PLINK-compatible binary genotype files &lt;br /&gt;
** Files required: .bed, .bin, .fam &lt;br /&gt;
** Configuration Name: PLINK &lt;br /&gt;
&lt;br /&gt;
=== Configuration File === &lt;br /&gt;
Configuration file contains the run-time options including the software binaries and command line arguments.  A default configuration file is automatically loaded.  Users may specify their own configuration file specifying just the values different than the defaults.  The configuration file is not required if there are no values to override. &lt;br /&gt;
&lt;br /&gt;
Comments begin with a &amp;lt;code&amp;gt;#&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Format: KEY = value &lt;br /&gt;
&lt;br /&gt;
Where KEY is the item being set and value is its new value &lt;br /&gt;
&lt;br /&gt;
See [[#Command-Line Options|Command-Line Options]] for values that can be set either via command line or via configuration. &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings &lt;br /&gt;
&lt;br /&gt;
==== Required Settings ==== &lt;br /&gt;
See [[#Reference Files|Reference Files]] for the required reference file settings. &lt;br /&gt;
&lt;br /&gt;
See [[#Sequence Index File|Sequence Index File]] for how to set the index file either via command line options or via configuration. &lt;br /&gt;
&lt;br /&gt;
==== Turning Off Optional Steps==== &lt;br /&gt;
Quality Control steps can be disabled. &lt;br /&gt;
&lt;br /&gt;
To Disable QPLOT, set: &lt;br /&gt;
 RUN_QPLOT = 0 &lt;br /&gt;
&lt;br /&gt;
To Disable VerifyBamID, set: &lt;br /&gt;
 RUN_VERIFY_BAM_ID = 0 &lt;br /&gt;
&lt;br /&gt;
==== Optional Configurable Settings ==== &lt;br /&gt;
You may want to adjust the amount of memory/threads that are used: &lt;br /&gt;
&lt;br /&gt;
There are additional configurable settings, but these are the ones most likely to be adjusted. &lt;br /&gt;
&lt;br /&gt;
* BWA_THREADS = -t N &lt;br /&gt;
** Fill in the N with the number of threads you want BWA to run with, default is 1 &lt;br /&gt;
* BWA_MAX_MEM = 2000000000 &lt;br /&gt;
** Maximum amount of memory used by samtools sort after running bwa &lt;br /&gt;
* JAVA_MEM = -Xmx4g &lt;br /&gt;
** Set the maximum size of the java memory allocation pool.  Default is 4g, adjust that as necessary. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Running the Alignment Pipeline == &lt;br /&gt;
&lt;br /&gt;
=== Command-Line Options === &lt;br /&gt;
* help - print usage &lt;br /&gt;
* test OUTPUT_DIR - run the test example placing the output in a user specified OUTPUT_DIR.  No other options are required. &lt;br /&gt;
* out_dir OUTPUT_DIR - directory for the output &lt;br /&gt;
** May also be specified via OUT_DIR in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* conf CONFIG_FILE - configuration file &lt;br /&gt;
* index_file INDEX_FILE_NAME  - name of the index file &lt;br /&gt;
** May also be specified via INDEX_FILE in the configuration file &lt;br /&gt;
** Required to be set either via command-line or configuration &lt;br /&gt;
* ref_dir REFERENCE_DIR - value to set config key REF_DIR to, overriding other values, REF_DIR can then be used inside config files. &lt;br /&gt;
** May also be specified via REF_DIR in the configuration file &lt;br /&gt;
* fastq FASTQ_PATH - prefix path to the fastq files specified in the INDEX_FILE &lt;br /&gt;
** May also be specified via FASTQ in the configuration file &lt;br /&gt;
* keepTmp - Do not remove the temporary files (removed by default) &lt;br /&gt;
** May also be specified via KEEP_TMP in the configuration file &lt;br /&gt;
* numjobs N - Replace N with the number of jobs that should be run in parallel &lt;br /&gt;
&lt;br /&gt;
Note: Command-line options take priority over configuration file settings&lt;br /&gt;
&lt;br /&gt;
===Running the Alignment Pipeline=== &lt;br /&gt;
Run &amp;lt;code&amp;gt;gotcloud align&amp;lt;/code&amp;gt; with the appropriate command-line parameters. &lt;br /&gt;
&lt;br /&gt;
Example: &lt;br /&gt;
 gotcloud align --conf config.txt --outdir output &lt;br /&gt;
&lt;br /&gt;
This step generates 1 Makefile per sample in the output/Makefiles/ directory and then automatically runs them.  The Makefiles contain all of the information to run each sample. &lt;br /&gt;
&lt;br /&gt;
If you only want to generate the makefiles and not run them, use the &amp;lt;code&amp;gt;--dryrun&amp;lt;/code&amp;gt; option.  It will generate the Makefiles and print instructions for running the Makefiles. &lt;br /&gt;
&lt;br /&gt;
Each Makefile is independent and can be run in parallel and across a cloud. &lt;br /&gt;
&lt;br /&gt;
On success, you will see:&lt;br /&gt;
 Processing finished in nn secs with no errors reported &lt;br /&gt;
and should see the following subdirectories under the user specified output directory: &lt;br /&gt;
* bams/ &lt;br /&gt;
* Makefiles/ &lt;br /&gt;
* QCFiles/ (if all quality control is not disabled) &lt;br /&gt;
* tmp/ &lt;br /&gt;
&lt;br /&gt;
You should see a &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; for each Sample in the index file. &lt;br /&gt;
&lt;br /&gt;
If you do not see these &amp;lt;code&amp;gt;.OK&amp;lt;/code&amp;gt; files, then your Alignment Pipeline failed. &lt;br /&gt;
&lt;br /&gt;
On success, the bams/ directory contains the final BAMs and bais.&lt;br /&gt;
&lt;br /&gt;
If processing fails part way through, you can pick up where you left off by rerunning gotcloud or the make command.&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7519</id>
		<title>QPLOT</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7519"/>
		<updated>2013-06-17T17:45:17Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Input files */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Introduction =&lt;br /&gt;
&lt;br /&gt;
The qplot program calculates various summary statistics some of which are plotted in a PDF file. These statistics can be used to assess the sequencing quality of sequence reads mapped to the reference genome. The main statistics are empirical Phred scores which are calculated based on the background mismatch rate. Background mismatch rate is the rate that sequenced bases are different from the reference genome, EXCLUDING dbSNP positions. Other statistics include GC biases, insert size distribution, depth distribution, genome coverage, empirical Q20 count, and so on. &lt;br /&gt;
&lt;br /&gt;
In the following sections, we will guide you through: [[#Where to Find It |how to obtain qplot]], [[#Usage |how to use qplot]], [[#Built-in example |example outputs]], [[#anchorOfInteractiveQplot |interactive diagnostic plots]], and [[#Diagnose sequencing quality |real applications]] in which qplot has helped identify sequencing problems.&lt;br /&gt;
&lt;br /&gt;
= Where to Find It =&lt;br /&gt;
&lt;br /&gt;
You can obtain qplot in two ways: &lt;br /&gt;
&lt;br /&gt;
(1) Download the pre-compiled binary along with the source code as described in [[#Binary Download|Binary Download]]. &lt;br /&gt;
&lt;br /&gt;
(2) Download source code only and compile it on your own machine. Please follow the instruction in [[#Source Code Distribution|Source Code Distribution]] on fetching source code and building instructions.&lt;br /&gt;
&lt;br /&gt;
== Binary Download ==&lt;br /&gt;
&lt;br /&gt;
We have prepared a pre-compiled (under Ubuntu) qplot along with source code . You can download it from: [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot.20120602.tar.gz qplot.20120602.tar.gz (File Size: 1.7G)] &lt;br /&gt;
&lt;br /&gt;
The executable file is under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
In addition, we provided the necessary input files under qplot/data/ (NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100).&lt;br /&gt;
&lt;br /&gt;
You can also find an example BAM input file under qplot/example/chrom20.9M.10M.bam. It is taken from the 1000 Genome Project with sequencing reads aligned to chromosome 20 positions 8M to 9M.&lt;br /&gt;
&lt;br /&gt;
== Source Code Distribution ==&lt;br /&gt;
&lt;br /&gt;
We provide a source code only download in [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-source.20120602.tar.gz qplot-source.20120602.tar.gz]. Optionally, you can download example file and/or data file:&lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-example.tar.gz  example]: example input file, and expected outputs if you following the [[#Built-in example | direction]]. &lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-data.tar.gz resources data]: necessary input files for qplot, including NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100.&lt;br /&gt;
&lt;br /&gt;
You can put above file(s) in the same folder and follow these steps:&lt;br /&gt;
&lt;br /&gt;
* 1. Unarchive downloaded file&lt;br /&gt;
 tar zvxf qplot-source.20120602.tar.gz&lt;br /&gt;
&lt;br /&gt;
A new folder &#039;&#039;qplot&#039;&#039; will be created.&lt;br /&gt;
&lt;br /&gt;
* 2. Build libStatGen&lt;br /&gt;
 cd qplot&lt;br /&gt;
 make libStatGen&lt;br /&gt;
&lt;br /&gt;
This step will download a necessary software library [http://genome.sph.umich.edu/wiki/C%2B%2B_Library:_libStatGen libStatGen] and compile source code into a binary code library.&lt;br /&gt;
&lt;br /&gt;
* 3. Build qplot&lt;br /&gt;
 make all&lt;br /&gt;
&lt;br /&gt;
This step will then build qplot. Upon success, the executable qplot can be found under qplot/bin/.&lt;br /&gt;
&lt;br /&gt;
* 4. (Optional) unarchive example and/or data&lt;br /&gt;
 tar zvxf qplot-example.tar.gz&lt;br /&gt;
&lt;br /&gt;
An example file, &#039;&#039;chrom20.9M.10M.bam&#039;&#039;, will be extracted to qplot/example/. It contains ~1.1 million aligned Illumina sequencing reads of NA12878 from 1000 Genome Project. Example command line, &#039;&#039;cmd.sh&#039;&#039;, example outputs, &#039;&#039;qplot.pdf&#039;&#039;, &#039;&#039;qplot.stats&#039;&#039;, and &#039;&#039;qplot.R&#039;&#039; are also provided and will be extracted qplot/example/ as well. &lt;br /&gt;
&lt;br /&gt;
 tar zvxf qplot-data.tar.gz&lt;br /&gt;
&lt;br /&gt;
Three files will be extracted to qplot/data/: &#039;&#039;human.g1k.v37-bs.umfa&#039;&#039; is binary NCBI reference genome build 37; &#039;&#039;dbSNP130.UCSC.coordinates.tbl&#039;&#039; is dbSNP version 130; and &#039;&#039;human.g1k.w100.gc&#039;&#039; is pre-calculated GC content with windows size 100.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Please download source code from [[]], the building &lt;br /&gt;
{{ToolGitRepo|repoName=qplot|noDownload=}}&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Usage =&lt;br /&gt;
&lt;br /&gt;
== Command line ==&lt;br /&gt;
&lt;br /&gt;
After you obtain the qplot executable (either by compiling the source code or by downloading the pre-compiled binary file), you will find the executable file under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
Here is the qplot help page by invoking qplot without any command line arguments:&lt;br /&gt;
&lt;br /&gt;
  some_linux_host &amp;gt; qplot/bin/qplot&lt;br /&gt;
 &lt;br /&gt;
              References : --reference [/net/fantasia/home/zhanxw/software/qplot/data/human.g1k.v37.fa],&lt;br /&gt;
                           --dbsnp [/net/fantasia/home/zhanxw/software/qplot/data/dbSNP130.UCSC.coordinates.tbl],&lt;br /&gt;
                           --gccontent [/net/fantasia/home/zhanxw/software/qplot/data/human.g1k.w100.gc]&lt;br /&gt;
   Create gcContent file : --create_gc [], --winsize [100]&lt;br /&gt;
             Region list : --regions [], --invertRegion&lt;br /&gt;
            Flag filters : --read1_skip, --read2_skip, --paired_skip,&lt;br /&gt;
                           --unpaired_skip&lt;br /&gt;
          Dup and QCFail : --dup_keep, --qcfail_keep&lt;br /&gt;
         Mapping filters : --minMapQuality [0.00]&lt;br /&gt;
      Records to process : --first_n_record [-1]&lt;br /&gt;
        Lanes to process : --lanes []&lt;br /&gt;
   Read group to process : --readGroup []&lt;br /&gt;
      Input file options : --noeof&lt;br /&gt;
            Output files : --plot [], --stats [], --Rcode [], --xml []&lt;br /&gt;
             Plot labels : --label [], --bamLabel []&lt;br /&gt;
&lt;br /&gt;
== Input files ==&lt;br /&gt;
&lt;br /&gt;
qplot runs on the input BAM/SAM file(s) specified on the command-line after all other parameters.&lt;br /&gt;
&lt;br /&gt;
Additionally, three (3) precomputed files are required. &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--reference&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The reference genome is the same as karma reference genome. If the index files do not exist, qplot will create the index files &#039;&#039;&#039;automatically&#039;&#039;&#039; using the input reference fasta file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--dbsnp&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This file has two columns. First column is the chromosome name which must be consistent with the reference created above. Second column is 1-based SNP position. If you want to create your own dbSNP data from downloaded UCSC dbSNP file, one way to do it is: &amp;lt;code&amp;gt;cat dbsnp_129_b36.rod|grep &amp;quot;single&amp;quot; | awk &#039;$4-$3==1&#039; |cut -f2,4 &amp;gt; dbSNP_129_b36.tbl&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--gccontent&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Although GC content can be calculated on the fly each time, it is much more efficient to load a precomputed GC content from a file. To generate the file, use the following command:&lt;br /&gt;
 qplot --reference reference.fa --windowsize winsize --create_gc reference.gc&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Note&#039;&#039;: Before running qplot, it is critical to check how the chromosome names are coded. Some BAM/SAM files use just numbers, others use chr + numbers. &#039;&#039;&#039;You need to make sure that the chromosome names from the reference and dbSNP are consistent with the BAM/SAM files.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Parameters ==&lt;br /&gt;
&lt;br /&gt;
Some of the command line parameters are described here, but most are self explanatory.&lt;br /&gt;
&lt;br /&gt;
*Flag filter&lt;br /&gt;
&lt;br /&gt;
By default all reads are processed. If it is desired to check only the first read of a pair, use &amp;lt;code&amp;gt;--read2_skip&amp;lt;/code&amp;gt; to ignore the second read. And so on.&lt;br /&gt;
&lt;br /&gt;
*Duplication and QCFail&lt;br /&gt;
&lt;br /&gt;
By default reads marked as duplication and QCFail are ignored but can be retained by &lt;br /&gt;
 --dup_keep &lt;br /&gt;
or &lt;br /&gt;
 --qcfail_keep&lt;br /&gt;
&lt;br /&gt;
*Records to process &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--first_n_record&amp;lt;/code&amp;gt; option followed by a number, &#039;&#039;&#039;n&#039;&#039;&#039;, will enable qplot to read the first &#039;&#039;&#039;n&#039;&#039;&#039; reads to test the bam files and verify it works.&lt;br /&gt;
&lt;br /&gt;
* Lanes to process (only works for Illumina sequences)&lt;br /&gt;
&lt;br /&gt;
If the input bam files have more than one lane and only some of them need to be checked, use something like &amp;lt;code&amp;gt;--lanes 1,3,5&amp;lt;/code&amp;gt; to specify that only lanes 1, 3, and 5 need to be checked.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTE&#039;&#039;&#039; In order for this to work, the lane info has to be encoded in the read name such that the lane number is the second field with the delimiter &amp;quot;:&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
* Read group to process : &lt;br /&gt;
&lt;br /&gt;
Read group option can restrict qplot to process a subset of reads. For example, if BAM contain the following @RG tags:&lt;br /&gt;
&lt;br /&gt;
 @RG	ID:UM0348_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
&lt;br /&gt;
If specify nothing or not using &amp;quot;--readGroup&amp;quot;, QPLOT by default will process all reads; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348&amp;quot;, then only read group UM0348_1, UM_0348_2, UM_0348_3, UM_0348_4 will be processed; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348_1&amp;quot;, then only one read group UM0348_1 will be processed.&lt;br /&gt;
&lt;br /&gt;
* Input file options :&lt;br /&gt;
&lt;br /&gt;
BAM files are compress by BGZF algorithm and it should contain EOF by default. QPLOT will by default stop working when it does not found a valid EOF tag inside BAM files. &lt;br /&gt;
However, you can force QPLOT to continue process using --noeof. But you should be award the input files may be corrupted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Mapping filters&lt;br /&gt;
&lt;br /&gt;
Qplot will exclude reads with lower mapping qualities than the user specified parameter, &amp;lt;code&amp;gt;--minMapQuality&amp;lt;/code&amp;gt;. By default, mapped reads with all mapping quality will be included in the analysis.&lt;br /&gt;
&lt;br /&gt;
*Region list&lt;br /&gt;
&lt;br /&gt;
If the interest of qplot is a list of regions, e.g. exons, this can be achieved by providing a list of regions. The regions should be in the form of &amp;quot;chr start end label&amp;quot; each line in the file (NOTE: &#039;&#039;start&#039;&#039; and &#039;&#039;end&#039;&#039; position are inclusive and they follow the convention of [http://genome.ucsc.edu/FAQ/FAQformat#format1 BED file]). &lt;br /&gt;
In order for this option to work, within each chromosome (contig) the regions have to be sorted by starting position, and also the input bam files have to be sorted. &lt;br /&gt;
For example, you can create a text file, region.txt like following:&lt;br /&gt;
&lt;br /&gt;
 1 100 500 region_A&lt;br /&gt;
 1 600 800 region_B&lt;br /&gt;
 2 100 300 region_C&lt;br /&gt;
 &lt;br /&gt;
Then specifying &amp;lt;code&amp;gt; --regions region.txt&amp;lt;/code&amp;gt; enables qplot to calculate various statistics out of sequenced bases only within the above 3 regions.&lt;br /&gt;
&lt;br /&gt;
Qplot also provides the &amp;lt;code&amp;gt;--invertRegion&amp;lt;/code&amp;gt; option. Enabling this option tells qplot to operate on those sequence bases that are outside the given region.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Plot labels&lt;br /&gt;
&lt;br /&gt;
Two kinds of labels are enabled. &amp;lt;code&amp;gt;--label&amp;lt;/code&amp;gt; is the label for the plot (default is empty) which is appended to the title of each subplot. &amp;lt;code&amp;gt;--bamLabels&amp;lt;/code&amp;gt; followed by a column separated list of labels provides the labels for each input SAM/BAM file, e.g. sample ID (default is numbers 1, 2, ... until the number of input bam files). For example:&lt;br /&gt;
 --label Run100 --bamLabels s1,s2,s3,s4,s5,s6,s7,s8&lt;br /&gt;
&lt;br /&gt;
== Output files ==&lt;br /&gt;
&lt;br /&gt;
There are three (optional) output files.&lt;br /&gt;
* &amp;lt;code&amp;gt;--plot &#039;&#039;qa.pdf&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a PDF file named &#039;&#039;qa.pdf&#039;&#039; containing 2 pages each with 4 figures. The plot is generated using Rscript.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--stats &#039;&#039;qa.stats&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a text file named &#039;&#039;qa.stats&#039;&#039; containing various summary statistics for each input BAM/SAM file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--Rcode &#039;&#039;qa.R&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate &#039;&#039;qa.R&#039;&#039; which is the R code used for plotting the figures in the &#039;&#039;qa.pdf&#039;&#039; file. If Rscript is not installed in the system, you can use the qa.R to generate the figures on other machines, or extract plotting data from each run and combine multiple runs together to generate more comprehensive plots (See [[#Example | Example]]).&lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
Qplot can generate diagnostic graphs, related R code, and summary statistics for each SAM/BAM file.&lt;br /&gt;
&lt;br /&gt;
== Built-in example ==&lt;br /&gt;
&lt;br /&gt;
In the pre-compiled binary download, you will find a subdirectory named examples. We provide a sample file from the 1000 Genome project, it contains aligned reads on chromosome 20 from position 8 Mbp to 9Mbp. You can invoke qplot using the following command line:&lt;br /&gt;
&lt;br /&gt;
 ../bin/qplot --reference ../data/human.g1k.v37.umfa --dbsnp ../data/dbSNP130.UCSC.coordinates.tbl --gccontent ../data/human.g1k.w100.gc --plot qplot.pdf --stats qplot.stats --Rcode qplot.R --label &amp;quot;chr20:9M-10M&amp;quot; chrom20.9M.10M.bam&lt;br /&gt;
&lt;br /&gt;
Sample outputs are listed below:&lt;br /&gt;
&lt;br /&gt;
1) Figure: [[Media:qplot.pdf | qplot.pdf]]&lt;br /&gt;
&lt;br /&gt;
2) Summary statistics:&lt;br /&gt;
 Stats\BAM       chrom20.9M.10M.bam&lt;br /&gt;
 TotalReads(e6)  1.11&lt;br /&gt;
 MappingRate(%)  97.24&lt;br /&gt;
 MapRate_MQpass(%)       97.24&lt;br /&gt;
 TargetMapping(%)        0.00&lt;br /&gt;
 ZeroMapQual(%)  2.39&lt;br /&gt;
 MapQual&amp;lt;10(%)   2.86&lt;br /&gt;
 PairedReads(%)  83.76&lt;br /&gt;
 ProperPaired(%) 71.34&lt;br /&gt;
 MappedBases(e9) 0.04&lt;br /&gt;
 Q20Bases(e9)    0.04&lt;br /&gt;
 Q20BasesPct(%)  88.63&lt;br /&gt;
 MeanDepth       42.22&lt;br /&gt;
 GenomeCover(%)  0.03&lt;br /&gt;
 EPS_MSE 1.81&lt;br /&gt;
 EPS_Cycle_Mean  18.71&lt;br /&gt;
 GCBiasMSE       0.01&lt;br /&gt;
 ISize_mode      137&lt;br /&gt;
 ISize_medium    184&lt;br /&gt;
 DupRate(%)      5.90&lt;br /&gt;
 QCFailRate(%)   0.00&lt;br /&gt;
 BaseComp_A(%)   29.9&lt;br /&gt;
 BaseComp_C(%)   20.1&lt;br /&gt;
 BaseComp_G(%)   20.2&lt;br /&gt;
 BaseComp_T(%)   29.8&lt;br /&gt;
 BaseComp_O(%)   0.1&lt;br /&gt;
&lt;br /&gt;
== Gallery of examples ==&lt;br /&gt;
&lt;br /&gt;
Here we show qplot can be applied in various sequencing scenarios. Also users can customize statistics generated by qplot to their needs.&lt;br /&gt;
&lt;br /&gt;
* Whole genome sequencing with 24-multiplexing&lt;br /&gt;
&lt;br /&gt;
With a customized script, we aggregated 24 bar-coded samples in the same graph.&lt;br /&gt;
The graph will help compare sequencing quality between samples. &lt;br /&gt;
&lt;br /&gt;
[[Media: qplot.Pool.9847.pdf | QPlot of 24 samples(PDF) ]]&lt;br /&gt;
&lt;br /&gt;
* Interactive qplot &lt;br /&gt;
&lt;br /&gt;
&amp;lt;span id=&amp;quot;anchorOfInteractiveQplot&amp;quot;&amp;gt;&amp;lt;/span&amp;gt;&lt;br /&gt;
Qplot can be interactive. In the following example, you can use mouse scroll to zoom in and zoom out on each graph and pan to a certain part of the graph.&lt;br /&gt;
By presenting qplot data on a web page, users can easily identify problematic sequencing samples. Users of qplot can customize its outputs into web page format greatly easing the data exploring process.&lt;br /&gt;
&lt;br /&gt;
[http://www-personal.umich.edu/~zhanxw/qplot.Pool.9847.html  QPlot of 24 samples(HTML) ]&lt;br /&gt;
&lt;br /&gt;
== Diagnose sequencing quality ==&lt;br /&gt;
&lt;br /&gt;
Qplot is designed and implemented for the need of checking sequencing quality. &lt;br /&gt;
Besides the example of analyzing RNA-seq data as shown in our manuscript, &lt;br /&gt;
here we demonstrate two additional scenarios in which qplot can help identify problems after obtaining sequencing data. &lt;br /&gt;
&lt;br /&gt;
* Base quality distributed abnormally&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBaseQual.pdf | Example of qplot helping to identify wrong phred base quality]]&lt;br /&gt;
&lt;br /&gt;
By checking the first graph &amp;quot;Empirical vs reported Phred score&amp;quot;, we found reported base qualities are shifted to the right.&lt;br /&gt;
In this particular example, &#039;33&#039; was incorrectly added to all base qualities. &lt;br /&gt;
When such data used in variant calling, we may increase false positive SNP variants.&lt;br /&gt;
&lt;br /&gt;
* Bar-coded samples&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBarCoding.pdf | Example of qplot identifying the effect of ignoring bar-coding]]&lt;br /&gt;
&lt;br /&gt;
By checking &amp;quot;Empirical phred score by cycle&amp;quot; (top right graph on the first page), we noticed the empirical qualities in the first several cycles are abnormally low. This phenomenon leads us to hypothesize that the first several bases have different properties. Further investigation confirmed that this sequencing was done using bar-coded DNA samples, but the analysis did not properly de-multiplex each sample.&lt;br /&gt;
&lt;br /&gt;
= Contact =&lt;br /&gt;
&lt;br /&gt;
Questions and requests should be sent to Bingshan Li ([mailto:bingshan@umich.edu bingshan@umich.edu]) or Xiaowei Zhan ([mailto:zhanxw@umich.edu zhanxw@umich.edu]) or Goncalo Abecasis ([mailto:goncalo@umich.edu goncalo@umich.edu])&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7518</id>
		<title>QPLOT</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=QPLOT&amp;diff=7518"/>
		<updated>2013-06-17T17:44:58Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Input files */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Introduction =&lt;br /&gt;
&lt;br /&gt;
The qplot program calculates various summary statistics some of which are plotted in a PDF file. These statistics can be used to assess the sequencing quality of sequence reads mapped to the reference genome. The main statistics are empirical Phred scores which are calculated based on the background mismatch rate. Background mismatch rate is the rate that sequenced bases are different from the reference genome, EXCLUDING dbSNP positions. Other statistics include GC biases, insert size distribution, depth distribution, genome coverage, empirical Q20 count, and so on. &lt;br /&gt;
&lt;br /&gt;
In the following sections, we will guide you through: [[#Where to Find It |how to obtain qplot]], [[#Usage |how to use qplot]], [[#Built-in example |example outputs]], [[#anchorOfInteractiveQplot |interactive diagnostic plots]], and [[#Diagnose sequencing quality |real applications]] in which qplot has helped identify sequencing problems.&lt;br /&gt;
&lt;br /&gt;
= Where to Find It =&lt;br /&gt;
&lt;br /&gt;
You can obtain qplot in two ways: &lt;br /&gt;
&lt;br /&gt;
(1) Download the pre-compiled binary along with the source code as described in [[#Binary Download|Binary Download]]. &lt;br /&gt;
&lt;br /&gt;
(2) Download source code only and compile it on your own machine. Please follow the instruction in [[#Source Code Distribution|Source Code Distribution]] on fetching source code and building instructions.&lt;br /&gt;
&lt;br /&gt;
== Binary Download ==&lt;br /&gt;
&lt;br /&gt;
We have prepared a pre-compiled (under Ubuntu) qplot along with source code . You can download it from: [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot.20120602.tar.gz qplot.20120602.tar.gz (File Size: 1.7G)] &lt;br /&gt;
&lt;br /&gt;
The executable file is under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
In addition, we provided the necessary input files under qplot/data/ (NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100).&lt;br /&gt;
&lt;br /&gt;
You can also find an example BAM input file under qplot/example/chrom20.9M.10M.bam. It is taken from the 1000 Genome Project with sequencing reads aligned to chromosome 20 positions 8M to 9M.&lt;br /&gt;
&lt;br /&gt;
== Source Code Distribution ==&lt;br /&gt;
&lt;br /&gt;
We provide a source code only download in [http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-source.20120602.tar.gz qplot-source.20120602.tar.gz]. Optionally, you can download example file and/or data file:&lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-example.tar.gz  example]: example input file, and expected outputs if you following the [[#Built-in example | direction]]. &lt;br /&gt;
&lt;br /&gt;
[http://www.sph.umich.edu/csg/zhanxw/software/qplot/qplot-data.tar.gz resources data]: necessary input files for qplot, including NCBI human genome build v37, dbSNP 130, and pre-computed GC file with windows size 100.&lt;br /&gt;
&lt;br /&gt;
You can put above file(s) in the same folder and follow these steps:&lt;br /&gt;
&lt;br /&gt;
* 1. Unarchive downloaded file&lt;br /&gt;
 tar zvxf qplot-source.20120602.tar.gz&lt;br /&gt;
&lt;br /&gt;
A new folder &#039;&#039;qplot&#039;&#039; will be created.&lt;br /&gt;
&lt;br /&gt;
* 2. Build libStatGen&lt;br /&gt;
 cd qplot&lt;br /&gt;
 make libStatGen&lt;br /&gt;
&lt;br /&gt;
This step will download a necessary software library [http://genome.sph.umich.edu/wiki/C%2B%2B_Library:_libStatGen libStatGen] and compile source code into a binary code library.&lt;br /&gt;
&lt;br /&gt;
* 3. Build qplot&lt;br /&gt;
 make all&lt;br /&gt;
&lt;br /&gt;
This step will then build qplot. Upon success, the executable qplot can be found under qplot/bin/.&lt;br /&gt;
&lt;br /&gt;
* 4. (Optional) unarchive example and/or data&lt;br /&gt;
 tar zvxf qplot-example.tar.gz&lt;br /&gt;
&lt;br /&gt;
An example file, &#039;&#039;chrom20.9M.10M.bam&#039;&#039;, will be extracted to qplot/example/. It contains ~1.1 million aligned Illumina sequencing reads of NA12878 from 1000 Genome Project. Example command line, &#039;&#039;cmd.sh&#039;&#039;, example outputs, &#039;&#039;qplot.pdf&#039;&#039;, &#039;&#039;qplot.stats&#039;&#039;, and &#039;&#039;qplot.R&#039;&#039; are also provided and will be extracted qplot/example/ as well. &lt;br /&gt;
&lt;br /&gt;
 tar zvxf qplot-data.tar.gz&lt;br /&gt;
&lt;br /&gt;
Three files will be extracted to qplot/data/: &#039;&#039;human.g1k.v37-bs.umfa&#039;&#039; is binary NCBI reference genome build 37; &#039;&#039;dbSNP130.UCSC.coordinates.tbl&#039;&#039; is dbSNP version 130; and &#039;&#039;human.g1k.w100.gc&#039;&#039; is pre-calculated GC content with windows size 100.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;!-- Please download source code from [[]], the building &lt;br /&gt;
{{ToolGitRepo|repoName=qplot|noDownload=}}&lt;br /&gt;
--&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Usage =&lt;br /&gt;
&lt;br /&gt;
== Command line ==&lt;br /&gt;
&lt;br /&gt;
After you obtain the qplot executable (either by compiling the source code or by downloading the pre-compiled binary file), you will find the executable file under qplot/bin/qplot. &lt;br /&gt;
&lt;br /&gt;
Here is the qplot help page by invoking qplot without any command line arguments:&lt;br /&gt;
&lt;br /&gt;
  some_linux_host &amp;gt; qplot/bin/qplot&lt;br /&gt;
 &lt;br /&gt;
              References : --reference [/net/fantasia/home/zhanxw/software/qplot/data/human.g1k.v37.fa],&lt;br /&gt;
                           --dbsnp [/net/fantasia/home/zhanxw/software/qplot/data/dbSNP130.UCSC.coordinates.tbl],&lt;br /&gt;
                           --gccontent [/net/fantasia/home/zhanxw/software/qplot/data/human.g1k.w100.gc]&lt;br /&gt;
   Create gcContent file : --create_gc [], --winsize [100]&lt;br /&gt;
             Region list : --regions [], --invertRegion&lt;br /&gt;
            Flag filters : --read1_skip, --read2_skip, --paired_skip,&lt;br /&gt;
                           --unpaired_skip&lt;br /&gt;
          Dup and QCFail : --dup_keep, --qcfail_keep&lt;br /&gt;
         Mapping filters : --minMapQuality [0.00]&lt;br /&gt;
      Records to process : --first_n_record [-1]&lt;br /&gt;
        Lanes to process : --lanes []&lt;br /&gt;
   Read group to process : --readGroup []&lt;br /&gt;
      Input file options : --noeof&lt;br /&gt;
            Output files : --plot [], --stats [], --Rcode [], --xml []&lt;br /&gt;
             Plot labels : --label [], --bamLabel []&lt;br /&gt;
&lt;br /&gt;
== Input files ==&lt;br /&gt;
&lt;br /&gt;
qplot runs on the input BAM/SAM file(s) specified on the command-line after all other parameters.&lt;br /&gt;
&lt;br /&gt;
Additionally, three (3) precomputed files are required. &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--reference&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The reference genome is the same as karma reference genome. If the index files do not exist, qplot will create the index files &#039;&#039;&#039;automatically&#039;&#039;&#039; using the input reference fasta file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--dbsnp&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This file has two columns. First column is the chromosome name which must be consistent with the reference created above. Second column is 1-based SNP position. If you want to create your own dbSNP data from downloaded UCSC dbSNP file, one way to do it is: &amp;lt;code&amp;gt;cat dbsnp_129_b36.rod|grep &amp;quot;single&amp;quot; | awk &#039;$4-$3==1&#039; |cut -f2,4 &amp;gt; dbSNP_129_b36.tbl&amp;lt;/code&amp;gt; &lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--gccontent&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Although GC content can be calculated on the fly each time, it is much more efficient to load a precomputed GC content from a file. To generate the file, use the following command:&lt;br /&gt;
 qplot --refence reference.fa --windowsize winsize --create_gc reference.gc&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Note&#039;&#039;: Before running qplot, it is critical to check how the chromosome names are coded. Some BAM/SAM files use just numbers, others use chr + numbers. &#039;&#039;&#039;You need to make sure that the chromosome names from the reference and dbSNP are consistent with the BAM/SAM files.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Parameters ==&lt;br /&gt;
&lt;br /&gt;
Some of the command line parameters are described here, but most are self explanatory.&lt;br /&gt;
&lt;br /&gt;
*Flag filter&lt;br /&gt;
&lt;br /&gt;
By default all reads are processed. If it is desired to check only the first read of a pair, use &amp;lt;code&amp;gt;--read2_skip&amp;lt;/code&amp;gt; to ignore the second read. And so on.&lt;br /&gt;
&lt;br /&gt;
*Duplication and QCFail&lt;br /&gt;
&lt;br /&gt;
By default reads marked as duplication and QCFail are ignored but can be retained by &lt;br /&gt;
 --dup_keep &lt;br /&gt;
or &lt;br /&gt;
 --qcfail_keep&lt;br /&gt;
&lt;br /&gt;
*Records to process &lt;br /&gt;
&lt;br /&gt;
The &amp;lt;code&amp;gt;--first_n_record&amp;lt;/code&amp;gt; option followed by a number, &#039;&#039;&#039;n&#039;&#039;&#039;, will enable qplot to read the first &#039;&#039;&#039;n&#039;&#039;&#039; reads to test the bam files and verify it works.&lt;br /&gt;
&lt;br /&gt;
* Lanes to process (only works for Illumina sequences)&lt;br /&gt;
&lt;br /&gt;
If the input bam files have more than one lane and only some of them need to be checked, use something like &amp;lt;code&amp;gt;--lanes 1,3,5&amp;lt;/code&amp;gt; to specify that only lanes 1, 3, and 5 need to be checked.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;NOTE&#039;&#039;&#039; In order for this to work, the lane info has to be encoded in the read name such that the lane number is the second field with the delimiter &amp;quot;:&amp;quot;.&lt;br /&gt;
&lt;br /&gt;
* Read group to process : &lt;br /&gt;
&lt;br /&gt;
Read group option can restrict qplot to process a subset of reads. For example, if BAM contain the following @RG tags:&lt;br /&gt;
&lt;br /&gt;
 @RG	ID:UM0348_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0348_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_1:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_2:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_3:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
 @RG	ID:UM0360_4:1	PL:ILLUMINA	LB:M5390	SM:M5390	CN:UM&lt;br /&gt;
&lt;br /&gt;
If specify nothing or not using &amp;quot;--readGroup&amp;quot;, QPLOT by default will process all reads; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348&amp;quot;, then only read group UM0348_1, UM_0348_2, UM_0348_3, UM_0348_4 will be processed; &lt;br /&gt;
If specify &amp;quot;--readGroup UM0348_1&amp;quot;, then only one read group UM0348_1 will be processed.&lt;br /&gt;
&lt;br /&gt;
* Input file options :&lt;br /&gt;
&lt;br /&gt;
BAM files are compress by BGZF algorithm and it should contain EOF by default. QPLOT will by default stop working when it does not found a valid EOF tag inside BAM files. &lt;br /&gt;
However, you can force QPLOT to continue process using --noeof. But you should be award the input files may be corrupted.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Mapping filters&lt;br /&gt;
&lt;br /&gt;
Qplot will exclude reads with lower mapping qualities than the user specified parameter, &amp;lt;code&amp;gt;--minMapQuality&amp;lt;/code&amp;gt;. By default, mapped reads with all mapping quality will be included in the analysis.&lt;br /&gt;
&lt;br /&gt;
*Region list&lt;br /&gt;
&lt;br /&gt;
If the interest of qplot is a list of regions, e.g. exons, this can be achieved by providing a list of regions. The regions should be in the form of &amp;quot;chr start end label&amp;quot; each line in the file (NOTE: &#039;&#039;start&#039;&#039; and &#039;&#039;end&#039;&#039; position are inclusive and they follow the convention of [http://genome.ucsc.edu/FAQ/FAQformat#format1 BED file]). &lt;br /&gt;
In order for this option to work, within each chromosome (contig) the regions have to be sorted by starting position, and also the input bam files have to be sorted. &lt;br /&gt;
For example, you can create a text file, region.txt like following:&lt;br /&gt;
&lt;br /&gt;
 1 100 500 region_A&lt;br /&gt;
 1 600 800 region_B&lt;br /&gt;
 2 100 300 region_C&lt;br /&gt;
 &lt;br /&gt;
Then specifying &amp;lt;code&amp;gt; --regions region.txt&amp;lt;/code&amp;gt; enables qplot to calculate various statistics out of sequenced bases only within the above 3 regions.&lt;br /&gt;
&lt;br /&gt;
Qplot also provides the &amp;lt;code&amp;gt;--invertRegion&amp;lt;/code&amp;gt; option. Enabling this option tells qplot to operate on those sequence bases that are outside the given region.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* Plot labels&lt;br /&gt;
&lt;br /&gt;
Two kinds of labels are enabled. &amp;lt;code&amp;gt;--label&amp;lt;/code&amp;gt; is the label for the plot (default is empty) which is appended to the title of each subplot. &amp;lt;code&amp;gt;--bamLabels&amp;lt;/code&amp;gt; followed by a column separated list of labels provides the labels for each input SAM/BAM file, e.g. sample ID (default is numbers 1, 2, ... until the number of input bam files). For example:&lt;br /&gt;
 --label Run100 --bamLabels s1,s2,s3,s4,s5,s6,s7,s8&lt;br /&gt;
&lt;br /&gt;
== Output files ==&lt;br /&gt;
&lt;br /&gt;
There are three (optional) output files.&lt;br /&gt;
* &amp;lt;code&amp;gt;--plot &#039;&#039;qa.pdf&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a PDF file named &#039;&#039;qa.pdf&#039;&#039; containing 2 pages each with 4 figures. The plot is generated using Rscript.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--stats &#039;&#039;qa.stats&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate a text file named &#039;&#039;qa.stats&#039;&#039; containing various summary statistics for each input BAM/SAM file.&lt;br /&gt;
&lt;br /&gt;
* &amp;lt;code&amp;gt;--Rcode &#039;&#039;qa.R&#039;&#039;&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Qplot will generate &#039;&#039;qa.R&#039;&#039; which is the R code used for plotting the figures in the &#039;&#039;qa.pdf&#039;&#039; file. If Rscript is not installed in the system, you can use the qa.R to generate the figures on other machines, or extract plotting data from each run and combine multiple runs together to generate more comprehensive plots (See [[#Example | Example]]).&lt;br /&gt;
&lt;br /&gt;
= Example =&lt;br /&gt;
&lt;br /&gt;
Qplot can generate diagnostic graphs, related R code, and summary statistics for each SAM/BAM file.&lt;br /&gt;
&lt;br /&gt;
== Built-in example ==&lt;br /&gt;
&lt;br /&gt;
In the pre-compiled binary download, you will find a subdirectory named examples. We provide a sample file from the 1000 Genome project, it contains aligned reads on chromosome 20 from position 8 Mbp to 9Mbp. You can invoke qplot using the following command line:&lt;br /&gt;
&lt;br /&gt;
 ../bin/qplot --reference ../data/human.g1k.v37.umfa --dbsnp ../data/dbSNP130.UCSC.coordinates.tbl --gccontent ../data/human.g1k.w100.gc --plot qplot.pdf --stats qplot.stats --Rcode qplot.R --label &amp;quot;chr20:9M-10M&amp;quot; chrom20.9M.10M.bam&lt;br /&gt;
&lt;br /&gt;
Sample outputs are listed below:&lt;br /&gt;
&lt;br /&gt;
1) Figure: [[Media:qplot.pdf | qplot.pdf]]&lt;br /&gt;
&lt;br /&gt;
2) Summary statistics:&lt;br /&gt;
 Stats\BAM       chrom20.9M.10M.bam&lt;br /&gt;
 TotalReads(e6)  1.11&lt;br /&gt;
 MappingRate(%)  97.24&lt;br /&gt;
 MapRate_MQpass(%)       97.24&lt;br /&gt;
 TargetMapping(%)        0.00&lt;br /&gt;
 ZeroMapQual(%)  2.39&lt;br /&gt;
 MapQual&amp;lt;10(%)   2.86&lt;br /&gt;
 PairedReads(%)  83.76&lt;br /&gt;
 ProperPaired(%) 71.34&lt;br /&gt;
 MappedBases(e9) 0.04&lt;br /&gt;
 Q20Bases(e9)    0.04&lt;br /&gt;
 Q20BasesPct(%)  88.63&lt;br /&gt;
 MeanDepth       42.22&lt;br /&gt;
 GenomeCover(%)  0.03&lt;br /&gt;
 EPS_MSE 1.81&lt;br /&gt;
 EPS_Cycle_Mean  18.71&lt;br /&gt;
 GCBiasMSE       0.01&lt;br /&gt;
 ISize_mode      137&lt;br /&gt;
 ISize_medium    184&lt;br /&gt;
 DupRate(%)      5.90&lt;br /&gt;
 QCFailRate(%)   0.00&lt;br /&gt;
 BaseComp_A(%)   29.9&lt;br /&gt;
 BaseComp_C(%)   20.1&lt;br /&gt;
 BaseComp_G(%)   20.2&lt;br /&gt;
 BaseComp_T(%)   29.8&lt;br /&gt;
 BaseComp_O(%)   0.1&lt;br /&gt;
&lt;br /&gt;
== Gallery of examples ==&lt;br /&gt;
&lt;br /&gt;
Here we show qplot can be applied in various sequencing scenarios. Also users can customize statistics generated by qplot to their needs.&lt;br /&gt;
&lt;br /&gt;
* Whole genome sequencing with 24-multiplexing&lt;br /&gt;
&lt;br /&gt;
With a customized script, we aggregated 24 bar-coded samples in the same graph.&lt;br /&gt;
The graph will help compare sequencing quality between samples. &lt;br /&gt;
&lt;br /&gt;
[[Media: qplot.Pool.9847.pdf | QPlot of 24 samples(PDF) ]]&lt;br /&gt;
&lt;br /&gt;
* Interactive qplot &lt;br /&gt;
&lt;br /&gt;
&amp;lt;span id=&amp;quot;anchorOfInteractiveQplot&amp;quot;&amp;gt;&amp;lt;/span&amp;gt;&lt;br /&gt;
Qplot can be interactive. In the following example, you can use mouse scroll to zoom in and zoom out on each graph and pan to a certain part of the graph.&lt;br /&gt;
By presenting qplot data on a web page, users can easily identify problematic sequencing samples. Users of qplot can customize its outputs into web page format greatly easing the data exploring process.&lt;br /&gt;
&lt;br /&gt;
[http://www-personal.umich.edu/~zhanxw/qplot.Pool.9847.html  QPlot of 24 samples(HTML) ]&lt;br /&gt;
&lt;br /&gt;
== Diagnose sequencing quality ==&lt;br /&gt;
&lt;br /&gt;
Qplot is designed and implemented for the need of checking sequencing quality. &lt;br /&gt;
Besides the example of analyzing RNA-seq data as shown in our manuscript, &lt;br /&gt;
here we demonstrate two additional scenarios in which qplot can help identify problems after obtaining sequencing data. &lt;br /&gt;
&lt;br /&gt;
* Base quality distributed abnormally&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBaseQual.pdf | Example of qplot helping to identify wrong phred base quality]]&lt;br /&gt;
&lt;br /&gt;
By checking the first graph &amp;quot;Empirical vs reported Phred score&amp;quot;, we found reported base qualities are shifted to the right.&lt;br /&gt;
In this particular example, &#039;33&#039; was incorrectly added to all base qualities. &lt;br /&gt;
When such data used in variant calling, we may increase false positive SNP variants.&lt;br /&gt;
&lt;br /&gt;
* Bar-coded samples&lt;br /&gt;
&lt;br /&gt;
[[Media: WrongBarCoding.pdf | Example of qplot identifying the effect of ignoring bar-coding]]&lt;br /&gt;
&lt;br /&gt;
By checking &amp;quot;Empirical phred score by cycle&amp;quot; (top right graph on the first page), we noticed the empirical qualities in the first several cycles are abnormally low. This phenomenon leads us to hypothesize that the first several bases have different properties. Further investigation confirmed that this sequencing was done using bar-coded DNA samples, but the analysis did not properly de-multiplex each sample.&lt;br /&gt;
&lt;br /&gt;
= Contact =&lt;br /&gt;
&lt;br /&gt;
Questions and requests should be sent to Bingshan Li ([mailto:bingshan@umich.edu bingshan@umich.edu]) or Xiaowei Zhan ([mailto:zhanxw@umich.edu zhanxw@umich.edu]) or Goncalo Abecasis ([mailto:goncalo@umich.edu goncalo@umich.edu])&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7508</id>
		<title>Tutorial: GotCloud</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7508"/>
		<updated>2013-06-16T23:03:49Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* Frequently Asked Questions (FAQs) */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= GotCloud Tutorial =&lt;br /&gt;
In this tutorial, we illustrate some of the essential steps in the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
For a background on GotCloud and Sequence Analysis Pipelines, see [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
While GotCloud can run on a cluster of machines or instances, this tutorial is just a small test that just runs on the machine the commands are run on.&lt;br /&gt;
&lt;br /&gt;
GotCloud and this basic tutorial were presented at the [http://ibg.colorado.edu/dokuwiki/doku.php?id=workshop:2013:announcement 2013 IBG Workshop].  It was presented in two sessions.  On Wednesday an overview was presented with steps for running the tutorial data: [[Media:IBG2013GotCloud.pdf|IBG2013GotCloud.pdf]].  On Friday more detail on the input files and what goes into generating the input files was presented: [[Media:GotCloudIBGWorkshop2013Friday.pdf|GotCloudIBGWorkshop2013Friday.pdf]].&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;This tutorial is in the process of being updated for gotcloud version 1.06 (April 17. 2013).&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== STEP 1 : Setup GotCloud ==&lt;br /&gt;
&lt;br /&gt;
[[GotCloud]] has been developed and tested on Linux Ubuntu 12.10 and 12.04.2 LTS but has not been tested on other Linux operating systems. It is not available for Windows. If you do not have your own set of machines to run on, GotCloud is also available for Ubuntu running on the Amazon Elastic Compute Cloud, see [[Amazon_Snapshot]] for more information.&lt;br /&gt;
&lt;br /&gt;
We will use 3 different directories for this tutorial:&lt;br /&gt;
# path to the directory where gotcloud is installed, default ~/gotcloud/&lt;br /&gt;
# path to the directory where the example data is installed, default ~/gotcloudExample&lt;br /&gt;
# path to your output directory, default ~/gotcloudTutorialOut/&lt;br /&gt;
&lt;br /&gt;
If the directories specified above do not reflect the directories you would like to use, replace their occurrances in the instructions below with the appropriate paths.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Step 1a: Install GotCloud ===&lt;br /&gt;
In order to run this tutorial, you need to make sure you have GotCloud installed on your system.  &lt;br /&gt;
&lt;br /&gt;
If you have root and would like to install gotcloud on your system, follow: [[GotCloud#Install_GotCloud_Software| root access installation instructions]]&lt;br /&gt;
&lt;br /&gt;
Otherwise, you can install it in your own directory:&lt;br /&gt;
# Change to the directory where you want gotcloud/ installed&lt;br /&gt;
# Download the gotcloud tar from the ftp site.&lt;br /&gt;
# Extract the tar&lt;br /&gt;
# Build (compile) the source&lt;br /&gt;
#* Note: as the source builds, many messages will scroll through your terminal.  You may even see some warnings.  These messages are normal and expected.  As long as the build does not end with an error, you have successfully built the source.&lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloud_latest.tgz  # Download&lt;br /&gt;
 tar xf gotcloud_latest.tgz     # Extracts into gotcloud/&lt;br /&gt;
 cd ~/gotcloud/src; make         # Build source&lt;br /&gt;
 &lt;br /&gt;
GotCloud requires the following tools to be installed.&lt;br /&gt;
You can run ~/gotcloud/scripts/check_requirements.sh&lt;br /&gt;
...TBD – put in required programs/tools.&lt;br /&gt;
* java (java-common default-jre on ubuntu)&lt;br /&gt;
* make (make on ubuntu)&lt;br /&gt;
* libssl (libssl0.9.8 on ubuntu)&lt;br /&gt;
* gcc 4.4 or newer&lt;br /&gt;
&lt;br /&gt;
=== Step 1b: Install Example Dataset ===&lt;br /&gt;
Our dataset consists of 60 individuals from Great Britain (GBR) sequenced by the 1000 Genomes Project. These individuals have been sequenced to an average depth of about 4x.&lt;br /&gt;
&lt;br /&gt;
To conserve time and disk-space, our analysis will focus on a small region on chromosome 20, 42900000 - 43200000. &lt;br /&gt;
&lt;br /&gt;
The tutorial will run the alignment pipeline on 2 of the individuals (HG00096, HG00100).  The fastqs used for this step are reduced to reads that fall into our target region.&lt;br /&gt;
&lt;br /&gt;
The tutorial will then used previously aligned/mapped reads for the full 60 individuals to generate a list of polymorphic sites and estimate accurate genotypes at each of these sites. &lt;br /&gt;
&lt;br /&gt;
The example dataset we&#039;ll be using is available at: ftp://share.sph.umich.edu/gotcloud/gotcloudExample.tgz &lt;br /&gt;
&lt;br /&gt;
# Change directory to where you want to install the Tutorial data &lt;br /&gt;
# Download the dataset tar from the ftp site &lt;br /&gt;
# Extract the tar &lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloudExample_latest.tgz  # Download &lt;br /&gt;
 tar xvf gotcloudExample_latest.tgz    # Extracts into gotcloudExample/&lt;br /&gt;
&lt;br /&gt;
== STEP 2 : Run GotCloud Alignment Pipeline == &lt;br /&gt;
The first step in processing next generation sequence data is mapping the reads to the reference genome, generating per sample BAM files. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Align the fastqs to the reference genome &lt;br /&gt;
#* handles both single &amp;amp; paired end &lt;br /&gt;
# Merge the results from multiple fastqs into 1 file per sample &lt;br /&gt;
# Mark Duplicate Reads are marked &lt;br /&gt;
# Recalibrate Base Qualities &lt;br /&gt;
&lt;br /&gt;
This processing results in 1 BAM file per sample. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline also includes Quality Control (QC) steps: &lt;br /&gt;
# Visualization of various quality measures (QPLOT) &lt;br /&gt;
# Screen for sample contamination &amp;amp; swap (VerifyBamID) &lt;br /&gt;
&lt;br /&gt;
Run the alignment pipeline (the example aligns 2 samples) : &lt;br /&gt;
 ~/gotcloud/gotcloud align --conf ~/gotcloudExample/[[#Alignment Configuration File|GBR2align.conf]] --outdir [[#Alignment Output Directory|~/gotcloudTutorialOut]] --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the alignment pipeline (about 1-3 minutes), you will see the following message: &lt;br /&gt;
 Processing finished in nn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The final BAM files produced by the alignment pipeline are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/bams&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* BAM (.bam) files - 1 per sample&lt;br /&gt;
** HG00096.recal.bam &lt;br /&gt;
** HG00100.recal.bam &lt;br /&gt;
* BAM index files (.bai) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.bai &lt;br /&gt;
** HG00100.recal.bam.bai &lt;br /&gt;
* BAM checksum files (.md5) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.md5 &lt;br /&gt;
** HG00100.recal.bam.md5 &lt;br /&gt;
* Indicator files that the step completed successfully:&lt;br /&gt;
** HG00096.recal.bam.done &lt;br /&gt;
** HG00100.recal.bam.done &lt;br /&gt;
&lt;br /&gt;
The Quality Control (QC) files are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/QCFiles&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* VerifyBamID output files:&lt;br /&gt;
** HG00096.genoCheck.depthRG &lt;br /&gt;
** HG00096.genoCheck.depthSM &lt;br /&gt;
** HG00096.genoCheck.selfRG &lt;br /&gt;
** HG00096.genoCheck.selfSM &lt;br /&gt;
** HG00100.genoCheck.depthRG &lt;br /&gt;
** HG00100.genoCheck.depthSM &lt;br /&gt;
** HG00100.genoCheck.selfRG &lt;br /&gt;
** HG00100.genoCheck.selfSM &lt;br /&gt;
&lt;br /&gt;
* VerifyBamID step completion files – 1 per sample&lt;br /&gt;
** HG00096.genoCheck.done &lt;br /&gt;
** HG00100.genoCheck.done &lt;br /&gt;
&lt;br /&gt;
* QPLOT output files&lt;br /&gt;
** HG00096.qplot.R &lt;br /&gt;
** HG00096.qplot.stats &lt;br /&gt;
** HG00100.qplot.R &lt;br /&gt;
** HG00100.qplot.stats &lt;br /&gt;
&lt;br /&gt;
* QPLOT step completion files – 1 per sample&lt;br /&gt;
** HG00096.qplot.done &lt;br /&gt;
** HG00100.qplot.done &lt;br /&gt;
&lt;br /&gt;
For information on the VerifyBamID output, see: [[Understanding VerifyBamID output]] &lt;br /&gt;
&lt;br /&gt;
For information on the QPLOT output, see: [[Understanding QPLOT output]]&lt;br /&gt;
&lt;br /&gt;
== STEP 3 : Run GotCloud Variant Calling Pipeline == &lt;br /&gt;
The next step is to analyze BAM files by calling SNPs and generating a VCF file containing the variant calls. &lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Filter out reads with low mapping quality &lt;br /&gt;
# Per Base Alignment Quality Adjustment (BAQ) &lt;br /&gt;
# Resolve overlapping paired end reads &lt;br /&gt;
# Generate genotype likelihood files &lt;br /&gt;
# Perform variant calling &lt;br /&gt;
# Extract features from variant sites &lt;br /&gt;
# Perform variant filtering &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
To speed variant calling, each chromosome is broken up into smaller regions which are processed separately.  While initially split by sample, the per sample data gets merged and is processed together for each region.  These regions are later merged to result in a single Variant Call File (VCF) per chromosome.  For the tutorial all of the data falls within a single region.&lt;br /&gt;
&lt;br /&gt;
Run the variant calling pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud snpcall --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --region 20:42900000-43200000 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the variant calling pipeline (about 3-4 minutes), you will see the following message: &lt;br /&gt;
  Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
On SNP Call success, the VCF files of interest are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered*&lt;br /&gt;
&lt;br /&gt;
This gives you the following files:&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.vcf.gz &#039;&#039;&#039; - vcf for whole chromosome after it has been run through hardfilters and SVM filters and marked with PASS/FAIL including per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf.norm.log - log file&lt;br /&gt;
* chr20.filtered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
* chr20.filtered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
* chr20.filtered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
&lt;br /&gt;
Also in the ~/gotcloudTutorialOut/vcfs/chr20 directory are intermediate files:&lt;br /&gt;
* the whole chromosome variant calls prior to any filtering: &lt;br /&gt;
** chr20.merged.sites.vcf - without per sample genotypes&lt;br /&gt;
** chr20.merged.stats.vcf &lt;br /&gt;
** chr20.merged.vcf - including per sample genotypes&lt;br /&gt;
** chr20.merged.vcf.OK - indicator that the step completed successfully&lt;br /&gt;
* the hardfiltered (pre-svm filtered) variant calls:&lt;br /&gt;
** chr20.filtered.vcf.gz - vcf for whole chromosome after it has been run through hard filters&lt;br /&gt;
** chr20.hardfiltered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.log - log file&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
* 40000001.45000000 subdirectory contains the data for just that region.&lt;br /&gt;
&lt;br /&gt;
The ~/gotcloudTutorialOut/split/chr20 folder contains a VCF with just the sites that pass the filters.&lt;br /&gt;
 ls ~/gotcloudTutorialOut/split/chr20/&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.PASS.vcf.gz &#039;&#039;&#039; – vcf of just sites that pass all filters&lt;br /&gt;
* chr20.filtered.PASS.split.1.vcf.gz - intermediate file&lt;br /&gt;
* chr20.filtered.PASS.split.err - log file&lt;br /&gt;
* chr20.filtered.PASS.split.vcflist - list of intermediate files&lt;br /&gt;
* subset.OK &lt;br /&gt;
&lt;br /&gt;
In addition to the vcfs subdirectory, there are additional intermediate files/directories:&lt;br /&gt;
* glfs – holds genotype likelihood format [[GLF]] files split by chromosome, region, and sample&lt;br /&gt;
* pvcfs – holds intermediate vcf files split by chromosome and region&lt;br /&gt;
&lt;br /&gt;
Note: the tutorial does not produce a target directory, but if you run with targeted data, you may see that.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 4 : Run GotCloud Genotype Refinement Pipeline == &lt;br /&gt;
The next step is to perform genotype refinement using linkage disequilibrium information using [http://faculty.washington.edu/browning/beagle/beagle.html Beagle] &amp;amp; [[ThunderVCF]]. &lt;br /&gt;
&lt;br /&gt;
Run the LD-aware genotype refinement pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud ldrefine --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of this pipeline (about 10 minutes), you will see the following message: &lt;br /&gt;
 Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The output from the beagle step of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
The output from the thunderVcf (final step) of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 5 : Run GotCloud Association Analysis Pipeline (EPACTS) == &lt;br /&gt;
&lt;br /&gt;
We will assume that the EPACTS are installed in the following directory&lt;br /&gt;
 setenv EPACTS /path/to/epacts&lt;br /&gt;
(If you need to install EPACTS, please refer to the documentation at [[EPACTS#Installation_Details]])&lt;br /&gt;
&lt;br /&gt;
 $EPACTS/epacts single --vcf ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered.vcf.gz --ped ~/gotcloudExample/test.GBR60.ped \\&lt;br /&gt;
    --out ~/gotcloudTutorialOut/epacts --test q.linear --run 1 --top 1 --chr 20&lt;br /&gt;
&lt;br /&gt;
Upon successful run, you will see files starting with ~/gotcloudTutorialOut/epacts&lt;br /&gt;
 ls ~/gotcloudTutorialOut/epacts*&lt;br /&gt;
&lt;br /&gt;
To see the top associated variants, you can run&lt;br /&gt;
 less ~/gotcloudTutorialOut/epacts.epacts.top5000&lt;br /&gt;
&lt;br /&gt;
To see the locus-zoom like plot, you can type the following command (assuming GNU gnuplot 4.2 or higher version was installed)&lt;br /&gt;
 xpdf ~/gotcloudTutorialOut/epacs.zoom.20.42987877.pdf&lt;br /&gt;
&lt;br /&gt;
Click [[Media:EPACTS TEST.zoom.20.42987877.pdf | Exampe LocusZoom PDF]] to see the expected output pdf&lt;br /&gt;
&lt;br /&gt;
= Frequently Asked Questions (FAQs) =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;I ran the tutorial example successfully, how can I run it with my real sequence data?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Congratulations for your successful run of your [[GotCloud]] Tutorial. Please see [[#Tutorial Inputs]] section to prepare your own input files for your sequence data. You will need to specify the FASTQ files associated with its sample names as explained. In addition, you will need to download the full reference and resource file across whole genome (the Tutorial contains only chr20 portion to make it compact) See [[#Alignment Configuration File]] section for the detailed information. Also, please refer to the original documentation of [[GotCloud]] for more detailed guide on installation beyond the scope of the tutorial.&lt;br /&gt;
&lt;br /&gt;
= Input Files for GotCloud Tutorial = &lt;br /&gt;
&lt;br /&gt;
This section describes the input files needed for the GotCloud tutorial. You don&#039;t need to know this detail to run the tutorial, but if you&#039;re interested in understanding the structure of GotCloud pipeline and run with your own sample, this would be a good starting point&lt;br /&gt;
&lt;br /&gt;
== Alignment Pipeline == &lt;br /&gt;
=== List of Input Files needed for Alignment ===&lt;br /&gt;
The command-line inputs to the tutorial alignment pipeline are: &lt;br /&gt;
# [[#Alignment Configuration File|Configuration File (--conf)]] &lt;br /&gt;
#* Specifies the configuration file to use when running&lt;br /&gt;
# [[#Alignment Output Directory|Output Directory (--outdir)]] &lt;br /&gt;
#* Directory where the output should be placed.&lt;br /&gt;
&lt;br /&gt;
Additional information required to run the alignment pipeline:&lt;br /&gt;
# [[#Alignment FASTQ Index File|Index file of FASTQs]]&lt;br /&gt;
# [[#Alignment Reference Files|Alignment Reference Files]]&lt;br /&gt;
For the tutorial, these values are specified in the configuration file.&lt;br /&gt;
&lt;br /&gt;
=== Alignment Configuration File === &lt;br /&gt;
The configuration file contains KEY = VALUE settings that override defaults and set specific values for the given run. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt; &lt;br /&gt;
INDEX_FILE = GBR60fastq.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_DIR = chr20Ref &lt;br /&gt;
AS = NCBI37 &lt;br /&gt;
FA_REF = $(REF_DIR)/human_g1k_v37_chr20.fa &lt;br /&gt;
DBSNP_VCF =  $(REF_DIR)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF = $(REF_DIR)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&amp;lt;/pre&amp;gt; &lt;br /&gt;
&lt;br /&gt;
This configuration file sets: &lt;br /&gt;
* [[#Alignment FASTQ Index File|INDEX_FILE]] - file containing the fastqs to be processed as well as the read group information for these fastqs.&lt;br /&gt;
* Reference Information: see [[#Alignment Reference Files|Alignment Reference Files]] for more information&lt;br /&gt;
** AS - assembly value to put in the BAM &lt;br /&gt;
&lt;br /&gt;
The index file and chromosome 20 references used in this tutorial are included with the example data under the ~/gotcloudExample directory.  The tutorial uses chromosome 20 only references in order to speed the processing time.  &lt;br /&gt;
&lt;br /&gt;
When running with your own data, you will need to update the:&lt;br /&gt;
* Index File to contain the information for your own fastq files&lt;br /&gt;
** See [[#Alignment FASTQ Index File|Alignment FASTQ Index File]] for more information on the contents of the index file.&lt;br /&gt;
* The reference files to be whole genome references&lt;br /&gt;
** If you are just running chromosome 20, you can use the tutorial references&lt;br /&gt;
** Whole genome reference files can be downloaded from [[GotCloudReference]].&lt;br /&gt;
** See [[#Alignment Reference Files|Alignment Reference Files]] for more information.&lt;br /&gt;
&lt;br /&gt;
Note: It is recommended that you use absolute paths (full path names, like “/home/mktrost/gotcloudReference” rather than just “gotcloudReference”).  This example does not use absolute paths in order to be flexible to where the data is installed, but using relative paths requires it to be run from the correct directory.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
&lt;br /&gt;
Reference files are required for running both the alignment and variant calling pipelines.  &lt;br /&gt;
&lt;br /&gt;
The configuration keys for setting these are:&lt;br /&gt;
* FA_REF - Genome sequence reference files (needed for both pipelines)&lt;br /&gt;
* DBSNP_VCF – DBSNP site VCF file (needed for both pipelines)&lt;br /&gt;
* HM3_VCF - HAPMAP site VCF file (needed for both pipelines)&lt;br /&gt;
* INDEL_PREFIX - Indel sites file (need for variant calling pipeline)&lt;br /&gt;
&lt;br /&gt;
The tutorial configuration file is setup to point to the required chromosome 20 reference files which are included with the tutorial example data in ~/gotcloudExample/chr20Ref/. &lt;br /&gt;
&lt;br /&gt;
If you are running more than just chromosome 20, you will need whole genome reference files which can be downloaded from [[GotCloudReference]].&lt;br /&gt;
&lt;br /&gt;
The configuration settings for these files are setup in the default configuration so do not need to be specified.  You just need to set REF_DIR in your configuration file to the path where you installed your reference files.&lt;br /&gt;
&lt;br /&gt;
To learn more about the reference files that are required, see [[GotCloud: Reference Files]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Alignment Output Directory === &lt;br /&gt;
This setting tells the alignment pipeline where to write the output and intermediate files. &lt;br /&gt;
&lt;br /&gt;
The output directory will be created if it doesn&#039;t already exist and will contain the following Directories/files: &lt;br /&gt;
* bams - directory containing the final bams/bai files &lt;br /&gt;
** HG00096.OK - indicates that this sample completed alignment processing &lt;br /&gt;
** HG00100.OK - indicates that this sample completed alignment processing &lt;br /&gt;
* failLogs - directory containing logs from steps that failed &lt;br /&gt;
** this directory is only created if an error is detected&lt;br /&gt;
* Makefiles - directory containing the makefiles with commands for processing each sample &lt;br /&gt;
** biopipe_HG00096.Makefile – commands for processing sample HG00096&lt;br /&gt;
** biopipe_HG00100.Makefile – commands for processing sample HG00100&lt;br /&gt;
** biopipe_HG00096.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
** biopipe_HG00100.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
* QCFiles - directory containing the QC Results &lt;br /&gt;
** &lt;br /&gt;
* tmp - directory containing temporary alignment files &lt;br /&gt;
** bwa.sai.t – contains temporary files for the 1st step of bwa that generates sai files from fastq files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
**** HG0096&lt;br /&gt;
** alignment.bwa – contains temporary files for the 2nd step of bwa that generates BAM files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
** alignment.pol - contains temporary files for the polish bam step that cleans up BAM files&lt;br /&gt;
** alignment.dedup – contains temporary files for the deduping step&lt;br /&gt;
*** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
*** *.metrics – metrics files for the deduping step&lt;br /&gt;
** alignment.recal – contains temporary files for the deduping step&lt;br /&gt;
*** *.log – contains information about the recalibration step&lt;br /&gt;
*** *.qemp – contains the recalibration tables used for recalibrating each BAM file&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Alignment FASTQ Index File=== &lt;br /&gt;
&lt;br /&gt;
There are four fastq files in {ROOT_DIR}/test/align/fastq/Sample_1 and four fastq files in {ROOT_DIR}/test/align/fastq/Sample_2, both in paired-end format.  Normally, we would need to build an index file for these files. Conveniently, an index file (indexFile.txt) already exists for the automatic test samples.  It can be found in {ROOT_DIR}/test/align/, and contains the following information in tab-delimited format: &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME FASTQ1                           FASTQ2                           RGID   SAMPLE    LIBRARY CENTER PLATFORM &lt;br /&gt;
 Sample1    fastq/Sample_1/File1_R1.fastq.gz fastq/Sample_1/File1_R2.fastq.gz RGID1  SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample1    fastq/Sample_1/File2_R1.fastq.gz fastq/Sample_1/File2_R2.fastq.gz RGID1a SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File1_R1.fastq.gz fastq/Sample_2/File1_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File2_R1.fastq.gz fastq/Sample_2/File2_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
&lt;br /&gt;
If you are in the {ROOT_DIR}/test/align directory, you can use this file as-is.  If you prefer, you can create a new index file and change the MERGE_NAME, RGID, SAMPLE, LIBRARY, CENTER, or PLATFORM values. It is recommended that you do not modify existing files in {ROOT_DIR}/test/align. &lt;br /&gt;
&lt;br /&gt;
If you want to run this example from a different directory, make sure the FASTQ1 and FASTQ2 paths are correct.  That is, each of the FASTQ1 and FASTQ2 entry in the index file should look like the following: &lt;br /&gt;
&lt;br /&gt;
 {ROOT_DIR}/test/align/fastq/Sample_1/File1_R1.fastq.gz &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to run this example from a different directory, but do not want to edit the index file, you can create a relative path to the test fastq files so their path agrees with that listed in the index file: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/align/fastq fastq &lt;br /&gt;
&lt;br /&gt;
This will create a symbolic link to the test fastq directory from your current directory. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Mapping_Pipeline#Sequence_Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Analyzing a Sample== &lt;br /&gt;
&lt;br /&gt;
Using UMAKE, you can analyze BAM files by calling SNPs, and generate a VCF file containing the results.  Once again, we can analyze BAM files used in the automatic test.  For this example, we have 60 BAM files, which can be found in {ROOT_DIR}/test/umake/bams.  These contain sequence information for a targeted region in chromosome 20. &lt;br /&gt;
&lt;br /&gt;
In addition to the BAM files, you will need three files to run UMAKE: an index file, a configuration file, and a bed file (needed to analyze BAM files from targeted/exome sequencing). &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
===Index file=== &lt;br /&gt;
&lt;br /&gt;
First, you need a list of all the BAM files to be analyzed. Conveniently, the a test index file (umake_test.index) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
 NA12272 ALL     bams/NA12272.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 NA12004 ALL     bams/NA12004.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 ... &lt;br /&gt;
 NA12874 ALL     bams/NA12874.mapped.LS454.ssaha2.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
&lt;br /&gt;
You can use this file directly if you change your current directory to {ROOT_DIR}/test/umake/. &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to copy and use this index file to a different directory, you can create a symbolic link to the bams folder as follows: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/umake/bams bams &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===BED file=== &lt;br /&gt;
&lt;br /&gt;
This file contains a single line: &lt;br /&gt;
&lt;br /&gt;
 chr20   20000050        20300000 &lt;br /&gt;
&lt;br /&gt;
You can copy this to the current directory and use it as-is. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Targeted.2FExome_Sequencing_Settings|targeted/exome sequencing settings]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Configuration file=== &lt;br /&gt;
&lt;br /&gt;
A configuration file (umake_test.conf) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
CHRS = 20 &lt;br /&gt;
BAM_INDEX = GBR60bam.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_ROOT = chr20Ref &lt;br /&gt;
# &lt;br /&gt;
REF = $(REF_ROOT)/human_g1k_v37_chr20.fa &lt;br /&gt;
INDEL_PREFIX = $(REF_ROOT)/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
DBSNP_VCF =  $(REF_ROOT)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF =  $(REF_ROOT)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 CHRS = 20 &lt;br /&gt;
 TEST_ROOT = $(UMAKE_ROOT)/test/umake &lt;br /&gt;
 BAM_INDEX = $(TEST_ROOT)/umake_test.index &lt;br /&gt;
 OUT_PREFIX = umake_test &lt;br /&gt;
 REF_ROOT = $(TEST_ROOT)/ref &lt;br /&gt;
 # &lt;br /&gt;
 REF = $(REF_ROOT)/karma.ref/human.g1k.v37.chr20.fa &lt;br /&gt;
 INDEL_PREFIX = $(REF_ROOT)/indels/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
 DBSNP_PREFIX =  $(REF_ROOT)/dbSNP/dbsnp_135_b37.rod &lt;br /&gt;
 HM3_PREFIX =  $(REF_ROOT)/HapMap3/hapmap3_r3_b37_fwd.consensus.qc.poly &lt;br /&gt;
 # &lt;br /&gt;
 RUN_INDEX = TRUE        # create BAM index file &lt;br /&gt;
 RUN_PILEUP = TRUE       # create GLF file from BAM &lt;br /&gt;
 RUN_GLFMULTIPLES = TRUE # create unfiltered SNP calls &lt;br /&gt;
 RUN_VCFPILEUP = TRUE    # create PVCF files using vcfPileup and run infoCollector &lt;br /&gt;
 RUN_FILTER = TRUE       # filter SNPs using vcfCooker &lt;br /&gt;
 RUN_SPLIT = TRUE        # split SNPs into chunks for genotype refinement &lt;br /&gt;
 RUN_BEAGLE = FALSE  # BEAGLE - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 RUN_SUBSET = FALSE  # SUBSET FOR THUNDER - MAY BE SET WITH BEAGLE STEP TOGETHER &lt;br /&gt;
 RUN_THUNDER = FALSE # THUNDER - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 ############################################################################### &lt;br /&gt;
 WRITE_TARGET_LOCI = TRUE  # FOR TARGETED SEQUENCING ONLY -- Write loci file when performing pileup &lt;br /&gt;
 UNIFORM_TARGET_BED = $(TEST_ROOT)/umake_test.bed # Targeted sequencing : When all individuals has the same target. Otherwise, comment it out &lt;br /&gt;
 OFFSET_OFF_TARGET = 50 # Extend target by given # of bases &lt;br /&gt;
 MULTIPLE_TARGET_MAP =  # Target per individual : Each line contains [SM_ID] [TARGET_BED] &lt;br /&gt;
 TARGET_DIR = target    # Directory to store target information &lt;br /&gt;
 SAMTOOLS_VIEW_TARGET_ONLY = TRUE # When performing samtools view, exclude off-target regions (may make command line too long) &lt;br /&gt;
&lt;br /&gt;
If you are running this from a different directory, you will want to change some of the lines as follows: &lt;br /&gt;
&lt;br /&gt;
 BAM_INDEX = {CURRENT_DIR}/umake_test.index &lt;br /&gt;
 UNIFORM_TARGET_BED = {CURRENT_DIR}/umake_test.bed &lt;br /&gt;
&lt;br /&gt;
where {CURRENT_DIR} is the absolute path to the directory that contains the index and bed files.  &lt;br /&gt;
&lt;br /&gt;
An additional option can be added in the configuration file: &lt;br /&gt;
&lt;br /&gt;
 OUT_DIR = {OUT_DIR} &lt;br /&gt;
&lt;br /&gt;
where {OUT_DIR} is the name of directory in which you want the output to be stored.  If you do not specify this in the configuration file, you will need to add an extra parameter when you run UMAKE in the next step. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Configuration_File|the configuration file]], [[Variant_Calling_Pipeline_(UMAKE)#Reference_Files|reference files]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Further Information== &lt;br /&gt;
&lt;br /&gt;
[[Mapping_Pipeline|Mapping (Alignment) Pipeline]] &lt;br /&gt;
&lt;br /&gt;
[[Variant_Calling_Pipeline_(UMAKE)|Variant Calling Pipeline (UMAKE)]]&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7501</id>
		<title>Tutorial: GotCloud</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Tutorial:_GotCloud&amp;diff=7501"/>
		<updated>2013-06-16T21:35:43Z</updated>

		<summary type="html">&lt;p&gt;Ben Lerch: /* STEP 2 : Run GotCloud Alignment Pipeline */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= GotCloud Tutorial =&lt;br /&gt;
In this tutorial, we illustrate some of the essential steps in the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
For a background on GotCloud and Sequence Analysis Pipelines, see [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
While GotCloud can run on a cluster of machines or instances, this tutorial is just a small test that just runs on the machine the commands are run on.&lt;br /&gt;
&lt;br /&gt;
GotCloud and this basic tutorial were presented at the [http://ibg.colorado.edu/dokuwiki/doku.php?id=workshop:2013:announcement 2013 IBG Workshop].  It was presented in two sessions.  On Wednesday an overview was presented with steps for running the tutorial data: [[Media:IBG2013GotCloud.pdf|IBG2013GotCloud.pdf]].  On Friday more detail on the input files and what goes into generating the input files was presented: [[Media:GotCloudIBGWorkshop2013Friday.pdf|GotCloudIBGWorkshop2013Friday.pdf]].&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;This tutorial is in the process of being updated for gotcloud version 1.06 (April 17. 2013).&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== STEP 1 : Setup GotCloud ==&lt;br /&gt;
&lt;br /&gt;
[[GotCloud]] has been developed and tested on Linux Ubuntu 12.10 and 12.04.2 LTS but has not been tested on other Linux operating systems. It is not available for Windows. If you do not have your own set of machines to run on, GotCloud is also available for Ubuntu running on the Amazon Elastic Compute Cloud, see [[Amazon_Snapshot]] for more information.&lt;br /&gt;
&lt;br /&gt;
We will use 3 different directories for this tutorial:&lt;br /&gt;
# path to the directory where gotcloud is installed, default ~/gotcloud/&lt;br /&gt;
# path to the directory where the example data is installed, default ~/gotcloudExample&lt;br /&gt;
# path to your output directory, default ~/gotcloudTutorialOut/&lt;br /&gt;
&lt;br /&gt;
If the directories specified above do not reflect the directories you would like to use, replace their occurrances in the instructions below with the appropriate paths.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Step 1a: Install GotCloud ===&lt;br /&gt;
In order to run this tutorial, you need to make sure you have GotCloud installed on your system.  &lt;br /&gt;
&lt;br /&gt;
If you have root and would like to install gotcloud on your system, follow: [[GotCloud#Install_GotCloud_Software| root access installation instructions]]&lt;br /&gt;
&lt;br /&gt;
Otherwise, you can install it in your own directory:&lt;br /&gt;
# Change to the directory where you want gotcloud/ installed&lt;br /&gt;
# Download the gotcloud tar from the ftp site.&lt;br /&gt;
# Extract the tar&lt;br /&gt;
# Build (compile) the source&lt;br /&gt;
#* Note: as the source builds, many messages will scroll through your terminal.  You may even see some warnings.  These messages are normal and expected.  As long as the build does not end with an error, you have successfully built the source.&lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloud_latest.tgz  # Download&lt;br /&gt;
 tar xf gotcloud_latest.tgz     # Extracts into gotcloud/&lt;br /&gt;
 cd ~/gotcloud/src; make         # Build source&lt;br /&gt;
 &lt;br /&gt;
GotCloud requires the following tools to be installed.&lt;br /&gt;
You can run ~/gotcloud/scripts/check_requirements.sh&lt;br /&gt;
...TBD – put in required programs/tools.&lt;br /&gt;
* java (java-common default-jre on ubuntu)&lt;br /&gt;
* make (make on ubuntu)&lt;br /&gt;
* libssl (libssl0.9.8 on ubuntu)&lt;br /&gt;
* gcc 4.4 or newer&lt;br /&gt;
&lt;br /&gt;
=== Step 1b: Install Example Dataset ===&lt;br /&gt;
Our dataset consists of 60 individuals from Great Britain (GBR) sequenced by the 1000 Genomes Project. These individuals have been sequenced to an average depth of about 4x.&lt;br /&gt;
&lt;br /&gt;
To conserve time and disk-space, our analysis will focus on a small region on chromosome 20, 42900000 - 43200000. &lt;br /&gt;
&lt;br /&gt;
The tutorial will run the alignment pipeline on 2 of the individuals (HG00096, HG00100).  The fastqs used for this step are reduced to reads that fall into our target region.&lt;br /&gt;
&lt;br /&gt;
The tutorial will then used previously aligned/mapped reads for the full 60 individuals to generate a list of polymorphic sites and estimate accurate genotypes at each of these sites. &lt;br /&gt;
&lt;br /&gt;
The example dataset we&#039;ll be using is available at: ftp://share.sph.umich.edu/gotcloud/gotcloudExample.tgz &lt;br /&gt;
&lt;br /&gt;
# Change directory to where you want to install the Tutorial data &lt;br /&gt;
# Download the dataset tar from the ftp site &lt;br /&gt;
# Extract the tar &lt;br /&gt;
&lt;br /&gt;
 cd ~&lt;br /&gt;
 wget ftp://share.sph.umich.edu/gotcloud/gotcloudExample_latest.tgz  # Download &lt;br /&gt;
 tar xvf gotcloudExample_latest.tgz    # Extracts into gotcloudExample/&lt;br /&gt;
&lt;br /&gt;
== STEP 2 : Run GotCloud Alignment Pipeline == &lt;br /&gt;
The first step in processing next generation sequence data is mapping the reads to the reference genome, generating per sample BAM files. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Align the fastqs to the reference genome &lt;br /&gt;
#* handles both single &amp;amp; paired end &lt;br /&gt;
# Merge the results from multiple fastqs into 1 file per sample &lt;br /&gt;
# Mark Duplicate Reads are marked &lt;br /&gt;
# Recalibrate Base Qualities &lt;br /&gt;
&lt;br /&gt;
This processing results in 1 BAM file per sample. &lt;br /&gt;
&lt;br /&gt;
The alignment pipeline also includes Quality Control (QC) steps: &lt;br /&gt;
# Visualization of various quality measures (QPLOT) &lt;br /&gt;
# Screen for sample contamination &amp;amp; swap (VerifyBamID) &lt;br /&gt;
&lt;br /&gt;
Run the alignment pipeline (the example aligns 2 samples) : &lt;br /&gt;
 ~/gotcloud/gotcloud align --conf ~/gotcloudExample/[[#Alignment Configuration File|GBR2align.conf]] --outdir [[#Alignment Output Directory|~/gotcloudTutorialOut]] --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the alignment pipeline (about 1-3 minutes), you will see the following message: &lt;br /&gt;
 Processing finished in nn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The final BAM files produced by the alignment pipeline are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/bams&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* BAM (.bam) files - 1 per sample&lt;br /&gt;
** HG00096.recal.bam &lt;br /&gt;
** HG00100.recal.bam &lt;br /&gt;
* BAM index files (.bai) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.bai &lt;br /&gt;
** HG00100.recal.bam.bai &lt;br /&gt;
* BAM checksum files (.md5) – 1 per sample&lt;br /&gt;
** HG00096.recal.bam.md5 &lt;br /&gt;
** HG00100.recal.bam.md5 &lt;br /&gt;
* Indicator files that the step completed successfully:&lt;br /&gt;
** HG00096.recal.bam.done &lt;br /&gt;
** HG00100.recal.bam.done &lt;br /&gt;
&lt;br /&gt;
The Quality Control (QC) files are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/QCFiles&lt;br /&gt;
In this directory you will see:&lt;br /&gt;
* VerifyBamID output files:&lt;br /&gt;
** HG00096.genoCheck.depthRG &lt;br /&gt;
** HG00096.genoCheck.depthSM &lt;br /&gt;
** HG00096.genoCheck.selfRG &lt;br /&gt;
** HG00096.genoCheck.selfSM &lt;br /&gt;
** HG00100.genoCheck.depthRG &lt;br /&gt;
** HG00100.genoCheck.depthSM &lt;br /&gt;
** HG00100.genoCheck.selfRG &lt;br /&gt;
** HG00100.genoCheck.selfSM &lt;br /&gt;
&lt;br /&gt;
* VerifyBamID step completion files – 1 per sample&lt;br /&gt;
** HG00096.genoCheck.done &lt;br /&gt;
** HG00100.genoCheck.done &lt;br /&gt;
&lt;br /&gt;
* QPLOT output files&lt;br /&gt;
** HG00096.qplot.R &lt;br /&gt;
** HG00096.qplot.stats &lt;br /&gt;
** HG00100.qplot.R &lt;br /&gt;
** HG00100.qplot.stats &lt;br /&gt;
&lt;br /&gt;
* QPLOT step completion files – 1 per sample&lt;br /&gt;
** HG00096.qplot.done &lt;br /&gt;
** HG00100.qplot.done &lt;br /&gt;
&lt;br /&gt;
For information on the VerifyBamID output, see: [[Understanding VerifyBamID output]] &lt;br /&gt;
&lt;br /&gt;
For information on the QPLOT output, see: [[Understanding QPLOT output]]&lt;br /&gt;
&lt;br /&gt;
== STEP 3 : Run GotCloud Variant Calling Pipeline == &lt;br /&gt;
The next step is to analyze BAM files by calling SNPs and generating a VCF file containing the variant calls. &lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline has multiple built-in steps to generate BAMs: &lt;br /&gt;
# Filter out reads with low mapping quality &lt;br /&gt;
# Per Base Alignment Quality Adjustment (BAQ) &lt;br /&gt;
# Resolve overlapping paired end reads &lt;br /&gt;
# Generate genotype likelihood files &lt;br /&gt;
# Perform variant calling &lt;br /&gt;
# Extract features from variant sites &lt;br /&gt;
# Perform variant filtering &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
To speed variant calling, each chromosome is broken up into smaller regions which are processed separately.  While initially split by sample, the per sample data gets merged and is processed together for each region.  These regions are later merged to result in a single Variant Call File (VCF) per chromosome.  For the tutorial all of the data falls within a single region.&lt;br /&gt;
&lt;br /&gt;
Run the variant calling pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud snpcall --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --region 20:42900000-43200000 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of the variant calling pipeline (about 3-4 minutes), you will see the following message: &lt;br /&gt;
  Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
On SNP Call success, the VCF files of interest are: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered*&lt;br /&gt;
&lt;br /&gt;
This gives you the following files:&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.vcf.gz &#039;&#039;&#039; - vcf for whole chromosome after it has been run through hardfilters and SVM filters and marked with PASS/FAIL including per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
* chr20.filtered.sites.vcf.norm.log - log file&lt;br /&gt;
* chr20.filtered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
* chr20.filtered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
* chr20.filtered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
&lt;br /&gt;
Also in the ~/gotcloudTutorialOut/vcfs/chr20 directory are intermediate files:&lt;br /&gt;
* the whole chromosome variant calls prior to any filtering: &lt;br /&gt;
** chr20.merged.sites.vcf - without per sample genotypes&lt;br /&gt;
** chr20.merged.stats.vcf &lt;br /&gt;
** chr20.merged.vcf - including per sample genotypes&lt;br /&gt;
** chr20.merged.vcf.OK - indicator that the step completed successfully&lt;br /&gt;
* the hardfiltered (pre-svm filtered) variant calls:&lt;br /&gt;
** chr20.filtered.vcf.gz - vcf for whole chromosome after it has been run through hard filters&lt;br /&gt;
** chr20.hardfiltered.sites.vcf - vcf for whole chromosome after it has been run through filters and marked with PASS/FAIL without the per sample genotypes&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.log - log file&lt;br /&gt;
** chr20.hardfiltered.sites.vcf.summary - summary of filters applied&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.OK - indicator that the filtering completed successfully&lt;br /&gt;
** chr20.hardfiltered.vcf.gz.tbi - index file for the vcf file&lt;br /&gt;
* 40000001.45000000 subdirectory contains the data for just that region.&lt;br /&gt;
&lt;br /&gt;
The ~/gotcloudTutorialOut/split/chr20 folder contains a VCF with just the sites that pass the filters.&lt;br /&gt;
 ls ~/gotcloudTutorialOut/split/chr20/&lt;br /&gt;
* &#039;&#039;&#039;chr20.filtered.PASS.vcf.gz &#039;&#039;&#039; – vcf of just sites that pass all filters&lt;br /&gt;
* chr20.filtered.PASS.split.1.vcf.gz - intermediate file&lt;br /&gt;
* chr20.filtered.PASS.split.err - log file&lt;br /&gt;
* chr20.filtered.PASS.split.vcflist - list of intermediate files&lt;br /&gt;
* subset.OK &lt;br /&gt;
&lt;br /&gt;
In addition to the vcfs subdirectory, there are additional intermediate files/directories:&lt;br /&gt;
* glfs – holds genotype likelihood format [[GLF]] files split by chromosome, region, and sample&lt;br /&gt;
* pvcfs – holds intermediate vcf files split by chromosome and region&lt;br /&gt;
&lt;br /&gt;
Note: the tutorial does not produce a target directory, but if you run with targeted data, you may see that.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 4 : Run GotCloud Genotype Refinement Pipeline == &lt;br /&gt;
The next step is to perform genotype refinement using linkage disequilibrium information using [http://faculty.washington.edu/browning/beagle/beagle.html Beagle] &amp;amp; [[ThunderVCF]]. &lt;br /&gt;
&lt;br /&gt;
Run the LD-aware genotype refinement pipeline: &lt;br /&gt;
 ~/gotcloud/gotcloud ldrefine --conf ~/gotcloudExample/[[GBR60vc.conf]] --outdir ~/gotcloudTutorialOut --numjobs 2 --baseprefix ~/gotcloudExample&lt;br /&gt;
&lt;br /&gt;
Upon successful completion of this pipeline (about 10 minutes), you will see the following message: &lt;br /&gt;
 Commands finished in nnn secs with no errors reported &lt;br /&gt;
&lt;br /&gt;
The output from the beagle step of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz ~/gotcloudTutorialOut/beagle/chr20/chr20.filtered.PASS.beagled.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
The output from the thunderVcf (final step) of the genotype refinement pipeline is found in: &lt;br /&gt;
 ls ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz ~/gotcloudTutorialOut/thunder/chr20/GBR/chr20.filtered.PASS.beagled.GBR.thunder.vcf.gz.tbi &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== STEP 5 : Run GotCloud Association Analysis Pipeline (EPACTS) == &lt;br /&gt;
&lt;br /&gt;
We will assume that the EPACTS are installed in the following directory&lt;br /&gt;
 setenv EPACTS /path/to/epacts&lt;br /&gt;
(If you need to install EPACTS, please refer to the documentation at [[EPACTS#Installation_Details]])&lt;br /&gt;
&lt;br /&gt;
 $EPACTS/epacts single --vcf ~/gotcloudTutorialOut/vcfs/chr20/chr20.filtered.vcf.gz --ped ~/gotcloudExample/test.GBR60.ped \\&lt;br /&gt;
    --out ~/gotcloudTutorialOut/epacts --test q.linear --run 1 --top 1 --chr 20&lt;br /&gt;
&lt;br /&gt;
Upon successful run, you will see files starting with ~/gotcloudTutorialOut/epacts&lt;br /&gt;
 ls ~/gotcloudTutorialOut/epacts*&lt;br /&gt;
&lt;br /&gt;
To see the top associated variants, you can run&lt;br /&gt;
 less ~/gotcloudTutorialOut/epacts.epacts.top5000&lt;br /&gt;
&lt;br /&gt;
To see the locus-zoom like plot, you can type the following command (assuming GNU gnuplot 4.2 or higher version was installed)&lt;br /&gt;
 xpdf ~/gotcloudTutorialOut/epacs.zoom.20.42987877.pdf&lt;br /&gt;
&lt;br /&gt;
Click [[Media:EPACTS TEST.zoom.20.42987877.pdf | Exampe LocusZoom PDF]] to see the expected output pdf&lt;br /&gt;
&lt;br /&gt;
= Frequently Asked Questions (FAQs) =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;I ran the tutorial example successfully, how can I run it with my real sequence data?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Congratulations for your successful run of your [[GotCloud]] Tutorial. Please see [[#Tutorial Inputs]] section to prepare your own input files for your sequence data. You will need to specify the FASTQ files associated with its sample names as explained. In addition, you will need to download the full reference and resource file across whole genome (the Tutorial contains only chr20 portion to make it compact) See [[#Alignment Configuration File]] section for the detailed information. Also, please refer to the original documentation of [[GotCloud]] for more detailed guide on installation beyond the scope of tutorial&lt;br /&gt;
&lt;br /&gt;
= Input Files for GotCloud Tutorial = &lt;br /&gt;
&lt;br /&gt;
This section describes the input files needed for the GotCloud tutorial. You don&#039;t need to know this detail to run the tutorial, but if you&#039;re interested in understanding the structure of GotCloud pipeline and run with your own sample, this would be a good starting point&lt;br /&gt;
&lt;br /&gt;
== Alignment Pipeline == &lt;br /&gt;
=== List of Input Files needed for Alignment ===&lt;br /&gt;
The command-line inputs to the tutorial alignment pipeline are: &lt;br /&gt;
# [[#Alignment Configuration File|Configuration File (--conf)]] &lt;br /&gt;
#* Specifies the configuration file to use when running&lt;br /&gt;
# [[#Alignment Output Directory|Output Directory (--outdir)]] &lt;br /&gt;
#* Directory where the output should be placed.&lt;br /&gt;
&lt;br /&gt;
Additional information required to run the alignment pipeline:&lt;br /&gt;
# [[#Alignment FASTQ Index File|Index file of FASTQs]]&lt;br /&gt;
# [[#Alignment Reference Files|Alignment Reference Files]]&lt;br /&gt;
For the tutorial, these values are specified in the configuration file.&lt;br /&gt;
&lt;br /&gt;
=== Alignment Configuration File === &lt;br /&gt;
The configuration file contains KEY = VALUE settings that override defaults and set specific values for the given run. &lt;br /&gt;
&lt;br /&gt;
&amp;lt;pre&amp;gt; &lt;br /&gt;
INDEX_FILE = GBR60fastq.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_DIR = chr20Ref &lt;br /&gt;
AS = NCBI37 &lt;br /&gt;
FA_REF = $(REF_DIR)/human_g1k_v37_chr20.fa &lt;br /&gt;
DBSNP_VCF =  $(REF_DIR)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF = $(REF_DIR)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&amp;lt;/pre&amp;gt; &lt;br /&gt;
&lt;br /&gt;
This configuration file sets: &lt;br /&gt;
* [[#Alignment FASTQ Index File|INDEX_FILE]] - file containing the fastqs to be processed as well as the read group information for these fastqs.&lt;br /&gt;
* Reference Information: see [[#Alignment Reference Files|Alignment Reference Files]] for more information&lt;br /&gt;
** AS - assembly value to put in the BAM &lt;br /&gt;
&lt;br /&gt;
The index file and chromosome 20 references used in this tutorial are included with the example data under the ~/gotcloudExample directory.  The tutorial uses chromosome 20 only references in order to speed the processing time.  &lt;br /&gt;
&lt;br /&gt;
When running with your own data, you will need to update the:&lt;br /&gt;
* Index File to contain the information for your own fastq files&lt;br /&gt;
** See [[#Alignment FASTQ Index File|Alignment FASTQ Index File]] for more information on the contents of the index file.&lt;br /&gt;
* The reference files to be whole genome references&lt;br /&gt;
** If you are just running chromosome 20, you can use the tutorial references&lt;br /&gt;
** Whole genome reference files can be downloaded from [[GotCloudReference]].&lt;br /&gt;
** See [[#Alignment Reference Files|Alignment Reference Files]] for more information.&lt;br /&gt;
&lt;br /&gt;
Note: It is recommended that you use absolute paths (full path names, like “/home/mktrost/gotcloudReference” rather than just “gotcloudReference”).  This example does not use absolute paths in order to be flexible to where the data is installed, but using relative paths requires it to be run from the correct directory.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
&lt;br /&gt;
Reference files are required for running both the alignment and variant calling pipelines.  &lt;br /&gt;
&lt;br /&gt;
The configuration keys for setting these are:&lt;br /&gt;
* FA_REF - Genome sequence reference files (needed for both pipelines)&lt;br /&gt;
* DBSNP_VCF – DBSNP site VCF file (needed for both pipelines)&lt;br /&gt;
* HM3_VCF - HAPMAP site VCF file (needed for both pipelines)&lt;br /&gt;
* INDEL_PREFIX - Indel sites file (need for variant calling pipeline)&lt;br /&gt;
&lt;br /&gt;
The tutorial configuration file is setup to point to the required chromosome 20 reference files which are included with the tutorial example data in ~/gotcloudExample/chr20Ref/. &lt;br /&gt;
&lt;br /&gt;
If you are running more than just chromosome 20, you will need whole genome reference files which can be downloaded from [[GotCloudReference]].&lt;br /&gt;
&lt;br /&gt;
The configuration settings for these files are setup in the default configuration so do not need to be specified.  You just need to set REF_DIR in your configuration file to the path where you installed your reference files.&lt;br /&gt;
&lt;br /&gt;
To learn more about the reference files that are required, see [[GotCloud: Reference Files]].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Alignment Output Directory === &lt;br /&gt;
This setting tells the alignment pipeline where to write the output and intermediate files. &lt;br /&gt;
&lt;br /&gt;
The output directory will be created if it doesn&#039;t already exist and will contain the following Directories/files: &lt;br /&gt;
* bams - directory containing the final bams/bai files &lt;br /&gt;
** HG00096.OK - indicates that this sample completed alignment processing &lt;br /&gt;
** HG00100.OK - indicates that this sample completed alignment processing &lt;br /&gt;
* failLogs - directory containing logs from steps that failed &lt;br /&gt;
** this directory is only created if an error is detected&lt;br /&gt;
* Makefiles - directory containing the makefiles with commands for processing each sample &lt;br /&gt;
** biopipe_HG00096.Makefile – commands for processing sample HG00096&lt;br /&gt;
** biopipe_HG00100.Makefile – commands for processing sample HG00100&lt;br /&gt;
** biopipe_HG00096.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
** biopipe_HG00100.Makefile.log – log file from running the associated Makefile&lt;br /&gt;
* QCFiles - directory containing the QC Results &lt;br /&gt;
** &lt;br /&gt;
* tmp - directory containing temporary alignment files &lt;br /&gt;
** bwa.sai.t – contains temporary files for the 1st step of bwa that generates sai files from fastq files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
**** HG0096&lt;br /&gt;
** alignment.bwa – contains temporary files for the 2nd step of bwa that generates BAM files&lt;br /&gt;
*** fastq&lt;br /&gt;
**** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
** alignment.pol - contains temporary files for the polish bam step that cleans up BAM files&lt;br /&gt;
** alignment.dedup – contains temporary files for the deduping step&lt;br /&gt;
*** *.done – indicator files that the step to generate the file completed&lt;br /&gt;
*** *.metrics – metrics files for the deduping step&lt;br /&gt;
** alignment.recal – contains temporary files for the deduping step&lt;br /&gt;
*** *.log – contains information about the recalibration step&lt;br /&gt;
*** *.qemp – contains the recalibration tables used for recalibrating each BAM file&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Alignment FASTQ Index File=== &lt;br /&gt;
&lt;br /&gt;
There are four fastq files in {ROOT_DIR}/test/align/fastq/Sample_1 and four fastq files in {ROOT_DIR}/test/align/fastq/Sample_2, both in paired-end format.  Normally, we would need to build an index file for these files. Conveniently, an index file (indexFile.txt) already exists for the automatic test samples.  It can be found in {ROOT_DIR}/test/align/, and contains the following information in tab-delimited format: &lt;br /&gt;
&lt;br /&gt;
 MERGE_NAME FASTQ1                           FASTQ2                           RGID   SAMPLE    LIBRARY CENTER PLATFORM &lt;br /&gt;
 Sample1    fastq/Sample_1/File1_R1.fastq.gz fastq/Sample_1/File1_R2.fastq.gz RGID1  SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample1    fastq/Sample_1/File2_R1.fastq.gz fastq/Sample_1/File2_R2.fastq.gz RGID1a SampleID1 Lib1    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File1_R1.fastq.gz fastq/Sample_2/File1_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
 Sample2    fastq/Sample_2/File2_R1.fastq.gz fastq/Sample_2/File2_R2.fastq.gz RGID2  SampleID2 Lib2    UM     ILLUMINA &lt;br /&gt;
&lt;br /&gt;
If you are in the {ROOT_DIR}/test/align directory, you can use this file as-is.  If you prefer, you can create a new index file and change the MERGE_NAME, RGID, SAMPLE, LIBRARY, CENTER, or PLATFORM values. It is recommended that you do not modify existing files in {ROOT_DIR}/test/align. &lt;br /&gt;
&lt;br /&gt;
If you want to run this example from a different directory, make sure the FASTQ1 and FASTQ2 paths are correct.  That is, each of the FASTQ1 and FASTQ2 entry in the index file should look like the following: &lt;br /&gt;
&lt;br /&gt;
 {ROOT_DIR}/test/align/fastq/Sample_1/File1_R1.fastq.gz &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to run this example from a different directory, but do not want to edit the index file, you can create a relative path to the test fastq files so their path agrees with that listed in the index file: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/align/fastq fastq &lt;br /&gt;
&lt;br /&gt;
This will create a symbolic link to the test fastq directory from your current directory. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Mapping_Pipeline#Sequence_Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Analyzing a Sample== &lt;br /&gt;
&lt;br /&gt;
Using UMAKE, you can analyze BAM files by calling SNPs, and generate a VCF file containing the results.  Once again, we can analyze BAM files used in the automatic test.  For this example, we have 60 BAM files, which can be found in {ROOT_DIR}/test/umake/bams.  These contain sequence information for a targeted region in chromosome 20. &lt;br /&gt;
&lt;br /&gt;
In addition to the BAM files, you will need three files to run UMAKE: an index file, a configuration file, and a bed file (needed to analyze BAM files from targeted/exome sequencing). &lt;br /&gt;
 &lt;br /&gt;
&lt;br /&gt;
===Index file=== &lt;br /&gt;
&lt;br /&gt;
First, you need a list of all the BAM files to be analyzed. Conveniently, the a test index file (umake_test.index) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
 NA12272 ALL     bams/NA12272.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 NA12004 ALL     bams/NA12004.mapped.ILLUMINA.bwa.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
 ... &lt;br /&gt;
 NA12874 ALL     bams/NA12874.mapped.LS454.ssaha2.CEU.low_coverage.20101123.chrom20.20000001.20300000.bam &lt;br /&gt;
&lt;br /&gt;
You can use this file directly if you change your current directory to {ROOT_DIR}/test/umake/. &lt;br /&gt;
&lt;br /&gt;
Alternately, if you want to copy and use this index file to a different directory, you can create a symbolic link to the bams folder as follows: &lt;br /&gt;
&lt;br /&gt;
 ln -s {ROOT_DIR}/test/umake/bams bams &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Index_File|the index file]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===BED file=== &lt;br /&gt;
&lt;br /&gt;
This file contains a single line: &lt;br /&gt;
&lt;br /&gt;
 chr20   20000050        20300000 &lt;br /&gt;
&lt;br /&gt;
You can copy this to the current directory and use it as-is. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Targeted.2FExome_Sequencing_Settings|targeted/exome sequencing settings]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
===Configuration file=== &lt;br /&gt;
&lt;br /&gt;
A configuration file (umake_test.conf) already exists in {ROOT_DIR}/test/umake/.  It contains the following information: &lt;br /&gt;
&lt;br /&gt;
CHRS = 20 &lt;br /&gt;
BAM_INDEX = GBR60bam.index &lt;br /&gt;
############ &lt;br /&gt;
# References &lt;br /&gt;
REF_ROOT = chr20Ref &lt;br /&gt;
# &lt;br /&gt;
REF = $(REF_ROOT)/human_g1k_v37_chr20.fa &lt;br /&gt;
INDEL_PREFIX = $(REF_ROOT)/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
DBSNP_VCF =  $(REF_ROOT)/dbsnp135_chr20.vcf.gz &lt;br /&gt;
HM3_VCF =  $(REF_ROOT)/hapmap_3.3.b37.sites.chr20.vcf.gz &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 CHRS = 20 &lt;br /&gt;
 TEST_ROOT = $(UMAKE_ROOT)/test/umake &lt;br /&gt;
 BAM_INDEX = $(TEST_ROOT)/umake_test.index &lt;br /&gt;
 OUT_PREFIX = umake_test &lt;br /&gt;
 REF_ROOT = $(TEST_ROOT)/ref &lt;br /&gt;
 # &lt;br /&gt;
 REF = $(REF_ROOT)/karma.ref/human.g1k.v37.chr20.fa &lt;br /&gt;
 INDEL_PREFIX = $(REF_ROOT)/indels/1kg.pilot_release.merged.indels.sites.hg19 &lt;br /&gt;
 DBSNP_PREFIX =  $(REF_ROOT)/dbSNP/dbsnp_135_b37.rod &lt;br /&gt;
 HM3_PREFIX =  $(REF_ROOT)/HapMap3/hapmap3_r3_b37_fwd.consensus.qc.poly &lt;br /&gt;
 # &lt;br /&gt;
 RUN_INDEX = TRUE        # create BAM index file &lt;br /&gt;
 RUN_PILEUP = TRUE       # create GLF file from BAM &lt;br /&gt;
 RUN_GLFMULTIPLES = TRUE # create unfiltered SNP calls &lt;br /&gt;
 RUN_VCFPILEUP = TRUE    # create PVCF files using vcfPileup and run infoCollector &lt;br /&gt;
 RUN_FILTER = TRUE       # filter SNPs using vcfCooker &lt;br /&gt;
 RUN_SPLIT = TRUE        # split SNPs into chunks for genotype refinement &lt;br /&gt;
 RUN_BEAGLE = FALSE  # BEAGLE - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 RUN_SUBSET = FALSE  # SUBSET FOR THUNDER - MAY BE SET WITH BEAGLE STEP TOGETHER &lt;br /&gt;
 RUN_THUNDER = FALSE # THUNDER - MUST SET AFTER FINISHING PREVIOUS STEPS &lt;br /&gt;
 ############################################################################### &lt;br /&gt;
 WRITE_TARGET_LOCI = TRUE  # FOR TARGETED SEQUENCING ONLY -- Write loci file when performing pileup &lt;br /&gt;
 UNIFORM_TARGET_BED = $(TEST_ROOT)/umake_test.bed # Targeted sequencing : When all individuals has the same target. Otherwise, comment it out &lt;br /&gt;
 OFFSET_OFF_TARGET = 50 # Extend target by given # of bases &lt;br /&gt;
 MULTIPLE_TARGET_MAP =  # Target per individual : Each line contains [SM_ID] [TARGET_BED] &lt;br /&gt;
 TARGET_DIR = target    # Directory to store target information &lt;br /&gt;
 SAMTOOLS_VIEW_TARGET_ONLY = TRUE # When performing samtools view, exclude off-target regions (may make command line too long) &lt;br /&gt;
&lt;br /&gt;
If you are running this from a different directory, you will want to change some of the lines as follows: &lt;br /&gt;
&lt;br /&gt;
 BAM_INDEX = {CURRENT_DIR}/umake_test.index &lt;br /&gt;
 UNIFORM_TARGET_BED = {CURRENT_DIR}/umake_test.bed &lt;br /&gt;
&lt;br /&gt;
where {CURRENT_DIR} is the absolute path to the directory that contains the index and bed files.  &lt;br /&gt;
&lt;br /&gt;
An additional option can be added in the configuration file: &lt;br /&gt;
&lt;br /&gt;
 OUT_DIR = {OUT_DIR} &lt;br /&gt;
&lt;br /&gt;
where {OUT_DIR} is the name of directory in which you want the output to be stored.  If you do not specify this in the configuration file, you will need to add an extra parameter when you run UMAKE in the next step. &lt;br /&gt;
&lt;br /&gt;
(More information about: [[Variant_Calling_Pipeline_(UMAKE)#Configuration_File|the configuration file]], [[Variant_Calling_Pipeline_(UMAKE)#Reference_Files|reference files]].) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==Further Information== &lt;br /&gt;
&lt;br /&gt;
[[Mapping_Pipeline|Mapping (Alignment) Pipeline]] &lt;br /&gt;
&lt;br /&gt;
[[Variant_Calling_Pipeline_(UMAKE)|Variant Calling Pipeline (UMAKE)]]&lt;/div&gt;</summary>
		<author><name>Ben Lerch</name></author>
	</entry>
</feed>