<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Tblackw</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Tblackw"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Tblackw"/>
	<updated>2026-09-24T11:16:42Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15060</id>
		<title>TOPMed Site Visit 2018</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15060"/>
		<updated>2018-09-28T14:14:36Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: Replaced content with &amp;quot; &amp;#039;&amp;#039;&amp;#039;Successfully occurred on Thursday September 13, 2018&amp;#039;&amp;#039;&amp;#039;&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
&#039;&#039;&#039;Successfully occurred on Thursday September 13, 2018&#039;&#039;&#039;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15056</id>
		<title>TOPMed Site Visit 2018</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15056"/>
		<updated>2018-09-04T15:43:55Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;One day:   Thursday, September 13, 2018,  8:30 am - 4:00 pm ?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Perhaps 5 1/2 total hours of presentations.&lt;br /&gt;
&lt;br /&gt;
= Draft Agenda =&lt;br /&gt;
&lt;br /&gt;
== Introduction and Overview of TOPMed data resources and services (30 minutes) == &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Goncalo Abecasis will present overview of IRC activies&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Introduce IRC personnel and their expertise&lt;br /&gt;
* Sequence for 130,000+ participants&lt;br /&gt;
* Variant calls and genotypes, phased and unphased&lt;br /&gt;
* Structural variant calls in progress&lt;br /&gt;
* BRAVO variant browser&lt;br /&gt;
* ENCORE analysis server&lt;br /&gt;
* TOPMed imputation reference panel&lt;br /&gt;
* Main developments in the past year, including security improvements, Manual, etc.&lt;br /&gt;
&lt;br /&gt;
== Characteristics of variants in data freeze 6 (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Hyun Min Kang and Jonathan LeFaive will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Overall numbers&lt;br /&gt;
* Differences by ancestry, study and sequencing center&lt;br /&gt;
* Allele frequencies of deleterious variants&lt;br /&gt;
* Genotype accuracy&lt;br /&gt;
* Compare harmonized versus sequencing center mappings&lt;br /&gt;
* Process for variant calling, genotyping, filtering, phasing and distribution&lt;br /&gt;
* Process for interim &#039;snapshot&#039; genotypes&lt;br /&gt;
&lt;br /&gt;
== Calling structural variants (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Our colleagues at Baylor will present this section. William Salerno?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Procedures and plans for structural variant calling&lt;br /&gt;
* Benefits of ensemble approach&lt;br /&gt;
* Initial results&lt;br /&gt;
* Data access mechanisms&lt;br /&gt;
* Anticipated data size&lt;br /&gt;
* Coordination with SNPs and indels&lt;br /&gt;
&lt;br /&gt;
== BRAVO variant browser helps to assess variant quality (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Daniel Taliun will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Purpose&lt;br /&gt;
* Which studies are included&lt;br /&gt;
* Usage / main features&lt;br /&gt;
* Both rare and common variants&lt;br /&gt;
* Improvements in the last year&lt;br /&gt;
* Access via an applications programming interface (API)&lt;br /&gt;
* View all information used in filtering&lt;br /&gt;
* Coordination with gnomAD&lt;br /&gt;
* Potential plans for PheWeb integration&lt;br /&gt;
&lt;br /&gt;
== Break (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== Summary of contract spending to date (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Denise Bianchi and Goncalo Abecasis will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Broad subdivisions:  personnel, cloud storage, cloud computing, hardware&lt;br /&gt;
* Divided between Task 1 and Task 2&lt;br /&gt;
&lt;br /&gt;
== Improved results from the latest TOPMed imputation panel (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Goncalo Abecasis will present this, unless Lukas Forer is available. Ketian Yu to prepare summaries of imputation quality&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
[[Media:nhlbi.4761.imputation.accuracy.2018aug31.pptx|&#039;&#039;&#039;(slides)&#039;&#039;&#039;]]&lt;br /&gt;
&lt;br /&gt;
* Principle of operation&lt;br /&gt;
* Measuring imputation quality in different populations&lt;br /&gt;
* Pushing the low frequency boundary&lt;br /&gt;
* Improved accuracy for African American and Latino samples&lt;br /&gt;
* Challenges and opportunities from collaboration and integration with NIH Commons / NHLBI Stage&lt;br /&gt;
&lt;br /&gt;
== Potential Population Genetics Update (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039; Check with Sebastian Zoellner&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Cloud access to TOPMed sequence data (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Tom Blackwell to take the lead on this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
[[Media:nhlbi.4768.fusera.slides.01.pdf|&#039;&#039;&#039;(slides)&#039;&#039;&#039;]]&lt;br /&gt;
&lt;br /&gt;
* NCBI&#039;s &#039;Fusera&#039; controlled access mechanism&lt;br /&gt;
* User perspective -- involves a Google or Amazon billing project&lt;br /&gt;
* What is needed for users to have a great overall experience?&lt;br /&gt;
* Education and training for users&lt;br /&gt;
&lt;br /&gt;
== Lunch (70 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== ENCORE analysis server (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Matthew Flickinger will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Principle of operation&lt;br /&gt;
* Releases only aggregate data summaries&lt;br /&gt;
* Visualizations help to assess results&lt;br /&gt;
* Data sharing and collaboration tools&lt;br /&gt;
* Capability to re-run previous jobs with new data&lt;br /&gt;
* SAIGE analysis gives accurate results in case-control studies&lt;br /&gt;
* Usage statistics comparing last 12-months to previous 12-months&lt;br /&gt;
* Some highlights from user survey results&lt;br /&gt;
&lt;br /&gt;
== Manuscript support (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Albert Vernon Smith will present overview for this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* How can and how is the IRC supporting TOPMed manuscripts and discoveries?&lt;br /&gt;
&lt;br /&gt;
* Overall TOPMed landmark paper&lt;br /&gt;
* Analysis of telomere length&lt;br /&gt;
* Mitochondrial DNA copy number&lt;br /&gt;
* Lipids analysis using TOPMed imputed genotypes&lt;br /&gt;
* UK BioBank with TOPMed imputation&lt;br /&gt;
* Context specific mutation rates&lt;br /&gt;
* Data sharing with Centers for Common Disease Genetics&lt;br /&gt;
&lt;br /&gt;
== Interactions with outside groups (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Albert Vernon Smith to take the lead on this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* NIH Data Commons &lt;br /&gt;
* NHLBI Data STAGE&lt;br /&gt;
* NHGRI Centers for Common Disease Genetics&lt;br /&gt;
* Global Alliance for Genomics and Health (GA4GH)&lt;br /&gt;
* NIMH Parkinsons Disease Consortium&lt;br /&gt;
&lt;br /&gt;
== Break (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== Future plans for the next Task Order (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== Feedback from NHLBI (40 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== Finish (3:50 pm) ==&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15055</id>
		<title>TOPMed Site Visit 2018</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=TOPMed_Site_Visit_2018&amp;diff=15055"/>
		<updated>2018-08-29T18:18:18Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: Created page with &amp;quot;&amp;#039;&amp;#039;&amp;#039;One day:   Thursday, September 13, 2018,  8:30 am - 3:00 pm ?&amp;#039;&amp;#039;&amp;#039;  Perhaps 4 1/3 total hours of presentations, if lunch is brought in.   (If you think I&amp;#039;m a bit optimistic o...&amp;quot;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;One day:   Thursday, September 13, 2018,  8:30 am - 3:00 pm ?&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
Perhaps 4 1/3 total hours of presentations, if lunch is brought in. &lt;br /&gt;
&lt;br /&gt;
(If you think I&#039;m a bit optimistic on the timings, I&#039;d have to agree with you.)&lt;br /&gt;
&lt;br /&gt;
= Draft Agenda =&lt;br /&gt;
&lt;br /&gt;
== Introduction and Overview of TOPMed data resources and services (30 minutes) == &lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Goncalo Abecasis will present overview of IRC activies&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Introduce IRC personnel and their expertise&lt;br /&gt;
* Sequence for 130,000+ participants&lt;br /&gt;
* Variant calls and genotypes, phased and unphased&lt;br /&gt;
* Structural variant calls in progress&lt;br /&gt;
* BRAVO variant browser&lt;br /&gt;
* ENCORE analysis server&lt;br /&gt;
* TOPMed imputation reference panel&lt;br /&gt;
* Main developments in the past year, including security improvements, Manual, etc.&lt;br /&gt;
&lt;br /&gt;
== Characteristics of variants in data freeze 6 (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Hyun Min Kang and Jonathan LeFaive will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Overall numbers&lt;br /&gt;
* Differences by ancestry, study and sequencing center&lt;br /&gt;
* Allele frequencies of deleterious variants&lt;br /&gt;
* Genotype accuracy&lt;br /&gt;
* Compare harmonized versus sequencing center mappings&lt;br /&gt;
* Process for variant calling, genotyping, filtering, phasing and distribution&lt;br /&gt;
* Process for interim &#039;snapshot&#039; genotypes&lt;br /&gt;
&lt;br /&gt;
== Calling structural variants (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Our colleagues at Baylor will present this section. William Salerno?&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Procedures and plans for structural variant calling&lt;br /&gt;
* Benefits of ensemble approach&lt;br /&gt;
* Initial results&lt;br /&gt;
* Data access mechanisms&lt;br /&gt;
* Anticipated data size&lt;br /&gt;
* Coordination with SNPs and indels&lt;br /&gt;
&lt;br /&gt;
== BRAVO variant browser helps to assess variant quality (20 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Daniel Taliun will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Purpose&lt;br /&gt;
* Which studies are included&lt;br /&gt;
* Usage / main features&lt;br /&gt;
* Both rare and common variants&lt;br /&gt;
* Improvements in the last year&lt;br /&gt;
* Access via an applications programming interface (API)&lt;br /&gt;
* View all information used in filtering&lt;br /&gt;
* Coordination with gnomAD&lt;br /&gt;
* Potential plans for PheWeb integration&lt;br /&gt;
&lt;br /&gt;
== Break (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== ENCORE analysis server (15 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Matthew Flickinger will present this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Principle of operation&lt;br /&gt;
* Releases only aggregate data summaries&lt;br /&gt;
* Visualizations help to assess results&lt;br /&gt;
* Data sharing and collaboration tools&lt;br /&gt;
* Capability to re-run previous jobs with new data&lt;br /&gt;
* SAIGE analysis gives accurate results in case-control studies&lt;br /&gt;
* Usage statistics comparing last 12-months to previous 12-months&lt;br /&gt;
* Some highlights from user survey results&lt;br /&gt;
&lt;br /&gt;
== Improved results from the latest TOPMed imputation panel (15 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Goncalo Abecasis will present this, unless Lukas Forrer is available. Ketian to prepare summaries of imputation quality&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* Principle of operation&lt;br /&gt;
* Measuring imputation quality in different populations&lt;br /&gt;
* Pushing the low frequency boundary&lt;br /&gt;
* Improved accuracy for African American and Latino samples&lt;br /&gt;
* Challenges and opportunities from collaboration and integration with NIH Commons / NHLBI Stage&lt;br /&gt;
&lt;br /&gt;
== Potential Population Genetics Update (15 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039; Check with Sebastian Zoellner&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
== Manuscript support (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Albert Vernon Smith will present overview for this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* How can and how is the IRC supporting TOPMed manuscripts and discoveries?&lt;br /&gt;
&lt;br /&gt;
* Overall TOPMed landmark paper&lt;br /&gt;
* Analysis of telomere length&lt;br /&gt;
* Mitochondrial DNA copy number&lt;br /&gt;
* Lipids analysis using TOPMed imputed genotypes&lt;br /&gt;
* UK BioBank with TOPMed imputation&lt;br /&gt;
* Context specific mutation rates&lt;br /&gt;
* Data sharing with Centers for Common Disease Genetics&lt;br /&gt;
&lt;br /&gt;
== Lunch (45 minutes) ==&lt;br /&gt;
&lt;br /&gt;
== Interactions with outside groups (30 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Albert Vernon Smith to take the lead on this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* NIH Data Commons &lt;br /&gt;
* NHLBI Data STAGE&lt;br /&gt;
* NHGRI Centers for Common Disease Genetics&lt;br /&gt;
* Global Alliance for Genomics and Health (GA4GH)&lt;br /&gt;
* NIMH Parkinsons Disease Consortium&lt;br /&gt;
&lt;br /&gt;
== Cloud access to TOPMed sequence data (15 minutes) ==&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Tom Blackwell to take the lead on this section&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* NCBI&#039;s &#039;Fusera&#039; controlled access mechanism&lt;br /&gt;
* User perspective -- involves a Google or Amazon billing project&lt;br /&gt;
* What is needed for users to have a great overall experience?&lt;br /&gt;
* Education and training for users&lt;br /&gt;
&lt;br /&gt;
== Feedback from NHLBI (60 minutes) ==&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1194</id>
		<title>Generic Exome Analysis Plan</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1194"/>
		<updated>2010-04-21T12:01:52Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This page outlines a generic plan for analysis of a whole exome sequencing project. The idea is that the points listed here might serve as a starting point for discussion of the analyses needed in a specific project.&lt;br /&gt;
&lt;br /&gt;
The initial version of this document was prepared with input from Shamil Sunayev, Ron Do and Goncalo Abecasis.&lt;br /&gt;
&lt;br /&gt;
= Read Mapping and Variant Calling =&lt;br /&gt;
&lt;br /&gt;
The first step in any analysis is to map sequence reads, callibrate base qualities, and call variants. Even at this stage, some simple quality metrics can be evaluated and will help identify potentially problematic samples.&lt;br /&gt;
&lt;br /&gt;
== Prior to Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Evaluate Base Composition Along Reads&lt;br /&gt;
: Calculate the proportion of A, C, G, T bases along each read. Flag runs with evidence of unusual patterns of base composition compared to the target genome.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Machine Quality Scores Along Reads&lt;br /&gt;
: Calculate average quality scores per position. Flag runs with evidence of unusual quality score distributions.&lt;br /&gt;
&lt;br /&gt;
; Calculate Number of Reads &lt;br /&gt;
: Calculate the input number of reads and number of bases for each sequenced sample&lt;br /&gt;
&lt;br /&gt;
== Read Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Map Reads with Appropriate Read Mapper&lt;br /&gt;
: Currently, [bio-bwa.sourceforge.net/bwa.shtm BWA] is a convenient, widely used read mapper.&lt;br /&gt;
&lt;br /&gt;
; Mark Duplicate Reads&lt;br /&gt;
: Duplicate reads, whether generated as PCR artifacts during library preparation or optical duplicates during image analysis and base-calling, can mislead variant calling algorithms. To avoid problems, one typically removes from consideration all reads that appear to map at exactly the same location (for paired ends, only reads for which both ends map to the same locations are excluded.)&lt;br /&gt;
&lt;br /&gt;
; Basic Mapping Statistics&lt;br /&gt;
: We should tally the overall proportion of mapped reads.&lt;br /&gt;
: We should also tally the proportion of reads that map&lt;br /&gt;
:* Inside the target regions&lt;br /&gt;
:* Near the target regions (defined as within 200bp of each target)&lt;br /&gt;
:* Elsewhere in the genome (defined as regions that are &amp;gt;200bp from each target)&lt;br /&gt;
&lt;br /&gt;
; Recalibrate Base Quality Scores&lt;br /&gt;
: Base quality scores can be updated by comparing sites that are unlikely to vary (such as those not currently reported as variants in dbSNP or in the most recent [[1000 Genome Project]] analyses.&lt;br /&gt;
&lt;br /&gt;
; Update Base Quality Score Metrics&lt;br /&gt;
: Generate new curves with base quality scores per position.&lt;br /&gt;
: Calculate the number of mapped bases that reach at least Q20. Potentially, calculate Q20 &#039;&#039;equivalent&#039;&#039; bases by summing the quality scores for bases with base quality &amp;gt;Q20 and dividing the total by 20.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Coverage as Function of GC Content&lt;br /&gt;
: For each target region, calculate read depth and also the proportion of GC bases in the reference genome. Flag samples where coverage varies strongly as a function of GC content.&lt;br /&gt;
&lt;br /&gt;
== Verify Sample Identities ==&lt;br /&gt;
&lt;br /&gt;
; Verify that Each Sequenced Sample Matches Prior Information &lt;br /&gt;
: We should verify that each sequenced sample matches previous genotypes for that sample. Ideally, this should be done using a likelihood based approach that can also identify potentially contaminated samples. &lt;br /&gt;
&lt;br /&gt;
; Adjudicate Mislabeled Samples&lt;br /&gt;
: If any samples that don&#039;t match prior genotype data are encountered, the read data for these samples should be compared to all available samples to identify potential sample mixups. &lt;br /&gt;
&lt;br /&gt;
; Identify Potentially Related Samples&lt;br /&gt;
: Perhaps this step should be done *after* variant calling?&lt;br /&gt;
&lt;br /&gt;
== Generate Variant Calls ==&lt;br /&gt;
&lt;br /&gt;
; Generate Initial Set of Variant Calls&lt;br /&gt;
: Variant calls should be generated taking into account all available samples simultaneously. We should consider variants calls that fall in target regions but also those that fall near the target. An open question is whether off target variant calls will be trustworthy.&lt;br /&gt;
&lt;br /&gt;
; Generate Linkage Disequilibrium Aware Set of Variant Calls&lt;br /&gt;
: A typical exome call set might include many sites where no variant is called due to low coverage. If we can integrate samples with previous [[GWAS]] data, we should be able to generate an update and much improved set of variant calls for each individual. Due to limitations in current calling methods, the quality of variant call sets is expected to increase substantially if multiple variant callers are used and their results are merged democratically.&lt;br /&gt;
&lt;br /&gt;
; Annotate Functional Impact of Each Variant&lt;br /&gt;
: Called variants should be annotated according to their potential function. At a minimum, we should distinguish synonymous, non-synonymous, conserved splice site, 5&#039;UTR, 3&#039;UTR and other variants. Ideally, we should also assess [[SIFT]] and [[PolyPhen]] scores for conserved variants.&lt;br /&gt;
&lt;br /&gt;
; Calculate Overall Frequency Spectrum&lt;br /&gt;
: Calculate observed frequency spectrum and compare to neutral expectations (which are that the number of variants should be roughly proportional to &#039;&#039;1/n&#039;&#039;, where &#039;&#039;n&#039;&#039; is the number of minor alleles).&lt;br /&gt;
&lt;br /&gt;
; Annotate Overall Variant Characteristics&lt;br /&gt;
: Calculate overall ratio of transitions to transversions, separately for coding and non-coding variants. Within coding variants, analyse synonymous and non-synonymous variants separately. &lt;br /&gt;
&lt;br /&gt;
; CpG Sites&lt;br /&gt;
: Calculate the rate of per base pair heterozygosity at potential CpG sites and compare this to other sites.&lt;br /&gt;
&lt;br /&gt;
; Tabulate for Each Sample&lt;br /&gt;
:* The number of synonymous and non-synonymous variants&lt;br /&gt;
:* The number of unique and shared variants&lt;br /&gt;
:* The number of transitions and transversions&lt;br /&gt;
&lt;br /&gt;
== Evaluate Variant Calls for Reference Samples ==&lt;br /&gt;
&lt;br /&gt;
; If Previously Sequenced Samples Available, Assess Variant Calls There&lt;br /&gt;
: Assess concordance rate at non-reference sites&lt;br /&gt;
&lt;br /&gt;
; If Duplicate Samples are Available, Assess Concordance&lt;br /&gt;
: Assess concordance rate at non-reference sites.&lt;br /&gt;
&lt;br /&gt;
; If Nuclear Families are Available, Assess Mendelian Consistency&lt;br /&gt;
: When reporting rates of Mendelian inconsistencies, report these not as a fraction of all sites, but as a fraction of sites with at least one non-reference call in the trio.&lt;br /&gt;
&lt;br /&gt;
== Variant Filters ==&lt;br /&gt;
&lt;br /&gt;
Initial sets of SNP calls invariably will include many false positives. The fraction of false positives among all variants called is likely to increase as more and more samples are sequenced. To keep the fraction of false positives under control, it is important to both apply an increasingly strict set of quality control filters but also to experimentally validate some newly discovered variants. Many of these filters are currently implemented in [[GATK]].&lt;br /&gt;
&lt;br /&gt;
; Mapping Quality Filter&lt;br /&gt;
: Consider removing variants at sites with low mapping quality scores. Even if the average mapping quality score for a site is high, consider removing variants at sites where a noticeable fraction of reads have low mapping quality scores.&lt;br /&gt;
&lt;br /&gt;
; Allele Balance Filter&lt;br /&gt;
: Among individuals who are assigned an heterozygous genotype, check the proportion of reads supporting each allele. Consider filtering out variants where one of the alleles accounts for &amp;lt;30% of reads.&lt;br /&gt;
&lt;br /&gt;
; Local Realignment Filter&lt;br /&gt;
: Consider removing single nucleotide variants at sites where local realignment of all covering sequence reads suggests that a length polymorphism is present in the population.&lt;br /&gt;
&lt;br /&gt;
; Read Depth&lt;br /&gt;
: Consider filtering out variants at sites where total read depth is unusually low. Unfortunately, capture protocols introduce very large amounts of variation in sequencing depth and it is usually not possible to accurately filter out sites sequenced at very high depth. &lt;br /&gt;
&lt;br /&gt;
= Special Considerations for Admixed Samples =&lt;br /&gt;
&lt;br /&gt;
; Estimate Local Ancestry Using GWAS Data&lt;br /&gt;
: For studies that include admixed samples, we should estimate local ancestry using GWAS data. If GWAS are not available, it is strongly recommended that these data should be generated. In principle, local ancestry estimates can be generated even before exome sequencing is complete.&lt;br /&gt;
&lt;br /&gt;
; Estimate Global Ancestry Covariates Using PCA or MDS Analysis&lt;br /&gt;
&lt;br /&gt;
= Initial Association Analyses =&lt;br /&gt;
&lt;br /&gt;
We anticipate that, at least early on, the initial association analysis of whole exome datasets in the context of complex trait association studies will focus on identifying and resolving quality control issues that might result in unexpected artifacts.&lt;br /&gt;
&lt;br /&gt;
== Initial Single SNP Tests ==&lt;br /&gt;
&lt;br /&gt;
In principle, these tests only have power for common variants. In practice, particularly when permutation based methods are used to assess significance, there should be little loss in power by testing all variants using single SNP tests. These tests should include:&lt;br /&gt;
&lt;br /&gt;
; Logistic Regression Based Tests for Discrete Traits&lt;br /&gt;
: Discrete outcomes should be evaluated using logistic regression. It is important to include appropriate covariates. For most traits, these might include age and sex and, potentially, principal components of ancestry. &lt;br /&gt;
&lt;br /&gt;
; Linear Regression Based Tests for Quantitative Traits&lt;br /&gt;
: For quantitative traits that are not strongly selected, linear regression based tests can also be used. Again, it is important to include an appropriate set of covariates. For many quantitative traits, it may be a very good idea to normalize traits to minimize the impact of outliers on association results.&lt;br /&gt;
&lt;br /&gt;
; Using genotypes as outcomes for Selected Quantitative Traits&lt;br /&gt;
: For both discrete and quantitative traits, analysis can be repeated using genotypes (scored as 0, 1 and 2) as outcomes and phenotypes as predictors.&lt;br /&gt;
&lt;br /&gt;
== Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial single SNP tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
== Burden Tests ==&lt;br /&gt;
&lt;br /&gt;
The same analyses that were originally carried for single variants should be carried out for groups of rare variants. In principle, one could simply use the presence of a rare variant (or a particular class of rare variant, such as a non-synonymous variant or a newly discovered variant) as a predictor and repeat the logistic regression, linear regression or genotype regression described above. For an initial pass, I think the precise form of this analysis is not critical, because the next step is to...&lt;br /&gt;
&lt;br /&gt;
== More Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial burden tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
= Visualize Results =&lt;br /&gt;
&lt;br /&gt;
A number of displays will likely be useful. Probably these should include:&lt;br /&gt;
&lt;br /&gt;
* Manhattan Plots&lt;br /&gt;
* [[LocusZoom]] Plots&lt;br /&gt;
* Q-Q Plots&lt;br /&gt;
&lt;br /&gt;
= Think You Are Done? =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;No way!!!&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
== Indels and Structural Variants ==&lt;br /&gt;
&lt;br /&gt;
You still need a plan to call and evaluate short insertions and deletions as well as larger structural variants.&lt;br /&gt;
&lt;br /&gt;
== Pathway Based Analyses ==&lt;br /&gt;
&lt;br /&gt;
Carry out analyses that include groups of genes with similar biological function (for example, according to [[Gene Ontology]] or [[Kyoto Encyclopedia of Genes and Genomes]] annotations.&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1193</id>
		<title>Generic Exome Analysis Plan</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1193"/>
		<updated>2010-04-21T11:54:41Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This page outlines a generic plan for analysis of a whole exome sequencing project. The idea is that the points listed here might serve as a starting point for discussion of the analyses needed in a specific project.&lt;br /&gt;
&lt;br /&gt;
The initial version of this document was prepared with input from Shamil Sunayev, Ron Do and Goncalo Abecasis.&lt;br /&gt;
&lt;br /&gt;
= Read Mapping and Variant Calling =&lt;br /&gt;
&lt;br /&gt;
The first step in any analysis is to map sequence reads, callibrate base qualities, and call variants. Even at this stage, some simple quality metrics can be evaluated and will help identify potentially problematic samples.&lt;br /&gt;
&lt;br /&gt;
== Prior to Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Evaluate Base Composition Along Reads&lt;br /&gt;
: Calculate the proportion of A, C, G, T bases along each read. Flag runs with evidence of unusual patterns of base composition compared to the target genome.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Machine Quality Scores Along Reads&lt;br /&gt;
: Calculate average quality scores per position. Flag runs with evidence of unusual quality score distributions.&lt;br /&gt;
&lt;br /&gt;
; Calculate Number of Reads &lt;br /&gt;
: Calculate the input number of reads and number of bases for each sequenced sample&lt;br /&gt;
&lt;br /&gt;
== Read Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Map Reads with Appropriate Read Mapper&lt;br /&gt;
: Currently, [bio-bwa.sourceforge.net/bwa.shtm BWA] is a convenient, widely used read mapper.&lt;br /&gt;
&lt;br /&gt;
; Mark Duplicate Reads&lt;br /&gt;
: Duplicate reads, whether generated as PCR artifacts during library preparation or optical duplicates during image analysis and base-calling, can mislead variant calling algorithms. To avoid problems, one typically removes from consideration all reads that appear to map at exactly the same location (for paired ends, only reads for which both ends map to the same locations are excluded.)&lt;br /&gt;
&lt;br /&gt;
; Basic Mapping Statistics&lt;br /&gt;
: We should tally the overall proportion of mapped reads.&lt;br /&gt;
: We should also tally the proportion of reads that map&lt;br /&gt;
:* Inside the target regions&lt;br /&gt;
:* Near the target regions (defined as within 200bp of each target)&lt;br /&gt;
:* Elsewhere in the genome (defined as regions that are &amp;gt;200bp from each target)&lt;br /&gt;
&lt;br /&gt;
; Recalibrate Base Quality Scores&lt;br /&gt;
: Base quality scores can be updated by comparing sites that are unlikely to vary (such as those not currently reported as variants in dbSNP or in the most recent [[1000 Genome Project]] analyses.&lt;br /&gt;
&lt;br /&gt;
; Update Base Quality Score Metrics&lt;br /&gt;
: Generate new curves with base quality scores per position.&lt;br /&gt;
: Calculate the number of mapped bases that reach at least Q20. Potentially, calculate Q20 &#039;&#039;equivalent&#039;&#039; bases by summing the quality scores for bases with base quality &amp;gt;Q20 and dividing the total by 20.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Coverage as Function of GC Content&lt;br /&gt;
: For each target region, calculate read depth and also the proportion of GC bases in the reference genome. Flag samples where coverage varies strongly as a function of GC content.&lt;br /&gt;
&lt;br /&gt;
== Verify Sample Identities ==&lt;br /&gt;
&lt;br /&gt;
; Verify that Each Sequenced Sample Matches Prior Information &lt;br /&gt;
: We should verify that each sequenced sample matches previous genotypes for that sample. Ideally, this should be done using a likelihood based approach that can also identify potentially contaminated samples. &lt;br /&gt;
&lt;br /&gt;
; Adjudicate Mislabeled Samples&lt;br /&gt;
: If any samples that don&#039;t match prior genotype data are encountered, the read data for these samples should be compared to all available samples to identify potential sample mixups. &lt;br /&gt;
&lt;br /&gt;
; Identify Potentially Related Samples&lt;br /&gt;
: Perhaps this step should be done *after* variant calling?&lt;br /&gt;
&lt;br /&gt;
== Generate Variant Calls ==&lt;br /&gt;
&lt;br /&gt;
; Generate Initial Set of Variant Calls&lt;br /&gt;
: Variant calls should be generated taking into account all available samples simultaneously. We should consider variants calls that fall in target regions but also those that fall near the target. An open question is whether off target variant calls will be trustworthy.&lt;br /&gt;
&lt;br /&gt;
; Generate Linkage Disequilibrium Aware Set of Variant Calls&lt;br /&gt;
: A typical exome call set might include many sites where no variant is called due to low coverage. If we can integrate samples with previous [[GWAS]] data, we should be able to generate an update and much improved set of variant calls for each individual. Due to limitations in current calling methods, the quality of variant call sets is expected to increase substantially if multiple variant callers are used and their results are merged democratically.&lt;br /&gt;
&lt;br /&gt;
; Annotate Functional Impact of Each Variant&lt;br /&gt;
: Called variants should be annotated according to their potential function. At a minimum, we should distinguish synonymous, non-synonymous, conserved splice site, 5&#039;UTR, 3&#039;UTR and other variants. Ideally, we should also assess [[SIFT]] and [[PolyPhen]] scores for conserved variants.&lt;br /&gt;
&lt;br /&gt;
; Calculate Overall Frequency Spectrum&lt;br /&gt;
: Calculate observed frequency spectrum and compare to neutral expectations (which are that the number of variants should be roughly proportional to &#039;&#039;1/n&#039;&#039;, where &#039;&#039;n&#039;&#039; is the number of minor alleles).&lt;br /&gt;
&lt;br /&gt;
; Annotate Overall Variant Characteristics&lt;br /&gt;
: Calculate overall ratio of transitions to transversions, separately for coding and non-coding variants. Within coding variants, analyse synonymous and non-synonymous variants separately. &lt;br /&gt;
&lt;br /&gt;
; CpG Sites&lt;br /&gt;
: Calculate the rate of per base pair heterozygosity at potential CpG sites and compare this to other sites.&lt;br /&gt;
&lt;br /&gt;
; Tabulate for Each Sample&lt;br /&gt;
:* The number of synonymous and non-synonymous variants&lt;br /&gt;
:* The number of unique and shared variants&lt;br /&gt;
:* The number of transitions and transversions&lt;br /&gt;
&lt;br /&gt;
== Evaluate Variant Calls for Reference Samples ==&lt;br /&gt;
&lt;br /&gt;
; If Previously Sequenced Samples Available, Assess Variant Calls There&lt;br /&gt;
: Assess concordance rate at non-reference sites&lt;br /&gt;
&lt;br /&gt;
; If Duplicate Samples are Available, Assess Concordance&lt;br /&gt;
: Assess concordance rate at non-reference sites.&lt;br /&gt;
&lt;br /&gt;
; If Nuclear Families are Available, Assess Mendelian Consistency&lt;br /&gt;
: When reporting rates of Mendelian inconsistencies, report these not as a fraction of all sites, but as a fraction of sites with at least one non-reference call in the trio.&lt;br /&gt;
&lt;br /&gt;
== Variant Filters ==&lt;br /&gt;
&lt;br /&gt;
Initial sets of SNP calls invariably will include many false positives. The fraction of false positives among all variants called is likely to increase as more and more samples are sequenced. To keep the fraction of false positives under control, it is important to both apply an increasingly strict set of quality control filters but also to experimentally validate some newly discovered variants. Many of these filters are currently implemented in [[GATK]].&lt;br /&gt;
&lt;br /&gt;
; Mapping Quality Filter&lt;br /&gt;
: Consider removing variants at sites with low mapping quality scores. Even if the average mapping quality score for a site is high, consider removing variants at sites where a noticeable fraction of reads have low mapping quality scores.&lt;br /&gt;
&lt;br /&gt;
; Allele Balance Filter&lt;br /&gt;
: Among individuals who are assigned an heterozygous genotype, check the proportion of reads supporting each allele. Consider filtering out variants where one of the alleles accounts for &amp;lt;30% of reads.&lt;br /&gt;
&lt;br /&gt;
; Local Realignment Filter&lt;br /&gt;
: Consider removing single nucleotide variants at sites where local realignment of all covering sequence reads suggests that a length polymorphism is present in the population.&lt;br /&gt;
&lt;br /&gt;
; Read Depth&lt;br /&gt;
: Consider filtering out variants at sites where total read depth is unusually low. Unfortunately, capture protocols introduce very large amounts of variation in sequencing depth and it is usually not possible to accurately filter out sites sequenced at very high depth. &lt;br /&gt;
&lt;br /&gt;
= Special Considerations for Admixed Samples =&lt;br /&gt;
&lt;br /&gt;
; Estimate Local Ancestry Using GWAS Data&lt;br /&gt;
: For studies that include admixed samples, we should estimate local ancestry using GWAS data. If GWAS are not available, it is strongly recommended that these data should be generated. In principle, local ancestry estimates can be generated even before exome sequencing is complete.&lt;br /&gt;
&lt;br /&gt;
; Estimate Global Ancestry Covariates Using PCA or MDS Analysis&lt;br /&gt;
&lt;br /&gt;
= Initial Association Analyses =&lt;br /&gt;
&lt;br /&gt;
We anticipate that, at least early on, the initial association analysis of whole exome datasets in the context of complex trait association studies will focus on identifying and resolving quality control issues that might result in unexpected artifacts.&lt;br /&gt;
&lt;br /&gt;
== Initial Single SNP Tests ==&lt;br /&gt;
&lt;br /&gt;
In principle, these tests only have power for common variants. In practice, particularly when permutation based methods are used to assess significance, there should be little loss in power by testing all variants using single SNP tests. These tests should include:&lt;br /&gt;
&lt;br /&gt;
; Logistic Regression Based Tests for Discrete Traits&lt;br /&gt;
: Discrete outcomes should be evaluated using logistic regression. It is important to include appropriate covariates. For most traits, these might include age and sex and, potentially, principal components of ancestry. &lt;br /&gt;
&lt;br /&gt;
; Linear Regression Based Tests for Quantitative Traits&lt;br /&gt;
: For quantitative traits that are not strongly selected, linear regression based tests can also be used. Again, it is important to include an appropriate set of covariates. For many quantitative traits, it may be a very good idea to normalize traits to minimize the impact of outliers on association results.&lt;br /&gt;
&lt;br /&gt;
; Using genotypes as outcomes for Selected Quantitative Traits&lt;br /&gt;
: For both discrete and quantitative traits, analysis can be repeated using genotypes (scored as 0, 1 and 2) as outcomes and phenotypes as predictors.&lt;br /&gt;
&lt;br /&gt;
== Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial single SNP tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
== Burden Tests ==&lt;br /&gt;
&lt;br /&gt;
The same analyses that were originally carried for single variants should be carried out for groups of rare variants. In principle, one could simple use the presence of a rare variant (or a particular class of rare variant, such as a non-synonymous variant or a newly discovered variant) as a predictor and repeat the logistic regression, linear regression or genotype regression described above. For an initial pass, I think the precise form of this analysis is not critical, because the next step is to...&lt;br /&gt;
&lt;br /&gt;
== More Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial burden tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
= Visualize Results =&lt;br /&gt;
&lt;br /&gt;
A number of displays will likely be useful. Probably these should include:&lt;br /&gt;
&lt;br /&gt;
* Manhattan Plots&lt;br /&gt;
* [[LocusZoom]] Plots&lt;br /&gt;
* Q-Q Plots&lt;br /&gt;
&lt;br /&gt;
= Think You Are Done? =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;No way!!!&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
== Indels and Structural Variants ==&lt;br /&gt;
&lt;br /&gt;
You still need a plan to call and evaluate short insertions and deletions as well as larger structural variants.&lt;br /&gt;
&lt;br /&gt;
== Pathway Based Analyses ==&lt;br /&gt;
&lt;br /&gt;
Carry out analyses that include groups of genes with similar biological function (for example, according to [[Gene Ontology]] or [[Kyoto Encyclopedia of Genes and Genomes]] annotations.&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1192</id>
		<title>Generic Exome Analysis Plan</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Generic_Exome_Analysis_Plan&amp;diff=1192"/>
		<updated>2010-04-21T11:51:20Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This page outlines a generic plan for analysis of a whole exome sequencing project. The idea is that the points listed here might serve as a starting point for discussion of the analyses needed in a specific project.&lt;br /&gt;
&lt;br /&gt;
The initial version of this document was prepared with input from Shamil Sunayev, Ron Do and Goncalo Abecasis.&lt;br /&gt;
&lt;br /&gt;
= Read Mapping and Variant Calling =&lt;br /&gt;
&lt;br /&gt;
The first step in any analysis is to map sequence reads, callibrate base qualities, and call variants. Even at this stage, some simple quality metrics can be evaluated and will help identify potentially problematic samples.&lt;br /&gt;
&lt;br /&gt;
== Prior to Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Evaluate Base Composition Along Reads&lt;br /&gt;
: Calculate the proportion of A, C, G, T bases along each read. Flag runs with evidence of unusual patterns of base composition compared to the target genome.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Machine Quality Scores Along Reads&lt;br /&gt;
: Calculate average quality scores per position. Flag runs with evidence of unusual quality score distributions.&lt;br /&gt;
&lt;br /&gt;
; Calculate Number of Reads &lt;br /&gt;
: Calculate the input number of reads and number of bases for each sequenced sample&lt;br /&gt;
&lt;br /&gt;
== Read Mapping ==&lt;br /&gt;
&lt;br /&gt;
; Map Reads with Appropriate Read Mapper&lt;br /&gt;
: Currently, [bio-bwa.sourceforge.net/bwa.shtm BWA] is a convenient, widely used read mapper.&lt;br /&gt;
&lt;br /&gt;
; Mark Duplicate Reads&lt;br /&gt;
: Duplicate reads, whether generated as PCR artifacts during library preparation or optical duplicates during image analysis and base-calling, can mislead variant calling algorithms. To avoid problems, one typically removes from consideration all reads that appear to map at exactly the same location (for paired ends, only reads for which both ends map to the same locations are excluded.)&lt;br /&gt;
&lt;br /&gt;
; Basic Mapping Statistics&lt;br /&gt;
: We should tally the overall proportion of mapped reads.&lt;br /&gt;
: We should also tally the proportion of reads that map&lt;br /&gt;
:* Inside the target regions&lt;br /&gt;
:* Near the target regions (defined as within 200bp of each target)&lt;br /&gt;
:* Elsewhere in the genome (defined as regions that are &amp;gt;200bp from each target)&lt;br /&gt;
&lt;br /&gt;
; Recalibrate Base Quality Scores&lt;br /&gt;
: Base quality scores can be updated by comparing sites that are unlikely to vary (such as those not currently reported as variants in dbSNP or in the most recent [[1000 Genome Project]] analyses.&lt;br /&gt;
&lt;br /&gt;
; Update Base Quality Score Metrics&lt;br /&gt;
: Generate new curves with base quality scores per position.&lt;br /&gt;
: Calculate the number of mapped bases that reach at least Q20. Potentially, calculate Q20 &#039;&#039;equivalent&#039;&#039; bases by summing the quality scores for bases with base quality &amp;gt;Q20 and dividing the total by 20.&lt;br /&gt;
&lt;br /&gt;
; Evaluate Coverage as Function of GC Content&lt;br /&gt;
: For each target region, calculate read depth and also the proportion of GC bases in the reference genome. Flag samples where coverage varies strongly as a function of GC content.&lt;br /&gt;
&lt;br /&gt;
== Verify Sample Identities ==&lt;br /&gt;
&lt;br /&gt;
; Verify that Each Sequenced Sample Matches Prior Information &lt;br /&gt;
: We should verify that each sequenced sample matches previous genotypes for that sample. Ideally, this should be done using a likelihood based approach that can also identify potentially contaminated samples. &lt;br /&gt;
&lt;br /&gt;
; Adjudicate Mislabeled Samples&lt;br /&gt;
: If any samples that don&#039;t match prior genotype data are encountered, the read data for these samples should be compared to all available samples to identify potential sample mixups. &lt;br /&gt;
&lt;br /&gt;
; Identify Potentially Related Samples&lt;br /&gt;
: Perhaps this step should be done *after* variant calling?&lt;br /&gt;
&lt;br /&gt;
== Generate Variant Calls ==&lt;br /&gt;
&lt;br /&gt;
; Generate Initial Set of Variant Calls&lt;br /&gt;
: Variant calls should be generated taking into account all available samples simultaneously. We should consider variants calls that fall in target regions but also those that fall near the target. An open question is whether off target variant calls will be trustworthy.&lt;br /&gt;
&lt;br /&gt;
; Generate Linkage Disequilibrium Aware Set of Variant Calls&lt;br /&gt;
: A typical exome call set might include many sites where no variant is called due to low coverage. If we can integrate samples with previous [[GWAS]] data, we should be able to generate an update and much improved set of variant calls for each individual. Due to limitations in current calling methods, the quality of variant call sets is expected to increase substantially if multiple variant callers are used and their results are merged democratically.&lt;br /&gt;
&lt;br /&gt;
; Annotate Functional Impact of Each Variant&lt;br /&gt;
: Called variants should be annotated according to their potential function. At a minimum, we should distinguish synonymous, non-synonymous, conserved splice site, 5&#039;UTR, 3&#039;UTR and other variants. Ideally, we should also assess [[SIFT]] and [[PolyPhen]] scores for conserved variants.&lt;br /&gt;
&lt;br /&gt;
; Calculate Overall Frequency Spectrum&lt;br /&gt;
: Calculate observed frequency spectrum and compare to neutral expectations (which are that the number of variants should be roughly proportional to &#039;&#039;1/n&#039;&#039;, where &#039;&#039;n&#039;&#039; is the number of minor alleles).&lt;br /&gt;
&lt;br /&gt;
; Annotate Overall Variant Characteristics&lt;br /&gt;
: Calculate overall ratio of transitions to transversions, separately for coding and non-coding variants. Within coding variants, analyse synonymous and non-synonymous variants separately. &lt;br /&gt;
&lt;br /&gt;
; CpG Sites&lt;br /&gt;
: Calculate the rate of per base pair heterozygosity at potential CpG sites and compare this to other sites.&lt;br /&gt;
&lt;br /&gt;
; Tabulate for Each Sample&lt;br /&gt;
:* The number of synonymous and non-synonymous variants&lt;br /&gt;
:* The number of unique and shared variants&lt;br /&gt;
:* The number of transitions and transversions&lt;br /&gt;
&lt;br /&gt;
== Evaluate Variant Calls for Reference Samples ==&lt;br /&gt;
&lt;br /&gt;
; If Previously Sequenced Samples Available, Assess Variant Calls There&lt;br /&gt;
: Assess concordance rate at non-reference sites&lt;br /&gt;
&lt;br /&gt;
; If Duplicate Samples are Available, Assess Concordance&lt;br /&gt;
: Assess concordance rate at non-reference sites.&lt;br /&gt;
&lt;br /&gt;
; If Nuclear Families are Available, Assess Mendelian Consistency&lt;br /&gt;
: When reporting rates of Mendelian inconsistencies, report these not as a fraction of all sites, but as a fraction of sites with at least one non-reference call in the trio.&lt;br /&gt;
&lt;br /&gt;
== Variant Filters ==&lt;br /&gt;
&lt;br /&gt;
Initial sets of SNP calls invariably will include many false positives. The fraction of false positives among all variants called is likely to increase as more and more samples are sequenced. To keep the fraction of false positives under control, it is important to both apply an increasingly strict set of quality control filters but also to experimentally validate some newly discovered variants. Many of these filters are currently implemented in [[GATK]].&lt;br /&gt;
&lt;br /&gt;
; Mapping Quality Filter&lt;br /&gt;
: Consider removing variants at sites with low mapping quality scores. Even if the average mapping quality score for a site is high, consider removing variants at sites where a noticeable fraction of reads have low mapping quality scores.&lt;br /&gt;
&lt;br /&gt;
; Allele Balance Filter&lt;br /&gt;
: Among individuals who are assigned an heterozygous genotype, check the proportion of reads supporting each allele. Consider filtering out variants where one of the alleles accounts for &amp;lt;30% of reads.&lt;br /&gt;
&lt;br /&gt;
; Local Realignment Filter&lt;br /&gt;
: Consider removing single nucleotide variants at sites where local realignment of all covering sequence reads suggests that a short indel polymorphism is present in the population.&lt;br /&gt;
&lt;br /&gt;
; Read Depth&lt;br /&gt;
: Consider filtering out variants at sites where total read depth is unusually low. Unfortunately, capture protocols introduce very large amounts of variation in sequencing depth and it is usually not possible to accurately filter out sites sequenced at very high depth. &lt;br /&gt;
&lt;br /&gt;
= Special Considerations for Admixed Samples =&lt;br /&gt;
&lt;br /&gt;
; Estimate Local Ancestry Using GWAS Data&lt;br /&gt;
: For studies that include admixed samples, we should estimate local ancestry using GWAS data. If GWAS are not available, it is strongly recommended that these data should be generated. In principle, local ancestry estimates can be generated even before exome sequencing is complete.&lt;br /&gt;
&lt;br /&gt;
; Estimate Global Ancestry Covariates Using PCA or MDS Analysis&lt;br /&gt;
&lt;br /&gt;
= Initial Association Analyses =&lt;br /&gt;
&lt;br /&gt;
We anticipate that, at least early on, the initial association analysis of whole exome datasets in the context of complex trait association studies will focus on identifying and resolving quality control issues that might result in unexpected artifacts.&lt;br /&gt;
&lt;br /&gt;
== Initial Single SNP Tests ==&lt;br /&gt;
&lt;br /&gt;
In principle, these tests only have power for common variants. In practice, particularly when permutation based methods are used to assess significance, there should be little loss in power by testing all variants using single SNP tests. These tests should include:&lt;br /&gt;
&lt;br /&gt;
; Logistic Regression Based Tests for Discrete Traits&lt;br /&gt;
: Discrete outcomes should be evaluated using logistic regression. It is important to include appropriate covariates. For most traits, these might include age and sex and, potentially, principal components of ancestry. &lt;br /&gt;
&lt;br /&gt;
; Linear Regression Based Tests for Quantitative Traits&lt;br /&gt;
: For quantitative traits that are not strongly selected, linear regression based tests can also be used. Again, it is important to include an appropriate set of covariates. For many quantitative traits, it may be a very good idea to normalize traits to minimize the impact of outliers on association results.&lt;br /&gt;
&lt;br /&gt;
; Using genotypes as outcomes for Selected Quantitative Traits&lt;br /&gt;
: For both discrete and quantitative traits, analysis can be repeated using genotypes (scored as 0, 1 and 2) as outcomes and phenotypes as predictors.&lt;br /&gt;
&lt;br /&gt;
== Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial single SNP tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
== Burden Tests ==&lt;br /&gt;
&lt;br /&gt;
The same analyses that were originally carried for single variants should be carried out for groups of rare variants. In principle, one could simple use the presence of a rare variant (or a particular class of rare variant, such as a non-synonymous variant or a newly discovered variant) as a predictor and repeat the logistic regression, linear regression or genotype regression described above. For an initial pass, I think the precise form of this analysis is not critical, because the next step is to...&lt;br /&gt;
&lt;br /&gt;
== More Q-Q Plots ==&lt;br /&gt;
&lt;br /&gt;
After carrying out initial burden tests, generate Q-Q plots for each analysis. Verify that Q-Q plots are reasonable and that genomic control value is close to 1.0. If not, refine sample and variant filters as needed.&lt;br /&gt;
&lt;br /&gt;
= Visualize Results =&lt;br /&gt;
&lt;br /&gt;
A number of displays will likely be useful. Probably these should include:&lt;br /&gt;
&lt;br /&gt;
* Manhattan Plots&lt;br /&gt;
* [[LocusZoom]] Plots&lt;br /&gt;
* Q-Q Plots&lt;br /&gt;
&lt;br /&gt;
= Think You Are Done? =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;No way!!!&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
== Indels and Structural Variants ==&lt;br /&gt;
&lt;br /&gt;
You still need a plan to call and evaluate short insertions and deletions as well as larger structural variants.&lt;br /&gt;
&lt;br /&gt;
== Pathway Based Analyses ==&lt;br /&gt;
&lt;br /&gt;
Carry out analyses that include groups of genes with similar biological function (for example, according to [[Gene Ontology]] or [[Kyoto Encyclopedia of Genes and Genomes]] annotations.&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Examples_of_Read_Mapping_with_Karma_and_BWA&amp;diff=165</id>
		<title>Examples of Read Mapping with Karma and BWA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Examples_of_Read_Mapping_with_Karma_and_BWA&amp;diff=165"/>
		<updated>2009-12-11T15:06:48Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt; #  Some instructions for read mapping and variant calling using the &lt;br /&gt;
 #  University of Michigan tools and procedures.  Please view this as source.&lt;br /&gt;
&lt;br /&gt;
 	#  -- Paul Anderson and Tom Blackwell, December 11, 2009  --&lt;br /&gt;
&lt;br /&gt;
 #  These instructions will cover three components:&lt;br /&gt;
&lt;br /&gt;
 #  (1)  read mapping using karma&lt;br /&gt;
 #  (2)  read mapping using bwa&lt;br /&gt;
 #  (3)  samtools processing from .sam to .glf&lt;br /&gt;
&lt;br /&gt;
 #  These instructions are specifically with reference to the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set that has been distributed.  Please &lt;br /&gt;
 #  see some general discussion at the beginning of item 2.  Everything &lt;br /&gt;
 #  that&#039;s not commented out should be runnable code, although you &lt;br /&gt;
 #  will need to alter the directory and file names.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (1)  Command line procedure for karma read mapping of the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set.  Paul Anderson writes:  &lt;br /&gt;
&lt;br /&gt;
 #  Start with four subdirectories:  &lt;br /&gt;
&lt;br /&gt;
 #  indiv  containing single and paired end Illumina .fastq sequence files, &lt;br /&gt;
 #  ab     containing AB SOLiD color space reads as in the test data set, &lt;br /&gt;
 #  k.ref  containing the gzipped human genome reference sequence,  and &lt;br /&gt;
 #  k.out  an empty directory for the resulting karma .sam and .stats files.&lt;br /&gt;
&lt;br /&gt;
 #  First, build karma&#039;s binary word index files:&lt;br /&gt;
&lt;br /&gt;
cd     k.ref&lt;br /&gt;
zcat   human_g1k_v37.fasta.gz  &amp;gt;  human_g1k_v37.fa&lt;br /&gt;
karma  --reference human_g1k_v37.fa  --createIndex  --occurrenceCutoff 5000&lt;br /&gt;
&lt;br /&gt;
 #  then map Illumina paired end or single end reads as, for example:&lt;br /&gt;
&lt;br /&gt;
cd	../k.out&lt;br /&gt;
karma	--reference ../k.ref/human_g1k_v37.fa  --pairedReads   --maxInsert 2000	  \&lt;br /&gt;
	../indiv/SRR010941_1.recal.fastq.gz  ../indiv/SRR010941_2.recal.fastq.gz&lt;br /&gt;
&lt;br /&gt;
karma	--reference ../k.ref/human_g1k_v37.fa  --maxInsert 2000        	 	  \&lt;br /&gt;
	../indiv/SRR010936.recal.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  The read lengths in the Illumina data I saw are long enough so that the &lt;br /&gt;
 #  only parameter worth tweaking might be  --occurrenceCutoff  in the index &lt;br /&gt;
 #  structures.  I will need to study the mapping performance to see if this &lt;br /&gt;
 #  makes any appreciable difference in speed.&lt;br /&gt;
&lt;br /&gt;
 #  LS 454 fastq files can also be mapped using the same index, but the &lt;br /&gt;
 #  mapping success rate will be much lower.&lt;br /&gt;
&lt;br /&gt;
 #  For AB Solid data, a color space reference and set of index files must be &lt;br /&gt;
 #  created.  For the lengths of reads I see, we will create two -- one with a &lt;br /&gt;
 #  12-mer index, and the other with a 15-mer.&lt;br /&gt;
&lt;br /&gt;
cd    ../k.ref&lt;br /&gt;
ln -s    human_g1k_v37.fa  human_g1k_v37_12CS.fa&lt;br /&gt;
ln -s    human_g1k_v37.fa  human_g1k_v37_15CS.fa&lt;br /&gt;
karma  --reference human_g1k_v37_12CS.fa  --colorSpace  --createIndex  --wordSize 12&lt;br /&gt;
karma  --reference human_g1k_v37_15CS.fa  --colorSpace  --createIndex  --wordSize 15&lt;br /&gt;
&lt;br /&gt;
 #  Then, short color space reads (25/26-mers) should be mapped using the 12-mer index:&lt;br /&gt;
&lt;br /&gt;
cd	../k.out&lt;br /&gt;
karma	--reference   ../k.ref/human_g1k_v37.fa   	\&lt;br /&gt;
	--csreference ../k.ref/human_g1k_v37_12CS.fa	\&lt;br /&gt;
	--colorSpace  --pairedReads  --maxInsert 2000	\&lt;br /&gt;
	../ab/TG150_1.color.space.fastq.gz     	 	\&lt;br /&gt;
	../ab/TG150_2.color.space.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  To map longer color space reads (&amp;gt;30-mer), use the bigger index:&lt;br /&gt;
&lt;br /&gt;
karma	--reference   ../k.ref/human_g1k_v37.fa     	\&lt;br /&gt;
	--csreference ../k.ref/human_g1k_v37_15CS.fa	\&lt;br /&gt;
	--colorSpace  --pairedReads  --maxInsert 2000	\&lt;br /&gt;
	../ab/TG152_1.color.space.fastq.gz    	 	\&lt;br /&gt;
	../ab/TG152_2.color.space.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  You can use the shorter reference to map any length reads, but it is slower.  &lt;br /&gt;
 #  The 15-mer color space index (human_g1k_v37_15CS.fa) is a good compromise &lt;br /&gt;
 #  for mapping reads 30 or more bases, and for machines with 24GB of RAM.&lt;br /&gt;
&lt;br /&gt;
 #  The maxInsert of 2000 is somewhat arbitrary - you want a number that includes &lt;br /&gt;
 #  the bulk of the paired end read insert distances you are interested in.  More &lt;br /&gt;
 #  is slower, so you don&#039;t want an arbitrarily large value.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (2)  Command line procedure for bwa read mapping of the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set.&lt;br /&gt;
&lt;br /&gt;
 #  As a general principle, one wants to keep all sequence data separated &lt;br /&gt;
 #  by sequencing technology and by individual throughout the entire mapping &lt;br /&gt;
 #  and variant calling process.  Different technologies may have different &lt;br /&gt;
 #  requirements for read mapping and different characteristics for variant &lt;br /&gt;
 #  calling.  This can be done either using the directory structure or with a &lt;br /&gt;
 #  file tracking database.  &lt;br /&gt;
&lt;br /&gt;
 #  As a further complication, the Broad Institute Illumina sequencing runs &lt;br /&gt;
 #  in the current test data set benefit from removing leading and trailing Ns &lt;br /&gt;
 #  before read mapping with bwa.  All Illumina data for the CEU trio that &lt;br /&gt;
 #  have SRR identifiers are Broad Institute runs.  (There are no Broad &lt;br /&gt;
 #  runs for the YRI trio.  There, SRR identifiers will indicate either Beijing &lt;br /&gt;
 #  or Wash U runs.)  I have used exactly the same bwa command lines with &lt;br /&gt;
 #  or without N-trimming.&lt;br /&gt;
&lt;br /&gt;
 #  In what follows, I will use somewhat simplified directory and file names.&lt;br /&gt;
 #  At this stage, it&#039;s easiest to have the command line prompt remain at top &lt;br /&gt;
 #  level in the directory structure, and refer to all files using relative &lt;br /&gt;
 #  pathnames, relative to the location of the prompt.  At the start, suppose &lt;br /&gt;
 #  that there are two subdirectories:  &#039;bwa.ref&#039;  containing the gzipped human &lt;br /&gt;
 #  genome reference sequence, and  &#039;indiv&#039;  containing many .fastq files for &lt;br /&gt;
 #  the target individual generated by the appropriate sequencing technology.  &lt;br /&gt;
 #  The bwa program has an inconvenient habit of writing to std.err.  I routinely &lt;br /&gt;
 #  redirect that to a log file so that I can keep on working while bwa runs in &lt;br /&gt;
 #  the background.  &lt;br /&gt;
&lt;br /&gt;
 #  In general, bwa read mapping requires initial indexing of the genome &lt;br /&gt;
 #  reference sequence, followed by two passes for each .fastq file.  The &lt;br /&gt;
 #  software components of this process are:  &lt;br /&gt;
&lt;br /&gt;
 #  bwa  index  --  Do this only when a new version of the genome reference &lt;br /&gt;
 #                  sequence is released.  Takes just under two hours.&lt;br /&gt;
 #  bwa   aln   --  Run this on every .fastq file individually.  Highly &lt;br /&gt;
 #                  variable timings -- takes between 10 minutes and many &lt;br /&gt;
 #                  hours per .fastq file.  Better data runs quicker.&lt;br /&gt;
 #  bwa  samse  --  Converts a single .fastq / .sai pair to .sam alignment &lt;br /&gt;
 #                  format.  Usually under 1 minute per file.&lt;br /&gt;
 #  bwa  sampe  --  Converts paired end .fastq / .sai pairs (four files total) &lt;br /&gt;
 #                  to a single .sam alignment file.  1 - 2 minutes per run.  &lt;br /&gt;
 #                  (In general, the CEU trio chromosome 20 test data set &lt;br /&gt;
 #                  contains VERY small .fastq files.  Chromosome 20 is just &lt;br /&gt;
 #                  over 2% of the entire genome.)&lt;br /&gt;
&lt;br /&gt;
 #  Sample command lines.  This is csh syntax.  &lt;br /&gt;
&lt;br /&gt;
 #  The parameter string  &amp;quot; -n 0.002  -M 7  -R 25 &amp;quot;  used in bwa aln below &lt;br /&gt;
 #  is just my initial guess.  I have made an entire run using this string.  &lt;br /&gt;
 #  I might make another run using Heng Li&#039;s default values for all three &lt;br /&gt;
 #  parameters and compare the results, but I have no answers from this yet.  &lt;br /&gt;
&lt;br /&gt;
bwa index  -a bwtsw  -p	bwa.ref/ncbi.v37.ref	 	 	 	\&lt;br /&gt;
	 	 	bwa.ref/human_g1k_v37.fasta.gz  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
mkdir bwa.sai bwa.sam&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa aln  -n 0.002  -M 7  -R 25  bwa.ref/ncbi.v37.ref	\&lt;br /&gt;
	    indiv/SRRx_1.fastq.gz  &amp;gt;  bwa.sai/SRRx_1.sai ) &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  ... and the same for  SRRx_2.fastq.gz  or  SRRx.fastq.gz.&lt;br /&gt;
 #  Then, depending on whether these are paired end reads or not, either:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa sampe  bwa.ref/ncbi.v37.ref	 	 	\&lt;br /&gt;
	    bwa.sai/SRRx_1.sai	   bwa.sai/SRRx_2.sai	 	\&lt;br /&gt;
	    indiv/SRRx_1.fastq.gz  indiv/SRRx_2.fastq.gz	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.pair.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  or, for single end reads:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa samse  bwa.ref/ncbi.v37.ref	 	 	\&lt;br /&gt;
	    bwa.sai/SRRx.sai	   indiv/SRRx.fastq.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.single.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  For LS 454 sequence data, bwa read mapping is a one-step process.  &lt;br /&gt;
 #  This uses the same genome reference sequence index as for Illumina &lt;br /&gt;
 #  data above, and does not use paired end information in the mapping.  &lt;br /&gt;
 #  (The directory  bwa.sam  shown here should be separate from that &lt;br /&gt;
 #  created above for Illumina data.  Similarly for AB SOLiD below.)  &lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa dbwtsw  bwa.ref/ncbi.v37.ref  indiv/SRRx.fastq.gz  \&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  For AB SOLiD data, one must build a separate index structure for &lt;br /&gt;
 #  the genome reference sequence, and (completely undocumented) one &lt;br /&gt;
 #  must rewrite all of the .fastq files replacing  0,1,2,3,&amp;quot;.&amp;quot;  with &lt;br /&gt;
 #  A,C,G,T,N  in every line of color space sequence and omitting the &lt;br /&gt;
 #  first two characters from each line of converted sequence and from &lt;br /&gt;
 #  the base call quality strings.  (An awk script does this conversion &lt;br /&gt;
 #  really quickly.)&lt;br /&gt;
&lt;br /&gt;
bwa index  -a bwtsw  -p	bwa.ref/ncbi.color.ref  -c	 	 	\&lt;br /&gt;
	 	 	bwa.ref/human_g1k_v37.fasta.gz  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
mkdir bwa.sai bwa.sam&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa aln -n 0.002 -M 7 -R 25 -o 0 -c  bwa.ref/ncbi.color.ref	\&lt;br /&gt;
	    indiv/TG145_1.cvt.gz  &amp;gt;  bwa.sai/TG145_1.sai ) &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  ... and the same for  TG145_2.cvt.gz  or  TG145.cvt.gz.&lt;br /&gt;
 #  Then, depending on whether these are paired end reads or not, either:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa sampe  bwa.ref/ncbi.color.ref	 	 	\&lt;br /&gt;
	    bwa.sai/TG145_1.sai	   bwa.sai/TG145_2.sai	 	\&lt;br /&gt;
	    indiv/TG145_1.cvt.gz   indiv/TG145_2.cvt.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.pair.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  or:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa samse  bwa.ref/ncbi.color.ref	 	 	\&lt;br /&gt;
	    bwa.sai/TG145.sai	   indiv/TG145.cvt.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.single.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (3)  Here is the process which builds the .glf format files used for &lt;br /&gt;
 #  SNP calling out of .sam format files for individual sequencing runs &lt;br /&gt;
 #  produced using either read mapping algorithm.  This will use only &lt;br /&gt;
 #  samtools utilities and contains nothing specific to either read mapper.  &lt;br /&gt;
&lt;br /&gt;
 #  For illustration, I will assume two of the directories left over from &lt;br /&gt;
 #  bwa read mapping in the preceding section.  These are:  &lt;br /&gt;
&lt;br /&gt;
 #  bwa.ref  containing files  human_g1k_v37.fasta.gz, human_g1k_v37.fasta.fai, and &lt;br /&gt;
 #  bwa.sam  containing all of the single-end and paired-end .sam files generated &lt;br /&gt;
 #  	 	 	for one individual and sequencing technology.&lt;br /&gt;
&lt;br /&gt;
 #  The full set of command lines is:  (Note csh syntax again -- and this time, &lt;br /&gt;
 #     I will change directories for convenience in file naming.)&lt;br /&gt;
&lt;br /&gt;
mkdir  bwa.bam  indiv.bam&lt;br /&gt;
cd     bwa.bam&lt;br /&gt;
set    minq=17&lt;br /&gt;
&lt;br /&gt;
foreach  file  ( ../bwa.sam/*.sam )&lt;br /&gt;
   set unsorted=`basename $file .sam`.nosort.bam&lt;br /&gt;
   samtools view -bhuS  -o $unsorted  -q $minq  $file	&amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
   samtools sort  $unsorted  `basename $file .sam`	&amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
   rm $unsorted&lt;br /&gt;
   end&lt;br /&gt;
&lt;br /&gt;
unset  minq&lt;br /&gt;
&lt;br /&gt;
 #  This is the point where one would insert steps to recalibrate &lt;br /&gt;
 #  base call quality values, check genotype identity or remove &lt;br /&gt;
 #  duplicate sequence reads.  Each of these steps is specific to &lt;br /&gt;
 #  an individual sequencing run.&lt;br /&gt;
&lt;br /&gt;
cd  ../indiv.bam&lt;br /&gt;
samtools  merge  person.bam  ../bwa.bam/*.bam&lt;br /&gt;
samtools  index  person.bam&lt;br /&gt;
samtools  view   person.bam  20 | samtools pileup	\&lt;br /&gt;
	  -f ../bwa.ref/human_g1k_v37.fasta.gz	 	\&lt;br /&gt;
	  -t ../bwa.ref/human_g1k_v37.fasta.fai	 	\&lt;br /&gt;
	  -g  -  &amp;gt;  person.glf&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Examples_of_Read_Mapping_with_Karma_and_BWA&amp;diff=164</id>
		<title>Examples of Read Mapping with Karma and BWA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Examples_of_Read_Mapping_with_Karma_and_BWA&amp;diff=164"/>
		<updated>2009-12-11T14:55:06Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: Created page with &amp;#039; #  Some instructions for read mapping and variant calling using the   #  University of Michigan tools and procedures.   	#  -- Paul Anderson and Tom Blackwell, December 11, 2009…&amp;#039;&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt; #  Some instructions for read mapping and variant calling using the &lt;br /&gt;
 #  University of Michigan tools and procedures.&lt;br /&gt;
&lt;br /&gt;
 	#  -- Paul Anderson and Tom Blackwell, December 11, 2009  --&lt;br /&gt;
&lt;br /&gt;
 #  These instructions will cover three components:&lt;br /&gt;
&lt;br /&gt;
 #  (1)  read mapping using karma&lt;br /&gt;
 #  (2)  read mapping using bwa&lt;br /&gt;
 #  (3)  samtools processing from .sam to .glf&lt;br /&gt;
&lt;br /&gt;
 #  These instructions are specifically with reference to the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set that has been distributed.  Please &lt;br /&gt;
 #  see some general discussion at the beginning of item 2.  Everything &lt;br /&gt;
 #  that&#039;s not commented out should be runnable code, although you &lt;br /&gt;
 #  will need to alter the directory and file names.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (1)  Command line procedure for karma read mapping of the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set.  Paul Anderson writes:  &lt;br /&gt;
&lt;br /&gt;
 #  Start with four subdirectories:  &lt;br /&gt;
&lt;br /&gt;
 #  indiv  containing single and paired end Illumina .fastq sequence files, &lt;br /&gt;
 #  ab     containing AB SOLiD color space reads as in the test data set, &lt;br /&gt;
 #  k.ref  containing the gzipped human genome reference sequence,  and &lt;br /&gt;
 #  k.out  an empty directory for the resulting karma .sam and .stats files.&lt;br /&gt;
&lt;br /&gt;
 #  First, build karma&#039;s binary word index files:&lt;br /&gt;
&lt;br /&gt;
cd     k.ref&lt;br /&gt;
zcat   human_g1k_v37.fasta.gz  &amp;gt;  human_g1k_v37.fa&lt;br /&gt;
karma  --reference human_g1k_v37.fa  --createIndex  --occurrenceCutoff 5000&lt;br /&gt;
&lt;br /&gt;
 #  then map Illumina paired end or single end reads as, for example:&lt;br /&gt;
&lt;br /&gt;
cd	../k.out&lt;br /&gt;
karma	--reference ../k.ref/human_g1k_v37.fa  --pairedReads   --maxInsert 2000	  \&lt;br /&gt;
	../indiv/SRR010941_1.recal.fastq.gz  ../indiv/SRR010941_2.recal.fastq.gz&lt;br /&gt;
&lt;br /&gt;
karma	--reference ../k.ref/human_g1k_v37.fa  --maxInsert 2000        	 	  \&lt;br /&gt;
	../indiv/SRR010936.recal.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  The read lengths in the Illumina data I saw are long enough so that the &lt;br /&gt;
 #  only parameter worth tweaking might be  --occurrenceCutoff  in the index &lt;br /&gt;
 #  structures.  I will need to study the mapping performance to see if this &lt;br /&gt;
 #  makes any appreciable difference in speed.&lt;br /&gt;
&lt;br /&gt;
 #  LS 454 fastq files can also be mapped using the same index, but the &lt;br /&gt;
 #  mapping success rate will be much lower.&lt;br /&gt;
&lt;br /&gt;
 #  For AB Solid data, a color space reference and set of index files must be &lt;br /&gt;
 #  created.  For the lengths of reads I see, we will create two -- one with a &lt;br /&gt;
 #  12-mer index, and the other with a 15-mer.&lt;br /&gt;
&lt;br /&gt;
cd    ../k.ref&lt;br /&gt;
ln -s    human_g1k_v37.fa  human_g1k_v37_12CS.fa&lt;br /&gt;
ln -s    human_g1k_v37.fa  human_g1k_v37_15CS.fa&lt;br /&gt;
karma  --reference human_g1k_v37_12CS.fa  --colorSpace  --createIndex  --wordSize 12&lt;br /&gt;
karma  --reference human_g1k_v37_15CS.fa  --colorSpace  --createIndex  --wordSize 15&lt;br /&gt;
&lt;br /&gt;
 #  Then, short color space reads (25/26-mers) should be mapped using the 12-mer index:&lt;br /&gt;
&lt;br /&gt;
cd	../k.out&lt;br /&gt;
karma	--reference   ../k.ref/human_g1k_v37.fa   	\&lt;br /&gt;
	--csreference ../k.ref/human_g1k_v37_12CS.fa	\&lt;br /&gt;
	--colorSpace  --pairedReads  --maxInsert 2000	\&lt;br /&gt;
	../ab/TG150_1.color.space.fastq.gz     	 	\&lt;br /&gt;
	../ab/TG150_2.color.space.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  To map longer color space reads (&amp;gt;30-mer), use the bigger index:&lt;br /&gt;
&lt;br /&gt;
karma	--reference   ../k.ref/human_g1k_v37.fa     	\&lt;br /&gt;
	--csreference ../k.ref/human_g1k_v37_15CS.fa	\&lt;br /&gt;
	--colorSpace  --pairedReads  --maxInsert 2000	\&lt;br /&gt;
	../ab/TG152_1.color.space.fastq.gz    	 	\&lt;br /&gt;
	../ab/TG152_2.color.space.fastq.gz&lt;br /&gt;
&lt;br /&gt;
 #  You can use the shorter reference to map any length reads, but it is slower.  &lt;br /&gt;
 #  The 15-mer color space index (human_g1k_v37_15CS.fa) is a good compromise &lt;br /&gt;
 #  for mapping reads 30 or more bases, and for machines with 24GB of RAM.&lt;br /&gt;
&lt;br /&gt;
 #  The maxInsert of 2000 is somewhat arbitrary - you want a number that includes &lt;br /&gt;
 #  the bulk of the paired end read insert distances you are interested in.  More &lt;br /&gt;
 #  is slower, so you don&#039;t want an arbitrarily large value.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (2)  Command line procedure for bwa read mapping of the CEU trio &lt;br /&gt;
 #  chromosome 20 test data set.&lt;br /&gt;
&lt;br /&gt;
 #  As a general principle, one wants to keep all sequence data separated &lt;br /&gt;
 #  by sequencing technology and by individual throughout the entire mapping &lt;br /&gt;
 #  and variant calling process.  Different technologies may have different &lt;br /&gt;
 #  requirements for read mapping and different characteristics for variant &lt;br /&gt;
 #  calling.  This can be done either using the directory structure or with a &lt;br /&gt;
 #  file tracking database.  &lt;br /&gt;
&lt;br /&gt;
 #  As a further complication, the Broad Institute Illumina sequencing runs &lt;br /&gt;
 #  in the current test data set benefit from removing leading and trailing Ns &lt;br /&gt;
 #  before read mapping with bwa.  All Illumina data for the CEU trio that &lt;br /&gt;
 #  have SRR identifiers are Broad Institute runs.  (There are no Broad &lt;br /&gt;
 #  runs for the YRI trio.  There, SRR identifiers will indicate either Beijing &lt;br /&gt;
 #  or Wash U runs.)  I have used exactly the same bwa command lines with &lt;br /&gt;
 #  or without N-trimming.&lt;br /&gt;
&lt;br /&gt;
 #  In what follows, I will use somewhat simplified directory and file names.&lt;br /&gt;
 #  At this stage, it&#039;s easiest to have the command line prompt remain at top &lt;br /&gt;
 #  level in the directory structure, and refer to all files using relative &lt;br /&gt;
 #  pathnames, relative to the location of the prompt.  At the start, suppose &lt;br /&gt;
 #  that there are two subdirectories:  &#039;bwa.ref&#039;  containing the gzipped human &lt;br /&gt;
 #  genome reference sequence, and  &#039;indiv&#039;  containing many .fastq files for &lt;br /&gt;
 #  the target individual generated by the appropriate sequencing technology.  &lt;br /&gt;
 #  The bwa program has an inconvenient habit of writing to std.err.  I routinely &lt;br /&gt;
 #  redirect that to a log file so that I can keep on working while bwa runs in &lt;br /&gt;
 #  the background.  &lt;br /&gt;
&lt;br /&gt;
 #  In general, bwa read mapping requires initial indexing of the genome &lt;br /&gt;
 #  reference sequence, followed by two passes for each .fastq file.  The &lt;br /&gt;
 #  software components of this process are:  &lt;br /&gt;
&lt;br /&gt;
 #  bwa  index	#  Do this only when a new version of the genome reference &lt;br /&gt;
	 	#  sequence is released.  Takes just under two hours.&lt;br /&gt;
 #  bwa   aln	#  Run this on every .fastq file individually.  Highly &lt;br /&gt;
	 	#  variable timings -- takes between 10 minutes and many &lt;br /&gt;
	 	#  hours per .fastq file.  Better data runs quicker.&lt;br /&gt;
 #  bwa  samse	#  Converts a single .fastq / .sai pair to .sam alignment &lt;br /&gt;
	 	#  format.  Usually under 1 minute per file.&lt;br /&gt;
 #  bwa  sampe	#  Converts paired end .fastq / .sai pairs (four files total) &lt;br /&gt;
	 	#  to a single .sam alignment file.  1 - 2 minutes per run.  &lt;br /&gt;
	 	#  (In general, the CEU trio chromosome 20 test data set &lt;br /&gt;
	 	#  contains VERY small .fastq files.  Chromosome 20 is just &lt;br /&gt;
	 	#  over 2% of the entire genome.)&lt;br /&gt;
&lt;br /&gt;
 #  Sample command lines.  This is csh syntax.  &lt;br /&gt;
&lt;br /&gt;
 #  The parameter string  &amp;quot; -n 0.002  -M 7  -R 25 &amp;quot;  used in bwa aln below &lt;br /&gt;
 #  is just my initial guess.  I have made an entire run using this string.  &lt;br /&gt;
 #  I might make another run using Heng Li&#039;s default values for all three &lt;br /&gt;
 #  parameters and compare the results, but I have no answers from this yet.  &lt;br /&gt;
&lt;br /&gt;
bwa index  -a bwtsw  -p	bwa.ref/ncbi.v37.ref	 	 	 	\&lt;br /&gt;
	 	 	bwa.ref/human_g1k_v37.fasta.gz  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
mkdir bwa.sai bwa.sam&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa aln  -n 0.002  -M 7  -R 25  bwa.ref/ncbi.v37.ref	\&lt;br /&gt;
	    indiv/SRRx_1.fastq.gz  &amp;gt;  bwa.sai/SRRx_1.sai ) &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  ... and the same for  SRRx_2.fastq.gz  or  SRRx.fastq.gz.&lt;br /&gt;
 #  Then, depending on whether these are paired end reads or not, either:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa sampe  bwa.ref/ncbi.v37.ref	 	 	\&lt;br /&gt;
	    bwa.sai/SRRx_1.sai	   bwa.sai/SRRx_2.sai	 	\&lt;br /&gt;
	    indiv/SRRx_1.fastq.gz  indiv/SRRx_2.fastq.gz	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.pair.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  or, for single end reads:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa samse  bwa.ref/ncbi.v37.ref	 	 	\&lt;br /&gt;
	    bwa.sai/SRRx.sai	   indiv/SRRx.fastq.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.single.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  For LS 454 sequence data, bwa read mapping is a one-step process.  &lt;br /&gt;
 #  This uses the same genome reference sequence index as for Illumina &lt;br /&gt;
 #  data above, and does not use paired end information in the mapping.  &lt;br /&gt;
 #  (The directory  bwa.sam  shown here should be separate from that &lt;br /&gt;
 #  created above for Illumina data.  Similarly for AB SOLiD below.)  &lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa dbwtsw  bwa.ref/ncbi.v37.ref  indiv/SRRx.fastq.gz  \&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  For AB SOLiD data, one must build a separate index structure for &lt;br /&gt;
 #  the genome reference sequence, and (completely undocumented) one &lt;br /&gt;
 #  must rewrite all of the .fastq files replacing  0,1,2,3,&amp;quot;.&amp;quot;  with &lt;br /&gt;
 #  A,C,G,T,N  in every line of color space sequence and omitting the &lt;br /&gt;
 #  first two characters from each line of converted sequence and from &lt;br /&gt;
 #  the base call quality strings.  (An awk script does this conversion &lt;br /&gt;
 #  really quickly.)&lt;br /&gt;
&lt;br /&gt;
bwa index  -a bwtsw  -p	bwa.ref/ncbi.color.ref  -c	 	 	\&lt;br /&gt;
	 	 	bwa.ref/human_g1k_v37.fasta.gz  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
mkdir bwa.sai bwa.sam&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa aln -n 0.002 -M 7 -R 25 -o 0 -c  bwa.ref/ncbi.color.ref	\&lt;br /&gt;
	    indiv/TG145_1.cvt.gz  &amp;gt;  bwa.sai/TG145_1.sai ) &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  ... and the same for  TG145_2.cvt.gz  or  TG145.cvt.gz.&lt;br /&gt;
 #  Then, depending on whether these are paired end reads or not, either:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa sampe  bwa.ref/ncbi.color.ref	 	 	\&lt;br /&gt;
	    bwa.sai/TG145_1.sai	   bwa.sai/TG145_2.sai	 	\&lt;br /&gt;
	    indiv/TG145_1.cvt.gz   indiv/TG145_2.cvt.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.pair.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
 #  or:&lt;br /&gt;
&lt;br /&gt;
( nice +20  bwa samse  bwa.ref/ncbi.color.ref	 	 	\&lt;br /&gt;
	    bwa.sai/TG145.sai	   indiv/TG145.cvt.gz	 	\&lt;br /&gt;
	      &amp;gt;  bwa.sam/SRRx.single.sam )  &amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
 #  (3)  Here is the process which builds the .glf format files used for &lt;br /&gt;
 #  SNP calling out of .sam format files for individual sequencing runs &lt;br /&gt;
 #  produced using either read mapping algorithm.  This will use only &lt;br /&gt;
 #  samtools utilities and contains nothing specific to either read mapper.  &lt;br /&gt;
&lt;br /&gt;
 #  For illustration, I will assume two of the directories left over from &lt;br /&gt;
 #  bwa read mapping in the preceding section.  These are:  &lt;br /&gt;
&lt;br /&gt;
 #  bwa.ref  containing files  human_g1k_v37.fasta.gz, human_g1k_v37.fasta.fai, and &lt;br /&gt;
 #  bwa.sam  containing all of the single-end and paired-end .sam files generated &lt;br /&gt;
 #  	 	 	for one individual and sequencing technology.&lt;br /&gt;
&lt;br /&gt;
 #  The full set of command lines is:  (Note csh syntax again -- and this time, &lt;br /&gt;
 #     I will change directories for convenience in file naming.)&lt;br /&gt;
&lt;br /&gt;
mkdir  bwa.bam  indiv.bam&lt;br /&gt;
cd     bwa.bam&lt;br /&gt;
set    minq=17&lt;br /&gt;
&lt;br /&gt;
foreach  file  ( ../bwa.sam/*.sam )&lt;br /&gt;
   set unsorted=`basename $file .sam`.nosort.bam&lt;br /&gt;
   samtools view -bhuS  -o $unsorted  -q $minq  $file	&amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
   samtools sort  $unsorted  `basename $file .sam`	&amp;gt;&amp;gt;&amp;amp; logfile&lt;br /&gt;
   rm $unsorted&lt;br /&gt;
   end&lt;br /&gt;
&lt;br /&gt;
unset  minq&lt;br /&gt;
&lt;br /&gt;
 #  This is the point where one would insert steps to recalibrate &lt;br /&gt;
 #  base call quality values, check genotype identity or remove &lt;br /&gt;
 #  duplicate sequence reads.  Each of these steps is specific to &lt;br /&gt;
 #  an individual sequencing run.&lt;br /&gt;
&lt;br /&gt;
cd  ../indiv.bam&lt;br /&gt;
samtools  merge  person.bam  ../bwa.bam/*.bam&lt;br /&gt;
samtools  index  person.bam&lt;br /&gt;
samtools  view   person.bam  20 | samtools pileup	\&lt;br /&gt;
	  -f ../bwa.ref/human_g1k_v37.fasta.gz	 	\&lt;br /&gt;
	  -t ../bwa.ref/human_g1k_v37.fasta.fai	 	\&lt;br /&gt;
	  -g  -  &amp;gt;  person.glf&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Main_Page&amp;diff=163</id>
		<title>Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Main_Page&amp;diff=163"/>
		<updated>2009-12-11T14:54:37Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Sequence Analysis Tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;!--        BANNER ACROSS TOP OF PAGE        --&amp;gt;&lt;br /&gt;
{| style=&amp;quot;width:100%; background:#fcfcfc; margin-top:1.2em; border:1px solid #ccc;&amp;quot; |&lt;br /&gt;
 | style=&amp;quot;width:100%; text-align:center; white-space:nowrap; color:#000;&amp;quot; | &lt;br /&gt;
&amp;lt;div style=&amp;quot;font-size:162%; border:none; margin:0; padding:.1em; color:#000;&amp;quot;&amp;gt;Abecasis Group Wiki&amp;lt;/div&amp;gt;&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
[[Image:2009.08_Group_Retreat_Photo.jpg|400px|center|Group Photo]]&lt;br /&gt;
&lt;br /&gt;
== Welcome! ==&lt;br /&gt;
&lt;br /&gt;
Welcome to our brand new wiki!&lt;br /&gt;
&lt;br /&gt;
If you would like to contribute, [[Special:UserLogin|log-in]] or [http://csgwiki.sph.umich.edu/index.php?title=Special:UserLogin&amp;amp;type=signup create an account]. We recommend using your e-mail address or Michigan uniqname as your user id.&lt;br /&gt;
&lt;br /&gt;
For basic instructions, see [http://en.wikipedia.org/wiki/Wikipedia:Tutorial the Wikipedia Tutorial].&lt;br /&gt;
&lt;br /&gt;
== Sequence Analysis Tools ==&lt;br /&gt;
&lt;br /&gt;
We are developing several tools for the analysis of next generation sequence data. &lt;br /&gt;
&lt;br /&gt;
=== Read Mapping ===&lt;br /&gt;
&lt;br /&gt;
[[Karma|Karma]] - Our fast short read aligner&lt;br /&gt;
&lt;br /&gt;
[[Karma-colorspace|Karma-ColorSpace]] - QUICKSTART on mapping color space reads&lt;br /&gt;
&lt;br /&gt;
[[Examples|Examples]] - Sample command lines with discussion&lt;br /&gt;
&lt;br /&gt;
=== Variant Calling ===&lt;br /&gt;
&lt;br /&gt;
[[glfSingle]] - Variant calling for a single, deeply sequenced individual&lt;br /&gt;
&lt;br /&gt;
[[glfTrio]] - Variant calling for a single, deeply sequenced nuclear family with two parents and one child&lt;br /&gt;
&lt;br /&gt;
[[glfMultiples]] -- Variant calling for multiple, unrelated individuals&lt;br /&gt;
&lt;br /&gt;
=== Variant Annotation ===&lt;br /&gt;
&lt;br /&gt;
[[vcfCodingSnps]] -- Annotate coding variants in a VCF file.&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=160</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=160"/>
		<updated>2009-12-02T20:19:38Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
* A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
&lt;br /&gt;
* A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
&lt;br /&gt;
* Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
&lt;br /&gt;
* Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
&lt;br /&gt;
* Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]]) &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).  Suppose that NCBI36.fa is a FASTA file which contains the nucleotide sequences for all chromosomes.&lt;br /&gt;
&lt;br /&gt;
The command to invoke is:&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a packed binary sequence file and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; With a .csfastq &amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam format named &amp;amp;nbsp; &amp;quot;single.sam&amp;quot;&amp;amp;nbsp; derived from the .fastq&amp;amp;nbsp; file name.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable and will produce multiple .sam output files, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq  single.2.csfastq  single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads  pair.1.csfastq  pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.1.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files, etc. and write output files&amp;amp;nbsp; &amp;quot;pair.1.sam&amp;quot;, &amp;quot;pair.3.sam&amp;quot;, etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads  pair.1.csfastq  pair.2.csfastq  pair.3.csfastq  pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= Additional Information =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 colors or longer (including the primer base).&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
The length of the index words influences mapping performance.&amp;amp;nbsp; Using short  index words increases the number of calculation cycles for a single read and duplications of a single word.&amp;amp;nbsp; On the other side, long index words require much larger memory.&amp;amp;nbsp; Please also keep in mind that appropriate size is related to your hardware architecture.&amp;amp;nbsp; For practical purposes, with at least 20 Gb of RAM, we find that a size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are&amp;amp;nbsp; &#039;&#039;single.sam&#039;&#039;&amp;amp;nbsp; and&amp;amp;nbsp; &#039;&#039;pair.1.sam&#039;&#039;&amp;amp;nbsp; and they conform to the .sam format specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=116</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=116"/>
		<updated>2009-11-20T02:09:00Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Map Color Space Reads */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq  single.2.csfastq  single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads  pair.1.csfastq  pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.1.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads  pair.1.csfastq  pair.2.csfastq  pair.3.csfastq  pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
The length of the index words influences mapping performance.&amp;amp;nbsp; Using short  index words increases the number of calculation cycles for a single read and duplications of a single word.&amp;amp;nbsp; On the other side, long index words require much larger memory.&amp;amp;nbsp; Please also keep in mind that appropriate size is related to your hardware architecture.&amp;amp;nbsp; For practical purposes, with at least 20 Gb of RAM, we find that a size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are&amp;amp;nbsp; &#039;&#039;single.sam&#039;&#039;&amp;amp;nbsp; and&amp;amp;nbsp; &#039;&#039;pair.1.sam&#039;&#039;&amp;amp;nbsp; and they conform to the .sam format specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=115</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=115"/>
		<updated>2009-11-20T02:07:59Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Map Color Space Reads */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.1.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
The length of the index words influences mapping performance.&amp;amp;nbsp; Using short  index words increases the number of calculation cycles for a single read and duplications of a single word.&amp;amp;nbsp; On the other side, long index words require much larger memory.&amp;amp;nbsp; Please also keep in mind that appropriate size is related to your hardware architecture.&amp;amp;nbsp; For practical purposes, with at least 20 Gb of RAM, we find that a size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are&amp;amp;nbsp; &#039;&#039;single.sam&#039;&#039;&amp;amp;nbsp; and&amp;amp;nbsp; &#039;&#039;pair.1.sam&#039;&#039;&amp;amp;nbsp; and they conform to the .sam format specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=114</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=114"/>
		<updated>2009-11-20T02:06:48Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* A Complete Example */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
The length of the index words influences mapping performance.&amp;amp;nbsp; Using short  index words increases the number of calculation cycles for a single read and duplications of a single word.&amp;amp;nbsp; On the other side, long index words require much larger memory.&amp;amp;nbsp; Please also keep in mind that appropriate size is related to your hardware architecture.&amp;amp;nbsp; For practical purposes, with at least 20 Gb of RAM, we find that a size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are&amp;amp;nbsp; &#039;&#039;single.sam&#039;&#039;&amp;amp;nbsp; and&amp;amp;nbsp; &#039;&#039;pair.1.sam&#039;&#039;&amp;amp;nbsp; and they conform to the .sam format specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=113</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=113"/>
		<updated>2009-11-20T02:03:46Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Choose an appropriate size for word index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
The length of the index words influences mapping performance.&amp;amp;nbsp; Using short  index words increases the number of calculation cycles for a single read and duplications of a single word.&amp;amp;nbsp; On the other side, long index words require much larger memory.&amp;amp;nbsp; Please also keep in mind that appropriate size is related to your hardware architecture.&amp;amp;nbsp; For practical purposes, with at least 20 Gb of RAM, we find that a size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=112</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=112"/>
		<updated>2009-11-20T01:58:42Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Auxiliary tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script&amp;amp;nbsp; &#039;&#039;solid2csfastq.py&#039;&#039;&amp;amp;nbsp; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=111</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=111"/>
		<updated>2009-11-20T01:57:57Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Auxiliary tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script &#039;&#039;solid2csfastq.py&#039;&#039; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=110</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=110"/>
		<updated>2009-11-20T01:57:07Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Auxiliary tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA and quality files, usually named&amp;amp;nbsp; XXX.csfasta&amp;amp;nbsp; and&amp;amp;nbsp; XXX\_QV.qual.&amp;amp;nbsp; We provide a script &#039;&#039;solid2csfastq.py&#039;&#039; which converts these into a single color space FASTQ file named&amp;amp;nbsp; XXX.csfastq.&amp;amp;nbsp; We believe that a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=109</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=109"/>
		<updated>2009-11-20T01:54:08Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Auxiliary tools */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
The ABI SOLiD platform generates separate FASTA files (e.g. XXX.csfasta) and quality files (e.g. XXX\_QV.qual).&amp;amp;nbsp; We provide a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert these into a single color space FASTQ file (e.g. XXX.csfastq).&amp;amp;nbsp; We believe a single color space FASTQ file simplifies post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=108</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=108"/>
		<updated>2009-11-20T01:51:06Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Minimum read length requirement */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
Keep in mind that KARMA requires color space reads that are at least twice as long as the index word size plus two (including the leading primer base).&amp;amp;nbsp; (For nucleotide space, the minimum read length is twice the word size.)&amp;amp;nbsp; For example, KARMA uses an index word size of 15 by default, so it will only map color space reads that are 32 base pairs or longer.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=107</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=107"/>
		<updated>2009-11-20T01:43:56Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Map Color Space Reads */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfastq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named&amp;amp;nbsp; &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;amp;nbsp;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored in two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a .sam&amp;amp;nbsp; file named&amp;amp;nbsp; &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end read files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=106</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=106"/>
		<updated>2009-11-20T01:38:12Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Map Color Space Reads */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
KARMA expects valid color space FASTQ files as input.&amp;amp;nbsp; We often use the suffix .csfastq to distinguish these from nucleotide space reads.&amp;amp;nbsp; For a .csfatq&amp;amp;nbsp; file of single end color space reads named &amp;amp;nbsp; single.csfastq, &amp;amp;nbsp; invoke the command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
This command line specifies both the nucleotide and color space reference sequences (and the word indexes, invisibly).&amp;amp;nbsp; The output will be written to a file in .sam&amp;amp;nbsp; format named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single.1.csfastq single.2.csfastq single.3.csfastq&lt;br /&gt;
&lt;br /&gt;
For paired end color space reads, use the option &amp;quot;--pairedReads&amp;quot;.&amp;amp;nbsp; Suppose the paired end reads are stored as two files,&amp;amp;nbsp; pair.1.csfastq&amp;amp;nbsp; and&amp;amp;nbsp; pair.2.csfastq.&amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq&lt;br /&gt;
&lt;br /&gt;
The mapping results will be stored in a SAM file named &amp;quot;pair.sam&amp;quot;, which contains reads from both files.&amp;amp;nbsp; If multiple paired end reads files are specified on the command line, KARMA will pair the 1st and 2nd files, 3rd and 4th files and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair.1.csfastq pair.2.csfastq pair.3.csfastq pair.4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=105</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=105"/>
		<updated>2009-11-20T01:19:11Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Build Binary Reference Genome and Word Index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files, one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not exceed half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=104</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=104"/>
		<updated>2009-11-20T01:17:24Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Build Binary Reference Genome and Word Index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, one also needs to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not be longer than half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=103</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=103"/>
		<updated>2009-11-20T01:16:43Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Build Binary Reference Genome and Word Index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039;&amp;amp;nbsp; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, we also need to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is used.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not be longer than half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=102</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=102"/>
		<updated>2009-11-20T01:14:42Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Build Binary Reference Genome and Word Index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).&amp;amp;nbsp;  Suppose that &amp;amp;nbsp; NCBI36.fa &amp;amp;nbsp; is a FASTA file which contains nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, we also need to build color space versions of both the genome reference sequence (option: --createReference) and the word index files (option: --createIndex).&amp;amp;nbsp;  The same nucleotide FASTA file is needed.&amp;amp;nbsp;  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.&amp;amp;nbsp;  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
When building the index files one can set the word length for indexing.&amp;amp;nbsp;  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.&amp;amp;nbsp;  Shorter index words will decrease the memory footprint at the cost of increased run time.&amp;amp;nbsp;  However, the word length must not be longer than half the length of the color space reads you intend to map, minus 1.&amp;amp;nbsp;  (See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.)&amp;amp;nbsp;  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=101</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=101"/>
		<updated>2009-11-20T01:05:07Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Build Binary Reference Genome and Word Index */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
First, build a binary version of the genome reference sequence as nucleotides (option: --createReference).  Suppose that NCBI36.fa is a FASTA file which contains the nucleotide sequences for all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
(To let KARMA map nucleotide space reads, one would use instead &#039;&#039;--createIndex&#039;&#039; to create both a binary sequence and the word index files.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Second, we also need to build a binary version of the genome reference sequence (option: --createReference) and the word index files (option: --createIndex) in color space.  The same nucleotide FASTA file is needed.  However, to avoid naming conflicts among the resulting binary files, we suggest appending &amp;quot;CS&amp;quot; to the base file name for clarity.  The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
An important parameter is the word length for indexing.  We recommend N = 15 (the default value) for the human genome on a machine with at least 20 Gb of RAM.  Shorter words will decrease the memory footprint at the cost of increased run time.  However, the word length must not be longer than half the length of the color space reads you intend to map, minus 1.  See [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion.  Specify ``--wordSize N`` in order to use &#039;&#039;N&#039;&#039; as the word size.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=100</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=100"/>
		<updated>2009-11-20T00:33:11Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Overview */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at a speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize the input data requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as nucleotides (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]}) &lt;br /&gt;
*&amp;amp;nbsp; A binary conversion of the genome reference sequence as colors plus word indices in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for a description) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads longer than a minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example demonstrating the whole procedure from building the word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; First, we need to build binary reference genome (option: --createReference)&amp;lt;br&amp;gt; &amp;amp;nbsp; (To let KARMA map nucleotide space reads, you need to use &#039;&#039;--createIndex&#039;&#039; to create the word index file.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; in nucleotide space. Assume NCBI36.fa is a FASTA file contains sequences of all chromosomes.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Second, we need to build binary reference genome (option: --createReference) and word index (option: --createIndex)&amp;lt;br&amp;gt;&amp;amp;nbsp; in color space. The same FASTA file is needed. However, to avoid naming conflicts, we suggest using word &amp;quot;CS&amp;quot; &amp;lt;br&amp;gt;&amp;amp;nbsp; appending to the base file name for clarity. The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; An important parameter is the size of words for indexing.&amp;lt;br&amp;gt; &amp;amp;nbsp; We recommand 15 (default value) for human reference genome.&amp;lt;br&amp;gt; &amp;amp;nbsp; Specifiy ``--wordSize N`` if you like to use &#039;&#039;N&#039;&#039; as word size.&amp;lt;br&amp;gt; &amp;amp;nbsp; Typically you will observe performance change (see [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion).&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Note, multiple chromosomes are supported.&amp;lt;br&amp;gt; &amp;amp;nbsp; In current version, KARMA can take one FASTA file which contains sequences of all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=99</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=99"/>
		<updated>2009-11-20T00:14:42Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Overview */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at the speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize software requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; Binary reference genome in nucleotide space (see [[#Input_file_requirement|Input file requirement]]}) &lt;br /&gt;
*&amp;amp;nbsp; Binary reference genome and word index in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in valid color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for file specification) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads are longer than minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine without using more memory than running a single instance.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example reviewing the whole procedure from building word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; First, we need to build binary reference genome (option: --createReference)&amp;lt;br&amp;gt; &amp;amp;nbsp; (To let KARMA map nucleotide space reads, you need to use &#039;&#039;--createIndex&#039;&#039; to create the word index file.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; in nucleotide space. Assume NCBI36.fa is a FASTA file contains sequences of all chromosomes.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Second, we need to build binary reference genome (option: --createReference) and word index (option: --createIndex)&amp;lt;br&amp;gt;&amp;amp;nbsp; in color space. The same FASTA file is needed. However, to avoid naming conflicts, we suggest using word &amp;quot;CS&amp;quot; &amp;lt;br&amp;gt;&amp;amp;nbsp; appending to the base file name for clarity. The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; An important parameter is the size of words for indexing.&amp;lt;br&amp;gt; &amp;amp;nbsp; We recommand 15 (default value) for human reference genome.&amp;lt;br&amp;gt; &amp;amp;nbsp; Specifiy ``--wordSize N`` if you like to use &#039;&#039;N&#039;&#039; as word size.&amp;lt;br&amp;gt; &amp;amp;nbsp; Typically you will observe performance change (see [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion).&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Note, multiple chromosomes are supported.&amp;lt;br&amp;gt; &amp;amp;nbsp; In current version, KARMA can take one FASTA file which contains sequences of all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=98</id>
		<title>Karma-colorspace</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma-colorspace&amp;diff=98"/>
		<updated>2009-11-20T00:13:14Z</updated>

		<summary type="html">&lt;p&gt;Tblackw: /* Overview */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;= Overview =&lt;br /&gt;
&lt;br /&gt;
KARMA (K-tuple Alignment with Rapid Matching Algorithm) is able to map 35 bp single end color space reads at the speed of approximately &amp;lt;math&amp;gt;1.2-2.0 \times 10^9&amp;lt;/math&amp;gt; reads per hour using Intel Xeon X760 2.66GHz and 128G memory.&lt;br /&gt;
&lt;br /&gt;
We summarize software requirements as following:&lt;br /&gt;
&lt;br /&gt;
*&amp;amp;nbsp; Binary reference genome in nucleotide space (see [[#Input_file_requirement|Input file requirement]]}) &lt;br /&gt;
*&amp;amp;nbsp; Binary reference genome and word index in color space (see [[#Build_Binary_Reference_Genome_and_Word_Index|Build Binary Reference Genome and Word Index]]) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads in valid color space FASTQ format (see [[#Input_file_requirement|Input file requirement]] for file specification) &lt;br /&gt;
*&amp;amp;nbsp; Color space reads are longer than minimum length requirement. (see [[#Minimum_read_length_requirement|Minimum read length requirement]]) &lt;br /&gt;
*&amp;amp;nbsp; Specify color space parameter when starting KARMA (see [[#Map_Color_Space_Reads|Map Color Space Reads]])&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
Please note the hardware requirements for KARMA are:&lt;br /&gt;
&lt;br /&gt;
*20G memory.  By using shared memory for the word index tables, multiple instances of KARMA can run on one machine yet consume about the same amount of memory as running one process.&amp;lt;br&amp;gt; &lt;br /&gt;
*30G disk space &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; We show a complete example reviewing the whole procedure from building word index to mapping color space reads in [[#A_Complete_Example|A Complete Example]].&lt;br /&gt;
&lt;br /&gt;
= Build Binary Reference Genome and Word Index&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; First, we need to build binary reference genome (option: --createReference)&amp;lt;br&amp;gt; &amp;amp;nbsp; (To let KARMA map nucleotide space reads, you need to use &#039;&#039;--createIndex&#039;&#039; to create the word index file.)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; in nucleotide space. Assume NCBI36.fa is a FASTA file contains sequences of all chromosomes.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Second, we need to build binary reference genome (option: --createReference) and word index (option: --createIndex)&amp;lt;br&amp;gt;&amp;amp;nbsp; in color space. The same FASTA file is needed. However, to avoid naming conflicts, we suggest using word &amp;quot;CS&amp;quot; &amp;lt;br&amp;gt;&amp;amp;nbsp; appending to the base file name for clarity. The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; An important parameter is the size of words for indexing.&amp;lt;br&amp;gt; &amp;amp;nbsp; We recommand 15 (default value) for human reference genome.&amp;lt;br&amp;gt; &amp;amp;nbsp; Specifiy ``--wordSize N`` if you like to use &#039;&#039;N&#039;&#039; as word size.&amp;lt;br&amp;gt; &amp;amp;nbsp; Typically you will observe performance change (see [[#Choose_an_appropriate_size_for_word_index|Choose an appropriate size for word index]] for more discussion).&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Note, multiple chromosomes are supported.&amp;lt;br&amp;gt; &amp;amp;nbsp; In current version, KARMA can take one FASTA file which contains sequences of all chromosomes.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Map Color Space Reads =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; KARMA takes valid color space FASTQ files inputs.&amp;lt;br&amp;gt; &amp;amp;nbsp; We usually use suffix .csfastq to distinguish it from nucleotide space reads.&amp;lt;br&amp;gt; &amp;amp;nbsp; For single end color space read, we can invoke command:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;single.sam&amp;quot;.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Multiple input files are also acceptable, e.g.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   single1.csfastq single2.csfastq single3.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; For paired end color space reads, option &amp;quot;--pairedReads&amp;quot; is requires.&amp;lt;br&amp;gt; &amp;amp;nbsp; Suppose the paired end reads are stored in file, pair1.csfastq and pair2.csfastq.&amp;lt;br&amp;gt; &amp;amp;nbsp; The command to invoke is:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Mapping results are store in a SAM file named &amp;quot;pair1.sam&amp;quot;, which contains reads from both files.&amp;lt;br&amp;gt; &amp;amp;nbsp;&amp;lt;br&amp;gt; &amp;amp;nbsp; Similarly multiple paired end reads files can be specified in command line, and KARMA will pair 1st and 2rd file, 3rd and 4th file and etc.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq pair3.csfastq pair4.csfastq&lt;br /&gt;
&lt;br /&gt;
= &amp;lt;br&amp;gt; Additional Information&amp;lt;br&amp;gt; =&lt;br /&gt;
&lt;br /&gt;
== Input file requirement ==&lt;br /&gt;
&lt;br /&gt;
KARMA requires input files in color space FASTQ format. The length of each read (which includes the leading primer base) should equal the length of its quality string. An example of a valid color space FASTQ file follows:&lt;br /&gt;
&lt;br /&gt;
  @Chromosome_20_048435095_Genome_2757096147&lt;br /&gt;
  A02232200222021320012102212311002212&lt;br /&gt;
  +&lt;br /&gt;
  !!1111111111111111111111111111111111&lt;br /&gt;
&lt;br /&gt;
== Minimum read length requirement ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; Keep in mind that the requirement of minimum color space read length for KARMA is twice the size of word plus two (including leading primer).&amp;lt;br&amp;gt; &amp;amp;nbsp; (For nucleotide space, the minimum length requirement is twice the word size.)&amp;lt;br&amp;gt; &amp;amp;nbsp; For example, KARMA use word size of 15 by default, so it will try to map color space reads that are longer than 32 base pairs.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Auxiliary tools ==&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; ABI SOLiD platform generated FASTA file (e.g. XXX.csfasta) and quality file (e.g. XXX\_QV.qual) separately. We wrote a script, &#039;&#039;solid2csfastq.py&#039;&#039;, to convert it to color space FASTQ file(e.g. XXX.csfastq). We believe a single color space FASTQ file will simplify post processing.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Choose an appropriate size for word index ==&lt;br /&gt;
&lt;br /&gt;
Size for word index is sensitive to mapping performance. A small size of word index will increase the number of calculation cycles for a single read and duplications of a single word. On the other side, a big size will require much larger memory. Please also keep in mind that appropriate size is related to your hardware architecture. For practically purpose, we found size of 15 is optimal.&lt;br /&gt;
&lt;br /&gt;
= A Complete Example =&lt;br /&gt;
&lt;br /&gt;
&amp;amp;nbsp; A wrap-up message for quick start mapping color space reads.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Building binary genome reference and word index:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --createReference --reference NCBI36.fa&lt;br /&gt;
   ln -s NCBI36.fa NCBI36CS.fa&lt;br /&gt;
   karma --colorSpace --createReference --createIndex --reference NCBI36CS.fa&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Mapping color space reads:&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace single.csfastq&lt;br /&gt;
   karma --reference NCBI36.fa --csReference NCBI36CS.fa --colorSpace \&lt;br /&gt;
   --pairedReads pair1.csfastq pair2.csfastq&lt;br /&gt;
&lt;br /&gt;
The output files are &#039;&#039;single.sam&#039;&#039; and &#039;&#039;pair1.sam&#039;&#039; and they conform SAM specification.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;/div&gt;</summary>
		<author><name>Tblackw</name></author>
	</entry>
</feed>