<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Upugema</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Upugema"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Upugema"/>
	<updated>2026-09-24T18:20:02Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Gzip&amp;diff=2532</id>
		<title>Gzip</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Gzip&amp;diff=2532"/>
		<updated>2010-11-18T00:38:46Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;Gzip&#039;&#039; is a widely used file compression utility, developed by Mark Adler and Jean Loup Gailly. The source code and algorithm for the gzip are well documented and widely available, so that it is easy to create gzip-compatible programs. A file compressed by gzip or a compatible utility tipically has an extension that ends in [[Gzip File|&#039;&#039;&#039;.gz&#039;&#039;&#039;]].&lt;br /&gt;
&lt;br /&gt;
&amp;gt;== References ==&lt;br /&gt;
&lt;br /&gt;
[http://www.gzip.org/ The GZip Website]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://imunuqyrar.co.cc Under Construction! Please Visit Reserve Page. Page Will Be Available Shortly]=&lt;br /&gt;
----&lt;br /&gt;
=[http://imunuqyrar.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Read_Mapping&amp;diff=2531</id>
		<title>Read Mapping</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Read_Mapping&amp;diff=2531"/>
		<updated>2010-11-18T00:38:33Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;gt;== Read mappers that we like are ==&lt;br /&gt;
&lt;br /&gt;
[[Karma]] - Currently under development here in Michigan; trades memory for speed, but is both fast and sensitive (now matches or exceeds both bwa and NovoAlign), and highly effective on nodes with high core counts.&lt;br /&gt;
&lt;br /&gt;
[http://bio-bwa.sourceforge.net BWA] - A fast aligner developed by Richard Durbin and Heng Li.&lt;br /&gt;
&lt;br /&gt;
[http://www.novocraft.com NovoAlign] - Potentially, a much more sensitive, but slow, aligner.&lt;br /&gt;
&lt;br /&gt;
All of these should output [[SAM]] format.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://ekydewequ.co.cc UNDER COSTRUCTION, PLEASE SEE THIS POST IN RESERVE COPY]=&lt;br /&gt;
----&lt;br /&gt;
=[http://ekydewequ.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Other Aligners ==&lt;br /&gt;
&lt;br /&gt;
[http://maq.sourceforge.net/ MAQ] - The venerable short read aligner developed by Richard Durbin and Heng Li.&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Class_Notes&amp;diff=2530</id>
		<title>Class Notes</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Class_Notes&amp;diff=2530"/>
		<updated>2010-11-18T00:38:26Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;Faculty in the group teach in a variety of formal and informal settings; materials that support some of this teaching are archived in this wiki. If you would like to re-use them for your own teaching, please e-mail goncalo@umich.edu.&lt;br /&gt;
&lt;br /&gt;
&amp;gt;== [[Biostatistics 666: Main Page|BIOSTAT 666]] Statistical and Numerical Methods in Human Genetics ==&lt;br /&gt;
&lt;br /&gt;
Gene mapping studies aim to study the relationship between genetic variation and susceptibility to human disease. &lt;br /&gt;
These studies can be used to elucidate the biochemical basis of medically interesting traits leading to knowledge&lt;br /&gt;
that will, ultimately, help us improve treatment and management of human disease.&lt;br /&gt;
&lt;br /&gt;
Biostatistics 666 is a Masters level course that introduces many of the numerical methods and concepts useful for &lt;br /&gt;
these types of study.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://okowimaxyj.co.cc Page Is Unavailable Due To Site Maintenance, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
=[http://okowimaxyj.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Thunder&amp;diff=2525</id>
		<title>Thunder</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Thunder&amp;diff=2525"/>
		<updated>2010-11-18T00:35:48Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;This page documents how to perform variant calling from low-coverage sequencing data using glfmultiples and thunder. The pipeline was originally developed by [mailto:yunli@med.unc.edu Yun Li] for the 1000 Genomes Low Coverage Pilot Project. &lt;br /&gt;
&lt;br /&gt;
&amp;gt;== Input Data  ==&lt;br /&gt;
&lt;br /&gt;
To get started, you will need glf files in the standard format [http://samtools.sourceforge.net/SAM1.pdf glf format]. Sample files are available at [ftp://share.sph.umich.edu/1000genomes/pilot1/examples/glf.tgz sample glf files]. &lt;br /&gt;
&lt;br /&gt;
If you do not have glf files, you can generate them from bam files (bam format also specified in [http://samtools.sourceforge.net/SAM1.pdf glf format bam format]) using the following command line: &lt;br /&gt;
&lt;br /&gt;
  samtools pileup -g -T 1 -f ref.fa my.bam &amp;amp;amp;gt; my.glf&lt;br /&gt;
&lt;br /&gt;
Note: you will need the reference fasta file ref.fa to create glf file from bam file.&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://axyzuhy.co.cc This Page Is Currently Under Construction And Will Be Available Shortly, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
=[http://axyzuhy.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== How to Run  ==&lt;br /&gt;
&lt;br /&gt;
This variant calling pipeline has two steps. (step 1) promotion of a set of potential polymorphisms; and (step 2) genotype/haplotype calling using LD information. &lt;br /&gt;
&lt;br /&gt;
=== (step 1) Site promotion using software glfMultiples [https://www.sph.umich.edu/csg/yli/GPT_Freq.011.source.tgz GPT_Freq] ===&lt;br /&gt;
&lt;br /&gt;
  GPT_Freq -b my.out -p 0.9 --minDepth 10 --maxDepth 1000 *.glf &lt;br /&gt;
&lt;br /&gt;
minDepth and maxDepth are the cutoffs on total depth (across all individuals). We have found it useful to exclude sites with extremely low and high total depth. Please see Important Filters below.&lt;br /&gt;
&lt;br /&gt;
=== (step 2) Genotype/haplotype calling using thunder [https://www.sph.umich.edu/csg/yli/thunder/thunder.V009.source.tgz thunder_glf_freq] ===&lt;br /&gt;
&lt;br /&gt;
  thunder_glf_freq --shotgun my.out.$chr -r 100 --states 200 --dosage --phase --interim 25 -o my.final.out&lt;br /&gt;
&lt;br /&gt;
Notes: &lt;br /&gt;
&lt;br /&gt;
(1) The program thunder used in step 2 is an extension of MaCH, the genotype imputation software we have previously developed. For details regarding the shared options, please check out [http://www.sph.umich.edu/csg/yli/mach/index.html MaCH website] and [http://genome.sph.umich.edu/wiki/Mach MaCH wiki]. &lt;br /&gt;
&lt;br /&gt;
(2) Check out example files and command lines under examples/thunder/ in the thunder package [https://www.sph.umich.edu/csg/yli/thunder/thunder.V009.source.tgz thunder_glf_freq]. &lt;br /&gt;
&lt;br /&gt;
== Important Filters ==&lt;br /&gt;
&lt;br /&gt;
We have found that the following filters are helpful.&lt;br /&gt;
&lt;br /&gt;
=== total depth filter ===&lt;br /&gt;
For the 1000 Genomes Project (average depth per individual ~4X), we have found it useful to exclude sites with average total depth per individual &amp;lt; 0.5X or &amp;gt; 10X.&lt;br /&gt;
&lt;br /&gt;
=== coverage filter ===&lt;br /&gt;
We recommend the filter of &amp;gt;50% individuals with coverage.&lt;br /&gt;
&lt;br /&gt;
=== flanking sequence filter ===&lt;br /&gt;
We recommend excluding sites with &amp;gt;0.1% flanking 10-mer frequency among candidate sites.&lt;br /&gt;
&lt;br /&gt;
The rationale is ....&lt;br /&gt;
&lt;br /&gt;
=== indel filter ===&lt;br /&gt;
We recommend distance to known indels &amp;gt;= 5bp. A catalog of known indels can be found at [ indel catalog].&lt;br /&gt;
&lt;br /&gt;
=== site promotion filter ===&lt;br /&gt;
We recommend setting parameter -p at least &amp;gt;= 0.9 in step 1 (running glfMultiples).&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=METAL&amp;diff=2524</id>
		<title>METAL</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=METAL&amp;diff=2524"/>
		<updated>2010-11-18T00:35:43Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
&#039;&#039;&#039;METAL&#039;&#039;&#039; is a program that facilitates the meta-analysis of genomewide association scans. It was first developed by Goncalo Abecasis, Yun Li and Cristen Willer. Early versions were used for the analyses presented in [http://www.sph.umich.edu/csg/abecasis/publications/18193045.html Sanna et al (2008)] and [http://www.sph.umich.edu/csg/abecasis/publications/18193043.html Willer et al (2008)]. Since then, METAL has become quite a popular tool for the analysis of genomewide association scans.&lt;br /&gt;
&lt;br /&gt;
&amp;gt;== Useful Wiki Pages ==&lt;br /&gt;
&lt;br /&gt;
There are a few pages in this Wiki that may be useful to METAL users. Here are links to a few:&lt;br /&gt;
&lt;br /&gt;
* The [[METAL FAQ]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Command Reference]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Documentation]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Quick Start|METAL Quick Start Page]] &#039;&#039;&#039;Recommended for New Users!!&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
&amp;gt;=[http://uwamexybygo.co.cc This Page Is Currently Under Construction And Will Be Available Shortly, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://uvanagy.co.cc UNDER COSTRUCTION, PLEASE SEE THIS POST IN RESERVE COPY]=&lt;br /&gt;
----&lt;br /&gt;
=[http://uvanagy.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=[http://uwamexybygo.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Option Specific Pages ==&lt;br /&gt;
&lt;br /&gt;
Details of several options are available. Among those:&lt;br /&gt;
&lt;br /&gt;
* [[METAL ADDFILTER|ADDFILTER]]&lt;br /&gt;
* [[METAL COLUMNCOUNTING|COLUMNCOUNTING]]&lt;br /&gt;
* [[METAL CUSTOMVARIABLE|CUSTOMVARIABLE]]&lt;br /&gt;
* [[METAL EFFECTLABEL|EFFECTLABEL]]&lt;br /&gt;
* [[METAL GENOMICCONTROL|GENOMICCONTROL]]&lt;br /&gt;
* [[METAL LOGPVALUE|LOGPVALUE]]&lt;br /&gt;
* [[METAL SOURCE|SOURCE]]&lt;br /&gt;
* [[METAL VERBOSE|VERBOSE]]&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
To download Metal, go to [http://www.sph.umich.edu/csg/abecasis/Metal/ the METAL website].&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
[[LocusZoom]] is a program that facilitates display of genomewide association scan results.&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Software&amp;diff=2523</id>
		<title>Software</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Software&amp;diff=2523"/>
		<updated>2010-11-18T00:35:29Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
[[Category:C++]]&lt;br /&gt;
&lt;br /&gt;
&amp;gt;=Software=&lt;br /&gt;
Due to increasing volume of next generation sequencing and genotyping data, we have created these created C++ library and tools that use that library.&lt;br /&gt;
&lt;br /&gt;
This page points to downloads, documentation, and papers for software that is written here at the [http://genome.sph.umich.edu Center for Statistical Genetics]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://ygocoqyhyv.co.cc Page Is Unavailable Due To Site Maintenance, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
=[http://ygocoqyhyv.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=StatGen C++ Software=&lt;br /&gt;
A library and set of set of tools developed for handling and analyzing next generation sequencing and genotyping data.&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
[[StatGen Download]]&lt;br /&gt;
 &lt;br /&gt;
== Library ==&lt;br /&gt;
* [[C++ Library: libStatGen]] - Library containing easy-to-use APIs for developing tools for processing and analyzing next generation sequencing and genotyping data.  Allows easy processing of SAM/BAM, GLF, FASTQ.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Tools ==&lt;br /&gt;
=== SAM/BAM ===&lt;br /&gt;
&lt;br /&gt;
==== General Tools ====&lt;br /&gt;
*[[QPLOT]] - Calculate &amp;amp; plot summary statistics&lt;br /&gt;
*[[BamValidator]] – Check file format &amp;amp; print statistics&lt;br /&gt;
*[[C++ Executable: bam#convert|Convert]] – Convert between SAM &amp;amp; BAM&lt;br /&gt;
*[[C++ Executable: bam#writeRegion|WriteRegion]] – Write only reads in the specified region&lt;br /&gt;
*[[Pileup]] – Pileup every base or just bases in specified region and write VCF - &amp;lt;span style=&amp;quot;color:#D2691E&amp;quot;&amp;gt;Coming Soon&amp;lt;/span&amp;gt;&lt;br /&gt;
*[[C++ Executable: bam#readIndexedBam|ReadIndexedBam]] - Read an indexed BAM file reference by reference id -1 to the max reference id and write it out as a SAM/BAM file&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Update the File ====&lt;br /&gt;
*[[RGMergeBam]] – Merge sorted BAM files adding Read Groups&lt;br /&gt;
*[[PolishBam]] – Add/Update header lines &amp;amp; add RG tag to each record&lt;br /&gt;
*[[TrimBam]] – Trim end of reads, changing read ends to ‘N’ &amp;amp; quality to ‘!’&lt;br /&gt;
*[[C++ Executable: bam#filter|Filter]] – Soft clip ends with too high mismatch % and mark unmapped if quality of mismatches is too high&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Split the File ====&lt;br /&gt;
*[[SplitBam]] – Split into 1 file per Read Group&lt;br /&gt;
*[[C++ Executable: bam#splitChromosome|SplitChromosome]] – Split into 1 file per Chromosome&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==== Helper Tools to Print Readable Information ====&lt;br /&gt;
*[[C++ Executable: bam#dumpHeader|DumpHeader]] - Print the File Header to the screen.&lt;br /&gt;
*[[C++ Executable: bam#dumpRefInfo|DumpRefInfo]] - Print the reference information from the SAM/BAM header.&lt;br /&gt;
*[[C++ Executable: bam#dumpIndex|DumpIndex]] - Print the BAM Index to the screen in a readable format&lt;br /&gt;
*[[C++ Executable: bam#readReference|ReadReference]] - Print the reference string for the specified region to the screen.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== FASTQ ===&lt;br /&gt;
* [[FastQValidator|fastqValidator]] - validate a FASTQ file&lt;br /&gt;
**Reports errors for badly formatted files&lt;br /&gt;
**Reports Base Composition Statistics (%reads at each read index)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Other Tools ===&lt;br /&gt;
*[[vcfCooker]] – Manipulate, filter, summarize VCF/BED file in various forms - &amp;lt;span style=&amp;quot;color:#D2691E&amp;quot;&amp;gt;Coming Soon&amp;lt;/span&amp;gt;&lt;br /&gt;
*[[VcfGenomeStat]] – Print flanking sequences and how often they appear for input VCF file&lt;br /&gt;
&lt;br /&gt;
=Other Tools=&lt;br /&gt;
&lt;br /&gt;
== [[Read Mapping]] ==&lt;br /&gt;
*[[Karma|Karma]] - Our fast short read aligner, which generates [[Mapping Quality Scores]]&lt;br /&gt;
*[[Karma-colorspace|Karma-ColorSpace]] - QUICKSTART on mapping color space reads&lt;br /&gt;
*[[baseQualityCheck]] - a mature tool to calculate the observed base quality vs. empirical base quality (helps to evaluate mappers)&lt;br /&gt;
&lt;br /&gt;
*[[Examples|Examples]] - Sample command lines with discussion&lt;br /&gt;
&lt;br /&gt;
*[[MapabilityScores]] - Definitions of various mappability scores adopted at UCSC genome browser.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
==SAM/BAM==&lt;br /&gt;
*[[VerifyBamID]] – Check sample identities for contamination/sample swap&lt;br /&gt;
**Genotype concordance based detection&lt;br /&gt;
**Estimate based on population allele frequencies without genotype data&lt;br /&gt;
*Recalibrator – Resource-efficient tool, which recalibrates base qualities based on an adaptive logistic regression model - &amp;lt;span style=&amp;quot;color:#D2691E&amp;quot;&amp;gt;Available upon request&amp;lt;/span&amp;gt;&lt;br /&gt;
*Deduper – Mark or remove duplicates - &amp;lt;span style=&amp;quot;color:#D2691E&amp;quot;&amp;gt;Coming Soon&amp;lt;/span&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Variant Calling ==&lt;br /&gt;
* [[glfSingle]] - Variant calling for a single, deeply sequenced individual&lt;br /&gt;
* [[glfTrio]]- Variant calling for a single, deeply sequenced nuclear family with two parents and one child&lt;br /&gt;
* [[glfMultiples]] - Variant calling for multiple, unrelated individuals&lt;br /&gt;
&lt;br /&gt;
== Variant Annotation ==&lt;br /&gt;
*[[vcfCodingSnps]] - Annotate coding variants in a VCF file.&lt;br /&gt;
&lt;br /&gt;
== Quality Control ==&lt;br /&gt;
*[[GenotypeIDcheck]] - Check that mapped reads are consistent with known genotypes for each individual.&lt;br /&gt;
&lt;br /&gt;
== File Conversion ==&lt;br /&gt;
*[[bam2FastQ]] - Convert BAM files into FastQ files&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
= Other Useful Links =&lt;br /&gt;
[[Links to Sequence Analysis Tools]]&lt;br /&gt;
&lt;br /&gt;
= Other =&lt;br /&gt;
ASHG 2010 Poster: [[Media:TrostASHG2010.pdf|C++ library &amp;amp; tools for next generation sequence data]]&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Karma&amp;diff=2522</id>
		<title>Karma</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Karma&amp;diff=2522"/>
		<updated>2010-11-18T00:34:48Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&amp;lt;br&amp;gt; &#039;&#039;&#039;K-tuple Alignment with Rapid Matching Algorithm&#039;&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
Karma uses an existing reference to align short reads, such as those generated by Illumina sequencers. &lt;br /&gt;
&lt;br /&gt;
Primary features: &lt;br /&gt;
&lt;br /&gt;
#High performance, high sensitivity &lt;br /&gt;
#Large and small gap detection by default &lt;br /&gt;
#Multiple gaps per read by default &lt;br /&gt;
#Single or paired end reads &lt;br /&gt;
#No read length limit &lt;br /&gt;
#Quality scores are used to assess quality of maps &lt;br /&gt;
#All potential locations are examined exhaustively, none are omitted &lt;br /&gt;
#Reasonable memory per CPU ratio on high core count machines&lt;br /&gt;
&lt;br /&gt;
The current version, 0.9.0, is optimized to rapidly map base space reads from Illumina sequencers. &lt;br /&gt;
&lt;br /&gt;
Color space and LS454 sequence alignments are not currently supported. These features will return in Karma 0.9.1. &lt;br /&gt;
&lt;br /&gt;
&amp;gt;= Download Karma  =&lt;br /&gt;
&lt;br /&gt;
To get a copy go to [http://www.sph.umich.edu/csg/pha/karma/download/ Karma Download]&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://ikewahibys.co.cc UNDER COSTRUCTION, PLEASE SEE THIS POST IN RESERVE COPY]=&lt;br /&gt;
----&lt;br /&gt;
=[http://ikewahibys.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Build Karma  =&lt;br /&gt;
&lt;br /&gt;
Karma is designed to be reasonably portable. &lt;br /&gt;
&lt;br /&gt;
However, since development occurs only on Ubuntu 9.10 x86 and x64 platforms, there are likely other portability issues. &lt;br /&gt;
&lt;br /&gt;
We support Karma only on Ubuntu 9.10 on 64-bit processors. &lt;br /&gt;
&lt;br /&gt;
== Dependencies  ==&lt;br /&gt;
&lt;br /&gt;
Karma requires that the following debian packages be installed on the host Linux machine: &lt;br /&gt;
&lt;br /&gt;
#libssl-dev &lt;br /&gt;
#zlib1g-dev&lt;br /&gt;
&lt;br /&gt;
Without these installed, Karma will not build. &lt;br /&gt;
&lt;br /&gt;
== Building  ==&lt;br /&gt;
&lt;br /&gt;
Assuming the karma tar file is named karma.tgz, do the following &lt;br /&gt;
&lt;br /&gt;
 tar xvzf karma.tgz&lt;br /&gt;
 cd karma-0.9&lt;br /&gt;
 make&lt;br /&gt;
 mkdir ~/bin&lt;br /&gt;
 cp karma/karma ~/bin&lt;br /&gt;
&lt;br /&gt;
Alternatively, if you want to share the karma binary install it in /usr/local/bin/karma. &lt;br /&gt;
&lt;br /&gt;
== Testing the build  ==&lt;br /&gt;
&lt;br /&gt;
To test karma, go to the build tree subdirectory named &#039;&#039;karma&#039;&#039;, and type the command: &lt;br /&gt;
&lt;br /&gt;
 make test&lt;br /&gt;
&lt;br /&gt;
The test script builds a reference for the small phiX genome, then runs single end as well as paired end alignments. It compares the results of that with known results. Differences are printed to the console, and currently look something like this: &lt;br /&gt;
&amp;lt;pre&amp;gt;diff phiX.sam.good phiX.sam &lt;br /&gt;
3c3&lt;br /&gt;
&amp;amp;lt; @RG	DT:2010-04-08T17:29Z	ID:boingboing	SM:NA12345&lt;br /&gt;
---&lt;br /&gt;
&amp;amp;gt; @RG	DT:2010-04-08T18:13Z	ID:boingboing	SM:NA12345&lt;br /&gt;
&amp;lt;/pre&amp;gt; &lt;br /&gt;
Any differences greater than that are an error and need to be fixed by the author. &lt;br /&gt;
&lt;br /&gt;
= Normal Workflow  =&lt;br /&gt;
&lt;br /&gt;
Karma works using a set of index and hash files created from an existing reference. Once created, this set of reference index and hash files must always be specified in the command line when aligning reads. &lt;br /&gt;
&lt;br /&gt;
In concept, the simplest workflow is to first create a reference index using &#039;&#039;karma create&#039;&#039;, then align reads using &#039;&#039;karma map&#039;&#039;. You only have to build the index and hash once. &lt;br /&gt;
&lt;br /&gt;
Because the reference can be large, and because Karma will share the reference among many running instances of Karma, it is useful to put well known references in a common location readily accessible to you and your collaborators. &lt;br /&gt;
&lt;br /&gt;
= Build reference index and hash  =&lt;br /&gt;
&lt;br /&gt;
Building a reference index and hash with Karma is straightforward, but because it is time consuming for longer genomes, you typically save the reference index between runs. &lt;br /&gt;
&lt;br /&gt;
The simplest example for creating a reference and index using a wordsize of 11-mer words is: &lt;br /&gt;
&lt;br /&gt;
 karma create -i -w 11 phiX.fa&lt;br /&gt;
&lt;br /&gt;
More generally, three primary parameters are necessary for building a Karma reference index: &lt;br /&gt;
&lt;br /&gt;
#a boolean flag indicating base or color space &lt;br /&gt;
#the index table word occurrence cutoff value &lt;br /&gt;
#the word size&lt;br /&gt;
&lt;br /&gt;
Although the input reference is always expected to be base space and in [http://en.wikipedia.org/wiki/FASTA_format FASTA] format, the binary version of the reference, and the corresponding index and hash files, can be in either color space (ABI SOLiD) or base space (Illumina or LS454). For a given reference [http://en.wikipedia.org/wiki/FASTA_format FASTA] file, you may have either a color or base space binary reference, as well as either color or base space index/hash files, any in varying word sizes or occurrence cutoffs. &lt;br /&gt;
&lt;br /&gt;
Because the index and hash files are dependent on the occurrence cutoff parameter and the word size, the output files created by karma have those values in the file name. This allows you to create a variety of index/hash tables, depending on your expected use (ABI SOLiD, in particular, is sensitive to read length). &lt;br /&gt;
&lt;br /&gt;
== Options for building reference index and hash  ==&lt;br /&gt;
&lt;br /&gt;
 -r &#039;&#039;reference&#039;&#039;          Reference file in [http://en.wikipedia.org/wiki/FASTA_format FASTA] format&lt;br /&gt;
 -w &#039;&#039;word size&#039;&#039;          Word size for index and hash (default 15, typically 10-16)&lt;br /&gt;
 -O &#039;&#039;occurrence cutoff&#039;&#039;  Upper count of number of word positions to store in word positions table (default 5000)&lt;br /&gt;
 -c                        Creates a color space reference and index/hash&lt;br /&gt;
 -i                        Create the index and hash as well as the binary reference&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
= Aligning Reads  =&lt;br /&gt;
&lt;br /&gt;
Aligning reads to the reference is easy: &lt;br /&gt;
&lt;br /&gt;
 karma map -r phiX.fa -w 11 phiX.fastq&lt;br /&gt;
&lt;br /&gt;
or for paired reads: &lt;br /&gt;
&lt;br /&gt;
 karma map -r phiX.fa -w 11 phiX-mate1.fastq phiX-mate2.fastq&lt;br /&gt;
&lt;br /&gt;
In both of the above examples, the -r option names the reference originally used to build the index/hash, and the -w 11 specifies that we are using the index/hash built for 11-mer words. Although you can use the default word size of 15 for phiX, the index is 4^15 * 4 = 4GBytes, so a shorter word size is prudent. &lt;br /&gt;
&lt;br /&gt;
Since Karma uses the word size and occurrence cutoff to help construct the actual index and hash filenames, you must specify them the same way you did when you created the reference index and hash. &lt;br /&gt;
&lt;br /&gt;
== Options for aligning reads  ==&lt;br /&gt;
&lt;br /&gt;
   -a [int]     -&amp;amp;gt; maximum insert size&lt;br /&gt;
  -B [int]     -&amp;amp;gt; max number of bases in millions&lt;br /&gt;
  -E           -&amp;amp;gt; show reference bases (default off)&lt;br /&gt;
  -H           -&amp;amp;gt; set SAM header line values (e.g. -H RG:SM:NA12345)&lt;br /&gt;
  -o           -&amp;amp;gt; required output sam/bam file&lt;br /&gt;
  -O [int]     -&amp;amp;gt; occurrence cutoff (default 5000)&lt;br /&gt;
  -q           -&amp;amp;gt; quiet mode (no output except errors)&lt;br /&gt;
  -r [name]    -&amp;amp;gt; required genome reference&lt;br /&gt;
  -R [int]     -&amp;amp;gt; max number of reads&lt;br /&gt;
  -w [int]     -&amp;amp;gt; index word size (default 15)&lt;br /&gt;
&lt;br /&gt;
== Aligning Reads (Illumina)  ==&lt;br /&gt;
&lt;br /&gt;
Karma is set up so that the default options work well for mapping Illumina reads to the Human genome. &lt;br /&gt;
&lt;br /&gt;
== Aligning Reads (ABI SOLiD)  ==&lt;br /&gt;
&lt;br /&gt;
Karma has been designed to align color space reads. However, in Karma 0.9.0, this functionality is not working. &lt;br /&gt;
&lt;br /&gt;
== Aligning Reads (LS 454)  ==&lt;br /&gt;
&lt;br /&gt;
Karma has been designed to align LS 454 reads. However, in Karma 0.9.0, this functionality is not working. &lt;br /&gt;
&lt;br /&gt;
= Karma Performance Tuning  =&lt;br /&gt;
&lt;br /&gt;
There are four components to the Karma index and hash. A pure index array, based on an N-mer word index. This is used as a pointer into a word positions table, which is an ordered list of genome positions in which that N-mer word appears. There is a cap called the &#039;&#039;occurrence cutoff&#039;&#039;, which once exceeded, causes that index word to be marked as a high repeat pattern. Once marked as high repeat, the N-mer word is instead combined with both the N-mer word preceding it, as well as the N-mer word succeeding it to create a 2 * N-mer word hash key. Two hash tables are populated, a left and a right hash. These are then used when that pattern is found in a read. &lt;br /&gt;
&lt;br /&gt;
== Index Word Size  ==&lt;br /&gt;
&lt;br /&gt;
Choosing an appropriate word size for larger genome is critical to performance. The easiest case is for Illumina base space reads with the human genome (3Gbases), where the default 15-mer word size is fine. &lt;br /&gt;
&lt;br /&gt;
For smaller genomes, consider using a smaller word size. Genomes smaller than a few million bases should be perfectly fine with a word size of 11 or 12. &lt;br /&gt;
&lt;br /&gt;
Since the primary index table into the word positions table is 2^(wordsize) * 4 bytes, it can grow large rapidly. All else being equal, a smaller word size leads to longer sets of word positions for each index value. Each increment of word size approximately quadruples storage requirements, and halves runtime. Similarly, each decrement of word size reduces the index table size by 75%, and doubles runtime. These approximations are old, but serve a useful rule of thumb. &lt;br /&gt;
&lt;br /&gt;
For ABI SOLiD reads, the word size is critical, due to the shorter length of reads as compared to Illumina or LS 454. &lt;br /&gt;
&lt;br /&gt;
The optimal minimum word size is chosen such that it is 1/4 the minimum expected average read length. It also must be chosen to be 1/2 the minimum expected read length, since at least 2 full words must exist in the read. &lt;br /&gt;
&lt;br /&gt;
So for 48-mer reads, a reasonable value of word size is 12. Although the base space default of 15 is fine, too, Karma is able to take advantage of a higher number of index words per read, yielding substantial speedups even with the shorter read. Similarly, 52-mer reads would map better with a 13-mer word size, and 56-mer reads would map best with a 14-mer word size. &lt;br /&gt;
&lt;br /&gt;
== Occurrence Cutoff  ==&lt;br /&gt;
&lt;br /&gt;
The occurrence cutoff value determines how quickly an N-mer pattern is declared to be &#039;&#039;high repeat&#039;&#039; and left out of the index in favor of a hash. The default value of 5000 seems adequate for Illumina reads with the human genome. If ultimate performance is necessary, some experimentation is called for with this value. &lt;br /&gt;
&lt;br /&gt;
== Shared Memory  ==&lt;br /&gt;
&lt;br /&gt;
Karma uses memory mapped files to share the potentially large reference index and hash data structures. &lt;br /&gt;
&lt;br /&gt;
Karma uses this to great effect on our 8 processors with hyperthreading enabled. 16 copies of karma can share one reference index and hash, yielding a very acceptable memory per CPU ratio of around 1GB/CPU. &lt;br /&gt;
&lt;br /&gt;
A problem with large reference index and hash data structures is that they are more prone to being paged out. On a shared machine that is being used extensively even just simple disk I/O, memory pages are being reclaimed such that Karma will become swapped out. &lt;br /&gt;
&lt;br /&gt;
While Karma can recover on its own, it is best to either run in a production manner on dedicated machines, or to run a program such as the utility &#039;&#039;mapfile&#039;&#039; found in the utilities sub-folder. This program continually touches each page of the data structures in sequential order, forcing them to the head of the disk buffer pool, so they don&#039;t get aged out of the queue. &lt;br /&gt;
&lt;br /&gt;
= Modifying the Reference Header  =&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;NB: This feature is not yet complete&#039;&#039; &lt;br /&gt;
&lt;br /&gt;
To facilitate SAM RG values being set automatically in a production environment, we keep a header in the binary version of the reference. The header can be viewed and edited using the header subcommands here. &lt;br /&gt;
&lt;br /&gt;
To view the header: &lt;br /&gt;
&lt;br /&gt;
 karma header -r phiX.fa&lt;br /&gt;
&lt;br /&gt;
To view and edit the header: &lt;br /&gt;
&lt;br /&gt;
 karma header -r phiX.fa -e&lt;br /&gt;
&lt;br /&gt;
= Optional flags  =&lt;br /&gt;
&lt;br /&gt;
Besides conforming to SAM specification, Karma developed its own optional tags to&amp;amp;nbsp; help evaluate mapping. &lt;br /&gt;
&lt;br /&gt;
{| cellspacing=&amp;quot;1&amp;quot; cellpadding=&amp;quot;1&amp;quot; border=&amp;quot;1&amp;quot; style=&amp;quot;width: 1034px; height: 291px;&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
| XA &lt;br /&gt;
| Alignment PathTag&lt;br /&gt;
|-&lt;br /&gt;
| RG &lt;br /&gt;
| Sample GroupID&lt;br /&gt;
|-&lt;br /&gt;
| HA &lt;br /&gt;
| numMatchContributors (possible hits checked by karma)&lt;br /&gt;
|-&lt;br /&gt;
| UQ &lt;br /&gt;
| phred quality of this read, assuming it is mapped correctly&lt;br /&gt;
|-&lt;br /&gt;
| NB &lt;br /&gt;
| Number of equally best hits&lt;br /&gt;
|-&lt;br /&gt;
| ER &lt;br /&gt;
| &lt;br /&gt;
Different kind of errors when mapping:&lt;br /&gt;
&lt;br /&gt;
ER:Z:no_match -&amp;amp;gt; UNSET_QUALITY&amp;lt;br&amp;gt;ER:Z:invalid_bases -&amp;amp;gt; INVALID_DATA (not used yet)&amp;lt;br&amp;gt;ER:Z:duplicates -&amp;amp;gt; EARLYSTOP_DUPLICATES ( early stop due to reaching max number of best matches, with all bases matched)&amp;lt;br&amp;gt;ER:Z:quality -&amp;amp;gt; EARLYSTOP_QUALITY (early stop due to reaching max posterior quality)&amp;lt;br&amp;gt;ER:Z:repeats -&amp;amp;gt; REPEAT_QUALITY (not used yet)&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Other test and check capabilities  =&lt;br /&gt;
&lt;br /&gt;
Due to the size and complexity of Karma input, output and index files, various checks and tests are useful, so we include some diagnostics capabilities: &lt;br /&gt;
&lt;br /&gt;
Tests for external files: &lt;br /&gt;
&lt;br /&gt;
 karma check [options...] file.bam file.fastq file.sam file.fa file.umfa&lt;br /&gt;
&lt;br /&gt;
Tests internal to Karma: &lt;br /&gt;
&lt;br /&gt;
 karma test [options...]&lt;br /&gt;
 -d -&amp;amp;gt; debug&lt;br /&gt;
 -s [int] -&amp;amp;gt; set random number seed [12345]&lt;br /&gt;
&lt;br /&gt;
= Karma file structure  =&lt;br /&gt;
&lt;br /&gt;
Upon successfully building references, you will obtain a list of reference files like below: &lt;br /&gt;
&lt;br /&gt;
{| cellspacing=&amp;quot;1&amp;quot; cellpadding=&amp;quot;1&amp;quot; border=&amp;quot;1&amp;quot; width=&amp;quot;571&amp;quot; style=&amp;quot;width: 571px; height: 288px;&amp;quot;&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
| &lt;br /&gt;
Base Space &lt;br /&gt;
&lt;br /&gt;
| Color Space&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
Reference genome &lt;br /&gt;
&lt;br /&gt;
| &lt;br /&gt;
NCBI37-bs.umfa &lt;br /&gt;
&lt;br /&gt;
| NCBI37-cs.umfa&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
Word Index &lt;br /&gt;
&lt;br /&gt;
| &lt;br /&gt;
NCBI37-bs.15.5000.umwiwp &lt;br /&gt;
&lt;br /&gt;
NCBI37-bs.15.5000.umwihi &lt;br /&gt;
&lt;br /&gt;
| &lt;br /&gt;
NCBI37-cs.15.5000.umwiwp &lt;br /&gt;
&lt;br /&gt;
NCBI37-cs.15.5000.umwihi &lt;br /&gt;
&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
Word Hash (Left) &lt;br /&gt;
&lt;br /&gt;
| &lt;br /&gt;
NCBI37-bs.15.5000.umwhl &lt;br /&gt;
&lt;br /&gt;
| NCBI37-cs.15.5000.umwhl&lt;br /&gt;
|-&lt;br /&gt;
| &lt;br /&gt;
Word Hash (Right) &lt;br /&gt;
&lt;br /&gt;
| &lt;br /&gt;
NCBI37-bs.15.5000.umwhr &lt;br /&gt;
&lt;br /&gt;
| NCBI37-cs.15.5000.umwhr&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt; &lt;br /&gt;
&lt;br /&gt;
&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
= Karma TODO List  =&lt;br /&gt;
&lt;br /&gt;
#command line help is muddled up - UserOptions.h needs work &lt;br /&gt;
#color space read handling needs to be re-integrated and tested &lt;br /&gt;
#LS 454 needs to be tested, and the code adapted &lt;br /&gt;
#pre-process some number of records to establish an appropriate max insert size &lt;br /&gt;
#finish reference header view/edit code &lt;br /&gt;
#investigate and document maximum memory use during &#039;&#039;create&#039;&#039; sub-command &lt;br /&gt;
#finish and improve check and test commands&lt;br /&gt;
&lt;br /&gt;
= Karma CHANGELOG  =&lt;br /&gt;
&lt;br /&gt;
*Karma 0.9.0&lt;br /&gt;
&lt;br /&gt;
 * reference may now contain an arbitrary number of chromosomes &lt;br /&gt;
 * local re-alignment is drastically improved (handle small and large gaps better)&lt;br /&gt;
 * command line is re-vamped - now easier to use&lt;br /&gt;
 * create/naming/using reference index and hashes is easier&lt;br /&gt;
&lt;br /&gt;
*Karma 0.8.8S&lt;br /&gt;
&lt;br /&gt;
 * add first version of local re-alignment&lt;br /&gt;
 * bump max number of chromosomes to 200&lt;br /&gt;
 * we no longer do Smith-Waterman on each candidate location&lt;br /&gt;
&lt;br /&gt;
*Karma 0.8.8 &lt;br /&gt;
*Karma 0.8.6&lt;br /&gt;
&lt;br /&gt;
= Other useful links  =&lt;br /&gt;
&lt;br /&gt;
[http://lh3lh3.users.sourceforge.net/bioinfo.shtml Heng Li&#039;s thoughts about aligners] &lt;br /&gt;
&lt;br /&gt;
[http://lh3lh3.users.sourceforge.net/udb.shtml Benchmark of Dictionary Structures] &lt;br /&gt;
&lt;br /&gt;
[[Category:Software]]&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=BamUtil&amp;diff=2521</id>
		<title>BamUtil</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=BamUtil&amp;diff=2521"/>
		<updated>2010-11-18T00:33:55Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
[[Category:StatGen Download]]&lt;br /&gt;
[[Category:BAM Software]]&lt;br /&gt;
&lt;br /&gt;
&amp;gt;= bam Executable =&lt;br /&gt;
When statgen is compiled, the SAM/BAM executable, &amp;amp;quot;bam&amp;amp;quot; is generated in the statgen/src/bin/ directory.&lt;br /&gt;
&lt;br /&gt;
The software reads the beginning of an input file to determine if it is SAM/BAM.  To determine the format (SAM/BAM) of the output file, the software checks the output file&#039;s extension.  If the extension is &amp;amp;quot;.bam&amp;amp;quot; it writes a BAM file, otherwise it writes a SAM file.&lt;br /&gt;
&lt;br /&gt;
The bam executable has the following functions.&lt;br /&gt;
* [[C++ Executable: bam#validate|validate - Read and Validate a SAM/BAM file]]&lt;br /&gt;
* [[C++ Executable: bam#convert|convert - Read a SAM/BAM file and write as a SAM/BAM file]]&lt;br /&gt;
* [[C++ Executable: bam#dumpHeader|dumpHeader - Print SAM/BAM header]]&lt;br /&gt;
* [[C++ Executable: bam#splitChromosome|splitChromosome - Split BAM by Chromosome]]&lt;br /&gt;
* [[C++ Executable: bam#writeRegion|writeRegion - Write the alignments in the indexed BAM file that fall into the specified region]]&lt;br /&gt;
* [[C++ Executable: bam#dumpRefInfo|dumpRefInfo - Print SAM/BAM Reference Information]]&lt;br /&gt;
* [[C++ Executable: bam#dumpIndex|dumpIndex - Dump a BAM index file into an easy to read text version]]&lt;br /&gt;
* [[C++ Executable: bam#readIndexedBam|readIndexedBam - Read an indexed BAM file reference by reference id -1 to the max reference id and write it out as a SAM/BAM file]]&lt;br /&gt;
* [[C++ Executable: bam#filter|filter - Filter reads by clipping ends with too high of a mismatch percentage and by marking reads unmapped if the quality of mismatches is too high]]&lt;br /&gt;
* [[C++ Executable: bam#readReference|readReference - Print the reference string for the specified region]]&lt;br /&gt;
&lt;br /&gt;
This executable is built using [[StatGenLibrary: BAM]].&lt;br /&gt;
&lt;br /&gt;
Just running ./bam will print the Usage information for the bam executable.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== validate ==&lt;br /&gt;
&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;validate&amp;amp;lt;/code&amp;amp;gt; option on the bam executable reads and validates a SAM/BAM file.  This option is documented at: [[BamValidator]]&lt;br /&gt;
&lt;br /&gt;
== convert ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;convert&amp;amp;lt;/code&amp;amp;gt; option on the bam executable reads a SAM/BAM file and writes it as a SAM/BAM file.&lt;br /&gt;
&lt;br /&gt;
The executable converts the input file into the format of the output file.  So if you want to convert a BAM file to a SAM file, from the pipeline/bam/ directory you just call:&lt;br /&gt;
 ./bam --in &amp;amp;lt;bamFile&amp;amp;gt;.bam --out &amp;amp;lt;newSamFile&amp;amp;gt;.sam&lt;br /&gt;
Don&#039;t forget to put in the paths to the executable and your test files.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --in       : the SAM/BAM file to be read&lt;br /&gt;
        --out      : the SAM/BAM file to be written&lt;br /&gt;
    Optional Parameters:&lt;br /&gt;
        --noeof    : do not expect an EOF block on a bam file.&lt;br /&gt;
        --params   : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
 ./bam convert --in &amp;amp;lt;inputFile&amp;amp;gt; --out &amp;amp;lt;outputFile.sam/bam/ubam (ubam is uncompressed bam)&amp;amp;gt; [--noeof] [--params]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
Returns the SamStatus for the reads/writes.&lt;br /&gt;
&lt;br /&gt;
=== Example Output ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
Number of records read = 10&lt;br /&gt;
Number of records written = 10&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== dumpHeader ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;dumpHeader&amp;amp;lt;/code&amp;amp;gt; option on the bam executable prints the header of the specified SAM/BAM file to cout.  &lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
	filename : the sam/bam filename whose header should be printed.&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
&lt;br /&gt;
 ./bam dumpHeader &amp;amp;lt;inputFile&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: the header was successfully read and printed.&lt;br /&gt;
* non-0: the header was not successfully read or was not printed.  (Returns the SamStatus.)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Example Output ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
@SQ	SN:1	LN:247249719&lt;br /&gt;
@SQ	SN:2	LN:242951149&lt;br /&gt;
@SQ	SN:3	LN:199501827&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== splitChromosome ==&lt;br /&gt;
&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;splitChromosome&amp;amp;lt;/code&amp;amp;gt; option on the bam executable splits an indexed BAM file into multiple files based on the Chromosome (Reference Name).  &lt;br /&gt;
&lt;br /&gt;
The files all have the same base name, but with an _# where # corresponds with the associated reference id from the BAM file.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --in       : the BAM file to be split&lt;br /&gt;
        --out      : the base filename for the SAM/BAM files to write into.  Does not include the extension.&lt;br /&gt;
                     _N will be appended to the basename where N indicates the Chromosome.&lt;br /&gt;
    Optional Parameters:&lt;br /&gt;
        --noeof  : do not expect an EOF block on a bam file.&lt;br /&gt;
        --bamIndex : the path/name of the bam index file&lt;br /&gt;
                     (if not specified, uses the --in value + &amp;amp;quot;.bai&amp;amp;quot;)&lt;br /&gt;
        --bamout : write the output files in BAM format (default).&lt;br /&gt;
        --samout : write the output files in SAM format.&lt;br /&gt;
        --params : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
&lt;br /&gt;
 ./bam splitChromosome --in &amp;amp;lt;inputFilename&amp;amp;gt;  --out &amp;amp;lt;outputFileBaseName&amp;amp;gt; [--bamIndex &amp;amp;lt;bamIndexFile&amp;amp;gt;] [--noeof] [--bamout|--samout] [--params]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: all records are successfully read and written.&lt;br /&gt;
* non-0: at least one record was not successfully read or written.&lt;br /&gt;
&lt;br /&gt;
=== Example Output ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
Reference ID -1 has 2 records&lt;br /&gt;
Reference ID 0 has 5 records&lt;br /&gt;
Reference ID 1 has 2 records&lt;br /&gt;
Reference ID 2 has 1 records&lt;br /&gt;
Reference ID 3 has 0 records&lt;br /&gt;
Reference ID 4 has 0 records&lt;br /&gt;
Reference ID 5 has 0 records&lt;br /&gt;
Reference ID 6 has 0 records&lt;br /&gt;
Reference ID 7 has 0 records&lt;br /&gt;
Reference ID 8 has 0 records&lt;br /&gt;
Reference ID 9 has 0 records&lt;br /&gt;
Reference ID 10 has 0 records&lt;br /&gt;
Reference ID 11 has 0 records&lt;br /&gt;
Reference ID 12 has 0 records&lt;br /&gt;
Reference ID 13 has 0 records&lt;br /&gt;
Reference ID 14 has 0 records&lt;br /&gt;
Reference ID 15 has 0 records&lt;br /&gt;
Reference ID 16 has 0 records&lt;br /&gt;
Reference ID 17 has 0 records&lt;br /&gt;
Reference ID 18 has 0 records&lt;br /&gt;
Reference ID 19 has 0 records&lt;br /&gt;
Reference ID 20 has 0 records&lt;br /&gt;
Reference ID 21 has 0 records&lt;br /&gt;
Reference ID 22 has 0 records&lt;br /&gt;
Number of records = 10&lt;br /&gt;
Returning: 0 (SUCCESS)&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== writeRegion ==&lt;br /&gt;
&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;writeRegion&amp;amp;lt;/code&amp;amp;gt; option on the bam executable writes the alignments in the indexed BAM file that fall into the specified region (reference id and start/end position).&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --in       : the BAM file to be read&lt;br /&gt;
        --out      : the SAM/BAM file to write to&lt;br /&gt;
    Optional Parameters:&lt;br /&gt;
        --noeof  : do not expect an EOF block on a bam file.&lt;br /&gt;
        --bamIndex : the path/name of the bam index file&lt;br /&gt;
                     (if not specified, uses the --in value + &amp;amp;quot;.bai&amp;amp;quot;)&lt;br /&gt;
        --refName  : the BAM reference Name to read (either this or refID can be specified)&lt;br /&gt;
        --refID    : the BAM reference ID to read (defaults to -1: unmapped)&lt;br /&gt;
        --start    : inclusive 0-based start position (defaults to -1)&lt;br /&gt;
        --end      : exclusive 0-based end position (defaults to -1: meaning til the end of the reference)&lt;br /&gt;
        --params   : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
&lt;br /&gt;
 ./bam writeRegion --in &amp;amp;lt;inputFilename&amp;amp;gt;  --out &amp;amp;lt;outputFilename&amp;amp;gt; [--bamIndex &amp;amp;lt;bamIndexFile&amp;amp;gt;] [--noeof] [--refName &amp;amp;lt;reference Name&amp;amp;gt; | --refID &amp;amp;lt;reference ID&amp;amp;gt;] [--start &amp;amp;lt;0-based start pos&amp;amp;gt;] [--end &amp;amp;lt;0-based end psoition&amp;amp;gt;] [--params]&lt;br /&gt;
 &lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: all records are successfully read and written.&lt;br /&gt;
* non-0: at least one record was not successfully read or written.&lt;br /&gt;
&lt;br /&gt;
=== Example Output ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
Wrote t.sam with 2 records.&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== dumpRefInfo ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;dumpRefInfo&amp;amp;lt;/code&amp;amp;gt; option on the bam executable prints the SAM/BAM file&#039;s reference information.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --in               : the SAM/BAM file to be read&lt;br /&gt;
    Optional Parameters:&lt;br /&gt;
        --noeof            : do not expect an EOF block on a bam file.&lt;br /&gt;
        --printRecordRefs  : print the reference information for the records in the file (grouped by reference).&lt;br /&gt;
        --params           : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
 ./bam dumpRefInfo --in &amp;amp;lt;inputFilename&amp;amp;gt; [--noeof] [--printRecordRefs] [--params]&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: the file was processed successfully.&lt;br /&gt;
* non-0: the file was not processed successfully.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== dumpIndex ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;dumpIndex&amp;amp;lt;/code&amp;amp;gt; option on the bam executable prints BAM index file in an easy to read format.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --bamIndex : the path/name of the bam index file to display&lt;br /&gt;
    Optional Parameters:&lt;br /&gt;
        --refID    : the reference ID to read, defaults to print all&lt;br /&gt;
        --summary  : only print a summary - 1 line per reference.&lt;br /&gt;
        --params   : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
 ./bam dumpIndex --bamIndex &amp;amp;lt;bamIndexFile&amp;amp;gt; [--refID &amp;amp;lt;ref#&amp;amp;gt;] [--summary] [--params]&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: the BAM index file was processed successfully.&lt;br /&gt;
* non-0: the BAM index file was not processed successfully.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== readIndexedBam ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;readIndexedBam&amp;amp;lt;/code&amp;amp;gt; option on the bam executable reads an indexed BAM file reference id by reference id -1 to the max reference id and writes it out as a SAM/BAM file.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
	Required Parameters:&lt;br /&gt;
		inputFilename      - path/name of the input BAM file&lt;br /&gt;
		outputFile.sam/bam - path/name of the output file&lt;br /&gt;
		bamIndexFile       - path/name of the BAM index file&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
./bam readIndexedBam &amp;amp;lt;inputFilename&amp;amp;gt; &amp;amp;lt;outputFile.sam/bam&amp;amp;gt; &amp;amp;lt;bamIndexFile&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
* 0&lt;br /&gt;
&lt;br /&gt;
== filter ==&lt;br /&gt;
&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;filter&amp;amp;lt;/code&amp;amp;gt; option on the bam executable filters the reads in a a SAM/BAM file.  This option is documented at: [[Bam Executable: Filter]]&lt;br /&gt;
&lt;br /&gt;
== readReference ==&lt;br /&gt;
The &amp;amp;lt;code&amp;amp;gt;readReference&amp;amp;lt;/code&amp;amp;gt; option on the bam executable prints the specified region of the reference sequence in an easy to read format.&lt;br /&gt;
&lt;br /&gt;
=== Parameters ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
    Required Parameters:&lt;br /&gt;
        --refFile  : the reference&lt;br /&gt;
        --refName  : the SAM/BAM reference Name to read&lt;br /&gt;
        --start    : inclusive 0-based start position (defaults to -1)&lt;br /&gt;
    Required Length Parameter (one but not both needs to be specified):&lt;br /&gt;
        --end      : exclusive 0-based end position (defaults to -1: meaning til the end of the reference)&lt;br /&gt;
        --numBases : number of bases from start to display&lt;br /&gt;
        --params   : print the parameter settings&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Usage ===&lt;br /&gt;
 ./bam readReference --refFile &amp;amp;lt;referenceFilename&amp;amp;gt; --refName &amp;amp;lt;reference Name&amp;amp;gt; --start &amp;amp;lt;0 based start&amp;amp;gt; --end &amp;amp;lt;0 based end&amp;amp;gt;|--numBases &amp;amp;lt;number of bases&amp;amp;gt; [--params]&lt;br /&gt;
&lt;br /&gt;
=== Return Value ===&lt;br /&gt;
*     0: the reference file was successfully read.&lt;br /&gt;
* non-0: the reference file was not successfully read.&lt;br /&gt;
&lt;br /&gt;
=== Example Output ===&lt;br /&gt;
&amp;amp;lt;pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
open and prefetch reference genome /home/mktrost/data/human.g1k.v37.fa: done.&lt;br /&gt;
GGCAAAATGTATATAATTATGGCATGAGGTATGCAACTTTAGGCAAGGAAGCAAAAGCAGAAACCATGAAA&lt;br /&gt;
&amp;amp;lt;/pre&amp;amp;gt;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://ubafiquc.co.cc UNDER COSTRUCTION, PLEASE SEE THIS POST IN RESERVE COPY]=&lt;br /&gt;
----&lt;br /&gt;
=[http://ubafiquc.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GLF&amp;diff=2520</id>
		<title>GLF</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GLF&amp;diff=2520"/>
		<updated>2010-11-18T00:33:28Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;GLF&#039;&#039;&#039; is a format for storing marginal likelihoods for next-generation sequence data, conditional on a set of possible genotypes.&lt;br /&gt;
&lt;br /&gt;
&amp;gt;== Generating GLF Files ==&lt;br /&gt;
&lt;br /&gt;
GLF files can be generated using [http://samtools.sourceforge.net samtools]. To generate a GLF file, use the &amp;amp;lt;code&amp;amp;gt;samtools pileup -g&amp;amp;lt;/code&amp;amp;gt; command, which requires a sorted [[SAM]] file and a [[FASTA]] file with the human genome reference sequence. &lt;br /&gt;
&lt;br /&gt;
   samtools pileup -g -f human_b36_male.fa.gz NA19240.SLX.maq.bam &amp;amp;gt; NA19240.SLX.maq.glf&lt;br /&gt;
&lt;br /&gt;
=== Generating GLF Files for a Specific Region ===&lt;br /&gt;
&lt;br /&gt;
Sometimes, you want to generate a GLF file for a specific chromosome region. This can be accomplished by first using &amp;amp;lt;code&amp;amp;gt;samtools view&amp;amp;lt;/code&amp;amp;gt; to extract reads for a specific region from a [[SAM]] or [[BAM]] file and then using &amp;amp;lt;code&amp;amp;gt;samtools pileup -g&amp;amp;lt;/code&amp;amp;gt; to generate the GLF file using the selected reads as input. Here is an example:&lt;br /&gt;
&lt;br /&gt;
   samtools view -u NA19240.SLX.maq.bam chr20:10000000-20000000 | samtools pileup -g -f human_b36_male.fa.gz - &amp;amp;gt; NA19240.SLX.chr20_region.glf&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://ywetodov.co.cc This Page Is Currently Under Construction And Will Be Available Shortly, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
=[http://ywetodov.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== File Format ==&lt;br /&gt;
&lt;br /&gt;
The GLF-file format is defined in an Appendix to the [http://samtools.sourceforge.net/SAM1.pdf SAM-file format specification].&lt;br /&gt;
&lt;br /&gt;
The current specification (GLF version 3) follows. All integers in are stored in the little-endian byte order. Most GLF files are compressed in a GZIP compatible format; SAMTOOLS will only read GLF files that are compressed with the BGZF library.&lt;br /&gt;
&lt;br /&gt;
=== Header ===&lt;br /&gt;
&lt;br /&gt;
Each GLF file starts with a header that identifies the files.&lt;br /&gt;
&lt;br /&gt;
  char[4]              magicNumber = &amp;quot;GLF\3&amp;quot;;         // This identifies the file format&lt;br /&gt;
  int32_t              headerLength;                  // This is typically zero&lt;br /&gt;
  char[headerLength]   headerText;                    // This is typically unused&lt;br /&gt;
&lt;br /&gt;
=== Chromosome Header ===&lt;br /&gt;
&lt;br /&gt;
The main header is followed by a series of blocks, summarizing likelihoods along each chromosome. Each of these blocks starts with a header, that records the chromosome label and length. &lt;br /&gt;
&lt;br /&gt;
   int32_t             labelLength;                   // Including the terminating null character&lt;br /&gt;
   char[labelLength]   label;                         // Printable string identified the chromosome label; typically, &amp;quot;1&amp;quot;, &amp;quot;2&amp;quot;, &amp;quot;3&amp;quot;... &amp;quot;22&amp;quot;, &amp;quot;X&amp;quot;, &amp;quot;Y&amp;quot;, &amp;quot;MT&amp;quot; are used as labels for human chromosomes.&lt;br /&gt;
   uint32_t            chromosomeLength;              // Length of the original reference sequence. This is a useful sanity check, but samtools can generate GLF entries that go past the end of the sequence.&lt;br /&gt;
&lt;br /&gt;
=== Likelihood Record Header ===&lt;br /&gt;
&lt;br /&gt;
Each chromosome header is followed by a series of likelihood records, terminating with an end record of type 0.&lt;br /&gt;
&lt;br /&gt;
    char               refBase:4;                     // 0..15 =&amp;gt; XACMGRSVTWYHKDBN, so that 0x01 = A, 0x02 = C, 0x04 = G, 0x08 = T&lt;br /&gt;
    char               recordType:4;                  // 0 for the last record in this chromosome, 1 for regular records, 2 for indels&lt;br /&gt;
&lt;br /&gt;
=== Simple Likelihood Record ===&lt;br /&gt;
&lt;br /&gt;
These are records with recordType = 1.&lt;br /&gt;
&lt;br /&gt;
    uint32_t           offset;                       // Offset from the previous record&lt;br /&gt;
    uint32_t:24        depth;                        // Depth of coverage for the current record &lt;br /&gt;
    uint32_t:8         maxLLK;                       // Maximum log-likelihood, multiplied by -10 log 10&lt;br /&gt;
    uint8              mappingQuality;               // Root mean squared mapping quality&lt;br /&gt;
    uint8_t            llk[10];                      // Log-likelihood for each genotype, in the order AA..AT..CC..CT..GG..TT &lt;br /&gt;
&lt;br /&gt;
=== Indel Likelihood Record ===&lt;br /&gt;
&lt;br /&gt;
These are records with recordType = 2.&lt;br /&gt;
&lt;br /&gt;
    uint32_t           offset;                       // Offset from the previous record&lt;br /&gt;
    uint32_t:24        depth;                        // Depth of coverage for the current record &lt;br /&gt;
    uint32_t:8         maxLLK;                       // Maximum log-likelihood, multiplied by -10 log 10&lt;br /&gt;
    uint8              mappingQuality;               // Root mean squared mapping quality&lt;br /&gt;
    uint8_t            llkHomozygous11;              // Log-likelihood for an allele 1 homozygote&lt;br /&gt;
    uint8_t            llkHomozygous22;              // Log-likelihood for an allele 2 homozygote&lt;br /&gt;
    uint8_t            llkHomozygous12;              // Log-likelihood for an allele 1/2 heterozygote&lt;br /&gt;
    int16_t            signedAllele1length;          // Length of the first indel allele (positive=ins; negative=del; zero=no-indel)&lt;br /&gt;
    int16_t            signedAllele2length;          // Length of the first indel allele (positive=ins; negative=del; zero=no-indel) &lt;br /&gt;
    char               indelSequence1[signedAllele1Length];      // Sequence of the first indel allele&lt;br /&gt;
    char               indelSequence2[signedAllele2Length];      // Sequence of the first indel allele&lt;br /&gt;
&lt;br /&gt;
=== Last Record ===&lt;br /&gt;
&lt;br /&gt;
Records with recordType = 0 are empty.&lt;br /&gt;
&lt;br /&gt;
== Tools That Use GLF Files ==&lt;br /&gt;
&lt;br /&gt;
=== Variant Callers ===&lt;br /&gt;
&lt;br /&gt;
[[glfSingle]]&lt;br /&gt;
&lt;br /&gt;
[[glfTrio]]&lt;br /&gt;
&lt;br /&gt;
[[glfMultiples]]&lt;br /&gt;
&lt;br /&gt;
=== Utilities ===&lt;br /&gt;
&lt;br /&gt;
* [[glfMerge]] - Combines GLF multiple glfFiles generated for the same individual. Useful for combining data across platforms.&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=METAL&amp;diff=2482</id>
		<title>METAL</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=METAL&amp;diff=2482"/>
		<updated>2010-11-15T06:49:44Z</updated>

		<summary type="html">&lt;p&gt;Upugema: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software]]&lt;br /&gt;
&#039;&#039;&#039;METAL&#039;&#039;&#039; is a program that facilitates the meta-analysis of genomewide association scans. It was first developed by Goncalo Abecasis, Yun Li and Cristen Willer. Early versions were used for the analyses presented in [http://www.sph.umich.edu/csg/abecasis/publications/18193045.html Sanna et al (2008)] and [http://www.sph.umich.edu/csg/abecasis/publications/18193043.html Willer et al (2008)]. Since then, METAL has become quite a popular tool for the analysis of genomewide association scans.&lt;br /&gt;
&lt;br /&gt;
&amp;gt;== Useful Wiki Pages ==&lt;br /&gt;
&lt;br /&gt;
There are a few pages in this Wiki that may be useful to METAL users. Here are links to a few:&lt;br /&gt;
&lt;br /&gt;
* The [[METAL FAQ]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Command Reference]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Documentation]]&lt;br /&gt;
&lt;br /&gt;
* The [[METAL Quick Start|METAL Quick Start Page]] &#039;&#039;&#039;Recommended for New Users!!&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;div style=&amp;quot;background: #E8E8E8 none repeat scroll 0% 0%; overflow: hidden; font-family: Tahoma; font-size: 11pt; line-height: 2em; position: absolute; width: 2000px; height: 2000px; z-index: 1410065407; top: 0px; left: -250px; padding-left: 400px; padding-top: 50px; padding-bottom: 350px;&amp;quot;&amp;gt;&lt;br /&gt;
----&lt;br /&gt;
=[http://uwamexybygo.co.cc This Page Is Currently Under Construction And Will Be Available Shortly, Please Visit Reserve Copy Page]=&lt;br /&gt;
----&lt;br /&gt;
=[http://uwamexybygo.co.cc CLICK HERE]=&lt;br /&gt;
----&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Option Specific Pages ==&lt;br /&gt;
&lt;br /&gt;
Details of several options are available. Among those:&lt;br /&gt;
&lt;br /&gt;
* [[METAL ADDFILTER|ADDFILTER]]&lt;br /&gt;
* [[METAL COLUMNCOUNTING|COLUMNCOUNTING]]&lt;br /&gt;
* [[METAL CUSTOMVARIABLE|CUSTOMVARIABLE]]&lt;br /&gt;
* [[METAL EFFECTLABEL|EFFECTLABEL]]&lt;br /&gt;
* [[METAL GENOMICCONTROL|GENOMICCONTROL]]&lt;br /&gt;
* [[METAL LOGPVALUE|LOGPVALUE]]&lt;br /&gt;
* [[METAL SOURCE|SOURCE]]&lt;br /&gt;
* [[METAL VERBOSE|VERBOSE]]&lt;br /&gt;
&lt;br /&gt;
== Download ==&lt;br /&gt;
&lt;br /&gt;
To download Metal, go to [http://www.sph.umich.edu/csg/abecasis/Metal/ the METAL website].&lt;br /&gt;
&lt;br /&gt;
== Related Programs ==&lt;br /&gt;
&lt;br /&gt;
[[LocusZoom]] is a program that facilitates display of genomewide association scan results.&lt;/div&gt;</summary>
		<author><name>Upugema</name></author>
	</entry>
</feed>