<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Kleckner</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Kleckner"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Kleckner"/>
	<updated>2026-09-24T22:17:40Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14321</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14321"/>
		<updated>2016-07-12T19:51:45Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory.&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14309</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14309"/>
		<updated>2016-07-06T16:28:47Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory.&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14308</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14308"/>
		<updated>2016-07-06T16:27:27Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14307</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14307"/>
		<updated>2016-07-06T16:27:02Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in ones file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14306</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14306"/>
		<updated>2016-07-06T16:25:58Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14305</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14305"/>
		<updated>2016-07-06T15:58:36Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14304</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14304"/>
		<updated>2016-07-06T15:55:16Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14303</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14303"/>
		<updated>2016-07-06T15:55:04Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* Sample Filtering&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14302</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14302"/>
		<updated>2016-07-06T15:44:56Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* What is Complete */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14301</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14301"/>
		<updated>2016-07-06T15:37:03Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14287</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14287"/>
		<updated>2016-06-27T16:42:48Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14286</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14286"/>
		<updated>2016-06-27T16:41:54Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3839 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14285</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14285"/>
		<updated>2016-06-27T16:41:15Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14284</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14284"/>
		<updated>2016-06-27T16:39:30Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14283</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14283"/>
		<updated>2016-06-27T16:39:18Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330 &lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14282</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14282"/>
		<updated>2016-06-27T16:05:40Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Sample Name Conversion&#039;&#039;&#039;&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: &lt;br /&gt;
***/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
** The Sardinia samples had a Sardinia ID and a Sequencing Core ID. The conversion key can be found here:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14281</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14281"/>
		<updated>2016-06-27T15:47:02Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Sample Name Conversion&#039;&#039;&#039;&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: &lt;br /&gt;
***/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
** The Sardinia samples had a Sardinia ID and a Sequencing Core ID. The conversion key can be found here:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14280</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14280"/>
		<updated>2016-06-27T15:46:48Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;IN PROCESS&#039;&#039;&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Sample Name Conversion&#039;&#039;&#039;&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: &lt;br /&gt;
***/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
** The Sardinia samples had a Sardinia ID and a Sequencing Core ID. The conversion key can be found here:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14279</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14279"/>
		<updated>2016-06-27T15:46:18Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Sample Name Conversion&#039;&#039;&#039;&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: &lt;br /&gt;
***/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
** The Sardinia samples had a Sardinia ID and a Sequencing Core ID. The conversion key can be found here:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14278</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14278"/>
		<updated>2016-06-27T15:45:05Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Sample Name Conversion&#039;&#039;&#039;&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14277</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14277"/>
		<updated>2016-06-27T15:44:05Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14276</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14276"/>
		<updated>2016-06-27T15:41:46Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS: &lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14275</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14275"/>
		<updated>2016-06-27T15:41:03Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged and sorted Indel and Beagled SNP VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED TOGETHER using Beagle4:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14274</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14274"/>
		<updated>2016-06-27T15:32:53Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14273</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14273"/>
		<updated>2016-06-27T15:32:36Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14272</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14272"/>
		<updated>2016-06-27T15:31:28Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14271</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14271"/>
		<updated>2016-06-27T15:31:00Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14270</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14270"/>
		<updated>2016-06-27T15:29:29Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14269</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14269"/>
		<updated>2016-06-27T15:29:07Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
*** Overall results from all of the below filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14268</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14268"/>
		<updated>2016-06-27T15:28:45Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**** Overall results from all of the below filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
**** VCFs of Indels after filtering can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14267</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14267"/>
		<updated>2016-06-27T15:00:45Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14266</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14266"/>
		<updated>2016-06-27T15:00:15Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**# (1) AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# (2) the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# (3) at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# (4) at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# (5) the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14265</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14265"/>
		<updated>2016-06-27T14:59:43Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
**# (1) AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# (2) the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# (3) at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# (4) at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# (5) the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14264</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14264"/>
		<updated>2016-06-27T14:58:31Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
### (1) AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
### (2) the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
### (3) at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
### (4) at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
### (5) the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14263</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14263"/>
		<updated>2016-06-27T14:57:18Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an Indel filtering strategy composed of many levels. &lt;br /&gt;
*** (1) AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
*** (2) the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
*** (3) at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
*** (4) at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
*** (5) the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
* Sample Name Conversion&lt;br /&gt;
** The CSCT samples had two different names.... a numeric and an alphanumeric. The conversion key can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/SampleNameChanges.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14262</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14262"/>
		<updated>2016-06-27T14:24:23Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
&lt;br /&gt;
* Pedigree (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* SNPCall Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* IndelCall Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14261</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14261"/>
		<updated>2016-06-27T14:04:52Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Locations of Files for Current Data Freeze of 3840 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22385)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze&lt;br /&gt;
** &lt;br /&gt;
&lt;br /&gt;
* QC Summary&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* mtDNA Copy Number Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples had &lt;br /&gt;
** By chromosome discrepancy figures can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** Overall discrepancy can be found:&lt;br /&gt;
*** IN PROCESS&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14260</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14260"/>
		<updated>2016-06-27T13:35:46Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3840 Samples===&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze&lt;br /&gt;
** &lt;br /&gt;
&lt;br /&gt;
* QC for all 3840 samples&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* Summary of Copy Number for 3840 samples&lt;br /&gt;
**/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** Individual copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14259</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14259"/>
		<updated>2016-06-27T13:31:33Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files ===&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* List of paths to BAMs used in this data freeze&lt;br /&gt;
** &lt;br /&gt;
&lt;br /&gt;
* QC for all 3840 samples&lt;br /&gt;
**&lt;br /&gt;
&lt;br /&gt;
* Summary of Copy Number for 3840 samples&lt;br /&gt;
**/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** Individual copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14258</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14258"/>
		<updated>2016-06-27T13:25:43Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files ===&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
* Recalibrated and deduced BAM Files (can be found in the following locations)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/Pula_Final/BAMS/by_sample/*.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/recal_20150210/*.recal.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/recal_20150330/*.recal.bam&lt;br /&gt;
** /net/mrtoad/stuff.from.csgspare/gpistis/Deduped_recalibrated_BAMS/*.uniq.dedup.recal.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/bams/decoy_ref/by_sample/*.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/RealigningCarlosSamples/Giorgio_Carlo_merged_20150412/*.CG.recal.merged.bam&lt;br /&gt;
*** some BAM files were generated by Carlo and Giorgio separately for the same individual. These BAMs were merged, recalibrated, and deduped and can be found here&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14257</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14257"/>
		<updated>2016-06-27T13:25:22Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Status as of June 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files ===&lt;br /&gt;
&lt;br /&gt;
* List of Sample Numbers&lt;br /&gt;
**&lt;br /&gt;
* Recalibrated and deduced BAM Files (can be found in the following locations)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/Pula_Final/BAMS/by_sample/*.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/recal_20150210/*.recal.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/recal_20150330/*.recal.bam&lt;br /&gt;
** /net/mrtoad/stuff.from.csgspare/gpistis/Deduped_recalibrated_BAMS/*.uniq.dedup.recal.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/bams/decoy_ref/by_sample/*.bam&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/RealigningCarlosSamples/Giorgio_Carlo_merged_20150412/*.CG.recal.merged.bam (some BAM files were generated by Carlo and Giorgio separately for the same individual. These BAMs were merged, recalibrated, and deduped and can be found here)&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14256</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14256"/>
		<updated>2016-06-27T13:19:39Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of June 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files ===&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14255</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14255"/>
		<updated>2016-06-27T13:19:06Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
=== Status as of June 2016 ===&lt;br /&gt;
&lt;br /&gt;
== Locations of Files ==&lt;br /&gt;
&lt;br /&gt;
== What is Complete ==&lt;br /&gt;
&lt;br /&gt;
== Future Directions ==&lt;br /&gt;
&lt;br /&gt;
== Key References == &lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SeqShop:_Variant_Calling_and_Filtering_for_SNPs_Practical,_December_2014&amp;diff=13080</id>
		<title>SeqShop: Variant Calling and Filtering for SNPs Practical, December 2014</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SeqShop:_Variant_Calling_and_Filtering_for_SNPs_Practical,_December_2014&amp;diff=13080"/>
		<updated>2015-03-30T23:12:57Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Sequnce Alignment Files: BAM Files */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;==Introduction==&lt;br /&gt;
Main Workshop wiki page: [[SeqShop: December 2014]]&lt;br /&gt;
&lt;br /&gt;
See the [[Media:Dec2014 SeqShop - GotCloud snpcall.pdf|introductory slides]] for an intro to this tutorial.&lt;br /&gt;
&lt;br /&gt;
== Goals of This Session ==&lt;br /&gt;
* What we want to learn &lt;br /&gt;
** How to generate filtered variant calls for SNPs from BAMs&lt;br /&gt;
** Basic variant call file format (VCF)&lt;br /&gt;
** How to examine the variants at particular genomic positions&lt;br /&gt;
** How to evaluate the quality of SNP calls&lt;br /&gt;
&lt;br /&gt;
== Setup in person at the SeqShop Workshop ==&lt;br /&gt;
&#039;&#039;This section is specifically for the SeqShop Workshop computers.&#039;&#039;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:600px&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;If you are not running during the SeqShop Workshop, please skip this section.&#039;&#039;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
{{SeqShopLogin}}&lt;br /&gt;
&lt;br /&gt;
=== Setup your run environment===&lt;br /&gt;
This is the same setup you did for the previous tutorial, but you need to redo it each time you log in.&lt;br /&gt;
&lt;br /&gt;
This will setup some environment variables to point you to&lt;br /&gt;
* [[GotCloud]] program&lt;br /&gt;
* Tutorial input files&lt;br /&gt;
* Setup an output directory&lt;br /&gt;
** It will leave your output directory from the previous tutorial in tact.&lt;br /&gt;
 source /net/seqshop-server/home/mktrost/seqshop/setup.txt&lt;br /&gt;
* You won&#039;t see any output after running &amp;lt;code&amp;gt;source&amp;lt;/code&amp;gt;&lt;br /&gt;
** It silently sets up your environment&lt;br /&gt;
** If you want to view the detail of the setup, type&lt;br /&gt;
 less /net/seqshop-server/home/mktrost/seqshop/setup.txt&lt;br /&gt;
and press &#039;q&#039; to finish.&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:200px&amp;quot;&amp;gt;&lt;br /&gt;
View setup.txt&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:setup.png|500px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Setup when running on your own outside of the SeqShop Workshop ==&lt;br /&gt;
&#039;&#039;This section is specifically for running on your own outside of the SeqShop Workshop.&#039;&#039;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible&amp;quot; style=&amp;quot;width:600px&amp;quot;&amp;gt;&lt;br /&gt;
&#039;&#039;If you are running during the SeqShop Workshop, please skip this section.&#039;&#039;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&lt;br /&gt;
This tutorial builds on the alignment tutorial, if you have not already, please first run that tutorial: [[SeqShop:_Sequence_Mapping_and_Assembly_Practical|Alignment Tutorial]]&lt;br /&gt;
&lt;br /&gt;
{{SeqShopRemoteEnv}}&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Examining GotCloud SnpCall Input files ==&lt;br /&gt;
=== Sequence Alignment Files: BAM Files ===&lt;br /&gt;
Per sample BAM files contain sequence reads that are mapped to positions in the genome.&lt;br /&gt;
&lt;br /&gt;
For a reminder on how to look at/read BAM files, see: [[SeqShop:_Sequence_Mapping_and_Assembly_Practical#BAM_Files|SeqShop Aligment: BAM Files]]&lt;br /&gt;
&lt;br /&gt;
For this tutorial, we will use the 4 BAMs produced in the [[SeqShop: Sequence Mapping and Assembly Practical]] as well as with 58 BAMs that were pre-aligned to that 1MB region of chromosome 22.&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
Reference files can be downloaded with GotCloud or from other sources.&lt;br /&gt;
* For this practical, I already downloaded them for you.&lt;br /&gt;
* See [[GotCloud: Genetic Reference and Resource Files]] for more information on downloading/generating reference files&lt;br /&gt;
&lt;br /&gt;
For GotCloud snpcall, you need:&lt;br /&gt;
# Reference genome FASTA file&lt;br /&gt;
#* Contains the reference base for each position of each chromosome&lt;br /&gt;
#** Used to compare bases in sequence reads to the reference positions they mapped to&lt;br /&gt;
#** Used to identify SNPs/variations in the sequence reads&lt;br /&gt;
#* Additional information on the FASTA format: http://en.wikipedia.org/wiki/FASTA_format&lt;br /&gt;
# VCF (variant call format) files with chromosomes/positions&lt;br /&gt;
#* indel - contains known insertions &amp;amp; deletions to help with filtering&lt;br /&gt;
#* omni - used as likely true positives for SVM filtering&lt;br /&gt;
#* hapmap - used as likely true positives for SVM filtering and for generating summary statistics&lt;br /&gt;
#* dbsnp - used for generating summary statistics&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
We looked at them yesterday, but you can take another look at the chromosome 22 reference files included for this tutorial:&lt;br /&gt;
 ls ${SS}/ref22&lt;br /&gt;
&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:200px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;View Screenshot&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:RefDir.png|700px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== GotCloud BAM List File ===&lt;br /&gt;
The [[GotCloud:_Variant_Calling_Pipeline#BAM_List_File|BAM list file]] points GotCloud to the BAM files&lt;br /&gt;
* generated by the alignment pipeline&lt;br /&gt;
&lt;br /&gt;
Look at the BAM list file the alignment pipeline generated&lt;br /&gt;
 cat ${OUT}/bam.list&lt;br /&gt;
&lt;br /&gt;
;What is the path to the BAM file for sample HG00640?&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:200px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Answer:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;/net/seqshop-server/home/YourUserName/out/bams/HG00640.recal.bam&amp;lt;/li&amp;gt;&lt;br /&gt;
[[File:BamindexNew.png|500px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
The alignment pipeline only processed 4 samples, but for snpcall, we want to run on 62 samples.&lt;br /&gt;
* The other 58 samples were already aligned:&lt;br /&gt;
 ls ${SS}/bams&lt;br /&gt;
&lt;br /&gt;
Look at the BAM list file for those BAMs:&lt;br /&gt;
 less ${SS}/bams/bam.list&lt;br /&gt;
&lt;br /&gt;
Remember, use &amp;lt;code&amp;gt;&#039;q&#039;&amp;lt;/code&amp;gt; to exit out of &amp;lt;code&amp;gt;less&amp;lt;/code&amp;gt;&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
;Do you notice a difference between this list and yours?&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:550px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Answer:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;It doesn&#039;t have a full path to the BAM file, while your list has /home/...&amp;lt;/li&amp;gt;&lt;br /&gt;
[[File:BamList1.png|300px]]&lt;br /&gt;
&amp;lt;li&amp;gt;That&#039;s ok, we will use the &amp;lt;code&amp;gt;--base_prefix ${SS}&amp;lt;/code&amp;gt; command-line option to prefix the BAM paths&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Alternatively, we could have set BAM_PREFIX in &amp;lt;code&amp;gt;gotcloud.conf&amp;lt;/code&amp;gt; to the path to the BAMs&lt;br /&gt;
&amp;lt;pre&amp;gt;BAM_PREFIX = /net/seqshop-server/home/mktrost/seqshop/example&amp;lt;/pre&amp;gt; &amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;NOTE: the conf file can&#039;t interpret ${SS} environment variables or &#039;~&#039;, so you would have to specify the full path&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;We just used the command-line option for this tutorial since this path will vary by user when running outside the workshop.&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
We need to add these BAMs to our list&lt;br /&gt;
* Append the bam.list from the pre-aligned BAMs to the one you generated from the alignment pipeline&lt;br /&gt;
** &#039;&#039;&#039;Be sure to do this command just once&#039;&#039;&#039;&lt;br /&gt;
 cat ${SS}/bams/bam.list &amp;gt;&amp;gt; ${OUT}/bam.list&lt;br /&gt;
* &amp;quot;&amp;gt;&amp;gt;&amp;quot; will append to the file that follows it&lt;br /&gt;
** Check that your BAM list is the correct size&lt;br /&gt;
**:&amp;lt;pre&amp;gt;wc -l ${OUT}/bam.list&amp;lt;/pre&amp;gt;&lt;br /&gt;
*** &amp;lt;code&amp;gt;wc -l&amp;lt;/code&amp;gt; counts the number of lines in the file&lt;br /&gt;
*** Should be 62&lt;br /&gt;
&lt;br /&gt;
Verify your BAM list contains the additional BAMs&lt;br /&gt;
  less ${OUT}/bam.list&lt;br /&gt;
&lt;br /&gt;
Remember, use &amp;lt;code&amp;gt;&#039;q&#039;&amp;lt;/code&amp;gt; to exit out of &amp;lt;code&amp;gt;less&amp;lt;/code&amp;gt;&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
;Do you see both sets of BAMs?&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:500px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Annotated Screenshot:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;If not, let me know&amp;lt;/li&amp;gt;&lt;br /&gt;
[[File:BamList2.png|400px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== GotCloud Configuration File ===&lt;br /&gt;
We will use the same configuration file as we used yesterday in GotCloud Align.&lt;br /&gt;
&lt;br /&gt;
See [[SeqShop:_Sequence_Mapping_and_Assembly_Practical#GotCloud Configuration File|SeqShop: Alignment: GotCloud Configuration File]] for more details&lt;br /&gt;
* Note we want to limit snpcall to just chr22 so the configuration already has &amp;lt;code&amp;gt;CHRS = 22&amp;lt;/code&amp;gt; (default was 1-22 &amp;amp; X).&lt;br /&gt;
&lt;br /&gt;
For more information on configuration, see: [[GotCloud:_Variant_Calling_Pipeline#Configuration_File|GotCloud snpcall: Configuration File]]&lt;br /&gt;
* Contains information on how to configure for exome/targeted sequencing&lt;br /&gt;
&lt;br /&gt;
== Run GotCloud SnpCall ==&lt;br /&gt;
[[File:SnpcallDiagramNew.png|500px]]&lt;br /&gt;
&lt;br /&gt;
Now that we have all of our input files, we need just a simple command to run:&lt;br /&gt;
* When running at home if you don&#039;t have 6 CPUs, reduce the --numjobs setting (it will take longer to run).&lt;br /&gt;
 ${GC}/gotcloud snpcall --conf ${SS}/gotcloud.conf --numjobs 6 --region 22:36000000-37000000 --base_prefix ${SS} --outdir ${OUT}&lt;br /&gt;
* &amp;lt;code&amp;gt;${GC}/gotcloud&amp;lt;/code&amp;gt; runs GotCloud&lt;br /&gt;
* &amp;lt;code&amp;gt;snpcall&amp;lt;/code&amp;gt; tells GotCloud you want to run the snpcall pipeline.&lt;br /&gt;
* &amp;lt;code&amp;gt;--conf&amp;lt;/code&amp;gt; tells GotCloud the name of the configuration file to use.&lt;br /&gt;
** The configuration for this test was downloaded with the seqshop input files.&lt;br /&gt;
* --numjobs tells GotCloud how many jobs to run in parallel&lt;br /&gt;
** Depends on your system&lt;br /&gt;
* --region 22:36000000-37000000&lt;br /&gt;
** The sample files are just a small region of chromosome 22, so to save time, we tell GotCloud to ignore the other regions&lt;br /&gt;
* &amp;lt;code&amp;gt;--base_prefix&amp;lt;/code&amp;gt; tells GotCloud the prefix to append to relative paths.&lt;br /&gt;
** The Configuration file cannot read environment variables, so we need to tell GotCloud the path to the input files, ${SS}&lt;br /&gt;
** Alternatively, gotcloud.conf could be updated to specify the full paths&lt;br /&gt;
* &amp;lt;code&amp;gt;--outdir&amp;lt;/code&amp;gt; tells GotCloud where to write the output.&lt;br /&gt;
** This could be specified in gotcloud.conf, but to allow you to use the ${OUT} to change the output location, it is specified on the command-line&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:500px&amp;quot;&amp;gt;&lt;br /&gt;
Curious if it started running properly?  Check out this screenshot:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:SnpcallStartNew.png|550px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
This should take about 5-8 minutes to run.&lt;br /&gt;
* It should end with a line like: &amp;lt;code&amp;gt;Commands finished in 402 secs with no errors reported&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
If you cancelled GotCloud part way through, just rerun your GotCloud command and it will pick up where it left off.&lt;br /&gt;
&lt;br /&gt;
If you want to understand more detailed step of GotCloud SNP calling, here is a schematic picture with a little bit more details&lt;br /&gt;
&lt;br /&gt;
[[File:Gotcloudoverview.png|600px]]&lt;br /&gt;
&lt;br /&gt;
== Examining GotCloud SnpCall Output ==&lt;br /&gt;
Let&#039;s look at the output directory:&lt;br /&gt;
 ls ${OUT}&lt;br /&gt;
&lt;br /&gt;
;Do you see any new files or directories?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:350px&amp;quot;&amp;gt;&lt;br /&gt;
* View Annotated Screenshot:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:gcsnpcallOutNew.png|500px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let&#039;s look at the vcfs directory:&lt;br /&gt;
 ls ${OUT}/vcfs&lt;br /&gt;
Just a &amp;lt;code&amp;gt;chr22&amp;lt;/code&amp;gt; directory, so look inside of there:&lt;br /&gt;
 ls ${OUT}/vcfs/chr22&lt;br /&gt;
;Can you identify the final filtered VCF and the associated summary file?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:350px&amp;quot;&amp;gt;&lt;br /&gt;
* Answer &amp;amp; annotated directory listing:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Filtered VCF (SVM &amp;amp; hard filters): chr22.filtered.vcf.gz&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Summary file: chr22.filtered.sites.vcf.summary&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
[[File:vcfsout.png|600px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Now, let&#039;s look in the split directory for the VCF with just the passing variants:&lt;br /&gt;
 ls ${OUT}/split/chr22&lt;br /&gt;
;Which file do you think is the one you want?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:350px&amp;quot;&amp;gt;&lt;br /&gt;
* Answer:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;chr22.filtered.PASS.vcf.gz&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
[[File:splitOut.png|600px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Filtering Summary Statistics ===&lt;br /&gt;
&lt;br /&gt;
 cat ${OUT}/vcfs/chr22/chr22.filtered.sites.vcf.summary&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:250px&amp;quot;&amp;gt;&lt;br /&gt;
View Screenshot&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:filterSumNew.png|700px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;To understand how to interpret the filtering summary statistics, please refer to [[Understanding vcf-summary output]]&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
=== Filtered VCF ===&lt;br /&gt;
&lt;br /&gt;
Let&#039;s look at the filtered sites file.&lt;br /&gt;
 less -S ${OUT}/vcfs/chr22/chr22.filtered.sites.vcf&lt;br /&gt;
&lt;br /&gt;
* Scroll down until you find some variants.&lt;br /&gt;
** Use space bar to jump a full page&lt;br /&gt;
** Use down arrow to move down one line&lt;br /&gt;
* Scroll right: lots of info fields, but no per sample genotype information&lt;br /&gt;
&lt;br /&gt;
;What is the first filtered out variant that you find &amp;amp; what filter did it fail?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:250px&amp;quot;&amp;gt;&lt;br /&gt;
* Answer:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;It failed SVM filter&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
[[File:SvmFiltNew.png|550px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Remember, use &amp;lt;code&amp;gt;&#039;q&#039;&amp;lt;/code&amp;gt; to exit out of &amp;lt;code&amp;gt;less&amp;lt;/code&amp;gt;&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Now, let&#039;s look at the filtered file with genotypes.&lt;br /&gt;
 zless -S ${OUT}/vcfs/chr22/chr22.filtered.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Scroll down until you find some variants.&lt;br /&gt;
** Use space bar to jump a full page&lt;br /&gt;
** Use down arrow to move down one line&lt;br /&gt;
* Scroll right until you should see per sample genotype information&lt;br /&gt;
&lt;br /&gt;
Remember, use &amp;lt;code&amp;gt;&#039;q&#039;&amp;lt;/code&amp;gt; to exit out of &amp;lt;code&amp;gt;less&amp;lt;/code&amp;gt;&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:250px&amp;quot;&amp;gt;&lt;br /&gt;
* View annotated screenshot:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:SvmFiltGLNew.png|550px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Passing SNPs ===&lt;br /&gt;
&lt;br /&gt;
Let&#039;s look at the file of just the pass sites:&lt;br /&gt;
 zless -S ${OUT}/split/chr22/chr22.filtered.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Scroll down: they all look like they &amp;lt;code&amp;gt;PASS&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Remember, use &#039;q&#039; to exit out of less&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
Let&#039;s check if they are all PASS.&lt;br /&gt;
 zcat ${OUT}/split/chr22/chr22.filtered.PASS.vcf.gz |grep -v &amp;quot;^#&amp;quot;| cut -f 7| grep -v &amp;quot;PASS&amp;quot;&lt;br /&gt;
It will return nothing since there are no non-passing variants in this file.&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:450px&amp;quot;&amp;gt;&lt;br /&gt;
;Want an explanation of this command?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
* zcat ...: uncompress the zipped VCF&lt;br /&gt;
* &#039;|&#039; : this takes the output of one command and sends it as input to the next&lt;br /&gt;
* grep -v &amp;quot;^#&amp;quot; : exclude any lines that start with &amp;quot;#&amp;quot; - headers&lt;br /&gt;
* cut -f 7 : extract the FILTER column (the 7th column)&lt;br /&gt;
* grep -v &amp;quot;PASS&amp;quot; : exclude any rows that have a &amp;quot;PASS&amp;quot; in the FILTER column&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Compare that to the filtered file we looked at before:&lt;br /&gt;
 zcat ${OUT}/vcfs/chr22/chr22.filtered.vcf.gz |grep -v &amp;quot;^#&amp;quot;| cut -f 7| grep -v &amp;quot;PASS&amp;quot;&lt;br /&gt;
;Do you see any filters?&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:450px&amp;quot;&amp;gt;&lt;br /&gt;
*Answer&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
* Yes&lt;br /&gt;
** It should have scrolled and you should see filters like:&lt;br /&gt;
*** INDEL5;SVM&lt;br /&gt;
*** INDEL5&lt;br /&gt;
*** SVM&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== GotCloud Genotype Refinement ==&lt;br /&gt;
To improve the quality of the genotypes, we run a genotype refinement pipeline.&lt;br /&gt;
&lt;br /&gt;
This pipeline runs [http://faculty.washington.edu/browning/beagle/beagle.html Beagle] &amp;amp; thunder.&lt;br /&gt;
&lt;br /&gt;
=== Genotype Refinement Input ===&lt;br /&gt;
The GotCloud genotype refinement pipeline takes as input ${OUT}/split/chr22/chr22.filtered.PASS.vcf.gz (the VCF file of PASS&#039;ing SNPs from snpcall).&lt;br /&gt;
&lt;br /&gt;
The bam list and the configuration file we used for GotCloud snpcall will tell GotCloud genotype refinement everything it needs to know, so no new input files need to be prepared.&lt;br /&gt;
&lt;br /&gt;
Note: the configuration file overrides the THUNDER command to make it go faster than the default settings so the tutorial will run faster:&lt;br /&gt;
[[File:thunderConf.png|600px]]&lt;br /&gt;
&lt;br /&gt;
=== Running GotCloud Genotype Refinement ===&lt;br /&gt;
Since everything is setup, just run the following command (very similar to snpcall).&lt;br /&gt;
 ${GC}/gotcloud ldrefine --conf ${SS}/gotcloud.conf --numjobs 6 --region 22:36000000-37000000 --base_prefix ${SS} --outdir ${OUT}&lt;br /&gt;
&lt;br /&gt;
* Beagle will take about 1-3 minutes to complete&lt;br /&gt;
* Thunder will automatically run and will take another 2-4 minutes&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:350px&amp;quot;&amp;gt;&lt;br /&gt;
When completed, it should look like this:&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:GcldrefineOutNew.png]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Genotype Refinement Output ===&lt;br /&gt;
&lt;br /&gt;
; What&#039;s new in the output directory?&lt;br /&gt;
&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:500px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Answer&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
:&amp;lt;pre&amp;gt;ls ${OUT}&amp;lt;/pre&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;beagle&amp;lt;/code&amp;gt; directory : Beagle output&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;thunder&amp;lt;/code&amp;gt; directory : Thunder output&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;umake.beagle.*&amp;lt;/code&amp;gt; : Contain the configuration &amp;amp; steps used in GotCloud beagle&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;umake.thunder.*&amp;lt;/code&amp;gt; : Contain the configuration &amp;amp; steps used in GotCloud thunder&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Let&#039;s take a look at that interesting location we found in the [[SeqShop:_Sequence_Mapping_and_Assembly_Practical#Accessing_BAMs_by_Position|alignment tutorial]] : chromosome 22, positions 36907000-36907100&lt;br /&gt;
&lt;br /&gt;
Use tabix to extract that from the VCFs:&lt;br /&gt;
 ${GC}/bin/tabix ${OUT}/thunder/chr22/ALL/thunder/chr22.filtered.PASS.beagled.ALL.thunder.vcf.gz 22:36907000-36907100 |less -S&lt;br /&gt;
&lt;br /&gt;
Remember, type &#039;q&#039; to quit less.&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
;Are there any variants in this region?&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:500px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Answer:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Yes!&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Positions:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;36907001&amp;lt;/code&amp;gt;; Ref: T, Alt: C - that&#039;s what we saw before&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;36907098&amp;lt;/code&amp;gt;; Ref: T, Alt: C - that&#039;s what we saw before&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
;What is HG00551&#039;s genotype at these positions?&lt;br /&gt;
#First check which sample number HG00551 is:&lt;br /&gt;
 zcat ${OUT}/thunder/chr22/ALL/thunder/chr22.filtered.PASS.beagled.ALL.thunder.vcf.gz |grep &amp;quot;#CHROM&amp;quot;&lt;br /&gt;
* That will help you figure out it&#039;s genotype.&lt;br /&gt;
* Rerun the tabix command and scroll to find HG00551&#039;s genotype:&lt;br /&gt;
 ${GC}/bin/tabix ${OUT}/thunder/chr22/ALL/thunder/chr22.filtered.PASS.beagled.ALL.thunder.vcf.gz 22:36907000-36907100 |less -S&lt;br /&gt;
&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:500px&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;Answer:&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
&amp;lt;ul&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;It is the first sample&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;0|1&amp;lt;/code&amp;gt;: Heterozygous (although low GQ - quality)&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;li&amp;gt;&amp;lt;code&amp;gt;1|1&amp;lt;/code&amp;gt;; Homozygous Alt (C)&amp;lt;/li&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/ul&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Remember, type &#039;q&#039; to quit less.&lt;br /&gt;
 q&lt;br /&gt;
&lt;br /&gt;
=== Did I find interesting variants? ===&lt;br /&gt;
&lt;br /&gt;
The region we selected contains &#039;&#039;APOL1&#039;&#039; gene, which is known to play an important role in kidney diseases such as nephrotic syndrome. One of the non-synonymous risk allele, &amp;lt;code&amp;gt;rs73885139&amp;lt;/code&amp;gt; located at position &amp;lt;code&amp;gt;22:36661906&amp;lt;/code&amp;gt; increases the risk of nephrotic syndrome by &amp;gt;2-folds. Let&#039;s see if we found the interesting variant by looking at the VCF file by position.&lt;br /&gt;
&lt;br /&gt;
 ${GC}/bin/tabix ${OUT}/vcfs/chr22/chr22.filtered.vcf.gz 22:36661906 | head -1 &lt;br /&gt;
&lt;br /&gt;
Did you see a variant at the position?&lt;br /&gt;
&lt;br /&gt;
 22	36661906	.	A	G	23	PASS	DP=409;MQ=59;NS=62;AN=124;AC=2;AF=0.013847;&lt;br /&gt;
 AB=0.4169;AZ=-0.3525;FIC=-0.0089;SLRT=-0.0071;LBS=36,36,0,0,1,1,0,0;OBS=145,191,0,0,3,2,0,0;STR=-0.040;&lt;br /&gt;
 STZ=-0.740;CBR=0.008;CBZ=0.144;IOR=0.000;IOZ=-1.370;AOI=-5.614;AOZ=-4.243;LQR=0.178;MQ0=0.000;&lt;br /&gt;
 MQ10=0.000;MQ20=0.000;MQ30=0.000;SVM=1.45191	GT:DP:GQ:PL	0/0:4:28:0,12,65	&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Let&#039;s check the sequence data to confirm that the variant really exists&lt;br /&gt;
&lt;br /&gt;
 ${GC}/bin/samtools tview ${SS}/bams/HG01242.recal.bam ${SS}/ref22/human.g1k.v37.chr22.fa&lt;br /&gt;
&lt;br /&gt;
* Type &#039;g&#039; to go to a specific position&lt;br /&gt;
* Type 22:36661906 to move to the position&lt;br /&gt;
* Press arrows to move between positions&lt;br /&gt;
* Press &#039;b&#039; if you want to color by base quality&lt;br /&gt;
* Press &#039;?&#039; for more help&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:250px&amp;quot;&amp;gt;&lt;br /&gt;
View Screenshot&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
[[File:SamtoolstviewsnpNew.png|600px]]&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
=== Improvements ===&lt;br /&gt;
&lt;br /&gt;
Let&#039;s get some information on the BEAGLE VCF:&lt;br /&gt;
&lt;br /&gt;
 perl ${GC}/scripts/bed-diff.pl --vcf1 ${SS}/ref22/1kg.omni.chr22.36Mb.vcf.gz --vcf2 ${OUT}/beagle/chr22/chr22.filtered.PASS.beagled.ALL.vcf.gz --out ${OUT}/diffs/bedDiff.beagle&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
Look at the results:&lt;br /&gt;
 more ${OUT}/diffs/bedDiff.beagle.summary&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:400px&amp;quot;&amp;gt;&lt;br /&gt;
*Results&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
 OVERALL:	43588	44293	0.9841&lt;br /&gt;
 NREF-EITHER:	19644	20349	0.9654&lt;br /&gt;
 NMAJ-EITHER:	14560	15265	0.9538&lt;br /&gt;
&lt;br /&gt;
 HOMREF:	23944	91	0	0.9962&lt;br /&gt;
 HET:	355	11936	172	0.9577&lt;br /&gt;
 HOMALT:	6	81	7708	0.9888&lt;br /&gt;
&lt;br /&gt;
 HOMMAJ:	29028	112	4	0.9960&lt;br /&gt;
 HET:	380	11936	147	0.9577&lt;br /&gt;
 HOMMIN:	2	60	2624	0.9769&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
Now, let&#039;s see if it improved after running Thunder VCF:&lt;br /&gt;
 perl ${GC}/scripts/bed-diff.pl --vcf1 ${SS}/ref22/1kg.omni.chr22.36Mb.vcf.gz --vcf2 ${OUT}/thunder/chr22/ALL/thunder/chr22.filtered.PASS.beagled.ALL.thunder.vcf.gz --out ${OUT}/diffs/bedDiff.thunder&lt;br /&gt;
&lt;br /&gt;
Look at the results:&lt;br /&gt;
 more ${OUT}/diffs/bedDiff.thunder.summary&lt;br /&gt;
&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible mw-collapsed&amp;quot; style=&amp;quot;width:400px&amp;quot;&amp;gt;&lt;br /&gt;
*Results&lt;br /&gt;
&amp;lt;div class=&amp;quot;mw-collapsible-content&amp;quot;&amp;gt;&lt;br /&gt;
 OVERALL:	43711	44293	0.9869&lt;br /&gt;
 NREF-EITHER:	19777	20359	0.9714&lt;br /&gt;
 NMAJ-EITHER:	14715	15297	0.9620&lt;br /&gt;
&lt;br /&gt;
 HOMREF:	23934	101	0	0.9958&lt;br /&gt;
 HET:	272	12084	107	0.9696&lt;br /&gt;
 HOMALT:	5	97	7693	0.9869&lt;br /&gt;
&lt;br /&gt;
 HOMMAJ:	28996	145	3	0.9949&lt;br /&gt;
 HET:	272	12084	107	0.9696&lt;br /&gt;
 HOMMIN:	2	53	2631	0.9795&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&amp;lt;/div&amp;gt;&lt;br /&gt;
&lt;br /&gt;
There is an improvement.&lt;br /&gt;
&lt;br /&gt;
== What is GotCloud snpcall doing? ==&lt;br /&gt;
To run GotCloud, you really just needed a single command.&lt;br /&gt;
&lt;br /&gt;
Well, that one command runs many steps.  Here is a diagram of all the steps.&lt;br /&gt;
&lt;br /&gt;
[[File:SnpcallPipeline.jpg|1200px]]&lt;br /&gt;
 &lt;br /&gt;
Aren&#039;t you glad you didn&#039;t have to configure &amp;amp; run each one yourself?&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Return to Workshop Wiki Page ==&lt;br /&gt;
Return to main workshop wiki page: [[SeqShop: December 2014]]&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Variant_Calling_Pipeline&amp;diff=13079</id>
		<title>GotCloud: Variant Calling Pipeline</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Variant_Calling_Pipeline&amp;diff=13079"/>
		<updated>2015-03-30T22:54:59Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Running the Automatic Test */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]]&lt;br /&gt;
&lt;br /&gt;
The Variant Calling Pipeline (previously called &#039;UMAKE&#039;) makes genotype calls from recalibrated BAM files. These genotype calls are output into [http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41 VCF (Variant Call Format) files].&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Running the GotCloud Variant Calling Pipeline ==&lt;br /&gt;
&lt;br /&gt;
The variant calling pipeline (umake) is run using &amp;lt;code&amp;gt;gotcloud snpcall&amp;lt;/code&amp;gt; and &amp;lt;code&amp;gt;gotcloud ldrefine&amp;lt;/code&amp;gt;. &lt;br /&gt;
&lt;br /&gt;
===Running the Automatic Test===&lt;br /&gt;
&lt;br /&gt;
The automatic test runs the variant calling pipeline on a small test set and checks the results against expected results validating that GotCloud is installed correctly.&lt;br /&gt;
&lt;br /&gt;
*Run &amp;lt;code&amp;gt;snpcall&amp;lt;/code&amp;gt; pipeline test:&lt;br /&gt;
 gotcloud snpcall --test OUTPUT_DIR&lt;br /&gt;
** Where OUTPUT_DIR is the directory where you want to store the test results&lt;br /&gt;
** If you see &amp;lt;code&amp;gt;Successfully ran the test case, congratulations!&amp;lt;/code&amp;gt;, then you are ready to run snpcall on your own samples.&lt;br /&gt;
*Run &amp;lt;code&amp;gt;ldrefine&amp;lt;/code&amp;gt; pipeline test:&lt;br /&gt;
 gotcloud ldrefine --test OUTPUT_DIR&lt;br /&gt;
** Where &amp;lt;code&amp;gt;OUTPUT_DIR&amp;lt;/code&amp;gt; is the directory where you want to store the test results&lt;br /&gt;
** If you see &amp;lt;code&amp;gt;Successfully ran the test case, congratulations!&amp;lt;/code&amp;gt;, then you are ready to run ldrefine on your own samples.&lt;br /&gt;
&lt;br /&gt;
== Overview of Variant Calling Pipeline Steps ==&lt;br /&gt;
Here is an overview of the Variant Calling Pipeline:&lt;br /&gt;
&lt;br /&gt;
[[File: umakeSteps.png]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
For more information on the filters applied during the Variant Calling Pipeline, see, [[GotCloud: Filters]].&lt;br /&gt;
&lt;br /&gt;
== Input Data==&lt;br /&gt;
* [[#BAM Files|Aligned/Processed/Recalibrated BAM files]]&lt;br /&gt;
* [[#BAM List File|BAM list file containing Sample IDs &amp;amp; BAM file names]]&lt;br /&gt;
* [[#Reference Files|Reference files]]&lt;br /&gt;
* (Optional) [[#Configuration File|Configuration file to override default options]]&lt;br /&gt;
&lt;br /&gt;
=== BAM Files ===&lt;br /&gt;
The BAM files need to be duplicate-marked and base-quality recalibrated in order to obtain high quality SNP calls. Generating these BAM files from original FASTQs is automatically done as part of the [[Alignment Pipeline]] of GotCloud.&lt;br /&gt;
&lt;br /&gt;
=== BAM List File ===&lt;br /&gt;
* Automatically created when running the GotCloud [[Alignment Pipeline]]&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
The path to the BAM List file is defaulted to the &amp;lt;code&amp;gt;outputDirectory/bam.list&amp;lt;/code&amp;gt;.  It can be overridden by setting &amp;lt;code&amp;gt;--bamlist&amp;lt;/code&amp;gt;, &amp;lt;code&amp;gt;--bam_list&amp;lt;/code&amp;gt;, or &amp;lt;code&amp;gt;--list&amp;lt;/code&amp;gt; on the command-line or by setting BAM_LIST in your configuration file to the path to the BAM List File.  See [[#Required_Options|Required Options]] for more information.&lt;br /&gt;
&lt;br /&gt;
=== Reference Files ===&lt;br /&gt;
See [[GotCloud: Genetic Reference and Resource Files]] for detailed information about the multiple required reference files for the variant calling pipeline, including:&lt;br /&gt;
* How to obtain default references&lt;br /&gt;
* Configuration keys &amp;amp; default values&lt;br /&gt;
* How to generate your own references&lt;br /&gt;
* How to point GotCloud to your reference files&lt;br /&gt;
&lt;br /&gt;
Required Reference File Types:&lt;br /&gt;
* [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]]&lt;br /&gt;
* [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF Files|DBSNP VCF Files]]&lt;br /&gt;
* [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF Files|HapMap3 VCF Files]]&lt;br /&gt;
* [[GotCloud: Genetic Reference and Resource Files#OMNI VCF Files|OMNI VCF Files]]&lt;br /&gt;
* [[GotCloud: Genetic Reference and Resource Files#INDEL VCF File(s)|INDEL VCF File(s)]]&lt;br /&gt;
&lt;br /&gt;
=== Configuration File ===&lt;br /&gt;
{{:GotCloud: Configuration}}&lt;br /&gt;
&lt;br /&gt;
See [[#Variant Calling Command-line Options/Configuration Settings|Variant Calling Command-line Options/Configuration Settings]] for more information on Configuration options.&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam index file in path/freeze5&lt;br /&gt;
 CHRS = 20 22&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 INDEL_PREFIX = $(REF_DIR)/1kg.pilot_release.merged.indels.sites.hg19&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Variant Calling Command-line Options/Configuration Settings ==&lt;br /&gt;
{{:GotCloud: Variant Calling Options}}&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Use Cases &amp;amp; Recommended Settings ==&lt;br /&gt;
=== Single Sample Processing ===&lt;br /&gt;
To run single sample processing we recommend adding the following settings to your configuration file:&lt;br /&gt;
 UNIT_CHUNK = 20000000&lt;br /&gt;
 MODEL_GLFSINGLE = TRUE&lt;br /&gt;
 MODEL_SKIP_DISCOVER = FALSE&lt;br /&gt;
 MODEL_AF_PRIOR = TRUE&lt;br /&gt;
 VCF_EXTRACT = $(REF_DIR)/snpOnly.vcf.gz&lt;br /&gt;
 EXT = $(REF_DIR)/ALL.chrCHR.phase3.combined.sites.unfiltered.vcf.gz $(REF_DIR)/chrCHR.filtered.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
Explanation of these settings:&lt;br /&gt;
* &amp;lt;code&amp;gt;UNIT_CHUNK&amp;lt;/code&amp;gt; - since this is only 1 sample, process larger regions at a time than default&lt;br /&gt;
* &amp;lt;code&amp;gt;MODEL_GLFSINGLE&amp;lt;/code&amp;gt; - single sample, so model glfsingle&lt;br /&gt;
* &amp;lt;code&amp;gt;MODEL_SKIP_DISCOVER&amp;lt;/code&amp;gt; - do not skip the variant discovery step&lt;br /&gt;
* &amp;lt;code&amp;gt;MODEL_AF_PRIOR&amp;lt;/code&amp;gt; - use AF prior for genotyping&lt;br /&gt;
* &amp;lt;code&amp;gt;VCF_EXTRACT&amp;lt;/code&amp;gt; - VCF file to use for extracting the site information to genotype&lt;br /&gt;
**  This file is included in the latest reference release: [[GotCloud:_Genetic_Reference_and_Resource_Files#hs37d5-db142|hs37d5-db142]]&lt;br /&gt;
* &amp;lt;code&amp;gt;EXT&amp;lt;/code&amp;gt; - VCF reference files to use for the external filtering&lt;br /&gt;
** These files are included in the latest reference release: [[GotCloud:_Genetic_Reference_and_Resource_Files#hs37d5-db142|hs37d5-db142]]&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Running ==&lt;br /&gt;
&lt;br /&gt;
Running variant calling is straightforward:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;code&amp;gt;&lt;br /&gt;
 &#039;&#039;&#039;gotcloud snpcall --conf vc.conf --numjobs 2&lt;br /&gt;
 &#039;&#039;&#039;gotcloud ldrefine --conf vc.conf --numjobs 2&lt;br /&gt;
&amp;lt;/code&amp;gt;&lt;br /&gt;
&lt;br /&gt;
* Replace &amp;lt;code&amp;gt;vc.conf&amp;lt;/code&amp;gt; with the path/name of the user&#039;s configuration file&lt;br /&gt;
** If you are not overriding any defaults, you can alternatively specify &amp;lt;code&amp;gt;--list path/bam.list&amp;lt;/code&amp;gt; replacing &amp;lt;code&amp;gt;path/bam.list&amp;lt;/code&amp;gt; with the path/name of your BAM list file.&lt;br /&gt;
* Replace &amp;lt;code&amp;gt;2&amp;lt;/code&amp;gt; following &amp;lt;code&amp;gt;--numjobs&amp;lt;/code&amp;gt; with the number of jobs to be run in parallel&lt;br /&gt;
* If &amp;lt;code&amp;gt;OUT_DIR&amp;lt;/code&amp;gt; is not defined in the configuration file, add &amp;lt;code&amp;gt;--outdir&amp;lt;/code&amp;gt; followed by the path to the user&#039;s desired output directory.&lt;br /&gt;
&lt;br /&gt;
=== Running on a Cluster ===&lt;br /&gt;
See [[#Cluster Configuration|Cluster Configuration]] for information on how to configure GotCloud to run on a cluster.&lt;br /&gt;
&lt;br /&gt;
== Results ==&lt;br /&gt;
&lt;br /&gt;
If there is a failure, you should see a message like: &lt;br /&gt;
 make: *** [...] Error 1&lt;br /&gt;
Where ... is filled in with other text indicating what step failed.&lt;br /&gt;
&lt;br /&gt;
On SNP Call success, you should see the following output sub-directories under your output directory:&lt;br /&gt;
* glfs with a bams &amp;amp; samples subdirectory&lt;br /&gt;
* pvcfs with a subdirectory per chromosome and then per region&lt;br /&gt;
* &#039;&#039;&#039;split&#039;&#039;&#039; with a subdirectory per chromosome&lt;br /&gt;
* &#039;&#039;&#039;vcfs&#039;&#039;&#039; with a subdirectory per chromosome&lt;br /&gt;
* (optionally your target directory)&lt;br /&gt;
&lt;br /&gt;
Under the &#039;&#039;&#039;vcf/chrXX&#039;&#039;&#039; directory, there should be:&lt;br /&gt;
* chrXX.filtered.sites.vcf&lt;br /&gt;
* chrXX.filtered.sites.vcf.norm.log&lt;br /&gt;
* chrXX.filtered.sites.vcf.summary&lt;br /&gt;
* &#039;&#039;&#039;chrXX.filtered.vcf.gz&#039;&#039;&#039; - final filtered variant call file&lt;br /&gt;
* chrXX.filtered.vcf.gz.OK&lt;br /&gt;
* chrXX.filtered.vcf.gz.tbi&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf.log&lt;br /&gt;
* chrXX.hardfiltered.sites.vcf.summary&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz.OK&lt;br /&gt;
* chrXX.hardfiltered.vcf.gz.tbi&lt;br /&gt;
* chrXX.merged.sites.vcf&lt;br /&gt;
* chrXX.merged.stats.vcf&lt;br /&gt;
* chrXX.merged.vcf&lt;br /&gt;
* chrXX.merged.vcf.OK&lt;br /&gt;
&lt;br /&gt;
The .merged.vcf is the merged together versions of the separate regions in the same chromosome.&lt;br /&gt;
&lt;br /&gt;
The filtered is the merged.vcf after it has been run through filters and is marked with PASS/FAIL.&lt;br /&gt;
&lt;br /&gt;
Under the &#039;&#039;&#039;split/chrXX&#039;&#039;&#039; directory, there should be:&lt;br /&gt;
* chrXX.filtered.PASS.split.[N].vcf.gz&lt;br /&gt;
* chrXX.filtered.PASS.split.err&lt;br /&gt;
* chrXX.filtered.PASS.split.vcflist&lt;br /&gt;
* &#039;&#039;&#039;chrXX.filtered.PASS.gz&#039;&#039;&#039; - final variant call file with only PASS variants&lt;br /&gt;
* subset.OK&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=VerifyBamID&amp;diff=13042</id>
		<title>VerifyBamID</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=VerifyBamID&amp;diff=13042"/>
		<updated>2015-03-20T01:25:08Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Column information in the output files */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;[[Category:Software|VerifyBamID]]&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;verifyBamID&#039;&#039;&#039; is a software that verifies whether the reads in particular file match previously known genotypes for an individual (or group of individuals), and checks whether the reads are contaminated as a mixture of two samples. &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; can detect sample contamination and swaps when external genotypes are available. When external genotypes are not available, &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; still robustly detects sample swaps.&lt;br /&gt;
&lt;br /&gt;
== Download verifyBamID  ==&lt;br /&gt;
&lt;br /&gt;
To get a copy of verifyBamId, go to: https://github.com/statgen/verifyBamID/releases&lt;br /&gt;
&lt;br /&gt;
Select the latest release and download in one of 3 ways:&lt;br /&gt;
# Binary expected to run in Ubuntu x64 platform. In other platforms, please download the source distribution and build it.&lt;br /&gt;
#* verifyBamID.#.#.#.gz&lt;br /&gt;
#* You will need to run &amp;quot;gunzip&amp;quot; on the .gz file&lt;br /&gt;
# Souce Code including libStatGen (uses a fixed version of libStatGen)&lt;br /&gt;
#* verifyBamIDLibStatGen.#.#.#.tgz&lt;br /&gt;
#* Run &amp;quot;tar xvf&amp;quot; on this file.  Cd into the resulting directory &amp;amp; type make.&lt;br /&gt;
# Source Code without libStatGen (allows alternative/newer versions of libStatGen)&lt;br /&gt;
#* Source code (tar.gz) or Source code (zip)&lt;br /&gt;
#* You will need to download libStatGen separately if you do not already have it.&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
To get a copy of older releases go to the [http://www.sph.umich.edu/csg/kang/verifyBamID/download VerifyBamID Download] download page.&lt;br /&gt;
&lt;br /&gt;
== Join in verifyBamID mailing list ==&lt;br /&gt;
&lt;br /&gt;
Please join in the [http://groups.google.com/group/verifybamid VerifyBamID Google Group] to ask / discuss / comment about verifyBamID.&lt;br /&gt;
&lt;br /&gt;
== What&#039;s new ==&lt;br /&gt;
&lt;br /&gt;
(2014/02/13)&lt;br /&gt;
* Put verifyBamID in github.&lt;br /&gt;
* Added PhoneHome/Version Checking to VerifyBamID&lt;br /&gt;
&lt;br /&gt;
(2012/06/20) &lt;br /&gt;
* Fixed a bug of incorrect estimate of contamination when --chip-full option was used (Thanks to Richard Smith)&lt;br /&gt;
* Fixed a bug of incorrect per-readgroup output in --chip-* parameter&lt;br /&gt;
&lt;br /&gt;
(2012/05/24) &lt;br /&gt;
* Fixed a bug of incorrect per-readgroup output (Thanks to Matthew Flickinger)&lt;br /&gt;
* &#039;&#039;&#039;(IMPORTANT)&#039;&#039;&#039; Add an option to remove either side of overlapping fragment. This option is turned on by default, and can be turned off usig --ignoreOverlapPair. If your sequence data has very short insert size, this update may increase the sensitivity of estimated contamination.&lt;br /&gt;
* Changes in the directory structure and Makefile&lt;br /&gt;
&lt;br /&gt;
(2012/05/18) The new release of verifyBamID have undergone major change since the last version (as of 2011 April). Here are the highlights&lt;br /&gt;
* The genotype / allele frequency file is now based on VCF format rather than PLINK format.&lt;br /&gt;
* The reference sequence information is no longer required&lt;br /&gt;
* Uses Brent&#039;s method for precise estimation of contamination parameters&lt;br /&gt;
* Generate the depth distribution statistics.&lt;br /&gt;
* Estimated reference-bias parameters (useful mostly for ABI SOLiD sequence data)&lt;br /&gt;
&lt;br /&gt;
== Build verifyBamID  ==&lt;br /&gt;
&lt;br /&gt;
The binary download of verifyBamID is available. You may use that version in Ubuntu 64-bit platform. &lt;br /&gt;
&lt;br /&gt;
If you download the source that includes libStatGen:&lt;br /&gt;
 tar xvf verifyBamIDLibStatGen.#.#.#.tgz&lt;br /&gt;
 cd verifyBamID_#.#.#&lt;br /&gt;
 make&lt;br /&gt;
 Executable: verifyBamID/bin/verifyBamID&lt;br /&gt;
&lt;br /&gt;
If you download the source without libStatGen:&lt;br /&gt;
 tar xvf verifyBamID-#.#.#.tar.gz&lt;br /&gt;
 cd verifyBamID-1.1.0&lt;br /&gt;
 make cloneLib (if ../libStatGen does not exist)&lt;br /&gt;
 make&lt;br /&gt;
 Executable: ./bin/verifyBamID&lt;br /&gt;
&lt;br /&gt;
Note that &#039;&#039;&#039;make cloneLib&#039;&#039;&#039; command will create a directory ../libStatGen under your verifyBamID directory, and &#039;&#039;&#039;make&#039;&#039;&#039; will create binary of verifyBamID under verifyBamID/bin/&lt;br /&gt;
&lt;br /&gt;
If you have a different version of libStatGen at that path, then skip the cloneLib step.  If the libStatGen you want to use is at a different location then update verifyBamID&#039;s Makefile.inc.  Replace: LIB_PATH_VERIFY_BAM_ID ?= $(LIB_PATH_GENERAL) with&lt;br /&gt;
 LIB_PATH_VERIFY_BAM_ID = /path/to/libStatGen&lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
verifyBamID is designed to be reasonably portable. &lt;br /&gt;
&lt;br /&gt;
However, since development occurs only on Ubuntu (9.10-13.10) x86 and x64 platforms, and later, there are likely other portability issues. &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Basic Usage ==&lt;br /&gt;
&lt;br /&gt;
A key step in any genetic analysis is to verify whether data being generated matches expectations. &#039;&#039;verifyBamID&#039;&#039; checks whether reads in a BAM file match previous genotypes for a specific sample. In addition, it detects possible sample mixture from population allele frequency only, which can be particularly useful when the genotype data is not available.&lt;br /&gt;
&lt;br /&gt;
Using a mathematical model that relates observed sequence reads to an hypothetical true genotype, &#039;&#039;verifyBamID&#039;&#039; tries to decide whether sequence reads match a particular individual or are more likely to be contaminated (including a small proportion of foreign DNA), derived from a closely related individual, or derived from a completely different individual.&lt;br /&gt;
&lt;br /&gt;
== Basic Usage Example ==&lt;br /&gt;
&lt;br /&gt;
Here is a typical command line:&lt;br /&gt;
&lt;br /&gt;
 verifyBamID --vcf [input.vcf] --bam [input.bam] --out [output.prefix] --verbose --ignoreRG&lt;br /&gt;
 &lt;br /&gt;
 where&lt;br /&gt;
 [input.bam] is a BAM (Binary Alignment Map) file of a sequence reads&lt;br /&gt;
 [input.vcf] is input VCF file containing individual genotypes or AF or AC/AN fields in the INFO field. gzipped VCF is also allowed.&lt;br /&gt;
 [outPrefix] is output prefix of output files - [outPrefix].{selfRG,selfSM,bestRG,bestSM,depthRG,depthSM} will be created.&lt;br /&gt;
&lt;br /&gt;
More detailed description of command line input is below&lt;br /&gt;
&lt;br /&gt;
== Preparing input files ==&lt;br /&gt;
&lt;br /&gt;
verifyBamID requires two input files - VCF file containing external genotypes or allele frequency information, and the BAM file.&lt;br /&gt;
&lt;br /&gt;
=== VCF input genotype file ===&lt;br /&gt;
&lt;br /&gt;
The input VCF file contains (1) external genotype information and/or (2) allele frequency information as AF entry or AC/AN entries in the INFO field. (See [http://www.1000genomes.org/wiki/Analysis/Variant%20Call%20Format/vcf-variant-call-format-version-41 | VCF specification] for further details). If neither information is provided, verifyBamID will not work properly.&lt;br /&gt;
&lt;br /&gt;
If external genotype information is provided, sequence+array method will identify contamination and sample swaps by comparing the concordance between the external genotypes and the sequence reads. Additionally, sequence-only method will provide additional contamination estimates by modeling the sequence reads as mixture of two unknown samples based on the allele frequency information in the VCF file.&lt;br /&gt;
&lt;br /&gt;
Input VCF file needs to meet several additional contraints need to meet in order to properly run verifyBamID.&lt;br /&gt;
* The VCF is assumed to be well-formed. For example, verifyBamID does not check whether REF allele actually matches with reference sequence.  &lt;br /&gt;
* The VCF should only contain SNPs. Current version of verifyBamID does not accept INDELs, MNPs, Structural Variations, or other complex variants.&lt;br /&gt;
* The individual IDs in the VCF file, must be identical with the individual identifier in the BAM file. Otherwise, --smID option can override the sample ID information of the BAM file to the ID that matches to the individual IDs in the VCF file.&lt;br /&gt;
* IMPORTANT : For targeted sequencing data, it is important to subselect the markers to only include on-target markers in the genotype file. Off-target markers are not likely to have multiple non-duplicated reads at the marker position, and it may create artifacts in the analysis due to overlapping fragments.&lt;br /&gt;
* Currently, verifyBamID takes only autosomal chromosomes as input VCF.&lt;br /&gt;
&lt;br /&gt;
An example input VCF file (without external genotype) is provided below. Note that AC and AC entries exists in the INFO field for the allele frequency information.&lt;br /&gt;
&lt;br /&gt;
 #CHROM	POS	ID	REF	ALT	QUAL	FILTER	INFO&lt;br /&gt;
 20	61651	SNP20-9651	C	A	.	PASS	CR=99.86851;GentrainScore=0.7055;HW=0.077647716;AN=2180;AC=11&lt;br /&gt;
 20	63231	SNP20-11231	T	G	.	PASS	CR=99.93036;GentrainScore=0.7837;HW=0.035481825;AN=2182;AC=275&lt;br /&gt;
 20	63244	rs6139074	A	C	.	PASS	CR=98.893394;GentrainScore=0.8001;HW=7.327299E-7;AN=2162;AC=501&lt;br /&gt;
 20	63799	rs1418258	C	T	.	PASS	CR=99.75217;GentrainScore=0.8170;HW=0.6653377;AN=2182;AC=881&lt;br /&gt;
&lt;br /&gt;
=== Input BAM file ===&lt;br /&gt;
&lt;br /&gt;
verifyBamID requires a sorted, indexed, base quality recalibrated, and duplication-marked BAM file. It also requires to contain &amp;quot;@RG&amp;quot; header lines to annotation different readGroups (sequencing runs and lanes). The SM tag in the &amp;quot;@RG&amp;quot; header should match with one of the genotyped sample. Otherwise, verifyBamID may not be able to test whether the sequenced sample matches with genotyped sample, but will try to detect sample mixture from allele frequency, and will try to detect the best-matching sample among the genotyped sample.&lt;br /&gt;
&lt;br /&gt;
== What the default option does ==&lt;br /&gt;
&lt;br /&gt;
The default option of &#039;&#039;&#039;verifyBamID&#039;&#039;&#039; is the recommended setting for the most sequencing studies to provide a rapid and informative response. The default option provides the following features:&lt;br /&gt;
* --free-mix is turned on for estimating contamination using sequence-only method&lt;br /&gt;
* --chip-mix is turned on for estimating contamination or swap using sequence+array method, if the external genotype file is provided in the VCF&lt;br /&gt;
* --self is turnd on : The default option does not try to compare the sequence reads to identify the best matching individual (which is possible with --best option). It only compares with the external genotypes from the same individual to the sequenced individual.&lt;br /&gt;
* --maxDepth 20 is used without --precise option : The default option is intended for whole genome low coverage sequencing. For the targeted exome sequencing, --maxDepth 1000 and --precise is recommended.&lt;br /&gt;
* --ignoreRG is not a default option, but a recommended option, when you want to check the contamination for the entire BAM rather than examining each read group separately. This option will increase the computational efficiency especially in the case whether the sequence reads are multiplexed across many sequencing runs.&lt;br /&gt;
&lt;br /&gt;
== Interpreting output files ==&lt;br /&gt;
&lt;br /&gt;
See also [[Understanding VerifyBamID output]].&lt;br /&gt;
&lt;br /&gt;
=== Output files ===&lt;br /&gt;
When verifyBamID runs successfully, the following sets of files may be generated.&lt;br /&gt;
* [outPrefix].selfSM - Per-sample statistics describing how well the sample matches to the annotated sample.&lt;br /&gt;
* [outPrefix].depthSM - The depth distribution of the sequence reads per sample&lt;br /&gt;
* [outPrefix].selfRG - Per-readGroup statistics describing how well each lane matches to the annotated sample. (available only without --ignoreRG option)&lt;br /&gt;
* [outPrefix].depthRG - The depth distribution of the sequence reads per readGroup. (available only without --ignoreRG option)&lt;br /&gt;
* [outPrefix].bestSM - Per-sample best-match statistics with best-matching sample among the genotyped sample (available only with --best option)&lt;br /&gt;
* [outPrefix].bestRG - Per-readgroup best-match statistics with best-matching sample among the genotyped sample (available only with --best and without --ignoreRG option)&lt;br /&gt;
&lt;br /&gt;
=== Column information in the output files ===&lt;br /&gt;
The .selfSM/.selfRG/.bestSM/.bestRG files have the following 19 columns per sample, or per readgroup (lane). &lt;br /&gt;
&lt;br /&gt;
# SEQ_SM : Sample ID of the sequenced sample. Obtained from @RG header / SM tag in the BAM file&lt;br /&gt;
# RG : ReadGroup ID of sequenced lane. For [outPrefix].selfSM and [outPrefix].bestSM, these values are &amp;quot;ALL&amp;quot;&lt;br /&gt;
# CHIP_ID : Sample ID compared to in the genotype file. For [outPrefix].selfRG and [outPrefix].selfSM, these values should be identical to [SEQ_SM] or &amp;quot;NA&amp;quot; if the genotype of sequenced samples are unavailable. For [outPrefix].bestRG and [outPrefix].bestSM, these values should be the ID of best-matching sample among the genotype files compared to.&lt;br /&gt;
# # SNPs : # of SNPs passing the criteria from the VCF file&lt;br /&gt;
# # READS : Total # of reads loaded from the BAM file&lt;br /&gt;
# # AVG_DP : Average sequencing depth at the sites in the VCF file&lt;br /&gt;
# FREEMIX : Sequence-only estimate of contamination (0-1 scale)&lt;br /&gt;
# FREELK1 : Maximum log-likelihood of the sequence reads given estimated contamination under sequence-only method&lt;br /&gt;
# FREELK0 : Log-likelihood of the sequence reads given no contamination under sequence-only method&lt;br /&gt;
# FREE_RH : Estimated reference bias parameter Pr(refBase|HET) (when --free-refBias or --free-full is used)&lt;br /&gt;
# FREE_RA : Estimated reference bias parameter Pr(refBase|HOMALT) (when --free-refBias or --free-full is used)&lt;br /&gt;
# CHIPMIX : Sequence+array estimate of contamination (NA if the external genotype is unavailable) (0-1 scale)&lt;br /&gt;
# CHIPLK1 : Maximum log-likelihood of the sequence reads given estimated contamination under sequence+array method (NA if the external genotypes are unavailable)&lt;br /&gt;
# CHIPLK0 : Log-likelihood of the sequence reads given no contamination under sequence+array method (NA if the external genotypes are unavailable)&lt;br /&gt;
# CHIP_RH : Estimated reference bias parameter Pr(refBase|HET) (when --chip-refBias or --chip-full is used)&lt;br /&gt;
# CHIP_RA : Estimated reference bias parameter Pr(refBase|HOMALT) (when --chip-refBias or --chip-full is used)&lt;br /&gt;
# DPREF : Depth (Coverage) of HomRef site (based on the genotypes of (SELF_SM/BEST_SM), passing mapQ, baseQual, maxDepth thresholds.&lt;br /&gt;
# RDPHET : DPHET/DPREF, Relative depth at Heterozygous site.&lt;br /&gt;
# RDPALT : DPHET/DPREF, Relative depth at HomAlt site.&lt;br /&gt;
&lt;br /&gt;
=== A guideline to interpret output files ===&lt;br /&gt;
&lt;br /&gt;
verifyBamID provides a series of information that is informative to determine whether the sample is possibly contaminated or swapped, but there is no single criteria that works for every circumstances. There are a few unmodeled factor in the estimation of [SELF-IBD]/[BEST-IBD] and [%MIX], so please note that the MLE estimation may not always exactly match to the true amount of contamination. Here we provide a guideline to flag potentially contaminated/swapped samples &lt;br /&gt;
&lt;br /&gt;
*  Each sample or lane can be checked in this way. When [CHIPMIX] &amp;gt;&amp;gt; 0.02 and/or [FREEMIX] &amp;gt;&amp;gt; 0.02, meaning 2% or more of non-reference bases are observed in reference sites, we recommend to examine the data more carefully for the possibility of contamination.&lt;br /&gt;
* We recommend to check each lane for the possibility of sample swaps. When [CHIPMIX] ~ 1 AND [FREEMIX] ~ 0, then it is possible that the sample is swapped with another sample. When [CHIPMIX] ~ 0 in .bestSM file, [CHIP_ID] might be actually the swapped sample. Otherwise, the swapped sample may not exist in the genotype data you have compared. &lt;br /&gt;
* When genotype data is not available but allele-frequency-based estimates of [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, then it is possible that the sample is contaminated with other sample. We recommend to use per-sample data rather than per-lane data for checking this for low coverage data, because the inference will be more confident when there are large number of bases with depth 2 or higher.&lt;br /&gt;
&lt;br /&gt;
== Command Line Options ==&lt;br /&gt;
&lt;br /&gt;
 The following parameters are available.  Ones with &amp;quot;[]&amp;quot; are in effect:&lt;br /&gt;
 &lt;br /&gt;
 Available Options&lt;br /&gt;
                             Input Files : --vcf [], --bam [], --subset [],&lt;br /&gt;
                                           --smID []&lt;br /&gt;
                    VCF analysis options : --genoError [1.0e-03],&lt;br /&gt;
                                           --minAF [0.01],&lt;br /&gt;
                                           --minCallRate [0.50]&lt;br /&gt;
   Individuals to compare with chip data : --site, --self, --best&lt;br /&gt;
          Chip-free optimization options : --free-none, --free-mix [ON],&lt;br /&gt;
                                           --free-refBias, --free-full&lt;br /&gt;
          With-chip optimization options : --chip-none, --chip-mix [ON],&lt;br /&gt;
                                           --chip-refBias, --chip-full&lt;br /&gt;
                    BAM analysis options : --ignoreRG, --ignoreOverlapPair,&lt;br /&gt;
                                           --noEOF, --precise, --minMapQ [10],&lt;br /&gt;
                                           --maxDepth [20], --minQ [13],&lt;br /&gt;
                                           --maxQ [40], --grid [0.05]&lt;br /&gt;
                 Modeling Reference Bias : --refRef [1.00], --refHet [0.50],&lt;br /&gt;
                                           --refAlt [0.00]&lt;br /&gt;
                          Output options : --out [], --verbose&lt;br /&gt;
                               PhoneHome : --noPhoneHome,&lt;br /&gt;
                                           --phoneHomeThinning [50]&lt;br /&gt;
&lt;br /&gt;
Each option provides the following features:&lt;br /&gt;
* --vcf : specify required VCF file&lt;br /&gt;
* --bam : specify required BAM file (indexed with .bam.bai or .bai file)&lt;br /&gt;
* --subset : list of individual IDs to calculate the allele frequency. All individuals are used if unspecified&lt;br /&gt;
* --smID : If the individual ID in the BAM file and VCF file does not match, substitute the BAM file&#039;s ID into the specified argument&lt;br /&gt;
* --genoError : error rate of the external genotype file&lt;br /&gt;
* --minAF : minimum allele frequency of the markers to include&lt;br /&gt;
* --minAF : minimum call rate of the markers to include&lt;br /&gt;
* --site : If set, use only site information in the VCF and do not compare with the actual genotypes&lt;br /&gt;
* --self : Only compare the ID-matching individuals between the VCF and BAM file&lt;br /&gt;
* --best : Find the best matching individuals (.bestSM and .bestRG files will be produced). This option is substantially longer than the default option&lt;br /&gt;
* --free-none : Do not perform sequence-only method to estimate parameters&lt;br /&gt;
* --free-mix : (default) Estimate contamination using sequence-only method with Brent&#039;s single dimensional optimization.&lt;br /&gt;
* --free-refBias : Estimate the reference bias parameters using sequence-only method with Simplex method&lt;br /&gt;
* --free-full : Estimate both reference bias parameters and the contamination parameters using sequence-only method&lt;br /&gt;
* --chip-none : Do not perform sequence+array method to estimate parameters&lt;br /&gt;
* --free-mix : (default) Estimate contamination using sequence+array method with Brent&#039;s single dimensional optimization.&lt;br /&gt;
* --free-refBias : Estimate the refernece bias parameters using sequence+array method with Simplex method&lt;br /&gt;
* --free-full : Estimate both reference bias parameters and the contamination parameters using sequence+array method&lt;br /&gt;
* --ignoreRG : ignore the read grouup level comparison and compare samples only (recommended for an expedited run)&lt;br /&gt;
* --ignoreOverlapPair : ignore overlapping pair end fragment covering the same base. Disabling this option may decrease the sensitivity of the method when the insert size is short (with slight gain in the computational speed)&lt;br /&gt;
* --noEOF : do not check the EOF marker of the BAM file (for earlier version of BAM)&lt;br /&gt;
* --precise : calculate the likelihood in log-scale for high-depth data (recommended when --maxDepth is greater than 20. Can be a little bit slower)&lt;br /&gt;
* --minMapQ : minimum mapping quality of the sequence reads to compare&lt;br /&gt;
* --minQ : minimum base quality to include&lt;br /&gt;
* --maxQ : maximum base quality to cap&lt;br /&gt;
* --grid : the grid interval to search the optimum before running Brent&#039;s algorithm.&lt;br /&gt;
* --refRef : Initial Pr(refBase|HOMREFGeno) parameter&lt;br /&gt;
* --refHet : Initial Pr(refBase|HETGeno) parameter&lt;br /&gt;
* --refAlt : Initial Pr(refBase|HOMALTGeno) parameter&lt;br /&gt;
* --out : output file prefix (required)&lt;br /&gt;
* --verbose : print the progress of the method on the screeen&lt;br /&gt;
{{PhoneHomeParameters|hdr=====|bullet=1}}&lt;br /&gt;
&lt;br /&gt;
== Principle of Operation ==&lt;br /&gt;
&lt;br /&gt;
Each read group in a BAM file is evaluated independently. This means that in file with multiple read groups, problems will be flagged at the read group level (a plus). However, it also means that it might be hard to discern the correct assignment of read groups with very little data.&lt;br /&gt;
&lt;br /&gt;
For each aligned base that overlaps a known genotype, we calculate the probability the probability that it was derived from a particular known genotype. This comparison considers only bases that overlap previously known genotypes and that meet the base quality and mapping quality thresholds.&lt;br /&gt;
&lt;br /&gt;
Each individual in a pedigree has a different combination of genotypes, and bamGenotypeCheck will systematically search for the individual whose genotypes best match the observed read data.&lt;br /&gt;
&lt;br /&gt;
For more about the technical details, see the page [[Verifying Sample Identities - Implementation]]&lt;br /&gt;
&lt;br /&gt;
== Reference ==&lt;br /&gt;
&lt;br /&gt;
Please cite the following paper:&lt;br /&gt;
&lt;br /&gt;
G. Jun, M. Flickinger, K. N. Hetrick, Kurt, J. M. Romm, K. F. Doheny, G. Abecasis, M. Boehnke,and H. M. Kang, &#039;&#039;Detecting and Estimating Contamination of Human DNA Samples in Sequencing and Array-Based Genotype Data&#039;&#039;, American journal of human genetics doi:10.1016/j.ajhg.2012.09.004 (volume 91 issue 5 pp.839 - 848) &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== Contamination in Array Data ==&lt;br /&gt;
&lt;br /&gt;
[[VerifyIDintensity]] or [[BAFRegress]] can estimate sample contamination from Illumina genotype array data.&lt;br /&gt;
&lt;br /&gt;
== Acknowledgements ==&lt;br /&gt;
&lt;br /&gt;
VerifyBamID is a result from collaborative effort by Hyun Min Kang, Goo Jun, Matthew Flickinger, Mary Kate Wing, and Goncalo Abecasis. Please email to Hyun Min Kang [[mailto:hmkang@umich.edu| hmkang@umich.edu ]] for any questions.&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13041</id>
		<title>GotCloud: Alignment Sub-Pipelines</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13041"/>
		<updated>2015-03-19T01:37:18Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Example Command Line */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== List of Alignment Sub-Pipelines == &lt;br /&gt;
&lt;br /&gt;
===recab=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM.&lt;br /&gt;
&lt;br /&gt;
===recabQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline does everything that *recab* does (takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM). It then goes the next step to perform quality control (running qplot and verifyBamID).&lt;br /&gt;
&lt;br /&gt;
===bamQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file and its index file (.bai) and performs quality control (running qplot and verifyBamID). It differs from *bamQC_createIndex* in that it requires that the user already have .bai files for the recalibrated BAM files. &lt;br /&gt;
 &lt;br /&gt;
===bamQC_createIndex=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file, creates an index file for it, and performs quality control (running qplot and verifyBamID). It differs from *bamQC* in that it does not require that the user already have a .bai file for the recalibrated BAM file. &lt;br /&gt;
&lt;br /&gt;
== recab ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recab|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recab=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recab* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recab* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recab|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF Files|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recab --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== recabQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recabQC|BAM_LIST]] file)&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recabQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recabQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039; - contains merge results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039; - contains recalibration results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recabQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais, while the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recabQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recabQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC|BAM_LIST File]])&lt;br /&gt;
* BAI file for each subject&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full path to BAM file)&lt;br /&gt;
# BAI File (preferable to have full path to BAI file)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC_createIndex ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# creates a BAI file for any BAM that is missing it&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC_createIndex|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC_createIndex=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full paths to BAM files)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC_createIndex* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
* A BAI file with the exact same path and name as the BAM file that was input, with *.bai on the end&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC_createIndex* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC_createIndex|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC_createIndex --numjobs &amp;lt;N&amp;gt;&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13040</id>
		<title>GotCloud: Alignment Sub-Pipelines</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13040"/>
		<updated>2015-03-19T01:36:09Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Command-Line and Configuration Options */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== List of Alignment Sub-Pipelines == &lt;br /&gt;
&lt;br /&gt;
===recab=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM.&lt;br /&gt;
&lt;br /&gt;
===recabQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline does everything that *recab* does (takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM). It then goes the next step to perform quality control (running qplot and verifyBamID).&lt;br /&gt;
&lt;br /&gt;
===bamQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file and its index file (.bai) and performs quality control (running qplot and verifyBamID). It differs from *bamQC_createIndex* in that it requires that the user already have .bai files for the recalibrated BAM files. &lt;br /&gt;
 &lt;br /&gt;
===bamQC_createIndex=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file, creates an index file for it, and performs quality control (running qplot and verifyBamID). It differs from *bamQC* in that it does not require that the user already have a .bai file for the recalibrated BAM file. &lt;br /&gt;
&lt;br /&gt;
== recab ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recab|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recab=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recab* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recab* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recab|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF Files|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recab --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== recabQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recabQC|BAM_LIST]] file)&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recabQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recabQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039; - contains merge results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039; - contains recalibration results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recabQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais, while the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recabQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recabQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC|BAM_LIST File]])&lt;br /&gt;
* BAI file for each subject&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full path to BAM file)&lt;br /&gt;
# BAI File (preferable to have full path to BAI file)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC_createIndex ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# creates a BAI file for any BAM that is missing it&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC_createIndex|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC_createIndex=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full paths to BAM files)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC_createIndex* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
* A BAI file with the exact same path and name as the BAM file that was input, with *.bai on the end&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC_createIndex* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC_createIndex|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13039</id>
		<title>GotCloud: Alignment Sub-Pipelines</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13039"/>
		<updated>2015-03-19T01:35:58Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* Inputs */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== List of Alignment Sub-Pipelines == &lt;br /&gt;
&lt;br /&gt;
===recab=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM.&lt;br /&gt;
&lt;br /&gt;
===recabQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline does everything that *recab* does (takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM). It then goes the next step to perform quality control (running qplot and verifyBamID).&lt;br /&gt;
&lt;br /&gt;
===bamQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file and its index file (.bai) and performs quality control (running qplot and verifyBamID). It differs from *bamQC_createIndex* in that it requires that the user already have .bai files for the recalibrated BAM files. &lt;br /&gt;
 &lt;br /&gt;
===bamQC_createIndex=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file, creates an index file for it, and performs quality control (running qplot and verifyBamID). It differs from *bamQC* in that it does not require that the user already have a .bai file for the recalibrated BAM file. &lt;br /&gt;
&lt;br /&gt;
== recab ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recab|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recab=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recab* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recab* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recab|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF Files|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recab --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== recabQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recabQC|BAM_LIST]] file)&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recabQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recabQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039; - contains merge results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039; - contains recalibration results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recabQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais, while the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recabQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recabQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC|BAM_LIST File]])&lt;br /&gt;
* BAI file for each subject&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full path to BAM file)&lt;br /&gt;
# BAI File (preferable to have full path to BAI file)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC_createIndex ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# creates a BAI file for any BAM that is missing it&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC_createIndex|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC_createIndex=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full paths to BAM files)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC_createIndex* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
* A BAI file with the exact same path and name as the BAM file that was input, with *.bai on the end&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC_createIndex* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13038</id>
		<title>GotCloud: Alignment Sub-Pipelines</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=GotCloud:_Alignment_Sub-Pipelines&amp;diff=13038"/>
		<updated>2015-03-19T01:35:44Z</updated>

		<summary type="html">&lt;p&gt;Kleckner: /* BAM_LIST File for bamQC */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;br /&gt;
Back to parent: [[GotCloud]] &lt;br /&gt;
&lt;br /&gt;
&lt;br /&gt;
== List of Alignment Sub-Pipelines == &lt;br /&gt;
&lt;br /&gt;
===recab=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM.&lt;br /&gt;
&lt;br /&gt;
===recabQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline does everything that *recab* does (takes in a list of bam files for each sample, merges the BAMs for samples that have multiple BAMs, dedups and recalibrates, and then indexes the recalibrated BAM). It then goes the next step to perform quality control (running qplot and verifyBamID).&lt;br /&gt;
&lt;br /&gt;
===bamQC=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file and its index file (.bai) and performs quality control (running qplot and verifyBamID). It differs from *bamQC_createIndex* in that it requires that the user already have .bai files for the recalibrated BAM files. &lt;br /&gt;
 &lt;br /&gt;
===bamQC_createIndex=== &lt;br /&gt;
&lt;br /&gt;
This sub-pipeline takes in a single, recalibrated BAM file, creates an index file for it, and performs quality control (running qplot and verifyBamID). It differs from *bamQC* in that it does not require that the user already have a .bai file for the recalibrated BAM file. &lt;br /&gt;
&lt;br /&gt;
== recab ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recab|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recab=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recab* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recab* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais.&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recab|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF Files|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recab --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== recabQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# merge BAMs for samples that have multiple BAMs&lt;br /&gt;
# dedup and recalibrate&lt;br /&gt;
# index the recalibrated BAM&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Bam files (stored in a [[#BAM_LIST File for recabQC|BAM_LIST]] file)&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for recabQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File 1 (preferable to have full paths to BAM files)&lt;br /&gt;
# BAM File 2 (if more than 1 BAM per sample)&lt;br /&gt;
:...&lt;br /&gt;
&lt;br /&gt;
: # BAM File N (if more than 1 BAM per sample)&lt;br /&gt;
 [SAMPLE_ID]    [COMMA SEPARATED POPULATION LABELS] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE1] [BAM_FILE2] ...&lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** multiple BAMs per individual may be provided, but should all be on the same line of the list file&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *recabQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
*&#039;&#039;&#039;recab/mergedBams/&#039;&#039;&#039; - contains merge results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.merged.bam&#039;&#039; - a merged BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.log&#039;&#039; - merge log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.merged.bam.OK&#039;&#039; - temp file indicating the merge step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;recab/&#039;&#039;&#039; - contains recalibration results&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam&#039;&#039; - a merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.recal.bam.bai&#039;&#039; - an indexed version of the  merged, recalibrated, and deduped BAM file&#039;&#039;&#039;&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.metrics&#039;&#039; - dedup &amp;amp; recalibration log&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.qemp&#039;&#039; - recalibration tables&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.done&#039;&#039; - temp file indicating the recalibration step completed successfully&lt;br /&gt;
** &#039;&#039;*/SAMPLE.recal.bam.bai.done&#039;&#039; - temp file indicating the indexing step completed successfully&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *recabQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the recab/ folder contains the final BAMs and bais, while the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for recabQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name recabQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC|BAM_LIST File]])&lt;br /&gt;
* BAI file for each subject&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full path to BAM file)&lt;br /&gt;
# BAI File (preferable to have full path to BAI file)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] [BAI_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== bamQC_createIndex ==&lt;br /&gt;
*What it does: &lt;br /&gt;
# creates a BAI file for any BAM that is missing it&lt;br /&gt;
# qplot&lt;br /&gt;
# verifyBamID&lt;br /&gt;
&lt;br /&gt;
====Inputs====&lt;br /&gt;
* Single merged, recalibrated, and deduped BAM file for each subject (stored in a [[#BAM_LIST File for bamQC|BAM_LIST File]])&lt;br /&gt;
* Reference files&lt;br /&gt;
* (Optional) configuration file to override default options&lt;br /&gt;
&lt;br /&gt;
=====BAM_LIST File for bamQC_createIndex=====&lt;br /&gt;
* Each line of the BAM list file represents a single individual&lt;br /&gt;
&lt;br /&gt;
Columns:&lt;br /&gt;
# sample id&lt;br /&gt;
# comma separated population labels (optional column)&lt;br /&gt;
# BAM File (preferable to have full paths to BAM files)&lt;br /&gt;
&lt;br /&gt;
 [SAMPLE_ID] [COMMA SEPARATED POPULATION LABELS] [BAM_FILE] &lt;br /&gt;
or&lt;br /&gt;
 [SAMPLE_ID] [BAM_FILE] &lt;br /&gt;
&lt;br /&gt;
* Notes:&lt;br /&gt;
** tab delimited&lt;br /&gt;
** population label is optional - it will default to &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt;&lt;br /&gt;
*** only used by Thunder (part of ldrefine pipeline)&lt;br /&gt;
*** if all samples are from the same population, population label can be skipped or you can just specify &amp;lt;code&amp;gt;ALL&amp;lt;/code&amp;gt; for the population label for each sample.&lt;br /&gt;
&lt;br /&gt;
====Outputs====&lt;br /&gt;
Upon successful completion of the *bamQC_createIndex* sub-pipeline, you should see the following files/subdirectories under the user specified output directory:&lt;br /&gt;
* A BAI file with the exact same path and name as the BAM file that was input, with *.bai on the end&lt;br /&gt;
* &#039;&#039;&#039;QCFiles/&#039;&#039;&#039; - contains quality control results &lt;br /&gt;
** VerifyBamID Output - see [[VerifyBamID#A_guideline_to_interpret_output_files|VerifyBamID: A guideline to interpret output files]] for more information&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthRG&#039;&#039; - depth distribution of the sequence reads per read group&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.depthSM&#039;&#039; - depth distribution of the sequence reads per sample&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.err&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.log&#039;&#039; - log file&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.OK&#039;&#039; - temp file indicating the VerifyBAMID step completed successfully&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.genoCheck.selfRG&#039;&#039; - per-readGroup statistics describing how well each lane matches to the annotated sample&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.genoCheck.selfSM&#039;&#039; - main output file containing the contamination estimate; per-sample statistics describing how well the sample matches to the annotated sample&#039;&#039;&#039;&lt;br /&gt;
**** Check the &#039;FREEMIX&#039; column for genotype-free estimate of contamination 0-1 scale, the lower, the better&lt;br /&gt;
**** If [FREEMIX] &amp;gt;= 0.03 and [FREELK1]-[FREELK0] is large, possible contamination&lt;br /&gt;
** Qplot Output - see: [[QPLOT#Diagnose_sequencing_quality|QPLOT: Diagnose sequencing quality]] for more info on how to use QPLOT results&lt;br /&gt;
*** &#039;&#039;*/SAMPLE.qplot.OK&#039;&#039; - temp file indicating the qplot step completed successfully&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.R&#039;&#039; - Rscript that can be used to generate the pdf graphs&#039;&#039;&#039;&lt;br /&gt;
*** &#039;&#039;&#039;&#039;&#039;*/SAMPLE.qplot.stats&#039;&#039; - sample statistics&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
You should see .done and .OK files for each SAMPLE in the index file. If you do not see the .done and .OK files, then your *bamQC_createIndex* sub-pipeline failed.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;On success, the QCFiles/ folder contains the quality control output&#039;&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
===Command-Line and Configuration Options===&lt;br /&gt;
&lt;br /&gt;
*Required Options&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --list/--bam_list/--bamlist &#039;&#039;file&#039;&#039; || BAM_LIST || path to the [[#BAM_LIST File for bamQC|BAM_LIST File]] || $(OUT_DIR)/bam.list&lt;br /&gt;
|-&lt;br /&gt;
| --numjobs &#039;&#039;#&#039;&#039; || || number of jobs to run in parallel || 0 (generate Makefile of steps, but do not run)&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
*Common Options&lt;br /&gt;
&lt;br /&gt;
{| class=&amp;quot;wikitable&amp;quot; style=&amp;quot;margin: 1em 1em 1em 0; background-color: #f9f9f9; border: 1px #aaa solid; border-collapse: collapse;&amp;quot; border=&amp;quot;1&amp;quot;&lt;br /&gt;
! Command-line Flag !! Configuration Key !! Value Description !! Default Value&lt;br /&gt;
|-&lt;br /&gt;
| --outdir &#039;&#039;path&#039;&#039; || OUT_DIR || output directory ||&lt;br /&gt;
|-&lt;br /&gt;
| --conf &#039;&#039;file&#039;&#039; || || configuration file to use ||&lt;br /&gt;
|-&lt;br /&gt;
|  || REF_DIR || where the reference/resource files are stored || gotcloud.ref subdirectory within the base GotCloud directory&lt;br /&gt;
|-&lt;br /&gt;
| || REF || [[GotCloud: Genetic Reference and Resource Files#Reference fasta Files|Reference fasta Files]] || $(REF_DIR)/human.g1k.v37.fa&lt;br /&gt;
|-&lt;br /&gt;
| || DBSNP_VCF || [[GotCloud: Genetic Reference and Resource Files#DBSNP VCF File|DBSNP VCF Files]] || $(REF_DIR)/dbsnp_135.b37.vcf.gz&lt;br /&gt;
|-&lt;br /&gt;
| || HM3_VCF || [[GotCloud: Genetic Reference and Resource Files#HapMap3 VCF File|HapMap3 VCF Files]] || $(REF_DIR)/hapmap_3.3.b37.sites.vcf.gz&lt;br /&gt;
|}&lt;br /&gt;
&lt;br /&gt;
==== Example Configuration File ====&lt;br /&gt;
Example configuration file where reference files happen to be stored in /path/reference, and bam list file is stored in in path/freeze5&lt;br /&gt;
 BAM_LIST = /path/freeze5.bam.list&lt;br /&gt;
 OUT_DIR = /path/freeze5/output&lt;br /&gt;
 REF_DIR = /path/reference/&lt;br /&gt;
 REF = $(REF_DIR)/hs37d5.fa&lt;br /&gt;
 DBSNP_VCF = $(REF_DIR)/dbsnp_135.b37.sites.vcf.gz&lt;br /&gt;
 HM3_VCF = $(REF_DIR)/hapmap3_r3_b37.sites.vcf.gz&lt;br /&gt;
&lt;br /&gt;
==== Example Command Line ====&lt;br /&gt;
 gotcloud pipe –-name bamQC --numjobs &amp;lt;N&amp;gt;&lt;/div&gt;</summary>
		<author><name>Kleckner</name></author>
	</entry>
</feed>