<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Sarahgag</id>
	<title>Genome Analysis Wiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="http://genome.sph.umich.edu/w/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=Sarahgag"/>
	<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/wiki/Special:Contributions/Sarahgag"/>
	<updated>2026-09-24T06:02:14Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.43.1</generator>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15059</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15059"/>
		<updated>2018-09-08T14:24:07Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Locations of Files for Current Data Freeze of 3839 Samples */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2018 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual. For association analyses, 22358 has been removed.&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;PheWAS: GWAS on 120 traits for Visit 1&#039;&#039;&#039;&lt;br /&gt;
** Results: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/GOODheader/R2sameasPapers&lt;br /&gt;
** PheWeb (previous release: http://pheweb.irgb.cnr.it); current release: http://sardinia-pheweb.sph.umich.edu&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. &amp;lt;b&amp;gt;--&amp;gt; Complete &amp;lt;/b&amp;gt;&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15058</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15058"/>
		<updated>2018-09-08T14:19:31Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Status as of July 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2018 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. &amp;lt;b&amp;gt;--&amp;gt; Complete &amp;lt;/b&amp;gt;&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15057</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=15057"/>
		<updated>2018-09-08T14:19:07Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. &amp;lt;b&amp;gt;--&amp;gt; Complete &amp;lt;/b&amp;gt;&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14976</id>
		<title>Biostatistics 666: Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14976"/>
		<updated>2017-11-29T20:03:57Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Class Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Objective ==&lt;br /&gt;
&lt;br /&gt;
Gene mapping studies study the relationship between genetic variation and susceptibility to human disease. These studies can be used to elucidate the biochemical basis of medically interesting traits leading to knowledge that will, ultimately, help us improve treatment and management of human disease. Biostatistics 666 is a Masters level course that introduces many of the concepts, statistical models and numerical methods useful for studies.&lt;br /&gt;
&lt;br /&gt;
For additional information, see also [[Biostatistics 666: Core Competencies|Core Competencies in Biostatistics Program covered by this course]].&lt;br /&gt;
&lt;br /&gt;
== Target Audience ==&lt;br /&gt;
&lt;br /&gt;
Students in Biostatistics 666 should be comfortable with simple algebra and, ideally, have previous exposure to maximum likelihood. Previous knowledge of Genetics is helpful, but not required. Most students registering for the course are Master or Doctoral students in Human Genetics, Bioinformatics, Statistics or Biostatistics.&lt;br /&gt;
&lt;br /&gt;
== Scheduling ==&lt;br /&gt;
&lt;br /&gt;
Final grades will take into account performance in written in-class assessments, take home problem sets, as well as participation in class and any class projects.&lt;br /&gt;
&lt;br /&gt;
== Class Notes ==&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Contemporary Human Genetics]] - [[Media:666.2017.01 - Introductory Lecture.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Hardy-Weinberg Equilibrium]] - [[Media:666.2017.02 - Hardy Weinberg Equilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Disequilibrium]] - [[Media:666.2017.03 - Linkage Disequilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the Coalescent]] - [[Media:666.04 - The Coalescent.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Variation in the Coalescent]] - [[Media:666.2017.05_-_Coalescent_and_Allele_Counts.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Recombination and Migration in the Coalescent]] - [[Media:666.2017.06_-_More_Advanced_Coalescent_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Maximum Likelihood Allele Frequency Estimation]] - [[Media:666.2017.07 - Maximum Likelihood Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the E-M Algorithm]] - [[Media:666.2017.08 - The E-M Algorithm.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Estimation]] - [[Media:666.2017.09_-_Haplotype_Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Haplotype Estimation]] - [[Media:666.2017.10_-_Advanced_Haplotyping.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Power of Genomewide Association Studies]] - [[Media:666.2010.01_-_Power_of_Genomewide_Studies.pdf ‎|PDF]]&lt;br /&gt;
&lt;br /&gt;
Biostatistics 666: Kinship Coefficients and Relatedness - [[Media:666.2017.12_-_Kinship_Coefficients.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Variance Component Analyses]] - [[Media:666.2017.13_-_Variance_Component_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Analysis in Sibling Pairs]] - [[Media:666.2017.14_-_Affected_Sibling_Pairs.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Genotype Imputation]] - [[Media:666.2017.15_-_Genotype_Imputation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Whole Genome Sequencing]] - [[Media:666.2017.16_-_Introduction_to_Sequencing.pptx|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Functional Genomics (Guest lecture)]] - [[Media:Biostat666FunctionalGenomics_v4(2017).pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
Lectures below have not yet been updated for 2017.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Association Tests]] - [[Media:666.12.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Association Tests in Structured Populations]] - [[Media:666.13.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Multipoint Analysis in Sibling Pairs]] - [[Media:666.16.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Relationship Checking]] - [[Media:666.17.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Low Pass Sequence Data]] - [[Media:666.2010.05.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Copy Number Using Sequence Data]] - [[Media:666.2011.02.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to De Novo Assembly]] - [[Media:666.2012.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Rare Variant Burden Tests]] - [[Media:666.2011.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Likelihood Calculations for Large Pedigrees]] - [[Media:666.24.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Older Lectures ==&lt;br /&gt;
&lt;br /&gt;
These lectures were used in previous years, but have been replaced in the most recent edition of the course.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introductory Lecture]] - [[Media:666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Course Introduction and Hardy Weinberg Equilibrium]] - [[Media:2011.666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Changing Population Size]] - [[Media:666.07.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Computation with the Coalescent]] - [[Media:666.08.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Copy Number Variation]] - [[Media:666.Zollner.CNV.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Tests for Pairs of Individuals]] - [[Media:666.14.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Possible Triangle Constraint]] - [[Media:666.15.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Lander-Green Algorithm]] - [[Media:666.18.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Applications of the Lander-Green Algorithm]] - [[Media:666.19.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Problem Sets ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet 1.pdf|Problem Set 1]] -- [[Media:666.Worksheet_1.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_2.pdf|Problem Set 2]] -- [[Media:666.Worksheet_2.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_3.pdf|Problem Set 3]] -- [[Media:666.Worksheet_3.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_4.pdf|Problem Set 4]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_5.pdf|Problem Set 5]]&lt;br /&gt;
&lt;br /&gt;
== Sample Mid-Term ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.SampleMidterm.pdf|Sample Mid-Term]]&lt;br /&gt;
&lt;br /&gt;
== Office Hours ==&lt;br /&gt;
&lt;br /&gt;
For the 2017 Fall Term, office hours are tentatively schedule for Friday afternoons at 3pm. I will provide free coffee to anyone who turns up.&lt;br /&gt;
&lt;br /&gt;
== Standards of Academic Conduct ==&lt;br /&gt;
&lt;br /&gt;
The following is an extract from the School of Public Health&#039;s Student Code of Conduct [http://www.sph.umich.edu/academics/policies/conduct.html]:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Student academic misconduct includes behavior involving plagiarism, cheating, fabrication, falsification of records or official documents, intentional misuse of equipment or materials, and aiding and abetting the perpetration of such acts. The preparation of reports, papers, and examinations, assigned on an individual basis, must represent each student’s own effort. Reference sources should be indicated clearly. The use of assistance from other students or aids of any kind during a written examination, except when the use of books or notes has been approved by an instructor, is a violation of the standard of academic conduct.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
In the context of this course, any work you hand-in should be your own and any material that is a transcript (or interpreted transcript) of work by others must be clearly labeled as such.&lt;br /&gt;
&lt;br /&gt;
== Course History ==&lt;br /&gt;
&lt;br /&gt;
This course was started by Ken Lange and Mike Boehnke and is typically taught every year. &lt;br /&gt;
&lt;br /&gt;
Goncalo Abecasis taught it in the following academic years:&lt;br /&gt;
&lt;br /&gt;
* 2001/2002 (jointly with [http://www.unm.edu/~anthro/people_faculty_jeff_long.html Jeff Long], who is now at the University of New Mexico)&lt;br /&gt;
* 2002/2003&lt;br /&gt;
* 2003/2004&lt;br /&gt;
* 2004/2005&lt;br /&gt;
* 2005/2006&lt;br /&gt;
* 2006/2007&lt;br /&gt;
* 2009/2010 &lt;br /&gt;
* 2010/2011&lt;br /&gt;
* 2011/2012&lt;br /&gt;
* 2012/2013&lt;br /&gt;
* 2017/2018&lt;br /&gt;
&lt;br /&gt;
He last taught it in the 2017/2018 academic year. For previous course notes, see [[http://csg.sph.umich.edu/abecasis/class Goncalo&#039;s older class notes]].&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14975</id>
		<title>Biostatistics 666: Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14975"/>
		<updated>2017-11-29T20:02:44Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Class Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Objective ==&lt;br /&gt;
&lt;br /&gt;
Gene mapping studies study the relationship between genetic variation and susceptibility to human disease. These studies can be used to elucidate the biochemical basis of medically interesting traits leading to knowledge that will, ultimately, help us improve treatment and management of human disease. Biostatistics 666 is a Masters level course that introduces many of the concepts, statistical models and numerical methods useful for studies.&lt;br /&gt;
&lt;br /&gt;
For additional information, see also [[Biostatistics 666: Core Competencies|Core Competencies in Biostatistics Program covered by this course]].&lt;br /&gt;
&lt;br /&gt;
== Target Audience ==&lt;br /&gt;
&lt;br /&gt;
Students in Biostatistics 666 should be comfortable with simple algebra and, ideally, have previous exposure to maximum likelihood. Previous knowledge of Genetics is helpful, but not required. Most students registering for the course are Master or Doctoral students in Human Genetics, Bioinformatics, Statistics or Biostatistics.&lt;br /&gt;
&lt;br /&gt;
== Scheduling ==&lt;br /&gt;
&lt;br /&gt;
Final grades will take into account performance in written in-class assessments, take home problem sets, as well as participation in class and any class projects.&lt;br /&gt;
&lt;br /&gt;
== Class Notes ==&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Contemporary Human Genetics]] - [[Media:666.2017.01 - Introductory Lecture.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Hardy-Weinberg Equilibrium]] - [[Media:666.2017.02 - Hardy Weinberg Equilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Disequilibrium]] - [[Media:666.2017.03 - Linkage Disequilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the Coalescent]] - [[Media:666.04 - The Coalescent.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Variation in the Coalescent]] - [[Media:666.2017.05_-_Coalescent_and_Allele_Counts.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Recombination and Migration in the Coalescent]] - [[Media:666.2017.06_-_More_Advanced_Coalescent_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Maximum Likelihood Allele Frequency Estimation]] - [[Media:666.2017.07 - Maximum Likelihood Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the E-M Algorithm]] - [[Media:666.2017.08 - The E-M Algorithm.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Estimation]] - [[Media:666.2017.09_-_Haplotype_Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Haplotype Estimation]] - [[Media:666.2017.10_-_Advanced_Haplotyping.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Power of Genomewide Association Studies]] - [[Media:666.2010.01_-_Power_of_Genomewide_Studies.pdf ‎|PDF]]&lt;br /&gt;
&lt;br /&gt;
Biostatistics 666: Kinship Coefficients and Relatedness - [[Media:666.2017.12_-_Kinship_Coefficients.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Variance Component Analyses]] - [[Media:666.2017.13_-_Variance_Component_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Analysis in Sibling Pairs]] - [[Media:666.2017.14_-_Affected_Sibling_Pairs.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Genotype Imputation]] - [[Media:666.2017.15_-_Genotype_Imputation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Whole Genome Sequencing]] - [[Media:666.2017.16_-_Introduction_to_Sequencing.pptx|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Functional Genomics]] - [[Media:Biostat666FunctionalGenomics_v4(2017).pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
Lectures below have not yet been updated for 2017.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Association Tests]] - [[Media:666.12.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Association Tests in Structured Populations]] - [[Media:666.13.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Multipoint Analysis in Sibling Pairs]] - [[Media:666.16.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Relationship Checking]] - [[Media:666.17.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Low Pass Sequence Data]] - [[Media:666.2010.05.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Copy Number Using Sequence Data]] - [[Media:666.2011.02.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to De Novo Assembly]] - [[Media:666.2012.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Rare Variant Burden Tests]] - [[Media:666.2011.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Likelihood Calculations for Large Pedigrees]] - [[Media:666.24.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Older Lectures ==&lt;br /&gt;
&lt;br /&gt;
These lectures were used in previous years, but have been replaced in the most recent edition of the course.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introductory Lecture]] - [[Media:666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Course Introduction and Hardy Weinberg Equilibrium]] - [[Media:2011.666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Changing Population Size]] - [[Media:666.07.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Computation with the Coalescent]] - [[Media:666.08.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Copy Number Variation]] - [[Media:666.Zollner.CNV.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Tests for Pairs of Individuals]] - [[Media:666.14.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Possible Triangle Constraint]] - [[Media:666.15.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Lander-Green Algorithm]] - [[Media:666.18.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Applications of the Lander-Green Algorithm]] - [[Media:666.19.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Problem Sets ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet 1.pdf|Problem Set 1]] -- [[Media:666.Worksheet_1.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_2.pdf|Problem Set 2]] -- [[Media:666.Worksheet_2.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_3.pdf|Problem Set 3]] -- [[Media:666.Worksheet_3.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_4.pdf|Problem Set 4]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_5.pdf|Problem Set 5]]&lt;br /&gt;
&lt;br /&gt;
== Sample Mid-Term ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.SampleMidterm.pdf|Sample Mid-Term]]&lt;br /&gt;
&lt;br /&gt;
== Office Hours ==&lt;br /&gt;
&lt;br /&gt;
For the 2017 Fall Term, office hours are tentatively schedule for Friday afternoons at 3pm. I will provide free coffee to anyone who turns up.&lt;br /&gt;
&lt;br /&gt;
== Standards of Academic Conduct ==&lt;br /&gt;
&lt;br /&gt;
The following is an extract from the School of Public Health&#039;s Student Code of Conduct [http://www.sph.umich.edu/academics/policies/conduct.html]:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Student academic misconduct includes behavior involving plagiarism, cheating, fabrication, falsification of records or official documents, intentional misuse of equipment or materials, and aiding and abetting the perpetration of such acts. The preparation of reports, papers, and examinations, assigned on an individual basis, must represent each student’s own effort. Reference sources should be indicated clearly. The use of assistance from other students or aids of any kind during a written examination, except when the use of books or notes has been approved by an instructor, is a violation of the standard of academic conduct.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
In the context of this course, any work you hand-in should be your own and any material that is a transcript (or interpreted transcript) of work by others must be clearly labeled as such.&lt;br /&gt;
&lt;br /&gt;
== Course History ==&lt;br /&gt;
&lt;br /&gt;
This course was started by Ken Lange and Mike Boehnke and is typically taught every year. &lt;br /&gt;
&lt;br /&gt;
Goncalo Abecasis taught it in the following academic years:&lt;br /&gt;
&lt;br /&gt;
* 2001/2002 (jointly with [http://www.unm.edu/~anthro/people_faculty_jeff_long.html Jeff Long], who is now at the University of New Mexico)&lt;br /&gt;
* 2002/2003&lt;br /&gt;
* 2003/2004&lt;br /&gt;
* 2004/2005&lt;br /&gt;
* 2005/2006&lt;br /&gt;
* 2006/2007&lt;br /&gt;
* 2009/2010 &lt;br /&gt;
* 2010/2011&lt;br /&gt;
* 2011/2012&lt;br /&gt;
* 2012/2013&lt;br /&gt;
* 2017/2018&lt;br /&gt;
&lt;br /&gt;
He last taught it in the 2017/2018 academic year. For previous course notes, see [[http://csg.sph.umich.edu/abecasis/class Goncalo&#039;s older class notes]].&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14974</id>
		<title>Biostatistics 666: Main Page</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=Biostatistics_666:_Main_Page&amp;diff=14974"/>
		<updated>2017-11-29T20:02:19Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Class Notes */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Objective ==&lt;br /&gt;
&lt;br /&gt;
Gene mapping studies study the relationship between genetic variation and susceptibility to human disease. These studies can be used to elucidate the biochemical basis of medically interesting traits leading to knowledge that will, ultimately, help us improve treatment and management of human disease. Biostatistics 666 is a Masters level course that introduces many of the concepts, statistical models and numerical methods useful for studies.&lt;br /&gt;
&lt;br /&gt;
For additional information, see also [[Biostatistics 666: Core Competencies|Core Competencies in Biostatistics Program covered by this course]].&lt;br /&gt;
&lt;br /&gt;
== Target Audience ==&lt;br /&gt;
&lt;br /&gt;
Students in Biostatistics 666 should be comfortable with simple algebra and, ideally, have previous exposure to maximum likelihood. Previous knowledge of Genetics is helpful, but not required. Most students registering for the course are Master or Doctoral students in Human Genetics, Bioinformatics, Statistics or Biostatistics.&lt;br /&gt;
&lt;br /&gt;
== Scheduling ==&lt;br /&gt;
&lt;br /&gt;
Final grades will take into account performance in written in-class assessments, take home problem sets, as well as participation in class and any class projects.&lt;br /&gt;
&lt;br /&gt;
== Class Notes ==&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Contemporary Human Genetics]] - [[Media:666.2017.01 - Introductory Lecture.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Hardy-Weinberg Equilibrium]] - [[Media:666.2017.02 - Hardy Weinberg Equilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Disequilibrium]] - [[Media:666.2017.03 - Linkage Disequilibrium.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the Coalescent]] - [[Media:666.04 - The Coalescent.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Variation in the Coalescent]] - [[Media:666.2017.05_-_Coalescent_and_Allele_Counts.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Modeling Recombination and Migration in the Coalescent]] - [[Media:666.2017.06_-_More_Advanced_Coalescent_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Maximum Likelihood Allele Frequency Estimation]] - [[Media:666.2017.07 - Maximum Likelihood Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to the E-M Algorithm]] - [[Media:666.2017.08 - The E-M Algorithm.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Estimation]] - [[Media:666.2017.09_-_Haplotype_Estimation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Haplotype Estimation]] - [[Media:666.2017.10_-_Advanced_Haplotyping.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Power of Genomewide Association Studies]] - [[Media:666.2010.01_-_Power_of_Genomewide_Studies.pdf ‎|PDF]]&lt;br /&gt;
&lt;br /&gt;
Biostatistics 666: Kinship Coefficients and Relatedness - [[Media:666.2017.12_-_Kinship_Coefficients.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Variance Component Analyses]] - [[Media:666.2017.13_-_Variance_Component_Models.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Analysis in Sibling Pairs]] - [[Media:666.2017.14_-_Affected_Sibling_Pairs.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Genotype Imputation]] - [[Media:666.2017.15_-_Genotype_Imputation.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Whole Genome Sequencing]] - [[Media:666.2017.16_-_Introduction_to_Sequencing.pptx|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Functional Genomics]] - [[Media:Biostat666FunctionalGenomics_v4(2017).pdf]]&lt;br /&gt;
&lt;br /&gt;
Lectures below have not yet been updated for 2017.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Haplotype Association Tests]] - [[Media:666.12.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Association Tests in Structured Populations]] - [[Media:666.13.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Multipoint Analysis in Sibling Pairs]] - [[Media:666.16.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Relationship Checking]] - [[Media:666.17.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Low Pass Sequence Data]] - [[Media:666.2010.05.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Analysis of Copy Number Using Sequence Data]] - [[Media:666.2011.02.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introduction to De Novo Assembly]] - [[Media:666.2012.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Rare Variant Burden Tests]] - [[Media:666.2011.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Likelihood Calculations for Large Pedigrees]] - [[Media:666.24.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Older Lectures ==&lt;br /&gt;
&lt;br /&gt;
These lectures were used in previous years, but have been replaced in the most recent edition of the course.&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Introductory Lecture]] - [[Media:666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Course Introduction and Hardy Weinberg Equilibrium]] - [[Media:2011.666.01.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Changing Population Size]] - [[Media:666.07.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Advanced Coalescent, Computation with the Coalescent]] - [[Media:666.08.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Copy Number Variation]] - [[Media:666.Zollner.CNV.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Linkage Tests for Pairs of Individuals]] - [[Media:666.14.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Possible Triangle Constraint]] - [[Media:666.15.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: The Lander-Green Algorithm]] - [[Media:666.18.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
[[Biostatistics 666: Applications of the Lander-Green Algorithm]] - [[Media:666.19.pdf|PDF]]&lt;br /&gt;
&lt;br /&gt;
== Problem Sets ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet 1.pdf|Problem Set 1]] -- [[Media:666.Worksheet_1.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_2.pdf|Problem Set 2]] -- [[Media:666.Worksheet_2.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_3.pdf|Problem Set 3]] -- [[Media:666.Worksheet_3.Solution.pdf|Solution]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_4.pdf|Problem Set 4]]&lt;br /&gt;
&lt;br /&gt;
[[Media:666.Worksheet_5.pdf|Problem Set 5]]&lt;br /&gt;
&lt;br /&gt;
== Sample Mid-Term ==&lt;br /&gt;
&lt;br /&gt;
[[Media:666.SampleMidterm.pdf|Sample Mid-Term]]&lt;br /&gt;
&lt;br /&gt;
== Office Hours ==&lt;br /&gt;
&lt;br /&gt;
For the 2017 Fall Term, office hours are tentatively schedule for Friday afternoons at 3pm. I will provide free coffee to anyone who turns up.&lt;br /&gt;
&lt;br /&gt;
== Standards of Academic Conduct ==&lt;br /&gt;
&lt;br /&gt;
The following is an extract from the School of Public Health&#039;s Student Code of Conduct [http://www.sph.umich.edu/academics/policies/conduct.html]:&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;Student academic misconduct includes behavior involving plagiarism, cheating, fabrication, falsification of records or official documents, intentional misuse of equipment or materials, and aiding and abetting the perpetration of such acts. The preparation of reports, papers, and examinations, assigned on an individual basis, must represent each student’s own effort. Reference sources should be indicated clearly. The use of assistance from other students or aids of any kind during a written examination, except when the use of books or notes has been approved by an instructor, is a violation of the standard of academic conduct.&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
In the context of this course, any work you hand-in should be your own and any material that is a transcript (or interpreted transcript) of work by others must be clearly labeled as such.&lt;br /&gt;
&lt;br /&gt;
== Course History ==&lt;br /&gt;
&lt;br /&gt;
This course was started by Ken Lange and Mike Boehnke and is typically taught every year. &lt;br /&gt;
&lt;br /&gt;
Goncalo Abecasis taught it in the following academic years:&lt;br /&gt;
&lt;br /&gt;
* 2001/2002 (jointly with [http://www.unm.edu/~anthro/people_faculty_jeff_long.html Jeff Long], who is now at the University of New Mexico)&lt;br /&gt;
* 2002/2003&lt;br /&gt;
* 2003/2004&lt;br /&gt;
* 2004/2005&lt;br /&gt;
* 2005/2006&lt;br /&gt;
* 2006/2007&lt;br /&gt;
* 2009/2010 &lt;br /&gt;
* 2010/2011&lt;br /&gt;
* 2011/2012&lt;br /&gt;
* 2012/2013&lt;br /&gt;
* 2017/2018&lt;br /&gt;
&lt;br /&gt;
He last taught it in the 2017/2018 academic year. For previous course notes, see [[http://csg.sph.umich.edu/abecasis/class Goncalo&#039;s older class notes]].&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=File:Biostat666FunctionalGenomics_v4(2017).pdf&amp;diff=14973</id>
		<title>File:Biostat666FunctionalGenomics v4(2017).pdf</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=File:Biostat666FunctionalGenomics_v4(2017).pdf&amp;diff=14973"/>
		<updated>2017-11-29T20:01:42Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14679</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14679"/>
		<updated>2017-04-04T19:21:35Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. &amp;lt;b&amp;gt;--&amp;gt; Complete &amp;lt;/b&amp;gt;&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there) &amp;lt;b&amp;gt;--&amp;gt; GWASs on 120 Visit 1 traits Complete &amp;lt;/b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14678</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14678"/>
		<updated>2017-04-04T19:21:20Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. &amp;lt;b&amp;gt;--&amp;gt; Complete&amp;lt;\b&amp;gt;&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there) &amp;lt;b&amp;gt;--&amp;gt; GWASs on 120 Visit 1 traits Complete&amp;lt;\b&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14677</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14677"/>
		<updated>2017-04-04T19:20:33Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. /b/--&amp;gt; Complete/\b/&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there) *--&amp;gt; GWASs on 120 Visit 1 traits Complete*&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14676</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14676"/>
		<updated>2017-04-04T19:19:58Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Future Directions */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory. *--&amp;gt; Complete*&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there) *--&amp;gt; GWASs on 120 Visit 1 traits Complete*&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14373</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14373"/>
		<updated>2016-08-04T12:51:26Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: &lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory.&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
* Merged files: /net/wonderland/home/csidore/Epacts_vcfs/Michelle_panel/newpanel.chr*.vcf.gz&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
	<entry>
		<id>http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14372</id>
		<title>SardiNIA</title>
		<link rel="alternate" type="text/html" href="http://genome.sph.umich.edu/w/index.php?title=SardiNIA&amp;diff=14372"/>
		<updated>2016-08-03T02:55:35Z</updated>

		<summary type="html">&lt;p&gt;Sarahgag: /* Status as of July 2016 */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;== Project Leaders ==&lt;br /&gt;
&lt;br /&gt;
* David Schlessinger (National Institutes on Aging, Baltimore)&lt;br /&gt;
* Manuela Uda (National Research Center, Cagliari, Italy)&lt;br /&gt;
* Goncalo Abecasis (University of Michigan, Ann Arbor)&lt;br /&gt;
&lt;br /&gt;
== Genomewide Association Study ==&lt;br /&gt;
&lt;br /&gt;
We carried out an initial genomewide association by genotyping 1405 samples with Affymetrix 500K SNP arrays. Then, because SardiNIA samples are closely related to each other, we were able to use these genotypes to impute the genomes of many close relatives who were genotyped with Affymetrix 10K SNP arrays. Typically, our SardiNIA GWAS analysis thus include approximately 4300 genotyped or imputed individuals. To learn more about the approach see Chen and Abecasis (2007) and Scuteri et al (2007).&lt;br /&gt;
&lt;br /&gt;
=== Planned Updates ===&lt;br /&gt;
&lt;br /&gt;
Several ongoing experiments are expected to gradually improve our GWAS data. First, we expect to integrate Affymetrix 1M SNP chip genotypes into our analyses. Second, we plan to genotype all sampled individuals with the Metabochip, which includes 200,000 SNPs. This will enable us to fill in genotypes for relatives of those genotyped with denser arrays more accurately, because we will more precisely identify shared stretches of chromosome.&lt;br /&gt;
&lt;br /&gt;
== Medical Sequencing Project ==&lt;br /&gt;
&lt;br /&gt;
We are sequencing the genomes of 1,000 individuals to learn about the genetics of blood lipid levels and personality.&lt;br /&gt;
&lt;br /&gt;
== Status as of July 2016 ==&lt;br /&gt;
&lt;br /&gt;
=== Locations of Files for Current Data Freeze of 3839 Samples===&lt;br /&gt;
&lt;br /&gt;
NOTICE: We identified late in the process that two of the samples (22855 and 22358)  were actually the same individual. They both should be the same individual 22855. Therefore, there are 3840 sample IDs in each of the files below, but only 22855 should move on to later processes. In future data freezes with this data, these two sequencing sets should be merged into a single 22855 individual.&lt;br /&gt;
&lt;br /&gt;
* List of &#039;&#039;&#039;Sample Numbers&#039;&#039;&#039;&lt;br /&gt;
** The following file contains three columns: [SampleID used in these analyses] [ID supplied by CSCT or Sardinia or other project] [Sequencing core ID (if different)]:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/sampleIDConversion.txt&lt;br /&gt;
&lt;br /&gt;
* List of paths to &#039;&#039;&#039;BAMs&#039;&#039;&#039; used in this data freeze (Index file)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_index_20150504.index&lt;br /&gt;
** 401 of these samples have some new BAM contribution since the previous data freeze... their BAMs can be found here: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/newsamples.index&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Pedigree&#039;&#039;&#039; (Not too helpful -- used for SNPCall)&lt;br /&gt;
** All Samples: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/FINAL_ped_20150510.ped&lt;br /&gt;
** Disjoint Trios: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/Pedigree_Fall15_DataFreeze_Triplets.ped&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;QC&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCStats.txt&lt;br /&gt;
** For a list of paths to all of the separate QPLOT files for each sample, see the file: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/GeneratingQCDistributions/QCFileListFinal.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;SNPCall&#039;&#039;&#039; Results (Produced with Gotcloud SNPCall and phased using Beagle4)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPCALL/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;IndelCall&#039;&#039;&#039; Results (Produced with Gotcloud Indel)&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/indel/final/all.genotypes.sites.vcf.gz&lt;br /&gt;
** SNPEff and VEP declarations of Indel types can be found:&lt;br /&gt;
*** SNPEff: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/snpEff/*&lt;br /&gt;
*** VEP: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/VEP/*&lt;br /&gt;
** We used an &#039;&#039;&#039;Indel filtering strategy&#039;&#039;&#039; composed of many levels. &lt;br /&gt;
**# AC must be 1 or greater -- eliminate Indels with AC=0&lt;br /&gt;
**# the Indel should overlap with a VNTR region or overlap with another Indel. We used Adrian&#039;s annotate indels program to identify such overlaps. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/All.annotated.sites.vcf.gz&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/AnnotateIndels/Overlaps.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative AD field (we define &amp;quot;informative&amp;quot; to mean that the sample has U/(R+A+U)&amp;lt;0.50). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter2_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# at least 50% of the Indels should have informative PL field (we define &amp;quot;informative&amp;quot; to mean that the PL field for the sample is anything BUT ././. or 0/0/0). Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/filter1_output.txt&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/INDEL_filtering_final.txt&lt;br /&gt;
**# the Indel needs to have BF_LRE_LUD (a Bayes factor comparing a. related &amp;amp; HWE to b. unrelated &amp;amp; HWD) &amp;gt; -10. Higher BF_LRE_LUD should indicate a better Indel. We used Hyun&#039;s MiLK program to obtain BF_LRE_LUD values. Results for Indels on all chromosomes can be found here:&lt;br /&gt;
**#* /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/MiLK_filtering/all.genotypes.milk.sites.vcf&lt;br /&gt;
*** Overall results from all of the below filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/FinalIndelFilteringStatistics.txt&lt;br /&gt;
*** VCFs of Indels after filtering can be found here: &lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.*.PASS.vcf.gz&lt;br /&gt;
**** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/INDELCALL/filterIndels/FINAL/all.genotypes.PASS.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Merged Indel VCF with beagles SNP VCF, then sorted to get the following VCFS:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/chr*.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* Indel and SNP VCFs that have been combined AND BEAGLED AGAIN TOGETHER using Beagle4. &#039;&#039;&#039;These are the latest VCFS&#039;&#039;&#039;:&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/SNPINDEL/beagle4/beagle4_chr*/beagle4/chr*/chr*.filtered.PASS.beagled.vcf.gz&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;mtDNA Copy Number&#039;&#039;&#039; Summary&lt;br /&gt;
** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/sardiniaCopyNumber_include.txt&lt;br /&gt;
** By sample copy number results by sample can be found: /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/MTDNA_COPYNUMBER/Coverage_IncludingUnpairedReads/*.CopyNumber.noRand.500000.*.txt&lt;br /&gt;
&lt;br /&gt;
* &#039;&#039;&#039;Filtering Samples&#039;&#039;&#039;: Discrepancy Figures&lt;br /&gt;
** We compared chip data from previous work to the sequencing data produced now in hopes of identifying which samples have reliable sequencing data. This comparison was done for each chromosome separately and then combined into an overall discrepancy figure. Only 3189 of the 3840 samples (3188 of 3839 if you throw out sample with two sets of data) had chip data.&lt;br /&gt;
** By chromosome discrepancy figures for each sample can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/*.diff.discordance_matrix&lt;br /&gt;
** By chromosome discrepancy figures for all samples in one file can be found:&lt;br /&gt;
*** /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt&lt;br /&gt;
&lt;br /&gt;
=== What is Complete ===&lt;br /&gt;
* SNP Call&lt;br /&gt;
** 24,901,469 SNPs passed filters&lt;br /&gt;
*** 16,822,922 are in dbSNP (67.6%)&lt;br /&gt;
*** %Known Ts/Tv - 2.24&lt;br /&gt;
*** %Novel Ts/Tv - 1.95&lt;br /&gt;
* InDel Call&lt;br /&gt;
** 1,194,945 passed filters&lt;br /&gt;
&lt;br /&gt;
=== Future Directions ===&lt;br /&gt;
* NOTE: For sample filtering below, you will need to finish chromosome 1 for me. I have it currently running. Once it is done in a few days, you will need to run the command &#039;python calculateConcordance_onefile.py&#039; while in the &#039;/net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr1/&#039; directory.&lt;br /&gt;
* &#039;&#039;&#039;Sample Filtering&#039;&#039;&#039;&lt;br /&gt;
** We did not do any filtering of samples (based on dupRate, genome coverage, mapping rate, proper paired, mean depth, or any other QPLOT stats) prior to SNP and Indel calling. Because of this, we want to do this filtering now. 3,188 or 3,839 samples have genome chip data from a few years ago. For these, we could look at the non-reference concordance between the chip genotypes and the sequencing genotypes and declare &#039;bad&#039; samples to be those that fall below a certain threshold, such as 98% non-ref concordance. However, since the remaining 651 samples do not have chip data, this is not an option for them. Therefore, we decided on the following strategy instead: &lt;br /&gt;
**# Calculate non-reference concordance for the 3,188 samples that have chip data. &lt;br /&gt;
**# Create a prediction model using QPLOT statistics as predictors of non-reference concordance. Either do so on all of the 3,188 samples and look at R^2 (likely inflated from overfitting) or use cross-validation (test and training set) to give a measure of external predictive power. &lt;br /&gt;
**# If reasonable predictive power/R^2, use the prediction model to estimate the non-reference concordance amongst the 651 samples that do not have chip data. Also use the prediction model to estimate the non-reference concordance among the 3,188 samples that do have chip data.&lt;br /&gt;
**# Set a cut-off for &#039;good&#039; versus &#039;bad&#039; samples based on the estimated non-reference concordance and use it to filter samples.&lt;br /&gt;
** NOTE: The number of positions for which the chip data give 0/0 and sequencing gives 0/0, chip data gives 0/0 and sequencing gives 0/1, chip data gives 0/0 and sequencing gives 1/1, chip gives 0/1 and sequencing gives 0/0, etc. etc. BY chromosome can be found in the files /net/sardinia/progenia/SardiNIA/VariantCalling_20150330/filteringSamples/chr*/all.discordance.summary.txt. These can be used to calculate overall non-reference concordance across all chromosomes. &lt;br /&gt;
* &#039;&#039;&#039;Mitochondrial Depth Analysis&#039;&#039;&#039;&lt;br /&gt;
* &#039;&#039;&#039;Telomere Length Analysis&#039;&#039;&#039;&lt;br /&gt;
** Investigate the associations between telomere length (an indicator of aging) and variants. Likely interesting in Sardinia population because Sardinians have longer lifespans &amp;amp; centenarians.&lt;br /&gt;
* &#039;&#039;&#039;Phenotype Study&#039;&#039;&#039;&lt;br /&gt;
** Likely will not yield much because not many additional samples since Carlo&#039;s last data freeze (3,514 samples there)&lt;br /&gt;
&lt;br /&gt;
== New Updates ==&lt;br /&gt;
* The duplicate person in the sequencing files was removed, and those files were used to impute the phased genotyping files.&lt;br /&gt;
* Imputation was conducted using Minimac3. SNPs were imputed, and also indels were imputed.&lt;br /&gt;
&lt;br /&gt;
== Key References ==&lt;br /&gt;
&lt;br /&gt;
If you are looking to learn about the project, I strongly recommend that you read the following papers &lt;br /&gt;
&lt;br /&gt;
* Pilia G, Chen WM, Scuteri A, Orru M, Albai G, Dei M, Lai S, Usala G, Lai M, Loi P, Mameli C, Vacca L, Deiana M, Olla N, Masala M, Cao A, Najjar SS, Terracciano A, Nedorezov T, Sharov A, Zonderman AB, Abecasis GR, Costa P, Lakatta E and Schlessinger D (2006). Heritability of Cardiovascular and Personality Traits in 6,148 Sardinians. PLoS Genet 2:1207-1223 [[http://www.sph.umich.edu/csg/abecasis/publications/16934002.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Scuteri A, Sanna S, Chen WM, Uda M, Albai G, Strait J, Najjar S, Nagarajah R, Orru M, Usala G, Dei M, Lai S, Maschio A, Busonero F, Mulas A, Ehret GB, Fink AA, Weder A, Cooper R, Galan P, Chakravarti A, Schlessinger D, Cao A, Lakatta E and Abecasis GR (2007). Genome Wide Association Scan shows Genetic Variants in the FTO gene are Associated with Obesity Related Traits PLoS Genetics 3:1200-10 [[http://www.sph.umich.edu/csg/abecasis/publications/PLOS-Obesity-Scan.html Abstract and PDF]]&lt;br /&gt;
&lt;br /&gt;
* Chen WM and Abecasis GR (2007). Family-based association tests for genomewide association scans. Am J Hum Genet 81:913-26 [[http://www.sph.umich.edu/csg/abecasis/publications/17924335.html Abstract and PDF]]&lt;/div&gt;</summary>
		<author><name>Sarahgag</name></author>
	</entry>
</feed>