RAREMETALWORKER METHOD: Difference between revisions

From Genome Analysis Wiki
Jump to navigationJump to search
Shuang Feng (talk | contribs)
Shuang Feng (talk | contribs)
Line 17: Line 17:
== Single Variant Score Tests ==
== Single Variant Score Tests ==
Our single variant association test is the score test using linear mixed model, treating single variants as fixed effects. The alternative model is:
Our single variant association test is the score test using linear mixed model, treating single variants as fixed effects. The alternative model is:


<math> \mathbf{y}=\mathbf{X}\boldsymbol{\beta}+\gamma_i(\mathbf{G_i}-\bar{\mathbf{G_i}})+\mathbf{g}+\boldsymbol{\varepsilon} </math>.
<math> \mathbf{y}=\mathbf{X}\boldsymbol{\beta}+\gamma_i(\mathbf{G_i}-\bar{\mathbf{G_i}})+\mathbf{g}+\boldsymbol{\varepsilon} </math>.


In this model, the scalar parameter <math>\gamma_i</math> is to measure the additive genetic effect of the <math>i^{th}</math> variant. As usual, the score statistic for testing <math>H_0:\gamma_i=0</math> is:
In this model, the scalar parameter <math>\gamma_i</math> is to measure the additive genetic effect of the <math>i^{th}</math> variant. As usual, the score statistic for testing <math>H_0:\gamma_i=0</math> is:


<math> U_i=(\mathbf{G_i}-\mathbf{\bar{G_i}} )^T \boldsymbol{\Omega}^(-1)(\mathbf{y}-\mathbf{X}\boldsymbol{\beta}) </math>
<math> U_i=(\mathbf{G_i}-\mathbf{\bar{G_i}} )^T \boldsymbol{\Omega}^(-1)(\mathbf{y}-\mathbf{X}\boldsymbol{\beta}) </math>

Revision as of 11:38, 11 March 2014

Brief Introduction

RAREMETALWORKER(RMW) generates single variant association results from score test, together with summary statistics and covariance matrices of the score statistics.

In the following sections, we will go through the methods behind RWM including statistic model, handling sample relatedness, and definitions of statistics in the output.

Modeling Relatedness

we use a variance component model to handle familial relationships. In a sample of n individuals, we model the observed phenotype vector (𝐲) as a sum of covariate effects (specified by a design matrix 𝐗 and a vector of covariate effects β), additive genetic effects (modeled in vector 𝐠) and non-shared environmental effects (modeled in vector ε). Thus the null model is:

𝐲=𝐗β+𝐠+ε


We assume that genetic effects are normally distributed, with mean 𝟎 and covariance 𝐊σg2 where the matrix 𝐊 summarizes kinship coefficients between sampled individuals and σg2 is a positive scalar describing the genetic contribution to the overall variance. We assume that non-shared environmental effects are normally distributed with mean 𝟎 and covariance 𝐈σe2, where 𝐈 is the identity matrix.

To estimate 𝐊, we either use known pedigree structure to define 𝐊 or else use the empirical estimator 𝐊=1l∑i=1l(Gi−2fi𝟏)(Gi−2fi𝟏)4fi(1−fi), where l is the count of variants, Gi and fi are the genotype vector and estimated allele frequency for the ith variant, respectively. Each element in Gi encodes the minor allele count for one individual. Model parameters β^, σg2^ and σe2^, are estimated using maximum likelihood and the efficient algorithm described in Lippert et. al. For convenience, let the estimated covariance matrix of 𝐲 be Ω=2σg2𝐊+σe2𝐈.

Single Variant Score Tests

Our single variant association test is the score test using linear mixed model, treating single variants as fixed effects. The alternative model is:

𝐲=𝐗β+γi(G𝐢−G𝐢¯)+𝐠+ε.

In this model, the scalar parameter γi is to measure the additive genetic effect of the ith variant. As usual, the score statistic for testing H0:γi=0 is:

Ui=(G𝐢−Gi¯)TΩ(−1)(𝐲−𝐗β)

And the variance-covariance matrix of these statistics is (see Appendix A for details):

                       V=(G-G ̅ )^T (Ω ̂^(-1)-Ω ̂^(-1) X(X^T Ω ̂^(-1) X)^(-1) X^T Ω ̂^(-1) )(G-G ̅).

Under the null, test statistics T_i=(U_i^2)/V_ii is asymptotically distributed as chi-squared with one degree of freedom.

Summary Statistics

Covariance Matrices

Chromosome X