Most variants identified by GWAS are located outside of coding regions. Is this because many GWAS hits are simply tagging the true causal variants, which were not directly identified by the GWAS?
More generally, I am interested in whether there is any biological relationship between the tag SNP and the causal SNP. For example, could a tag SNP itself lie within a regulatory element in a non-coding region and influence the function or expression of a coding causal variant (or gene) within the same LD block? Or is the tag SNP typically just a statistical proxy with no functional role of its own?
Great questions. I think most GWAS hits falling outside coding regions likely reflects something real: individual differences in complex traits seem to be driven more by when and how much genes are expressed than by differences in protein structure itself.
As for whether the tag SNP itself has a functional role: usually we just don’t know. The tag SNP is typically the variant that was measured and happens to be correlated with whatever causal variant is nearby. They travel together in the genome due to LD, making them statistically interchangeable. Fine-mapping and functional studies can sometimes identify the likely causal variant, but more about that next week!