Showing posts with label GWAS. Show all posts
Showing posts with label GWAS. Show all posts

Wednesday, June 19, 2013

GTEx Community Meeting - notes

Yesterday I attended the GTEx (Genotype-Tissue Expression project) Community Meeting held at the Broad Institute in Cambridge, MA, who hosts the GTEx portal. This meeting offered opportunities to GTEx researchers and those scientists not part of this NIH Common Fund program to engage in dialog regarding new aspects of the GTEx project. An overview of the project is here. A main impetus for GTEx is many GWAS signals linking genotype to disease phenotype have a role in regulation of gene expression. Thus, learning more about gene expression can assist in the interpretation of GWAS results. Below are my notes from this one-day meeting.

Simona Volpi – NIH. See commonfund.nih.gov/gtex for details. Samples are from biobanked tissues. This may make it difficult to engage in challenge experiments. Goal is to establish a database of genotype-gene expression relationships. Goal is to collect from 900 donors. They use PAXgene, alcohol-based fixative for the tissues in 0.2 – 0.5 gram aliquots. Tissue processing includes histopathologic review, FPPE paraffin embedding, RNA extraction.  A U01 RFA seeking application to propose ways to enhance GTEx is being formed and will soon be announced.
  • BMI range for donors is greater than 18.5 and less than 35.
  • Cause of death of donors: 34% cerebrovascular, 13% cardiac, 22% respiratory, 21% from accidents (transportation and non-transportation).
  • RFA-RM-12-009 – eGTEx RFA from NIH – perhaps this is closed. They are working on liberalizing the access policy.
  • Data will be housed in dbGaP. Need to apply for access to get a lot of info, but some basic info is available at the GTEx portal at the Broad.

Kristin Ardlie – LDACC – Laboratory, data analysis and coordinating center. There are 47 tissues: 35 PAXgene tissues + blood + 11 frozen brain sub-regions. Blood is collected and processed pre-mortem. Goal by Jan 2014: 9534 RNA samples from 430+ donors. Goal for RNA-seq is 50 million aligned reads and no less than 15 million. Input is 200 ng total RNA of RNAs w a RIN of 6.0 or higher. RIN = RNA integrity score. Skeletal muscle and lung have high RINs, but pancreas, adipose and other enzyme-rich tissues have lower RINs and quicker decay post-ischemic time (or post-mortem interval). 

Manolis Dermitzakis uses a FDR of 5% based on Storey to identify eQTL SNPs. They use a 1 MB window around TSS but also a 100 kbp window. Using the established 15 PEER factors (to account for population ancestry structure), about 800 or so eQTL for adipose are expected. Overall, there are about 6200 eQTL genes for subcutaneous adipose. Skeletal muscle gets toward 7800 eQTL genes. He urges caution when seeing overlap between eQTL and GWAS hits because these are almost certain as data increase in volume, so take into consideration the effect size (20% increase of mRNA and protein is not the same in terms of biological consequences as a 2-fold increase). Can they borrow power from a “related” tissue to look at a specific tissue? Estimates of tissue sharing are high for adipose; nerve is best at 0.92, artery is 0.91. About 0.56 is the probability for a SNP to be active in all nine tissues. Probability of being active in a single tissue is just 0.03. 

Roderic Guigo. There are about 15000 to 20000 expressed genes in most tissues, with blood less and testis more. Most tissues have 2000 to 3000 expressed lncRNAs, with testis having many more. 3820 genes are expressed in only one tissue. Most genes express about half of annotated isoforms in a given tissue. With two isoforms, the major isoform dominates with 90% of expression of that gene, ~40-50% of expression comes from the dominant mRNA isoform when there are 5 or so isoforms. Splicing QTL, SNPs affecting splicing pattern of the gene, but may or may not affect expression. Their group had to develop software to detect these in the RNA-seq data and also to account for the complex phenotype: isoforms and expression counts.

Mike Weale. Using arrays for eQTL studies can lead to generation of false positives. See Ramasamy et al 2013 Nucl Acids Res for an analysis of probe-dropping with better reference data. Something like 5.5% of probes map to ref genome SNPs but account for 90% of brain eQTL hits. See Trabzuni Hum MolGenet 21:4094 for the famous example of a false positive eQTL for MAPT. Their PiP finder tool is at bitly.com/pipfinder.

Barbara Englehardt. Replication of cis- and trans-eQTL across cell types. Her goal is to predict cis-eQTL as functional SNPs. This method is soon to be out in PLoS Genetics. Study size and replicate arrays account for >95% of the variability in fraction of genes showing an eQTL. They found no false positives when using replicate expression arrays, but false negatives persisted. Replcation strengthened cis-eQTL discovery.

Yaniv Erlich. STRs short tandem repeats of 2-6 bp. Sometimes these occur in promoter regions. Using 1000G data, they saw about 80% of STRs are polymorphic with MAF >1%. Many more STRs in introns than in exons and loss of heterozygozity with populations not of African origin. Looking for effects of STR variation on gene expression, they found 2673 eSTRs, but with replication in orthogonal data (use arrays when original data came from RNA-seq, eg) they found 81% of eSTRs showed the same direction of effect. Were they tagging SNPs? 77% of eSTR had same slope as before when conditioned on most likely cis-eQTL SNP, meaning that they were not tagging SNPs in most cases. They do not see any dose-dependence with the STRs, meaning a length effect on the expression effect. He speculates that the STRs create Z-DNA.

João Fadista. Prediction on individual level genotypes based on solely on GTEx gene expression. Assuming each gene has at least one cis-eQTL and 20,000 genes, there will be 320,000 combinations and this combination or pattern could be used to predict a person’s genotype. His examples will come from the Nordic Network for Islet Transplantation and includes other tissues/organs. 89 islet donors, 61% men, 5.8 ± 0.9% HbA1c levels. Found 136 eQTL. Only 22 of these had genotype prediction data in all GTEx samples, but could be sufficient to predict genotype: 322 is greater than current world population by more than 4-fold. See also work by Eric Schadt (Nat Genet, Bayesian method to predict individual SNP…) on their replication of liver and adipose eQTL. Why do this and jeopardize GTEx, asks M. Dermitzakis, and JF replies that a blood gene expression test combined with eQTL data can predict disease. M. Dermitzakis states that heritability of gene expression is about 0.3 and so predictions of tissue-specific gene expression will be limited.

Stephen Montgomery. He looks at allele-specific expression and GTEx data to discover eQTL. See Wei Sun on TReCASE tool in Biometrics from 2012. Even with 15 million RNA-seq reads, 52.2% of sites have depth less than 30 and so have lowered ability to confidently label as ASE (allele specific expression). Allelic ratios of the mRNAs are highly heritable, as seen by looking at a 3-generation, 17-member family. Such is seen across low and moderately expressed genes. Can ASEs say anything about deleterious variants? Looked at 10 tissues in one 25-yo Chinese male, and looked at deleterious sties (50), loss-of-function (LoF, stop-gain) sites (74) and ? (very few). The LoF variant is lowly expressed across the tissues, as reported by Dan MacArthur. 

Tuuli Lappalainan. Uses December release of GTEx data to look at ASEs. ASEs can recover from under-powered studies to identify eQTL. Master data to be released with upcoming paper: ≥ 8 RNA-seq reads over a site. Most analyses sampling are done to exactly 30 reads in order to avoid coverage issues. Note: only relatively highly expressed, perhaps ubiquitously expressed, genes can be analyzed. She’s examining distribution of allelic effects between individuals and between tissues. She wants to quantify regulatory variation in each tissue by looking across all tissues and samples. Thyroid has a large relative (to other tissues) proportion of cis-eQTL and ASEs unlike other tissues. Her data are progressing to descriptions of proxy tissues for eQTL analysis. She is asking, How likely is a second tissue in the same individual to show ASE? eQTL work is done in populations and now look at the individual and that person’s ASE effect.  Because of wide variation in expression levels and ASE effects across individuals, the variants are not great predictor of individual phenotypes even at the cellular level.

Manuel Rivas. Transcriptome analysis of the functional impact of putative loss of function variants. Looks to annotate exome resequencing data and rare variants to isoforms from RNA-seq. LMNA provides a nice example of a gene with muscle specific mRNA isoforms and thus only these two isoforms should be used in explaining muscle disorders as the other isoforms are not expressed in this tissue.  

Chris Fuller. GWAS variants as eQTL based on analysis of GTEx data. Sherlock is their tool, It uses all GWAS SNPs even those below genome-wide significance. It uses both cis- and trans-eQTL loci. Sherlock maps disease-SNP associations to disease-gene groups. Linkage is very important in this work. The stronger results come from cis-eQTL, as shown by looking at Crohn disease GWAS. He also implicates genes through trans-eQTL data. See sherlock.ucsf.edu. Many small GWAS may remain unpublished for lack of strong single-SNP results. Aggregating SNPs boosts statistical power. They then implicate a relevant leukemia gene (FLI1) though multiple (n=6) trans-eQTL SNPs.


Eric Gamazon. GTEx – Expanding on GWAS. Uses the Wellcome Trust Case Control Consortium and their 7 diseases, including CAD. Adipose cis-eQTL are enriched for Crohn disease, CAD, hypertension and rheumatoid arthritis variants. GTEx adipose eQTL improves HOMA-IR GWAS. What proportion of variability in expression is captured by eQTL? He claims that he can capture 30-50% of heritability from genome-wide markers SNPs (> 200,000) for type 1 diabetes and Crohn disease when using a small number of informative cis-eQTL GTEx SNPs (2883 SNPs for T1D, ~3000 for CD). He makes no claims about saying anything about causal variants with this approach. 

Jason Wright. Chasing causal loci: Genome engineering of a non-coding region of 9p21 to identify mechanisms of diabetes predisposition. Enhancer assay of a standard type was used to look at various 2-kbp sequences across the region. Little or no enhancer activity seen until they transfected rodent islet cells. Promoter regions of all three genes (CDKN2A, eg) physically interact with the 10-kbp region containing the risk haplotype. TALEN (TAL effector nucleases) genome engineering gives near isogenic cell lines with or without specific alleles; he did this to delete the entire risk region. It looks as though the 9p21 region affects expression of CDKN2A and not CDKN2B and CDKN2AB-AS by about 20% per risk allele and in cis. He is now trying CRISPR genome engineering to fine map the 10-kbp region.

Daniel MacArthur. Gene expression data and PPIs are used to inform the clinical exome sequencing. IBAS protein-protein interaction score and way to score placement within a PPI network.

Luke Ward. Systematic annotation of GTEx eQTL using ENCODE and Roadmap data. His slide on genetic variant, tissue/cell type, molecular phenotypes (histone methylation, eg) and organismal phenotypes (lipids, heart rate) slide is neat and while outlining potential for a druggable path could also be used to outline a nutrition path to retain and maintain health, as opposed to recovering health. HaploReg (http://nar.oxfordjournals.org/content/40/D1/D930.short) is the portal where the HepG2 enhancer variants can be found – these are relevant for blood lipids. 

GTEx & NIH panel. Audience participation in terms of data types/fields to make available and discussion of other tissues to sample.

Gad Getz. He gave a recap of the day's talks…


Friday, March 8, 2013

Paper of the Week: Mechanistic dissection of a heart disease risk variant

This week, my choice for Paper of the Week is one in which the mechanism by which a heart disease risk variant, initially observed in GWAS, has been determined. The paper is by Pu, Xiao et al. and published in the American Journal of Human Genetics 92:366-374.

ADAMTS7 encodes a metalloprotease which cleaves protein in order to activate them. One of its substrates is COMP, known as thrombospondin-5, which is known to be made by vascular smooth muscle cells (VSMC) and inhibit their migration. The authors noted that ADAMTS7 accumulated in smooth muscle cells in both carotid and coronary atherosclerotic plaques. The relevant genetic variant here is rs3825807, calling for a nonsynonymous A to G, leading to a substitution at amino acid 214, Ser to Pro, in the prodomain of the ADAMTS7 protease. VSMCs harboring the G/G genotype for rs3825807 had attenuated migratory ability, while conditioned media of VSMCs of the G/G genotype contained less of the cleaved form of COMP.


The authors write that "the results of our study indicate that rs3825807 has an effect on ADAMTS7 maturation, thrombospondin-5 cleavage, and VSMC migration, with the variant associated with protection from atherosclerosis and CAD (coronary artery disease) rendering a reduction in ADAMTS7 function." I find this a noteworthy study because it takes a GWAS hit for CAD risk and informs us as to how that allele can lead to plaque formation and atherosclerosis.

Friday, May 4, 2012

POTW: Uncovering the function of an intergenic SNP

My choice for Paper Of The Week this week is a report from a few weeks back (digging through the pile...) in which a polymorphism conferring increased risk of renal cell carcinoma is investigated for allele-specific functions. The paper is "Common genetic variants at the 11q13.3 renal cancer susceptibility locus influence binding of HIF to an enhancer of cyclin D1 expression" by Schödel, et al. (Nature Genetics 44:420-425).

Although the authors had several clues that the risk SNPs would (likely) affect expression of CCND1 (cyclin D1) in a manner regulated by hypoxia-induced factors - namely, that HIFs were known to regulate CCND1 but from an unknown binding site and that CCND1 is an established oncogene, among others - they accumulated much new data to nail down the role of EPAS1 (HIF-2) in regulating CCND1 expression.

One nice aspect of this work is the authors' taking advantage of signals seen in a renal carcinoma cell line and not in a breast cancer cell line (serving then as control). For example, they looked at the epigenetic enhancer marks at the 11q13.3 susceptibility locus with FAIRE (ormaldehyde-assisted isolation of regulatory elements to identify regions of nucleosome occupancy), and EPAS1 binding as assessed by ChIP-qPCR. The use of pVHL-defective RCC cell lines verified the role of VHL (von Hippel–Lindau tumor suppressor) in this cancer and consequence of allele-specific expression of CCND1.

Taken together, the data presented show that the haplotype associating with reduced renal cell cancer risk hinders EPAS1 binding, "resulting in an allelic imbalance in cyclin D1 expression, thus affecting a link between hypoxia pathways and cell cycle control." This is nice work and a fine example of the approaches needed to develop a clear understanding of polymorphism and disease risk from a functional perspective.

Friday, March 16, 2012

POTW: Evolutionary constraints and the discovery of disease markers

My selection for Paper of the Week for 16 March 2012 is by Joel Dudley, et al. and published as a letter in Molecular Biology & Evolution. Its title is "Evolutionary meta-analysis of association studies reveals ancient constraints affecting disease marker discovery."

The authors examined over 5800 disease-associating variants, comparing the genomic neighborhood across a panel of species. This covered 230 different disease and disease risk phenotypes. Importantly, the authors demonstrate that there is a propensity to discover such disease SNPs at "conserved genomic positions, because the effect size (odds ratio) and allelic P-value of genetic association of a SNP relates strongly to the evolutionary conservation of their genomic position." This then allowed them to develop a new means to rank such association SNPs in which a conservation score, based on the evolutionary analysis, is incorporated into the P-value of the genotype-phenotype association.

As many GWAS SNPs alter gene expression - either through altered transcription factor binding or microRNA-mRNA interaction, and as such evolutionary mechanisms most likely involve a sensing or monitoring of the environment with concomitant changes in gene expression, this makes sense. In fact, the role of such types of SNPs (those under selective pressure) and their role in heart disease, was a topic on which we published in 2010.

The article by Dudley, et al. is really nice work and one whose insight we will use to inform our GWAS analysis.

Thursday, February 23, 2012

POTW: January, 2012 choices

Three papers published last month that I found to be of interest are listed here. These are:

The mystery of missing heritability: Genetic interactions create phantom heritability, by Zuk, et al. This addresses the missing heritability question, suggesting that "the total heritability may be much smaller and thus the proportion of heritability explained much larger."

Characterisation and discovery of novel miRNAs and moRNAs in JAK2V617F mutated SET2 cells, by Botoluzzi, et al. What interested me in this article was the generation of novel microRNAs that were induced by the cancerous state triggered by this JAK2 variant. This indicates to me that the microRNA realm is broad and rich with many as yet undiscovered relationships.

The PLoS One paper entitled "Genetic signatures of exceptional longevity in humans," by Sebastiani, et al. We here were very curious how this was different from the version retracted from Science and what findings are now reported. TOMM40 near APOE is indeed interesting.

Paper of the Week: cis-eQTL between normal and cancer tissue

My choice for Paper of the Week this week is "cis-Expression QTL Analysis of Established Colorectal Cancer Risk Variants in Colon Tumors and Adjacent Normal Tissue," by Loo, Cheng, Tiirikainen, Lum-Jones, Seifried, et al. (2012) appearing in PLoS One.

I find this article of interest because I feel that many GWAS hits for disease risk will serve to alter expression of a near or distant gene(s) in an allele-specific manner. This group looked at gene expression differences between tissues that were either colorectal tumors or their paired, adjacent normal tissue, and then associated those gene expression differences with allelic variation. Assaying 40 individuals was sufficient to identify 3 SNPs affecting expression of 4 genes: ATP5C1, DLGAP5, NOL3 and DDX28.

A link to this paper is here.

Citation:
Loo LWM, Cheng I, Tiirikainen M, Lum-Jones A, Seifried A, et al. (2012) cis-Expression QTL Analysis of Established Colorectal Cancer Risk Variants in Colon Tumors and Adjacent Normal Tissue. PLoS ONE 7(2): e30477.

Wednesday, October 26, 2011

The curious case of SNP rs2880301

"Wow! I've never seen anything like that before," my colleague Chao-Qiang Lai exclaimed when examining output from his analysis of genome-wide association (GWAS) data. He was looking for genetic markers influencing the level of triglyceride in serum as part of the GOLDN study. GOLDN is looking at the genetics of the response to lipid-lowering medication. The result of Chao's preliminary analysis indicated that SNP rs2880301 associated with TG levels with a p-value of 10-218. He showed me some data from the scan of the Affymetrix 6.0 genotyping chip and we postulated that we could be looking at some type of CNV (copy number variant) or deletion, but the lack of minor allele homozygotes troubled us.

What intrigued us right from the start was our colleagues at other institutions who are also analysis the GOLDN GWAS data did not report this SNP in their initial findings. The dbSNP entry for rs2880301 indicates a C to T variant with an allele frequency of 0.24 in the four primary HapMap populations from USA, Nigeria, China and Japan. No differences in allele frequency means no chance of positive (or negative) selection on this variant. No, none indeed as we were to learn later.

So, Chao dug deeper into his data and he and I shot ideas back and forth. After my suggestion to look at sex, he saw that when the SNP and sex are together in the same model, the analysis did not complete. Then, looking at the individual genotypes, he saw that all the men had genotype CT and all the women CC. This is from a total of just over 800 subjects.

OK, time for me to step in and see where this "SNP" maps in the genome. My first query was the flanking sequence supplied by Affymetrix. This 33-bp segment maps nearly perfectly to both chromosome 13, within intron 1 of the TPTE2 gene and agreeing with both dbSNP and Affyemtrix's annotation of the SNP, and curiously to a spot on the Y chromosome. (TPTE2 is a membrane-associated phosphatase which acts on the 3-position phosphate of inositol phospholipids and could be argues as relevant to TG biology.) The only residue not matching is the "polymorphic" base of the "SNP." A C is found on chr13 and a T is found on Y. Thus, the SNP becomes a marker of sex and Chao was right - it is a type of deletion (females carry no Y chromosome) - but a deletion he had not envisioned.

Is rs2880301 then a marker for gender? Not really. I compared the genomic regions where the homologous SNP sequences were found on both chromosomes, extending over 6 kbp in each direction. I saw a large region of sequence identity between 13 and Y - over 96% - for a ~5 kbp segment. Running RepeatMasker indicated that rs2880301 falls within an L1 LINE, a common repeat element. Thus, while it is intriguing that an array of repeats (70% of the 13-kbp segment of chr13 is masked by RepeatMasker) are conserved between chromosomes 13 and Y, and in order, SNP rs2880301 is not really a SNP. All subjects are C on chr13 and all Y chromosomes are T.

What we then had in our data were five genotypes: CC on chr13 for all women, CC on chr13 for all men and T on Y. Thus, the "allele frequencies" of C between 0.75 and 0.80 and T between 0.20 and 0.25 seen by us and others, including the HapMap data, roughly correspond to populations that are half to slightly more than half women.

Tuesday, September 13, 2011

Genetics of hypertension

An important paper describing genetic factors involved in blood pressure and heart disease was published this week. 29 loci were identified - 16 of these novel. This is impressive, but not so surprising as both diastolic and systolic blood pressure are complex, heritable traits. Many, many factors are at play here. In fact, the authors describe a risk score based on genotypes at 29 genome-wide significant variants, which was associated with hypertension, left ventricular wall thickness, stroke and coronary artery disease, but not kidney disease or kidney function.

Some time ago, I wrote about the genome of the SHR (spontaneous hypertensive rat). This rat and the FHH strain are models for hypertension. The genome of the SHR strain showed that 788 genes contained variants with respect to the reference rat genome. Whether any or all of those variants function in producing the hypertension phenotype is not clear, but reasoning went that these variants would be a logical place from which to build list of candidate genes.

Well, I thought it would be a fun and quick exercise to see how many of these 29 loci detected by genome-wide associations in nearly 70,000 individuals of European ancestry (with validation of top signals in up to 133,000 additional individuals of European descent) compare to those genes that possess genetic variants in either the FHH or SHR rat strains. This may give insight into the applicability of sequencing the genomes of specific model organism strains and into how many or how often genes are identified by both GWAS and whole-genome sequencing. To do this, I used the Rat Genome Browser hosted by the Medical College of Wisconsin.

Of the 29 human loci, 23 map in the rat genome within a QTL for blood pressure. This looks good, but consider that there are numerous blood pressure (BP) QTL mapped throughout the rat genome. In fact, the rat Adm gene maps within 27 different BP QTL and three other gene regions, Furin - Fes, Plekha7 and Mov10, map within more than 20 different BP QTL. Some of these QTL are large, spanning many genes, which means that fine-mapping is needed - such as a GWAS - to identify more precisely candidate loci.

Only 18 of the human BP genes identified in the paper contain SNPs in either the FHH or SHR strains. Often, the variants are shared in both strains. Both synonymous and nonsynonymous SNPs were noted, but synonymous far outnumbered those variants that altered the underlying amino acid sequence. No SNPs in gene control regions were noted, which may indeed be the case or a limitation of the data sources used here.

The human genes whose rat versions contain SNPs in the hypertensive-susceptible strains are:

SLC39A8
ATP2B1
GNAS - EDN3
MTHFR - NPPB
FGF5
CYP1A1 - ULK3
FURIN - FES
FLJ32810 - TMEM133
NPR3 - C5orf23
EBF1
PLCE1
BAT2-BAT5
ZNF652
TBX5 - TBX3
JAG1
GUCY1A3 - GUCY1B3
MECOM
ULK4


SNPs altering gene expression still need to be added to this analysis. Nonetheless, the numbers and types of genes that share genetic variation in hypertensive mammals (human, rat) is revealing. It is likely that the 788 identified genes with variation in the SHR rat are not all important for hypertension, but that strain does carry variants in 17 of these new BP genes. Or is that just 17?

Thursday, March 10, 2011

Genetics of coronary heart disease

Note: This is a guest-post, authored by geneticist and molecular biologist Dr. Chao-Qiang Lai; with edits added by LP.

Last week, Nature Genetics published three letters reporting results from genome-wide association studies (GWAS) for coronary heart disease (CAD). The studies reported a number of markers that reached the threshold of statistical significance for association to CAD with concomitant association to traditional biomarkers of disease risk, such as elevated LDL-cholesterol (LDL-C), elevated total cholesterol, decreased HDL-cholesterol (HDL-C), hypertension, obesity (as measured by elevated body mass index), or type 2 diabetes. However, the two larger and more highly powered GWAS (C4D Genetics Consortium, Schunkert, et al.) also identified CAD-associated variants that are not associated with traditional biomarkers. The third study is of interest because it examines CAD in Chinese populations, but beginning with a discovery set of 130 cases and 130 controls leaves it a bit under-powered. They report a unique association between a SNP in C6orf105 and CAD, which is not found in European or south Asian populations. Curiously, this gene has also been implicated in non-syndromic oral cleft.

There are many sources of CAD. Blood lipids are most commonly thought of as the prime source, but blood pressure in the form of hypertension is also a source. Traditional biomarkers such as LDL-C, HDL-C, triglycerides, and hypertension have been used almost as the sole surrogates for measuring the devolvement and progression of CAD over the course of some 50 years. Meta-analyses of GWAS based on over 100,000 subjects (22,233 cases and 64,762 controls from 14 GWAS) thus far have identified 23 genetic variants associating with CAD. The eye-opening aspect to this is these variants account for about 10% of CAD cases with the shocking observation that 17 of 23 confirmed loci appear to have no association with traditional markers. This observation then suggests two possible explanations.

One possibility is when we assume that the remainder of the CAD cases (90%) contribute to risk associated with traditional markers, such genetic factors cannot be detected based on current GWAS methodology. This is likely to be true because of to the effect sizes of these variants are too small, or their effects are camouflaged by gene-gene (GxG) and gene-environment (GxE) interactions or by epigenetic mechanisms.

This second possibility rests on the fundamental premise that all markers associating with CAD have more or less equal chance to be detected. It then follows that a majority of genetic factors that contribute to CAD has nothing to do with traditional markers. If this is indeed the case, it opens a new avenue to identify the new mechanism(s) and new biomarkers that lead to CAD. In fact, this possibility is supported by many observations. For example, 50% of those individuals who have CAD have low LDL-C (Braunwald & Shattuck; Ridker).

These genes – for example, ADAMTS7, PDGFD, ABO and PPAP2B – point to new mechanisms. While GxG, GxE and epigenetic interactions remain as viable contributors to CAD risk, the path to better understanding of the other component(s) to CAD risk will likely transit through metabolic profiling to identify the compounds that distinguish elevated from nominal risk. Furthermore, research will need to be conducted in model organisms based on these newly discovered genes, perhaps in pig as this is a good model for heart function and disease in human.

Wednesday, February 2, 2011

Synonymous SNPs are not so synonymous

Early this week, an excellent paper by Brest, Darfeuille-Michaud, Hofman, et al. in Nature Genetics provides a prime example of going beyond genome-wide association studies (GWAS) to dissect the functional consequences of a genetic variant associated with disease risk. In so doing, the authors provide another case of synonynous SNPs not being so synonymous.

Here are what I find to be the key points of the research presented in this report:

1. The exonic SNP c.313C>T (rs10065172) is in perfect linkage disequilibrium (r2=1.0) with a deletion polymorphism of 20 kbp mapping upstream of the IRGM gene. This deletion has been strongly associated with Crohn's disease in several European populations or those with European ancestry. What is important here is a SNP can act as a tag or proxy for the deletion.

2. The c.313C>T variant alters codon 105 of the IRGM protein from CTG>TTG. Both codons call for leucine upon translation and so this SNP is classified as synonymous. The authors speculate that there could be allele-specific consequences to protein expression. Based on two other reports from other groups, the authors decided to investigate whether allele-specific interactions between the IRGM transcript and a microRNA could be at play here. They observed a predicted binding between microRNA-196 (or miR-196, both miR-196A encoded by A1 and A2 genes and by miR-196B) that was affected by the variation at SNP c.313C>T. Importantly, they show that not only is the miR-196-IRGM interaction real but that expression of miR-196 is elevated in inflammatory epithelia from Crohn's sufferers. These results underscore the point that synonymous SNPs are not so synonymous. The different alleles can exhibit different functions that have health consequences.

3. From GWAS to function. Although this paper does not report original results from GWAS, it builds on those results in an important way. There are four key papers reporting GWAS results for IRGM and Crohn's disease. These papers are by Parkes et al (2007), the Wellcome Trust Case Control Consortium (2007), Barrett et al (2008) and Franke et al (2010). So, in just over three years from the initial discovery of association of this once rather unremarkable gene (only 5 papers were published on IRGM prior to the initial GWAS report of 2007, most reporting a role in autophagy), we now have a much deeper understanding how a synonymous variant leads to the disease condition.

This is very nice work indeed and can be held up as an example of the success of GWAS in laying a foundation for getting at the mechanism of a disease.