<!DOCTYPE art SYSTEM 'http://www.biomedcentral.com/xml/article.dtd'>
<art>
   <ui>1753-6561-3-S7-S110</ui>
   <ji>1753-6561</ji>
   <fm>
      <dochead>Proceedings</dochead>
      <bibl>
         <title>
            <p>Graphic analysis of population structure on genome-wide rheumatoid arthritis data</p>
         </title>
         <aug>
            <au id="A1" ca="yes">
               <snm>Zhang</snm>
               <fnm>Jun</fnm>
               <insr iid="I1"/>
               <email>junzhang@uchicago.edu</email>
            </au>
            <au id="A2">
               <snm>Weng</snm>
               <fnm>Chunhua</fnm>
               <insr iid="I2"/>
               <email>chunhua.weng@dbmi.columbia.edu</email>
            </au>
            <au id="A3">
               <snm>Niyogi</snm>
               <fnm>Partha</fnm>
               <insr iid="I3"/>
               <email>niyogi@cs.uchicago.edu</email>
            </au>
         </aug>
         <insg>
            <ins id="I1">
               <p>Department of Radiology, The University of Chicago, 5841 South Maryland Avenue, Chicago, Illinois 60637 USA</p>
            </ins>
            <ins id="I2">
               <p>Department of Biomedical Informatics, Columbia University, 622 West 168 Street, New York, New York 10032 USA</p>
            </ins>
            <ins id="I3">
               <p>Departments of Statistics and Computer Science, The University of Chicago, 1100 East 58<sup>th </sup>Street, Chicago, Illinois 60637 USA</p>
            </ins>
         </insg>
         <source>BMC Proceedings</source>
         <supplement>
            <title>
               <p>Genetic Analysis Workshop 16</p>
            </title>
            <note>Proceedings</note>
         </supplement>
         <conference>
            <title>
               <p>Genetic Analysis Workshop 16</p>
            </title>
            <location>St Louis, MO, USA</location>
            <date-range>17-20 September 2008</date-range>
            <url>http://www.gaworkshop.org/</url>
         </conference>
         <issn>1753-6561</issn>
         <pubdate>2009</pubdate>
         <volume>3</volume>
         <issue>Suppl 7</issue>
         <fpage>S110</fpage>
         <url>http://www.biomedcentral.com/1753-6561/3/S7/S110</url>
         <xrefbib>
            <pubid idtype="pmpid">20017975</pubid>
         </xrefbib>
      </bibl>
      <history>
         <pub>
            <date>
               <day>15</day>
               <month>12</month>
               <year>2009</year>
            </date>
         </pub>
      </history>
      <cpyrt>
         <year>2009</year>
         <collab>Zhang et al; licensee BioMed Central Ltd.</collab>
         <note>This is an open access article distributed under the terms of the Creative Commons Attribution License (<url>http://creativecommons.org/licenses/by/2.0</url>), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</note>
      </cpyrt>
      <abs>
         <sec>
            <st>
               <p>Abstract</p>
            </st>
            <p>Principal-component analysis (PCA) has been used for decades to summarize the human genetic variation across geographic regions and to infer population migration history. Reduction of spurious associations due to population structure is crucial for the success of disease association studies. Recently, PCA has also become a popular method for detecting population structure and correction of population stratification in disease association studies. Inspired by manifold learning, we propose a novel method based on spectral graph theory. Regarding each study subject as a node with suitably defined weights for its edges to close neighbors, one can form a weighted graph. We suggest using the spectrum of the associated graph Laplacian operator, namely, Laplacian eigenfunctions, to infer population structures instead of principal components (PCs). For the whole genome-wide association data for the North American Rheumatoid Arthritis Consortium (NARAC) provided by Genetic Workshop Analysis 16, Laplacian eigenfunctions revealed more meaningful structures of the underlying population than PCA. The proposed method has connection to PCA, and it naturally includes PCA as a special case. Our simple method is computationally fast and is suitable for disease studies at the genome-wide scale.</p>
         </sec>
      </abs>
   </fm>
   <bdy>
      <sec>
         <st>
            <p>Introduction</p>
         </st>
         <p>It is well known that unidentified population structure can cause spurious associations in genome-wide association studies <abbrgrp><abbr bid="B1">1</abbr><abbr bid="B2">2</abbr></abbrgrp>. Such associations typically occur when the disease frequency varies across subpopulations, thereby resulting in the oversampling of affected individuals from particular subpopulations. It is therefore critical to correctly infer population structure from genotypic data when performing genome-wide association studies. Though this topic has been extensively studied, the prevailing methods such as genomic control and structured association still have limitations <abbrgrp><abbr bid="B3">3</abbr></abbrgrp>. Recently, principal-component analysis (PCA) has been employed to summarize genetic background variation <abbrgrp><abbr bid="B4">4</abbr><abbr bid="B5">5</abbr></abbrgrp>. Price et al. <abbrgrp><abbr bid="B3">3</abbr></abbrgrp> suggested the inclusion of a few top PCs as covariates in a regression setting to correct for structure. However, there is concern about the interpretation of PCs. Recently, for instance, Novembre and Stephens <abbrgrp><abbr bid="B6">6</abbr></abbrgrp> showed that patterns (such as gradients and waves) appearing in the PC analysis of continuous genetic data sometimes resemble sinusoidal mathematical artifacts. These generally arise when PCs are applied to spatially correlated data. Nevertheless, PCA can provide evidence of major demographic migration events and is still widely used in many contexts for genetic data analysis.</p>
         <p>Here we propose a novel approach for detecting population structure inspired by graph theory. Unlike PCA, which uses all pairs of individuals, this method uses the idea of shrinkage and considers only close neighbors as measured by pairwise correlation. Therefore, it is robust to outliers and the results obtained can reveal the local dependence structures of population samples. We demonstrate our method, LAPSTRUCT, on the North American Rheumatoid Arthritis Consortium (NARAC) data provided by Genetic Analysis Workshop 16. Rheumatoid arthritis (RA) is a complex and chronic inflammatory joint disease with both genetic components and environmental factors. It has been observed that <it>PTPN22 </it>and <it>TRAF1-C5 </it>genes are associated with RA <abbrgrp><abbr bid="B7">7</abbr></abbrgrp>.</p>
      </sec>
      <sec>
         <st>
            <p>Methods</p>
         </st>
         <p>The NARAC study sample includes 868 cases ascertained at RA clinics and 1194 controls from the New York cancer study. The individuals from NARAC were genotyped with the Illumina 550 k single-nucleotide polymorphism (SNP) array in the whole genome, with total 545,080 SNPs. 507,246 SNPs passed quality control after removing SNPs with a departure from Hardy-Weinberg equilibrium (using <it>&#967;</it><sup>2 </sup>statistic) in controls significant at the 10<sup>-5 </sup>level, SNPs with genotype call rates &lt;90%, and SNPs with a minor allele frequency &lt;0.01. Each individual's affection status (unaffected as 0, affected as 1) was regarded as the phenotype. All 2026 individuals in the NARAC data were included in this analysis.</p>
         <p>First, let <it>g </it>denote the matrix of genotype (0, 0.5, 1) of individual <it>j </it>at SNP. We standardize each SNP <it>i </it>by subtracting the row mean <inline-formula><graphic file="1753-6561-3-S7-S110-i1.gif"/></inline-formula>, and then divide each entry by <inline-formula><graphic file="1753-6561-3-S7-S110-i2.gif"/></inline-formula>, where <it>p</it><sub><it>i </it></sub>is an estimate of the allele frequency at SNP <it>i </it>given by <inline-formula><graphic file="1753-6561-3-S7-S110-i3.gif"/></inline-formula>; all missing entries are excluded from the computation. Let <it>g </it>still denote the standardized genotype matrix, then <inline-formula><graphic file="1753-6561-3-S7-S110-i4.gif"/></inline-formula>. Then, for each pair of individuals <it>j </it>and <it>k</it>, we define the distance ||<it>v</it><sub><it>j </it></sub>- <it>v</it><sub><it>k</it></sub>|| = 1 - <it>C</it><sub><it>jk</it></sub>. Regard each individual <it>j </it>as a vertex <it>V</it><sub><it>j </it></sub>in a weighted graph G = (V, E), where <it>j </it>= 1 to N. Set the weight between individuals <it>j </it>and <it>k </it>to be a Gaussian kernel <inline-formula><graphic file="1753-6561-3-S7-S110-i5.gif"/></inline-formula> for <it>j </it>&#8800; <it>k </it>and ||<it>v </it>- <it>v</it><sub><it>k</it></sub>|| &lt;<it>&#949;</it>, <it>W</it><sub><it>jk </it></sub>= 0 for <it>j </it>&#8800; <it>k </it>and ||<it>v </it>- <it>v</it><sub><it>k</it></sub>|| > <it>&#949; </it>and <it>W</it><sub><it>jj </it></sub>= 1.0 for all <it>j</it>. Here, <it>&#949; </it>is a positive real number that measures the size of each subject's neighborhood in terms of correlations; that is, all individuals within distance <it>&#949; </it>are regarded as one's close neighbors.</p>
         <p>Cases and controls are regarded as vertices of a weighted graph and each vertex is connected to its close neighbors through edges according to their pairwise distances. This reflects the fact that distances between vertices that are far apart are relatively less important, and therefore need not be preserved if the sample size of the dataset is reasonably large. The eigenfunctions of the associated graph Laplacian operator on the graph are generalized geometric harmonic functions, which contain geometric structure information of the population dependence graph. The eigenvectors of the graph Laplacian are the first-order linear approximations of Laplacian eigenfunctions. Therefore, they are much more meaningful than the usual PCs as they relate to the intrinsic structure of the data.</p>
         <p>Let <it>D </it>be a diagonal matrix of size <it>N &#215; N </it>with entries <inline-formula><graphic file="1753-6561-3-S7-S110-i6.gif"/></inline-formula>, which is a natural measure on the vertices. The Laplacian matrix on graph <it>G </it>is defined as <it>L </it>= <it>W</it>-<it>D</it>. Note that <it>L </it>is a symmetric and positive semidefinite matrix, and we restrict to the normalized version <it>D</it><sup>-1</sup><it>L</it>, which is not symmetric anymore. The eigenfunctions of the normalized equation <it>Le </it>= <it>&#955;e </it>are denoted by <it>e</it><sub><it>j </it></sub>= (<it>e</it><sub><it>j</it>1</sub>, ..., <it>e</it><sub><it>eN</it></sub>)<sup><it>T </it></sup>for each <it>j</it>, ranked according to the increasing of their corresponding eigenvalues, i.e., <it>&#955;</it><sub>0 </sub>&#8804; <it>&#955;</it><sub>1 </sub>&#8804; <it>&#955;</it><sub>2 </sub>&#8804; &#8943;. It is easy to see that 0 is always an eigenvalue with constant eigenvector consisting of all 1 values. These eigenfunctions generalize the low frequency Fourier harmonics on a manifold approximated by the graph <it>G</it>. The Laplacian eigenmap with first <it>n </it>(usually small, 2 or 3) eigenvectors is defined as <it>f</it>: <it>k </it>&#8594; (<it>e</it><sub>1<it>k</it></sub>, <it>e</it><sub>2<it>k</it></sub>, &#8943;, <it>e</it><sub><it>nk</it></sub>) for individual <it>k </it>to achieve dimension reduction. Note the situation here is different from PCA, where one takes the PCs corresponding to the largest eigenvalues that account for the largest amount of variation in the data.</p>
         <p>The Laplacian eigenmap has the important locality preserving property, that is, the distance between a pair of subjects in the Laplacian eigenmap reflects their degree of correlation. The more they are correlated, the closer together they are mapped. Immediately, Laplacian eigenmap leads to cluster-like structures for subjects who either come from the same discrete subpopulation or share more common ancestry in an admixed population. Therefore, we suggest using Laplacian eigenvectors instead of PCs to study population structure. Next we follow Price et al. <abbrgrp><abbr bid="B3">3</abbr></abbrgrp> to regress genotypes and phenotypes on the top ten Laplacian eigenvectors for each individual and compute the adjusted <it>&#967;</it><sup>2 </sup>statistic of the residuals.</p>
      </sec>
      <sec>
         <st>
            <p>Results</p>
         </st>
         <p>The PC map (Figure <figr fid="F1">1a</figr>) depicts the European population structure similar to the map previously published by Price et al. <abbrgrp><abbr bid="B3">3</abbr></abbrgrp>. The Laplacian eigenmap (Figure <figr fid="F1">1b</figr>) shows the compact trend from center to bottom right and a long tail-like trend to the left. Surprisingly, these two trends are remarkably separated in the unnormalized version of Laplacian eigenmap (Figure <figr fid="F1">1c</figr>). We compared the results for two SNPs that have been reported to be associated with RA (see Table <tblr tid="T1">1</tblr>). The results are consistent with the prevailing principal-components-based approach, EIGENSTRAT.</p>
         <tbl id="T1">
            <title>
               <p>Table 1</p>
            </title>
            <caption>
               <p>Association testing results for genes <it>PTPN22 </it>and <it>TRAF1-C5 </it>by EIGENSTRAT and LAPSTRUCT</p>
            </caption>
            <tblbdy cols="4">
               <r>
                  <c ca="left">
                     <p>
                        <b>SNP</b>
                     </p>
                  </c>
                  <c ca="center">
                     <p>
                        <b>Chromosome</b>
                     </p>
                  </c>
                  <c ca="center">
                     <p>
                        <b>EIGENSTRAT</b>
                     </p>
                  </c>
                  <c ca="center">
                     <p>
                        <b>LAPSTRUCT</b>
                     </p>
                  </c>
               </r>
               <r>
                  <c cspan="4">
                     <hr/>
                  </c>
               </r>
               <r>
                  <c ca="left">
                     <p>rs2476601</p>
                  </c>
                  <c ca="center">
                     <p>1</p>
                  </c>
                  <c ca="center">
                     <p>26.74 (2.33 &#215; 10<sup>-7</sup>)</p>
                  </c>
                  <c ca="center">
                     <p>33.72 (6.36 &#215; 10<sup>-9</sup>)</p>
                  </c>
               </r>
               <r>
                  <c ca="left">
                     <p>rs3761847</p>
                  </c>
                  <c ca="center">
                     <p>9</p>
                  </c>
                  <c ca="center">
                     <p>27.57 (1.52 &#215; 10<sup>-7</sup>)</p>
                  </c>
                  <c ca="center">
                     <p>25.39 (4.68 &#215; 10<sup>-7</sup>)</p>
                  </c>
               </r>
            </tblbdy>
         </tbl>
         <fig id="F1">
            <title>
               <p>Figure 1</p>
            </title>
            <caption>
               <p>Population structures</p>
            </caption>
            <text>
               <p><b>Population structures</b>. Detected by PCA: a, Laplacian; b, its unnormalized version; c, both with <it>&#949; </it>= 1.0.</p>
            </text>
            <graphic file="1753-6561-3-S7-S110-1"/>
         </fig>
      </sec>
      <sec>
         <st>
            <p>Discussion</p>
         </st>
         <p>By setting a constant weight for each pair of individuals and sufficiently large <it>&#949; </it>to include all individuals into everyone's neighborhood, the proposed approach naturally includes PCA as a special case. This fact follows from the observation below. If all weights <it>W</it><sub><it>ij </it></sub>are equal, say, <inline-formula><graphic file="1753-6561-3-S7-S110-i7.gif"/></inline-formula>, where <it>N </it>is the total number of individuals, then <inline-formula><graphic file="1753-6561-3-S7-S110-i8.gif"/></inline-formula> and <inline-formula><graphic file="1753-6561-3-S7-S110-i9.gif"/></inline-formula>, where <it>e </it>= (1, ..., 1)<sup><it>T</it></sup>. Let <it>g </it>= (<it>g</it><sub>1</sub>, ..., <it>g</it><sub><it>N</it></sub>)<sup><it>T </it></sup>denote the genotype data of all individuals, where each <it>g</it><sub><it>i </it></sub>stands for the genotype vector for the <it>i</it><sup>th </sup>individual and let <it>&#956; </it>denote the sample mean vector of genotypes. Then one has <inline-formula><graphic file="1753-6561-3-S7-S110-i10.gif"/></inline-formula>. Because <inline-formula><graphic file="1753-6561-3-S7-S110-i11.gif"/></inline-formula> is the sample covariance matrix of the individuals, the Laplacian eigenfunctions equal the PCs.</p>
         <p>In general, for sufficiently large <it>&#949;</it>, the top Laplacian eigenfunctions describe global variations instead of local dependence structures, and they numerically approximate to the top PCs. As <it>&#949; </it>decreases, the Laplacian eigenmap describes the local dependence structures at different scales. When <it>&#949; </it>becomes so small that each subject's neighborhood shrinks to itself, Laplacian eigenmap cannot detect any structure. In practice, the successful use of the proposed algorithm requires a method to choose effective <it>&#949; </it>to make the graph connected and maintain valid type 1 error for association studies. Similar to the PCA approach for association testing, a method to choose the eigenvector dimension is also required for optimal performance.</p>
         <p>We have introduced a novel method for population structure detection that preserves local dependence structures. The Laplacian eigenmap naturally leads to population clusters according to the degree of pairwise correlation among individuals. In our example for testing for association between RA and SNPs, the Laplacian eigenmap method resulted in less noise than the PCA method and detected the same associations between SNPs and RA as the PCA method.</p>
      </sec>
      <sec>
         <st>
            <p>List of abbreviations used</p>
         </st>
         <p>NARAC: North American Rheumatoid Arthritis Consortium; PC: Principalcomponent; PCA: Principal component analysis; RA: Rheumatoidarthritis; SNP: Single-nucleotide polymorphism.</p>
      </sec>
      <sec>
         <st>
            <p>Competing interests</p>
         </st>
         <p>The authors declare that they have no competing interests.</p>
      </sec>
      <sec>
         <st>
            <p>Authors' contributions</p>
         </st>
         <p>JZ and PN designed the algorithm. JZ analyzed the data. JZ, CW, and PN wrote the manuscript.</p>
      </sec>
   </bdy>
   <bm>
      <ack>
         <sec>
            <st>
               <p>Acknowledgements</p>
            </st>
            <p>JZ is grateful to Matthew Stephens for his interest and great advice to improve the presentation of our findings. The Genetic Analysis Workshops are supported by NIH grant R01 GM031575 from the National Institute of General Medical Sciences.</p>
            <p>This article has been published as part of <it>BMC Proceedings </it>Volume 3 Supplement 7, 2009: Genetic Analysis Workshop 16. The full contents of the supplement are available online at <url>http://www.biomedcentral.com/1753-6561/3?issue=S7</url>.</p>
         </sec>
      </ack>
      <refgrp>
         <bibl id="B1">
            <title>
               <p>The effects of human population structure on large genetic association studies</p>
            </title>
            <aug>
               <au>
                  <snm>Marchini</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Cardon</snm>
                  <fnm>LR</fnm>
               </au>
               <au>
                  <snm>Phillips</snm>
                  <fnm>NS</fnm>
               </au>
               <au>
                  <snm>Donnelly</snm>
                  <fnm>P</fnm>
               </au>
            </aug>
            <source>Nat Genet</source>
            <pubdate>2004</pubdate>
            <volume>36</volume>
            <fpage>512</fpage>
            <lpage>517</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/ng1337</pubid>
                  <pubid idtype="pmpid" link="fulltext">15052271</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B2">
            <title>
               <p>Assessing the impact of population stratification on genetic association studies</p>
            </title>
            <aug>
               <au>
                  <snm>Freedman</snm>
                  <fnm>ML</fnm>
               </au>
               <au>
                  <snm>Reich</snm>
                  <fnm>D</fnm>
               </au>
               <au>
                  <snm>Penney</snm>
                  <fnm>KL</fnm>
               </au>
               <au>
                  <snm>McDonald</snm>
                  <fnm>GJ</fnm>
               </au>
               <au>
                  <snm>Mignault</snm>
                  <fnm>AA</fnm>
               </au>
               <au>
                  <snm>Patterson</snm>
                  <fnm>N</fnm>
               </au>
               <au>
                  <snm>Gabriel</snm>
                  <fnm>SB</fnm>
               </au>
               <au>
                  <snm>Topol</snm>
                  <fnm>EJ</fnm>
               </au>
               <au>
                  <snm>Smoller</snm>
                  <fnm>JW</fnm>
               </au>
               <au>
                  <snm>Pato</snm>
                  <fnm>CN</fnm>
               </au>
               <au>
                  <snm>Pato</snm>
                  <fnm>MT</fnm>
               </au>
               <au>
                  <snm>Petryshen</snm>
                  <fnm>TL</fnm>
               </au>
               <au>
                  <snm>Kolonel</snm>
                  <fnm>LN</fnm>
               </au>
               <au>
                  <snm>Lander</snm>
                  <fnm>ES</fnm>
               </au>
               <au>
                  <snm>Sklar</snm>
                  <fnm>P</fnm>
               </au>
               <au>
                  <snm>Henderson</snm>
                  <fnm>B</fnm>
               </au>
               <au>
                  <snm>Hirschhorn</snm>
                  <fnm>JN</fnm>
               </au>
               <au>
                  <snm>Altshuler</snm>
                  <fnm>D</fnm>
               </au>
            </aug>
            <source>Nat Genet</source>
            <pubdate>2004</pubdate>
            <volume>36</volume>
            <fpage>388</fpage>
            <lpage>393</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/ng1333</pubid>
                  <pubid idtype="pmpid" link="fulltext">15052270</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B3">
            <title>
               <p>Principal components analysis corrects for stratification in genome-wide association studies</p>
            </title>
            <aug>
               <au>
                  <snm>Price</snm>
                  <fnm>AL</fnm>
               </au>
               <au>
                  <snm>Patterson</snm>
                  <fnm>N</fnm>
               </au>
               <au>
                  <snm>Plenge</snm>
                  <fnm>RM</fnm>
               </au>
               <au>
                  <snm>Weinblatt</snm>
                  <fnm>ME</fnm>
               </au>
               <au>
                  <snm>Shadick</snm>
                  <fnm>NA</fnm>
               </au>
               <au>
                  <snm>Reich</snm>
                  <fnm>D</fnm>
               </au>
            </aug>
            <source>Nat Genet</source>
            <pubdate>2006</pubdate>
            <volume>38</volume>
            <fpage>904</fpage>
            <lpage>909</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/ng1847</pubid>
                  <pubid idtype="pmpid" link="fulltext">16862161</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B4">
            <title>
               <p>Association mapping, using a mixture model for complex traits</p>
            </title>
            <aug>
               <au>
                  <snm>Zhu</snm>
                  <fnm>X</fnm>
               </au>
               <au>
                  <snm>Zhang</snm>
                  <fnm>S</fnm>
               </au>
               <au>
                  <snm>Zhao</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Cooper</snm>
                  <fnm>RS</fnm>
               </au>
            </aug>
            <source>Genet Epidemiol</source>
            <pubdate>2002</pubdate>
            <volume>23</volume>
            <fpage>181</fpage>
            <lpage>196</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1002/gepi.210</pubid>
                  <pubid idtype="pmpid" link="fulltext">12214310</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B5">
            <title>
               <p>Qualitative semi-parametric test for genetic associations in case-control designs under structured populations</p>
            </title>
            <aug>
               <au>
                  <snm>Chen</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Zhu</snm>
                  <fnm>X</fnm>
               </au>
               <au>
                  <snm>Zhao</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Zhang</snm>
                  <fnm>S</fnm>
               </au>
            </aug>
            <source>Ann Hum Genet</source>
            <pubdate>2003</pubdate>
            <volume>67</volume>
            <fpage>250</fpage>
            <lpage>264</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1046/j.1469-1809.2003.00036.x</pubid>
                  <pubid idtype="pmpid" link="fulltext">12914577</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B6">
            <title>
               <p>Interpreting principal component analyses of spatial population genetic variation</p>
            </title>
            <aug>
               <au>
                  <snm>Novembre</snm>
                  <fnm>J</fnm>
               </au>
               <au>
                  <snm>Stephens</snm>
                  <fnm>M</fnm>
               </au>
            </aug>
            <source>Nat Genet</source>
            <pubdate>2008</pubdate>
            <volume>40</volume>
            <fpage>646</fpage>
            <lpage>649</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="doi">10.1038/ng.139</pubid>
                  <pubid idtype="pmpid" link="fulltext">18425127</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
         <bibl id="B7">
            <title>
               <p>TRACF1-C5 as a risk locus for rheumatoid arthritis--a genome-wide study</p>
            </title>
            <aug>
               <au>
                  <snm>Plenge</snm>
                  <fnm>RM</fnm>
               </au>
               <au>
                  <snm>Seielstad</snm>
                  <fnm>M</fnm>
               </au>
               <au>
                  <snm>Padyukov</snm>
                  <fnm>L</fnm>
               </au>
               <au>
                  <snm>Lee</snm>
                  <fnm>AT</fnm>
               </au>
               <au>
                  <snm>Remmers</snm>
                  <fnm>EF</fnm>
               </au>
               <au>
                  <snm>Ding</snm>
                  <fnm>B</fnm>
               </au>
               <au>
                  <snm>Liew</snm>
                  <fnm>A</fnm>
               </au>
               <au>
                  <snm>Khalili</snm>
                  <fnm>H</fnm>
               </au>
               <au>
                  <snm>Chandrasekaran</snm>
                  <fnm>A</fnm>
               </au>
               <au>
                  <snm>Davies</snm>
                  <fnm>LR</fnm>
               </au>
               <au>
                  <snm>Li</snm>
                  <fnm>W</fnm>
               </au>
               <au>
                  <snm>Tan</snm>
                  <fnm>AK</fnm>
               </au>
               <au>
                  <snm>Bonnard</snm>
                  <fnm>C</fnm>
               </au>
               <au>
                  <snm>Ong</snm>
                  <fnm>RT</fnm>
               </au>
               <au>
                  <snm>Thalamuthu</snm>
                  <fnm>A</fnm>
               </au>
               <au>
                  <snm>Pettersson</snm>
                  <fnm>S</fnm>
               </au>
               <au>
                  <snm>Liu</snm>
                  <fnm>C</fnm>
               </au>
               <au>
                  <snm>Tian</snm>
                  <fnm>C</fnm>
               </au>
               <au>
                  <snm>Chen</snm>
                  <fnm>WV</fnm>
               </au>
               <au>
                  <snm>Carulli</snm>
                  <fnm>JP</fnm>
               </au>
               <au>
                  <snm>Beckman</snm>
                  <fnm>EM</fnm>
               </au>
               <au>
                  <snm>Altshuler</snm>
                  <fnm>D</fnm>
               </au>
               <au>
                  <snm>Alfredsson</snm>
                  <fnm>L</fnm>
               </au>
               <au>
                  <snm>Criswell</snm>
                  <fnm>LA</fnm>
               </au>
               <au>
                  <snm>Amos</snm>
                  <fnm>CI</fnm>
               </au>
               <au>
                  <snm>Seldin</snm>
                  <fnm>MF</fnm>
               </au>
               <au>
                  <snm>Kastner</snm>
                  <fnm>DL</fnm>
               </au>
               <au>
                  <snm>Klareskog</snm>
                  <fnm>L</fnm>
               </au>
               <au>
                  <snm>Gregersen</snm>
                  <fnm>PK</fnm>
               </au>
            </aug>
            <source>N Engl J Med</source>
            <pubdate>2007</pubdate>
            <volume>357</volume>
            <fpage>1199</fpage>
            <lpage>209</lpage>
            <xrefbib>
               <pubidlist>
                  <pubid idtype="pmcid">2636867</pubid>
                  <pubid idtype="pmpid" link="fulltext">17804836</pubid>
                  <pubid idtype="doi">10.1056/NEJMoa073491</pubid>
               </pubidlist>
            </xrefbib>
         </bibl>
      </refgrp>
   </bm>
</art>
