<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>EIGENSTRAT on PopGen Blog</title>
    <link>https://popgenblog.com/tags/eigenstrat/</link>
    <description>Recent content in EIGENSTRAT on PopGen Blog</description>
    <generator>Hugo -- 0.148.2</generator>
    <language>en-us</language>
    <lastBuildDate>Fri, 17 Apr 2026 13:55:43 +0200</lastBuildDate>
    <atom:link href="https://popgenblog.com/tags/eigenstrat/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Downloading and Converting AADR v66</title>
      <link>https://popgenblog.com/posts/aadr-v66-download-and-conversion/</link>
      <pubDate>Fri, 17 Apr 2026 13:55:43 +0200</pubDate>
      <guid>https://popgenblog.com/posts/aadr-v66-download-and-conversion/</guid>
      <description>&lt;p&gt;Recently, in April 2026, new AADR versions were released on &lt;a href=&#34;https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/FFIDCW&#34;&gt;Harvard Dataverse&lt;/a&gt;. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable.&lt;/p&gt;</description>
    </item>
    <item>
      <title>How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)</title>
      <link>https://popgenblog.com/posts/awk-subset-populations-genetics/</link>
      <pubDate>Mon, 10 Nov 2025 15:25:02 +0100</pubDate>
      <guid>https://popgenblog.com/posts/awk-subset-populations-genetics/</guid>
      <description>&lt;p&gt;In an earlier post, &lt;a href=&#34;https://popgenblog.com/posts/plink-pca-tutorial/&#34;&gt;PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)&lt;/a&gt;, I showed the manual way to build a subset from the &lt;code&gt;.ind/.fam&lt;/code&gt;. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using &lt;code&gt;awk&lt;/code&gt; that generates a PLINK &lt;code&gt;--keep&lt;/code&gt; file automatically from a list of populations.&lt;/p&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li&gt;Prepare a list of populations to keep:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Create a text file (e.g. &lt;code&gt;pops&lt;/code&gt;) in the same directory as your reference &lt;code&gt;.ind&lt;/code&gt; and &lt;code&gt;.fam&lt;/code&gt;. Put one population label per line:&lt;/p&gt;</description>
    </item>
    <item>
      <title>Converting EIGENSTRAT/PACKEDANCESTRYMAP to PACKEDPED</title>
      <link>https://popgenblog.com/posts/convert-eigenstrat-to-packedped/</link>
      <pubDate>Tue, 29 Jul 2025 15:30:00 +0000</pubDate>
      <guid>https://popgenblog.com/posts/convert-eigenstrat-to-packedped/</guid>
      <description>&lt;p&gt;The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style &lt;code&gt;.geno/.snp/.ind&lt;/code&gt; dataset. This naming can be confusing: the &lt;code&gt;.snp&lt;/code&gt; and &lt;code&gt;.ind&lt;/code&gt; files are the usual EIGENSTRAT metadata files, but the &lt;code&gt;.geno&lt;/code&gt; file may either be plain-text EIGENSTRAT or binary PACKEDANCESTRYMAP.&lt;/p&gt;
&lt;p&gt;PACKEDPED format allows for easier downstream processing using the &lt;strong&gt;PLINK&lt;/strong&gt; toolset. With PLINK, it becomes straightforward to extract sample subsets, filter SNPs, and perform a wide range of analyses.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
