<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Awk on PopGen Blog</title>
    <link>https://popgenblog.com/tags/awk/</link>
    <description>Recent content in Awk on PopGen Blog</description>
    <generator>Hugo -- 0.148.2</generator>
    <language>en-us</language>
    <lastBuildDate>Wed, 19 Aug 2026 22:17:00 +0200</lastBuildDate>
    <atom:link href="https://popgenblog.com/tags/awk/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Convert 23andMe, AncestryDNA, MyHeritage &amp; FTDNA Raw DNA to PLINK (BED/BIM/FAM)</title>
      <link>https://popgenblog.com/posts/raw-dna-to-plink/</link>
      <pubDate>Wed, 19 Aug 2026 22:17:00 +0200</pubDate>
      <guid>https://popgenblog.com/posts/raw-dna-to-plink/</guid>
      <description>&lt;p&gt;To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (&lt;code&gt;.bed&lt;/code&gt;, &lt;code&gt;.bim&lt;/code&gt;, &lt;code&gt;.fam&lt;/code&gt;), you will have to first convert the raw file to 23andMe format. You can then convert it with PLINK 1.9 using &lt;code&gt;--23file&lt;/code&gt;.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;converting-raw-dna-to-23andme-format-with-awk&#34;&gt;Converting Raw DNA to 23andMe Format with AWK&lt;/h2&gt;
&lt;p&gt;Windows users can use WSL to access &lt;code&gt;awk&lt;/code&gt;; see &lt;a href=&#34;https://popgenblog.com/posts/download-ancient-modern-dna-aadr/&#34;&gt;How to Download the AADR Dataset (Linux &amp;amp; WSL)&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If your DNA file is already in 23andMe format, skip this section.&lt;/p&gt;</description>
    </item>
    <item>
      <title>How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)</title>
      <link>https://popgenblog.com/posts/awk-subset-populations-genetics/</link>
      <pubDate>Mon, 10 Nov 2025 15:25:02 +0100</pubDate>
      <guid>https://popgenblog.com/posts/awk-subset-populations-genetics/</guid>
      <description>&lt;p&gt;In an earlier post, &lt;a href=&#34;https://popgenblog.com/posts/plink-pca-tutorial/&#34;&gt;PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)&lt;/a&gt;, I showed the manual way to build a subset from the &lt;code&gt;.ind/.fam&lt;/code&gt;. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using &lt;code&gt;awk&lt;/code&gt; that generates a PLINK &lt;code&gt;--keep&lt;/code&gt; file automatically from a list of populations.&lt;/p&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li&gt;Prepare a list of populations to keep:&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Create a text file (e.g. &lt;code&gt;pops&lt;/code&gt;) in the same directory as your reference &lt;code&gt;.ind&lt;/code&gt; and &lt;code&gt;.fam&lt;/code&gt;. Put one population label per line:&lt;/p&gt;</description>
    </item>
    <item>
      <title>How to Run ADMIXTURE (Unsupervised): Full Tutorial &amp; Python Plotting Script</title>
      <link>https://popgenblog.com/posts/admixture-unsupervised/</link>
      <pubDate>Sat, 02 Aug 2025 22:47:12 +0200</pubDate>
      <guid>https://popgenblog.com/posts/admixture-unsupervised/</guid>
      <description>&lt;p&gt;In this post, I’ll demonstrate how to estimate ancestry proportions using one of the most widely used tools in population genetics: &lt;a href=&#34;https://dalexander.github.io/admixture/download.html&#34;&gt;ADMIXTURE&lt;/a&gt;. ADMIXTURE is a model-based clustering algorithm that estimates individual ancestry proportions and ancestral allele frequencies from multilocus SNP genotypes.&lt;/p&gt;
&lt;hr&gt;
&lt;h3 id=&#34;preparing-the-dataset&#34;&gt;Preparing the Dataset&lt;/h3&gt;
&lt;p&gt;Download the appropriate ADMIXTURE binary and either place it in your dataset directory or make it globally accessible. For this run, I included a subset of West Asian populations along with a few adjacent populations (around 150 samples in total). Linkage Disequilibrium (LD) pruning was applied beforehand. If you&amp;rsquo;re unsure how to prune your dataset, refer to the previous post.&lt;/p&gt;</description>
    </item>
    <item>
      <title>SmartPCA Tutorial: How to Run PCA on Genetic Data</title>
      <link>https://popgenblog.com/posts/smartpca-tutorial/</link>
      <pubDate>Wed, 30 Jul 2025 20:47:41 +0200</pubDate>
      <guid>https://popgenblog.com/posts/smartpca-tutorial/</guid>
      <description>&lt;p&gt;This post is a continuation of the previous one, where I demonstrated how to perform PCA with PLINK. While PLINK’s PCA is great for quick, exploratory analysis, smartpca (part of the EIGENSOFT toolset) is particularly common in population-genetic and ancient-DNA studies.&lt;/p&gt;
&lt;p&gt;Smartpca can be compiled from the EIGENSOFT source or installed through conda. I covered the installation process in this earlier post: &lt;a href=&#34;https://popgenblog.com/posts/convert-eigenstrat-to-packedped/&#34;&gt;From EIGENSTRAT to PACKEDPED&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;As before, I’ll use a small subset. The focus here is on the technical process. One key difference in this post is that I’ll perform Linkage Disequilibrium (LD) pruning, which reduces redundancy between correlated SNPs before PCA.&lt;/p&gt;</description>
    </item>
    <item>
      <title>PLINK PCA Tutorial: Running PCA in PLINK (Commands &#43; Output)</title>
      <link>https://popgenblog.com/posts/plink-pca-tutorial/</link>
      <pubDate>Tue, 29 Jul 2025 16:00:00 +0000</pubDate>
      <guid>https://popgenblog.com/posts/plink-pca-tutorial/</guid>
      <description>&lt;p&gt;In this post, I’ll demonstrate how to perform a PCA on a PLINK dataset.
Before we begin, we need to prepare a subset of samples we&amp;rsquo;re interested in analyzing.&lt;/p&gt;
&lt;p&gt;To do this, we’ll extract sample information from the &lt;code&gt;.fam&lt;/code&gt; file.
But first, we need to identify the samples of interest. For example, those from a specific population such as Sardinians.&lt;/p&gt;
&lt;p&gt;The easiest way is to open the corresponding &lt;code&gt;.ind&lt;/code&gt; file and look at the population column, which is the third column in each row. Open the file in a text editor, and search for the population name, in this case, Sardinian.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
