[{"content":"Ringbauer et al. argued in their 2025 paper, \u0026ldquo;Punic people were genetically diverse with almost no Levantine ancestors\u0026rdquo;, that Punic people had almost no Levantine ancestors. Their conclusion seems to be based mainly on qpAdm rotation results.\nThe source choices were distal. Instead of using Iron Age groups from around the Mediterranean, including the newly sequenced Akhziv Phoenician group, the study used Bronze Age groups. It also used Neolithic Ganj Dareh as a source.\nThe ADMIXTURE run did not give meaningful results, as they stated themselves. ADMIXTURE does not seem to be the best choice here, even leaving aside how meaningful its inferred components usually are.\nThe PCA-based ancestry estimate is also questionable, since it treats a sample’s position along a Sicily Bronze Age to North Africa cline as a direct ancestry proportion, at least as I understand their method.\nI modeled the Punic samples using F4Mix with a set of Iron Age Mediterranean reference populations, including the Phoenicians from Akhziv. Substantial Phoenician-like ancestry appears across several groups:\nMost notably, nearly every Carthage samples scores substantial Phoenician-like ancestry here (41.9% Akhziv on average), compared with only one sample in the study\u0026rsquo;s Carthago models. As a separate plausibility check, several Selinunte samples score Iron Age Greek ancestry. Selinunte was founded as a Greek colony.\nThese models might not be perfect, and F4Mix does not provide a formal fit test here. However, most of the studies tested qpAdm source combinations were rejected, and ADMIXTURE does not test model fit either. The F4Mix results are generally coherent, although some samples have noticeably worse chi-squared values, particularly both Cádiz samples, as well as individual samples from Tharros and Lilybaeum. The high-Akhziv Lilybaeum sample (I21856) has a higher chi-square, but it is not one of the poorly fitted outliers in the run.\nI ran one- and two-source qpAdm rotations for the ten individuals with the highest Akhziv weights (including those for which the optimizer did not terminate successfully) in F4Mix. Akhziv appears in the best feasible model for five of them. The chart shows each model’s source weights and p-value: the two p-values highlighted in red indicate models below the 0.05 threshold, so those models might be incomplete.\n","permalink":"https://popgenblog.com/posts/punic-levantine-ancestry/","summary":"Ringbauer et al. argued in their 2025 paper, \u0026ldquo;Punic people were genetically diverse with almost no Levantine ancestors\u0026rdquo;, that Punic people had almost no Levantine ancestors. Their conclusi","tags":[],"title":"The Mediterranean Punic World Was Genetically Diverse, With Substantial Levantine Ancestry"},{"content":"f4-statistics can be used to test asymmetries in allele sharing between populations. They measure the covariance between allele-frequency differences across two pairs of populations.\nTheory and formula An f4-statistic is the average, across SNPs, of the product of the allele-frequency differences between two pairs of populations: f4(A,B;C,D)=Ei[(pA,i−pB,i)(pC,i−pD,i)] f_4(A,B;C,D)=\\mathbb{E}_i\\left[(p_{A,i}-p_{B,i})(p_{C,i}-p_{D,i})\\right] f4​(A,B;C,D)=Ei​[(pA,i​−pB,i​)(pC,i​−pD,i​)]Multiplying out results in: f4(A,B;C,D)=Ei[pA,ipC,i−pA,ipD,i−pB,ipC,i+pB,ipD,i] f_4(A,B;C,D)=\\mathbb{E}_i\\left[p_{A,i}p_{C,i}-p_{A,i}p_{D,i}-p_{B,i}p_{C,i}+p_{B,i}p_{D,i}\\right] f4​(A,B;C,D)=Ei​[pA,i​pC,i​−pA,i​pD,i​−pB,i​pC,i​+pB,i​pD,i​]In the formula, AAA, BBB, CCC, and DDD are populations: pA,ip_{A,i}pA,i​, pB,ip_{B,i}pB,i​, pC,ip_{C,i}pC,i​, and pD,ip_{D,i}pD,i​ are their respective allele frequencies at SNP iii, and Ei\\mathbb{E}_iEi​ is the average across SNPs.\nThe expanded form can now also be written as shared-drift covariances:\nf4(A,B;C,D)=Cov⁡(A,C)−Cov⁡(A,D)−Cov⁡(B,C)+Cov⁡(B,D) f_4(A,B;C,D)=\\operatorname{Cov}(A,C)-\\operatorname{Cov}(A,D)-\\operatorname{Cov}(B,C)+\\operatorname{Cov}(B,D) f4​(A,B;C,D)=Cov(A,C)−Cov(A,D)−Cov(B,C)+Cov(B,D)Cov⁡(X,Y)\\operatorname{Cov}(X,Y)Cov(X,Y) stands for the shared drift covariance between populations XXX and YYY. A larger covariance indicates that the two populations share more allele-frequency changes and more shared genetic drift, while a smaller covariance indicates less shared drift.\n1. General interpretation The covariance terms can be grouped into two sets of cross-pairs:\nf4(A,B;C,D)=[Cov⁡(A,C)+Cov⁡(B,D)]−[Cov⁡(A,D)+Cov⁡(B,C)]. f_4(A,B;C,D)= \\left[\\operatorname{Cov}(A,C)+\\operatorname{Cov}(B,D)\\right] -\\left[\\operatorname{Cov}(A,D)+\\operatorname{Cov}(B,C)\\right]. f4​(A,B;C,D)=[Cov(A,C)+Cov(B,D)]−[Cov(A,D)+Cov(B,C)].Thus, without an outgroup assumption:\nWhen f4\u0026gt;0f_4\u0026gt;0f4​\u0026gt;0, the combined shared-drift covariance of (A,C)(A,C)(A,C) and (B,D)(B,D)(B,D) is greater than that of (A,D)(A,D)(A,D) and (B,C)(B,C)(B,C). When f4\u0026lt;0f_4\u0026lt;0f4​\u0026lt;0, the combined shared-drift covariance of (A,D)(A,D)(A,D) and (B,C)(B,C)(B,C) is greater than that of (A,C)(A,C)(A,C) and (B,D)(B,D)(B,D). When f4≈0f_4\\approx0f4​≈0, there is no detectable difference between the two pairings. These are relative statements about the two sums of covariance terms; without further assumptions, the statistic cannot identify which pair is responsible for the imbalance. For this reason, f4f_4f4​-statistics are often used with one population specified as a suitable outgroup, as described below.\n2. Interpretation with an outgroup When AAA is a suitable outgroup, it is expected to share approximately the same drift with CCC and DDD, so:\nCov⁡(A,C)≈Cov⁡(A,D). \\operatorname{Cov}(A,C)\\approx\\operatorname{Cov}(A,D). Cov(A,C)≈Cov(A,D).These terms cancel each other in expectation, leaving:\nf4(A,B;C,D)≈Cov⁡(B,D)−Cov⁡(B,C). f_4(A,B;C,D)\\approx\\operatorname{Cov}(B,D)-\\operatorname{Cov}(B,C). f4​(A,B;C,D)≈Cov(B,D)−Cov(B,C).Thus, the sign indicates the direction of excess allele sharing:\nf4\u0026lt;0 (negative)⇒B and C share more driftf4\u0026gt;0 (positive)⇒B and D share more driftf4≈0 or non-significant⇒no detectable difference in B’s affinity to C and D \\boxed{\\begin{aligned} f_4 \u0026amp;\u0026lt; 0\\; (\\text{negative}) \u0026amp;\u0026amp;\\Rightarrow B\\text{ and }C\\text{ share more drift} \\\\ f_4 \u0026amp;\u0026gt; 0\\; (\\text{positive}) \u0026amp;\u0026amp;\\Rightarrow B\\text{ and }D\\text{ share more drift} \\\\ f_4 \u0026amp;\\approx 0\\;\\text{or non-significant} \u0026amp;\u0026amp;\\Rightarrow \\text{no detectable difference in }B\\text{’s affinity to }C\\text{ and }D \\end{aligned}} f4​f4​f4​​\u0026lt;0(negative)\u0026gt;0(positive)≈0or non-significant​​⇒B and C share more drift⇒B and D share more drift⇒no detectable difference in B’s affinity to C and D​​This also shows that an f4-statistic measures relative excess shared drift: the sign shows which pair shares more drift, and the estimate shows how large the difference is.\nHowever, the estimate alone does not indicate whether the observed difference is statistically distinguishable from zero. This is assessed using the ZZZ-score:\nZ=f4^SE⁡(f4^) Z=\\frac{\\widehat{f_4}}{\\operatorname{SE}(\\widehat{f_4})} Z=SE(f4​​)f4​​​In this formula, the hat indicates a value estimated from the observed SNP data: f4^\\widehat{f_4}f4​​ is the f4f_4f4​-estimate, and SE⁡(f4^)\\operatorname{SE}(\\widehat{f_4})SE(f4​​) is its standard error. In AdmixPy, the standard error is estimated using a block jackknife. The sign of ZZZ is the same as the sign of f4^\\widehat{f_4}f4​​, while ∣Z∣\\lvert Z\\rvert∣Z∣ measures how many standard errors the estimate lies from zero. The ZZZ-score is a useful summary metric because it represents the estimate relative to its standard error. It is also usually on a more convenient numerical scale than the small decimal values of the f4f_4f4​-estimates themselves, which makes results easier to compare.\nFor example, Z=−3Z=-3Z=−3 means that the f4f_4f4​-estimate is three standard errors below zero and corresponds to a two-sided ppp-value of approximately 0.00270.00270.0027. More negative ZZZ-scores provide stronger evidence that BBB shares more drift with CCC than with DDD.\nThere is no definitive significance threshold. A value of ∣Z∣≥3\\lvert Z\\rvert\\ge3∣Z∣≥3 is often used as a convention, lower absolute ZZZ-scores could still be considered depending on the aim of the analysis.\nf4-Examples with AdmixPy In this previous post, I gave already some general absolute drift share examples in \u0026ldquo;admixtools2\u0026rdquo;, usually against only a single axis which mostly focus on sign convention.\nThe population names in the examples are dataset-specific. Replace them with the exact labels in the third column of your .ind file.\nBelow, I use AdmixPy. For installation instructions, see Introducing AdmixPy. If you already have AdmixPy installed, update it to the latest version first.\nAfter activating your virtual environment, start a Python REPL by entering python in the terminal. Then import AdmixPy and set the dataset prefix:\nimport admixpy as ap prefix = \u0026#34;v66_compatibility\u0026#34; 1. Which Northeast Asian source best represents the non-ANE ancestry in Tarim EMBA1? The Early Bronze Age Tarim Basin population, known from the well-preserved Tarim mummies, derived most of its ancestry from Ancient North Eurasians but also had a Northeast Asian-related contribution. To test which reference groups share additional drift with Tarim EMBA1 relative to Afontova Gora, run:\nap.f4( prefix, \u0026#34;Chimp\u0026#34;, [ \u0026#34;Russia_Vologda_Mesolithic\u0026#34;, \u0026#34;Turkey_Epipaleolithic\u0026#34;, \u0026#34;Iran_BeltCave_Mesolithic\u0026#34;, \u0026#34;China_TianyuanCave_UP\u0026#34;, \u0026#34;China_AmurRiverBasin_N\u0026#34;, \u0026#34;Russia_PrimorskyKrai_AmurRiver_N\u0026#34;, \u0026#34;USA_WA_Kennewick_8800BP\u0026#34;, \u0026#34;China_AmurRiverBasin_Mesolithic\u0026#34;, ], \u0026#34;Tarim_EMBA1\u0026#34;, \u0026#34;AfontovaGora_UP\u0026#34;, ) I use Chimp as a outgroup here, although an African outgroup such as Mbuti would also work for this example. For pop2, I pass a Python list with the groups to test in position BBB. Tarim EMBA1 is in position CCC, and the ANE-related Afontova Gora group in position DDD. Swapping these last two populations would reverse the signs, as would swapping the outgroup from position AAA to BBB.\nThis returns:\npop1 pop2 pop3 pop4 est se z p n 0 Chimp Russia_Vologda_Mesolithic Tarim_EMBA1 AfontovaGora_UP 0.000195698 0.000486525 0.4 0.688 164434 1 Chimp Turkey_Epipaleolithic Tarim_EMBA1 AfontovaGora_UP 0.000531405 0.000673312 0.79 0.43 141828 2 Chimp Iran_BeltCave_Mesolithic Tarim_EMBA1 AfontovaGora_UP -0.000203625 0.000895436 -0.23 0.82 63899 3 Chimp China_TianyuanCave_UP Tarim_EMBA1 AfontovaGora_UP -0.00158086 0.000593296 -2.66 0.008 150232 4 Chimp China_AmurRiverBasin_N Tarim_EMBA1 AfontovaGora_UP -0.00393877 0.000541723 -7.27 3.57e-13 153428 5 Chimp Russia_PrimorskyKrai_AmurRiver_N Tarim_EMBA1 AfontovaGora_UP -0.00353571 0.000539207 -6.56 5.48e-11 164427 6 Chimp USA_WA_Kennewick_8800BP Tarim_EMBA1 AfontovaGora_UP 7.71892e-05 0.000917437 0.08 0.933 62992 7 Chimp China_AmurRiverBasin_Mesolithic Tarim_EMBA1 AfontovaGora_UP -0.00385563 0.000615036 -6.27 3.63e-10 132437 As expected, the West Eurasian and Kennewick reference populations yield ZZZ-scores near zero. None shares significantly more drift with Tarim EMBA1 than with the Ancient North Eurasian Afontova Gora group.\nThe three Amur-related populations in position BBB have significant negative ZZZ-scores. Tianyuan also has a negative result (Z=−2.66Z=-2.66Z=−2.66) that approaches the conventional significance threshold and points in the direction of excess East Asian-related affinity in Tarim EMBA1. As explained earlier, a negative result in this arrangement indicates excess affinity between BBB (pop2) and CCC (pop3).\nDifferences in SNP coverage (column n) can affect f4f_4f4​-estimates and their ZZZ-scores. The SNP counts for the significant statistics above are not highly imbalanced, but I nevertheless reran the tests with allsnps=False for those three groups. This forces all three statistics to use the same intersecting set of SNPs:\nap.f4( prefix, \u0026#34;Chimp\u0026#34;, [ \u0026#34;China_AmurRiverBasin_N\u0026#34;, \u0026#34;Russia_PrimorskyKrai_AmurRiver_N\u0026#34;, \u0026#34;China_AmurRiverBasin_Mesolithic\u0026#34;, ], \u0026#34;Tarim_EMBA1\u0026#34;, \u0026#34;AfontovaGora_UP\u0026#34;, allsnps=False, ) This results in:\npop1 pop2 pop3 pop4 est se z p n 0 Chimp China_AmurRiverBasin_N Tarim_EMBA1 AfontovaGora_UP -0.00390121 0.000559719 -6.97 3.17e-12 130128 1 Chimp Russia_PrimorskyKrai_AmurRiver_N Tarim_EMBA1 AfontovaGora_UP -0.0037304 0.000592335 -6.3 3.02e-10 130128 2 Chimp China_AmurRiverBasin_Mesolithic Tarim_EMBA1 AfontovaGora_UP -0.00387299 0.000616228 -6.28 3.28e-10 130128 Amur River Basin Neolithic still has the largest absolute ZZZ-score, although the differences among the three results are modest.\nThese findings could be used for a qpAdm model. I would test Amur River Basin Neolithic alongside Afontova Gora as source populations and place Amur River Basin Mesolithic among the right populations to anchor the Northeast Asian source.\nThis approach can also be extended into a proximity ranking. Instead of comparing Tarim EMBA1 with a single population in pop4, one can pass a list of candidate populations and use a less redundant, broader set of relevant references in pop2. The resulting statistics can then be summarised with pandas and numpy, both of which are already installed as AdmixPy dependencies. Candidates whose f4f_4f4​-statistics show the smallest overall deviations from zero are the most similar to the target relative to that reference panel: the candidate and target have the most symmetric relationships to the populations in pop2.\n2. Phylogenetic placement and unrooted pairings An outgroup is not required. Under a simple tree without admixture, f4(A,B;C,D)f_4(A,B;C,D)f4​(A,B;C,D) is expected to be zero when AAA and BBB lie on one side of the internal split and CCC and DDD lie on the other.\nFor example:\nap.f4(prefix, \u0026#34;Andaman_100BP\u0026#34;, \u0026#34;Onge\u0026#34;, \u0026#34;Papuan\u0026#34;, \u0026#34;Han\u0026#34;) pop1 pop2 pop3 pop4 est se z p n 0 Andaman_100BP Onge Papuan Han 0.000207602 0.000267998 0.77 0.439 697025 This near-zero result represents the unrooted split (Andaman,Onge)∣(Papuan,Han)(\\text{Andaman},\\text{Onge})\\mid(\\text{Papuan},\\text{Han})(Andaman,Onge)∣(Papuan,Han). Placing the closely related Andaman and Onge in positions AAA and CCC instead:\nap.f4(prefix, \u0026#34;Andaman_100BP\u0026#34;, \u0026#34;Papuan\u0026#34;, \u0026#34;Onge\u0026#34;, \u0026#34;Han\u0026#34;) returns:\npop1 pop2 pop3 pop4 est se z p n 0 Andaman_100BP Papuan Onge Han 0.00971223 0.00033594 28.91 8.79e-184 697025 This large positive value shows that the combined covariance of the Andaman–Onge and Papuan–Han pairings is greater than that of the Andaman–Han and Papuan–Onge pairings. So it supports the same unrooted split as the first result.\nThe outgroup-based interpretation would be misleading here because Andaman is closely related to Onge and does not share approximately equal drift with Onge and Han. If Andaman were incorrectly treated as an outgroup, the second result might be read as evidence that Papuan shares more drift with Han than with Onge. Replacing Andaman with Mbuti tests that interpretation:\nap.f4(prefix, \u0026#34;Mbuti\u0026#34;, \u0026#34;Papuan\u0026#34;, \u0026#34;Onge\u0026#34;, \u0026#34;Han\u0026#34;) pop1 pop2 pop3 pop4 est se z p n 0 Mbuti Papuan Onge Han -0.000253514 0.00027654 -0.92 0.359 706613 With a suitable outgroup, there is no significant evidence that Papuan shares more drift with Han than with Onge. The large value in the previous arrangement was therefore driven by the Andaman–Onge relationship. This direct arrangement tests whether Han is closer to Onge or Papuan:\nap.f4(prefix, \u0026#34;Mbuti\u0026#34;, \u0026#34;Han\u0026#34;, \u0026#34;Onge\u0026#34;, \u0026#34;Papuan\u0026#34;) returns:\npop1 pop2 pop3 pop4 est se z p n 0 Mbuti Han Onge Papuan -0.00216992 0.00030816 -7.04 1.9e-12 706613 With Mbuti as the outgroup, the significantly negative value indicates that Han shares more drift with Onge than with Papuan (Z=−7.04Z=-7.04Z=−7.04). This does not contradict the earlier unrooted split: Papuan and Han can lie on the same side of that split without forming a clade in the rooted tree. A topology in which Papuan branches first, then Han, with Andaman and Onge as sisters, would result in the same unrooted split when Mbuti is excluded.\nIn qpGraph, these results support testing an internal node joining Andaman 100BP and Onge, with Han joining their lineage before Papuan. Placing these branches under a shared East Eurasian ancestral node would still require additional testing rather than following directly from these results. Tianyuan and Ust-Ishim might be useful reference groups for testing that deeper placement.\n3. Detecting deviations from a population tree with f4f_4f4​-statistics As mentioned above, under a simple bifurcating tree without admixture, one of the three possible arrangements of four populations should result in a non-significant f4f_4f4​-statistic around zero. If none does, their relationships cannot be represented by a single unrooted split. The following three arrangements test Kotias Klde Mesolithic, Samara Yamnaya, Barcin Neolithic, and Vologda Mesolithic.\nap.f4( prefix, \u0026#34;Georgia_KotiasKlde_Mesolithic\u0026#34;, \u0026#34;Turkey_Barcin_Neolithic-DG\u0026#34;, \u0026#34;Russia_Samara_EBA_Yamnaya\u0026#34;, \u0026#34;Russia_Vologda_Mesolithic\u0026#34;, ) ap.f4( prefix, \u0026#34;Georgia_KotiasKlde_Mesolithic\u0026#34;, \u0026#34;Russia_Samara_EBA_Yamnaya\u0026#34;, \u0026#34;Turkey_Barcin_Neolithic-DG\u0026#34;, \u0026#34;Russia_Vologda_Mesolithic\u0026#34;, ) ap.f4( prefix, \u0026#34;Georgia_KotiasKlde_Mesolithic\u0026#34;, \u0026#34;Russia_Vologda_Mesolithic\u0026#34;, \u0026#34;Turkey_Barcin_Neolithic-DG\u0026#34;, \u0026#34;Russia_Samara_EBA_Yamnaya\u0026#34;, ) To cover the three permutations, one population can remain fixed in position AAA, while each of the other three should be placed once in position BBB. The other two populations occupy positions CCC and DDD.\nThese return:\ntest est se z p n 0 0.000879912 0.000115676 7.61 2.81e-14 1241213 1 0.00364597 0.000167812 21.73 1.15e-104 1241213 2 0.00276606 0.000168548 16.41 1.59e-60 1241213 The three statistics are permutations of the same quartet. They indicate that no single unrooted split fits these four populations, meaning that at least one population is admixed. Based on prior knowledge, the likely admixed population here is the Bronze Age Yamnaya group.\n4. Outgroups and ancestry-pole contrasts in a population with sub-Saharan African admixture To compare the modern BedouinB group with the Early Medieval Bedouin group from Tell Qarassa along several ancestry axes, I first use Mbuti in position AAA, as in the first example:\nap.f4( prefix, \u0026#34;Mbuti\u0026#34;, [ \u0026#34;Turkey_Central_Boncuklu_PPN\u0026#34;, \u0026#34;Iran_GanjDareh_N\u0026#34;, \u0026#34;Jordan_PPNB\u0026#34;, \u0026#34;Dinka\u0026#34;, \u0026#34;Yoruba\u0026#34;, \u0026#34;Morocco_Iberomaurusian\u0026#34;, \u0026#34;Natufian\u0026#34;, ], \u0026#34;Syria_TellQarassa_EarlyMedieval\u0026#34;, \u0026#34;BedouinB\u0026#34;, ) This returns:\npop1 pop2 pop3 pop4 est se z p n 0 Mbuti Turkey_Central_Boncuklu_PPN Syria_TellQarassa_EarlyMedieval BedouinB -0.00163561 0.000199886 -8.18 2.78e-16 1216503 1 Mbuti Iran_GanjDareh_N Syria_TellQarassa_EarlyMedieval BedouinB -0.000982149 0.000189909 -5.17 2.32e-07 1210579 2 Mbuti Jordan_PPNB Syria_TellQarassa_EarlyMedieval BedouinB -0.00144038 0.000208195 -6.92 4.57e-12 917580 3 Mbuti Dinka Syria_TellQarassa_EarlyMedieval BedouinB -0.00031521 0.000223465 -1.41 0.158 705778 4 Mbuti Yoruba Syria_TellQarassa_EarlyMedieval BedouinB -4.26945e-05 7.75272e-05 -0.55 0.582 1232719 5 Mbuti Morocco_Iberomaurusian Syria_TellQarassa_EarlyMedieval BedouinB -0.000821848 0.000187372 -4.39 1.15e-05 1055293 6 Mbuti Natufian Syria_TellQarassa_EarlyMedieval BedouinB -0.00175053 0.000209975 -8.34 7.63e-17 711464 The negative results for the Near Eastern references indicate that they share more drift with Tell Qarassa than with BedouinB. African admixture could explain the reduced Near Eastern affinity of BedouinB. However, BedouinB is not significantly closer to either Dinka or Yoruba. For a predominantly West Eurasian population, this combination suggests ancestry that shifts BedouinB away from the Near Eastern profile. Sub-Saharan African admixture is a plausible explanation, although the Mbuti-based tests alone do not seem to indicate that.\nA more sensitive contrast can be constructed by replacing Mbuti with Ust-Ishim:\nap.f4( prefix, \u0026#34;Russia_UstIshim_IUP\u0026#34;, [ \u0026#34;Turkey_Central_Boncuklu_PPN\u0026#34;, \u0026#34;Iran_GanjDareh_N\u0026#34;, \u0026#34;Jordan_PPNB\u0026#34;, \u0026#34;Dinka\u0026#34;, \u0026#34;Yoruba\u0026#34;, \u0026#34;Morocco_Iberomaurusian\u0026#34;, \u0026#34;Natufian\u0026#34;, ], \u0026#34;Syria_TellQarassa_EarlyMedieval\u0026#34;, \u0026#34;BedouinB\u0026#34;, ) Now the returned table is:\npop1 pop2 pop3 pop4 est se z p n 0 Russia_UstIshim_IUP Turkey_Central_Boncuklu_PPN Syria_TellQarassa_EarlyMedieval BedouinB -0.000452437 0.000229414 -1.97 0.049 1216469 1 Russia_UstIshim_IUP Iran_GanjDareh_N Syria_TellQarassa_EarlyMedieval BedouinB 0.000233737 0.000230595 1.01 0.311 1210551 2 Russia_UstIshim_IUP Jordan_PPNB Syria_TellQarassa_EarlyMedieval BedouinB -0.000211399 0.000258228 -0.82 0.413 917557 3 Russia_UstIshim_IUP Dinka Syria_TellQarassa_EarlyMedieval BedouinB 0.00138816 0.000400902 3.46 0.000535 705767 4 Russia_UstIshim_IUP Yoruba Syria_TellQarassa_EarlyMedieval BedouinB 0.00118368 0.000234436 5.05 4.44e-07 1232680 5 Russia_UstIshim_IUP Morocco_Iberomaurusian Syria_TellQarassa_EarlyMedieval BedouinB 0.000333061 0.000242472 1.37 0.17 1055263 6 Russia_UstIshim_IUP Natufian Syria_TellQarassa_EarlyMedieval BedouinB -0.000703807 0.000270607 -2.6 0.009 711444 Ust-Ishim is approximately equally related to the West Eurasian ancestry shared by Tell Qarassa and BedouinB. The additional African ancestry in BedouinB nevertheless reduces its overall affinity to Ust-Ishim. Placing Dinka or Yoruba on the opposite side of the comparison makes this African shift easier to detect. The significant positive results place BedouinB closer to the sub-Saharan African side of the contrast than Tell Qarassa. The Near Eastern references now also remain closer to zero. Boncuklu and especially Natufian still show a tendency toward greater affinity with Tell Qarassa. Because of the imbalance in SNP counts across the SSA groups, it could be worth rerunning them separately with allsnps=Falsehere.\nThis ancestry-pole setup is mainly useful when the two populations differ along an African–non-African axis. For tests among Eurasian populations with no detectable SSA ancestry, an African population will usually be approximately symmetric to both, making the standard outgroup approach sufficient.\nIn another post, I’ll use qpAdm to model the examples shown here.\n","permalink":"https://popgenblog.com/posts/interpreting-f4-statistics-admixpy/","summary":"f4-statistics can be used to test asymmetries in allele sharing between populations. They measure the covariance between allele-frequency differences across two pairs of populations.\nTheory and formul","tags":[],"title":"Interpreting f4-Statistics with AdmixPy"},{"content":"Among the unreleased samples in the Akbari et al. dataset is an individual (IID: I25719) rumoured to be from Kish, dated in the supplementary table to the Early Dynastic III period, around 2450 BCE. His Y-DNA haplogroup is G-FT19860, a downstream branch of G2a, one of the main paternal lineages associated with Anatolian Neolithic farmers, and his mtDNA haplogroup is R0a1a.\nThe sample has decent coverage: 598,059 \u0026ldquo;compatibility\u0026rdquo; SNPs were called, 645,472 were missing.\nThis would make the individual the first Bronze Age lower Mesopotamian aDNA sample till date.\nA two-source qpAdm model for sample I25719 results in 71.4 ± 5.3% Nevalı Çori PPN-related ancestry and 28.6 ± 5.3% Ganj Dareh Neolithic-related ancestry, with a very good model fit (p = 0.753).\nReplacing Nevalı Çori with Çayönü still gives a good-fitting model, although the Zagros Neolithic-related contribution becomes less significant. A single-source Çayönü model also passes borderline in this right-group setup, with a p-value of about 0.07. I would not put too much weight on that, however, adding a few more right groups could make the model fail.\nThese results strongly suggest that this potential Sumerian sample ultimately derived most of his ancestry from Nevalı Çori-like Neolithic populations. This continuity could be mediated through still-unsampled Chalcolithic populations from the southern Mesopotamian region.\nI also compared I25719 directly with Nevalı Çori, by running a series of f4(Chimp,X;Nevalı C¸ori,I25719.AG)f_4(\\text{Chimp}, X; \\text{Nevalı Çori}, I25719.AG)f4​(Chimp,X;Nevalı C¸​ori,I25719.AG) statistics across relevant reference populations:\nPositive values indicate excess affinity of I25719 to the reference population, while negative values indicate excess affinity of Nevalı Çori. The pattern is consistent with the qpAdm model, with I25719 remaining relatively close to Nevalı Çori overall, while Nevalı Çori shows a greater western Anatolian affinity and I25719 a slight shift toward the eastern Fertile Crescent. The deviations are even smaller when Çayönü is used as the comparison population, although, as stated above, the qpAdm decomposition is weaker.\nIn f4-based compatibility rankings, I25719 does not show significant differentiation from neighboring Mesopotamian groups, like the Nemrik 9 Middle Assyrian-period individual (I6441, northern Iraq, ~1350 BCE), followed by the classical-period individual from Bahrain (AS_EMT). Kish had a substantial East Semitic presence during the Early Dynastic period, so whether this individual was specifically Sumerian cannot be definitely established without more context. However, I would not expect core Sumerian-speaking populations to have been genetically that different from this sample.\n","permalink":"https://popgenblog.com/posts/kish-i25719-neolithic-mesopotamian-ancestry/","summary":"Among the unreleased samples in the Akbari et al. dataset is an individual (IID: I25719) rumoured to be from Kish, dated in the supplementary table to the Early Dynastic III period, around 2450 BCE. H","tags":[],"title":"A Potential Early Dynastic Kish G2a Individual (I25719) with Upper Mesopotamian Neolithic Ancestry"},{"content":"Among the Akbari et al. dataset, there is a previously unreported individual (Sample IID: I41584/I41584_preQC) potentially from Anatolia. His terminal Y-DNA subclade is Y148982/Y106006 (TMRCA ~3100 BCE according to FTDNA), a lineage that also includes most modern West Asian and Near Eastern V1636 samples. His maternal haplogroup is HV29d1, also according to FTDNA.\nThe genotype PCA makes it look as if the closest ancient groups to this sample are Bronze Age Aegean groups, however, the same is the case for one of the Old Hittite samples from Kaman. Against all the coloured ancient samples of this PCA across all generated PCs make this unreleased sample closest to an Ovaören individual, though, still notably distant followed by Aegean samples.\nSince a genotype PCA is not very useful for assessing genetic affinity and ancestry modelling, I also did an f4-based compatibility ranking. In this analysis, the sample shows a distinctly Bronze Age Anatolian affinity:\nThe sample was tested against 11 ancient diverse and relevant reference populations using the 57,886 SNPs shared across all candidate groups. The closest match is Karum-period Kaman, followed by Bronze Age Ovaören and then Old Hittite Kaman.\nWith qpAdm, the sample can be successfully modelled as 87.0 ± 3.4% Old Hittite Kaman, 7.8 ± 3.8% Progress-2 Eneolithic Steppe, and 5.2 ± 1.6% Iron Gates Mesolithic. Dropping Progress-2 within this set-up makes the model still pass with a p-value of 0.34 (chisq 10.13).\nReplacing the Kaman Old Hittite-period main source with the Karum-period group still produces a passing model, but with slightly worse fit metrics. Replacing the Bronze Age Anatolian core source with Aegean proxies, makes the model fail within the same set-up.\nI41584 can also be modelled successfully with a single antiquity/medieval Anatolian source, with a similar fit. Later-period Anatolian references also tend to outperform the closest Bronze Age groups in the subsequent f4 compatibility rankings I did.\nBut even if the sample is from a later period, its profile still appears to be predominantly derived from Bronze Age Anatolians, which would make an earlier presence of the lineage among preceding Anatolian speaking groups likely. The lineage itself is already attested in Chalcolithic Arslantepe, indicating that its presence in Anatolia long predates historical periods.\nSee also: The Genetic Origins of the Proto-Anatolians\nNote: The supplementary table lists I41584 with a date mean of 800 BP. The individual is therefore medieval, not Bronze Age. But the ancestry described above remains unchanged.\n","permalink":"https://popgenblog.com/posts/anatolian-v1636-candidate/","summary":"Among the Akbari et al. dataset, there is a previously unreported individual (Sample IID: I41584/I41584_preQC) potentially from Anatolia. His terminal Y-DNA subclade is Y148982/Y106006 (TMRCA ~3100 BC","tags":[],"title":"A Potential Bronze Age Anatolian-Derived V1636 Sample (I41584)"},{"content":"Ust-Ishim is a sample identified from only a femur bone pulled out of eroding sand on the banks of the Irtysh River in western Siberia in 2008. Later analysis established that the bone belonged to a man estimated to have lived around 45,000 years ago. His genome is one of the earliest Upper Paleolithic WGS genomes.\nNeither Clearly East- nor West Eurasian What makes this sample particularly interesting is that he does not fall clearly into either the East or West Eurasian category. The initial Fu et al. study placed him before, or approximately at, the separation of subsequent eastern and western Eurasian populations, which include all modern Eurasian populations. Later graph modelling reached essentially the same conclusion; the best fit placed Ust-Ishim slightly towards the West Eurasian branch, but the uncertainty overlapped the East-West split.\nThis can also be tested with f4f_4f4​-statistics:\nThese tests show no significant excess affinity between Ust-Ishim and either Tianyuan or several Upper Paleolithic West Eurasian groups.\nShared Drift Between Ust-Ishim and Modern Populations In an outgroup f3f_3f3​-statistic test, Native Southern American groups like Surui and Karitiana share the most drift with Ust-Ishim out of the modern populations shown.\nDoes Ancient North Eurasian Ancestry Explain the Affinity? Using Mal’ta and Afontova Gora as references indicates that these two ANE-related groups cannot fully explain the excess affinity between Ust-Ishim with Surui and Karitiana. While these f4f_4f4​-statistics fall below the ∣Z∣≥3|Z| \\ge 3∣Z∣≥3 threshold used in some ancient-DNA studies, they point in the same direction and corroborate the affinity.\nReplacing the ANE references with the previous groups, Kostenki, Tianyuan, and Sunghir, results in significant deviations even under this stricter threshold.\n","permalink":"https://popgenblog.com/posts/ust-ishim-early-eurasian-ancestry/","summary":"Ust-Ishim is a sample identified from only a femur bone pulled out of eroding sand on the banks of the Irtysh River in western Siberia in 2008. Later analysis established that the bone belonged to a m","tags":[],"title":"Ust-Ishim: A 45,000-Year-Old Genome at the East–West Eurasian Split"},{"content":"To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (.bed, .bim, .fam), you will have to first convert the raw file to 23andMe format. You can then convert it with PLINK 1.9 using --23file.\nConverting Raw DNA to 23andMe Format with AWK Windows users can use WSL to access awk; see How to Download the AADR Dataset (Linux \u0026amp; WSL).\nIf your DNA file is already in 23andMe format, skip this section.\nBefore running these commands, extract the raw DNA file from its compressed archive (e.g.,.zip or .gz).\nFTDNA to 23andMe The assumed FTDNA format is:\nRSID,CHROMOSOME,POSITION,RESULT \u0026#34;rs4477212\u0026#34;,\u0026#34;1\u0026#34;,\u0026#34;82154\u0026#34;,\u0026#34;--\u0026#34; Convert it with:\nawk -F\u0026#39;[,\\t ]+\u0026#39; \u0026#39;BEGIN{OFS=\u0026#34;\\t\u0026#34;} NR\u0026gt;1 \u0026amp;\u0026amp; NF\u0026gt;=4 {gsub(/\u0026#34;/,\u0026#34;\u0026#34;); gsub(/\\r/,\u0026#34;\u0026#34;,$4); print $1,$2,$3,$4}\u0026#39; input.csv \u0026gt; 23file.txt AncestryDNA to 23andMe The assumed AncestryDNA format is tab-separated with five columns. Comment lines beginning with # will be skipped by the command:\nrsid\tchromosome\tposition\tallele1\tallele2 rs3094315\t1\t752566\t0\t0 Convert the two allele columns into a single genotype column. Missing genotypes represented as 0 are converted to the 23andMe missing-genotype notation --:\nawk -F\u0026#39;\\t\u0026#39; \u0026#39;BEGIN{OFS=\u0026#34;\\t\u0026#34;} !/^#/ \u0026amp;\u0026amp; $1!=\u0026#34;rsid\u0026#34; \u0026amp;\u0026amp; NF\u0026gt;=5 { gsub(/\\r/,\u0026#34;\u0026#34;,$5) chr=$2 if (chr==\u0026#34;23\u0026#34;) chr=\u0026#34;X\u0026#34; else if (chr==\u0026#34;24\u0026#34;) chr=\u0026#34;Y\u0026#34; else if (chr==\u0026#34;25\u0026#34;) chr=\u0026#34;X\u0026#34; else if (chr==\u0026#34;26\u0026#34;) chr=\u0026#34;MT\u0026#34; genotype=(($4==\u0026#34;0\u0026#34; || $5==\u0026#34;0\u0026#34;) ? \u0026#34;--\u0026#34; : $4 $5) print $1,chr,$3,genotype }\u0026#39; input.txt \u0026gt; 23file.txt MyHeritage to 23andMe The assumed MyHeritage format is:\nRSID,CHROMOSOME,POSITION,RESULT rs12564807,1,734462,-- Convert it with:\nawk -F\u0026#39;,\u0026#39; \u0026#39;BEGIN{OFS=\u0026#34;\\t\u0026#34;} $1!=\u0026#34;RSID\u0026#34; \u0026amp;\u0026amp; NF\u0026gt;=4 {gsub(/^[ \\t]+|[ \\t\\r]+$/,\u0026#34;\u0026#34;,$1); gsub(/^[ \\t]+|[ \\t\\r]+$/,\u0026#34;\u0026#34;,$4); print $1,$2,$3,$4}\u0026#39; input.txt \u0026gt; 23file.txt The resulting 23file.txt should contain four tab-separated columns:\nrsid\tchromosome\tposition\tgenotype A header is not required by PLINK.\nConverting 23andMe Format to PLINK For this step, use PLINK 1.9. PLINK 2 currently lists --23file, but it is not implemented.\nYou can download PLINK 1.9 from the PLINK 1.9 page and select the appropriate binary for your operating system.\nBefore running the command, make sure the PLINK executable is available in your system PATH. This allows you to run PLINK simply by typing plink in the terminal. Alternatively, move the executable to the directory containing your DNA file and run it from there on Linux, macOS, or WSL:\n./plink --23file 23file.txt 0 sample1 1 1 --make-bed --out sample If PLINK is in your PATH, use:\nplink --23file 23file.txt 0 sample1 1 1 --make-bed --out sample The relevant arguments are:\n23file.txt = input file in 23andMe format 0 = family ID (FID) sample1 = individual ID (IID) 1 = sex: 1 = male, 2 = female, 0 = unknown 1 = 1 means unaffected/control and 2 means affected/case sample = output filename prefix After conversion, you will have the standard PLINK binary dataset:\nsample.bed sample.bim sample.fam These files can then be used for downstream PLINK analyses or merged with other compatible datasets. To merge with the AADR, you can convert them with PLINK 2 to PACKEDANCESTRYMAP or TGENO and merge them using mergeit.\nImportant: This procedure only converts the file format. It does not handle strand orientation. If you plan to merge the result with AADR or another PLINK dataset, you might need to handle or exclude SNPs with strand or allele conflicts. If you prefer an all-in-one solution, see: Convert Raw DNA Files to EIGENSTRAT for ADMIXTOOLS and Merge with AADR.\n","permalink":"https://popgenblog.com/posts/raw-dna-to-plink/","summary":"To convert raw DNA data from 23andMe, AncestryDNA, MyHeritage, or FamilyTreeDNA (FTDNA) to PLINK binary format (.bed, .bim, .fam), you will have to first convert the raw file to 23andMe format. You ca","tags":["PLINK","awk"],"title":"Convert 23andMe, AncestryDNA, MyHeritage \u0026 FTDNA Raw DNA to PLINK (BED/BIM/FAM)"},{"content":"I was inferring genetic traits of ancient individuals, among them the \u0026ldquo;Cheddar Man\u0026rdquo;, whose pigmentation phenotype I thought was well established from his genotype.\nHowever, most of the 58 trait markers I was checking for could not be called reliably (with MAPQ ≥30 and base quality ≥30). Most markers had no reads at all, several others were supported by only a single read. This included markers like HERC2/OCA2 rs12913832 for eye colour, likewise SLC24A5 and SLC45A2 used to infer skin pigmentation.\nSo I switched to Loschbour, a Mesolithic hunter-gatherer whose remains were discovered at a rock shelter in Luxembourg’s Müllerthal region. His published genome has about 22× coverage, making him a much more suitable candidate for this type of analysis.\nRock shelter in Luxembourg’s Müllerthal region where the Mesolithic Loschbour remains were discovered. Photo: Cayambe / Wikimedia Commons, CC BY-SA 3.0.\nAppearance and some genetic traits of the Loschbour individual Eye colour: likely blue Marker Loschbour genotype Read support HERC2/OCA2 rs12913832 G/G 21/21 reads Loschbour was G/G at rs12913832, with all 21 reads supporting G. This strongly implies blue eyes.\nHair colour and texture Hair colour Gene / marker Genotype TYR rs1042602 C/C EXOC2 rs4959270 A/A SLC45A2 rs28777 C/A TYRP1 rs683 A/A SLC24A4 rs2402130 G/A KITLG rs12821256 T/T PIGU/ASIP rs2378249 A/A HERC2 rs12913832 G/G OCA2 rs1800407 C/C SLC45A2 rs16891982 C/C IRF4 rs12203592 T/T The HIrisPlex hair-colour markers were well covered, with 9 to 23 reads per site.\nThe major MC1R red-hair variants rs1805007, rs1805008, and rs1805009 were absent, likewise the blond-derived allele at KITLG rs12821256. Loschbour was also ancestral C/C at SLC45A2 rs16891982. Together, these results favor dark hair.\nIn the 2014 HIrisPlex analysis reported in \u0026ldquo;Ancient human genomes suggest three ancestral populations for present-day Europeans\u0026rdquo;, the probability of dark hair was 97.8%: 57.9% for black and 41.3% for brown. A later HIrisPlex-S reanalysis in \u0026ldquo;Ancient genomes indicate population replacement in Early Neolithic Britain\u0026rdquo; (2019) inferred probabilities of 53.2% for brown hair and 46.3% for black.\nHair texture At TCHH rs11803731, Loschbour was heterozygous:\nMarker Genotype Read support TCHH rs11803731 A/T 8 reads: 4 A, 4 T The T allele has been associated with a greater likelihood of straight hair.\nThe derived EDAR V370A allele, associated with thick straight hair and several other traits in some East Asian populations, was absent.\nBeard density The beard-density markers showed mixed tendencies.\nGene / marker Genotype Reported tendency EDAR rs365060 C/C Thicker beard growth rs117717824 G/G Thinner beard growth LNX1 rs4864809 Heterozygous Intermediate PREP rs6901317 Heterozygous Intermediate Skin pigmentation Loschbour carried the ancestral state at both SLC24A5 rs1426654 and SLC45A2 rs16891982, with strong read support: 20/20 reads for SLC24A5 and 17/17 reads for SLC45A2.\nThe derived alleles, which Loschbour did not have, are associated with lighter pigmentation in later Western Eurasians. Two markers, however, are not enough to reconstruct Loschbour\u0026rsquo;s exact skin tone.\nThe same HIrisPlex-S analysis, incorporating a larger set of pigmentation variants, assigned Loschbour an 89.3% probability of having an “intermediate” skin tone.\nFor comparison, I applied the scoring method used by the YSEQ Phenotype Predictor to the phenotype markers. Since I no longer had the BAM at this point, I could only score the markers I had already extracted. Using the nine available markers (out of 12 used by the predictor) resulted in scores of 38.2% Light/pale, 37.2% Moderate and 24.6% Dark/olive.\nFreckling: strong genetic predisposition Marker Genotype Read support IRF4 rs12203592 T/T 21 reads Loschbour was homozygous T/T at IRF4 rs12203592. The T allele at this marker has been strongly associated with freckling and increased sun sensitivity, with T/T showing the strongest tendency.\nHeight: shorter stature I also analysed the genome using the PGS002804 polygenic score for height (based on the Yengo et al. height score).\nThe score uses 1,099,005 variants. For Loschbour:\n1,089,338 variants had confident genotype calls (call rate of 99.12%). The usable reads had a mean depth of 20.3×. Matching these calls against the 1000 Genomes reference panel left 1,083,087 variants for the comparison. Loschbour\u0026rsquo;s matched raw height PGS was 0.9671. In comparison with 503 present-day Europeans from the 1000 Genomes Project, this was 3.23 standard deviations below the mean. No individual in that panel had a lower score.\nReference population N Mean PGS SD Loschbour Z-score Empirical percentile All Europeans (EUR) 503 3.7185 0.8510 −3.23 \u0026lt;0.2% CEU 99 4.2754 0.7537 −4.39 \u0026lt;1.0% GBR 91 4.0883 0.6331 −4.93 \u0026lt;1.1% FIN 99 3.6075 0.8153 −3.24 \u0026lt;1.0% IBS 107 2.9628 0.7023 −2.84 \u0026lt;0.9% TSI 107 3.7470 0.6788 −4.10 \u0026lt;0.9% Relative to this present-day European reference panel, the score points to a strong genetic tendency toward shorter adult stature.\nHis skeleton has been estimated at approximately 1.60 m (5′3″) tall, which is in line with the polygenic height score.\nABO blood group rs8176719: the O-frameshift Eight high-quality reads covered this site, all with the GRCh37 reference/deletion state, supporting the frameshift allele resulting in blood group O.\nA/B-discriminating markers Marker Observation rs8176746 11/11 reads = G rs8176747 13/13 reads = C Neither marker carried the common B variant.\nTherefore, Loschbour was very likely blood group O, with an O/O genotype.\nRhesus blood group The mean sequencing depth across RHD and RHCE was:\nRHD mean depth = 3.988× RHCE mean depth = 7.840× RHD / RHCE = 0.509 Using the common whole-genome sequencing copy-number approximation,\n(RHD coverage / RHCE coverage) × 2 the estimated RHD copy number is:\n≈ 1.02 copies Several RHD exons also had sequence coverage. The estimated copy number of about one fits an RHD-hemizygous genotype: one chromosome with RHD and one without it.\nThis means Loschbour was probably RhD-positive, with possibly a D/d configuration.\nOther common Rh antigens At the major C/c-associated marker in RHCE (rs676785), all 13 reads carried G, supporting the c state.\nAt rs609320, the 9 C and 4 G reads indicate E/e heterozygosity.\nThus, the probable Rh phenotype is D+ C− c+ E+ e+. This is less certain than the type O result.\nOther inferred traits Lactose digestion At MCM6/LCT rs4988235, Loschbour lacked the lactase-persistence allele, with 19 reads supporting the non-persistence state. rs182549 likewise carried the non-persistence variant. Together, these markers indicate lactase non-persistence.\nEarwax and apocrine secretion (ABCC11 rs17822931) Loschbour lacked the ABCC11 variant responsible for dry earwax and reduced apocrine secretion. The expected phenotype is therefore wet-type earwax with normal apocrine secretion.\nBitter-taste perception All three TAS2R38 markers had good coverage. The result is consistent with relatively strong sensitivity to bitter compounds.\nConclusion These results support several commonly reported traits in Loschbour, including light eyes, dark hair, lactase non-persistence, and blood group O.\nFor skin pigmentation, the ancestral SLC24A5 and SLC45A2 states alone would suggest relatively darker pigmentation, but broader analyses seem to place Loschbour closer to an intermediate category.\n","permalink":"https://popgenblog.com/posts/loschbour-phenotype/","summary":"I was inferring genetic traits of ancient individuals, among them the \u0026ldquo;Cheddar Man\u0026rdquo;, whose pigmentation phenotype I thought was well established from his genotype.\nHowever, most of the 58 ","tags":[],"title":"Genetic Traits of Loschbour: Appearance, Height, Blood Type, and More"},{"content":"Last week I published F4Mix, a tool for fitting modern and ancient DNA samples against a pool of source populations, usually ancient ones. F4Mix estimates, for each target, the non-negative mixture of reference populations whose covariance-aware f4 profile best matches it. This makes it useful for testing every sample against the same sources.\nWith a proper setup, the tool gives meaningful results, and can reveal both substructure and clear outliers within a site.\nAn example run: The files relevant for running F4Mix are run_model.py for setting up the model configuration, and optionally plot_weights.py for plotting the results. The dataset path and setup are hardcoded, and need to be adjusted directly in the files.\nThe core setup is similar to a qpAdm run: you define the dataset prefix (using the same formats supported by AdmixPy), the targets, the sources, the references (right groups), and an outgroup. The model fitting is done automatically.\nThe quality of results depends on the selected sources, the right groups (which should anchor ancestry axes for the sources), and the usable SNP count of each sample. Currently, the default warns when the minimum effective SNP count of any f4-statistic drops below 50,000. This does not necessarily mean that results below that threshold are unusable. Historical plausibility and the output metrics can be used to verify the results. The example run script also contains a section to exclude specific samples by their IIDs. Comment those lines out if the samples to exclude are not present in any population of the specific run.\n","permalink":"https://popgenblog.com/posts/f4mix/","summary":"Last week I published F4Mix, a tool for fitting modern and ancient DNA samples against a pool of source populations, usually ancient ones. F4Mix estimates, for each target, the non-negative mixture of","tags":[],"title":"F4Mix: Sample-Wise Ancestry Fitting with f4 Statistics"},{"content":"f3-statistics are used to test if populations are admixed or to measure shared genetic drift between two populations relative to an outgroup.\nThis post explains the theory behind admixture f3-statistics and shows how to run admixture f3 tests with AdmixPy.\nIf you want to skip the theoretical part, you can jump to Running admixture f3-statistics in AdmixPy.\nWhat is an f3-statistic? For three populations, the statistic is written as:\nf3(A;B,C)=Ei[(pA,i−pB,i)(pA,i−pC,i)] f_3(A;B,C)=\\mathbb{E}_i\\left[(p_{A,i}-p_{B,i})(p_{A,i}-p_{C,i})\\right] f3​(A;B,C)=Ei​[(pA,i​−pB,i​)(pA,i​−pC,i​)]Here, AAA is in the target position. Populations BBB and CCC are the reference populations. The values pA,ip_{A,i}pA,i​, pB,ip_{B,i}pB,i​, and pC,ip_{C,i}pC,i​ are the allele frequencies in populations AAA, BBB, and CCC, respectively, at SNP iii. The expectation is an average across SNPs.\nThe statistic multiplies two allele-frequency differences. It therefore assesses whether AAA differs from BBB and CCC in the same direction.\nAt a given SNP, the product is positive when the allele frequency in AAA is higher than in both reference populations or lower than in both. It is negative when the allele frequency in AAA lies between those of BBB and CCC. The f3-statistic averages these products across SNPs.\nf3 derived from f2 The f2-statistic is the mean squared allele-frequency difference between two populations:\nf2(X,Y)=E[(pX−pY)2] f_2(X,Y)=E[(p_X-p_Y)^2] f2​(X,Y)=E[(pX​−pY​)2]Expanding the three pairwise f2 distances gives:\nf3(A;B,C)=f2(A,B)+f2(A,C)−f2(B,C)2 \\boxed{ f_3(A;B,C)= \\frac{f_2(A,B)+f_2(A,C)-f_2(B,C)}{2} } f3​(A;B,C)=2f2​(A,B)+f2​(A,C)−f2​(B,C)​​Hence, the value is negative when:\nf2(B,C)\u0026gt;f2(A,B)+f2(A,C) f_2(B,C)\u0026gt;f_2(A,B)+f_2(A,C) f2​(B,C)\u0026gt;f2​(A,B)+f2​(A,C)This means that the two references are farther from each other than their combined distances to the target.\nExample, the three f2 distances are:\nf2(A,B)=0.04,f2(A,C)=0.05,f2(B,C)=0.12 f_2(A,B)=0.04, \\qquad f_2(A,C)=0.05, \\qquad f_2(B,C)=0.12 f2​(A,B)=0.04,f2​(A,C)=0.05,f2​(B,C)=0.12Then:\nf3(A;B,C)=0.04+0.05−0.122=−0.015 f_3(A;B,C)=\\frac{0.04+0.05-0.12}{2}=-0.015 f3​(A;B,C)=20.04+0.05−0.12​=−0.015The target occupies an intermediate position between the two references in allele-frequency space. That is the basic geometry behind an admixture f3-statistic.\nf3 can also be written as an f4-statistic with a repeated population:\nf3(A;B,C)=f4(A,B;A,C) f_3(A;B,C)=f_4(A,B;A,C) f3​(A;B,C)=f4​(A,B;A,C)This shows that the f3 admixture test is a special case of an f4-statistic. f4 statistics can also be used to test admixture using four populations, a topic I will possibly cover in another post.\nWhy a negative f3 statistic detects admixture Suppose target AAA formed through admixture between BBB and CCC:\npA=αpB+(1−α)pC+εA, p_A=\\alpha p_B+(1-\\alpha)p_C+\\varepsilon_A, pA​=αpB​+(1−α)pC​+εA​,where α\\alphaα is the ancestry proportion from BBB, and εA\\varepsilon_AεA​ represents drift accumulated after admixture.\nAssuming this later drift is independent of the difference between the sources:\nf3(A;B,C)=dA−α(1−α)f2(B,C) \\boxed{ f_3(A;B,C)=d_A-\\alpha(1-\\alpha)f_2(B,C) } f3​(A;B,C)=dA​−α(1−α)f2​(B,C)​where\ndA=E[εA2]. d_A=E[\\varepsilon_A^2]. dA​=E[εA2​].The admixture term is negative, whereas subsequent drift in the target contributes positively. Therefore, f3 becomes negative when\nα(1−α)f2(B,C)\u0026gt;dA. \\alpha(1-\\alpha)f_2(B,C)\u0026gt;d_A. α(1−α)f2​(B,C)\u0026gt;dA​.This also explains why an admixed population does not necessarily produce a negative result. The estimate can stay positive when the ancestry contribution from one source is small, the sources are closely related (resulting in a small f2(B,C)f_2(B, C)f2​(B,C) distance; hence a smaller negative admixture term), or the target had substantial drift after admixture.\nThe tree interpretation Consider a simple population tree with no admixture:\nA | | a | M / \\ b / \\ c / \\ B C Here, MMM is the branch point connecting the three populations. The values aaa, bbb, and ccc represent the genetic drift accumulated along each branch.\nOn an additive tree, the f2f_2f2​ distance between two populations is the sum of the branch lengths along the path connecting them. The path from AAA to BBB contains branches aaa and bbb, so\nf2(A,B)=a+b. f_2(A,B)=a+b. f2​(A,B)=a+b.Likewise,\nf2(A,C)=a+c f_2(A,C)=a+c f2​(A,C)=a+cand\nf2(B,C)=b+c. f_2(B,C)=b+c. f2​(B,C)=b+c.Now substitute these distances into the f2f_2f2​ formula:\nf3(A;B,C)=f2(A,B)+f2(A,C)−f2(B,C)2. f_3(A;B,C) =\\frac{f_2(A,B)+f_2(A,C)-f_2(B,C)}{2}. f3​(A;B,C)=2f2​(A,B)+f2​(A,C)−f2​(B,C)​.f3(A;B,C)=(a+b)+(a+c)−(b+c)2=2a+b+c−b−c2=2a2=a \\begin{aligned} f_3(A;B,C) \u0026amp;=\\frac{(a+b)+(a+c)-(b+c)}{2} \\\\ \u0026amp;=\\frac{2a+b+c-b-c}{2} \\\\ \u0026amp;=\\frac{2a}{2} \\\\ \u0026amp;=a \\end{aligned} f3​(A;B,C)​=2(a+b)+(a+c)−(b+c)​=22a+b+c−b−c​=22a​=a​Therefore, f3(A;B,C)f_3(A;B,C)f3​(A;B,C) equals the drift along the branch connecting AAA to MMM. Because a branch length cannot be negative, f3f_3f3​ cannot be negative on a strictly additive population tree. A significantly negative estimate therefore rejects this tree topology, and indicates that AAA has ancestry related to both BBB and CCC.\nSignificance and Z-scores The Z-score measures how many standard errors the estimate lies from zero:\nZ=f^3SE⁡(f^3) Z=\\frac{\\widehat f_3}{\\operatorname{SE}(\\widehat f_3)} Z=SE(f​3​)f​3​​For an admixture test, Z\u0026lt;−3Z\u0026lt;-3Z\u0026lt;−3 is usually used to indicate a significantly negative f3 estimate.\nAdmixPy also reports a two-sided p-value:\np=2Φ(−∣Z∣) p=2\\Phi(-|Z|) p=2Φ(−∣Z∣)A Z-score of −3-3−3 means that the f3 estimate is three standard errors below zero. AdmixPy reports a two-sided p-value of approximately 0.00270.00270.0027 for this result. This indicates that such an extreme estimate would be unlikely if the true f3 value were zero. For the directional hypothesis f3\u0026lt;0f_3\u0026lt;0f3​\u0026lt;0, the one-sided p-value is Φ(Z)\\Phi(Z)Φ(Z).\nRunning admixture f3-statistics in AdmixPy For installation instructions, see Introducing AdmixPy. Start Python (or set up a Python file) in the folder containing your genotype files. Then import the package:\nimport admixpy as ap Set the dataset prefix:\nprefix = \u0026#34;v66_compatibility\u0026#34; Run one admixture f3-statistic with:\nap.f3(prefix, pop1=\u0026#34;Japanese\u0026#34;, pop2=\u0026#34;Japan_Chiba_HG_Jomon\u0026#34;, pop3=\u0026#34;China_Shandong_Dinggong_LN\u0026#34;) This returns:\npop1 pop2 pop3 est se z p n 0 Japanese Japan_Chiba_HG_Jomon China_Shandong_Dinggong_LN -0.00926283 0.000606282 -15.28 1.07e-52 980456 The significantly negative result (z \u0026lt; -3) suggests that Japanese are admixed between Jomon Hunter-Gatherer-related ancestry and mainland East Asian Neolithic-farmer-related ancestry.\nAnother example which produces a significantly negative estimate:\nap.f3(prefix, pop1=\u0026#34;Morocco_EN\u0026#34;, pop2=\u0026#34;Morocco_Iberomaurusian\u0026#34;, pop3=\u0026#34;Spain_EN\u0026#34;) Result:\npop1 pop2 pop3 est se z p n 0 Morocco_EN Morocco_Iberomaurusian Spain_EN -0.0197945 0.00112779 -17.55 5.79e-69 879916 The significantly negative result suggests that Morocco Neolithic is admixed between Iberomaurusian-related ancestry and ancestry related to ANF-derived groups, here Spain Early Neolithic in particular.\nThese are simpler Neolithic examples. Often, even a population with expected admixture does not produce a significantly negative f3-statistic for reasons mentioned above.\nUnexpected population pairs can produce significantly negative results. Testing a wide range of combinations rather than only what seems directly plausible is useful. aDNA studies often report admixture f3-statistics as large supplementary tables with many combinations of target and reference populations.\nAn example:\npop1 pop2 pop3 est se z p n 0 Karelian Norwegian Korean -0.00281868 0.000471994 -5.97 2.35e-09 274624 Here, the negative f3 value is statistically significant (z=-5.97). This implies that Karelians are admixed between ancestries related to northern Europeans and eastern Eurasians. Korean is used here only for demonstration; better sources for the Uralic-associated East Asian ancestry would be available.\nAdmixPy returns the following columns:\nest: the estimated f3-statistic; se: its block-jackknife standard error; z: the estimate divided by its standard error; p: the two-sided p-value; n: the number of contributing SNPs. f3 does not quantify ancestry proportions. Methods such as qpAdm are needed to estimate ancestry proportions.\nTesting several source pairs AdmixPy can run many f3 tests at once. To do that, create a pandas data frame with one row for each population triplet and columns named pop1, pop2, and pop3. The same target can then be tested against several source pairs, or different targets can be tested against different references:\nimport pandas as pd tests = pd.DataFrame({ \u0026#34;pop1\u0026#34;: [\u0026#34;Target\u0026#34;, \u0026#34;Target\u0026#34;, \u0026#34;Target\u0026#34;], \u0026#34;pop2\u0026#34;: [\u0026#34;Reference_A\u0026#34;, \u0026#34;Reference_A\u0026#34;, \u0026#34;Reference_B\u0026#34;], \u0026#34;pop3\u0026#34;: [\u0026#34;Reference_B\u0026#34;, \u0026#34;Reference_C\u0026#34;, \u0026#34;Reference_C\u0026#34;], }) ap.f3(prefix, tests).sort_values(\u0026#34;z\u0026#34;) The call returns one result for each row in tests. Sorting by z places the most negative statistics first.\nUnnormalized and normalized f3 statistics in AdmixPy For direct genotype data, ap.f3(...) returns a normalized statistic by default. It divides the corrected f3 numerator by a sample-size adjusted estimate of the target heterozygosity. Use outgroupmode=True to return the unnormalized f3 numerator. Passing apply_corr=False keeps SNPs with fewer than two allele observations, the resulting estimate is sampling-biased.\nSetting apply_corr=False can be necessary when a pseudohaploid singleton is used as the target (pop1) or as a repeated source because a finite-sample correction cannot be computed in those cases. A singleton used only as a distinct source can still be analyzed with apply_corr=True, even when pseudohaploid.\nWhen ap.f3(...) is given an f2 cache or F2Blocks object, it returns unnormalized f3 values; outgroupmode has no effect in that case. This is equivalent in scale to direct-genotype f3 with outgroupmode=True.\nFinite-sample correction Population allele frequencies are estimated from a finite number of observed alleles. This creates sampling noise.\nThe target AAA appears in both allele-frequency differences. For three distinct populations, AdmixPy subtracts this correction at each SNP:\nδA,i=p^A,i(1−p^A,i)cA,i−1, \\delta_{A,i}= \\frac{\\widehat p_{A,i}(1-\\widehat p_{A,i})} {c_{A,i}-1}, δA,i​=cA,i​−1p​A,i​(1−p​A,i​)​,where cA,ic_{A,i}cA,i​ is the observed allele count in population AAA at SNP iii. The corrected SNP value is:\ng^i=(p^A,i−p^B,i)(p^A,i−p^C,i)−δA,i \\widehat g_i= (\\widehat p_{A,i}-\\widehat p_{B,i}) (\\widehat p_{A,i}-\\widehat p_{C,i})-\\delta_{A,i} g​i​=(p​A,i​−p​B,i​)(p​A,i​−p​C,i​)−δA,i​Because the correction contains cA,i−1c_{A,i}-1cA,i​−1, it requires at least two observed alleles in the repeated population. A pseudohaploid singleton provides only one observed allele, so this correction is undefined when that population is repeated. A singleton used only as a distinct source does not need this repeated-population correction.\nBlock jackknifing for the standard error Nearby SNPs are correlated through linkage disequilibrium. They cannot be treated as fully independent observations. Doing so would make the standard error too small.\nAdmixPy uses a delete-one-block jackknife. It leaves out one SNP block at a time and recalculates the statistic. For equal blocks, the standard error is:\nSE⁡2=B−1B∑b=1B(f^3,−b−f‾3,−)2 \\operatorname{SE}^2= \\frac{B-1}{B} \\sum_{b=1}^{B} (\\widehat f_{3,-b}-\\overline{f}_{3,-})^2 SE2=BB−1​b=1∑B​(f​3,−b​−f​3,−​)2This formula describes the simpler resampling=\u0026quot;nominal_blocks\u0026quot; setting, where blocks are weighted by their nominal sizes. AdmixPy’s default resampling=\u0026quot;pairwise_counts\u0026quot; uses a generalized count-weighted jackknife based on the number of usable SNPs in each block.\nAdmixPy reads genetic distances and physical positions from the .snp or .bim files. By default, blgsize defines blocks using a genetic-map distance of 0.05 cM. If no usable genetic map is available, AdmixPy falls back to 2-Mb physical blocks. Setting blgsize to 100 or more uses physical positions directly, with the value interpreted in base pairs. If neither genetic nor physical positions are available, each chromosome is treated as one block.\nSummary f3-statistics compress three pairwise population relationships into one value.\nWith the target in the first position, a significantly negative value is evidence of admixture. It shows that the target lies between the two references in allele-frequency space.\nWith a deep outgroup in the first position, the statistic measures shared drift between the other two populations. Larger values indicate more shared drift relative to that outgroup. Outgroup-f3 will be covered in a separate post in more detail.\n","permalink":"https://popgenblog.com/posts/admixture-f3-statistics/","summary":"f3-statistics are used to test if populations are admixed or to measure shared genetic drift between two populations relative to an outgroup.\nThis post explains the theory behind admixture f3-statisti","tags":[],"title":"Testing for Admixture with f3-Statistics in AdmixPy"},{"content":"Yes. For two qpAdm models with the same target, the same right groups, and the same settings, the model with the higher p-value is the better statistical fit.\nqpAdm calculates a covariance-weighted discrepancy between the observed and fitted f4-statistics. The p-value reflects how well the model explains the used f4-statistics. A higher p-value means the discrepancy between the observed and fitted values is less unusual under the model.\nThis does not mean that the model with the highest p-value for a target is automatically the best one, because qpAdm results depend on the selected right groups. Uninformative right-groups can lack the power to detect a bad model, while overly restrictive ones can make a plausible model appear to fit badly. Therefore, p-values are more comparable when models for the same target are compared using the same groups and settings. Models with different numbers of sources are also comparable since p-values account for different degrees of freedom. Z-scores can be used to assess whether an additional source is justified.\n","permalink":"https://popgenblog.com/posts/higher-p-values-better-qpadm/","summary":"Yes. For two qpAdm models with the same target, the same right groups, and the same settings, the model with the higher p-value is the better statistical fit.\nqpAdm calculates a covariance-weighted di","tags":[],"title":"Are Higher qpAdm P-Values Better?"},{"content":"A Vahaduo generated PCA model for Sardinians gives:\n82.8% Barcin Neolithic 11.6% Loschbour 5.6% Yamnaya Distance: 3.4303% Ganj Dareh was included in the sources but gets a weight of zero. This seems to imply that Sardinians don\u0026rsquo;t have any eastern-Farmer related ancestry.\nWhen Sardinians are modelled with qpAdm using Barcin Neolithic, Loschbour, Yamnaya, and Ganj Dareh, the model fits well:\n68.6% Barcin Neolithic 11.9% Loschbour 10.2% Yamnaya 9.4% Ganj Dareh Neolithic p = 0.769 When Ganj Dareh is dropped, the model fails (p=1.18×10−12p = 1.18 \\times 10^{-12}p=1.18×10−12).\nThis does not neccessarily mean that Sardinians derive exactly 9.4% of their ancestry directly from Ganj Dareh. But it shows that, with Barcin Neolithic as the core Anatolian source, the other three populations do not reproduce the target’s f4-profile.\nThe qpAdm result also passes when a more eastern-shifted Anatolian farmer source is used:\n75.1% Çatalhöyük EN 12.8% Loschbour 12.1% Yamnaya p = 0.685 There is no longer a need for a separate Ganj Dareh source. The genetic profile of Çatalhöyük is able to absorb the CIHG-related affinity that was previously represented by a mixture of Barcin and Ganj Dareh.\nThis example therefore distinguishes three separate questions:\nIs an ancestry dimension missing from the proposed model?\nWhich sampled population is the best available proxy for it?\nWhich historical population was the actual source?\nqpAdm can formally address the first question and can help with the second. It cannot guarantee the third.\nPCA distance minimisation does not even formally test the first question. Its zero coefficient means only that a source was unnecessary for obtaining the closest point in that particular PCA (used here: G25). It cannot definitely tell that the drift represented by that group is absent.\nHence, the closest PCA fit might look geometrically plausible while still being an incomplete admixture model.\nDisclaimer: qpAdm results are always conditional on the chosen right populations. A different right set can change the result. That said, with a reasonable selection of right groups, the models should be revealing, especially at this time depth. The right groups used here are:\nright = [ \u0026#34;Chimp\u0026#34;, \u0026#34;Turkey_Epipaleolithic\u0026#34;, \u0026#34;Georgia_KotiasKlde_Mesolithic\u0026#34;, \u0026#34;Russia_Vologda_Mesolithic\u0026#34;, \u0026#34;Switzerland_Epipaleolithic\u0026#34;, \u0026#34;Iran_BeltCave_Mesolithic\u0026#34;, ] With decent right groups, qpAdm can detect when a proposed source combination is missing an ancestry dimension. Removing Belt Cave and Klde Mesolithic in this example specifically would be tuning the right groups around the wanted result. But even after removing Belt Cave Mesolithic and Kotias Klde Mesolithic, the three-source model still fails decisively (p=1.82×10−7)(p = 1.82 \\times 10^{-7})(p=1.82×10−7).\n","permalink":"https://popgenblog.com/posts/best-pca-fit-wrong-admixture-model/","summary":"A Vahaduo generated PCA model for Sardinians gives:\n82.8% Barcin Neolithic 11.6% Loschbour 5.6% Yamnaya Distance: 3.4303% Ganj Dareh was included in the sources but gets a weight of zero. This seems t","tags":[],"title":"Why the Best PCA Fit May Still Be the Wrong Admixture Model"},{"content":"This post covers how to run pairwise f2-statistics and FST in AdmixPy. They are simple to interpret, and are also useful computationally. Once f2 blocks have been computed and cached, many downstream analyses can reuse them without repeatedly reading and converting the original genotype data.\nWhat does f2 measure? The f2-statistic quantifies allele-frequency differentiation between two populations, AAA and BBB, and is defined as:\nf2(A,B)=E[(pA−pB)2] f_2(A, B) = E[(p_A - p_B)^2] f2​(A,B)=E[(pA​−pB​)2]where pAp_ApA​ and pBp_BpB​ are the allele frequencies of populations AAA and BBB at a SNP, and the squared allele-frequency differences are averaged across SNPs.\nThis makes f2 a compact, distance-like summary of population differentiation. Because of sampling variation, bias-corrected estimates can sometimes be slightly negative.\nBias-corrected f2 A basic calculation of f2 would simply average the squared allele-frequency differences:\n(pA−pB)2 (p_A - p_B)^2 (pA​−pB​)2But allele frequencies are estimated from finite samples. If a population has only one or a few observed alleles at a SNP, the observed allele frequency can be different from the true population frequency because of sampling noise. Without correction, this sampling noise inflates f2.\nAdmixPy uses an allele-count-aware finite-sample correction:\nf^2(A,B)=E[(pA−pB)2−pA(1−pA)cA−1−pB(1−pB)cB−1], \\hat f_2(A,B)=E\\left[ (p_A-p_B)^2 -\\frac{p_A(1-p_A)}{c_A-1} -\\frac{p_B(1-p_B)}{c_B-1} \\right], f^​2​(A,B)=E[(pA​−pB​)2−cA​−1pA​(1−pA​)​−cB​−1pB​(1−pB​)​],Here, cAc_AcA​ and cBc_BcB​ are the observed allele counts for populations AAA and BBB. With the bias correction default, SNP values with cA≤1c_A\\leq1cA​≤1 or cB≤1c_B\\leq1cB​≤1 are excluded.\nWith bias correction enabled, AdmixPy requires at least two observed alleles in each population at a SNP. Values with (c\u0026lt;2) are excluded with a warning. Setting apply_corr=False retains them, but explicitly returns a sampling-biased estimate.\nDiploid and pseudohaploid samples Finite-sample correction is based on the number of observed alleles. Ploidy therefore directly affects estimates for populations represented by a single sample or only a few samples.\nA diploid individual contributes two alleles per SNP:\nc=2 c = 2 c=2Its observed allele frequency can be:\np∈{0,0.5,1} p \\in \\{0, 0.5, 1\\} p∈{0,0.5,1}A pseudohaploid individual contributes only one allele per SNP:\nc=1 c = 1 c=1Its observed allele frequency can only be:\np∈{0,1} p \\in \\{0, 1\\} p∈{0,1}For larger population groups, the effect is reduced because the allele-frequency estimate is averaged across more samples. For example, a pseudohaploid population represented by many individuals still has many observed alleles in total, even though each individual contributes only one allele per SNP. In that case, pseudohaploid calling is still taken into account.\nf2 and FST are population-level summaries. Bias-corrected f2 and FST can be calculated for a single diploid sample, although the estimates may be noisy. A population represented by one pseudohaploid sample has only one observed allele per called SNP, so bias-corrected estimates cannot be calculated for those SNPs.\nRunning pairwise f2 in AdmixPy For AdmixPy setup instructions, refer to the previous post: Introducing AdmixPy: f-statistics, qpAdm, and qpWave in Python. After activating the virtual environment, start a Python REPL by entering python in the terminal.\nThen import AdmixPy:\nimport admixpy as ap Since we imported admixpy as ap, pairwise f2-statistics can now be run with ap.f2().\nFor this example, I will compare a Late Neolithic population from Shandong against a set of modern East Asian populations:\n# Adjust to your dataset prefix or path prefix = \u0026#34;ho_v66\u0026#34; pops = [\u0026#34;Han\u0026#34;, \u0026#34;Japanese\u0026#34;, \u0026#34;Korean\u0026#34;, \u0026#34;Mongol\u0026#34;, \u0026#34;Tibetan\u0026#34;] Now run:\nap.f2(prefix, pop1=\u0026#34;China_Shandong_Dinggong_LN\u0026#34;, pop2=pops) The result:\npop1 pop2 est se n 0 China_Shandong_Dinggong_LN Han 0.00220414 0.000125263 206614 1 China_Shandong_Dinggong_LN Japanese 0.00413708 0.000136411 206614 2 China_Shandong_Dinggong_LN Korean 0.00460585 0.000190963 206614 3 China_Shandong_Dinggong_LN Mongol 0.00462775 0.000142786 206614 4 China_Shandong_Dinggong_LN Tibetan 0.00379421 0.000132498 206614 est is calculated by weighting each block estimate by the number of usable SNPs for that population pair. n is the total number of contributing SNPs, and se is the block-jackknife standard error, calculated by leaving out one block at a time.\nThe interpretation is simple: lower f2 means the two populations are more similar in allele-frequency space, while higher f2 means they are more differentiated. In this example, the Han average has the lowest f2 value with China_Shandong_Dinggong_LN, suggesting that among the listed comparison populations, Han are closest to this Shandong Late Neolithic group in allele-frequency space.\nThis should not be interpreted as an ancestry model. The result only says that, among this set of populations, Han show the lowest pairwise allele-frequency differentiation from China_Shandong_Dinggong_LN.\nBecause China_Shandong_Dinggong_LN contains multiple samples, it provides enough observed alleles for bias correction at more SNPs than a single pseudohaploid sample population would.\nEffect of the finite-sample correction For comparison, the same analysis can be run without the finite-sample correction (apply_corr=False):\npop1 pop2 est se n 0 China_Shandong_Dinggong_LN Han 0.0226191 0.000127416 206955 1 China_Shandong_Dinggong_LN Japanese 0.0251914 0.000140491 206955 2 China_Shandong_Dinggong_LN Korean 0.0334269 0.000201275 206955 3 China_Shandong_Dinggong_LN Mongol 0.0255762 0.000140268 206955 4 China_Shandong_Dinggong_LN Tibetan 0.0244494 0.000134095 206955 Without the correction, the estimates are approximately 5.5 to 10 times higher. Han still has the lowest f2, but Korean and Mongol switch ranks.\nBlocks and standard errors Like other ADMIXTOOLS-style statistics, AdmixPy estimates uncertainty using SNP blocks. SNPs are grouped by chromosome and genetic position. AdmixPy computes f2 separately for each block, then combines the blocks into the final estimate.\nThe standard error is calculated using a leave-one-block-out jackknife. In each jackknife replicate, one block is left out and the statistic is recomputed from the remaining blocks. The variation across these leave-one-block-out estimates gives the standard error.\nIf standard errors were computed as if every SNP were independent, they would usually be too small. Block jackknifing gives a more realistic estimate of uncertainty by accounting for correlation among nearby SNPs.\nUsing get_f2() and the f2 cache The main helper is get_f2(). It accepts genotype data, an in-memory F2Blocks object, or an on-disk f2 cache.\nFor example, you can compute f2 blocks from a TGENO prefix:\nblocks = ap.get_f2(prefix, pops=[\u0026#34;Chimp\u0026#34;, \u0026#34;Sardinian\u0026#34;, \u0026#34;Orcadian\u0026#34;, \u0026#34;Norwegian\u0026#34;, \u0026#34;Mongol\u0026#34;]) AdmixPy automatically detects pseudohaploid samples by default and adjusts their allele counts when computing allele frequencies. In most cases this does not need to be changed. It can be changed by adding the adjust_pseudohaploid parameter; False treats all samples as diploid, while an integer changes how many SNPs are checked during pseudohaploid detection.\nIf you want to reuse these blocks later, you can write them to disk with write_f2():\nap.write_f2(blocks, \u0026#34;f2_cache\u0026#34;) This creates a folder called f2_cache. Inside it, the block-level pairwise estimates, per-pair usable-SNP counts, and block lengths are saved. Then you can load the folder again with get_f2():\nblocks = ap.get_f2(\u0026#34;f2_cache\u0026#34;) The loaded blocks object can be passed to functions like ap.f2(), ap.f3(), or ap.f4().\nCached f2 blocks can be reused for f3, f4, qpWave, and qpAdm, but they cannot reproduce direct-genotype allsnps=True calculations when different statistics require different SNP intersections.\nThe default maxmiss=0 keeps only SNPs observed in every selected population, which can be restrictive for groups with high missingess. With finite-sample correction, SNPs with fewer than two allele observations are then excluded pairwise. Setting maxmiss=1 allows pairs to use their available SNP overlap, improving estimates for high-missingness populations. remove_na=False preserves blocks lacking data for some pairs instead of discarding them globally.\nWith many aDNA populations, setting remove_na=False - Trueis the default - is often needed for building a usable cache. When coverage varies across populations, blocks can lack estimates for some population pairs. Keeping those blocks nevertheless preserves the available pairwise data instead of discarding the entire block.\nRunning FST in AdmixPy AdmixPy also provides pairwise Hudson-style FST through ap.fst():\nap.fst(prefix, pop1=\u0026#34;China_Shandong_Dinggong_LN\u0026#34;, pop2=pops) Result:\npop1 pop2 est se n 0 China_Shandong_Dinggong_LN Han 0.00863094 0.000492654 273728 1 China_Shandong_Dinggong_LN Japanese 0.0162544 0.000530885 273728 2 China_Shandong_Dinggong_LN Korean 0.017967 0.000742235 273728 3 China_Shandong_Dinggong_LN Mongol 0.0176329 0.000540258 273728 4 China_Shandong_Dinggong_LN Tibetan 0.0147032 0.000511363 273728 For each block (b), AdmixPy calculates Hudson-style FST as:\nFST,b=∑s∈b[(pA−pB)2−pA(1−pA)cA−1−pB(1−pB)cB−1]∑s∈b[(pA−pB)2+pA(1−pA)+pB(1−pB)] F_{ST,b}= \\frac{ \\sum_{s\\in b}\\left[ (p_A-p_B)^2- \\frac{p_A(1-p_A)}{c_A-1}- \\frac{p_B(1-p_B)}{c_B-1} \\right] }{ \\sum_{s\\in b}\\left[ (p_A-p_B)^2+p_A(1-p_A)+p_B(1-p_B) \\right] } FST,b​=∑s∈b​[(pA​−pB​)2+pA​(1−pA​)+pB​(1−pB​)]∑s∈b​[(pA​−pB​)2−cA​−1pA​(1−pA​)​−cB​−1pB​(1−pB​)​]​With the bias corrected default, SNPs with cA≤1c_A\\leq1cA​≤1 or cB≤1c_B\\leq1cB​≤1 are excluded. Setting apply_corr=False removes the two correction terms.\nBy default, fst_aggregation=\u0026quot;block_ratios\u0026quot; averages these block-level estimates using their usable SNP counts. fst_aggregation=\u0026quot;pooled_components\u0026quot; instead pools the numerators and denominators across blocks before taking the ratio. The results can differ when genetic variation varies among blocks.\nFST is closely related to f2, but it is normalized by the amount of genetic variation in the compared populations. AdmixPy uses a Hudson-style estimator: the numerator is the same finite-sample-corrected allele-frequency difference used for f2, while the denominator is the expected pairwise allele difference between the two populations.\nBy default, f2 uses only sites polymorphic across populations, while FST keeps them. To compare f2 and FST on the same sites, pass poly_only=True to both calls or poly_only=False.\nf2 and FST often give a similar ordering of comparison populations. In this example, Han again has the lowest value relative to China_Shandong_Dinggong_LN among the listed populations. The difference is scale and interpretation: f2 is an unnormalized measure of allele-frequency differentiation, while FST asks how large that differentiation is relative to the total allele-frequency variation available at the SNPs being compared.\nSummary At the pairwise level, f2 and FST summarize how differentiated two populations are in allele-frequency space. Lower estimates generally mean more similar allele frequencies; higher estimates mean more differentiation.\nAdmixPy uses a finite-sample correction based on observed allele counts, which is relevant for single-sample populations and pseudohaploid ancient DNA. It also computes statistics by genomic blocks and uses a block jackknife for standard errors.\nA practical advantage is caching. By storing block-level f2 values in an F2Blocks object or an on-disk f2 cache, AdmixPy makes repeated f2, f3, f4, qpWave, and qpAdm analyses much faster. Instead of repeatedly reading and converting genotype data, you can compute the shared foundation once and reuse it across many tests.\nFor exploratory population-genetic analysis, this makes f2 both a useful statistic and a useful workflow primitive.\n","permalink":"https://popgenblog.com/posts/f2-statistics/","summary":"This post covers how to run pairwise f2-statistics and FST in AdmixPy. They are simple to interpret, and are also useful computationally. Once f2 blocks have been computed and cached, many downstream ","tags":[],"title":"Pairwise f2 Statistics and FST in AdmixPy"},{"content":"I recently published AdmixPy on GitHub, a fast implementation of f-statistics, qpAdm, and qpWave in Python that runs on Linux, macOS, and Windows. It works directly on the new AADR TGENO distribution format and is notably faster than ADMIXTOOLS 2 and simpler to set up. Supported input formats: EIGENSTRAT (.geno/.snp/.ind), packed AncestryMap (.geno/.snp/.ind), TGENO (.tgeno/.snp/.ind), and SNP-major PLINK binary (.bed/.bim/.fam).\nAdmixPy is implemented in Python and depends only on NumPy, SciPy, and pandas. Installation is handled through pip, and it should behave the same on every platform.\nSetup AdmixPy requires Python 3.10 or newer and runs on Linux, macOS, and Windows.\nIt is recommended to install AdmixPy in a virtual environment. Create one in your working directory:\npython3 -m venv venv source venv/bin/activate Then install or upgrade to the latest release from PyPI:\npython -m pip install --upgrade admixpy Verify the installation:\npython -c \u0026#34;import admixpy; print(admixpy.__version__)\u0026#34; Alternative: Installing from source Clone the repository and enter it:\ngit clone https://github.com/system0x7/admixpy.git cd admixpy Create and activate a virtual environment:\npython3 -m venv venv source venv/bin/activate On Windows, activate with venv\\Scripts\\activate instead.\nInstall the package:\npython -m pip install --upgrade pip python -m pip install -e . Verify the install:\npython -c \u0026#34;import admixpy; print(admixpy.__file__)\u0026#34; You should see a path ending in admixpy/__init__.py. If you get an ImportError, double-check that the virtual environment is activated.\nUsage The main functions are:\nadmixpy.f2(data, pop1, pop2) admixpy.fst(data, pop1, pop2) admixpy.f3(data, pop1, pop2, pop3) admixpy.f4(data, pop1, pop2, pop3, pop4) admixpy.qpwave(data, left, right) admixpy.qpadm(data, target, left, right) data can be a genotype dataset prefix or precomputed f2 data. For PLINK input, population labels are read from the FID column of the .fam file.\nStart a Python REPL (after activating the venv) in the directory containing your AADR files and run an f4 statistic:\n\u0026gt;\u0026gt;\u0026gt; import admixpy as a \u0026gt;\u0026gt;\u0026gt; prefix = \u0026#34;v66_compatibility\u0026#34; \u0026gt;\u0026gt;\u0026gt; a.f4(prefix, \u0026#34;Chimp\u0026#34;, \u0026#34;Turkey_N\u0026#34;, \u0026#34;Sardinian\u0026#34;, \u0026#34;French\u0026#34;) Result:\npop1 pop2 pop3 pop4 est se z p n 0 Chimp Turkey_N Sardinian French -0.00138048 9.23816e-05 -14.94 1.72e-50 682551 The significantly negative (Z=−14.94Z=-14.94Z=−14.94) estimate with AAA as outgroup indicates that Anatolian Neolithic farmers share more drift with Sardinians than with French, reflecting the stronger Neolithic Farmer affinity in Sardinia. Follow-up posts will work through f-statistics, qpAdm and qpWave models on AADR data using AdmixPy.\nRelated AdmixPy articles Pairwise f2 Statistics and FST in AdmixPy Testing for Admixture with f3-Statistics in AdmixPy Interpreting f4-Statistics with AdmixPy ","permalink":"https://popgenblog.com/posts/admixpy/","summary":"I recently published AdmixPy on GitHub, a fast implementation of f-statistics, qpAdm, and qpWave in Python that runs on Linux, macOS, and Windows. It works directly on the new AADR TGENO distribution ","tags":[],"title":"Introducing AdmixPy: f-statistics, qpAdm, and qpWave in Python"},{"content":"Commercial raw DNA exports are not provided in the file formats normally used by ADMIXTOOLS, ADMIXTOOLS 2, AADR-based workflows, or PLINK. Files from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, and Living DNA are usually plain-text vendor exports, while downstream workflows often require PLINK PACKEDPED or EIGENSTRAT/PACKEDANCESTRYMAP files.\nEIGENSTRAT is often used loosely to refer to the .geno/.snp/.ind triplet. Strictly speaking, EIGENSTRAT is the plain-text version of that triplet; PACKEDANCESTRYMAP is the packed binary form of the same three files. ADMIXTOOLS and ADMIXTOOLS 2 work with either, but PACKEDANCESTRYMAP takes far less disk space and loads much faster, which is why it\u0026rsquo;s the practical default used here.\nThis toolkit converts commercial raw DNA exports into both PACKEDPED and PACKEDANCESTRYMAP, and can merge the result with an AADR dataset.\nIt takes one or more raw DNA exports for the same individual and writes:\n.bed .bim .fam .geno .snp .ind It can also merge the converted sample directly into an existing AADR-style dataset, including the newer AADR releases that use TGENO-style genotype storage.\nThe purpose is to reduce the amount of manual format handling normally required for this workflow. There is no need to compile convertf, no need to compile mergeit, no PLINK dependency for the raw conversion step, and no manual concatenation or conversion of several vendor files to a common 23andMe-style format first. The wrapper detects the input layouts, creates a merged master raw file for the same individual, writes the converted outputs, and keeps the result in a clear folder structure.\nThe toolkit is available for 12$.\nDownload: Raw DNA to AADR Toolkit - Convert 23andMe, Ancestry \u0026amp; More to PackedAncestryMap And Packedped\nThe free workflow If you want to analyze your own DNA with ADMIXTOOLS (qpAdm, qpWave, f-statistics) or run PCA and ADMIXTURE against a reference, you first need to convert your raw file and merge it into an AADR-style reference dataset.\nThe manual workflow can involve a lot of brittle steps:\nconvert the commercial raw file to 23andMe-style format use PLINK to convert it to PACKEDPED compile mergeit/convertf convert the latest AADR v66 dataset with convertf to PACKEDANCESTRYMAP by setting up a parameter file and waiting at least an hour convert your own PLINK-derived PACKEDPED dataset to PACKEDANCESTRYMAP set up another parameter file to merge both PACKEDANCESTRYMAP datasets hope that after several hours you have a working merged dataset In the best case, this works after a lot of manual setup and waiting.\nIn the worst case, all you get is a cryptic error message. You might then try converting the AADR dataset to PACKEDPED and merging with PLINK’s --bmerge, which can lead to another round of allele, SNP, and strand issues, followed by another conversion back to the PACKEDANCESTRYMAP format needed for ADMIXTOOLS.\nAdvantages of this Bundle A major advantage is that the bundle can automatically create a single master raw file from several commercial DNA exports for the same individual.\nFor example, if you have one AncestryDNA file, one FamilyTreeDNA file, and one MyHeritage file for the same person, you can pass all of them to the wrapper. The tool reads the different vendor formats, normalizes them, merges them by rsID, and writes one merged raw file before producing PACKEDPED and PACKEDANCESTRYMAP output.\nInstead of preparing several intermediate files by hand and moving between different tools, the wrapper does the format detection, raw-file merging, conversion, and optional AADR merge in one reproducible workflow.\nThe main advantages are:\naccepts raw DNA exports from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, and Living DNA directly, no manual conversion to 23andMe format first can merge multiple vendor exports for the same individual into a single master raw file (by rsID), handling concatenation, sorting, deduplication, and vendor-format cleanup automatically preserves the merged master file in 01_merged_raw/ for reuse writes both PLINK PACKEDPED (.bed/.bim/.fam) and PACKEDANCESTRYMAP (.geno/.snp/.ind) output no PLINK dependency, no need to compile ADMIXTOOLS, no manual convertf or mergeit parameter files merges the new sample directly into an AADR-style PACKEDANCESTRYMAP dataset, intersecting SNPs and checking allele compatibility automatically clear output folder structure and logs for easy inspection runs on Windows, Linux, and macOS with minimal dependencies (Python + NumPy) end-to-end in under 10 minutes Usage The workflow has two basic steps.\nFirst, create a merged master raw file and convert it to PACKEDPED and PACKEDANCESTRYMAP. You can pass a single --input for one vendor file, or multiple --input flags to merge several exports for the same individual:\npython tools/bundle_raw_convert.py \\ --input my_ancestry_raw.txt \\ --input my_ftdna_raw.csv \\ --input my_myheritage_raw.txt \\ --iid sample_001 \\ --gender M \\ --ind-label Sample \\ --out-dir output This reads the vendor files (one or more), merges them by rsID when multiple are given, and writes:\noutput/01_merged_raw/sample_001.merged.txt output/02_packedped/sample_001.bed output/02_packedped/sample_001.bim output/02_packedped/sample_001.fam output/03_packedancestrymap/sample_001.geno output/03_packedancestrymap/sample_001.snp output/03_packedancestrymap/sample_001.ind Second, merge the new sample into an existing AADR-style dataset:\npython tools/mergeit_fast.py \\ /path/to/aadr/v66 \\ output/03_packedancestrymap/sample_001 \\ output/04_aadr_merged/sample_001.aadr mergeit_fast.py automatically detects the genotype layout of the input datasets.\nIt then detects whether the genotype file is:\npacked SNP-major .geno with a GENO header packed sample-major .tgeno with a TGENO header plain-text SNP-major genotype data plain-text sample-major genotype data The merged output is always written as packed ancestry map, which can be used directly with ADMIXTOOLS 2 in R:\nmerged_prefix.geno merged_prefix.snp merged_prefix.ind So you do not need to manually convert every input to the same internal genotype layout before merging. The merger reads the supported .geno or .tgeno input layout, intersects SNPs, checks chromosome, position, and allele compatibility, applies allele flips where needed, and writes a packed ancestry map result.\nAfter downloading and unzipping the bundle, see the README.md file in the root directory. It explains the full workflow in more detail, including setting up Python, installing the requirements, downloading the AADR v66 files, running the conversion script, and merging your sample into the reference dataset.\n","permalink":"https://popgenblog.com/posts/raw-dna-to-admixtools/","summary":"Commercial raw DNA exports are not provided in the file formats normally used by ADMIXTOOLS, ADMIXTOOLS 2, AADR-based workflows, or PLINK. Files from 23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, a","tags":[],"title":"Convert Raw DNA Files to EIGENSTRAT for ADMIXTOOLS and Merge with AADR"},{"content":"The origins of the Proto-Anatolians are often treated as one of the more obscure problems, but the genetic data may be not that ambigous. Anatolian is regarded as the earliest-splitting branch of \u0026ldquo;Indo-European\u0026rdquo;, and its divergence is deep enough that some linguists distinguish a pre–Proto-Indo-European stage, sometimes called \u0026ldquo;Indo-Anatolian\u0026rdquo;, from the Proto-Indo-European reconstructed from the non-Anatolian branches. Under either framing, the relevant question is the same: whether the earlier Eneolithic steppe-related ancestry behind Yamnaya, particularly the Caucasus–Lower Volga (CLV) component, also moved south of the Caucasus into Anatolia. For this purpose, I use Progress-2 specifically as proxy for the north Caucasus-facing part of this Eneolithic steppe-related ancestry, since it sits directly at the northern end of the Caucasus and therefore serves as a good proxy for groups that may have passed through the region.\nLanguages are obviously not genetics, and ancient DNA does not identify speech communities by itself. But this caveat should not become a license to ignore genetic evidence whenever it points in a direction one does not prefer. A repeated, statistically supported ancestry pattern along a coherent geographic route is not proof of language, but it is evidence about movement and population formation. If competing scenarios are allowed to rest on linguistic reconstruction, archaeological interpretation, and historical inference, then genetic evidence should not be excluded. If anything, it is more grounded than any of these.\nIn this post, I will show possible evidence for an eastern route into Anatolia, through or around the Caucasus.\nThe Caucasus Route and Arrival from the East To test whether Eneolithic steppe-related ancestry appeared in the Southern Caucasus by the Chalcolithic, I will begin with f4f_4f4​-statistics in the following arrangement:\nf4(Outgroup,Progress-2;Neolithic baseline,Target) f_4(\\text{Outgroup}, \\text{Progress-2}; \\text{Neolithic baseline}, \\text{Target}) f4​(Outgroup,Progress-2;Neolithic baseline,Target)This test asks whether the target shares more alleles with Progress-2 than the Neolithic baseline does. A positive result would mean that the target is shifted toward Eneolithic north Caucasus steppe-related ancestry relative to that baseline.\nA significantly positive result in this arrangement would therefore suggest that the target cannot be explained as simply lying on the Southern Caucasus Neolithic cline, but instead carries additional affinity to Eneolithic steppe-related ancestry. This provides a first test of whether ancestry related to groups north of the Caucasus had already entered the Southern Caucasus by the Chalcolithic.\nWith Areni-1 Chalcolithic in the target position As a first example, I place Areni-1 Chalcolithic in the target position:\nTarget f4f_4f4​ SE Z P Areni-1 Chalcolithic 0.00219 0.000541 4.05 5.03e-5 This indicates that Areni-1 Chalcolithic shares significantly more alleles with Progress-2 than Mentesh Tepe Neolithic (6000-4000 BCE) does. So, it does not behave as a simple continuation of the Southern Caucasus Neolithic baseline, but instead shows excess affinity to Eneolithic steppe-related ancestry from north of the Caucasus.\nThis supports the first step of the eastern-route argument: before turning to Anatolia itself, we can already observe a detectable movement, or at least possible gene flow, linking Eneolithic steppe-related groups north of the Caucasus with populations on its southern side by the Chalcolithic.\nWith Arslantepe Late Chalcolithic in the target position As the next step, I place Arslantepe Late Chalcolithic (average of the samples: ART020, ART027, ART017) in the target position. Arslantepe is relevant here because it lies in eastern Anatolia, to the west of the Caucasus, and because a Late Chalcolithic individual (ART038) from the site carries Y-DNA haplogroup R1b-V1636. This is the same paternal lineage found among the men buried in the kurgans at Progress-2, the Eneolithic steppe-related population used here as the northern Caucasus-facing reference. Arslantepe is therefore an obvious test case for whether the ancestry affinity seen south of the Caucasus also becomes detectable in the Upper Euphrates region.\nFirst, using Çayönü as an eastern Anatolian Neolithic baseline:\nf4(Ju_hoan_North,Progress-2;C¸ayo¨nu¨ Neolithic,Arslantepe Late Chalcolithic) f_4(\\text{Ju\\_hoan\\_North}, \\text{Progress-2}; \\text{Çayönü Neolithic}, \\text{Arslantepe Late Chalcolithic}) f4​(Ju_hoan_North,Progress-2;C¸​ayo¨nu¨ Neolithic,Arslantepe Late Chalcolithic) Baseline Target f4f_4f4​ SE Z P Çayönü Neolithic Arslantepe Late Chalcolithic 0.00160 0.000390 4.11 3.94e-5 This result is clearly positive. Arslantepe Late Chalcolithic shares significantly more alleles with the Eneolithic steppe-related source than Çayönü Neolithic does, suggesting that it cannot be modeled as a simple continuation of the local eastern Anatolian Neolithic baseline.\nThese results should not be read as requiring a simple direct migration from Areni-1 to Arslantepe. Rather, Areni-1 shows that Progress-2-related ancestry was already present south of the Caucasus by the Chalcolithic. Arslantepe then shows that a related signal also appears farther west in eastern Anatolia. The point is therefore not that the same population moved step by step from Areni into Anatolia.\nQuantification of Eneolithic Steppe-related ancestry in Anatolia After demonstrating the gradual appearance of steppe-related allele frequencies along a Caucasus route, I will now try to formally quantify Eneolithic steppe-related ancestry with qpAdm in several northern Near Eastern groups relevant to this question.\nMy preference is to treat qpAdm as a convex ancestry-modeling problem. In this setup, the left populations are the proposed ancestry sources, while the right populations serve as external anchors that are differentially related to those sources. I\u0026rsquo;ll therefore prefer a well-constrained and stable right-population setup over heavy reliance on rotation, in which potential sources are gradually shifted into the right populations. Rotation is often defended as a stress test, and used carefully it can be exactly that. In practice, however, this can slide into an artificial optimisation, where increasingly close or related populations to other sources are moved to the right side until a preferred model becomes feasible. In my view, a carefully chosen right set should test the model rather than be tuned until the model passes.\nFor the following models, I use this right-population set:\nJu_hoan_North Iraq_PPNA Georgia_KotiasKlde_Mesolithic Russia_Vologda_Mesolithic Switzerland_Epipaleolithic Tajikistan_Mesolithic Turkey_Epipaleolithic Israel_Natufian Iran_BeltCave_Mesolithic In keeping with the convex framing, all of these right populations temporally precede the sources used on the left, but only immediately so rather than by a wide margin, which is what lets them function as informative anchors instead of distant outgroups.\nArslantepe Late Chalcolithic For Arslantepe Late Chalcolithic, I model the target as a two-way mixture of Çayönü PPN and Progress-2 Eneolithic Steppe.\nThe model is accepted with a good fit:\nTarget Source Weight SE Z Arslantepe Late Chalcolithic Çayönü PPN 0.866 0.0255 34.0 Arslantepe Late Chalcolithic Progress-2 Eneolithic Steppe 0.134 0.0255 5.25 Model f4f_4f4​ rank dof chisq P Çayönü PPN + Progress-2 Eneolithic Steppe 1 7 6.26 0.510 The model estimates Arslantepe Late Chalcolithic as approximately 86.6% Çayönü PPN-related and 13.4% Progress-2 Eneolithic Steppe-related, with the steppe-related component being clearly significant.\nThe popdrop results are also informative:\nDropped source dof chisq P Interpretation None 7 6.26 0.510 Full model accepted Progress-2 Eneolithic Steppe 8 33.6 4.81e-5 Steppe source required Çayönü PPN 8 672 7.87e-140 Local Anatolian source required When Progress-2 Eneolithic Steppe is removed, the model fails, which shows that the steppe-related source is not simply decorative but necessary for the fit. At the same time, the overwhelming failure after removing Çayönü PPN confirms that most of the ancestry remains local eastern Anatolian-related.\nThis result aligns with the earlierf4f_4f4​-statistic. Arslantepe Late Chalcolithic carries a mostly local eastern Anatolian ancestry profile, but with a significant Eneolithic steppe-related contribution. In quantitative terms, this contribution is modest, around 13%, but it is statistically required.\nArslantepe38 The same model can also be applied to ART038, the R1b-V1636 individual from the Arslantepe Royal Tomb:\nTarget Source Weight SE Z ART038 Çayönü PPN 0.885 0.0460 19.2 ART038 Progress-2 Eneolithic Steppe 0.115 0.0460 2.49 Model f4 rank dof chisq P Çayönü PPN + Progress-2 Eneolithic Steppe 1 7 4.48 0.723 Dropped source dof chisq P Interpretation None 7 4.48 0.723 Full model accepted Progress-2 Eneolithic Steppe 8 12.1 0.148 Çayönü-only model still accepted Çayönü PPN 8 363 1.24e-73 Local Anatolian source required The model estimates ART038 as about 88.5% Çayönü PPN-related and 11.5% Progress-2 Eneolithic Steppe-related. The steppe-related component approaches significance, with a Z-score of 2.49, and the addition of this source improves the fit strongly (p=0.723p=0.723p=0.723), lowering the chisq from 12.1 in the Çayönü-only model to 4.48 in the two-way model.\nAt the same time, the steppe source is not strictly required here, since the Çayönü-only model remains formally acceptable with p=0.148p=0.148p=0.148. This result should be regarded as tentative, especially because this is a single ancient individual rather than a population average, making the result more prone to ploidy-related artifacts, coverage issues, and individual-level variation. Still, given the improved fit, the direction of the estimate, and the paternal link to Progress-2 through R1b-V1636, including Eneolithic steppe-related ancestry in the model is reasonable, though not required in this individual case.\nTilbeşar Höyük (Gaziantep) Bronze Age, I14649 A further noteworthy sample is I14649 from Bronze Age Tilbeşar Höyük, represented here by the Gaziantep Bronze Age label. The site lies roughly 50 km west of Carchemish, the later Neo-Hittite capital. Historically, this sample predates written evidence for Anatolian speakers in the region, so it cannot be treated as linguistically identifiable in any direct sense.\nWhat makes this individual interesting is the apparent mobility of R1b-V1636-bearing groups around the time of, and shortly after, their appearance at Late Chalcolithic Arslantepe. By the Bronze Age, the same paternal lineage is found farther southwest at Tilbeşar Höyük, showing that this lineage was not confined to the Upper Euphrates zone. Rather than requiring a simple movement directly from Arslantepe to Tilbeşar, this may reflect a broader dispersal of related groups across central and southeastern Anatolia, including the northern Levantine frontier zone, a region that later also becomes relevant for Luwian-speaking groups.\nTarget Source Weight SE Z Turkey Southeast Gaziantep BA Turkey Central Tepecik Ciftlik Neolithic 0.832 0.0622 13.4 Turkey Southeast Gaziantep BA Progress-2 Eneolithic Steppe 0.168 0.0622 2.69 Model f4 rank dof chisq P Turkey Central Tepecik Ciftlik Neolithic + Russia_Eneolithic_Steppe 1 7 3.65 0.820 This gives an estimate of roughly 83.2% Tepecik-Çiftlik Neolithic-related and 16.8% Progress-2 Eneolithic Steppe-related ancestry. The model passes comfortably with p=0.820p=0.820p=0.820, and the steppe-related component is significant with Z=2.69Z=2.69Z=2.69.\nRemoving the steppe component reduces the model fit to p=0.218p=0.218p=0.218, so while it is not strictly required with these right populations, it still notably improves the fit.\nKalehöyük Old Hittite Period The same two-way model can also be applied to Old Hittite Period Kalehöyük:\nTarget Source Weight SE Z Kalehöyük Old Hittite Period Çayönü PPN 0.817 0.0315 26.0 Kalehöyük Old Hittite Period Progress-2 Eneolithic Steppe 0.183 0.0315 5.81 Model f4f_4f4​ rank dof chisq P Çayönü PPN + Progress-2 Eneolithic Steppe 1 7 3.08 0.878 Dropped source dof chisq P Interpretation None 7 3.08 0.878 Full model accepted Progress-2 Eneolithic Steppe 8 36.7 1.32e-5 Steppe source required Çayönü PPN 8 524 5.50e-108 Local Anatolian source required For the Old Hittite period average, the estimate rises to about 18.3% Progress-2 Eneolithic Steppe-related ancestry, and the fit is very strong with p=0.878p=0.878p=0.878. Again, the steppe-related source is required, since removing it causes the model to fail.\nThe Old Hittite period fit is especially noteworthy. A model using a rather eastern Anatolian Neolithic source such as Çayönü PPN might not be the first expectation for central Anatolia, yet it produces a good fit, given that the right-population set is fairly constrained.\nReplacing Çayönü PPN with a more central Anatolian Neolithic source, Tepecik-Çiftlik, does not improve the situation. In fact, the two-way model with Tepecik-Çiftlik and Eneolithic Steppe fails with p=0.00134p=0.00134p=0.00134, even though the estimated steppe-related ancestry remains around 12.2%. This makes the stronger Çayönü-based fit more notable.\nSummary Target Steppe-related estimate Z P-Value Arslantepe LC 13.4% 5.25 0.510 ART038 11.5% 2.49 0.723 Tilbeşar BA I14649 16.8% 2.69 0.820 Kalehöyük Old Hittite 18.3% 5.81 0.878 F4 PCA of Relevant Populations As a visual addition, the relevant Anatolian samples used above fall within the broader Anatolian Chalcolithic and Bronze Age cluster. They are also noticeably separated from the Kura-Araxes-related groups. A PCA is obviously only useful for showing broad affinities along the variance-maximizing components of the f4f_4f4​-statistics, not for proving a specific ancestry model. Still, the broad pattern is clear: the Anatolian samples used here fall within the Anatolian Chalcolithic and Bronze Age cluster and remain noticeably separated from the Kura-Araxes-related groups. Explaining the Progress-2-related ancestry therefore does not seem to require mass Kura-Araxes migration into Anatolia, as someone might object.\nThis separation is also supported by a direct f4f_4f4​-test. f4(Chimp,Turkey_Cayonu_PPN;Armenia_DzhoghazBerkaber_EBA_KuraAraxes,Arslantepe Late Chalcolithic)f_4(\\text{Chimp}, \\text{Turkey\\_Cayonu\\_PPN}; \\text{Armenia\\_DzhoghazBerkaber\\_EBA\\_KuraAraxes}, \\text{Arslantepe Late Chalcolithic})f4​(Chimp,Turkey_Cayonu_PPN;Armenia_DzhoghazBerkaber_EBA_KuraAraxes,Arslantepe Late Chalcolithic) is effectively zero and non-significant (f4=−0.000029, Z=−0.09, p=0.931f_4=-0.000029,\\ Z=-0.09,\\ p=0.931f4​=−0.000029, Z=−0.09, p=0.931).\nThus, Arslantepe does not show detectable excess affinity to Kura-Araxes groups relative to the local eastern Anatolian Çayönü baseline, while the same Arslantepe average does show excess affinity to Progress-2 in the earlier test, so the signal appears to be steppe-related rather than simply Kura-Araxes-related.\nConclusion These results point to a consistent pattern. Eneolithic steppe-related ancestry is first detectable south of the Caucasus and in eastern Anatolia, and later remains visible in several Chalcolithic and Bronze Age Anatolian contexts. The contribution is not large, and in some cases it is clearly diluted, but it is repeatedly detectable and often statistically required. Nor does the contribution need to be large.\nIf early Proto-Anatolian speakers first existed as a subculture within a mostly local Anatolian environment, rather than as a mass population already spreading across Anatolia, then even a modest ancestry component could be historically meaningful.\nThis does not mean that every individual carrying such ancestry was necessarily an Anatolian speaker, nor that a simple two-way qpAdm model is the final word on the ancestry of these populations. More intermediate sources and more regionally specific models may improve individual fits in some cases. A separate supporting clue comes from IBD evidence: Ovaören MA2213 has been reported to share a 15.2 cM segment with Vonyuchka-1 from the North Caucasus steppe, pointing to a genealogical connection across the same broad northern Caucasus and Anatolian interaction zone.\nThe eastern route is supported by a trail of genetic evidence moving from the Eneolithic steppe and northern Caucasus zone, through the Southern Caucasus, and into Anatolia.\nIn opposition, the Balkan route remains harder to reconcile with the genetic evidence. It would require the relevant ancestry to enter Anatolia from the west or northwest, yet the clearest indications discussed here appear first in the Southern Caucasus, eastern Anatolia, and later central and southeastern Anatolia.\n","permalink":"https://popgenblog.com/posts/genetic-origins-proto-anatolians/","summary":"The origins of the Proto-Anatolians are often treated as one of the more obscure problems, but the genetic data may be not that ambigous. Anatolian is regarded as the earliest-splitting branch of \u0026ldq","tags":[],"title":"The Genetic Origins of the Proto-Anatolians"},{"content":"Recently, in April 2026, new AADR versions were released on Harvard Dataverse. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when co-analyzing ancient DNA generated with different experimental setups. This is especially relevant when combining data produced with different capture reagents such as Agilent (AG), Twist (TW), and shotgun (SG), because these can introduce systematic differences that may affect downstream analyses. The compatibility panels were added to minimize that problem and make mixed-platform datasets more directly comparable.\nDownloading the New Compatibility 2M SNP Subset Below are the commands to download the latest AADR compatibility dataset. At the moment, this seems like the most sensible choice if you want to work with mixed-platform ancient DNA in ADMIXTOOLS, especially when combining data generated with different capture reagents or with shotgun data. If, instead, you need a broader set of modern samples, for example for PCA or ADMIXTURE, you should choose an _HO dataset, since these datasets include more modern individuals.\nwget -O aadr_v66p1_2m_compatibility.tgeno \u0026#34;https://dataverse.harvard.edu/api/access/datafile/13994522\u0026#34; wget -O aadr_v66p1_2m_compatibility.snp \u0026#34;https://dataverse.harvard.edu/api/access/datafile/13994517\u0026#34; wget -O aadr_v66p1_2m_compatibility.ind \u0026#34;https://dataverse.harvard.edu/api/access/datafile/13994523\u0026#34; # Optional: sample metadata / annotations wget -O aadr_v66p1_2m_compatibility.anno \u0026#34;https://dataverse.harvard.edu/api/access/datafile/13994521\u0026#34; Note: the newer AADR releases are now distributed in TGENO format. That works fine with the original ADMIXTOOLS implementation, but not directly with admixtools2 in R. If you want to use the R version, you first need to convert the dataset to PACKEDANCESTRYMAP format with convertf.\nOption 1: Using TGENO Directly with AdmixPy AdmixPy, my Python implementation of f-statistics, qpAdm, and qpWave, reads AADR TGENO files directly, so no conversion step is required. It also supports EIGENSTRAT, PACKEDANCESTRYMAP, and PLINK binary datasets, and runs on Linux, macOS, and Windows.\nSee Introducing AdmixPy for installation instructions and an f4 example.\nOption 2: Converting with PLINK 2 Since 1 July 2025, PLINK 2 supports TGENO input directly and can export EIGENSOFT data in PACKEDANCESTRYMAP format. This is a simpler and faster option than with convertf. You can download the appropritate PLINK 2 binary for your operating system on the PLINK 2 website. After making sure that the executable is available in your PATH, run:\n# PLINK 2 expects the genotype file to have a .geno extension # when --eigfile is used. The file itself remains TGENO format. ln -s aadr_v66p1_2m_compatibility.tgeno \\ aadr_v66p1_2m_compatibility.geno plink2 \\ --eigfile aadr_v66p1_2m_compatibility \\ --export eig \\ --out aadr_v66p1_2m_compatibility_packed This produces .geno, .snp, and .ind files in PACKEDANCESTRYMAP format with the specified output prefix.\nThis, however, replaces the population labels in the third column of the .ind file with phenotype information. Since the sample order is not changed during this conversion, you can simply replace the generated .ind file with the original AADR file to restore the population labels:\ncp aadr_v66p1_2m_compatibility.ind \\ aadr_v66p1_2m_compatibility_packed.ind Option 3: Converting TGENO to PACKEDANCESTRYMAP For use in admixtools2, you will need a recent enough build of convertf that supports TGENO input and can convert the current TGENO format to PACKEDANCESTRYMAP. Older binaries, including the EIGENSOFT version I compiled in a previous post alongside smartpca, will not work here.\nTo compile the current version on a Debian-based Linux system:\nsudo apt update sudo apt install -y \\ build-essential \\ gfortran \\ liblapack-dev \\ liblapacke-dev \\ libgsl-dev \\ libopenblas-dev git clone https://github.com/DReichLab/AdmixTools cd AdmixTools/src make clobber make LDLIBS=\u0026#34;-llapacke\u0026#34; install After compilation, the binaries will be inside:\nAdmixTools/bin You can either copy convertf somewhere in your $PATH:\nsudo cp ../bin/convertf /usr/local/bin/ or just move the binary into the dataset directory and run it locally from there. For rare use, that is enough.\nNext, create a parameter file called for example convert.par:\ngenotypename: aadr_v66p1_2m_compatibility.tgeno snpname: aadr_v66p1_2m_compatibility.snp indivname: aadr_v66p1_2m_compatibility.ind outputformat: PACKEDANCESTRYMAP genooutfilename: aadr_v66p1_2m_compatibility_packed.geno snpoutfilename: aadr_v66p1_2m_compatibility_packed.snp indoutfilename: aadr_v66p1_2m_compatibility_packed.ind Now we run:\nconvertf -p convert.par # If convertf is in the current directory instead: ./convertf -p convert.par This may take a while. After it finishes, you will have a PACKEDANCESTRYMAP version compatible with admixtools2.\nOne thing worth pointing out here: the TGENO files may appear to load in admixtools2 without throwing an obvious error, but the results are not correct.\nGeneral Thoughts on V66 My first impression of this SNP set is positive. The compatibility panels are a sensible and useful approach, especially for mixed-platform analyses, where they can help account for platform effects more effectively than relying on older SNP sets alone.\nWhat I like less is the current .ind file labeling. Population groupings seem more country-based than site-based in many cases, which feels like a step back compared to v62. The older naming was often easier to work with when you actually wanted archaeologically meaningful grouping rather than broader geographic bins.\nThat said, this is not a major problem. You can always relabel or regroup samples manually depending on the analysis with the annotation files shipped with the releases.\nIf you want to continue from dataset preparation to actual analysis, see: Running f4-Statistics with Admixtools in R and Running qpAdm in R: Testing and Interpreting Ancestry Models for ancestry modeling.\n","permalink":"https://popgenblog.com/posts/aadr-v66-download-and-conversion/","summary":"Recently, in April 2026, new AADR versions were released on Harvard Dataverse. Among the more important additions are the new compatibility datasets introduced for reducing platform-specific bias when","tags":["convertf","EIGENSTRAT"],"title":"Downloading and Converting AADR v66"},{"content":"I recently published dt, a modern data transformation tool designed to make the awk workflows commonly used on this blog more intuitive, expressive, and fast. Dt is written in Rust because it compiles to a single binary that runs anywhere, and it uses Polars for the actual data processing, giving you columnar operations that handle large files efficiently. The syntax uses explicit functions (filter(), select(), mutate()) chained together with pipes, making common transformations easier to read and modify. There\u0026rsquo;s also an interactive REPL that shows you the result after each operation, letting you build complex pipelines step-by-step, catch mistakes early, and undo errors with .undo.\nInstallation Getting started with dt is simple: run the command in your terminal based on your operating system:\nmacOS/Linux:\ncurl --proto \u0026#39;=https\u0026#39; --tlsv1.2 -LsSf https://github.com/system0x7/dt/releases/latest/download/data-transform-installer.sh | sh Windows:\npowershell -ExecutionPolicy ByPass -c \u0026#34;irm https://github.com/system0x7/dt/releases/latest/download/data-transform-installer.ps1 | iex\u0026#34; Via Cargo (any platform with Rust installed):\ncargo install data-transform The installers will add dt to your PATH automatically. After installation, verify it\u0026rsquo;s working:\ndt --version You can now try launching the interactive REPL by simply running dt in the terminal (.exit to leave).\nNote: When using the REPL, commands should be entered and run one at a time. If you are copying examples from this post, paste each line separately to ensure they execute correctly.\nHow dt Simplifies awk For a full reference visit the dt reference. To demonstrate how dt simplifies common workflows used in this blog, let\u0026rsquo;s revisit some of the used awk commands.\nExample 1: Filtering Samples by Population The awk version:\nawk -v OFS=\u0026#39;\\t\u0026#39; \u0026#39; FNR==1 {f++} f==1 {want[$1]=1; next} # populations to keep f==2 {if ($3 in want) keep[$1]=1; next} # reference.ind (IID in col1, pop in col3) f==3 {if ($2 in keep) print $1,$2} # reference.fam (FID col1, IID col2) \u0026#39; pops reference.ind reference.fam \u0026gt; pops.keep To do this in dt we load the three files separately first:\npops = read(\u0026#39;pops\u0026#39;, header=false) ind_file = read(\u0026#39;reference.ind\u0026#39;, header=false) fam_file = read(\u0026#39;reference.fam\u0026#39;, header=false) Here, pops is simply a one-column file: each row contains one unique population label to keep.\nNow we filter the ind file for the sample iids:\niids = ind_file | filter($3 in pops) | select($1) In words, this means: take the ind_file we previously loaded, filter its third column (note the dollar notation for accessing unnamed columns) based on the population labels in the loaded pops file, and select the first column (which contains the sample IIDs).\nIn the next step, we filter the reference fam file in a similar way:\nkeep_file = fam_file | filter($2 in iids) | select($1, $2) keep_file | write(\u0026#39;pops.keep\u0026#39;, header=false, delimiter=\u0026#39; \u0026#39;) This translates to: take the loaded fam_file, filter its second column (which contains the IIDs) based on the IIDs we want, and select (\u0026ldquo;keep\u0026rdquo;) only the first (FIDs) and second (IIDs) columns. The result is a PLINK-compatible keep file containing only samples belonging to population labels mentioned row-wise in the pops file.\nThe delimiter=' ' option writes the output as space-separated columns, which works with PLINK. A tab delimiter, delimiter='\\t', would work as well.\nThis last step can also be written more concisely by piping the selected table directly into the output:\nfam_file | filter($2 in iids) | select($1, $2) | write(\u0026#39;pops.keep\u0026#39;, header=false, delimiter=\u0026#39; \u0026#39;) Example 2: String Manipulation with Lookups Below is a rather obscure-looking awk command from this previous post: SmartPCA Tutorial: How to Run PCA on Genetic Data (EIGENSTRAT).\nawk -F\u0026#39;,\u0026#39; -v OFS=\u0026#39;,\u0026#39; \u0026#39; NR == FNR { line = $0 gsub(/^[[:space:]]+/, \u0026#34;\u0026#34;, line) if (line == \u0026#34;\u0026#34;) next split(line, a, /[[:space:]]+/) if (a[1] == \u0026#34;\u0026#34; || a[3] == \u0026#34;\u0026#34;) next label = a[3] sub(/\\.(AG|DG|HO|SG)$/, \u0026#34;\u0026#34;, label) pop[a[1]] = label next } { if (split($1, parts, \u0026#34;:\u0026#34;) \u0026gt;= 2 \u0026amp;\u0026amp; (parts[2] in pop)) $1 = pop[parts[2]] \u0026#34;:\u0026#34; parts[2] print } \u0026#39; data.ind smartpca.csv \u0026gt; smartpca_with_labels.csv In the dt REPL, we can do this step-by-step. First, we load the needed files:\npca = read(\u0026#39;smartpca.csv\u0026#39;, header=false) ind = read(\u0026#39;data.ind\u0026#39;, header=false) Next, we clean up the population labels in the ind file by removing the suffixes .AG, .DG, .HO, or .SG from the third column:\nlabels = ind | mutate($3 = replace($3, re(\u0026#39;\\.(AG|DG|HO|SG)$\u0026#39;), \u0026#39;\u0026#39;)) This creates a lookup table where the first column contains IDs and the third column contains cleaned population labels.\nNow we transform the PCA file. For each row, we extract the IID from the first column (the part after the colon, hence [1] instead of [0]), look up its population label in the transformed ind file where the third column contains population labels with their suffixes (.HO, .AG, etc.) removed, and reconstruct (\u0026ldquo;mutate\u0026rdquo;) the first column as label:IID:\nresult = pca | mutate($1 = lookup(labels, split($1, \u0026#39;:\u0026#39;)[1], on=$1, return=$3) + \u0026#39;:\u0026#39; + split($1, \u0026#39;:\u0026#39;)[1]) In plain language: split the first column by : to get the IID, look it up in our labels table to find the cleaned population label, then rebuild the column by concatenating the label with the IID using : as separator.\nThe split() function takes the column to split (here $1), the delimiter to split on (: in quotes), and which part to extract: [0] for the part before the delimiter or [1] for the part after. The lookup() function takes four arguments: the lookup table (labels), the key to search for (the extracted IID), the column in the lookup table to match against (on=$1), and the column to return (return=$3). The rest of the mutation concatenates strings together using the + operator, joining the looked-up label with : and the original IID.\nFinally, write the result:\nresult | write(\u0026#39;smartpca_with_labels.csv\u0026#39;, header=false) The syntax may look unfamiliar at first in the second example, but dt follows a consistent pattern. After a few times, the operations become intuitive. The pipeline structure makes your work self-documenting, and the REPL lets you verify each step interactively.\nIf you encounter bugs or have feature requests, you can report them on GitHub.\n","permalink":"https://popgenblog.com/posts/data-transform/","summary":"I recently published dt, a modern data transformation tool designed to make the awk workflows commonly used on this blog more intuitive, expressive, and fast. Dt is written in Rust because it compiles","tags":[],"title":"dt: A Modern awk Alternative for Common Data Workflows"},{"content":"In this post, I will build a transparent admixture-screening workflow from scratch in R using f4-statistics and constrained regression. The main advantage is automation: instead of hand-writing every candidate model, the script tests many 2-way, 3-way, and 4-way source combinations in one pass and ranks them by fit. ADMIXTOOLS 2 already includes batch tools such as qpadm_multi() and qpadm_rotate(), so the point is not that qpAdm cannot be automated. The point is that this custom workflow is compact, transparent, easy to modify, and useful for exploratory model search before you validate the strongest candidates more formally.\nqpAdm remains the standard formal framework for testing admixture models with left and right populations. It estimates admixture weights, computes model fit, optionally constrains weights to be non-negative, and uses block jackknife or bootstrap resampling to estimate uncertainty. The script below is a complementary f4-space screening tool, not a substitute for qpAdm or qpWave.\nThe idea If a target population can be approximated as a mixture of several source-related ancestries, then the target\u0026rsquo;s pattern of f4-statistics relative to a fixed outgroup and a fixed reference set should also be approximable as a weighted combination of the source f4-vectors. This relies on the same broad f-statistic framework that underlies ADMIXTOOLS methods more generally.\nIn this post I compute f4-statistics of the form:\nf4(Outgroup, Test; Ref1, RefX)\nfor the target and for each candidate source. Each population is therefore represented by a vector of allele-sharing relationships across the reference panel. If the target lies near the convex combination of some sources in this f4 space, that source set is a plausible ancestry model worth checking more carefully.\nOptimization objective and constraints The core fit is a constrained weighted least-squares problem:\nmin⁡w∑i(yi−(Xw)isei)2 \\min_w \\sum_i \\left(\\frac{y_i - (Xw)_i}{se_i}\\right)^2 wmin​i∑​(sei​yi​−(Xw)i​​)2subject to:\nwj≥0for all j w_j \\ge 0 \\quad \\text{for all } j wj​≥0for all jand\n∑jwj=1 \\sum_j w_j = 1 j∑​wj​=1Here, y is the target f4-vector, X is the matrix of source f4-vectors, and se contains the marginal standard errors. The objective therefore measures standardized residuals rather than ordinary Euclidean distance and is solved by quadratic programming under biologically meaningful constraints.\nUnlike qpAdm, the fit uses only marginal standard errors, not the full covariance matrix among f4-statistics. The weights are useful for screening, but the p-values should be treated as approximate fit scores rather than formal qpAdm-equivalent tests.\nThe role of the outgroup and references The outgroup anchors the statistic, while the reference populations define the coordinate system in which the target and sources are compared. This is why a model is never tested in isolation, but relative to the information provided by the chosen right-side references.\nHence, more references are not automatically better. A reference that is symmetrically related to all candidate sources adds little discriminatory information. Second, changing the right set can change whether a model looks acceptable, because the model is conditional on that set.\nWhat this script is good for This script is especially useful when you have a target, a stable right set, and a medium-sized panel of candidate sources, and you want to answer questions like:\nWhich 2-way, 3-way, or 4-way models are geometrically plausible? Which sources recur across many good-fitting models? Which sources look interchangeable? Which candidates are obviously poor proxies? Because the script builds one f4 basis and reuses it across all combinations, it can screen large model spaces quickly. That makes it a good first-pass tool before running competition tests, qpWave checks, or full qpAdm validation on the strongest candidates.\nWhat it is not This script does not test the rank structure of the full left-right f4 matrix the way qpWave does, and it does not use the covariance-aware qpAdm machinery. If a model looks good here, the right next step is still to check it with qpAdm.\nThat limitation is not the same as the limitations of PCA-distance models. PCA is mainly a visualization and dimensionality-reduction method, and its distances are measured in a reduced coordinate system whose geometry is optimized for variance, not ancestry modeling. This script instead fits mixtures directly in a space defined by allele-sharing asymmetries relative to explicit outgroups and reference populations. So it sits in an intermediate position: less rigorous than qpAdm because it lacks full covariance-aware model testing, but more interpretable and statistically anchored than PCA-based ancestry modeling.\nInstallation You will need admixtools, tidyverse, and quadprog:\ninstall.packages(\u0026#34;remotes\u0026#34;) install.packages(\u0026#34;tidyverse\u0026#34;) install.packages(\u0026#34;quadprog\u0026#34;) remotes::install_github(\u0026#34;uqrmaie1/admixtools\u0026#34;) library(admixtools) library(tidyverse) library(quadprog) Example setup Below I will model the Samara Yamnaya Early Bronze Age group.\n# Specify your target population target_name \u0026lt;- \u0026#34;Russia_Samara_EBA_Yamnaya.AG\u0026#34; # Specify the outgroup outgroup \u0026lt;- \u0026#34;Mbuti\u0026#34; # EIGENSTRAT prefix data \u0026lt;- \u0026#34;data\u0026#34; Candidate sources:\npotential_sources \u0026lt;- c( \u0026#34;Turkey_N\u0026#34;, \u0026#34;Turkey_Cayonu_PPN\u0026#34;, \u0026#34;Iran_HajjiFiruz_N\u0026#34;, \u0026#34;Georgia_KotiasKlde_Mesolithic\u0026#34;, \u0026#34;Russia_Samara_EN_Mesolithic\u0026#34;, \u0026#34;Tajikistan_Mesolithic\u0026#34;, \u0026#34;Hungary_EN_Starcevo-1\u0026#34;, \u0026#34;Azerbaijan_MenteshTepe_N_ShomutepeShulaveri\u0026#34;, \u0026#34;Russia_AfontovaGora_UP\u0026#34;, \u0026#34;Russia_YanaRiver_UP\u0026#34;, \u0026#34;Russia_Vologda_Mesolithic\u0026#34; ) Right populations (references):\nreferences \u0026lt;- c( \u0026#34;Turkey_Epipaleolithic\u0026#34;, \u0026#34;Russia_Malta_UP\u0026#34;, \u0026#34;Switzerland_Epipaleolithic\u0026#34;, \u0026#34;Georgia_Satsurblia_LateUP\u0026#34;, \u0026#34;Iran_BeltCave_Mesolithic\u0026#34;, \u0026#34;Iraq_PPNA\u0026#34; ) These references are intended to capture different axes of West Eurasian variation, including western hunter-gatherer-related ancestry, Caucasus-related ancestry, Ancient North Eurasian-related ancestry, early Iranian farmer-related ancestry, and Anatolian-related ancestry. The reference panel should contain populations that help distinguish among the candidate sources.\nIn retrospect, I would have included a broader set of reference populations: with only six references the script has five f4 entries to fit, so a 4-way model is left with just two degrees of freedom. When the fit is that loosely constrained, many source combinations will look good regardless of whether they are real, which feeds directly into the optimistic p-values discussed below. A wider reference panel would add f4 entries, tighten the fits, and make the ranking more meaningful.\nThe full script # --- 1. INPUT CHECKS --- stopifnot(length(references) \u0026gt;= 2) potential_sources \u0026lt;- unique(setdiff(potential_sources, target_name)) if (length(potential_sources) \u0026lt; 2) { stop(\u0026#34;You need at least two candidate sources.\u0026#34;) } # --- 2. INTERNAL HELPER FUNCTIONS --- solve_weights \u0026lt;- function(target_vec, source_mat, se_vec, tol = 1e-10) { row_ok \u0026lt;- is.finite(target_vec) \u0026amp; is.finite(se_vec) \u0026amp; se_vec \u0026gt; 0 \u0026amp; apply(source_mat, 1, function(z) all(is.finite(z))) y \u0026lt;- as.numeric(target_vec[row_ok]) X \u0026lt;- source_mat[row_ok, , drop = FALSE] se \u0026lt;- as.numeric(se_vec[row_ok]) if (length(y) == 0 || nrow(X) == 0) return(NULL) # Weighted least squares objective: # minimize sum(((y - X %*% w) / se)^2) y_w \u0026lt;- y / se X_w \u0026lt;- sweep(X, 1, se, \u0026#34;/\u0026#34;) # Skip rank-deficient combinations if (qr(X_w)$rank \u0026lt; ncol(X_w)) return(NULL) n \u0026lt;- ncol(X_w) Dmat \u0026lt;- crossprod(X_w) + diag(tol, n) dvec \u0026lt;- crossprod(X_w, y_w) # Constraints: sum(weights) = 1 and weights \u0026gt;= 0 Amat \u0026lt;- cbind(rep(1, n), diag(n)) bvec \u0026lt;- c(1, rep(0, n)) fit \u0026lt;- try(quadprog::solve.QP(Dmat, dvec, Amat, bvec, meq = 1), silent = TRUE) if (inherits(fit, \u0026#34;try-error\u0026#34;)) return(NULL) w \u0026lt;- fit$solution w[abs(w) \u0026lt; 1e-8] \u0026lt;- 0 if (any(w \u0026lt; -1e-6)) return(NULL) w \u0026lt;- pmax(w, 0) w \u0026lt;- w / sum(w) list(weights = w, keep = row_ok) } evaluate_model \u0026lt;- function(y, X, se, w) { y_pred \u0026lt;- as.numeric(X %*% w) residuals \u0026lt;- y - y_pred z_resid \u0026lt;- residuals / se chi_sq \u0026lt;- sum(z_resid^2) df_val \u0026lt;- length(y) - (length(w) - 1) p_val \u0026lt;- if (df_val \u0026gt; 0) pchisq(chi_sq, df_val, lower.tail = FALSE) else NA_real_ rss \u0026lt;- sum(residuals^2) max_abs_z \u0026lt;- max(abs(z_resid)) list( chi_sq = chi_sq, df = df_val, p_value = p_val, rss = rss, max_abs_z = max_abs_z ) } # --- 3. DATA PREPARATION --- cat(\u0026#34;Calculating f4 basis...\\n\u0026#34;) all_pops \u0026lt;- c(target_name, potential_sources) f4_results \u0026lt;- f4( data, pop1 = outgroup, pop2 = all_pops, pop3 = references[1], pop4 = references[-1], allsnps = FALSE ) # allsnps = FALSE keeps SNP selection consistent across these f4 entries. # This can reduce SNP counts, but it avoids mixing different SNP subsets # across the basis vectors. f4_matrix_wide \u0026lt;- f4_results %\u0026gt;% dplyr::select(pop2, pop4, est, se) %\u0026gt;% tidyr::pivot_wider(names_from = pop2, values_from = c(est, se)) y \u0026lt;- as.numeric(f4_matrix_wide[[paste0(\u0026#34;est_\u0026#34;, target_name)]]) target_se \u0026lt;- as.numeric(f4_matrix_wide[[paste0(\u0026#34;se_\u0026#34;, target_name)]]) x_full \u0026lt;- f4_matrix_wide %\u0026gt;% dplyr::select(starts_with(\u0026#34;est_\u0026#34;)) %\u0026gt;% dplyr::select(-all_of(paste0(\u0026#34;est_\u0026#34;, target_name))) colnames(x_full) \u0026lt;- gsub(\u0026#34;^est_\u0026#34;, \u0026#34;\u0026#34;, colnames(x_full)) x_full \u0026lt;- as.matrix(x_full) if (ncol(x_full) \u0026lt; 2) { stop(\u0026#34;Not enough candidate sources after preprocessing.\u0026#34;) } # --- 4. COMBINATORIAL SEARCH --- results_list \u0026lt;- list() source_sizes \u0026lt;- 2:min(4, ncol(x_full)) for (m in source_sizes) { cat(\u0026#34;Testing all\u0026#34;, m, \u0026#34;source combinations...\\n\u0026#34;) combos \u0026lt;- combn(colnames(x_full), m) for (i in seq_len(ncol(combos))) { curr_names \u0026lt;- combos[, i] x_sub \u0026lt;- x_full[, curr_names, drop = FALSE] fit \u0026lt;- solve_weights(y, x_sub, target_se) if (!is.null(fit)) { w \u0026lt;- fit$weights keep \u0026lt;- fit$keep y_use \u0026lt;- y[keep] se_use \u0026lt;- target_se[keep] x_use \u0026lt;- x_sub[keep, , drop = FALSE] stats \u0026lt;- evaluate_model(y_use, x_use, se_use, w) results_list[[length(results_list) + 1]] \u0026lt;- tibble::tibble( n_sources = length(w), model = paste(curr_names, collapse = \u0026#34; + \u0026#34;), chi_sq = stats$chi_sq, df = stats$df, rss = stats$rss, max_abs_z = stats$max_abs_z, p_value = stats$p_value, weights = paste( paste0(curr_names, \u0026#34;=\u0026#34;, round(w * 100, 1), \u0026#34;%\u0026#34;), collapse = \u0026#34;\\n \u0026#34; ) ) } } } if (length(results_list) == 0) { stop(\u0026#34;No valid models were found.\u0026#34;) } # --- 5. OUTPUT --- final_table \u0026lt;- dplyr::bind_rows(results_list) %\u0026gt;% dplyr::mutate( status = dplyr::case_when( is.na(p_value) ~ \u0026#34;NA\u0026#34;, p_value \u0026gt; 0.05 ~ \u0026#34;PASS\u0026#34;, TRUE ~ \u0026#34;FAIL\u0026#34; ) ) %\u0026gt;% dplyr::arrange(dplyr::desc(p_value), chi_sq) top20 \u0026lt;- final_table %\u0026gt;% dplyr::slice_head(n = 20) cat(\u0026#34;\\n--- TOP 20 MODELS (RANKED BY APPROXIMATE P-VALUE) ---\\n\\n\u0026#34;) for (i in seq_len(nrow(top20))) { cat(sprintf(\u0026#34;[%d]\\n\u0026#34;, i)) cat(\u0026#34;model: \u0026#34;, top20$model[i], \u0026#34;\\n\u0026#34;, sep = \u0026#34;\u0026#34;) cat(\u0026#34;weights:\\n\u0026#34;, top20$weights[i], \u0026#34;\\n\u0026#34;, sep = \u0026#34;\u0026#34;) cat(sprintf(\u0026#34;p_value: %.6g\\n\u0026#34;, top20$p_value[i])) cat(sprintf(\u0026#34;chi_sq: %.6g\\n\u0026#34;, top20$chi_sq[i])) cat(sprintf(\u0026#34;df: %s\\n\u0026#34;, top20$df[i])) cat(sprintf(\u0026#34;max_abs_z: %.6g\\n\u0026#34;, top20$max_abs_z[i])) cat(\u0026#34;status: \u0026#34;, top20$status[i], \u0026#34;\\n\\n\u0026#34;, sep = \u0026#34;\u0026#34;) } A note on allsnps = FALSE ADMIXTOOLS 2 documents that, when computing f4-statistics directly from genotype data, different f4 entries may otherwise be based on different SNP subsets if some data are missing. Setting allsnps = FALSE forces a common SNP set across the populations involved in the calculation, which is attractive here because the script is treating the f4 values as coordinates in one shared basis. The trade-off is that SNP counts can drop, sometimes a lot, when many populations are included.\nInterpreting the output The script returns a ranked table of source combinations and their fitted weights. The most useful columns are:\nmodel, the source combination weights, the fitted ancestry proportions under the simplex constraint p_value, an approximate fit score under the script\u0026rsquo;s diagonal weighting scheme max_abs_z, the largest standardised residual, showing the biggest unexplained deviation left by the model. Values near or above 3 are worth inspecting, because they indicate that at least one f4 entry is misfit by roughly three standard errors status, a rough pass/fail flag Because the script does not use the full covariance structure among f4-statistics, its p-values are usually more optimistic than formal qpAdm p-values. They are still useful for ranking and screening, but they should not be treated as the final arbiters of model validity.\nExample result In my Yamnaya example, the top model was:\n[1] model: Turkey_Cayonu_PPN + Georgia_KotiasKlde_Mesolithic + Russia_Vologda_Mesolithic weights: Turkey_Cayonu_PPN=19% Georgia_KotiasKlde_Mesolithic=25.4% Russia_Vologda_Mesolithic=55.7% p_value: 0.980515 chi_sq: 0.181524 df: 3 max_abs_z: 0.34791 status: PASS with fitted weights of roughly:\n55.7% Russia_Vologda_Mesolithic 25.4% Georgia_KotiasKlde_Mesolithic 19% Turkey_Cayonu_PPN When I ran the same model in qpAdm, I got:\ntarget left weight se z \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Russia_Samara_EBA_Yamnaya Turkey_Cayonu_PPN 0.180 0.0733 2.46 2 Russia_Samara_EBA_Yamnaya Georgia_KotiasKlde_Mesolithic 0.287 0.0651 4.40 3 Russia_Samara_EBA_Yamnaya Russia_Vologda_Mesolithic 0.533 0.0361 14.8 $rankdrop # A tibble: 3 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested \u0026lt;int\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 2 4 11.2 2.45e- 2 6 61.2 2.55e-11 2 1 10 72.4 1.51e-11 8 359. 1.09e-72 3 0 18 431. 2.53e-80 NA NA NA The fitted proportions are fairly close between the two approaches, which suggests that the screening script is capturing a real ancestry direction in f4 space. However, this qpAdm model does not formally pass: the rankdrop p-value for the 3-way model is 0.0245, below the conventional 0.05 threshold. So this source combination should be treated as a useful screening hit rather than as a validated admixture model.\nOne possible reason is that qpAdm can become overly conservative when the right set is larger than necessary or contains redundant reference populations. Harney et al. specifically caution that too many right populations can bias qpAdm p-values downward and lead to rejection of otherwise plausible models. In practice, this means the model may be worth re-evaluating with a thinner, more informative right set, chosen on methodological grounds rather than tuned ad hoc to force a pass.\nIn a later sensitivity check, thinning the right set (outgroup + references) by removing Iraq_PPNA and Iran_BeltCave_Mesolithic made the same 3-source qpAdm model marginally pass (p = 0.0606) under allsnps = FALSE, reinforcing the point that model fit is conditional on right-population choice. Removing Mbuti increased the fit further (p = 0.279).\n","permalink":"https://popgenblog.com/posts/deterministic-f4-solver/","summary":"In this post, I will build a transparent admixture-screening workflow from scratch in R using f4-statistics and constrained regression. The main advantage is automation: instead of hand-writing every ","tags":["ADMIXTOOLS"],"title":"Fast, Transparent f4-Based Admixture Screening in R"},{"content":"mergeit is part of the EIGENSOFT package and can be used to merge exactly two EIGENSTRAT/PACKEDANCESTRYMAP datasets.\nSetting up EIGENSOFT mergeit is part of the EIGENSOFT package. You can install it via conda:\nconda install -c bioconda eigensoft Alternatively, if you prefer to compile from source, see: From EIGENSTRAT to PACKEDPED.\nSetting Up A Parameter File Like other EIGENSOFT tools, mergeit requires a parameter file:\ngeno1: aadr.geno snp1: aadr.snp ind1: aadr.ind geno2: eigenstrat_output.geno snp2: eigenstrat_output.snp ind2: eigenstrat_output.ind genooutfilename: merged.geno snpoutfilename: merged.snp indoutfilename: merged.ind Save this as mergeit.par and run:\nmergeit -p mergeit.par This produces merged.geno, merged.snp, and merged.ind, ready for downstream analysis with ADMIXTOOLS.\n","permalink":"https://popgenblog.com/posts/mergeit-tutorial/","summary":"mergeit is part of the EIGENSOFT package and can be used to merge exactly two EIGENSTRAT/PACKEDANCESTRYMAP datasets.\nSetting up EIGENSOFT mergeit is part of the EIGENSOFT package. You can install it v","tags":[],"title":"How to Merge EIGENSTRAT Datasets Using mergeit"},{"content":"In this post, I\u0026rsquo;ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the previous post, we already know it\u0026rsquo;s aligned to hs37d5. However, if you\u0026rsquo;re starting with a BAM file, you\u0026rsquo;ll need to verify the reference genome first. I\u0026rsquo;ll start by showing how to check BAM headers to identify the reference genome.\nIdentifying the Reference Genome from BAM Headers Before processing any BAM file, you should verify which reference genome it was aligned against. This is critical because AADR compatibility requires hs37d5 specifically. BAMs aligned to other GRCh37-based references like hg19 are also compatible (since they share the same coordinate system, differing only in chromosome naming conventions), but hg38/GRCh38 BAMs would require realignment from FASTQs.\nTo identify the reference genome, check the BAM header:\nsamtools view -H ERR14088885.bam | head -25 The output shows:\n@HD\tVN:1.6\tSO:coordinate @SQ\tSN:1\tLN:249250621 @SQ\tSN:2\tLN:243199373 @SQ\tSN:3\tLN:198022430 ... @SQ\tSN:X\tLN:155270560 @SQ\tSN:Y\tLN:59373566 The @SQ lines show reference sequences with their lengths. The chromosome naming (SN:1 instead of SN:chr1) and the specific length of chromosome 1 (249250621 bp) confirm this BAM is aligned to hs37d5 or GRCh37. If we saw SN:chr1 with the same length, that would indicate hg19. Different lengths (like 248956422 for chr1) would indicate hg38.\nReference Genome Quick-Check If you are unsure which reference was used, compare your @SQ lengths to this table:\nReference Name Chr1 Length (bp) Chromosome Naming hs37d5 / GRCh37 249,250,621 1, 2, X, MT hg19 249,250,621 chr1, chr2, chrX, chrM hg38 / GRCh38 248,956,422 chr1, chr2, chrX, chrM Note: AADR compatibility requires the hs37d5 / GRCh37 coordinate system. If your BAM uses hg19, it\u0026rsquo;s compatible (same coordinates, just rename chromosomes by removing the \u0026ldquo;chr\u0026rdquo; prefix if needed). BAMs aligned to hg38/GRCh38 require realignment from FASTQs.\nWith the reference genome verified, we can proceed to genotype calling.\nCreating a BED File from AADR SNP Positions First, I\u0026rsquo;ll create a BED file containing the AADR SNP positions (not to be confused with EIGENSTRAT\u0026rsquo;s .bed format, this is a genomic intervals file). It tells samtools mpileup to only process the positions we care about, rather than scanning the entire genome. This requires the .snp file from the AADR EIGENSTRAT dataset. If you haven\u0026rsquo;t downloaded it yet, see: How to Download the AADR Dataset (Linux \u0026amp; WSL).\nRun the command:\n# Assuming the .snp file is named v62.0_HO_public.snp awk \u0026#39;{print $2, $4-1, $4}\u0026#39; v62.0_HO_public.snp \u0026gt; v62_positions.bed This extracts the chromosome (column 2) and position (column 4), converting from 1-based to 0-based coordinates by subtracting 1 from the start position. The resulting BED file contains three columns: chromosome, start (0-based), and end (0-based, exclusive).\nInstalling pileupCaller pileupCaller is part of the sequenceTools package and can be installed via conda or compiled from source. For this tutorial, I\u0026rsquo;ll use the conda installation as it\u0026rsquo;s the most straightforward method.\nIf you don\u0026rsquo;t have conda installed, install Miniconda first:\n# Download Miniconda installer # For other architectures (ARM, macOS, etc.), see: https://docs.conda.io/en/latest/miniconda.html wget https://repo.anaconda.com/miniconda/Miniconda3-latest-Linux-x86_64.sh # Run installer bash Miniconda3-latest-Linux-x86_64.sh # Follow prompts, then restart your terminal or run: source ~/.bashrc Then install pileupCaller and bamutil (used in the next step):\nconda install -c bioconda sequencetools bamutil # Test installation pileupCaller --help Trimming Damage-Prone Read Ends Ancient DNA damage (C→T deamination) concentrates in the first and last ~2bp of fragments. We need to remove these positions before genotyping to avoid damage-induced errors.\nTrim 2bp from both ends of each read:\nbam trimBam ERR14088885_filtered.bam ERR14088885_trimmed.bam 2 # Index the trimmed BAM samtools index ERR14088885_trimmed.bam The output confirms successful trimming:\nNumber of records read = 2689408 Number of records written = 2689408 All reads are retained but with terminal bases masked, removing most damage without discarding data.\nPseudohaploid Genotype Calling with pileupCaller Now we run samtools mpileup and pipe the output directly to pileupCaller to generate EIGENSTRAT format files:\n# -q 10: Minimum mapping quality # -Q 20: Minimum base quality samtools mpileup -R -B -q 10 -Q 20 \\ -l v62_positions.bed \\ -f hs37d5/hs37d5.fa.gz \\ ERR14088885_trimmed.bam | \\ pileupCaller --sampleNames ERR14088885 \\ --samplePopName Bashur_Hoyuk_EBA \\ --randomHaploid \\ -f v62.0_HO_public.snp \\ -e eigenstrat_output Reference and Input Files hs37d5/hs37d5.fa.gz: Reference genome (must be indexed with samtools faidx) ERR14088885_trimmed.bam: Your filtered BAM file v62_positions.bed: BED file of AADR SNP positions v62.0_HO_public.snp: AADR EIGENSTRAT SNP file Sample Identifiers:\n--sampleNames ERR14088885: Sample ID (typically matching BAM filename) --samplePopName Bashur_Hoyuk_EBA: Population/individual label (becomes third column in .ind file) What it does samtools mpileup generates pileup format at AADR positions only, showing all reads covering each site pileupCaller --randomHaploid randomly samples one high-quality allele per position, producing pseudohaploid genotypes that avoid aDNA damage bias Quality Filtering Parameters Mapping Quality (-q 10):\nAncient DNA fragments are often short, resulting in lower mapping quality scores. A threshold of 10 retains authentic ancient reads that stricter filters (e.g., 30) would discard.\nBase Quality (-Q 20):\nRequires minimum Phred score of 20 (99% base call accuracy).\nRandom Haploid Sampling (--randomHaploid):\nRandomly samples one allele per site rather than calling heterozygotes, which is problematic in damaged/low-coverage aDNA where allelic dropout and damage can create false heterozygote calls.\nNote: For high-coverage modern data, use stricter thresholds (-q 30 -Q 30) and consider diploid calling instead of --randomHaploid.\nOutput PileupCaller outputs three EIGENSTRAT files: eigenstrat_output.geno, .snp, .ind\nAfter a successful run, you should see output in your terminal similar to this:\nERR14088885\t584131\t20265\t7.743468502784479e-2\t7.743468502784479e-2\t7.716077386750575e-2 Interpreting the results:\nTotalSites: Total AADR SNP positions attempted (584,131 for HO dataset) NonMissingCalls: Positions with sufficient coverage for genotype calls (20,265) avgRawReads: Mean coverage of raw pileup input across all sites (~0.077x or 7.7%) avgDamageCleanedReads: Mean coverage after single-stranded damage removal (~0.077x) avgSampledFrom: Mean coverage after removing tri-allelic reads (~0.077x) This sample has ~3.47% of AADR positions covered. While this may seem low, it\u0026rsquo;s sufficient for ADMIXTOOLS analyses (f-statistics, qpAdm) and rough PCA projection with smartpca.\nIn the next post, I\u0026rsquo;ll show how to merge this sample with the AADR dataset using mergeit.\nNow that you have your EIGENSTRAT files (.geno, .snp, .ind), you might want to delete intermediate files:\n# Clean up intermediate files rm v62_positions.bed ERR14088885_filtered.bam ERR14088885_filtered.bam.bai Keep your trimmed BAM and BAI files if you\u0026rsquo;re interested in uniparental marker analysis (Y-DNA haplogroup assignment, mitochondrial analysis) or damage assessment with mapDamage. I\u0026rsquo;ll possibly cover Y-DNA inference using ybyra in a future post, so keep your trimmed BAMs if you plan to follow along.\n","permalink":"https://popgenblog.com/posts/pileup-to-eigenstrat/","summary":"In this post, I\u0026rsquo;ll cover pseudohaploid genotype calling using pileupCaller and converting the output to EIGENSTRAT format for use with ADMIXTOOLS. Since we just created this BAM ourselves in the","tags":[],"title":"Pseudohaploid Genotyping for Ancient DNA: BAM to EIGENSTRAT"},{"content":"This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workflow is based on the run I used for ERR14088885, from the Başur Höyük study PRJEB83032.\nAncient DNA needs a different alignment strategy from ordinary modern whole-genome data. The molecules are short, the ends may carry post-mortem damage, and paired reads often overlap because the DNA insert is shorter than the sequencing cycles. For this sample I therefore clean poly-G tails, trim adapters, merge overlapping paired-end reads, align the merged molecules with bwa aln, remove low-confidence alignments, and deduplicate using both observed ends of each molecule.\nThe commands below run on any Debian or Ubuntu Linux machine.\nHardware and software The example run used 12 threads and 32 GB of RAM. A machine with fewer cores will also work, but alignment will take longer. The samtools sort -m value is memory per thread: 1500M with 12 threads permits sorting to use about 18 GB, leaving room for the operating system and the other tools.\nIf you\u0026rsquo;re running this on Windows Subsystem for Linux, work exclusively within the Linux filesystem (~/ or /home/username/), not in Windows directories such as /mnt/c/. Accessing the Windows filesystem from WSL incurs substantial I/O overhead and can make alignment extremely slow. Keep all downloads, reference genomes, and BAM files in your Linux home directory.\nInstall the required packages:\nsudo apt-get update sudo apt-get install -y \\ adapterremoval bwa curl default-jre-headless fastp samtools wget The recorded run used fastp 0.24.0, AdapterRemoval 2.3.4, bwa 0.7.18, samtools 1.21, and DeDup 0.12.9.\nSet up the working directories Run the pipeline from a project directory with enough free space for the reference, FASTQs, and intermediate files:\nset -Eeuo pipefail PROJECT_DIR=$PWD SAMPLE=ERR14088885 DATA_DIR=${PROJECT_DIR}/data/${SAMPLE} REFERENCE_DIR=${PROJECT_DIR}/hs37d5 OUTPUT_DIR=${PROJECT_DIR}/results/${SAMPLE}_paired_merged POLYG_DIR=${OUTPUT_DIR}/fastp_polyg TRIM_DIR=${OUTPUT_DIR}/adapterremoval TMP_DIR=${OUTPUT_DIR}/tmp TOOL_DIR=${PROJECT_DIR}/tools THREADS=12 SORT_MEMORY=1500M JAVA_MEMORY=20G mkdir -p \\ \u0026#34;$DATA_DIR\u0026#34; \u0026#34;$REFERENCE_DIR\u0026#34; \u0026#34;$OUTPUT_DIR\u0026#34; \\ \u0026#34;$POLYG_DIR\u0026#34; \u0026#34;$TRIM_DIR\u0026#34; \u0026#34;$TMP_DIR\u0026#34; \u0026#34;$TOOL_DIR\u0026#34; Adjust THREADS, SORT_MEMORY, and JAVA_MEMORY for your machine.\nDownload and index hs37d5 I use hs37d5 because the downstream target dataset is AADR, which uses GRCh37 coordinates. The decoy sequences in hs37d5 also give reads from repetitive or non-reference regions somewhere more appropriate to align than the primary chromosomes.\nDownload the compressed reference and verify it before indexing:\ncd \u0026#34;$REFERENCE_DIR\u0026#34; wget -c \\ https://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/phase2_reference_assembly_sequence/hs37d5.fa.gz gzip -t hs37d5.fa.gz echo \u0026#39;d10eebe06c0dbbcb04253e3294d63efc hs37d5.fa.gz\u0026#39; | md5sum -c - # Keep the downloaded .gz file and write an uncompressed FASTA for the tools. gzip -dc hs37d5.fa.gz \u0026gt; hs37d5.fa Build the BWA and samtools indexes. This is a one-time step for the reference:\nbwa index hs37d5.fa samtools faidx hs37d5.fa test -s hs37d5.fa.fai for suffix in amb ann bwt pac sa; do test -s \u0026#34;hs37d5.fa.${suffix}\u0026#34; || { echo \u0026#34;Missing BWA index: hs37d5.fa.${suffix}\u0026#34; \u0026gt;\u0026amp;2 exit 1 } done Set the reference path for the remaining commands:\nREFERENCE=${REFERENCE_DIR}/hs37d5.fa Download and check the paired FASTQs ENA provides ERR14088885 as two compressed FASTQ files containing the paired-end reads: R1 and R2.\ncd \u0026#34;$DATA_DIR\u0026#34; wget -c \\ ftp://ftp.sra.ebi.ac.uk/vol1/fastq/ERR140/085/ERR14088885/ERR14088885_1.fastq.gz \\ ftp://ftp.sra.ebi.ac.uk/vol1/fastq/ERR140/085/ERR14088885/ERR14088885_2.fastq.gz R1=${DATA_DIR}/${SAMPLE}_1.fastq.gz R2=${DATA_DIR}/${SAMPLE}_2.fastq.gz gzip -t \u0026#34;$R1\u0026#34; gzip -t \u0026#34;$R2\u0026#34; Remove poly-G tails The reads were produced on a NovaSeq instrument. NovaSeq uses two-color chemistry, where a lack of signal is read as G; once a short molecule has been sequenced through, this can produce artificial poly-G tails. A scan of the first 100,000 pairs found terminal runs of at least ten Gs in 69.7% of R1 and 66.3% of R2.\nI use fastp only for poly-G trimming. Adapter trimming and quality handling are left to AdapterRemoval in the next step:\nPOLYG_R1=${POLYG_DIR}/${SAMPLE}_1.polyg.fastq.gz POLYG_R2=${POLYG_DIR}/${SAMPLE}_2.polyg.fastq.gz fastp \\ --in1 \u0026#34;$R1\u0026#34; \\ --in2 \u0026#34;$R2\u0026#34; \\ --out1 \u0026#34;$POLYG_R1\u0026#34; \\ --out2 \u0026#34;$POLYG_R2\u0026#34; \\ --thread \u0026#34;$THREADS\u0026#34; \\ --trim_poly_g \\ --poly_g_min_len 10 \\ --disable_adapter_trimming \\ --disable_quality_filtering \\ --disable_length_filtering \\ --dont_eval_duplication \\ --json \u0026#34;${POLYG_DIR}/${SAMPLE}.fastp.json\u0026#34; \\ --html \u0026#34;${POLYG_DIR}/${SAMPLE}.fastp.html\u0026#34; test -s \u0026#34;$POLYG_R1\u0026#34; test -s \u0026#34;$POLYG_R2\u0026#34; The input contained 32,987,452 read pairs. After poly-G cleanup, 32,949,986 pairs were written for the next stage. Mean read lengths changed from 150 bases for both paired-end reads to 122 bases for R1 and 131 bases for R2.\nTrim adapters and merge overlapping paired-end reads Short aDNA inserts are often sequenced from both ends with extensive overlap. AdapterRemoval can trim the adapter sequence and combine the two observations into one consensus read. This prevents an overlapping pair from counting the same molecule twice.\nAR_PREFIX=${TRIM_DIR}/${SAMPLE} COLLAPSED_FASTQ=${AR_PREFIX}.collapsed.gz AdapterRemoval \\ --file1 \u0026#34;$POLYG_R1\u0026#34; \\ --file2 \u0026#34;$POLYG_R2\u0026#34; \\ --basename \u0026#34;$AR_PREFIX\u0026#34; \\ --threads \u0026#34;$THREADS\u0026#34; \\ --gzip \\ --trimns \\ --trimqualities \\ --minquality 20 \\ --minlength 30 \\ --collapse \\ --collapse-deterministic \\ --preserve5p test -s \u0026#34;$COLLAPSED_FASTQ\u0026#34; These choices have specific jobs:\n--minlength 30 discards molecules shorter than 30 bases after trimming. --collapse merges paired-end reads with at least 11 bases of overlap, the AdapterRemoval default recorded in this run. --collapse-deterministic writes N with quality 0 when overlapping bases disagree and have equal quality, instead of resolving the tie randomly. --preserve5p protects the physical molecule ends used for aDNA duplicate detection. With AdapterRemoval 2.3.4, it also means collapsed reads are not quality-trimmed at either end. The .settings report is important: read it before deciding whether a merged-only pipeline is suitable for another library.\nless \u0026#34;${AR_PREFIX}.settings\u0026#34; For ERR14088885, AdapterRemoval processed 32,949,986 pairs and produced 31,355,473 full-length collapsed reads, a collapse rate of about 95.2%. The retained reads averaged 67 bases. Because this workflow sends only ${SAMPLE}.collapsed.gz to BWA, non-overlapping pairs and singletons do not enter the final BAM. That is a deliberate, conservative choice for this library, not a general rule for every aDNA dataset. If the collapse rate is low, align the non-collapsed pairs and singletons as separate branches rather than silently discarding most of the data.\nAlign the collapsed molecules with BWA The original version of this tutorial used bwa mem -k 16 -L 16,16 directly on the raw paired FASTQs. I replaced that step with bwa aln, using the settings from the completed run:\nSAI=${OUTPUT_DIR}/${SAMPLE}.collapsed.sai bwa aln \\ -t \u0026#34;$THREADS\u0026#34; \\ -l 1024 \\ -n 0.01 \\ -o 2 \\ \u0026#34;$REFERENCE\u0026#34; \\ \u0026#34;$COLLAPSED_FASTQ\u0026#34; \\ \u0026gt; \u0026#34;$SAI\u0026#34; test -s \u0026#34;$SAI\u0026#34; The parameters are commonly used for short, damaged ancient DNA:\n-l 1024 makes the seed longer than the reads, effectively disabling seeding. Damage near a read end can then contribute to the full-read alignment instead of causing the seed to fail. -n 0.01 is BWA\u0026rsquo;s floating-point edit-distance setting. BWA chooses the maximum edit distance by read length, using a 1% missing-alignment threshold under its error model. -o 2 allows up to two gap openings instead of the default one. On 12 threads, bwa aln processed the 31,355,473 collapsed reads in 2 hours 43 minutes. Runtime will vary with CPU speed, storage, read length, and sample size.\nConvert the SAI output to SAM, add a read group, retain mapped reads with MAPQ at least 30, and coordinate-sort the result:\nFILTERED_BAM=${OUTPUT_DIR}/${SAMPLE}.mapq30.sorted.bam READ_GROUP=\u0026#34;@RG\\\\tID:${SAMPLE}\\\\tSM:${SAMPLE}\\\\tLB:${SAMPLE}\\\\tPL:ILLUMINA\u0026#34; bwa samse -r \u0026#34;$READ_GROUP\u0026#34; \\ \u0026#34;$REFERENCE\u0026#34; \u0026#34;$SAI\u0026#34; \u0026#34;$COLLAPSED_FASTQ\u0026#34; \\ | samtools view --bam --with-header -F 4 -q 30 - \\ | samtools sort \\ -@ \u0026#34;$THREADS\u0026#34; \\ -m \u0026#34;$SORT_MEMORY\u0026#34; \\ -T \u0026#34;${TMP_DIR}/${SAMPLE}.sort\u0026#34; \\ -o \u0026#34;$FILTERED_BAM\u0026#34; - test -s \u0026#34;$FILTERED_BAM\u0026#34; samtools quickcheck -v \u0026#34;$FILTERED_BAM\u0026#34; Here, -F 4 removes unmapped records and -q 30 removes alignments below MAPQ 30. Filtering before duplicate removal keeps DeDup focused on alignments that will be used downstream.\nThis sample produced 219,005 mapped reads at MAPQ 30 or higher. That is only about 0.7% of the collapsed input, so it should not be mistaken for a typical mapping rate. A low rate like this deserves follow-up QC: confirm the sample identity and reference, inspect read composition and adapter reports, and quantify endogenous human DNA before drawing biological conclusions.\nRemove PCR duplicates I use DeDup because it was designed for short ancient DNA and can compare both ends of a merged molecule. Download the pinned release once:\nDEDUP_VERSION=0.12.9 DEDUP_JAR=${TOOL_DIR}/DeDup-${DEDUP_VERSION}.jar if [[ ! -s \u0026#34;$DEDUP_JAR\u0026#34; ]]; then curl --fail --location --retry 3 \\ --output \u0026#34;${DEDUP_JAR}.partial\u0026#34; \\ \u0026#34;https://github.com/apeltzer/DeDup/releases/download/${DEDUP_VERSION}/DeDup-${DEDUP_VERSION}.jar\u0026#34; mv \u0026#34;${DEDUP_JAR}.partial\u0026#34; \u0026#34;$DEDUP_JAR\u0026#34; fi java -jar \u0026#34;$DEDUP_JAR\u0026#34; -h \u0026gt;/dev/null Run it in merged-read mode:\njava \u0026#34;-Xmx${JAVA_MEMORY}\u0026#34; -jar \u0026#34;$DEDUP_JAR\u0026#34; \\ -m \\ -i \u0026#34;$FILTERED_BAM\u0026#34; \\ -o \u0026#34;$OUTPUT_DIR\u0026#34; DEDUP_BAM=${OUTPUT_DIR}/${SAMPLE}.mapq30.sorted_rmdup.bam test -s \u0026#34;$DEDUP_BAM\u0026#34; The -m flag tells DeDup that every input record is a merged molecule, so read-name prefixes are not required. Do not apply it to an ordinary single-end library: the observed 3′ read end is not necessarily the original molecule end in that case.\nFor this run, DeDup examined 219,005 mapped reads, removed 100,188 duplicates, and retained 118,817 reads. The reported duplication rate was 0.46.\nWrite and verify the final BAM Sort DeDup\u0026rsquo;s output once more, index it, and write two standard samtools reports:\nFINAL_BAM=${OUTPUT_DIR}/${SAMPLE}.final.bam samtools sort \\ -@ \u0026#34;$THREADS\u0026#34; \\ -m \u0026#34;$SORT_MEMORY\u0026#34; \\ -T \u0026#34;${TMP_DIR}/${SAMPLE}.final-sort\u0026#34; \\ -o \u0026#34;$FINAL_BAM\u0026#34; \\ \u0026#34;$DEDUP_BAM\u0026#34; samtools index -@ \u0026#34;$THREADS\u0026#34; \u0026#34;$FINAL_BAM\u0026#34; samtools flagstat -@ \u0026#34;$THREADS\u0026#34; \u0026#34;$FINAL_BAM\u0026#34; \\ \u0026gt; \u0026#34;${OUTPUT_DIR}/${SAMPLE}.flagstat.txt\u0026#34; samtools stats -@ \u0026#34;$THREADS\u0026#34; \u0026#34;$FINAL_BAM\u0026#34; \\ \u0026gt; \u0026#34;${OUTPUT_DIR}/${SAMPLE}.stats.txt\u0026#34; samtools quickcheck -v \u0026#34;$FINAL_BAM\u0026#34; test -s \u0026#34;${FINAL_BAM}.bai\u0026#34; cat \u0026#34;${OUTPUT_DIR}/${SAMPLE}.flagstat.txt\u0026#34; The final files are:\nresults/ERR14088885_paired_merged/ERR14088885.final.bam results/ERR14088885_paired_merged/ERR14088885.final.bam.bai Result The final BAM contains 118,817 mapped reads after MAPQ filtering and duplicate removal, including 237 reads on chromosome Y. PMDtools 0.60 gave a mean PMD score of -0.745 across the BAM, with 1,027 reads (0.86%) reaching PMD \u0026gt;= 3; none of the Y reads reached that threshold. The single reads supporting M694/CTS5611 and CTS7400/PF6469 scored -0.797 and -1.403, respectively, under the default double-stranded model.\nFor practical purposes, this BAM is unusable for reliable autosomal or Y-DNA analysis. PMD, end, and substitution filtering leaves only 22 v66 \u0026ldquo;compatibility\u0026rdquo; SNPs, all at depth one and without independent read support. YBYRA found an R1b-M269-like best-scoring path, but its score remained below the reporting threshold, so no formal Y-DNA haplogroup could be assigned. The original study also excluded this sample from downstream genetic analysis.\n","permalink":"https://popgenblog.com/posts/ancient-dna-alignment-bwa-tutorial/","summary":"This is the first post in a series on processing an ancient DNA sample for use with ADMIXTOOLS. Here I go from paired-end FASTQ files to a filtered, duplicate-removed BAM aligned to hs37d5. The workfl","tags":[],"title":"Processing Ancient DNA: From FASTQ to Aligned BAM"},{"content":"This post covers using qpAdm in R to test ancestry models and estimate admixture proportions. qpAdm builds on f4-statistics and provides a framework for evaluating whether proposed source populations can explain a target population\u0026rsquo;s genetic makeup.\nFor R and admixtools setup instructions on Debian/Ubuntu, see my previous post: Running f4-Statistics with Admixtools in R. Windows users can find R installation instructions on the R website.\nWhat is qpAdm? qpAdm is a method for testing ancestry models and estimating admixture proportions. It determines whether a target population can be modeled as a mixture of specified source populations (\u0026ldquo;left populations\u0026rdquo;), and if the model fits, calculates the contribution from each source. The method builds on f4-statistics (covered in my previous post) to evaluate these ancestry models.\nUnlike many ancestry estimation tools, qpAdm provides a formal statistical framework for model evaluation through a fit test. qpAdm estimates the admixture proportions and calculates a p-value; if the p-value \u0026gt; 0.05, the model is considered a pass. Like f4-statistics and other admixtools methods, qpAdm handles low-coverage ancient DNA samples well, making it particularly suitable for ancient DNA research.\nHow to run qpAdm-Models in R? For this tutorial, I\u0026rsquo;ll use as usual the AADR dataset, a curated collection of ancient and modern genomic samples in EIGENSTRAT format. Since admixtools works natively with EIGENSTRAT files, the dataset can be used directly. For download instructions, see: Download Ancient \u0026amp; Modern DNA (AADR Tutorial). If you want to also run qpAdm on your own raw DNA file merged into AADR, see the see the Raw DNA to AADR toolkit, which handles the conversion and merge end-to-end.\nThen, navigate to the folder containing your EIGENSTRAT dataset and start R:\ncd /path/to/aadr/dataset R Within R we load the admixtools library (refer to the previous f4-statistics post linked above for set up instructions if you are unsure how to set it up on Debian/Ubuntu) and set the file prefix:\nlibrary(admixtools) # Assuming the EIGENSTRAT dataset is named data.geno/.snp/.ind data = \u0026#34;data\u0026#34; When running qpAdm, you specify three components: a target population, source populations (left), and reference populations (right). I typically define these as separate variables before passing them to the qpadm() function.\nTo demonstrate, let\u0026rsquo;s test a specific admixture model.\nExample: Modeling Northern Italians as a Mixture of Neolithic and Bronze Age Populations First, specify the target population using the exact label from the third column of your .ind file, for this example:\ntarget = \u0026#34;Italian_North.HO\u0026#34; The qpAdm function will automatically include all samples assigned to this population label.\nNext, we will set an initial set of left populations.\nFor a Neolithic-Bronze Age model we will start with three main populations: Anatolian Neolithic Farmers, Bronze Age Steppe pastoralists and Paleolithic Hunter Gatherers from Italy:\nleft = c(\u0026#34;Turkey_Marmara_Barcin_N.SG\u0026#34;, \u0026#34;Russia_Samara_EBA_Yamnaya.AG\u0026#34;, \u0026#34;Italy_Epigravettian.SG\u0026#34;) Next, select the right populations (reference populations). These are crucial for model evaluation, as they provide the statistical contrasts qpAdm uses to resolve ancestry components. There\u0026rsquo;s no definitive method for choosing right populations, though 6-12 is typically a good range. The key is selecting populations that span different branches of the population tree relevant to your model and are differentially related to your sources (see the f4-statistics post for details). Ideally, right populations should not have received gene flow from your left populations (and target), as this can obscure the statistical signal.\nWhen selecting right populations, consider the temporal depth of your analysis. For instance, when modeling Iron Age samples, including Bronze Age right populations can improve resolution compared to using only deep prehistoric reference populations. Similarly, if you\u0026rsquo;re trying to distinguish between closely related source populations, choose right populations that are differentially related to those sources. This reveals drift imbalances that help resolve subtle ancestry differences.\nFor this example, I\u0026rsquo;ll use a smaller set to demonstrate the iterative process of model refinement. Mbuti serves as a deep outgroup, commonly used in f4-statistics as well. Turkey_Central_Pinarbasi_Epipaleolithic provides an anchor for the Anatolian Neolithic component since it predates Barcin farmers. Russia_MA1_UP serves as an Ancient North Eurasian (ANE) anchor, which is relevant for resolving Yamnaya ancestry. Similarly, Georgia_Satsurblia_LateUP acts as a Caucasus Hunter-Gatherer (CHG) anchor, another component important for Yamnaya. I\u0026rsquo;ve also included Jordan_PPNB as southern Near Eastern Neolithic anchor, which will be useful later when we test the Levantine source contribution. These populations help differentiate the ancestry sources by providing distinct statistical contrasts:\nright = c(\u0026#34;Mbuti.DG\u0026#34;, \u0026#34;Turkey_Central_Pinarbasi_Epipaleolithic.AG\u0026#34;, \u0026#34;Russia_MA1_UP.SG\u0026#34;, \u0026#34;Jordan_PPNB.AG\u0026#34;, \u0026#34;Georgia_Satsurblia_LateUP.SG\u0026#34;) Now run the model and save it into a variable called (for example) result:\nresult = qpadm(data, left, right, target) After running the model, display the results by entering result:\n# A tibble: 3 × 5 target left weight se z \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Italian_North.HO Turkey_Marmara_Barcin_N.SG 0.587 0.0649 9.04 2 Italian_North.HO Russia_Samara_EBA_Yamnaya.AG 0.381 0.0941 4.04 3 Italian_North.HO Italy_Epigravettian.SG 0.0320 0.153 0.209 This shows the estimated ancestry proportions (weights), their standard errors (se), and z-scores indicating statistical significance. The z-scores test whether each ancestry component is significantly different from zero.\nThe results show 58.7% (± 6.49%) Neolithic Farmer ancestry, 38.1% (± 9.41%) Bronze Age Steppe ancestry, and 3.2% Epigravettian hunter-gatherer ancestry. However, the Epigravettian component has a z-score of only 0.209, indicating it\u0026rsquo;s not statistically distinguishable from zero. The extremely high standard error (± 15.3%) relative to the estimate confirms this component is poorly resolved in the current model setup. Even before checking the overall model fit (p-value), this signals a problem with how the model resolves hunter-gatherer ancestry.\nThe rankdrop section shows the model fit statistics:\n# A tibble: 3 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested \u0026lt;int\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 2 2 7.67 2.16e- 2 4 34.9 4.94e- 7 2 1 6 42.5 1.44e- 7 6 627. 3.16e-132 3 0 12 670. 1.32e-135 NA NA NA The first row shows our three-source model (f4rank = 2). The p-value of 0.0216 falls below the 0.05 threshold, meaning the model is rejected. The proposed sources cannot explain the target\u0026rsquo;s ancestry with the current right population set. This combination of a failed p-value and poorly resolved Epigravettian component suggests we need to refine our approach by adding more right populations that can better resolve hunter-gatherer ancestry:\n# We include a sixth right pop to our right populations vector right[6] = \u0026#34;Switzerland_Bichon_Epipaleolithic.SG\u0026#34; After adding the Bichon Epipaleolithic population to our right populations (which should serve as anchor for the WHG ancestry), we re-run the model and see:\n# A tibble: 3 × 5 target left weight se z \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Italian_North.HO Turkey_Marmara_Barcin_N.SG 0.579 0.0203 28.5 2 Italian_North.HO Russia_Samara_EBA_Yamnaya.AG 0.368 0.0226 16.3 3 Italian_North.HO Italy_Epigravettian.SG 0.0525 0.00910 5.77 $rankdrop # A tibble: 3 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested \u0026lt;int\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 2 3 7.78 5.08e- 2 5 593. 6.31e-126 2 1 8 601. 1.53e-124 7 1750. 0 3 0 15 2351. 0 NA NA NA The Epigravettian component is now well-resolved with a z-score of 5.77 and a much smaller standard error (± 0.91%). The model now passes with p = 0.0508, just above the 0.05 threshold. However, this borderline p-value suggests there\u0026rsquo;s room for improvement in the fit. To improve the fit, we\u0026rsquo;ll test whether adding a fourth source, a Levantine population, improves the model:\nleft[4] = \u0026#34;Lebanon_Hellenistic.SG\u0026#34; After adding Hellenistic-era samples from Lebanon as a fourth source, we re-run qpAdm and see:\n# A tibble: 4 × 5 target left weight se z \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Italian_North.HO Turkey_Marmara_Barcin_N.SG 0.450 0.0524 8.59 2 Italian_North.HO Russia_Samara_EBA_Yamnaya.AG 0.325 0.0328 9.89 3 Italian_North.HO Italy_Epigravettian.SG 0.0615 0.0112 5.49 4 Italian_North.HO Lebanon_Hellenistic.SG 0.164 0.0654 2.50 $rankdrop # A tibble: 4 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested \u0026lt;int\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 3 2 2.17 3.37e- 1 4 32.4 1.58e- 6 2 2 6 34.6 5.20e- 6 6 482. 5.04e-101 3 1 12 517. 5.23e-103 8 1485. 1.86e-315 4 0 20 2003. 0 NA NA NA The Levantine source contributes a significant 16.4% (± 6.54%), and the model fit improves substantially. The p-value rises to 0.337, well above the acceptance threshold, and all four ancestry components show strong statistical significance (z \u0026gt;= 2.5). This four-way model successfully captures Northern Italian ancestry as a mixture of Anatolian Neolithic farmers, Bronze Age steppe populations, Western hunter-gatherers, and Eastern Mediterranean ancestry.\nTo better constrain the timing of the Eastern Mediterranean ancestry, I\u0026rsquo;ll test whether Bronze Age Levantine samples provide a better fit than Hellenistic-era samples. We can simply replace the Hellenistic source in our left populations vector:\n# Replacing Hellenistic Lebanon with Lebanon_MBA left[4] = \u0026#34;Lebanon_MBA.SG\u0026#34; After re-running qpAdm we see:\n# A tibble: 4 × 5 target left weight se z \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Italian_North.HO Turkey_Marmara_Barcin_N.SG 0.402 0.0559 7.19 2 Italian_North.HO Russia_Samara_EBA_Yamnaya.AG 0.301 0.0266 11.3 3 Italian_North.HO Italy_Epigravettian.SG 0.0747 0.00981 7.61 4 Italian_North.HO Lebanon_MBA.SG 0.222 0.0666 3.34 $rankdrop # A tibble: 4 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested \u0026lt;int\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;int\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 3 2 0.262 8.77e- 1 4 46.0 2.43e- 9 2 2 6 46.3 2.60e- 8 6 598. 5.37e-126 3 1 12 645. 3.12e-130 8 1840. 0 4 0 20 2485. 0 NA NA NA The Bronze Age Levantine model shows excellent fit with p = 0.877, much higher than the Hellenistic model. The MBA source contributes 22.2% (± 6.66%), slightly more than the Hellenistic source, and all components remain highly significant. This suggests the Eastern Mediterranean ancestry in Northern Italians is better modeled by Bronze Age Levantine populations.\nThis model provides a solid fit for Northern Italian ancestry. However, qpAdm modeling is inherently iterative. You could continue refining by adding more informative right populations, testing alternative sources (as we did with the Levantine component), or exploring different right population combinations to see which provides the strongest resolution. The goal is finding a model that both passes statistically and makes historical sense.\nOne common issue you might encounter is negative ancestry proportions for a source population. This typically indicates that source isn\u0026rsquo;t contributing ancestry to your target. The usual response is to remove or replace that source and re-run the model. However, negative values can sometimes result from missing sources in your left populations or insufficient resolution in your right populations. Adding the right populations or an additional source can sometimes turn negative contributions positive.\nUsing the allsnps = TRUE Parameter When running qpAdm models, it\u0026rsquo;s generally recommended to use the allsnps = TRUE parameter. This setting selects different SNPs for each f4-statistic that qpAdm computes, which was the standard behavior in the original ADMIXTOOLS qpAdm implementation. This approach is particularly useful when working with sparse genotype data, as it maximizes the available genetic information for each individual statistic rather than restricting all statistics to a common SNP set. Using this setting typically provides more stable estimates and better model resolution.\nTo run the same model with allsnps = TRUE, simply add this parameter to the qpadm() function:\nresult = qpadm(data, left, right, target, allsnps = TRUE) Using our final Northern Italian model with Bronze Age Levantine ancestry, here\u0026rsquo;s the comparison:\nWith allsnps = TRUE:\n# A tibble: 4 × 5 target left weight se z 1 Italian_North.HO Turkey_Marmara_Barcin_N.SG 0.370 0.0434 8.53 2 Italian_North.HO Russia_Samara_EBA_Yamnaya.AG 0.361 0.0294 12.3 3 Italian_North.HO Italy_Epigravettian.SG 0.0678 0.00910 7.45 4 Italian_North.HO Lebanon_MBA.SG 0.202 0.0581 3.48 $rankdrop # A tibble: 4 × 7 f4rank dof chisq p dofdiff chisqdiff p_nested 1 3 2 1.54 4.64e- 1 4 69.8 2.48e- 14 2 2 6 71.4 2.15e- 13 6 697. 3.04e-147 3 1 12 768. 1.12e-156 8 2583. 0 4 0 20 3351. 0 NA NA NA Notice the differences compared to the default settings shown earlier: the ancestry proportions shift (Anatolian Neolithic decreases slightly from 40.2% to 37.0%, while Steppe increases from 30.1% to 36.1%), standard errors generally decrease, and the p-value drops from 0.877 to 0.464 but still passes comfortably. The chi-squared (chisq) value of 1.54 indicates good fit to the data relative to the selected right populations. The model remains valid, and the reduced standard errors indicate more precise estimates when using the full SNP set.\nThe allsnps = TRUE setting is valuable even when working with modern populations that have complete SNP coverage. Without it, similar but clearly different models can pass with relatively comparable fit statistics, making the test less stringent. Using allsnps = TRUE makes models \u0026ldquo;harder\u0026rdquo; to pass, which helps ensure you\u0026rsquo;re accepting only models that truly fit the data well.\nSummary A successful qpAdm model should have: a low chi-square value (indicating minimal deviation between observed and expected f4-statistics) and a p-value above 0.05 (model passes), low standard errors on ancestry proportions, ideally highly statistically significant contributions from each source, and all components making historical sense. However, passing models don\u0026rsquo;t necessarily mean these are the actual ancestral populations, in some cases it can be proxies for similar populations.\nqpAdm is a powerful tool for testing ancestry models and estimating admixture proportions, but the results require careful interpretation. Multiple models may fit equally well, which is why comparing alternative source combinations and evaluating them against historical and archaeological context is essential.\nThe choice of right populations critically affects model outcomes, as demonstrated by how adding the Epipaleolithic Bichon population resolved the Epigravettian component. Start with a diverse set of reference populations that span relevant ancestral divergences, and iterate from there.\n","permalink":"https://popgenblog.com/posts/qpadm-tutorial-admixture-modeling/","summary":"This post covers using qpAdm in R to test ancestry models and estimate admixture proportions. qpAdm builds on f4-statistics and provides a framework for evaluating whether proposed source populations ","tags":["ADMIXTOOLS"],"title":"Running qpAdm with ADMIXTOOLS2 in R: Testing and Interpreting Ancestry Models"},{"content":"This post covers how to run f4-statistics using the admixtools package for R. Compared with the original ADMIXTOOLS workflow, the R implementation is more convenient for testing multiple population combinations because it can be used interactively, without repeatedly editing parameter files.\nFor more in-depth examples, see Interpreting f4-Statistics with AdmixPy.\nWhat are f4-statistics? F4-statistics measure asymmetry in allele sharing among four populations. For four populations AAA, BBB, CCC, and DDD, the statistic is written as:\nf4(A,B;C,D)=1n∑i=1n(ai−bi)(ci−di) f_4(A, B; C, D) = \\frac{1}{n} \\sum_{i=1}^{n} (a_i - b_i)(c_i - d_i) f4​(A,B;C,D)=n1​i=1∑n​(ai​−bi​)(ci​−di​)Here, aia_iai​, bib_ibi​, cic_ici​, and did_idi​ are the allele frequencies of populations AAA, BBB, CCC, and DDD at SNP iii, and nnn is the number of SNPs used in the calculation.\nMathematically, the statistic is an average product of allele-frequency differences. In population-genetic terms, it tests whether allele sharing is symmetric across the quartet: does one population share the same amount of drift with two comparison populations, or is there a detectable excess affinity on one side? Under a simple tree model with no admixture connecting the two sides of the comparison, the statistic is expected to be zero. A significantly positive or negative value indicates asymmetric allele sharing, usually implying that the four populations do not fit that simple tree.\nThis might sound more complicated than it is. In essence you\u0026rsquo;re testing whether populations form a simple tree or show deviations consistent with gene flow between populations.\nInterpreting f4-statistics (with A as deep outgroup) When using a deep outgroup in position AAA, such as Mbuti, the sign tells you which populations share more drift:\nPositive f4(A,B;C,D)f_4(A,B;C,D)f4​(A,B;C,D): population BBB shares more drift with DDD than with CCC Negative f4(A,B;C,D)f_4(A,B;C,D)f4​(A,B;C,D): population BBB shares more drift with CCC than with DDD Zero or non-significant: BBB shares similar amounts of drift with CCC and DDD, consistent with a simple tree Because f4-statistics are ordered, changing the population order can change the sign. Swapping the two populations in the first pair, or swapping the two populations in the second pair, reverses the sign:\nf4(B,A;C,D)=−f4(A,B;C,D) f_4(B,A;C,D) = -f_4(A,B;C,D) f4​(B,A;C,D)=−f4​(A,B;C,D)f4(A,B;D,C)=−f4(A,B;C,D) f_4(A,B;D,C) = -f_4(A,B;C,D) f4​(A,B;D,C)=−f4​(A,B;C,D)This does not change the underlying relationship being tested, but it changes whether the result is reported as positive or negative. In particular, swapping AAA and BBB reverses the sign and changes which population is being interpreted as sharing more drift with CCC or DDD.\nIf you are new to f4-statistics, the easiest convention is to place a deep outgroup, such as Mbuti, in position AAA, keep the population you want to interpret in position BBB, which corresponds to pop2 in admixtools2, and then compare whether it shares more drift with CCC or DDD.\nStatistical significance is assessed using z-scores. The conventional threshold is ∣z∣\u0026gt;3|z| \u0026gt; 3∣z∣\u0026gt;3, though ∣z∣\u0026gt;2|z| \u0026gt; 2∣z∣\u0026gt;2 can also be useful for exploratory analysis.\nSetting Up R and Admixtools First, install R and the required dependencies. For Ubuntu/Debian:\nsudo apt update -y sudo apt install -y r-base r-base-dev build-essential \\ libcurl4-openssl-dev libssl-dev libxml2-dev libgsl-dev Next, start R by running R in your terminal, then install admixtools:\ninstall.packages(\u0026#34;remotes\u0026#34;) remotes::install_github(\u0026#34;uqrmaie1/admixtools\u0026#34;) Running f4-statistics I\u0026rsquo;ll use the AADR dataset, the most comprehensive curated collection of ancient and modern genomic samples. Since Admixtools works with EIGENSTRAT files and AADR is distributed in this format, the dataset can be used directly without preprocessing. For download instructions, see: Download Ancient \u0026amp; Modern DNA (AADR Tutorial).\nIf you want to also analyze your own DNA (23andMe, AncestryDNA, FamilyTreeDNA, MyHeritage, Living DNA) alongside AADR, see the Raw DNA to AADR toolkit, which handles the conversion and merge end-to-end.\nThen, navigate to the directory containing your downloaded AADR dataset files (.geno, .snp, .ind) and start R:\n# move into the folder so the files can be loaded without specifying paths cd /path/to/aadr/dataset R Load the previously installed admixtools library and dataset:\nlibrary(admixtools) # Specify the file prefix (without extensions) # If your files are named v62.0_HO_public.geno/.snp/.ind # use only the prefix: data = \u0026#34;v62.0_HO_public\u0026#34; When specifying populations in f4-statistics, use the exact population labels from the third column of your .ind file, the function will automatically include all samples assigned to that population label.\nNow we can run f4-statistics. You\u0026rsquo;ll need specific hypotheses about population relationships to test. I\u0026rsquo;ll demonstrate three examples that illustrate different outcomes.\nIn the examples below, pop1 through pop4 correspond to populations A, B, C, and D respectively, where (A, B) and (C, D) form the two pairs being compared.\nExample 1: Do Yamnaya Share More Drift with EHG than WHG? For this test, we\u0026rsquo;ll use Mbuti as the outgroup (position A), Yamnaya in position BBB, a WHG-related sample (Italian Epigravettian) in position CCC, and an EHG-related sample (Latvia Mesolithic) in position DDD.\nRun the following in your R session:\nf4(data, pop1=\u0026#34;Mbuti.HO\u0026#34;, pop2=\u0026#34;Russia_Samara_EBA_Yamnaya.AG\u0026#34;, pop3=\u0026#34;Italy_Epigravettian_alt.AG.BY.AA\u0026#34;, pop4=\u0026#34;Latvia_Mesolithic_oEHG.AG\u0026#34;) The result shows:\npop1 pop2 pop3 pop4 est se z p n \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Mbuti.HO Russia_Samara_EBA_Ya… Ital… Latv… 0.00245 9.17e-4 2.67 0.00751 49477 The z-score of 2.67 falls just below the conventional significance threshold of 3, but the result is still informative. While some use a stricter threshold of |z| \u0026gt; 3, others consider |z| \u0026gt; 2 sufficient for exploratory analysis, particularly when the direction of the effect is theoretically motivated. The positive value indicates that Yamnaya shares more drift with EHG than with the WHG sample from Villabruna. The output also reports n, which is the number of SNPs used in the calculation. Here, roughly 49k SNPs contributed to the statistic. In general, larger SNP counts increase accuracy, making it easier to reach |z| \u0026gt; 3.\nRe-running a closely related version of the Yamnaya test on a denser dataset, and using Russia_Samara_EN_Mesolithic in place of the previous EHG-related population, gives a strongly significant result (est = 0.00210, z = 10.4, p = 2.87e-25, n = 843572). With Mbuti as a deep outgroup, the positive sign again indicates that Yamnaya shares more alleles with Russia_Samara_EN_Mesolithic (position DDD) than with Italy_NordEst_Epigravettian (position CCC). So this is not just suggestive anymore: it is strong evidence of asymmetric allele sharing in this quartet.\nExample 2: Do Anatolian Neolithic Farmers Share More Drift with Sardinians or French? f4(data, pop1=\u0026#34;Mbuti.HO\u0026#34;, pop2=\u0026#34;Turkey_Marmara_Barcin_N.AG\u0026#34;, pop3=\u0026#34;Sardinian.HO\u0026#34;, pop4=\u0026#34;French.HO\u0026#34;) The result shows:\npop1 pop2 pop3 pop4 est se z p n \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Mbuti.HO Turkey_Marmara_Ba… Sard… Fren… -0.00233 1.25e-4 -18.6 2.49e-77 155527 This result is highly significant with a z-score of -18.6. The negative value indicates that Anatolian Neolithic farmers (population BBB) share significantly more drift with Sardinians (population CCC) than with the French (population DDD). This is consistent with Sardinians retaining the highest proportion of Neolithic farmer ancestry in Europe.\nIn this population setup, that is naturally interpreted as Sardinians having more Anatolian farmer-related ancestry than French.\nBut the interpretation depends on the comparison populations. If Sardinians were replaced by Han Chinese, the statistic would become even more significant, but for a less informative reason: French share far more Anatolian farmer-related ancestry and allele-frequency drift with Anatolian Neolithic farmers than Han Chinese do. The much larger signal would mostly reflect the use of an extremely unequal distant comparator, not a test of Neolithic farmer ancestry differences within Europe.\nThe f4-statistic itself only measures allele-sharing asymmetry in the quartet. The historical meaning comes from choosing populations that isolate the contrast you want to test.\nExample 3: Comparing French, Scottish, and Norwegian Relationships f4(data, pop1=\u0026#34;Mbuti.HO\u0026#34;, pop2=\u0026#34;French.HO\u0026#34;, pop3=\u0026#34;Scottish.HO\u0026#34;, pop4=\u0026#34;Norwegian.HO\u0026#34;) The result shows:\npop1 pop2 pop3 pop4 est se z p n \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Mbuti.DG French.HO Scottish.HO Norwegian.HO 1.92e-4 2.05e-4 0.941 0.347 155583 The f4-statistic is non-significant (z = 0.941), meaning this particular ordered test does not detect excess affinity of French to either Scottish or Norwegian. With Mbuti as the outgroup and French in position B, the test only asks whether French share more drift with Scottish or with Norwegian, it does not address whether Scottish and Norwegian themselves share excess affinity with each other. As we\u0026rsquo;ll see when rotating the populations below, a non-significant result in one orientation does not rule out asymmetries that become visible in another.\nTo explore the relationship more completely, you can rotate which population is placed in position BBB. This changes the question being asked: instead of only asking whether French are closer to Scottish or Norwegian, you can also ask whether Scottish are closer to French or Norwegian, and whether Norwegian are closer to French or Scottish.\nFor example:\n# Are French closer to Scottish or Norwegian? f4(data, pop1=\u0026#34;Mbuti\u0026#34;, pop2=\u0026#34;French\u0026#34;, pop3=\u0026#34;Scottish\u0026#34;, pop4=\u0026#34;Norwegian\u0026#34;) # Are Scottish closer to French or Norwegian? f4(data, pop1=\u0026#34;Mbuti\u0026#34;, pop2=\u0026#34;Scottish\u0026#34;, pop3=\u0026#34;Norwegian\u0026#34;, pop4=\u0026#34;French\u0026#34;) # Are Norwegian closer to Scottish or French? f4(data, pop1=\u0026#34;Mbuti\u0026#34;, pop2=\u0026#34;Norwegian\u0026#34;, pop3=\u0026#34;Scottish\u0026#34;, pop4=\u0026#34;French\u0026#34;) Note: These rotated examples use AADR v66 labels, which no longer include the .HO suffix. The interpretation is unchanged.\nWhen running the f4-statistic with Norwegian in position BBB:\npop1 pop2 pop3 pop4 est se z p n \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Mbuti Norwegian Scottish French -0.000861 0.000195 -4.42 0.00000980 678870 The negative sign, with AAA as a deep outgroup, indicates that Norwegian shares significantly (|z| = 4.42) more drift (est = -0.000861) with Scottish than with French. This demonstrates why the population in position BBB matters: the earlier French-centered test asked whether French are closer to Scottish or Norwegian, while this test asks whether Norwegian is closer to Scottish or French. These are related questions, but they are not the same ordered comparison.\nWhen running the f4-statistic with Scottish in position BBB:\npop1 pop2 pop3 pop4 est se z p n \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;chr\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; \u0026lt;dbl\u0026gt; 1 Mbuti Scottish Norwegian French -0.000602 0.000258 -2.34 0.0195 678870 The negative sign points in the same direction: Scottish tendentially share more drift with Norwegian than with French. However, the result is below the stricter ∣z∣\u0026gt;3|z| \u0026gt; 3∣z∣\u0026gt;3 threshold (|z| = 2.34). In context, it is still consistent with the broader expectation that Scottish and Norwegian are more closely related to each other than either is to French.\nConclusion F4-statistics are excellent for detecting admixture signals and differential relatedness, but they should not be overinterpreted. A significant f4-statistic tells you that four populations do not fit a simple tree and that one pair shows excess allele sharing relative to the other, but it does not tell you the direction of gene flow, whether the connection is direct or indirect, or the ancestry proportions involved. For that kind of admixture modeling, qpAdm is the next step.\nAt the same time, f4-statistics are also useful for asking whether populations can be placed on a simple tree in a way that is consistent with the data. When an f4-statistic is significantly different from zero, that proposed tree-like relationship is rejected. For this reason, f4-statistics are also needed for qpGraph.\n","permalink":"https://popgenblog.com/posts/interpreting-f4-tests-admixtools/","summary":"This post covers how to run f4-statistics using the admixtools package for R. Compared with the original ADMIXTOOLS workflow, the R implementation is more convenient for testing multiple population co","tags":["ADMIXTOOLS"],"title":"How to Run and Interpret f4-Statistics in R: AADR Examples"},{"content":"Recently, MyHeritage quietly announced that all new DNA kits will be processed using low-pass whole-genome sequencing (WGS) instead of traditional genotyping arrays, for just $30 per kit. The coverage depth will be roughly 2x (as announced in their blog post), compared to the 30x used in clinical sequencing. It’s nowhere near diagnostic quality, but it is whole-genome data nevertheless.\nThe surprising part is the price. Getting an entire genome, even at shallow depth, for less than what MyHeritage charges to “unlock” an uploaded SNP file (~$35) seems almost too cheap. Below are a few of the issues I see: technical, economic, and data-security related.\nA History of Security Concerns Before examining the technical implications of this WGS announcement, it\u0026rsquo;s worth remembering MyHeritage\u0026rsquo;s track record with data protection. In June 2018, the company disclosed that ~92 million user accounts were compromised in a breach that occurred on October 26, 2017 MyHeritage, a breach they only learned about seven months later when a security researcher contacted them.\nThe breach exposed email addresses and passwords for every user who had signed up before the breach date. While MyHeritage publicly stated that the passwords were salted and hashed, some users later reported finding their MyHeritage credentials in third-party breach compilations or OSINT tools, possibly due to post-breach password cracking or data merging. Although the original MyHeritage disclosure referred only to salted hashes, these later appearances reinforced public doubts about the company’s earlier security practices.\nThe company failed to detect the intrusion on their own, relying instead on external notification after a seven-month delay. This raises a critical question: if MyHeritage couldn\u0026rsquo;t detect unauthorized access to email and password databases for over half a year, what confidence should users have in the security of far more sensitive whole-genome sequence data? The 2018 breach involved relatively low-value account credentials. The new WGS data represents something fundamentally different: immutable biometric information that, once leaked, cannot be changed like a password.\nThe Economics Don\u0026rsquo;t Add up MyHeritage is using Ultima Genomics platform, which can sequence human genomes for as little as $100 at 30x coverage and with the newer Solaris chemistry, that cost has dropped to $80 per genome at 30x. But there\u0026rsquo;s a critical detail: those prices are for 30x coverage. MyHeritage is offering 2x coverage.\nBreaking Down the Real Costs With the UG100 Solaris figures quoted, a 30x run corresponds to ~333 million reads. That scales to ~11 million reads per 1x and ~22 million reads for 2x. At $0.24 per million reads, the raw sequencing cost (reagents + instrument time) for 2x coverage is roughly $5–6 per sample. But that’s only the starting point. Once you factor in the rest of the pipeline, even at massive scale, the total cost rises quickly:\nLibrary preparation: Even with simplified workflows, prep kits and labor add $10-20 per sample Sample handling and logistics: Receiving kits, DNA extraction, quality control Data storage: Even at wholesale cloud storage rates, storing whole-genome data for millions of users isn\u0026rsquo;t trivial Bioinformatics processing: Alignment, variant calling, quality metrics Customer service, shipping, marketing, platform maintenance Even under generous assumptions, a lean 2x workflow still totals around $25–45 per kit. At a $30 retail price, MyHeritage is operating on razor-thin margins, or absorbing a loss in exchange for building a large-scale genomic dataset.\nA comparison MyHeritage charges approximately $35 to \u0026ldquo;unlock\u0026rdquo; analysis features for DNA data you upload from another testing service. You\u0026rsquo;ve already paid for the sequencing elsewhere, extracted your own DNA, shipped it to another company, and downloaded your raw data file. All MyHeritage has to do is run your existing SNP data through their algorithms and database. They charge $35 for that.\nBut they charge only $30 to:\nShip you a kit Extract and sequence your entire genome (even at low coverage) Store terabytes (at scale) of WGS data indefinitely Process it through their matching algorithms Provide all the same features as the $35 upload unlock The most plausible explanation: your genome data is worth more to them than the $30 you\u0026rsquo;re paying.\nWhat Are They Really Selling? It’s worth noting that MyHeritage publicly commits to not selling or licensing individual genomic data and claims to operate under GDPR-aligned standards. Whether those safeguards will hold as the dataset expands remains to be seen. From a business perspective, though, the pricing still looks like a classic loss-leader strategy. The $30 kit isn’t likely profitable on its own, it’s the cost of acquiring large-scale genomic data. The real products lie elsewhere:\nPharmaceutical partnerships and research licensing: Companies like 23andMe have secured deals worth hundreds of millions with pharmaceutical giants like GSK. \u0026ldquo;De-identified\u0026rdquo; genomic data is extraordinarily valuable for drug development, population genetics research, and biomarker discovery.\nAggregate data sales: Even if individual genomes are \u0026ldquo;anonymized,\u0026rdquo; aggregated WGS data has immense commercial value. Genomics companies are increasingly positioning themselves not as sequencing providers, but as data companies, the Facebooks and Googles of genetic data.\nLifetime monetization: A microarray test gives MyHeritage ~700,000 data points about you. Whole-genome sequencing gives them 3 billion. Even at 2x coverage, that\u0026rsquo;s dramatically more information about:\nRare genetic variants Structural variations Ancestry-informative markers Health-related associations Drug metabolism genes Future research applications we haven\u0026rsquo;t even invented yet You pay once. They own your complete genome forever. The lifetime value of that data, through research partnerships, database licensing, and future applications, far exceeds $30.\nIf their goal is just ancestry inferring and genealogy matching (which works perfectly fine with SNP arrays), why go through the expense of whole-genome sequencing at all?\nIn 10 years, when precision medicine takes off and rare variant databases become commercially viable, MyHeritage will have whole genomes on file, not just the SNPs that were relevant in 2024.\n","permalink":"https://popgenblog.com/posts/myheritage-30-euro-genome-analysis/","summary":"Recently, MyHeritage quietly announced that all new DNA kits will be processed using low-pass whole-genome sequencing (WGS) instead of traditional genotyping arrays, for just $30 per kit. The coverage","tags":[],"title":"MyHeritage's $30 WGS: Technical Analysis"},{"content":"In an earlier post, PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output), I showed the manual way to build a subset from the .ind/.fam. That works, but if you want to keep thousands of samples it gets tedious fast. Below is a one-liner using awk that generates a PLINK --keep file automatically from a list of populations.\nPrepare a list of populations to keep: Create a text file (e.g. pops) in the same directory as your reference .ind and .fam. Put one population label per line:\nNorwegian.HO Sardinian.DG Han.HO You can include as many as you need.\nAssumption: your .ind is standard EIGENSTRAT (col1 = IID and col3 = population), and your PLINK .fam corresponds to the same individuals (col1 = FID, col2 = IID). If you are working with the AADR, use the original .ind as downloaded.\nGenerate a PLINK --keep file with awk: awk -v OFS=\u0026#39;\\t\u0026#39; \u0026#39; FNR==1 {f++} f==1 {want[$1]=1; next} # populations to keep f==2 {if ($3 in want) keep[$1]=1; next} # reference.ind (IID in col1, pop in col3) f==3 {if ($2 in keep) print $1,$2} # reference.fam (FID col1, IID col2) \u0026#39; pops reference.ind reference.fam \u0026gt; pops.keep What it does:\nreads your population list (pops) scans the .ind and marks all matching IIDs emits FID\\tIID pairs for those IIDs by looking them up in .fam pops.keep is now ready for PLINK.\nRun PLINK to keep the subset: # Example plink --bfile reference --keep pops.keep --make-bed --out subset You now have a packed binary PLINK dataset containing all samples belonging to the population labels listed in pops.\nRelated tutorials How to Download the AADR Dataset (Linux \u0026amp; WSL) PLINK PCA Tutorial: Running PCA in PLINK Converting EIGENSTRAT to PACKEDPED Format Running ADMIXTURE in Unsupervised Mode: Full Tutorial \u0026amp; Python Plotting Script ","permalink":"https://popgenblog.com/posts/awk-subset-populations-genetics/","summary":"In an earlier post, PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output), I showed the manual way to build a subset from the .ind/.fam. That works, but if you want to keep thousands of samples","tags":["PLINK","EIGENSTRAT","awk"],"title":"How to Subset Genetic Samples by Population Labels with awk (Create PLINK --keep file)"},{"content":"This post is a short follow-up to the previous one on Estimating Ancestry Components Using ADMIXTURE. Here, we’ll explore supervised ADMIXTURE, a mode that allows you to explicitly define ancestral populations and infer the ancestry proportions of unassigned individuals based on those references.\nWhat Is a Supervised Run? In supervised mode, ADMIXTURE skips the component discovery step and instead uses user-defined groupings to represent ancestral components. The benefit: if you already have solid candidates for reference populations, you can use them to quickly infer ancestry proportions for target or admixed individuals.\nSupervised runs are much faster, especially with large datasets, since ADMIXTURE no longer has to determine the internal structure of each cluster.\nCreating a .pop File Supervised ADMIXTURE requires an additional file with the extension .pop. This file must:\nHave the exact same prefix as your .bed input file. Contain the same number of rows as your .fam file (one row per sample). Use labels (strings or integers) to assign each individual to a known ancestral population. Use - (or leave the line empty) for individuals whose ancestry you want ADMIXTURE to infer. For example, given a dataset admix.final.bed, create a file named admix.final.pop:\nNorthern_West_Asia Northern_West_Asia Northern_West_Asia - Northern_West_Asia In this example, the first three and the fifth samples are assigned to a component labeled Northern_West_Asia. The fourth sample (-) is unassigned and will be inferred based on the others. You can define as many distinct labels as components you want ADMIXTURE to model (Ks). The number of unique labels in the .pop file determines K, the number of ancestral populations you’ll specify in the run.\nRunning ADMIXTURE in Supervised Mode Once your .pop file is ready, you can start the supervised run with:\n# Assuming you used six distinct populations ./admixture --supervised admix.final.bed 6 ADMIXTURE will then use the population labels from your .pop file to build the ancestral components and infer proportions for unassigned individuals.\nWhen to Use Supervised Mode Supervised ADMIXTURE is especially useful when:\nYou have high-confidence reference populations (ancient samples or unadmixed modern groups). You want to test how certain samples relate to predefined clusters. You want faster runs compared to iterative unsupervised testing. It’s important to note that component quality is entirely in your hands here. ADMIXTURE won’t \u0026ldquo;fix\u0026rdquo; poor references. Ideally, your reference populations should be homogeneous, representative, and contain many individuals.\nVisualizing Results The output files are the same as in unsupervised runs:\n.Q: Ancestry proportions per individual. .P: Allele frequencies per component. You can reuse the same plotting script from the previous post. Since your .pop labels now directly correspond to components, interpreting the bar plot becomes more straightforward.\n","permalink":"https://popgenblog.com/posts/admixture-supervised-tutorial/","summary":"This post is a short follow-up to the previous one on Estimating Ancestry Components Using ADMIXTURE. Here, we’ll explore supervised ADMIXTURE, a mode that allows you to explicitly define ancestral po","tags":["ADMIXTURE"],"title":"Running ADMIXTURE in Supervised Mode"},{"content":"In this post, I’ll demonstrate how to estimate ancestry proportions using one of the most widely used tools in population genetics: ADMIXTURE. ADMIXTURE is a model-based clustering algorithm that estimates individual ancestry proportions and ancestral allele frequencies from multilocus SNP genotypes.\nPreparing the Dataset Download the appropriate ADMIXTURE binary and either place it in your dataset directory or make it globally accessible. For this run, I included a subset of West Asian populations along with a few adjacent populations (around 150 samples in total). Linkage Disequilibrium (LD) pruning was applied beforehand. If you\u0026rsquo;re unsure how to prune your dataset, refer to the previous post.\nEven after pruning, my dataset still contained over 140,000 SNPs. Since ADMIXTURE can be memory-intensive, especially given that multiple runs are often needed for interpretation, it\u0026rsquo;s a good idea to reduce the dataset further using PLINK’s --thin-count flag. This randomly subsamples SNPs and significantly speeds up ADMIXTURE runs.\nExample: thinning to 50k SNPs:\nplink --bfile admix.data.pruned --thin-count 50000 --make-bed --out admix.final This will produce a PACKEDPED-format dataset (.bed/.bim/.fam) with 50,000 randomly selected SNPs.\nRunning ADMIXTURE ADMIXTURE supports both unsupervised and supervised modes. In this post, I’ll focus on unsupervised runs. Given a user-defined number of ancestral populations (K), ADMIXTURE estimates individual ancestry proportions and ancestral allele frequencies by maximizing the likelihood of the observed genotypes. Basic run (for K = 6):\n# ADMIXTURE expects a .bed file as input # Running with 8 threads ./admixture -j8 admix.final.bed 6 There’s no definitive way to choose the right number of ancestral components (K); it usually involves a combination of testing and intuition. While ADMIXTURE supports cross-validation (--cv), in practice, evaluating the plausibility of inferred components tends to be more informative. Overfitting (too many Ks) often results in arbitrary splits of populations. For example, in my test dataset, K=7 caused the Mongol samples to split into two distinct components, probably reflecting internal structure, but for most purposes this was excessive, since there’s no clear indication of what that split represents for this dataset.\nFor my K=6 run, I expected the following components:\nNorthern West Asian component Divergent Zoroastrian component North African component (peaking in Mozabites) Sub-Saharan African component (peaking in Malawi) East Asian component (peaking in Mongols) Southern West Asian component (peaking in Yemeni Highlanders) A higher K does not guarantee the appearance of expected or known populations. Component formation depends on statistical divergence.\nOutput Files ADMIXTURE generates:\n.Q – ancestry proportions per individual (aligned with rows in the .fam file) .P – allele frequencies per component and SNP (aligned with the .bim file) ADMIXTURE components must be interpreted from their distribution across samples in the dataset. For example, if the first 10 individuals are Mozabites and all show a very high proportion of the same component, say \u0026gt;0.95, that component most likely reflects North African ancestry.\nPlotting a Bar Plot To interpret ADMIXTURE’s output visually, we’ll plot the .Q file, which contains ancestry proportions for each individual in the dataset. Each row corresponds to a sample in the .fam file, and each column to one of the inferred components (K).\nTo make the plot more readable, we’ll group samples by population using a labels file - one line per sample, aligned with the .fam file.\nIf you haven’t created a labels file for your current subset, see Performing PCA on Genetic Data Using PLINK for a quick method. In short:\n# Adjust the .ind to your prefix awk \u0026#39;NR==FNR {map[$1]=$3; next} {print map[$2]}\u0026#39; v62.0_HO_public.ind admix.final.fam \u0026gt; labels This command matches each individual in the .fam file to their population label from the .ind file (column 3), creating a label file that aligns row-wise with the ADMIXTURE output.\nOnce you have the .Q file and the labels file, you can use the Python script below to plot the ADMIXTURE proportions as a bar plot grouped by population. If you don\u0026rsquo;t have a Python environment with the required libraries yet, set one up in your ADMIXTURE output folder:\n# Create and activate a virtual environment python3 -m venv venv source venv/bin/activate # On Windows: venv\\Scripts\\activate Then install the dependencies:\npip3 install numpy pandas matplotlib Now save the script below as barplot.py, adjust the configuration parameters at the top:\nimport numpy as np import pandas as pd import matplotlib.pyplot as plt from collections import defaultdict # --- Configuration --- q_file = \u0026#34;admix.final.6.Q\u0026#34; # Path to your .Q file labels_file = \u0026#34;labels\u0026#34; # Path to your labels file figsize = (20, 6) label_fontsize = 10 min_group_size = 3 # Only show label if group size ≥ this bar_width = 1.0 # --- Load data --- Q = pd.read_csv(q_file, sep=r\u0026#34;\\s+\u0026#34;, header=None) labels = pd.read_csv(labels_file, header=None)[0] assert len(Q) == len(labels), \u0026#34;Mismatch between Q rows and labels\u0026#34; # --- Sort by labels to merge same-pop groups --- sorted_idx = labels.argsort() Q_sorted = Q.iloc[sorted_idx].reset_index(drop=True) labels_sorted = labels.iloc[sorted_idx].reset_index(drop=True) # --- Compute new plotting indices --- N, K = Q_sorted.shape x = np.arange(N) # --- Set colors --- if K \u0026lt;= 10: cmap = plt.get_cmap(\u0026#34;tab10\u0026#34;) colors = [cmap(i) for i in range(K)] elif K \u0026lt;= 20: cmap = plt.get_cmap(\u0026#34;tab20\u0026#34;) colors = [cmap(i) for i in range(K)] else: # continous map for K \u0026gt; 20 cmap = plt.get_cmap(\u0026#34;hsv\u0026#34;) colors = [cmap(i / K) for i in range(K)] # --- Plot bars --- fig, ax = plt.subplots(figsize=figsize) bottom = np.zeros(N) for k in range(K): col = Q_sorted.iloc[:, k].to_numpy() ax.bar( x, col, bottom=bottom, color=colors[k], width=bar_width, linewidth=0, ) bottom += col # --- Group and label if ≥ min_group_size --- group_positions = defaultdict(list) for idx, label in enumerate(labels_sorted): group_positions[label].append(idx) for label, indices in group_positions.items(): if len(indices) \u0026gt;= min_group_size: mid = (indices[0] + indices[-1]) / 2 ax.plot([mid, mid], [-0.03, 0], color=\u0026#34;black\u0026#34;, lw=0.8) ax.text( mid, -0.04, label, ha=\u0026#34;right\u0026#34;, va=\u0026#34;top\u0026#34;, fontsize=label_fontsize, rotation=45, rotation_mode=\u0026#34;anchor\u0026#34;, ) # --- Final styling tweaks --- ax.set_xlim(-0.5, N - 0.5) ax.set_ylim(-0.08, 1.05) ax.set_ylabel(\u0026#34;Ancestry\u0026#34;, fontsize=12) ax.set_xticks([]) ax.set_yticks(np.linspace(0, 1, 6)) ax.tick_params(axis=\u0026#34;both\u0026#34;, length=0) ax.spines[\u0026#34;top\u0026#34;].set_visible(False) ax.spines[\u0026#34;right\u0026#34;].set_visible(False) ax.spines[\u0026#34;bottom\u0026#34;].set_position(\u0026#34;zero\u0026#34;) plt.subplots_adjust(left=0.06, bottom=0.18) plt.show() # To save instead of displaying, comment out plt.show() and uncomment the line below: # plt.savefig(\u0026#34;admixture_plot.png\u0026#34;, dpi=300, bbox_inches=\u0026#34;tight\u0026#34;) Finally run it with:\npython3 barplot.py After running the script, you\u0026rsquo;ll have a visual representation of your ADMIXTURE results with individuals grouped by population and colored by their inferred ancestry components.\nHere’s the bar plot generated for my West Asian subset at K=6: ","permalink":"https://popgenblog.com/posts/admixture-unsupervised/","summary":"In this post, I’ll demonstrate how to estimate ancestry proportions using one of the most widely used tools in population genetics: ADMIXTURE. ADMIXTURE is a model-based clustering algorithm that esti","tags":["ADMIXTURE","awk"],"title":"How to Run ADMIXTURE (Unsupervised): Full Tutorial \u0026 Python Plotting Script"},{"content":"This post is a continuation of the previous one, where I demonstrated how to perform PCA with PLINK. While PLINK’s PCA is great for quick, exploratory analysis, smartpca (part of the EIGENSOFT toolset) is particularly common in population-genetic and ancient-DNA studies.\nSmartpca can be compiled from the EIGENSOFT source or installed through conda. I covered the installation process in this earlier post: From EIGENSTRAT to PACKEDPED.\nAs before, I’ll use a small subset. The focus here is on the technical process. One key difference in this post is that I’ll perform Linkage Disequilibrium (LD) pruning, which reduces redundancy between correlated SNPs before PCA.\nLD Pruning # Window size: 50 SNPs # Step size: 5 SNPs # LD threshold: r² \u0026gt; 0.2 will be pruned plink --bfile input --indep-pairwise 50 5 0.2 # Extract the pruned SNPs into a new dataset plink --bfile input \\ --extract plink.prune.in \\ --make-bed \\ --out final Note: If you’re working with low-coverage ancient samples, I would avoid estimating LD directly from a mixed modern + ancient dataset. Ancient samples usually have considerably more missing data, and often pseudo-haploid genotypes, which makes them poorly suited for estimating LD.\nA better approach is:\nCalculate the LD-pruning set using the modern reference samples. Apply the same SNP list to the ancient samples. Merge the resulting datasets. For example:\n# Step 1: Prune the modern samples plink --bfile modern \\ --indep-pairwise 50 5 0.2 # Step 2: Apply the same SNP list to both datasets plink --bfile modern \\ --extract plink.prune.in \\ --make-bed \\ --out pruned_modern plink --bfile ancient \\ --extract plink.prune.in \\ --make-bed \\ --out ancient_extracted # Step 3: Merge the datasets plink --bfile pruned_modern \\ --bmerge ancient_extracted \\ --make-bed \\ --out final This gives us our final PACKEDPED dataset:\nfinal.bed final.bim final.fam Preparing the PACKEDPED Dataset for Smartpca Smartpca can read PLINK\u0026rsquo;s binary PACKEDPED format directly, so there is no need to convert the dataset back to EIGENSTRAT or PACKEDANCESTRYMAP first.\nThe three files can be supplied directly to smartpca:\ngenotypename: final.bed snpname: final.bim indivname: final.fam A PLINK .fam file contains six columns:\nFID IID father mother sex phenotype For example:\n2032 TLA018.HO 0 0 1 2 2033 TLA019.HO 0 0 1 2 2034 TLA020.HO 0 0 1 2 When EIGENSOFT reads PED or PACKEDPED data, the sixth column can also be used as a population group label. This is useful because smartpca uses population labels to determine which samples define the PCA axes and which are projected.\nFor this PCA, I’ll keep the sixth column as 1 for reference samples and the samples I want to project are labelled P:\n2032 TLA018.HO 0 0 1 1 2033 TLA019.HO 0 0 1 P 2034 TLA020.HO 0 0 1 1 After replacing the phenotype column with population labels, I would treat this .fam as input for smartpca rather than as a normal PLINK phenotype file. If you need to preserve phenotype information for another analysis, keep a copy of the original .fam.\nCreating the Smartpca Parameter File Like most EIGENSOFT tools, smartpca uses a parameter file to specify its input, output and analysis options.\nCreate a file called pca_param:\ngenotypename: final.bed snpname: final.bim indivname: final.fam evecoutname: pca.evec evaloutname: pca.eval altnormstyle: NO numoutevec: 10 numoutlieriter: 0 lsqproject: YES poplistname: pca.poplist familynames: NO numthreads: 5 There are three options here that are particularly important for this workflow.\nlsqproject lsqproject: YES This enables projection using least-squares equations.\nWhen the PCs are calculated from higher-quality reference samples while lower-coverage or higher-missingness ancient samples are projected onto those axes.\nOn its own, however, lsqproject does not decide which samples define the PCA axes.\nFor that, we use poplistname.\npoplistname In the example above, all samples used to define the PCA axes have 1 in the sixth column of final.fam, while projected samples have P.\nIn PACKEDPED input, smartpca interprets the phenotype code 1 as the population label Control. The custom label P is kept as-is. So we create a population list containing only Control:\nRun:\necho Case \u0026gt; pca.poplist Smartpca will calculate the principal components using samples belonging to populations listed in pca.poplist.\nSince P is not in the list, those individuals do not influence the calculation of the axes and are instead projected onto them.\nIf your reference samples already have real population labels rather than a single Case label, you can use those directly. For example:\n1001 HGDP01001 0 0 1 Sardinian 1002 HGDP01002 0 0 2 Greek 1003 HGDP01003 0 0 1 Russian 2033 TLA019.HO 0 0 1 P Then generate the reference population list with:\nawk \u0026#39;$6!=\u0026#34;P\u0026#34;{print $6}\u0026#39; final.fam | sort -u \u0026gt; pca.poplist This would create something like:\nGreek Russian Sardinian familynames I also set:\nfamilynames: NO By default, EIGENSOFT combines the PLINK family ID and individual ID when reading PED/PACKEDPED data.\nFor example:\n11788 HGDP00553.HO would otherwise appear in the PCA output as:\n11788:HGDP00553.HO I don\u0026rsquo;t need the family IDs here, so disabling this behavior keeps the sample IDs as:\nHGDP00553.HO This also makes the plotting steps below considerably simpler.\nRunning Smartpca Now run smartpca with:\nsmartpca -p pca_param The two main output files are:\npca.eval pca.evec pca.eval contains the eigenvalues, while pca.evec contains the sample IDs, principal component coordinates and population labels.\nPreparing for Plotting To make the .evec output compatible with the plotting script from my previous tutorial, convert it to CSV and remove the final population column:\nawk \u0026#39;NR\u0026gt;1 { $NF=\u0026#34;\u0026#34; sub(/[ \\t]+$/, \u0026#34;\u0026#34;) gsub(/[ \\t]+/, \u0026#34;,\u0026#34;) print }\u0026#39; pca.evec \u0026gt; smartpca.csv The resulting file looks something like:\nHGDP00553.HO,0.0136,0.0104,0.0353,-0.0510,... HGDP00554.HO,0.0136,0.0107,0.0334,-0.0494,... HGDP00555.HO,0.0137,0.0104,0.0346,-0.0511,... At this point, you have two options:\nUse the Python plotting script*from my previous PLINK PCA tutorial Use Vahaduo for browser-based interactive visualization Option 1: Plotting with Python If you\u0026rsquo;ve already set up Python and created the plotting script from the PLINK tutorial, you can use it directly.\nThe Case and P labels in final.fam were only used to tell smartpca which samples should define the axes. They are not useful as population labels for plotting.\nSince the original dataset already contains population assignments in data.ind, we can recover them from there.\nCreate a labels file with:\nawk \u0026#39; NR==FNR { map[$1]=$3 next } { split($0, a, \u0026#34;,\u0026#34;) print map[a[1]] } \u0026#39; data.ind smartpca.csv \u0026gt; labels Then use:\nsmartpca.csv labels with the plotting script from my previous PLINK PCA tutorial.\nThis is the resulting plot for a European subset:\nNote: The apparent spread of some samples can be exaggerated when PCA is calculated from a very small reference subset. With a larger reference panel, the axes generally become more stable and individual outliers tend to appear less extreme.\nThis is one reason ancient-DNA studies commonly calculate PCA axes using larger, higher-quality reference panels and project ancient samples rather than allowing sparse ancient genotypes to influence the axes themselves.\nOption 2: Plotting with Vahaduo (No Coding Required) Vahaduo is a web-based tool that can generate interactive 2D and 3D PCA plots without requiring a Python setup.\nHowever, it requires population labels to be embedded directly in the CSV file.\nOur current smartpca.csv looks like this:\n... HGDP00553.HO,0.0136,0.0104,0.0353,-0.0510,... HGDP00554.HO,0.0136,0.0107,0.0334,-0.0494,... HGDP00555.HO,0.0137,0.0104,0.0346,-0.0511,... ... Vahaduo groups samples using the text before the : character, so we can prepend the original population label to each sample ID.\nThe following awk command uses the original data.ind file from the downloaded EIGENSOFT dataset:\nawk -F\u0026#39;,\u0026#39; -v OFS=\u0026#39;,\u0026#39; \u0026#39; NR == FNR { line = $0 gsub(/^[[:space:]]+/, \u0026#34;\u0026#34;, line) if (line == \u0026#34;\u0026#34;) next split(line, a, /[[:space:]]+/) if (a[1] == \u0026#34;\u0026#34; || a[3] == \u0026#34;\u0026#34;) next label = a[3] # Remove common dataset suffixes from population labels sub(/\\.(AG|DG|HO|SG)$/, \u0026#34;\u0026#34;, label) pop[a[1]] = label next } { if ($1 in pop) $1 = pop[$1] \u0026#34;:\u0026#34; $1 print } \u0026#39; data.ind smartpca.csv \u0026gt; smartpca_with_labels.csv This transforms the file into, for example:\n... Papuan:HGDP00553.HO,0.0136,0.0104,0.0353,-0.0510,... Papuan:HGDP00554.HO,0.0136,0.0107,0.0334,-0.0494,... Papuan:HGDP00555.HO,0.0137,0.0104,0.0346,-0.0511,... ... Using Vahaduo Go to https://vahaduo.github.io/custompca/ Navigate to the PCA tool Paste the contents of smartpca_with_labels.csv into the PCA Data section Click Plot PCA in the PCA plot section Vahaduo provides several convenient features:\nNo Python setup required Interactive plots Zooming and panning 3D visualization Hovering over individual samples Quick exploratory analysis Saving plots PCA-Based Mixture Models Vahaduo can also run a distance-minimizing heuristic that fits a target sample as a mixture of selected source groups, and the CSV produced above can be used with that feature.\nPCA-based models are dataset-dependent and measure proximity in a reduced-dimensional space. That proximity can be useful for exploration, but it cannot by itself precisely infer ancestry or admixture proportions.\n","permalink":"https://popgenblog.com/posts/smartpca-tutorial/","summary":"This post is a continuation of the previous one, where I demonstrated how to perform PCA with PLINK. While PLINK’s PCA is great for quick, exploratory analysis, smartpca (part of the EIGENSOFT toolset","tags":["PLINK","PCA","convertf","awk"],"title":"SmartPCA Tutorial: How to Run PCA on Genetic Data"},{"content":"In this post, I’ll demonstrate how to perform a PCA on a PLINK dataset. Before we begin, we need to prepare a subset of samples we\u0026rsquo;re interested in analyzing.\nTo do this, we’ll extract sample information from the .fam file. But first, we need to identify the samples of interest. For example, those from a specific population such as Sardinians.\nThe easiest way is to open the corresponding .ind file and look at the population column, which is the third column in each row. Open the file in a text editor, and search for the population name, in this case, Sardinian.\nYou should find entries like this:\n... HGDP01075.HO M Sardinian.HO HGDP01076.HO M Sardinian.HO HGDP01077.HO M Sardinian.HO ... These entries tell us which IIDs (individual IDs) belong to Sardinian samples. Now take note of these IIDs, for example: HGDP01075.HO.\nNext, open your .fam file (from the previously created PACKEDPED dataset). You’ll see something like:\n... 1664 HGDP01075.HO 0 0 1 2 1665 HGDP01076.HO 0 0 1 2 1666 HGDP01077.HO 0 0 1 2 ... Now match the IIDs from the .ind file to those in the .fam file. For each matching line in the .fam, copy the entire line into a new file. This new file will be used with PLINK\u0026rsquo;s --keep flag to retain only the samples of interest.\nNote: Since the .ind and .fam files are row-aligned, if you identify a block of consecutive samples in the .ind file belonging to the same population (e.g., 10 Sardinians in a row), you can safely copy the corresponding 10 lines from the .fam file without having to check each ID manually. PLINK\u0026rsquo;s --keep accepts either the full .fam row (all 6 columns) or just the first two columns (FID and IID).\nAfter selecting a subset of samples of interest, for example, several European populations, and saving them to a file (in this case named samples, without a file extension), we can use PLINK to generate a new dataset:\nplink --bfile data --keep samples --make-bed --out subset This will create a new binary PLINK dataset (subset.bed, .bim, .fam) containing only the selected individuals.\nRunning The PCA Now we can run the PCA on the subset:\n# --out specifies the output file prefix # The default number of principal components is 10, # so \u0026#34;10\u0026#34; could be omitted unless you want to change the number plink --bfile subset --pca 10 --out pca This will output two files: pca.eigenval and pca.eigenvec. The pca.eigenvec is what you’ll use for plotting or downstream analysis.\nTo keep only the IIDs and the principal components (and discard FIDs), you can use the following awk one-liner on Linux:\nawk \u0026#39;{$1=\u0026#34;\u0026#34;; sub(/^ /, \u0026#34;\u0026#34;); gsub(/ /, \u0026#34;,\u0026#34;); print}\u0026#39; pca.eigenvec \u0026gt; output.csv This will generate a CSV-output file where each line contains a sample ID followed by its principal component values.\nThe resulting file can be loaded into Python or R for downstream analysis. It’s also compatible with hobbyist tools like Vahaduo.\nPlotting the PCA The previously created CSV-file can be used to make a PCA-plot. However, since the file contains only sample IDs and principal component values, without population labels, we’ll first generate a separate label file to make the plot easier to interpret. To do this, we’ll match each sample in the .fam file to its corresponding population label from the original .ind file. The .ind file contains the population in the third column. We’ll extract that and output it in the same order as the .fam file, so the labels line up with the PCA values.\nHere’s a one-liner using awk:\nawk \u0026#39;NR==FNR {map[$1]=$3; next} {print map[$2]}\u0026#39; v62.0_HO_public.ind subset.fam \u0026gt; labels This will create a file called labels, with one population label per line, matching the sample order in your PCA output.\nSetting Up Python If you haven\u0026rsquo;t installed Python yet, here\u0026rsquo;s how to do it:\nOn Linux (Ubuntu/Debian):\nsudo apt update sudo apt install python3 python3-pip python3-venv Now set up up a virtual environment and install dependencies:\npython3 -m venv venv # Create venv named venv source venv/bin/activate # On Windows: venv\\Scripts\\activate pip install pandas matplotlib seaborn # Install libraries Creating the Plotting Script Below is a Python script that reads the eigenvector file (output.csv) and the corresponding label file (labels), and generates a PCA plot using pandas, seaborn, and matplotlib.\nSave the following code as pca_plot.py:\nimport pandas as pd import seaborn as sns import matplotlib.pyplot as plt # Load eigenvectors df = pd.read_csv(\u0026#34;output.csv\u0026#34;, header=None) df.rename(columns={0: \u0026#34;ID\u0026#34;}, inplace=True) # Load labels with open(\u0026#34;labels\u0026#34;) as f: labels = [line.strip() for line in f] if len(labels) != df.shape[0]: raise ValueError(\u0026#34;Mismatch between number of labels and samples\u0026#34;) df[\u0026#34;Label\u0026#34;] = labels # Rename eigenvector columns for i in range(1, df.shape[1] - 1): df.rename(columns={i: f\u0026#34;PC{i}\u0026#34;}, inplace=True) # Seaborn aesthetics # Alternatively remove this line for no grid sns.set(style=\u0026#34;whitegrid\u0026#34;, context=\u0026#34;talk\u0026#34;) fig, ax = plt.subplots(figsize=(12, 7)) # Plot sns.scatterplot( data=df, x=\u0026#34;PC1\u0026#34;, y=\u0026#34;PC2\u0026#34;, hue=\u0026#34;Label\u0026#34;, palette=\u0026#34;deep\u0026#34;, s=50, edgecolor=\u0026#34;black\u0026#34;, ax=ax ) # Title and labels ax.set_title(\u0026#34;PCA\u0026#34;, fontsize=18) ax.set_xlabel(\u0026#34;PC1\u0026#34;) ax.set_ylabel(\u0026#34;PC2\u0026#34;) # Legend outside the plot to the right ax.legend( title=\u0026#34;Population\u0026#34;, bbox_to_anchor=(1.02, 1), loc=\u0026#34;upper left\u0026#34;, borderaxespad=0, frameon=True ) # Adjust layout to make room for legend plt.tight_layout(rect=[0, 0, 0.85, 1]) plt.show() # To save instead of showing, comment out plt.show() and use: # plt.savefig(\u0026#34;pca_plot.png\u0026#34;, dpi=300) Finally run it with:\npython3 pca_plot.py Result The PCA plot below shows the structure observed in my test subset: ","permalink":"https://popgenblog.com/posts/plink-pca-tutorial/","summary":"In this post, I’ll demonstrate how to perform a PCA on a PLINK dataset. Before we begin, we need to prepare a subset of samples we\u0026rsquo;re interested in analyzing.\nTo do this, we’ll extract sample in","tags":["PLINK","PCA","awk"],"title":"PLINK PCA Tutorial: Running PCA in PLINK (Commands + Output)"},{"content":"The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style .geno/.snp/.ind dataset. This naming can be confusing: the .snp and .ind files are the usual EIGENSTRAT metadata files, but the .geno file may either be plain-text EIGENSTRAT or binary PACKEDANCESTRYMAP.\nPACKEDPED format allows for easier downstream processing using the PLINK toolset. With PLINK, it becomes straightforward to extract sample subsets, filter SNPs, and perform a wide range of analyses.\nDownloading PLINK I use PLINK 1.9. While there is a newer version (2.0), I prefer 1.9 because it includes several features that were deprecated or removed in the newer release.\nChoose and download the binary suitable for your operating system.\nIf you\u0026rsquo;re on Linux or WSL, you can make the binary globally accessible like this:\n# Make PLINK globally accessible # Run this from the directory where you downloaded the binary sudo cp plink /usr/local/bin/ # Test if PLINK works plink Installing EIGENSOFT Option 1: Install via conda If you have conda installed, this is the simplest method:\nconda install -c bioconda eigensoft If you don\u0026rsquo;t have conda yet:\n# Download Miniconda installer # For other architectures (ARM, macOS, etc.), see: https://docs.conda.io/en/latest/miniconda.html # Run installer bash Miniconda3-latest-Linux-x86_64.sh # Follow prompts, then restart your terminal or run: source ~/.bashrc Option 2: Compile from source If you prefer to compile manually, first install the required dependencies:\nsudo apt update sudo apt install -y build-essential gfortran liblapack-dev liblapacke-dev libgsl-dev libopenblas-dev Then clone the repository and compile:\ngit clone https://github.com/DReichLab/EIG cd EIG/src LDLIBS=\u0026#34;-llapacke\u0026#34; make This will generate several necessary binaries. The most relevant ones are: convertf and mergeit in the src folder, and smartpca in the eigensrc subdirectory. You can confirm the files were compiled with:\nls Make tools globally accessible To use convertf and mergeit from anywhere, move them to a global location like /usr/local/bin/:\nsudo cp convertf mergeit /usr/local/bin/ Accessing smartpca If the build was successful, you’ll find smartpca in the eigensrc directory:\n# Navigate into the eigensrc folder cd ~/EIG/src/eigensrc # Check files in folder ls Make it globally accessible:\nsudo cp smartpca /usr/local/bin/ I will provide examples on using smartpca for principal component analysis (PCA) in another post. Since it was compiled alongside the other EIGENSOFT tools, I’m just mentioning it here for completeness.\nConverting to PACKEDPED Now we can convert the dataset, whether its .geno file is plain-text EIGENSTRAT or PACKEDANCESTRYMAP, to PACKEDPED. Navigate into the folder or directory you store the dataset. Then with a text editor of your choice generate a file with the following content (adjust the input prefixes to those of your dataset):\ngenotypename: v62.0_HO_public.geno snpname: v62.0_HO_public.snp indivname: v62.0_HO_public.ind outputformat: PACKEDPED genotypeoutname: data.bed snpoutname: data.bim indivoutname: data.fam Save the file with any name you like, I named mine simply parameter (no file extension). This file tells convertf what to do. Start the conversion with:\nconvertf -p parameter The process may take a while, as the dataset is large.\nOnce complete, you’ll have a set of PLINK-compatible binary files: data.bed, data.bim, and data.fam, ready for downstream processing.\n","permalink":"https://popgenblog.com/posts/convert-eigenstrat-to-packedped/","summary":"The files downloaded in the previous blog post are distributed as an EIGENSTRAT-style .geno/.snp/.ind dataset. This naming can be confusing: the .snp and .ind files are the usual EIGENSTRAT metadata f","tags":["EIGENSTRAT","PACKEDPED","PLINK"],"title":"Converting EIGENSTRAT/PACKEDANCESTRYMAP to PACKEDPED"},{"content":" Note: This post uses an older AADR release and parts of it may now be outdated. For the latest AADR v66 download, including TGENO conversion and ADMIXTOOLS2 compatibility notes, see Downloading and Converting AADR v66.\nA Linux environment is unavoidable when it comes to bioinformatical data processing and preparation. You can use your favorite distribution.\nFor Windows users, the Windows Subsystem for Linux (WSL) provides a good alternative to dual booting or setting up a full virtual machine.\nInstalling WSL with Debian Open PowerShell as Administrator and run:\nwsl --install -d Debian Once installed, update the system:\nsudo apt update \u0026amp;\u0026amp; sudo apt upgrade -y Downloading A Genetic Dataset Before doing PCA, ADMIXTURE, qpAdm, etc, you need actual genotype data. A good and comprehensive resource is the Allen Ancient DNA Resource (AADR).\nIf you\u0026rsquo;re only interested in ancient samples, download the following three files from the AADR:\nv62.0_1240k_public.geno, .snp, and .ind.\n(If you want information on sample origins, you should also download the corresponding .anno file.)\nIf you\u0026rsquo;d like to include modern samples as well, which can be useful for personal genetic comparisons, download the same file types, but with the prefix v62.0_HO_public.\nSince these files are large, it\u0026rsquo;s best to download them using wget (over Linux) from the direct download links. This avoids browser interruptions.\nExample using wget: # Install wget sudo apt install wget # Download the files # Example for v62 wget -O v62.0_HO_public.geno \u0026#34;https://dataverse.harvard.edu/api/access/datafile/10537419\u0026#34; wget -O v62.0_HO_public.snp \u0026#34;https://dataverse.harvard.edu/api/access/datafile/10537421\u0026#34; wget -O v62.0_HO_public.ind \u0026#34;https://dataverse.harvard.edu/api/access/datafile/10537420\u0026#34; ","permalink":"https://popgenblog.com/posts/download-ancient-modern-dna-aadr/","summary":" Note: This post uses an older AADR release and parts of it may now be outdated. For the latest AADR v66 download, including TGENO conversion and ADMIXTOOLS2 compatibility notes, see Downloading and C","tags":[],"title":"How to Download the AADR Dataset (Linux \u0026 WSL)"},{"content":"You can reach me at contact [at] popgenetics [dot] dev\n(Just replace [at] with @ and [dot] with .)\n","permalink":"https://popgenblog.com/contact/","summary":"You can reach me at contact [at] popgenetics [dot] dev\n(Just replace [at] with @ and [dot] with .)\n","tags":[],"title":"Contact"},{"content":"Use a hosted analysis when you want a quick result, or use the toolkit when you want to keep the workflow on your own machine.\n01 HOSTED ANALYSIS Genostruct Explore 23andMe, AncestryDNA, FTDNA, and MyHeritage raw-data exports with free calculators and paid ancestry analyses. Paid reports include free previews.\nNo account required · Files are not retained\nOpen Genostruct ↗ Ancestry analyses → 02 DIY TOOLKIT Raw DNA to AADR Toolkit Convert commercial raw-data exports to PACKEDPED and PACKEDANCESTRYMAP, with optional AADR merging for ADMIXTOOLS, PCA, and ADMIXTURE.\n$12 · Includes workflow documentation\nGet the toolkit ↗ Read the usage guide → ","permalink":"https://popgenblog.com/products/","summary":"Use a hosted analysis when you want a quick result, or use the toolkit when you want to keep the workflow on your own machine.\n01 HOSTED ANALYSIS Genostruct Explore 23andMe, AncestryDNA, FTDNA, and My","tags":[],"title":"Products \u0026 tools"}]