Pixy: Unbiased estimation of nucleotide diversity and divergence in the presence of missing data.
Ontology highlight
ABSTRACT: Population genetic analyses often use summary statistics to describe patterns of genetic variation and provide insight into evolutionary processes. Among the most fundamental of these summary statistics are π and dXY , which are used to describe genetic diversity within and between populations, respectively. Here, we address a widespread issue in π and dXY calculation: systematic bias generated by missing data of various types. Many popular methods for calculating π and dXY operate on data encoded in the variant call format (VCF), which condenses genetic data by omitting invariant sites. When calculating π and dXY using a VCF, it is often implicitly assumed that missing genotypes (including those at sites not represented in the VCF) are homozygou
SUBMITTER: Korunes KL
PROVIDER: S-EPMC8044049 | biostudies-literature | 2021 May
REPOSITORIES: biostudies-literature
ACCESS DATA