Methodological foundations
Methods for individual traits are named systematically after their sources. The following descriptors indicate how they are used: Exploratory admits a broader but less certain evidence base; Strict uses the same sources and model with tighter quality gates; Minimal denotes a deliberately small comparison model; and PGS denotes a published polygenic score. All direct models use observed AADR markers without imputation. Only the Deep SNP scan explicitly evaluates positions outside the AADR panel.
Complex traits such as hair colour and skin pigmentation are analysed primarily through genome-wide association studies (GWAS). The relevant variants are usually single-nucleotide polymorphisms (SNPs): differences at individual DNA positions, each defined by one of the four bases A, T, C or G. Larger genomic structures can also affect a trait, but SNP effects are the principal basis used here and can be traced most transparently.
Genes and SNPs differ greatly in their association with a trait. In the strongly skewed effect-size distribution described here, the 5–20 largest-effect SNPs can carry more weight than thousands of individually weak associations combined. These direct models therefore evaluate a deliberately selected set of the strongest known signals rather than thousands of small effects, as a PGS does. Some traits, including height, are so polygenic that the weaker effects still make a material contribution; for those traits, PGS models are included as a benchmark.
The general procedure for these direct SNP models was as follows:
The literature was reviewed for relevant genetic markers, and the reported effect sizes were extracted.
Each marker was then checked for coverage in AADR. Markers absent from the panel were documented and re-evaluated against the raw sequencing data in the Deep SNP scan.
Among the markers covered by AADR, the most robust, well-replicated and large-effect candidates were selected and combined into a transparent, trait-specific evidence index. The calculation uses published effect sizes from sources with the greatest available transferability; markers whose effect direction or weight could not be harmonised reliably were excluded.
Unobserved loci are not treated as counter-evidence. Evidence that is too weak, non-independent or internally inconsistent is blocked so that it cannot distort the result.
Direct models, minimal models, and PGS
Genome-wide association studies (GWAS) and polygenic scores (PGS) are closely related but serve different purposes. A GWAS compares the genomes and measured traits of many modern participants to identify associated variants and estimate their effects. A PGS then combines effect estimates from such association studies into a model that assigns an aggregate score to an individual.
The Exploratory and Strict direct models use effect sizes derived in the same way. In a PGS method, by contrast, all SNPs and weights come from a jointly calibrated score. This improves internal consistency, but also introduces many weak markers whose calibration in modern populations may not transfer readily to populations that lived thousands of years ago. The source-based direct models focus on a small number of strong, transparently traceable markers; PGS models use a much broader evidence base that is more sensitive to temporal and population differences.
Lona-Durazo/Sulem (Minimal) and Morgan 2018 (Minimal) are deliberately not described as PGS models: both are small, GWAS-based direct-marker comparison models. The genuinely polygenic models are identified explicitly as Tanigawa 2022 (PGS), Privé 2022 (PGS) and Yengo 2022 (PGS).
