Methodological foundations
The methods for the individual traits are systematically named after their sources. The following descriptors are also used: Exploratory admits more evaluable but less certain evidence; Strict uses the same source basis with tighter quality release; Minimal denotes deliberately small comparison models; PGS denotes a published polygenic score. All direct models use observed AADR markers without imputation; only the Deep SNP scan explicitly evaluates markers that are absent from the AADR panel.
For complex traits (hair colour, skin pigmentation and so on) GWAS (“genome-wide association studies”) are used above all. In doing so we observe SNPs (“single nucleotide polymorphisms”), the smallest possible coding genetic units, which are always defined exactly by the four nucleic bases A, T, C, G. Larger gene structures also have an influence, but SNPs make the main contribution and are the easiest and most unambiguous to trace.
Among all the larger genes and SNPs there are in each case some that have a stronger influence and some that are only very weakly related to the trait. Following the power law, the 5-20 SNPs with the strongest effects often carry more weight than all the other thousands of weaker-effect SNPs together, each of which pushes towards the trait only marginally. In contrast to PGS, it is therefore not thousands of weak signals that are evaluated here, but specifically the strongest known signals. Some characteristics such as height are, however, so complex that the weaker signals still make a relevant contribution. In these cases PGS models are used as a benchmark.
The general procedure for these direct SNP models was as follows:
Es wurden aktuelle Publikationen nach relevanten genetischen Markern durchsucht und deren Effektstärke ermittelt.
Es wurde geprüft, inwieweit die AADR-Datengrundlage diese Marker tatsächlich abbildet. Diejenigen die nicht abgebildet waren, wurden dokumentiert und in der SNP-Tiefensuche nochmals an den Rohdaten berechnet.
Aus den in AADR vorhandenen Markern wurden die belastbarsten, gut reproduzierten und effektstarken Kandidaten herausgegriffen und je nach Merkmal zu einem transparenten Evidenzindex verrechnet. Bei der Berechnung wurde auf publizierte Effektstärken aus möglichst übertragbaren Quellen zurückgegriffen; Marker ohne sicher harmonisierbare Richtung oder Gewichtung wurden ausgeschlossen.
Nicht beobachtete Loci werden nicht als Gegenbeweis gewertet. Zu schwache, nicht unabhängige oder widersprüchliche Aussagen werden blockiert, damit sie das Ergebnis nicht verzerren.
Direct models, minimal models, and PGS
“Genome Wide Association Studies” (GWAS) and “Polygenic Scores” (PGS) cannot be separated from one another. GWAS examine the genomes of thousands of modern subjects and compare them with the traits that can actually be observed. GWAS is therefore the study by which genetic markers are found and weighted. PGS is the computational model with which effect sizes for the trait are then calculated for single individuals on the basis of GWAS.
The effect sizes of the exploratory and strict direct models are produced in the same way. In the PGS methods, by contrast, the SNPs that enter the calculation each come from a jointly calibrated score source. This increases internal consistency, but it brings in many weak markers whose modern calibration cannot readily be transferred to populations thousands of years old. The source-based direct models concentrate on a few strong and transparently traceable markers; PGS work on a considerably broader basis that is, however, more sensitive in temporal and population terms.
Lona-Durazo/Sulem (Minimal) and Morgan 2018 (Minimal) are deliberately not named as PGS: both are small, GWAS-oriented direct-marker comparison models. The genuine polygenic models, by contrast, are called Tanigawa 2022 (PGS), Privé 2022 (PGS) and Yengo 2022 (PGS).
