Debbie Kennett 🧬🌳

@debbiekennett.bsky.social

DNA, genetic genealogy, family history. forensics. Honorary Research Fellow at UCL. All views my own.

Nevertheless, our results are important because they apply to genealogical data *already accessible* to law enforcement in many countries, without relying on commercial providers. Our results should inform regulators and practitioners of FIGG, as well as quantify biobank privacy risks. 10/10

Key limitations: * Finding a match does not guarantee identification * Many relatives missed because of the limited time depth of the register * Data errors (clerical, adoption, non-paternity, gamete donation, relationship misclassification, MZ twins) * Generalizability to other countries 9/10

What is the probability of having *two* (or more) relatives in the database? This is important, as it is often the intersection of the family trees of multiple matches that leads to identification. The match rate naturally decreases, although not as much for the very large databases. 8/10

The probability of detecting multiple relatives of the target in the database. (A) The match rate when requiring at least two relatives in the database. All parameters are as in Figure 2. (B) The distribution of the number of relatives in the database. We only considered target individuals (alive or deceased) who have at least one relative in the database. The database size was 1% of the population (alive and >18) and we considered all 7th-degree or closer relatives. Target individuals already in the database count themselves towards the total number of relatives that they have there.

Unexpectedly, the match rate barely increases beyond 5th-degree (2nd cousins). This is probably because the common ancestors of third cousins usually lived before the country's foundation. Implication: whenever relatives are found in the database, they are usually not too distant. 7/10

Bild

Results: The match rate is the proportion of people with at least one relative (a "match") in the database. For a database covering 1% of the country, the match rate is ~25%. For a database covering 10% of the country, the rate is 68%. 6/10

The FIGG match rate using a national population register. We plot the proportion of the Israeli population (including deceased) who have at least one relative in the genomic database to whom they connect via a path in the population register of a maximal pre-defined length. The x-axis is the database size, in proportion of the living population over 18 years old. Each curve (legend) corresponds to a different maximal degree of relationship. For example, a first-degree relative is a parent or a full sibling, and a second-degree relative can be a half-sibling, a grandparent, or an aunt. Each data point is the average over 100 simulations. Standard deviations across simulations are shown as error bars, but are too small to be visible.

FIGG simulation: We designated a random subset of the population as the "database". For each person in the database, we used the register to track all of their relatives up to 7th degree (eg, 3rd cousins). This gave us the list of people who would have one or more relatives in the database. 5/10

Bild

Data: We used the entire Israeli national population register, documenting all parent-child relationships in all present and past citizens. After QC, the dataset covered 12.5 million people, among them 10.4 million alive. Summary stats look reasonable. 4/10

BildCharacteristics of the IPIA data. (A) The distribution of birth decades (1870s-2010s). (B) The average number of children per individual (including childless individuals or individuals who have not yet completed their families) vs the parental birth decade. Birth decades start at the indicated year (e.g., the 1960 birth decade corresponds to birth years 1960-1969). (C) The distribution of the number of children per individual. We combined all individuals with 15 or more children into a single category (see also Figure S1). The y-axis is logarithmic. (D) The distribution of the paternal (light blue) and maternal age (pink) at childbirth, also known as the generation interval. The distribution is over all births in the database. Numbers on the x-axis indicate the midpoints of 3-year bins (e.g., 25 corresponds to ages 24-26).

Our research question: Given a genomic database covering x% of the population of a country, what is the probability that a target person has one or more relatives in the database? Particularly, we focus on relatives from whom the target can be traced using available genealogical data. 3/10

In forensic investigative genetic genealogy (FIGG), law enforcement people upload the DNA of an unknown person (a criminal or unidentified remains) to a consumer genomics database (eg FTDNA), find relatives, and intersect their family trees until reaching an identification. 2/10

A schematic of FIGG. For a subset of the population, a genome-wide DNA profile is available in a “genomic database” (light blue box). The identity of a target person (peach box) is unknown, but we have access to their genome-wide genetic profile. By comparing the DNA profile of the target against that of all individuals in the database, genetic relatives of the target (“matches”; red arrows) can be identified (light green box inside the database). Once matches have been detected, the investigation focuses on their relatives (green and red arrows), given that the target person must be one of them. Thus, to be identifiable using FIGG, a target person must have at least one relative in the database. We require the relationship to be equal or closer to a prespecified degree, corresponding to the most distant relationship that can be confidently detected using genetic data (usually 7th degree, e.g., a third cousin). In practice, more distant relationships can also sometimes be detected.

New preprint! Investigative genetic genealogy has revolutionized forensic identification, with hundreds of cases solved to date. But for any given new case, how likely is genetic genealogy to succeed? We used a whole-country genealogy to find out! 1/10 www.biorxiv.org/content/10.6...

Forensic investigative genetic genealogy match rate estimated from a nation-wide population register

Forensic investigative genetic genealogy (FIGG) is a revolutionary method in forensic genetics, whereby genetic relatives of an unknown target person are detected in direct-to-consumer genomic databas...

biorxiv.org