AI Ethics

AI Genetic Genealogy: 60% Match Rate

AI genetic genealogy illustration of DNA strand connecting family tree branches
The AI Family Tree

By Stuart Kerr, Technology Correspondent, LiveAIWire

AI genetic genealogy is producing capabilities in family reconstruction and trait prediction that were science fiction a decade ago. Direct-to-consumer genetic testing companies have enrolled more than 40 million people in databases that, when combined with AI-driven analysis, can reconstruct family trees across multiple generations, identify previously unknown relatives, and in some cases resolve questions of biological parentage that individuals had no prior reason to suspect. The technology has reunited adoptees with biological families, solved cold-case crimes through distant relative matching, and enabled population-level research into genetic health risks.

It has also created a new class of privacy risk, enabled predictive inferences about individuals who never consented to testing, and produced a commercial ecosystem operating with regulatory frameworks that have not kept pace with its capabilities. Research published in Science demonstrated that genomic data from as few as 1.28 million individuals, representing roughly 1 per cent of a population of European ancestry, is sufficient to identify a third cousin or closer match for about 60 per cent of individuals in that population, effectively making most people within reach of the database identifiable even if they have never submitted their own DNA for testing.

What AI Genetic Genealogy Can Now Predict

The predictive reach of AI-driven genetic analysis extends significantly beyond ancestral matching. Polygenic risk scores, which aggregate the small effects of many genetic variants to estimate an individual’s relative risk of conditions including heart disease, type 2 diabetes, and certain cancers, are now generated routinely by consumer testing services and used in clinical settings with increasing frequency. The predictive power is genuine but bounded: polygenic risk scores capture genetic predisposition, not destiny, and their accuracy varies substantially across populations depending on the diversity of the training data from which they are derived.

That last point has significant equity implications. Most large genomic datasets have been assembled primarily from populations of European ancestry, with the result that polygenic risk score predictions are substantially more accurate for European-ancestry individuals than for those of African, South Asian, or other non-European ancestries. NIH-affiliated research on polygenic risk scores explicitly acknowledges this limitation, noting that applying risk scores developed on European-ancestry populations to other groups produces predictions of unknown and potentially misleading accuracy.

The Consent and Surveillance Problem

The most significant governance challenge in AI genetic genealogy is that the system generates privacy consequences for people who have not consented to it. When an individual submits DNA to a consumer testing service, they are not only sharing their own genomic information. They are sharing information about all their biological relatives, including relatives who have not consented to testing and may not know a database exists.

The question of who controls genetic information, and under what conditions it can be used for purposes beyond those for which it was collected, is not resolved by any current regulatory framework with sufficient clarity. As LiveAIWire’s analysis of how AI data systems create governance gaps found, the populations most affected by AI-generated data inferences are frequently those who had no meaningful opportunity to evaluate or contest the collection. In AI genetic genealogy, that category includes virtually every biological relative of anyone who has ever submitted a DNA sample.

What Responsible Governance Requires

The governance requirements for AI genetic genealogy are more demanding than for most AI applications because the data involved is immutable, highly sensitive, and consequential for biological relatives who are not party to the original consent transaction. Minimum requirements include clear limitations on secondary uses of genetic data beyond those explicitly consented to at collection, prohibition on law enforcement access to consumer genetic databases without judicial authorisation equivalent to that required for other forms of surveillance, and equity requirements for genetic AI tools used in clinical settings.

The technology’s capabilities are outpacing both public understanding and regulatory frameworks in ways that create risk accumulation. As LiveAIWire’s coverage of AI and moral consequence in high-stakes applications found, the systems with the greatest potential benefit often carry the greatest potential for harm when deployed without adequate governance, and AI genetic genealogy sits at an extreme on both dimensions.

The Equity Dimension

The equity implications of AI genetic genealogy extend beyond the consent problem to the unequal distribution of benefit and risk across populations. The populations with the largest genomic databases, primarily those of European ancestry enrolled in consumer testing services in high-income countries, benefit most from AI-driven genealogical reconstruction and receive the most accurate polygenic risk score predictions. The populations with the smallest database representation receive less accurate predictions and are simultaneously subject to law enforcement use of genomic databases in ways that have documented racial disparities in impact.

Several initiatives including the H3Africa consortium and the All of Us Research Program in the United States are building the data infrastructure that equity requires. As LiveAIWire’s coverage of how AI benefits and risks fall unevenly across populations found, the distributional effects of AI applications in high-stakes domains are not accidental. They are the predictable consequence of investment patterns that can be changed when the political will and institutional capacity to change them exist.

What You Should Know Before Testing

If you are considering direct-to-consumer genetic testing, understanding what you are consenting to beyond the immediate service is important. Most consumer DNA testing services include terms that allow the company to use your genetic data for research and product development, with varying degrees of opt-out provision. The relatives matched to your profile in their database have not all explicitly consented to being matched with you, because their data was submitted by relatives on your side of the match.

Practical protective steps include reading the terms governing data use before submission, using pseudonymous contact details if the service allows it, understanding the opt-out provisions for research use, and being aware that deleting your account may not result in deletion of your genetic data from all of the company’s databases and research partnerships. These limitations do not make genetic testing inadvisable, but they do make informed consent more complicated than the marketing of consumer testing services suggests.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.