By Stuart Kerr, Technology Correspondent, LiveAIWire
AI Embryo Selection has already reached a milestone that should make every prospective parent look twice: the first major randomised, double-blind trial of a deep-learning embryo-ranking system failed to prove it was noninferior to standard embryologist selection. The study involved 1,066 patients across 14 IVF clinics in Australia and Europe, making it one of the strongest real-world tests yet of whether an algorithm can help decide which embryo should be transferred first.
That result matters because two very different technologies are now being discussed under the same label. One form of AI embryo selection analyses embryo images and time-lapse videos to rank developmental potential. The other uses genomic data and polygenic scores to estimate future disease risk, and potentially nonmedical traits. The first is an attempt to make a familiar IVF decision more consistent. The second changes what the decision itself can mean.
Table of Contents
What AI Embryo Selection Actually Means
In a conventional IVF laboratory, embryologists examine developing embryos and judge which appear most likely to implant successfully. They look at morphology, developmental timing and other features that experienced specialists learn to interpret. The process is skilled, but it is also partly subjective. Two embryologists can look at the same embryo and disagree about its grade, and the same person may not grade an image identically on another day.
Image-based AI embryo selection tries to reduce that subjectivity. A model can analyse still images or sequences from a time-lapse incubator, detect patterns across many visual features and assign each embryo a score. In principle, this can make ranking faster and more reproducible while giving the embryologist an additional source of evidence rather than replacing them.
What This Means for You Before a Clinic Offers It
If you are going through IVF, the most important question about AI embryo selection is not simply whether a clinic uses AI. Ask what the system is being asked to do. An image-ranking tool that helps an embryologist decide which viable embryo to transfer first is fundamentally different from a genomic service that assigns embryos statistical risk scores for conditions that may not appear for decades.
You should also ask whether the exact system has been tested prospectively, whether it was validated at centres similar to yours and whether the clinician can override its recommendation. LiveAIWire has found the same distinction in the wider AI doctor dilemma and clinical oversight: a narrow, validated tool used as decision support deserves a different level of trust from a model whose output quietly becomes the decision.
The First Randomised Trial Delivered a Reality Check
The strongest evidence for image-based AI embryo selection comes from a 2024 trial published in Nature Medicine. Researchers compared a deep-learning system called iDAScore with standard morphology-based selection. The trial enrolled 1,066 women under 42 who had at least two early-stage blastocysts on day five and randomly assigned embryo selection to either the algorithm or the conventional approach.
The result was not the clean victory that an AI headline might suggest. Clinical pregnancy occurred in 46.5 percent of the AI group and 48.2 percent of the morphology group. The trial had been designed to show that the algorithm was not more than five percentage points worse than standard selection, but the confidence interval crossed that predefined margin. The researchers therefore reported that the study did not demonstrate noninferiority.
That does not mean the algorithm failed or that AI has no role in IVF. It means something more useful: once a promising model was tested in a rigorous randomised clinical setting, its apparent technical sophistication did not translate into proven superiority, or even statistically demonstrated noninferiority under the trial’s chosen threshold. That is exactly why trial evidence matters.
Why Image-Based AI Embryo Selection Still Matters
The trial does not erase the practical advantages of automated assessment. Time-lapse systems can observe development continuously without repeatedly removing embryos from an incubator. Algorithms can process more visual information than a human can practically track and can apply the same scoring logic every time. Even if pregnancy rates are not higher, greater consistency and reduced laboratory workload could still have value.
But those benefits need to be measured separately from the claim that the algorithm chooses a better embryo. The history of medical AI is full of systems that perform impressively on retrospective datasets and then look much more ordinary when tested prospectively. LiveAIWire’s analysis of AI cancer diagnosis evidence shows why randomised trials, external validation and patient outcomes matter more than a headline accuracy score.
The Reliability Problem Is More Uncomfortable Than Accuracy
A 2025 study raised a different concern about AI embryo selection. Researchers trained 50 replicate versions of the same type of image model, changing the random initialisation while keeping the basic architecture and task the same. They then compared how consistently those models ranked embryos belonging to the same patients.
The agreement was poor. Across one dataset the average Kendall’s W concordance score was about 0.36, where 1 represents complete agreement. On a second dataset from a different fertility centre it was about 0.34. More strikingly, the models sometimes ranked degenerate or arrested embryos above viable blastocysts even when a better embryo was available. Average critical error rates were 12.4 percent in one dataset and 17.3 percent in the other.
The researchers described substantial instability and inconsistency in the models they tested. This was a research stress test, not evidence that every commercial IVF algorithm makes these mistakes. Its importance is narrower and more profound: two models can achieve similar headline accuracy while producing different rankings for the actual decision that matters.
Genetic Embryo Screening Is a Different Revolution
Image-based AI embryo selection asks which embryo appears most likely to develop successfully. Polygenic embryo screening asks a different question: which embryo has the most favourable predicted genetic risk profile?
In genetic AI embryo selection, polygenic risk scores combine the estimated effects of many genetic variants associated with a disease or trait. Instead of looking for a single high-impact mutation, the model aggregates many small statistical associations. That can produce a relative risk estimate for conditions such as cardiovascular disease, diabetes or some cancers. It is not a diagnosis and it cannot tell parents that an embryo will or will not develop a particular condition.
This distinction is crucial because a score that works reasonably well for separating risk across a large adult population may be much less decisive when comparing a small number of embryos from the same two biological parents. Sibling embryos share much of their genetic background. The practical choice is not between the highest-risk person and lowest-risk person in a country. It may be between three embryos whose predicted absolute risks are only modestly different.
The Medical Establishment Says the Evidence Has Been Outrun
For genetic AI embryo selection, the American College of Medical Genetics and Genomics has taken an unusually clear position. In its statement on polygenic risk scores for embryo selection, ACMG concluded that clinical utility had not been proven and said the practice had moved too fast with too little evidence. It also warned that no standards comparable with adult polygenic risk testing exist for prenatal or embryo use.
The ACMG statement explains why this is not just a philosophical objection. Polygenic scores are probabilistic, environmental factors matter, genomic datasets remain uneven across ancestries, embryo biopsies involve their own technical limitations and a low score can create false reassurance because it does not capture every genetic or environmental route to disease.
In 2026, the American Society for Reproductive Medicine went further. Its Ethics and Practice Committees said PGT-P remains a nascent and unproven technology, is not recommended for clinical use and should not currently be offered as a clinical service. ASRM also said polygenic embryo testing should not be used for nonmedical trait selection, including traits such as height, intelligence or eye colour.
Disease Prevention and Trait Selection Are Not the Same Argument
The ethical debate becomes distorted when every use of genetic embryo screening is described as a quest for a designer baby. A prospective parent with a strong family history of a serious disease may be motivated by something much more ordinary: reducing the chance that their child experiences the same illness. That motivation deserves to be taken seriously even when the test itself remains uncertain.
Trait selection raises a different problem. Height, cognitive ability, appearance and behaviour are influenced by complex mixtures of genetics, development, environment and chance. Treating a polygenic prediction as a promise risks turning a probabilistic association into an expectation placed on a child before birth.
ASRM specifically warns that parental decisions based on such scores could affect how a child is raised and restrict what it calls the child’s opportunity for an open future. That connects with LiveAIWire’s broader reporting on children and AI developmental questions, where the central issue is not only what an algorithm predicts, but how adults change their behaviour because they believe the prediction.
The Ancestry Problem Cannot Be Fixed With Better Marketing
Polygenic scores depend on genome-wide association studies, and those databases have historically been weighted toward people of European ancestry. ACMG notes that predictive performance can fall when a score is applied to populations different from those used to build it. Environmental and socioeconomic differences can also change how genetic associations translate into real-world disease.
For AI embryo selection based on genetics, that creates an equity problem before the ethical argument even begins. A service may look equally precise on a glossy report while carrying different levels of uncertainty for different families. Unless a provider can show how the score performs for the ancestry and population relevant to the prospective parents, the number risks appearing more universal than the evidence supports.
More Data Does Not Remove the Trade-Offs
A parent may imagine that screening for more conditions simply produces a clearly superior embryo. The genetics does not cooperate with that intuition. The embryo with the lowest predicted risk for one disease may not have the lowest predicted risk for another. Some genetic variants also influence more than one trait, a phenomenon known as pleiotropy, so selecting in one direction can have consequences elsewhere that are difficult to predict.
There is also a practical limit. Many IVF cycles do not produce a large pool of transferable embryos. If only one or two viable embryos are available, a ranking system has little room to optimise anything. Pursuing additional IVF cycles solely to obtain more embryos for statistical selection adds cost, physical burden and medical risk. A technology designed to offer more choice can therefore create pressure to manufacture more choices before it becomes useful.
The Real Risk Is Automation Authority
The deepest connection between image ranking and genetic screening is not the data they use. It is the authority people may grant the output. A score with two decimal places looks objective. A ranked list looks decisive. Once an algorithm places embryo A above embryo B, declining its recommendation can feel irrational even when the difference is small or the model is poorly validated.
That is the same automation problem LiveAIWire has examined in high-stakes algorithmic decisions outside medicine. The danger is not that machines are incapable of useful prediction. It is that a prediction can quietly acquire more authority than its evidence deserves, especially when the person receiving it cannot inspect the model, the training data or the uncertainty behind the score.
Seven Questions to Ask Before Letting an Algorithm Rank Embryos
Before accepting AI embryo selection based on image ranking, ask what outcome the model predicts, whether it was tested prospectively, how many patients were included, whether validation occurred outside the developer’s own clinic, how often clinicians override its ranking, whether your clinic has measured its local performance and what happens when the algorithm disagrees with an experienced embryologist.
For polygenic screening, ask for absolute risk rather than only relative risk, how much the embryos actually differ from one another, how the score performs for your ancestry, whether a genetic counsellor independent of the testing company will interpret the result, what evidence connects the score to health outcomes in children born after selection and whether the clinic would make the same recommendation without the algorithmic ranking.
If those questions produce vague answers, the problem is not that you failed to understand the technology. It may be that the certainty suggested by the interface is greater than the certainty supported by the evidence.
AI Embryo Selection Is Not One Question With One Answer
The phrase AI embryo selection makes it sound as though a single technology has arrived to choose the future child. In reality, image-based embryo ranking and polygenic embryo screening sit at very different points on the evidence curve.
Image ranking has reached the stage where it can be tested in proper randomised trials, and the first major trial produced a sobering result rather than a triumph. That is a sign of scientific progress, not failure. It tells clinics and patients what still needs to be proven.
Polygenic screening reaches much further, from embryo viability into predictions about decades of future health and potentially nonmedical traits. Yet the professional bodies responsible for reproductive medicine and medical genetics say its clinical utility remains unproven. The more consequential the promise becomes, the more important it is to ask whether the evidence has kept pace.
Would you let AI choose your embryo? The evidence suggests a better question: exactly what is the algorithm choosing for, how was that claim tested, and who remains responsible when the score is wrong?
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.
