Machine-learning models found patterns in ten-minute speed dating conversations that appeared related to whether people wanted to exchange contact details, but the models could not show a clear advantage over a random baseline once uncertainty was taken seriously. That is the intriguing and cautionary result of a new Frontiers in Computer Science study.
The researchers analysed 619 complete decision pairs from real speed-dating conversations involving 147 Japanese participants. They divided each conversation into one-minute segments and extracted facial, audio, language and interaction features, then trained separate models for decisions made by female and male participants.
Speed dating gave the AI a minute-by-minute behavioural record
Rather than treating each date as one block, the team split the ten-minute interaction into ten intervals. The feature set included facial movements, head orientation, gaze, speech acoustics, parts of speech and measures of coordination between the two people.
In total, 365 features were extracted for each one-minute segment. The researchers used Random Forest classifiers and a nested participant-level evaluation designed to keep the person being tested out of the model-selection process. That is important in social prediction because a model can otherwise learn quirks associated with particular individuals rather than general behavioural patterns.
The source data came from the Multi-Modal Speed Dating corpus developed in earlier work. That corpus contains video, audio, transcripts and post-conversation decisions from genuine speed-dating interactions. The earlier dataset was described in the INTERACT 2023 proceedings, where the researchers explored whether personal characteristics known before a date could help predict later attraction scores.
The new study is more interested in what happens during the conversation itself. That makes the work feel closer to the popular idea that subtle behaviour might reveal chemistry before either person says what they think.
The headline result is weaker than a dating predictor
The highest observed weighted F1 score for female participants’ decisions was 0.6248 when all feature types were combined. For male participants’ decisions, the best observed result was 0.5379 using visual features.
Those numbers can sound impressive without a baseline. The researchers therefore compared the models with random predictions based on the class proportions in the training data. All eight model configurations produced numerically higher weighted F1 scores than the mean random baseline, but every 95% interval for the difference included zero.
In plain English, the study did not establish that the machine-learning models were reliably better than the chosen random baseline once variation among participants was included. The researchers describe the behavioural signals as potentially informative and call for independent data to see whether the patterns hold.
That is an important distinction because dating prediction is a perfect subject for overclaiming. A model can identify a pattern in one dataset without possessing a general ability to read attraction.
Brow movement appeared in the model, but it is not a secret dating rule
Among the features that ranked highly in the best-performing models were measures related to AU04, a facial action unit corresponding to brow lowering. The importance of time also varied, with the largest aggregate share appearing around minutes six to seven for the female-decision model and seven to eight for the male-decision model.
It would be a mistake to turn that into advice such as “watch the eyebrows in minute seven”. The authors found that feature rankings changed substantially across evaluation folds. They explicitly say larger independent datasets are needed to determine whether the observed patterns are stable.
The temporal result is better read as a reminder that social interaction unfolds. The same smile, pause or head movement can mean something different early in a conversation than it does after several minutes of rapport or discomfort.
LiveAIWire has covered AI systems that score human appearance. The speed-dating work is different because it focuses on behaviour between two people rather than static attractiveness, but both areas raise the same temptation to treat a noisy human judgement as if it were a stable numerical property.
Dating is difficult for models because the target is relational
Romantic interest is not a label hidden inside one person’s face. It depends on the other person, the interaction, expectations, culture, timing and a large amount of context that a short recorded conversation may not capture.
The dataset itself shows that decisions were not simply symmetrical. Of the 619 complete pairs, the researchers reported several combinations in which one person wanted to exchange contact details and the other did not. That is a reminder that the same conversation can produce two different judgements.
This is one reason a system that predicts one participant’s decision cannot be treated as a universal measure of “chemistry”. Even if performance improves, it would be predicting an outcome in a particular setting, not discovering an objective score for a relationship.
LiveAIWire has also examined how AI is changing dating apps through synthetic profiles and scams. Behavioural prediction would add another layer, potentially allowing platforms to infer interest from interaction rather than relying only on what users explicitly select.
A weak predictor can still reveal useful science
The study is valuable precisely because the model was not presented as magically accurate. Prediction performance and scientific insight are not the same thing. A model may be too unreliable to decide whether two people should meet again while still helping researchers test which behaviours deserve closer study.
The authors used interpretable features rather than relying entirely on opaque representations. That makes it possible to ask which parts of speech, facial movements or timing patterns contributed to a prediction, then challenge whether those patterns survive in another sample.
It also exposes how easy it is to mistake a model explanation for a human explanation. A feature can help a classifier divide examples without being the psychological reason a person made a decision. Brow lowering might correlate with many other aspects of a conversation, and the model cannot tell us what the participant consciously noticed.
That distinction is especially important when AI analyses faces. LiveAIWire’s background coverage of AI face-detection research shows how much technical processing can sit between pixels and a final label. Adding social meaning on top introduces another level of uncertainty.
The honest result is more interesting than a fake romance oracle
A sensational version of this story would say AI can tell whether a date went well by watching your face. The evidence does not support that. The models saw signals, but their advantage over a simple baseline was not established clearly enough to justify that claim.
What the work does show is that machine learning can turn a conversation into a structured record of changing behaviour, then test whether those patterns relate to later decisions. With larger and more diverse datasets, that could become a useful research instrument.
The caution is that dating is exactly the kind of intimate setting where a modest statistical signal can be transformed into an overconfident product. Before behavioural prediction is used to rank people, recommend partners or tell someone how another person feels, the standard of evidence needs to be much higher than a model that looked slightly better than random in one dataset.
Dating prediction is a difficult test because the label is personal
Many machine-learning tasks have an objective target: whether a component failed, whether an image contains an object or whether a transaction was fraudulent. Attraction is different. The outcome is produced by two people in a specific interaction, and even they may not be able to explain exactly why they did or did not want another date.
That makes the speed-dating setting scientifically interesting. The model is not trying to recover a hidden physical fact. It is trying to predict a human decision from behaviour that is subtle, socially shaped and partly idiosyncratic. A signal can be statistically useful across a group without being a dependable rule for an individual conversation.
The study’s uncertainty therefore matters as much as its feature rankings. Facial movement, voice and language can all carry information, but finding patterns after a conversation is easier than proving those patterns will predict the next group of people. The authors’ call for independent data is especially important in a domain where overconfident conclusions can quickly turn into simplistic claims about what attraction looks like.
A dating app would need a much higher bar than an academic demo
There is a large gap between showing that multimodal data contain some predictive structure and using a model to judge real users. A deployed system could influence who gets shown to whom, how a conversation is interpreted or whether somebody is encouraged to continue. Errors would therefore shape the social environment the model is supposed to predict.
There are also obvious consent questions. People may accept that a dating service processes profile information, while feeling differently about software analysing facial movements or vocal behaviour during a live interaction. The fact that a signal can be measured does not automatically mean it should be used for ranking people.
The study is more revealing when treated as evidence about the limits of behavioural prediction. AI can extract patterns from a short human encounter, yet the researchers could not establish that their models clearly beat the simple baseline once uncertainty was accounted for. That is a useful result in itself. Human chemistry may leave measurable traces, but turning those traces into a reliable prediction remains much harder than spotting them after the fact.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
