AI & Work

The Same Qualifications Got Different AI Hiring Decisions

Two identical job candidates receive different hiring decisions from an AI recruiter after one uses an optimised CV.
Same candidate, same qualifications, different result. Research shows how changes to a CV can alter the decisions made by AI hiring systems.

In a controlled synthetic CV benchmark, the same qualifications produced different AI hiring decisions about which candidate ranked higher when researchers changed only the way a résumé was written or laid out. The tested language model with the strongest validity score reversed 29.6 per cent of its matched pairwise rankings after presentation changes. Another model flipped in 41.4 per cent. The finding does not prove that a real employer rejected anyone unfairly, but it exposes a basic screening risk: a system can appear to assess competence while also responding strongly to presentation.

That matters because an applicant may reasonably assume that changing bullet points into paragraphs, adding detail or polishing the language should not alter how their evidence is ranked against another candidate. A recruiter may make the same assumption when buying an automated screening tool. The new research suggests both assumptions need testing rather than trust.

How the AI hiring decisions were tested

The preprint, submitted on 15 September 2026, built a benchmark around 17 occupations and 102 synthetic candidate profiles. For each occupation, the researchers created six candidates spanning qualified, borderline and underqualified cases. Every candidate was then represented in five forms: an original bullet-point résumé, a more verbose version, a paragraph-based structure, an AI-polished version and a version containing layout or text-extraction noise.

The job requirements were derived from the US Department of Labor’s O*NET occupational database. The researchers used occupation-specific rubrics to score each résumé and checked whether the rewrites preserved the evidence about the candidate. Of 510 résumé variants, 505 passed that validation. This matters because the experiment was designed to alter presentation without quietly giving a candidate a new qualification or removing an existing one.

The test produced 255 base candidate comparisons and 1,250 approved pair rows. The headline flip-rate analysis covered 1,000 comparisons between an original résumé and one of its perturbed forms. Open instruction-tuned language models made independent pass or fail decisions against the same occupational rubric. Conventional BM25 and TF-IDF text-retrieval baselines were included as reference points.

The most accurate model was not fully stable

Llama 3.1 8B Chat recorded the strongest validity score among the tested language models, at 0.781, while changing 29.6 per cent of matched rankings. Llama 3.2 3B Chat was slightly more stable, with a 28.5 per cent flip rate, but had a lower validity score of 0.551. Mistral 7B v0.3 changed its decision in 41.4 per cent of comparisons, while Gemma 2 2B Chat changed in 45.6 per cent.

The simpler retrieval baselines were far more stable, with flip rates of 4.0 per cent for BM25 and 5.2 per cent for TF-IDF, but their validity scores were lower. That creates an uncomfortable trade-off. A system that appears better at applying a detailed rubric may also be more sensitive to surface changes that should be irrelevant. Stability alone is not enough, because a consistently poor ranking is still poor. Validity alone is not enough either, because equivalent evidence should not reorder candidates so often.

A flip is also not automatically an error. If a model first ranked the candidates incorrectly and then corrected the order, the change improved the outcome. If it moved away from the benchmark’s intended competence order, it introduced an error. The study’s flip rate measures presentation sensitivity, not discrimination, legal unfairness or real-world harm. That distinction is essential when translating a laboratory result into hiring policy.

Why CV presentation can sway a language model

Large language models do not read a CV as a fixed list of certified facts. They process sequences of words whose order, emphasis and surrounding context can affect the response. A verbose résumé may repeat evidence and make it more prominent. Paragraphs may blur the boundaries between roles. Polished phrasing can sound more confident even when it contains no new achievement. Extraction noise can interrupt dates, headings or relationships that were obvious on the page.

This is different from the familiar problem of keyword matching, although the two can overlap. LiveAIWire has previously examined how AI recruitment tools can filter candidates before a human review. The new benchmark adds another question: does the same tool reach the same judgement when the underlying evidence survives but the surface form changes?

The answer cannot be inferred from a polished demo or a vendor’s average accuracy figure. It requires paired testing. An employer has to submit competence-equivalent versions of the same application, compare outcomes and examine which formatting changes trigger a different result. That is closer to a crash test than a beauty contest for model scores.

What this means for job applicants

Applicants should not have to reverse-engineer an invisible model, and this research does not reveal a universal format that always wins. The model with the highest validity still changed its answer often enough to make any one-size-fits-all trick unreliable. Adding words, converting everything into bullets or using AI polish could help in one context and hurt in another.

A practical response is to make evidence easy to identify without inflating it. State the role, dates, responsibilities and measurable results clearly. Use ordinary headings and a simple reading order. Check the text extracted from a PDF rather than judging only the visual page. Preserve a plain-text version in case a portal mangles columns or graphics. Most importantly, do not invent experience to satisfy suspected keywords.

Applicants should also keep ownership of the final document. Evidence that AI assistance can leave people with weaker unaided skills is a reminder to verify every rewrite rather than outsourcing professional judgement. A candidate needs to be able to explain any polished claim later, particularly if the next stage is an interview.

The finding also sits beside a broader change in recruitment. Some candidates already face AI-assisted interviews that can accelerate offers, while others use generative tools to rewrite applications. If both sides automate, presentation may become an escalating contest between optimisation tools rather than a clearer account of whether someone can do the job.

What employers should test before deployment

For employers, the immediate lesson is not to ban language models. It is to treat consistency under harmless transformation as a release requirement. A screening system should score the same candidate evidence in bullets and paragraphs, short and expanded descriptions, different templates and realistically imperfect extraction, then compare the resulting rankings. Results should be broken down by role and qualification band, because borderline candidates may be particularly vulnerable to small shifts.

Human review must also be more than a ceremonial approval box. Reviewers need the source evidence, the rubric and a clear way to challenge the model’s reasoning. If the system cannot show which requirement drove the recommendation, a person may simply inherit its confidence. Logging model version, prompt, input and output is necessary for investigating complaints and reproducing decisions after an update.

Testing should cover protected groups as well as formatting. Earlier research has shown why gender bias in AI systems requires direct measurement. Presentation sensitivity is a separate failure mode, but the two can interact if résumé styles, employment histories or language patterns correlate with age, disability, nationality, gender or access to professional coaching.

A procurement score needs more than average accuracy

A buyer evaluating an automated screener should ask for paired consistency results, not only a headline accuracy score. The test set should reflect the organisation’s actual roles and applicant documents. It should include scanned PDFs, career breaks, part-time work, non-standard job titles and qualifications from different countries. The acceptable threshold should be set before the supplier sees the test results.

There also needs to be a fallback. If text extraction fails, confidence falls or equivalent versions disagree, the application should move to a trained human rather than being silently discarded. Monitoring after launch should look for changes when a model, parser, prompt or job description is updated. A system that passed last quarter’s test may behave differently after any of those components change.

This is becoming a governance issue as well as a quality issue. The European Commission lists AI tools used for employment, including CV-sorting systems, as high-risk uses under the EU AI Act. The precise obligations and implementation dates depend on the system and jurisdiction, but the policy direction is clear: traceability, documentation, human oversight, robustness and accuracy belong in deployment plans, not marketing footnotes.

What the study does not establish

The candidates were synthetic, the occupations were controlled and the tested models were open models used in an experimental pipeline. The study did not audit a commercial applicant-tracking system, observe recruiters, measure interviews or follow real people through an employment process. It therefore cannot tell us how often presentation changes real hiring outcomes or whether any named employer uses a similarly sensitive system.

The paper is also a preprint, so its methods and conclusions have not yet completed journal peer review. Its evidence-preservation checks reduce one important source of confusion, but synthetic rewrites can never capture every cultural, professional or accessibility feature of a real CV. Replication with deployed products and consented applicant data would be needed before estimating real-world prevalence.

Even with those limits, the result identifies a testable weakness. Employers already ask whether an AI system agrees with human labels. They should also ask whether it agrees with itself when the facts stay the same. If presentation alone repeatedly changes which candidate ranks higher, the organisation has not merely found a formatting quirk. It has found a decision process that is not yet dependable enough to stand between a person and an opportunity.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.