AI & Society

Who’s Grading the Exam? AI in Professional Licensing and Certification Tests

AI licensing exams illustration of a certification document being reviewed alongside AI circuitry
AI can already pass the bar, CPA, and CFA exams. Here's how testing bodies are responding.

AI licensing exams are now a two-sided story that the professional bodies setting them are only beginning to reckon with. On one side, large language models have gone from failing simulated CPA and bar exams to passing both by comfortable margins, and can now clear all three parts of the notoriously difficult CFA exam in minutes.

On the other side, the organisations that write and grade these exams, the ones deciding who gets to call themselves a lawyer, an accountant, or a chartered financial analyst, are moving cautiously, sometimes reluctantly, toward using AI in the testing process itself. Those two trends are colliding in the same professions at the same time, and the outcome will decide what a professional credential is actually supposed to prove.

What AI Licensing Exams Reveal When AI Takes Them

The performance numbers are no longer a curiosity. Reporting compiled for Bloomberg Tax tracked the trend directly: as large language models have improved since ChatGPT’s 2022 public launch, CPA and bar exam pass rates for AI systems climbed from failing to comfortably passing, while human candidate performance over the same period slipped slightly, staying within historical norms but moving in the opposite direction from the machines.

A related CFA Institute finding cited in the same reporting found Level I human pass rates dropping to 43 percent in the same window that AI models were acing the exam outright.

What that gap means depends entirely on what the exam was actually built to measure. If a licensing exam mainly tests recall of rules, formulas, and structured legal or accounting doctrine, an AI system trained on enormous volumes of exactly that material has an inherent advantage no amount of human study time can close.

If the exam is meant to certify judgement, the ability to apply that knowledge to a messy, specific, real-world situation under pressure, the picture is considerably less settled, and it is precisely the question every major licensing body is now grappling with.

Why the Bar Exam’s Own Testing Body Is Moving Carefully

The National Conference of Bar Examiners, which writes bar exam content used across 54 US jurisdictions, has taken a notably cautious public position on AI in the testing process itself, even as it modernises nearly everything else about the exam. The NCBE’s next-generation bar exam, launching in July 2026, moves the entire test onto computers for the first time and partners with a new testing platform for grading infrastructure.

But reporting from the ABA Journal quotes Kara Smith, the NCBE’s chief product officer, drawing a clear line: artificial intelligence is not being used for grading at this time, and while the organisation is researching whether AI could eventually support human-conducted grading, “we feel strongly that AI will never replace humans in our development process.”

The NCBE’s own research arm has been more specific about what that support role might look like. In an official article published in The Bar Examiner, NCBE Chief Product Officer Kara Smith McWilliams describes ongoing “artificial intelligence (AI)-assisted scoring research to explore how AI models can provide secondary support to human grading through data-driven insights while preserving the integrity of expert judgment.”

The framing is deliberate: AI as a secondary check inside a process where two independent human graders already score every written response, not a replacement for either of them. The NCBE’s current system uses double grading with adjudication built in specifically to catch grader inconsistency, and the stated ambition for AI is to make that existing human process more efficient and more consistent, not to hand the decision to a model.

What This Approach Gets Right

The caution is not merely institutional conservatism. A licensing exam that certifies competence to practise law, medicine, or accountancy carries a public protection function that an ordinary school exam does not, and the NCBE’s insistence that AI support rather than replace human graders reflects a reasonable judgement about where the technology’s current strengths and weaknesses actually sit.

AI systems are demonstrably strong at pattern recognition across large volumes of structured, rule-based content, precisely the kind of material multiple-choice and formula-driven questions test. They are far less proven at the kind of holistic, context-sensitive judgement a human grader applies when assessing whether a candidate’s legal reasoning in an open-ended written response reflects genuine competence or a fluent-sounding approximation of it.

The Uncomfortable Question These Numbers Raise

If an AI system can pass the same exam a human candidate spends years studying for, the exam’s value as a signal changes, whether or not the testing body changes anything about how it is administered. The Bloomberg Tax analysis frames the dilemma precisely: if these exams primarily indicate technical proficiency, AI’s ability to pass them threatens the near-term relevance of the credential itself, since firms could plausibly deploy AI systems alongside less credentialed staff to do work once reserved for certified professionals.

If the exams instead indicate minimum competency, a baseline a working professional must clear rather than the full skill set the job requires, AI’s passing performance is less threatening and the credential’s value shifts toward what it demonstrates about the person who earned it: perseverance, the capacity to be retrained as tools change, and the judgement to know when an AI-generated answer is wrong.

That second framing is where most licensing bodies appear to be landing, at least implicitly, in how they are redesigning their exams. The NCBE’s own next-generation bar exam was explicitly built around testing lawyering skills and judgement in realistic practice scenarios rather than pure doctrinal recall, a redesign that predates the current wave of generative AI but that now serves, deliberately or not, as a harder target for a model that is strong on recall and weaker on situated professional judgement.

How This Differs From the Cheating Crisis Playing Out in Classrooms

It is worth being precise about what this piece is and is not describing. LiveAIWire’s earlier reporting on the AI exam cheating crisis documented a related but distinct problem: candidates using AI to cheat on remote, unsupervised professional exams, a problem serious enough that the Association of Chartered Certified Accountants ended remote invigilation entirely and sent candidates back to physical test centres after a study found 94 percent of AI-generated exam answers went undetected by human markers.

This piece is about a different question: not whether candidates are cheating with AI during an exam, but what it means for a credential’s value that AI can pass the exam legitimately, on its own, without needing to cheat at all, and how the bodies writing these exams are deciding whether and how to let AI touch the grading process itself. The two problems compound each other, since a testing body redesigning its exam to resist AI cheating is, at the same time, trying to design an exam AI cannot simply pass outright either.

What This Means for You

If you hold a professional credential or are studying for one, the practical implication of these trends is not that your license is about to become worthless. Every major testing body examined here, from the NCBE to the CFA Institute, is actively redesigning its exams around exactly the capabilities AI cannot yet reliably replicate: situated professional judgement, reasoning under unscripted pressure, and the accumulated experience that lets a professional recognise when a plausible-sounding answer, human or AI-generated, is actually wrong.

That redesign is likely to make credentialing exams harder in a different way, not easier, over the next several years.

For anyone evaluating whether a credential is still worth the study hours it demands, LiveAIWire’s broader reporting on the professions AI is creating found that roles requiring exactly the kind of judgement licensing bodies are now designing exams around, auditing AI-generated work, training and correcting AI systems, and applying professional expertise to catch AI’s mistakes, are among the fastest-growing categories of new work AI has produced. A credential that certifies the ability to catch a wrong answer is arguably becoming more valuable as AI generates more answers that need catching, not less.

The Same Recall-Versus-Judgement Problem Shows Up Everywhere AI Meets Testing

The tension between what a test measures and what AI can already fake is not unique to professional licensing. LiveAIWire’s coverage of the prompt engineering myth found a closely related pattern in a completely different context: skills that look impressive in a controlled demonstration, reciting a well-crafted prompt or reproducing a memorised answer, often turn out to be far less predictive of genuine capability than a messier, less quantifiable form of judgement developed through sustained practice. Licensing bodies redesigning their exams around unscripted reasoning and situated judgement are, in effect, making the same bet: that the skill worth certifying is the one that resists being reduced to a pattern a model can learn to imitate.

Who Gets Left Behind in the Redesign

The harder-to-answer question sits in access rather than exam design. LiveAIWire’s coverage of AI recruitment tools screening job candidates found that algorithmic hiring systems consistently disadvantage candidates without access to the specific preparation resources, coaching, and practice materials that better-resourced applicants can afford.

The same access gap applies directly to licensing exams moving toward oral examination, portfolio review, and supervised practical assessment: these formats are more expensive to administer and harder to access for candidates who cannot easily travel to a test centre, take unpaid leave for a placement, or afford intensive one-on-one coaching. A testing system built to resist AI more effectively risks becoming, at the same time, a testing system that is harder to access equitably, and the professional bodies making this transition will need to solve both problems at once rather than treating the second as a secondary concern.

The Honest Verdict

AI licensing exams sit at a genuine inflection point rather than a settled outcome. The technology can already pass the credentialing tests that once reliably separated qualified professionals from everyone else, and the bodies responsible for those tests are responding not by handing grading over to AI but by redesigning what the exams measure in the first place. Roughly 45 jurisdictions have already committed to adopting the NCBE’s redesigned exam, a scale of coordinated change across an entire profession that reflects how seriously the testing infrastructure itself is being rebuilt, not merely patched.

Whether that redesign succeeds in preserving the public trust these credentials exist to protect depends less on how sophisticated the AI systems become than on whether testing bodies keep human judgement at the centre of both the exam content and the grading process, exactly the line the NCBE has drawn, and exactly the line worth watching whether every other licensing body holds.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.