By Stuart Kerr, Technology Correspondent, LiveAIWire
AI comedy algorithms cleared a bar most people assumed was safely human only a couple of years ago: in a University of Southern California study published in PLOS ONE, nearly seventy percent of participants rated jokes written by ChatGPT as funnier than jokes written by ordinary people, against twenty-five percent who preferred the human version. In a second round of the same study, ChatGPT’s rewrites of Onion-style satirical headlines scored, on average, exactly as funny as headlines written by professional comedy writers. That is a genuinely startling result for a technology that, by most other measures, still struggles badly with humour.
The contradiction is not really a contradiction once you look at what each result is actually measuring. Short, punchy, pattern-based jokes, acronyms, one-liners, fill-in-the-blank gags, are exactly the kind of task large language models are built for. Longer-form, character-driven, personally rooted comedy is a different problem entirely, and that is where AI comedy algorithms keep failing in ways their creators have struggled to fix.
Table of Contents
What This Means If You Use AI to Write or Judge Comedy
If you are a comedian, a marketer, or just someone who asks a chatbot for a joke before a work presentation, the practical lesson from the research is to match the tool to the task. AI comedy algorithms are genuinely useful for short, structured, low-stakes humour: captions, headline rewrites, icebreakers, quick wordplay. They are far less reliable for anything that depends on lived experience, cultural specificity, or knowing exactly how far a joke can be pushed before it stops being funny and starts being something else. Treating a chatbot as a first-draft machine for structure, not as a substitute for a comedian’s judgement, is the distinction the evidence actually supports.
Drew Gorenz, the USC doctoral researcher who led the joke-rating study, was blunt about why the machine wins on some tasks and not others. Large language models, he said, “don’t know what it feels like to appreciate a good joke” and are “mostly just using pattern recognition” rather than anything like comic instinct. Even so, he was equally clear about the limits: “I don’t think it’s able to create a John Mulaney level joke,” he said, and stand-up delivery in particular loses most of its power once you strip it down to text on a page.
AI Comedy Algorithms and the Bias Problem Comedians Keep Flagging
The more serious findings come from a Google DeepMind study, titled “A Robot Walks Into a Bar,” that put twenty professional comedians through workshops at the Edinburgh Festival Fringe and asked them to write stand-up material with the help of chatbots. The verdict was blunt. One participant called the output “the most bland, boring thing, I stopped reading it, it was so bad.” Another described it as “cruise ship comedy material from the 1950s, but a bit less racist,” a line that captures the study’s central finding better than any statistic could.
That blandness is not an accident of training data volume. It is a direct consequence of how safety filtering works. The comedians in the study reported that the moderation systems built into mainstream chatbots consistently defaulted toward material centred on straight, white, mainstream perspectives, and struggled badly when pushed toward material reflecting other cultural viewpoints or experiences. Researchers described this as a form of algorithmic censorship rather than neutral safety, since the systems were not simply blocking harmful content, they were narrowing whose comic voice got represented at all.
The same over-caution showed up around dark humour specifically, which is a normal and long-established comic register that AI comedy algorithms handle badly. One comedian in the study said a chatbot refused to generate dark material because it wrongly inferred, from a joke involving mortality, that the comedian herself needed a crisis intervention rather than a punchline. That single example captures a pattern researchers found repeatedly: models trained to be cautious about harm frequently cannot tell the difference between a person in distress and a person doing their job.
Whose Jokes the Algorithm Learned From
Thomas Winters, one of the researchers behind the DeepMind study, put the underlying technical problem simply: humour depends on a “frame-shifting prerequisite” that draws on lived memory, cultural context, and split-second timing in ways current models are not built to replicate. A joke that lands in Mumbai and a joke that lands in Manchester are often not the same joke wearing different words, they depend on entirely different shared assumptions about what an audience already knows, and AI comedy algorithms trained overwhelmingly on English-language internet text default hard toward one cultural frame at the expense of every other.
That default has real consequences beyond awkward chatbot output. When platforms use similar language models to auto-generate or auto-moderate comedic content at scale, the same narrowing effect plays out across millions of pieces of content rather than one workshop transcript. A joke rooted in a specific community’s cultural reference points is statistically more likely to get flagged, softened, or simply never generated in the first place than a joke built from the more heavily represented material the model was trained on, a pattern that echoes the exposure work artists have been doing in the new wave of AI creative activism built around surfacing exactly this kind of bias.
The Ownership Question AI Comedy Cannot Avoid
Bias in what gets generated is only half the picture. The other half is whose comedic voice AI systems are allowed to imitate at all. In January 2024, the estate of the late comedian George Carlin filed a lawsuit against a media company accused of using AI to mimic his style and material without permission, a dispute that tested how far posthumous rights extend when a comedian’s entire recognisable voice, not just individual jokes, becomes trainable data. That case sits alongside broader unresolved questions about who consents when a deceased performer’s likeness is reconstructed, an issue we examined in our coverage of the ethics of digitally reviving the dead.
None of this means AI comedy algorithms have no place in comedy. Used as a brainstorming partner for structure and premises, the technology can genuinely speed up a writer’s process, in much the same accelerant role AI has already taken on behind the scenes in theatre production, and the USC results show it can even outperform amateur joke writers on narrow, text-based tasks. The problem is treating pattern-matched output as equivalent to a comedian’s judgement about context, audience, and exactly how far a joke should go, a judgement that depends on the kind of lived cultural specificity training data cannot substitute for.
What Comes Next for AI in Comedy
The comedians in the DeepMind study were not arguing for banning AI tools from their process. Several described real, if limited, uses for outlining structure or generating raw material to react against. What they were flagging, consistently, was that the current generation of chatbots optimises for broad inoffensiveness in a way that quietly erases the specific, risky, personal material that makes comedy worth watching in the first place. Fixing that is not primarily a technical problem. It is a question of whose humour gets treated as the neutral default setting, and that is a decision made by the people building these systems, not by the algorithms themselves.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, emerging technology, and their impact on business, society, and everyday life. LiveAIWire publishes original AI journalism every weekday at liveaiwire.com.