A group of university students recorded higher scores on a financial knowledge questionnaire after a week with an AI teaching chatbot. That sounds like the kind of result that could make AI financial literacy lessons attractive to schools and universities. But the same study recorded a striking deterioration on one inflation question, and its researchers warn that their small experiment cannot establish that the chatbot caused the overall improvement.
The open-access study in Discover Education, published in September 2026, followed 32 university students in Türkiye. Participants used a specially structured educational chatbot during a short learning period and answered questions before and afterwards. The researchers observed improvement in their overall questionnaire results, while carefully noting that there was no comparison group learning without the tool. The findings are a test of a possible teaching approach, not a verdict on AI tuition in general.
What the AI financial literacy experiment involved
Financial literacy means understanding everyday money concepts well enough to make informed decisions. That includes knowing how interest works, what inflation does to purchasing power, why an investment carries risk and how borrowing obligations develop over time. These ideas matter well beyond the classroom. Someone who misunderstands them may make an expensive decision even while confidently using financial language.
The study’s chatbot was not designed to issue personal investment instructions. It was built as an educational aid, with prompts and teaching structure intended to help students work through financial topics. Participants were told it was not a source of personalised financial advice. That distinction matters because explaining a concept and recommending an action with a person’s actual money involve different levels of information and responsibility.
The researchers measured a collection of topics including financial statements, basic economics, retirement and insurance, banking, investment, taxation and legislation. The overall average score increased during the experiment. But these were not results from a large randomised trial, and the questionnaire included agreement-rated knowledge statements rather than relying entirely on separate, objective problem-solving tasks. Measuring a learning tool requires deciding whether students know more, feel more certain or have simply become familiar with the questions.
According to the published analysis, the overall mean rose from 4.33 to 5.00 on the study’s response scale. That is an observed before-and-after change, not proof that the chatbot alone produced it. The same students were tested twice, over a short interval, and the study did not include a group receiving the same material through another teaching method. There is no way to rule out all the effects of repetition, independent learning or changes in confidence.
The inflation question that went the wrong way
One result cuts through the appealing average. Students became less accurate when responding to a false statement about inflation in Türkiye being below 10 per cent. The paper reports that 13 students moved towards a less accurate answer on that item, 16 were unchanged and three improved. The researchers also recorded a decline on a risk-and-return comparison. Those observations do not negate the overall improvement, but they show why a total score can conceal a mistake on a consequential concept.
An incorrect belief about inflation can alter how someone thinks about saving or a promised future payment. People may know the definition of inflation yet misunderstand the scale of price changes in their country. A teaching exercise must therefore do more than help a learner repeat an explanation. It has to check whether the learner can apply the idea to a real statement and spot when an apparently plausible claim is false.
The paper does not establish that the chatbot directly taught the incorrect inflation statement. Other factors may have contributed, and interaction logs were not gathered. Without a detailed record of what individual students asked and what they received, investigators cannot trace each misunderstanding back to a particular response. That uncertainty is exactly why the authors advise against treating the short pilot as causal proof of effectiveness.
Learning is not the same as liking an explanation
Chatbots have an obvious educational attraction. Students can ask a basic question privately, request a different explanation and return to a confusing topic without feeling they are slowing down a class. The tool can also offer worked examples and adapt its language to a learner’s immediate question. Those are potential advantages, not guarantees that the resulting knowledge will be correct or retained.
A student can finish a conversation feeling reassured while still holding an incorrect belief. Conversely, a difficult exercise may feel frustrating yet reveal an important gap. Good evaluation should capture performance on new tasks, not simply whether the exchange feels helpful or whether the same question seems easier on the second attempt. In money education, the strongest test is whether the learner can apply ideas accurately to unfamiliar decisions.
This concern extends to ordinary consumer finance. The OECD’s 2026 paper on AI and personal finance identifies opportunities for accessible information and tailored education, while highlighting risks from inaccurate outputs, bias, privacy and commercial influence. The OECD does not argue that every AI explanation is wrong. It points to the need for people to evaluate claims, sources and incentives when seeking help with real financial decisions.
What would make the next study stronger?
The authors outline ways to improve on the pilot. A comparison group could learn the same topics without the chatbot or with another educational intervention. Objective questions could test new problems rather than mainly agreement with familiar statements. A later assessment could show whether the apparent learning survives beyond the week of use. Usage records could also indicate which explanations students actually encountered and which activities seem associated with better results.
The sample matters. With 32 matched participants from one university, the study cannot tell us how the method would work across different ages, backgrounds or education systems. It also cannot identify reliable differences between subgroups. A programme intended for schoolchildren, adults dealing with debt or people approaching retirement would require its own testing and safeguards. One positive classroom pilot is not a licence to generalise across every financial learner.
The same applies to language. Financial terms can be deceptively familiar. The word interest can mean a cost, an earning or a motivation, depending on context. Inflation is not the same as the change in the price of one product, and a rising salary does not automatically mean higher spending power. An effective tutor must check which meaning the student has understood rather than simply producing a fluent explanation containing the right words.
How someone can use an AI tutor more carefully
A safer approach is to treat a chatbot as a practice partner rather than an answer authority. Ask it to explain compound interest, then work through an example yourself and verify the calculation against an independent educational resource. Ask it to supply a counterexample or explain why a tempting answer is wrong. If the question concerns taxes, a regulated product or the terms of a loan, check the relevant official source and the current rules in your jurisdiction.
When a student gets an answer wrong, the useful next step is not another enthusiastic summary. It is to find the precise misunderstanding. Did they confuse nominal and real returns? Did they fail to notice a time period? Did the exercise use information that had become outdated? The teaching conversation can then focus on the mistake. The research suggests why this level of feedback needs to be measured rather than assumed from a higher headline score.
LiveAIWire has covered AI advice that varied according to the user, highlighting another reason to test the actual output rather than trusting a product category. The quality of a financial answer can depend on the question, the information provided and the assumptions made by the system. An attractive explanation is not proof that the reasoning applies to the individual reader.
The distinction between an assistant and an adviser
Consumer-facing AI tools can be useful for learning vocabulary, producing practice questions and organising information for a discussion with a qualified professional. The threshold changes when a person asks what they personally should buy, borrow or sell. Real decisions require details about their circumstances, local rules and appetite for risk. A generic teaching chatbot is not responsible for verifying all of those facts, and it cannot guarantee a result.
Financial education also concerns habits and behaviour, not simply correct answers on a screen. A person may understand what a budget is while struggling to use one when income is irregular. Another may know how an investment works but need help deciding how much risk they can bear. A good learning programme should connect knowledge to decision-making while avoiding pressure to take actions a student does not understand.
LiveAIWire has also examined AI tools used in personal financial management, a broader consumer question than the narrow educational pilot. The new paper contributes a useful distinction: being given information about money is not the same as demonstrating sound understanding of it.
The experiment is therefore more interesting than a simple success story. Scores increased, but not every answer improved, and the design cannot identify the cause. That combination offers a valuable standard for the next wave of AI tutors. Judge them not by how friendly the conversation feels, nor by an average score alone, but by whether learners become more accurate, remain accurate later and can explain where their answers came from.
For further context, LiveAIWire’s reporting on AI bank lending decisions considers another setting where money-related predictions have consequences. In each case, the critical question is whether the system helps people arrive at reliable judgements, not simply whether it can generate a convincing response.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.
