How Often Is AI Wrong? What the 2026 Numbers Show
Last Updated: September 2026
How often AI gets things wrong is a question about your test, not about the model. An AI error rate is the share of outputs that fail a check on one task, under stated rules, against an answer someone agreed was right. Change the task and the figure moves. Change the person grading, and it moves again. A published rate tells you what one model did on somebody else’s test, on one day.
AI Smart Ventures has guided growing businesses through this exact mix-up, and the pattern repeats. An owner quotes a score read in a newsletter, then admits that nobody has counted what the team’s own tools put out last month. Those are two different questions, and only the second one can be acted on.
An uncounted error rate is not zero. It is unknown, and unknown is what stalls AI adoption in month three. One wrong figure that reaches a client does more harm than fifty caught in a draft, because the client is the one who finds it, and trust rebuilds slowly. Counting what gets out takes an hour a month. Guessing costs far more.
Key Takeaways
- The useful number is yours, not the model’s: count how many wrong outputs reached a client last month, then work back to the task that produced them.
- A rate without its test means nothing: a January 2026 study scored one model at 78.3% on a surgical exam, while a third of its answers leaned on made-up or misread sources.
- Sample, do not survey: grade 50 to 100 real outputs each month against a written standard, and log which errors were caught and which escaped.
- Zero in a small sample is not zero: read 100 outputs, find nothing wrong, and a true rate near three in a hundred still fits your evidence.
Those four points share one root. Being right is a property of a test, not of a tool, so the only rate that governs your risk is the one measured on your work, by your people, against your standard. That changes the buying question too. You stop asking which model is most accurate, and start asking which task you can check fast enough to trust.
What do published AI error rates actually measure?
A published error rate covers one model, on one set of questions, graded one way, on one date. Nothing else travels with the number. When the BBC and the EBU had journalists at public broadcasters grade more than 3,000 answers from four AI tools in October 2025, across 18 countries and 14 languages, 45% held at least one serious problem. Weak sourcing drove 31% of those. Plain errors of fact drove 20%.
How answers get graded matters as much as which questions get asked. A paper by Kalai and colleagues looked at ten popular tests in September 2025 and found that nearly all of them marked an answer simply right or wrong, with no credit for saying “I do not know”. A model that guesses scores as well as one that admits doubt. Two tools can post the same score while one of them makes things up far more often.
Why does the same model score so differently?
Because the task decides what counts as an error, and most tests measure only one kind. A January 2026 study in Cureus ran GPT-5 through 203 released orthopaedic exam questions. It got 78.3% right, which sounds strong. Yet 33% of its answers cited sources that were made up or did not back the claim, and that figure hit 50% among the answers it got wrong. One run, two very different rates.

Wording shifts the result too. Researchers at Northeastern and four other schools put 6,614 paired patient questions from clinical trial summaries to eight models in April 2026. Pairs that differed only in a positive or negative slant were far more likely to give opposite answers from the same evidence, and pushing back over several turns widened the gap. Your staff do not phrase requests like a test set, so a published rate never transfers to your desk.
How do you measure your own AI error rate?
Pull a sample, grade it against a written standard, and log what escaped. Fifty to a hundred real outputs a month is enough to start. You need a shared meaning for wrong, which is the part most teams skip. The BBC and the EBU put out a public list of AI answer faults and presented it as a shared test on 12 March 2026, covering accuracy, sourcing, context, and opinion dressed as fact. Borrow those four headings.
Run it as a fixed hour, not a project. That is practical AI. One person pulls last month’s outputs at random, reads each against the standard, and marks it clean, caught or escaped. Caught means a reviewer stopped it inside the team. Escaped means a client, a customer or a regulator could have seen it. Two people grading the first twenty together settles most arguments. After that the hour holds, and the number starts to mean something month over month.
- Sample at random: pull outputs by date or by ticket number, never by the ones somebody happens to remember.
- Write the standard first: one page saying what counts as wrong, agreed before anyone grades. Shared AI literacy begins there.
- Split caught from escaped: the caught count tells you the tool is imperfect, while the escaped count tells you the process is.
- Log the task, not just the total: record which job produced each error, because that is where the fix belongs.
Setting up that first month of grading is where AI implementation support repays the effort quickly. Bring the task your team runs through AI most often, and we will build the standard and the sample around it.
What does a good internal error rate look like?
There is no single number, but the shape is steady: a caught rate you can live with, and an escaped rate near zero on anything a client sees. Sort your tasks by what a mistake costs in rework and in trust. Internal drafting can take a high caught rate, since a person reads it anyway. A client-facing figure or a cited claim cannot. Set the target task by task, then check whether last month cleared it.
Small samples flatter you, and the maths behind that is old. A 1995 BMJ note by Eypasch and colleagues set out the rule of three: when an event has not happened across a run of trials, the highest rate that still fits, at 95% confidence, is about three divided by the number checked. Read a hundred outputs, find nothing wrong, and three in a hundred still fits what you saw. Report the number with that caveat attached.
| Task type | What to track | Sensible target |
|---|---|---|
| Internal drafts | caught errors per 100 outputs | high tolerance, light review |
| Research and cited claims | sources made up or misread | open every source before use |
| Client-facing work | escaped errors per 100 outputs | as near zero as the checks allow |
What should you do with the number you get?
Spend it. A measured rate settles three things: where review time goes, which tasks widen, and which ones stop. Put your reviewers on the tasks with the highest escaped rate, rather than spreading them evenly across everything the team does. Measure again after any model change or prompt change, because the figure is a snapshot and vendors ship updates often. Then share it inside the team, so people argue about evidence instead of hunches.
The number changes how you buy, too. Generic AI consultants tend to lead with a model or a platform, while a rate you measured yourself points at workflow optimization: a task and a checking step. Take it into vendor talks and ask how a tool would move your escaped rate on the one job that produces most of your errors. Watch it monthly for a quarter before reading a trend into it, since one bad month usually reflects who graded, not what changed.
Frequently Asked Questions
How often do AI projects fail?
The number moves each year, and each study counts failure its own way. S&P Global Market Intelligence asked more than 1,000 firms in North America and Europe, and reported in March 2025 that 42% had scrapped most of their AI projects, up from 17%. The average firm also dropped 46% of trials before they went live. Those are drop-out rates, not error rates.
What should I do if my AI gives a wrong answer?
Fix the output, then write the error down before you move on. Note the task, the date, what you asked for, and whether anyone outside saw it. That one line turns a bad afternoon into data you can count at month-end. Without a log, every wrong answer feels like the first, and you never learn whether this month beat the last.
What percentage of the time is AI wrong?
There is no single percentage, and any article handing you one has dropped the test it came from. In the BBC and EBU review of news answers, 45% held at least one serious problem. In a January 2026 exam study, one model got 78.3% of questions right. Same technology, different jobs, different graders. Ask which task, which rules, and which date before quoting a figure.
How often is AI wrong about the news?
More often than most readers assume, and sourcing is the bigger problem. Journalists at 22 public broadcasters graded over 3,000 answers from four AI tools, across 18 countries and 14 languages. Nearly half held at least one serious problem, about a third had weak or missing sourcing, and a fifth had real errors of fact. Treat any answer about the news as a lead.
Can AI be 100% accurate?
Not on open-ended writing work, and a claim of zero errors usually means nobody has measured. What you can reach is a very low escaped rate, where mistakes get caught before anyone outside the team sees them. That is a result of your process, not the model. Narrow tasks with clear right answers and a checking step come closest, which is why scope beats model choice.
How accurate is AI in medical diagnosis?
Exam scores and diagnosis are not the same thing, and the published figures measure the first. A January 2026 study in Cureus scored GPT-5 at 78.3% on 203 released orthopaedic exam questions, yet a third of its answers cited sources that were made up or did not back the claim. Exam questions arrive pre-cleaned. A patient does not, so read clinical figures as a ceiling.
How many outputs do I need to check each month?
Fifty to a hundred works for most teams, and doing it the same way each time matters more than doing more. Sample at random, grade against the same written standard, and hold the sample size steady so the trend means something. A clean sample of 100 still leaves room for a true rate near three in a hundred, so report that caveat rather than claiming zero.
Do AI error rates fall as models improve?
Some do, and the ones that matter to you may not move at all. Vendors ship updates often, and each one changes how a model behaves on your prompts in ways no ranking predicts. Run your sample again after any model or prompt change, then compare it with the same month last quarter. Gains you have not measured on your own work are a sales claim.
Should benchmark scores decide which AI tool we buy?
Use them to build a shortlist, never to decide. A benchmark tells you how a model did on public questions under test rules, which says little about how it will handle your records and your wording. Run a two-week trial on one real task, grade the output against your own standard, and compare escaped errors. That beats any published ranking, and it costs about an hour a week.
What does it take to start measuring AI accuracy?
A week of setup, then about an hour a month. You need one page saying what counts as wrong, a random sample of last month’s outputs, and one person who grades them. AI Smart Ventures works with growing businesses on this through AI advisory and capability building, so the habit outlasts the first month. Book a consultation to set your standard and run your first sample.
Executive Summary
Published AI error rates describe one model, on one test, graded one way, on one date. They run from a fifth of news answers holding errors of fact to a third of exam answers leaning on made-up sources, and none of it transfers to your desk. The number that governs your risk is the one you measure: how many wrong outputs reached a client last month. Pull 50 to 100 real outputs, grade them against a written standard, and split caught errors from escaped ones. Then move review time to where the escapes are.
What Should You Do Next?
Pick the one task your team runs through AI most often, and write a single page this week saying what counts as a wrong output for it. Pull 50 outputs from last month at random, grade them against that page, and mark each one clean, caught, or escaped. Put the escaped count somewhere visible, then repeat the same hour next month.
AI Smart Ventures offers AI Implementation for growing businesses, building checks into work they already run on AI. Schedule a consultation to set your standard and read your first month of results.
People Also Read
About the Author
Nicole A. Donnelly is the Founder of AI Smart Ventures and an AI Adoption Specialist with 20 years of experience as a founder and CEO and over a decade leading AI adoption initiatives. She helps businesses integrate artificial intelligence with clarity and confidence, driving innovation and sustainable growth. Nicole has trained over 20,217 professionals in Applied AI, delivered 624 workshops, and worked with close to 1,000 organizations across diverse industries.
Expertise: AI Transformation, AI Strategy, AI Implementation, AI Adoption, Applied AI, Marketing, Business Operations
Disclaimer: This content is for informational purposes only and does not constitute professional business or technology advice. Results vary based on industry, existing systems, and implementation commitment. Contact AI Smart Ventures for a consultation regarding your specific situation.


