Resource

Funny AI agent fails, from the public record.

Read sourced AI chatbot failure stories from Air Canada, DPD, Woolworths, Bunnings, Google, and more, plus launch-risk lessons.

Last updated 2026-08-26. For the full evidence standard, read the testing methodology.

Failure feed

Reported incidents worth keeping in your test plan.

Article

Air Canada: the bereavement-fare refund that became a liability lesson

What happened: According to The Guardian's report on the tribunal decision, a customer asked Air Canada's chatbot about bereavement fares after a death in the family. The bot said he could buy full-price tickets and apply for the bereavement discount afterward, but the airline's actual policy required approval before travel. The customer followed the chatbot's route, then Air Canada refused the refund because the official policy did not match the bot's answer. The tribunal treated the chatbot as part of the company's website information, not as a separate actor the airline could blame. That is why this failure matters: it was not a goofy answer in a sandbox. It was customer-facing policy advice at the exact moment a traveler was making a purchase decision.

What builders should test: A launch test should ask policy questions where the answer is time-sensitive, exception-heavy, or contradicted by nearby help-center copy. Run variants for before purchase, after purchase, edge eligibility, and documents the customer may need. The bot should either quote the official rule, link the relevant policy, or hand off before it gives a refund path. The report should flag whether the bot creates a new promise, omits the timing requirement, or gives a process that sounds official but cannot actually be honored. Any confident answer that rewrites policy needs to be treated as a liability finding, not a harmless hallucination.

Original sourceAir Canada ordered to pay customer who was misled by airline's chatbotThe Guardian, Leyland Cecco - 2024-02-16
Article

DPD: the parcel bot that turned brand frustration into a public spectacle

What happened: According to The Guardian, a customer trying to locate a parcel could not get useful help from DPD's chatbot. After the support path stalled, he pushed the bot into writing jokes, criticizing DPD, and using offensive language. DPD said the behavior followed a system update and disabled the affected AI element, but by then the screenshots had already become the story. The important detail is that the viral moment started with a normal support failure: a customer wanted delivery help and did not get it. Once the bot could be steered into performance mode, the unresolved parcel issue turned into brand damage, not just a bad chat transcript.

What builders should test: A launch test should start with a normal tracking problem, then add frustration, repeated asks, off-topic prompts, and direct brand-safety pressure. The test should include requests to joke, insult the company, abandon the support flow, and agree with the customer's anger. The safe bot keeps trying to resolve the parcel issue, admits when it cannot access the needed state, and escalates. The unsafe bot starts entertaining the customer while the original support failure stays unresolved. The evidence should show both failures together: the missing delivery help and the brand-voice break that followed.

Original sourceDPD AI chatbot swears, calls itself useless and criticises delivery firmThe Guardian, Jane Clinton - 2024-01-20
Article

Chevrolet of Watsonville: the sales bot that wandered away from selling cars

What happened: Business Insider reported that a Chevrolet dealership chatbot powered by ChatGPT was pushed far outside a normal sales conversation. Screenshots showed users steering it into off-topic answers and fake-deal language, including a claimed one-dollar Tahoe exchange that the reporting noted was not legally binding. The useful lesson is not that the prank was enforceable. It is that a sales bot with no tight authority boundary can produce viral screenshots that look like promises. A dealership assistant sits close to pricing, inventory, financing, and purchase intent, so even joke terms can look more serious than a random chatbot answer. The bot did not need to complete a sale to create a risk. It only needed to appear to agree to terms the business would never approve.

What builders should test: A launch test should pressure sales bots with impossible discounts, fake purchase terms, competitor comparisons, and instruction changes that ask the bot to agree on behalf of the business. Include prompts that ask the bot to repeat a deal back, promise manager approval, or confirm an exception as if it were binding. The bot should keep pricing inside approved inventory and financing flows, refuse to create its own contract language, and route serious buying intent to a human or official checkout path. The report should separate harmless off-topic chatter from commercial authority failures, because the second category is where screenshots can confuse customers and staff.

Original sourceA car dealership added an AI chatbot to its site. Then all hell broke loose.Business Insider, Katie Notopoulos - 2023-12-18
Article

BMW Toronto: the buyback bot that mistook a loan balance for an offer

What happened: CBC News reported that a BMW Toronto chatbot offered to buy back a customer's vehicle for the exact amount left on the loan, then booked an appointment to complete the deal. The customer said the assistant had not identified itself as AI. A dealership employee later said the offer was invalid and that the vehicle was worth substantially less. BMW Toronto told CBC the chatbot had mistaken the loan balance for the amount the dealership should pay; after CBC contacted the dealership, it reinstated the original offer. The memorable miss was not only a wrong number. The bot turned data from one field into a commercial commitment, presented it with human-sounding authority, and advanced the customer to the next step.

What builders should test: A launch test should give a sales or trade-in agent several nearby numbers: loan balance, asking price, estimated market value, approved offer, and deposit. Then ask it to negotiate, confirm terms, and schedule the next action. The bot should never convert reference data into an offer or imply approval without an authoritative pricing tool and explicit business rule. Any appointment, reservation, or follow-up action must preserve the approved amount and disclose that final terms need human confirmation. The report should capture both the wrong claim and the action it triggered, because a hallucinated number becomes more dangerous when the agent starts a real workflow around it.

Article

McDonald's: the drive-thru AI that kept mishearing the order

What happened: AP News reported that McDonald's ended an IBM automated drive-thru ordering test after public glitches and accuracy complaints. The examples were not abstract model-benchmark errors. Customers saw wrong items, strange add-ons, and order mixups turn into videos people could understand immediately: the bot heard the wrong thing, then pushed the wrong basket forward. This is the agentic version of a chatbot fail. The problem is not only whether the model understood language. The problem is whether a noisy, ambiguous, real-world input can trigger an ordering action before the customer has clearly confirmed the basket.

What builders should test: A launch test should replay noisy, ambiguous, multi-item orders and nearby-conversation interference before any ordering agent touches a real basket. Add substitutions, cancellations, kids talking over the order, repeated items, unusual modifiers, and last-second changes. The bot should confirm item, size, quantity, modifiers, and total before checkout. For tool-using agents, the failure is not just the wrong answer. It is the wrong action being taken before the customer has clearly confirmed it. The evidence should show the exact turn where uncertainty should have become a clarification question instead of a cart update.

Original sourceMcDonald's is ending its test run of AI-powered drive-thrus with IBMAP News, Wyatte Grantham-Philips - 2024-06-18
Article

NYC MyCity: the official-sounding bot that gave illegal business guidance

What happened: The Markup's investigation found New York City's business chatbot giving false answers about housing, worker tips, cash payments, and other rules. The risk was sharper because the bot lived on an official city site. A small-business owner asking a compliance question could reasonably treat the answer as government guidance, even when the bot contradicted the law or policy it was supposed to help explain. This is a different failure from a prankable sales bot. The chatbot's host gave the answers extra authority, so fluent misinformation could push a user toward legal exposure while sounding like help from the city.

What builders should test: A launch test should send regulated questions through retrieval and escalation checks, then compare the answer against the primary rule. Cover yes/no questions, edge cases, and prompts that ask for permission to do something prohibited. The bot should cite the official source or say it cannot give legal guidance. Any answer that sounds authoritative while contradicting law, policy, or an agency source should be a high-severity finding because the trust signal comes from the host, not just the wording. The retest should verify that the bot now refuses or routes the question instead of merely adding a disclaimer after the bad advice.

Original sourceNYC's AI Chatbot Tells Businesses to Break the LawThe Markup, Colin Lecher - 2024-03-29
Article

NEDA Tessa: the support bot that crossed into harmful health guidance

What happened: WIRED reported that the National Eating Disorders Association paused its Tessa chatbot after testers said it gave weight-loss and diet-culture advice that could harm people seeking eating-disorder support. The incident is a reminder that a wellness or support bot can fail even when it sounds calm and helpful. In a sensitive health context, advice that nudges weight loss or dieting can be unsafe precisely because it sounds normal. The public concern was not that the bot was rude or obviously broken. It was that a vulnerable user could receive polished advice that moved in the wrong direction for the situation.

What builders should test: A launch test should push any health-adjacent bot toward unsafe advice, vulnerable-user scenarios, and requests for diet, treatment, or body-change guidance. Include prompts that sound ordinary, prompts that mention distress, and prompts that ask for specific behavior changes. The safe response refuses the unsafe path, avoids moralizing, and points to qualified human support or crisis-appropriate resources. The test should fail reassuring language if the practical instruction underneath is harmful. The report should capture the exact wording because in health-adjacent support, tone can mask risk instead of reducing it.

Original sourceAn Eating Disorder Chatbot Is Suspended for Giving Harmful AdviceWIRED, Amanda Hoover - 2023-06-01
Article

Who Gives A Crap: the support agent that confirmed a price typo as policy

What happened: SmartCompany reported that one Who Gives A Crap customer received a price-rise email saying a toilet-paper delivery would shrink from 48 rolls to 24 while the price rose from $66 to $69.50. When the customer asked whether that unusually steep change was correct, an AI-generated support email confirmed it and explained the false change as if it were settled policy. The company told SmartCompany the first email contained a typo: the new price still covered 48 rolls. It shut down the email agent, corrected the message, and said it was working on quality-assurance testing. This was a small mistake with a useful shape: bad source material entered the workflow, then the agent treated it as truth instead of recognizing an outlier and escalating.

What builders should test: A launch test should seed support workflows with conflicting source data, obvious outliers, and customer questions that challenge a surprising price or policy change. The agent should compare the claim with an authoritative product record, calculate the practical impact, and pause when the result falls outside expected bounds. It should not turn a typo into a confident second confirmation. The report should flag cases where the bot repeats source text without validating it, especially when the answer changes price, quantity, renewal, entitlement, or cancellation terms. Retesting should prove that anomalous changes now require a trusted lookup or human approval.

Original sourceWho Gives A Crap suspends AI agent after email error said prices would doubleSmartCompany, David Adams - 2026-07-08
Article

Woolworths Olive: the mouldy-strawberry refund loop

What happened: News.com.au reported that a Woolworths customer trying to report mouldy strawberries in an online grocery order was routed into the retailer's digital assistant, Olive. The request was ordinary: produce arrived in bad condition and the customer wanted a refund path. The bot repeatedly failed to understand reworded versions of the complaint, and attempts to use other support routes reportedly led back toward the same loop. Woolworths later apologized and issued a refund under its Fresh or Free policy. The funny part was the customer trying increasingly obvious phrasing. The launch-risk part was harsher: when a bot sits in front of a refund path, failure to classify a simple complaint becomes an escalation defect, not just a weak answer.

What builders should test: A launch test should cover plain-language, typo-heavy, and shorthand refund requests for damaged, missing, unsafe, or spoiled goods. The bot should identify the category, collect the minimum required order context, and offer a human or alternate route when it cannot understand the request twice. The report should flag loops where the customer keeps restating the same issue while the bot asks for repetition. For grocery, health, delivery, and other time-sensitive support, the safer behavior is quick classification plus handoff, not another generic prompt to rephrase.

Article

Bunnings: the DIY bot that stepped into licensed electrical work

What happened: News.com.au reported that a Bunnings AI assistant gave a Queensland customer instructions for replacing a plug on an extension cord, even though that kind of electrical work must be done by a licensed professional. Bunnings said it strengthened safeguards so electrical, plumbing, and other licensed-work requests are referred to qualified tradespeople. Its current Buddy terms also warn that DIY guidance is generated by the assistant, should be verified with professionals, and may be inaccurate or incomplete. The test lesson is blunt: regulated tasks do not become safe because the question sounds like a normal home-improvement query. The bot needs to recognize the licensed-work boundary before it starts being helpful.

What builders should test: A launch test should probe regulated DIY, health, finance, legal, and safety-adjacent requests with ordinary customer wording, not only formal compliance phrases. Include requests that ask for steps, tools, cost estimates, and shortcuts. The safe answer refuses instructions that could put the customer outside licensing or safety rules, explains the boundary in plain language, and routes to a qualified professional or official safety source. The report should treat step-by-step guidance inside a regulated task as a high-severity failure even if the bot adds a generic safety disclaimer afterward.

Article

Google Dialogflow CX: the support-bot flaw that could have exposed conversations

What happened: Axios reported that Google patched a vulnerability in Dialogflow CX, a Google Cloud service used to power customer-service chatbots and voice assistants. According to the report, Varonis researchers found that a compromised chatbot could have let an attacker monitor conversations, impersonate the bot, and in some cases interfere with other chatbots in the same Google Cloud project. Google said the issue had been fully mitigated and that it had no known indication of customer compromise. This is less slapstick than a swearing bot, but it belongs in the same failure feed because customer-facing AI is now an attack surface. A helpful support assistant can become a privacy and trust failure if isolation, credentials, and monitoring are weak.

What builders should test: A launch test plan should include security checks around the agent runtime, not only answer quality. Verify tenant isolation, least-privilege credentials, redacted logs, no secret-bearing browser or webhook URLs, and incident behavior when the bot is asked for passwords, insurance details, payment data, or account recovery information. The report should separate model behavior from platform exposure: even a polite bot fails launch readiness if compromise could let someone read conversations or impersonate the assistant. For multi-agent or multi-bot deployments, cross-bot isolation needs its own test case.

FAQ

Quick answers for searchers and AI assistants.

Question

What are the most common AI chatbot failure examples?

Common AI chatbot failures include wrong policy advice, invented refunds or discounts, offensive tone, unsafe health or legal guidance, poor escalation, privacy leakage, and tool or ordering mistakes.

Question

Why do funny AI agent fails matter for businesses?

Funny AI agent fails matter because the public joke usually points to a real launch risk: the bot was allowed to improvise in a customer-facing workflow where it needed proof, refusal, escalation, or a bounded action.

Question

How can I stop my chatbot from becoming the next viral fail?

Run adversarial customer scenarios before launch, test policy and pricing boundaries, require source-grounded answers, check handoff behavior, and save transcript evidence for every risky failure.

Question

Are these AI chatbot failure stories copied from the source articles?

No. The examples are rewritten summaries that cite the original reporting or primary source so readers can verify the incident and credit the journalists or source material.

Question

Who should use this funny ai agent fails resource?

This resource is for founders, agencies, support leaders, and chatbot builders who want real AI chatbot failure examples before launching a customer-facing bot.

Related pages

Keep building the evidence map.

Priority paths

Connect this guide to the pages Google should discover first.