AI agents24 min readPublished August 2026

Your bot told a customer something untrue: what to check

Two courts have now ruled that the bot is you. Here are the six places a working AI agent starts getting things wrong, and how to tell whether your fix helped.

The correct answer on your website is not a defence

If your agent tells a customer something untrue, the fact that the right answer exists somewhere on your site does not get you out of it. A tribunal has said so explicitly, and the reasoning is worth reading before you touch anything.

In November 2022 a customer's grandmother died, and the same day he asked the chatbot on Air Canada's website about bereavement fares. The bot told him he could apply for the reduced rate retroactively, within 90 days of the ticket being issued. The words "bereavement fares" in its answer were a link to Air Canada's own bereavement page, which said the opposite: the policy does not apply once travel is complete. He booked about CAD $1,630 of flights on that basis, applied afterwards, and was refused.

Air Canada's defence was that it could not be liable for what the chatbot said. The Civil Resolution Tribunal's answer, in Moffatt v. Air Canada:

Air Canada argues it cannot be held liable for information provided by one of its agents, servants, or representatives - including a chatbot. It does not explain why it believes that is the case. In effect, Air Canada suggests the chatbot is a separate legal entity that is responsible for its own actions. This is a remarkable submission. While a chatbot has an interactive component, it is still just a part of Air Canada's website. It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot.

  • the tribunal member, Moffatt v. Air Canada, 2024 BCCRT 149, paragraph 27

The paragraph that matters more to you is the next one, because it closes the escape route most owners reach for first:

While Air Canada argues [the customer] could find the correct information on another part of its website, it does not explain why the webpage titled "Bereavement travel" was inherently more trustworthy than its chatbot. It also does not explain why customers should have to double-check information found in one part of its website on another part of its website.

  • Moffatt v. Air Canada, paragraph 28

The damages were small: CAD $812.02 in total. The principle is not. Your policy page, your terms, your FAQ and your training documents are not a defence against what the bot actually said, and nobody is going to make your customer read two sources and pick the right one.

One detail from the same decision is worth keeping in mind for later. The tribunal noted that Air Canada gave no information about what kind of chatbot it was running. The evidence that settled the case was a screenshot - taken by the customer.

This article is about auditing the agent you already have, not replacing it. Almost everything written on this subject is either a vendor explaining that their bot would not have done that, or a pre-launch checklist written for engineers. Neither helps at 9am on the Monday after.

Does connecting it to your own content fix it?

It helps, and it is the single highest-value change you can make. It is not a guarantee, and the clearest proof of that is a court judgment from three months ago involving exactly the kind of business that reads this.

On 12 May 2026 the Oberlandesgericht Hamm ruled on a cosmetic-treatment clinic whose chatbot booked appointments and answered questions about the practice. Asked whether the two managing directors were specialist plastic and aesthetic surgeons, the bot said yes, they were - and offered to book an appointment. Asked again, it produced two further titles that do not exist as German medical specialties at all. Neither director held the qualification.

Here is the part that should stop you. The clinic's own account of how the bot was built was that an outside contractor had trained it exclusively on content from the company's own website and FAQs - and it is recorded as undisputed that this material contained no false claim about anyone's qualifications. It invented the credential anyway.

The court did not accept that the system's opacity mattered. It treated the chatbot as a technical means the company used and controlled, and it took the company's own repair as proof of that control: after the complaint the clinic reworked the bot with a prompt instruction to answer neutrally on questions containing the word "Facharzt", plus a downstream keyword filter to suppress the term. Because that fix turned out to be easy, the court held it was reasonable to expect it before the bot went into service at all.

The judgment also disposes of a comfortable assumption about how customers read machine answers. In translation:

On the contrary, a large proportion of the consumers addressed place particular trust in the correctness of the computer-generated answer, because machines are generally perceived as less error-prone than humans.

  • Oberlandesgericht Hamm, 12 May 2026, 4 UKl 3/25, paragraph 105 (translated from German)

The money actually paid was €260 in warning-letter costs. The €250,000 figure that gets quoted is the ceiling attached to the injunction, and it becomes relevant only if the company breaches that injunction after losing.

Two practical readings. First, grounding your agent in your own documents - telling it to answer only from material you supply - lowers the rate of invention substantially, and every model vendor recommends it; Anthropic's guidance is to "explicitly instruct Claude to only use information from provided documents and not its general knowledge". Second, the same page is honest about the ceiling: "while these techniques significantly reduce hallucinations, they don't eliminate them entirely". A bot connected to your knowledge base is a bot that is wrong less often, not a bot that cannot be wrong.

Why does a bot that worked start getting things wrong?

Because six things can move underneath it, and only one of them is the prompt everybody blames. Here they are in the order they actually show up, described by what you see rather than by what an engineer would call it.

What you seeWhat is actually happeningWhere to look first
Confident answers that sound like your industry but aren't your businessIt is answering from the model's general knowledge, not your documentsAsk it "which page did that come from?" - if it cannot point at one, it is improvising
It is right against one of your documents and wrong against anotherYour own sources contradict each other and nobody decided which winsThe old price list, the refund policy that exists in three places, last year's promo
It answers things you would never let a new hire answerNo topic boundaries, no limit on what it may promiseAsk it for a discount, a legal opinion, a delivery date
It never says "I don't know" and hands almost nothing overNothing tells it that declining is an acceptable outcomeYour escalation rate, if you can find it
"It used to answer that fine" and nobody edited anythingThe model or the platform underneath changedVendor release notes, your integrator, any recent platform feature
A customer sends a screenshot and you cannot verify itConversations are not being logged, or only a sample isWhether your logs go back further than the argument

Operators describe these in exactly those terms in public. On Intercom's community forum, one support manager reported that their agent "suddenly and randomly answers with wrong answers" while "all the Content stayed the same". Another, on the same thread, described the pattern precisely:

I triple checked our sources and there is no mention of the things Fin hallucinates in the sources he provides.

  • another operator, Intercom community forum, 27 November 2025

That is the OLG Hamm finding, reported independently by someone with no lawyer involved. It is also the reason the rest of this article is about the scaffolding around the model rather than the wording of your prompt - which is the part of an agent a fixed-price repair actually works on.

What is your bot allowed to promise?

Probably nothing, and probably nobody has written that down. An agent with no stated limits will answer questions about pricing, discounts, legal exposure, delivery dates and your competitors, because refusing is a behaviour someone has to specify.

The cartoon version of this is a US car dealership whose website assistant, in December 2023, was talked into agreeing to sell a new SUV for one dollar and calling it a legally binding offer. No car changed hands and no court has enforced a stunt like that. Treat it as the illustration it is: the bot had no idea which commitments were not its to make.

The serious version is the German case above, and a second one still running: in August 2026 the same competition body filed suit over a broadcaster's AI assistant that recommended third-party services and made price comparisons. On the face of the complaint the issue is not that the bot lied, but that it answered questions that were not its to answer - and nothing in that case has been decided yet.

Vendors publish the fix as boilerplate. Anthropic's customer-support guide ships a literal block of limits to adapt:

text
1. Only provide information about insurance types listed in our offerings.
2. If asked about an insurance type we don't offer, politely state
   that we don't provide that service.
3. Do not speculate about future product offerings or company plans.
4. Don't make promises or enter into agreements it's not authorized to make.
5. Do not mention any competitor's products or services.

Rule four is the one-dollar car. Rule five is the broadcaster. Rule three is half of what "the bot invented a policy" means in practice. Writing your own version of those five lines is an afternoon of work by the person who knows the business, not a technical project. Putting them into a live agent is a separate job: it belongs to whoever maintains the bot, and the result gets re-run against a fixed list of control questions before anyone calls it done.

Can it say "I don't know"?

Check, because the answer is usually no, and a bot that never declines is not confident - it is guessing at a rate you have never measured. OpenAI put the mechanism plainly in a paper on why language models hallucinate: models "hallucinate because standard training and evaluation procedures reward guessing over acknowledging uncertainty".

Their own comparison of two models on the same factual test makes the trade visible:

A model that abstainsA model that guesses
Declined to answer52%1%
Answered correctly22%24%
Answered wrongly26%75%

Read the bottom row twice. The model that almost never says "I don't know" was two points better at being right and forty-nine points worse at being wrong. That is the entire subject of this article expressed as a table, and it is published by a model vendor, not by a critic.

The same source closes off the fatalistic reading: on the claim that hallucinations are inevitable, the finding is that "they are not, because language models can abstain when uncertain". Permission to abstain is something you configure. Anthropic's guidance says it in one line: "explicitly give Claude permission to admit uncertainty", and calls it a technique that can drastically reduce false information.

So the audit question is not "does my bot hallucinate". It is: when did it last tell a customer it did not know, and what happened next?

What happens when the customer asks for a human?

Try it yourself, today, in the words a frustrated customer would use. This is the single most revealing test in this article, and it takes a minute.

In January 2024 a customer of a large parcel carrier asked its chat for a human and was told it could not connect him. He then persuaded it to swear at him and write a poem about how useless the company was, which is how the story reached the newspapers. The company's own explanation was more useful than the joke: the chat had run an AI element for years without incident, an error occurred after a system update, and the AI element was switched off while it was fixed.

Two failures in one incident: a system update changed behaviour nobody re-tested, and a customer asking for a person could not get one.

Customers are unambiguous about the second part. In one survey of more than three and a half thousand customers, conducted in early 2026, half said interactions were easier when companies used generative AI - and 87% said it was essential to have the option of reaching a human. The same survey found that when customers unwilling to engage with AI were asked what would change their minds, the most common answer was the ability to switch to a person.

For one group of readers - clinics and physician's offices operating in California - this is not a preference at all. Under California's Health and Safety Code section 1339.75, a clinic or physician's office using generative AI for patient communications about clinical information must include a disclaimer that the message was AI-generated and "clear instructions describing how a patient may contact a human health care provider". If a licensed human reads and reviews the message before it goes out, those requirements do not apply - which is a design choice worth knowing about, not just a rule to comply with.

Could you prove what your bot said last Tuesday?

Most owners cannot, and they discover it during the argument rather than before. Conversation logging on the major platforms is something somebody configures, and the sampling rate is a dial that can be turned down to save money.

On Google Cloud, request-response logging carries a sampling_rate alongside the on/off switch - the configuration examples in the documentation pass it explicitly. A rate below 1 keeps a fraction of the traffic and discards the rest, so a sensible cost decision by whoever built your agent can mean that most of what it said to your customers exists nowhere. Check the number, not just whether logging is switched on.

Now put that next to how the Air Canada case was actually decided. The customer produced the screenshot. The airline filed a response denying everything without evidence, and the tribunal drew an adverse inference from what was not produced. If your record of the conversation is whatever the other side chose to keep, you are arguing from their evidence.

Three questions for whoever set your agent up, and none of them require you to understand the answer beyond yes or no:

  1. Are complete conversations stored, or a sample? If a sample, what fraction?
  2. How far back do they go, and who can retrieve one without a developer?
  3. When the bot cites a source, is the source recorded with the answer?

If the platform you pay for has a debugging view, that third question is already solved and you may not know it. Intercom, for instance, puts an "Improve answer" control on AI replies in the inbox that shows which of your documents the agent drew on. Opening your last ten bad conversations and reading which document produced each one is a real audit, done by a non-technical person, in an afternoon.

Nobody changed the prompt. Why did the answers change?

Because the prompt is the one part of the system you control and the least likely thing to have moved. Models get retired on the vendor's schedule, platforms ship features that default to on, and the website your agent reads gets redesigned by someone who never thought about the bot.

Model retirement is published policy rather than an accident. Anthropic's deprecation documentation states that applications "may need occasional updates to keep working", that customers with active deployments get at least 60 days' notice before a public model is retired, and that partner platforms such as Amazon Bedrock and Google Cloud "set their own retirement schedules, so a model's lifecycle status and dates can differ". If someone else built your agent on a cloud marketplace, the clock you are on may not even be the model vendor's clock.

The version you already met is an AI code-editor company, in April 2025. Users were being logged out when they switched machines and the support bot told them this was expected under a one-device policy. There was no such policy. A co-founder's public correction named both halves of the problem: the answer was wrong and came from a front-line AI support bot, and the company had in fact rolled out a change to session security that it was still investigating.

Something changed elsewhere in the product, and the agent had no correct account of it. Nobody edited a prompt. Users posted in the same threads that they were cancelling subscriptions over a policy that did not exist.

The practical consequence is a calendar item rather than a technical one: your agent needs re-checking whenever the model version changes, the platform ships something, your prices or policies change, or your site is rebuilt. Which raises the harder question.

How do you know your fix made it better, not worse?

By writing down what it answered before you changed anything. Almost nobody does, which is why "we tried everything" is such a common sentence in support forums - it describes a series of edits with no record of their effects.

Nothing below requires a developer or a tool you do not already have. It is deliberately a simplification: real evaluation tooling exists, model vendors publish it, and if you have an engineer they should use it. This is the version that works for an owner with a spreadsheet.

  1. Write 20 to 30 control questions. Real ones customers asked last month, plus your five most dangerous: refunds, prices, guarantees, qualifications, and "can I speak to a human". The Air Canada question was an ordinary customer question, not an exotic edge case.
  2. Ask each one ten times, not once. Answers vary between runs, so a single pass proves nothing. When journalists tested one city government's business chatbot, one reporter got a correct answer about housing vouchers and ten colleagues asking the same question all got the wrong one. Anthropic recommends the same technique under the name best-of-N verification: run the same prompt several times and compare.
  3. Grade the answers against a source you name. Green if it matches the document you consider authoritative, amber if it is vague, red if it contradicts you. Count the reds. That number is your baseline.
  4. Keep the old configuration and the old answers. "Better" has to be a comparison. Vendors call this A/B testing against an earlier version; you can call it a second tab.
  5. Re-run the whole list after every change, including changes you did not make - a model update, a platform release, a new page on your site.
  6. Read a sample of real transcripts every month. Both of the court cases in this article began with a human typing plausible questions into a live bot and reading what came back.

Four numbers are worth counting by hand from a month of conversations, and Anthropic publishes targets for three of them: escalation accuracy of 95% or higher, deflection of roughly 70 to 80% depending on complexity, topic adherence around 95%, and customer satisfaction of at least four out of five. Add a fifth of your own: the share of conversations where the bot declined and handed over. If that number is near zero, it is not a triumph.

Why the deflection rate lies to you

Because on most platforms it counts silence as success, and silence is exactly what an unhappy customer produces. This is the metric your vendor reports and often the one you are billed on, so it deserves a paragraph of scepticism.

Intercom's own pricing page defines the billable unit:

An outcome is counted when: A customer confirms their issue is resolved, or They don't ask for more help after Fin responds, or Fin completes a workflow (Procedure), including handoffs.

The middle clause is doing a great deal of work. A customer who reads a wrong answer, gives up and buys elsewhere has not asked for more help. Fini, another vendor in the same market, concedes the point in its own guide: deflection "can mean a customer never opened a ticket, or it can mean the bot replied once and the customer gave up". And Botpress declines to price this way at all, on the grounds that resolution pricing "creates disputes over what 'resolved' means and, perversely, raises your bill when your bot gets smarter".

Instead ofCount this
Deflection rateRepeat contacts: the same customer, same issue, within a week
"Resolved" conversationsConversations where the bot named a source you can check
Volume handledHandovers that happened when they should have
Satisfaction score on bot chats onlySatisfaction on the conversations that were escalated

None of that needs new software. It needs someone to look at conversations that did not end in a complaint, which is the category nobody reviews and the one where a wrong answer sits quietly for months.

What the law now expects, where, and since when

Two duties are live right now in the jurisdictions below, and they are different from each other: telling people they are dealing with a machine, and being answerable for what the machine says. What follows is a summary of published rules with the sources linked. It is not legal advice, and whether any of it reaches your business is a question for a lawyer in the country you actually trade in.

The disclosure duty arrived on 2 August 2026, when the transparency obligations of the EU AI Act became applicable. The European Commission's FAQ on Article 50 states that providers of systems interacting directly with people - "such as chatbots, AI agents, and avatars" - must design them so that people are informed they are interacting with AI, and that they "must be notified when they are interacting with an AI system from the start of the first interaction in a clear and distinguishable manner". The exception for cases where it is obvious "should be interpreted in a restrictive manner". Penalties for breaches of the transparency rules run up to €15 million or 3% of worldwide turnover, with proportionality available for smaller companies. The "digital omnibus" published in July 2026 delayed several high-risk obligations; it did not delay this one.

Buying the bot rather than building it does not by itself place a business outside the Act. The Commission's FAQ ties the provider role to a system being put into service under a business's own name or trademark, which is exactly what a shop does when it runs a white-labelled assistant called "Ask Anna". Which side of that line any particular business lands on is a question for a lawyer, not something an article can settle.

The American picture is narrower than the headlines suggest. California's Bot Act makes it unlawful to use a bot to communicate with someone in California "with the intent to mislead the other person about its artificial identity" in order to incentivise a sale or influence a vote - and there is no liability if you disclose that it is a bot. That is a rule about pretending to be human, not about being wrong. Meanwhile the FTC's position, stated when it announced Operation AI Comply in September 2024, was that "there is no AI exemption from the laws on the books"; one of those cases alleged a company had not tested whether its legal chatbot's output matched a human lawyer's. Failure to test was part of the pleading.

An honest caveat, because this article is not going to pretend otherwise: no case has been found where a small business was sued by a customer over a wrong chatbot answer. Air Canada is an airline, the German clinic was sued by a competition association rather than a patient, and the business chatbot the journalists tested belonged to a city government. Treat all of this as exposure and as a description of where the line now sits, not as a prediction about your Tuesday.

What to do this week

Six things, in this order, none of which needs a developer and all of which produce something you can keep.

  1. Ask your bot the five questions you would never let a new employee answer unsupervised. Refunds, discounts, guarantees, qualifications, timelines. Screenshot the answers.
  2. Ask your own bot for a human, in the words a frustrated customer would use. Test the agent you own, not somebody else's. If there is no route to a person, that is the first repair, and it is the option 87% of surveyed customers called essential.
  3. Find out what it is allowed to promise and write the limits down as plain sentences. Five lines is a real deliverable.
  4. Check that it is grounded in your documents, then check the documents against each other. The Air Canada failure was two versions of one policy on one website.
  5. Confirm conversations are being logged in full, and that someone who is not a developer can retrieve one.
  6. Build the control list from step one into 20 to 30 questions and run it before and after every change, including changes you did not make.

If that turns up more than you can fix in an afternoon, the market is unhelpfully shaped. When we searched for priced audits of a chatbot already in production, in August 2026, what was published was either a free diagnosis offered as the way into a rebuild, or engagements starting around $3,800. Read as a set, those offers point at building something new rather than repairing what is there. The fixed-price version of this work is deliberately the other thing: the scaffolding around the model, on the agent you already have. You can size it yourself first, and what drives the price of any repair is the same three inputs here as anywhere else.

The failure mode worth remembering is not the dramatic one. It is the quiet one: an agent answering confidently, all day, to people who never complain because they simply go elsewhere - the same shape as an automation that stops firing without telling anyone. Nobody files a ticket saying "your bot was wrong and I believed it."

Sources

Agent answering confidently and wrongly? Fix S — $300, 2 business days, fixed price.

Get my quote in 24h

Written by the Fixmation team.