Legal Translation

Can AI Translation Hallucinate? What Legal and Medical Teams Need to Know Before They Trust It

Sep 27, 2026
8 minutes
Can AI Translation Hallucinate

One in three ChatGPT-4.0 translations of pediatric discharge instructions into Haitian Creole contained a potentially clinically significant error, according to a 2024 study in Pediatrics. Professional translators made such errors in only 8.3% of cases. So yes: AI translation hallucination, where fluent output drifts from the source, can happen in legal translation and medical translation, not just in chatbot answers.

Not every error in that study was strictly a hallucination, and the wider evidence is mixed. A 2026 preprint found that frontier models preserved medical meaning well across eight languages. So here’s what this guide does: it explains why errors happen, compares medical and legal risk, and gives you a three-tier framework, the consequence ladder, for deciding which documents AI can handle.

What Is AI Translation Hallucination?

Definition: AI translation hallucination is output that reads fluently but is not faithful to the source text. The system adds content that was never there, drops content that was, changes a number or negation, or produces text with no real connection to the original. The fluency is what makes it dangerous.

It’s a specific case of hallucination in artificial intelligence, where a model produces confident output that isn’t grounded in its input. A 2025 taxonomy by Wu et al. sorts translation hallucinations into untranslated content, wrong target language, extraneous additions and repetition. It builds on a classic split from Guerreiro et al. (2023) between oscillatory hallucinations, which repeat a word or phrase, and fully detached ones.

The Four Failure Modes

Keep in mind that these examples are illustrations, not documented incidents.

Additions. The model inserts text the source never contained, such as an invented “unless agreed in writing” in a contract or a made-up reassurance in a patient leaflet.

Omissions. A clause or warning silently vanishes, like a missing indemnity or a dropped safety instruction.

Number and negation distortions. A “not” disappears, or “15 mg” becomes “50 mg”. An ICASSP 2025 Chinese-English study found error rates of up to 20% on large-unit numerals for Llama 3.1 8B, a small open model, so don’t generalise that figure to frontier tools.

Detached output. You get fluent text that has nothing to do with the source, or text delivered in the wrong language.

Why Fluent Output Hides the Problem

Fluency is the trap. In that same Pediatrics study, Google Translate and ChatGPT-4.0 performed about as well as professional translators for Spanish and Portuguese. A reviewer who reads only the target text sees polished prose and assumes it’s right.

Why Does AI Translation Hallucinate?

AI translation hallucinates because these systems predict plausible next words rather than checking meaning against the source. When text is unusually long or short, full of rare terms, ambiguous, or in a language with little training data, plausibility can beat faithfulness, and the model fills the gaps with fluent guesses.

How Neural and LLM Translation Work

Classic neural machine translation systems learn from huge collections of paired sentences, then generate target text piece by piece. Large language models do something similar after training on general text. Neither keeps a checklist of the source, so when they’re unsure, they just keep writing something that sounds right.

Low-Resource Languages and Detached Hallucinations

A 2026 preprint by Anyaegbuna et al. tested GPT-5.1, Claude Opus 4.5, Gemini 3 Pro and Kimi K2 on 704 medical translation pairs in eight languages, including Haitian Creole. Every model scored above 0.92 on a measure of how closely meaning matched the source, and there was no significant gap between high- and low-resource languages.

That sounds reassuring. But it’s still a preprint, and in our view a meaning-match score can miss a single flipped number. Detection is a separate problem, too: Benkirane et al. (EMNLP Findings 2024) found that hallucination detectors lost much of their edge on low-resource pairs across 16 language directions. So low-resource errors may also be harder to catch.

Stress tests show the ceiling. HalloMTBench (Wu et al., 2025) is built from deliberately hard cases, so it measures worst-case behaviour, not a real-world base rate. On that set, hallucination rates across 17 models ranged from about 33% (GPT-4o-mini) to 84% (Seed-X-PPO-7B).

Source Length, Rare Terminology and Missing Context

The HalloMTBench authors name source length, model scale, linguistic bias and language mixing as triggers, and the length effect cuts both ways: very short texts struggle as well as long ones. Legal and medical text adds defined terms, drug names and jurisdiction-specific concepts. And a model that doesn’t have the whole document can’t know what a term meant on page three.

How Is a Hallucination Different From an Ordinary Mistranslation?

A mistranslation is a wrong rendering of content that’s visibly there, such as a poor word choice, whereas a hallucination invents or detaches content with no basis in the source. Mistranslations tend to read awkwardly. Hallucinations read perfectly, which makes them far harder to catch, and omissions are harder still.

Mistranslation Hallucination Omission
How it looks Wrong or awkward word Fluent text not in source Content silently missing
Cause Ambiguity Plausibility over faithfulness Model skips content
Detectability Moderate Low Very low
Consequence Misread term Invented obligation or dose Missing warning or clause

 

Why Hallucinations Slip Past Bilingual Spot Checks

A spot check usually asks whether the target text sounds right. That catches a mistranslation, but it won’t catch an omission, because you can’t see what’s missing unless you compare against the source line by line. Added text often matches its surroundings, which makes hallucinations nearly as hard to spot. The fix is a side-by-side comparison by a bilingual reviewer.

What Does AI Translation Hallucination Look Like in Medical Translation?

In medical translation, AI translation hallucination looks like a fluent patient document that changes a dose, drops a warning or adds reassurance the source never gave. Discharge instructions, consent forms and leaflets are the most exposed, because patients act on them directly and rarely have anyone to check them.

Discharge Instructions, Consent Forms and Dosage Errors

These documents mix plain instructions with dense clinical detail. A consent form carries legal weight and medical risk at once, and its readers, by definition, can’t check the English original. The dangerous errors are the small ones: a dose off by a decimal place, milligrams swapped for milliliters, a dropped “do not take with”. Each of them looks perfectly normal on the page.

What the Haitian Creole Study and 2026 Validation Research Show

In the 2024 Pediatrics study (Brewster et al.), 20 standardised discharge instructions were translated into Spanish, Brazilian Portuguese and Haitian Creole. For Haitian Creole, potentially clinically significant errors appeared in 8.3% of professional translations, 23.3% of Google Translate output and 33.3% of ChatGPT-4.0 output.

The sample was small, the study tested ChatGPT-4.0, and no 2025 or 2026 replication has surfaced. So newer models may do better, and the 2026 preprint suggests the language gap may be narrowing. Still, neither study shows that machine output is safe unreviewed. The risk depends on language, model and document.

Regulatory Pressure: HHS Section 1557 and EU MDR Labelling

Regulators already expect human oversight. Under the 2024 HHS Section 1557 rule, 45 CFR 92.201(c)(3) requires a qualified human translator to review machine translation by a covered health entity when the text is critical to the rights, benefits or meaningful access of a person with limited English proficiency, when accuracy is essential, or when the source contains “complex, non-literal or technical language”. Confirm current enforcement with your compliance team.

In Europe, EU MDR Article 10(11) requires labels and instructions for use in the official languages of the Member State where a device is sold, and the Commission updated its language-requirements summary in August 2025. For this kind of content, professional medical translation services and ISO-certified translation give you a documented process, not just an output.

What Does AI Translation Hallucination Look Like in Legal Translation?

In legal translation, AI translation hallucination looks like one altered clause, date or negation that changes liability, meaning or admissibility. Contracts, witness statements, evidence and court filings are most at risk, because a single changed word can shift who owes what, or what a witness said.

Contracts, Witness Statements, Evidence and Court Filings

The familiar version is AI inventing case citations, which a February 2026 practitioner’s guide from the National Center for State Courts (NCSC) covers. Translation errors are quieter. Instead of a fake case, you get a real contract with a clause that says something slightly different.

A defined term has to be rendered identically every time, and a negation has to survive. A concept like “consideration” may also have no equivalent in the target legal system, so a model may substitute something that only sounds close.

Real cases show meaning going wrong, though none is proven hallucination. In March 2025, India’s Supreme Court criticised mistranslated filings, including a tribunal order where “reinstated” became “re-establishment”, and noted that machine translation alone is insufficient. Back in August 2024, a judge of the same court reportedly described an AI rendering of “leave granted” as “chhutti sweekaar” (holiday approved). These are anecdotes, not error rates, and we found no 2025 or 2026 ruling blaming a fabricated translation on AI hallucination.

Court and Bar Guidance, Including NCSC Resources

NCSC’s Machine Translation Guide (2025) says courts should never use machine translation for court events or to convey legal or procedural information. Its June 2025 article on AI in court translation adds that regular human review by certified translators is essential. Orange County’s Spanish translations were about 80% accurate and usable, yet every document still got professional review. That’s one county’s result, not a general rate.

A Wisconsin bill (SB 357), debated in March 2026, would have let courts use AI for interpretation, but it failed to pass when the session ended. The State Bar of Wisconsin and the ACLU of Wisconsin objected, and press coverage cited Stanford researchers documenting cases where “trial” became “test”.

Confidentiality and Privilege When Using Public AI Tools

Pasting a privileged contract into a public AI tool may expose confidential data. NCSC, writing for courts, recommends internal, firewalled AI over public cloud models, and the same logic applies to law firms. When accuracy, confidentiality and admissibility all matter, certified legal translation is the safer route.

How Can Legal and Medical Teams Decide Which Documents AI Can Translate?

Decide by consequence, not convenience. Ask who will rely on the translation, whether a wrong word could harm someone, whether the document is regulated, and whether the language is low-resource. Low-stakes internal material can use AI for gist, but high-stakes or regulated documents need a human translator and independent review.

The Consequence Ladder: A Three-Tier Risk Framework

Tier Approach Examples
1: Gist only Raw AI, labelled unverified Internal triage email, background reading
2: AI plus post-editing AI draft, full linguist review against the source Routine patient leaflet, internal policy
3: Full human plus independent review Human translator and second reviser Clinical trial consent form, contract, sworn affidavit

 

NCSC likewise suggests starting AI with low-risk material. And you should move any document up a rung if the language is low-resource, the content is regulated (an HHS Section 1557 trigger or EU MDR labelling), or the reader can’t check it.

Medical Versus Legal Risk at a Glance

Document Typical hallucination Who is hurt Tier
Discharge instructions Changed dose or unit, dropped “do not take with” Patient and family 2, or 3 if low-resource
Consent form Dropped warning, added reassurance Patient 3
Contract Invented or lost clause, flipped negation Signatory 3
Court filing or witness statement Altered term or date Party or witness 3

 

Triage Questions

Three of these four questions come straight from the HHS Section 1557 rule, and the fourth is our addition. Is the document critical to someone’s rights or care? Is accuracy essential? Is the language complex, non-literal or technical? Is the language pair low-resource? One yes pushes you toward a human. Keep in mind that the rule itself binds covered health entities only.

Red-Flag Checks

Before any Tier 2 document goes out, compare numbers, units and dates, every negation, dosages (with a clinician), names and defined terms, and paragraph or clause counts against the source.

In practice, back-translation into English looks reassuring, but the same model can repeat the same error in both directions. A documented translation quality management process is a stronger safeguard.

What Does a Safe AI-Plus-Human Translation Workflow Look Like?

A safe workflow pairs AI speed with qualified human review. A subject-matter linguist post-edits the machine draft against the source, an independent reviewer checks it again, terminology is controlled, and every step is logged, so no translation reaches a reader unverified.

Machine Translation Post-Editing by Subject-Matter Linguists

In machine translation post-editing (MTPE), a trained linguist compares the AI draft against the source and fixes errors. For legal and medical work, that linguist has to know the field. Otherwise they may smooth a sentence without noticing it changed a dose.

Independent Second Review, Terminology Control and Audit Trail

Add a second reviewer who didn’t do the first pass, and give both reviewers a glossary. NCSC recommends court-specific terminology and regular certified review, and an audit trail shows who changed what.

Certifications That Signal a Controlled Process

ISO 17100:2015 covers translation services, including an independent reviser. ISO 18587:2017 covers post-editing of machine translation, and a revised edition is at draft stage. ISO 13485:2016 covers quality management for medical devices. Confirm current versions on iso.org.

Questions to Ask a Translation Provider Before You Sign

Trust AI Translation Only as Far as You Can Verify It

Fluent is not the same as faithful. That’s why AI translation hallucination is so hard to catch in legal translation and medical translation.

The risk is manageable, though. Put every document on the consequence ladder, check numbers, negations and omissions every time, and put a qualified human behind anything where a wrong word could harm a patient or lose a case.

Not sure which tier your documents need? Request a quote for your documents and have a professional team review them through a process you can verify.

    Stay Ahead in Global Communication

    Translation insights and industry trends — delivered to your inbox every week.