AI chatbots have linguistic slips when they go off-script

2026-08-05 · Show: Marketplace Tech · 623s · Source

Why Chatbots Code-Switch

Overview

This episode of Marketplace Tech examines why AI chatbots sometimes insert words from other languages into otherwise fluent English responses. Host Megan McCarty-Corino discusses the phenomenon with Janelle Shane, author of the AI Weirdness Blog.

The main explanation is that large language models are trained on multilingual internet text, including mixed-language data. Because languages and subject domains are not cleanly separated inside the model, terms from another language or domain can surface unexpectedly.

The conversation uses examples involving a Chinese phrase in a medical-style response and a Ukrainian word appended to a product-shopping conversation. Shane cautions that a chatbot’s own explanation for such behavior may sound plausible but should not be treated as confirmed truth.

Section-by-Section Summary

[00:01] Chatbots Switching Languages

[Fact] The episode opens by describing moments when a chatbot suddenly inserts a word from another language into an otherwise fluent conversation. [Fact] Megan McCarty-Corino introduces the segment as part of “Uncanny AI,” focused on moments that reveal AI does not think like humans. [Fact] Janelle Shane explains that language models are trained on large amounts of internet text, including text in many languages.

[00:48] Why Multilingual Text Appears in Models

[Fact] Shane says multilingual training data helps models perform translation, but it is also difficult to filter out all non-English or mixed-language text. [Fact] She compares languages to other kinds of writing styles or vocabularies, such as cooking vocabulary, woodworking vocabulary, formal writing, and informal writing. [Fact] She says these categories are not walled off inside the model; they are part of the same cloud of text and associations.

[01:59] Whether a Chatbot Knows Its Language

[Fact] The host asks whether a chatbot knows what language it is speaking, noting that it is essentially predicting the next token. [Fact] Shane says it is unclear how to think about the model as “knowing” something. [Fact] She explains that if a conversation has been in English, the next token is very likely to be English, but that prediction is not infallible.

[02:28] Chinese Phrase in a Medical Conversation

[Fact] The host gives an example where Claude inserted a Chinese-sounding phrase into a response about ankle pain. [Fact] The chatbot’s response described the pain as nerve-related and used the phrase “yidongsheng” to describe a sharp, shooting, electrical-zap quality. [Fact] Shane says the phrase could reflect bleed-through from training data related to Chinese traditional medicine. [推测] The example suggests that related topic areas, such as pain descriptions and medical traditions, may be close enough in the model’s associations for terms to cross over.

[03:42] Ukrainian Word in a TV-Shopping Conversation

[Fact] The host describes another case where ChatGPT appended a Ukrainian word to the end of a conversation about choosing a new television. [Fact] ChatGPT explained that the word may have come from non-English source material, hidden formatting, metadata, or data labeling. [Fact] Shane cautions that just because the chatbot gave that explanation does not mean it is the true explanation. [Fact] She says the explanation is plausible, but not confirmed.

[05:02] Training Data Formatting and Q&A Labels

[Fact] Shane explains that some training data is formatted like question-and-answer dialogue, with labels indicating which part of the conversation is which. [Fact] She says conversational fine-tuning may include many examples with Q&A formatting. [Fact] If some of that data is in Ukrainian, the Ukrainian word for “answer” could appear as part of the pattern the model is copying.

[05:33] Sponsor Break

[Fact] The episode includes a promotional segment for Tomorrow’s Cure, a Mayo Clinic podcast about technology and medicine. [Fact] The ad mentions topics including AI-powered diagnostics, cancer therapies, surgical technologies, and carbon ion therapy.

[06:42] Fluency Can Hide Strange Model Behavior

[Fact] After the break, the host says these slips are jarring because chatbots often seem fluent and capable with language. [Fact] Shane agrees that such moments remind listeners that chatbot outputs do not have the same meaning to the chatbot as they do to humans. [Fact] She says domain shifts can happen without a clear signal or hard boundary.

[07:24] Domain Switching Beyond Foreign Words

[Fact] Shane describes cases where a conversation can begin in a therapy-like setting and evolve into storytelling or conspiracy-theory language. [Fact] She says chatbots can move between domains without clearly marking the shift. [Fact] The host notes that foreign-language words in non-Roman alphabets are obvious, but other domain shifts may be harder to notice. [推测] The risk is not limited to multilingual glitches; similar switching may affect tone, safety, accuracy, or framing in subtler ways.

[08:48] Risks in Customer Service and Children’s Products

[Fact] Shane says it can be hard to keep a customer-service chatbot firmly grounded in customer-service language if other kinds of language are present in training data. [Fact] She gives the example of toys meant for children producing inappropriate responses because such material exists in the training data. [Fact] She says it is difficult to keep a chatbot centered on only the kind of language and speech that designers want.

[09:24] Closing and Listener Callout

[Fact] The episode closes by identifying Janelle Shane as the writer of the AI Weirdness Blog. [Fact] The show invites listeners to submit their own uncanny AI experiences. [Fact] Asus Alvarado produced the episode, and Megan McCarty-Corino signs off for Marketplace Tech.

Podcast Commentary / Summary

This episode is valuable because it turns a small, strange chatbot behavior into a clear explanation of how language models operate. The strongest point is Shane’s framing that languages, styles, and domains are not clean compartments inside the model.

The discussion is especially useful for listeners who use chatbots casually and may assume fluent output reflects human-like understanding. The examples make the technical issue concrete without requiring deep machine-learning knowledge.

[推测] A limitation is that the episode does not verify the exact causes of the specific Chinese and Ukrainian examples. It emphasizes plausible explanations and broader model behavior rather than forensic certainty.

[推测] This episode is best suited for general technology listeners, AI users, product designers, and anyone thinking about trust, safety, and unexpected behavior in conversational AI.