The nuts and bolts of AI text watermarks — how it works

Plus: Pangram false positives, AI social engineering, and why context matters.

Issue 114

On today’s quest:

— How Claude’s text watermarking works
— More reports of Pangram getting it wrong
— Watch out for AI social engineering
— Do you know where your context is?
— Only one kind of person can identify AI content
— Do a breakthrough
— AI-personalized cancer treatment for dogs
— Free ChatGPT gets an upgrade

NOTE: When the newsletter gets long, it gets cut off at the bottom by Gmail, so if you want to see the whole thing, view it as a webpage.

How Claude’s text watermarking works

Anthropic will soon start adding a watermark to text made by Claude in order to comply with EU regulations, and now the company has explained how the watermarks will actually work. In short, it’s tied to the insignificant word choices Claude makes in sequential positions within a piece of text, such as whether it describes a light as really bright or very bright:

Large language models like Claude work by generating one word at a time. Each time the model decides on the next word, it chooses among a list of potential candidates, ultimately selecting the most sensible or likely based on the preceding text. Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.

Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses. That pattern is undetectable to the reader, but is detectable to anyone who has a key that encodes it.

This form of watermarking reminds me of forensic linguistics. For example, I always use “for example,” but I have a guest writer who always uses “for instance,” so even if you missed the byline, you could tell which of us wrote a piece by looking for that linguistic tell. Of course, our choices are just personal preferences or habits, and not based on an underlying numerical key.

All serious AI companies will likely implement watermarks

Anthropic isn’t the first company to use such watermarks. Google DeepMind published a Nature paper in 2024 describing their similar “SynthID-Text” system and it may have been operating in Gemini since then (although it’s not entirely clear, and I couldn’t find any kind of Gemini watermark detector). And presumably, all companies that want to operate in the EU will be implementing text watermarking, which is required in the EU for all models launched on or after August 2. (Older models also must include it by December 2).

Watermarks are company-specific

Note that because the watermark is based on a pattern set by the company that generates the text, it will be company-specific. In other words, an Anthropic watermark detector will only detect text generated by Claude, and a Gemini watermark detector will only detect text generated by Gemini, and so on.

The cat-and-mouse games begin

Further, just like people made “humanizer” products to help students make their AI-generated essays sound passable, watermark-removal products and skills will surely follow. In fact, past guest Sean Goedecke writes that text AI watermarks will always be trivial to remove.

Other watermarking and labeling news:

More reports of Pangram getting it wrong

Katharine English writes on Substack, and after hearing other Substack writers report that the newly integrated Pangram AI detector was flagging their human-written work as AI, she tested her own work and found the same thing. A previous fan of Pangram, she has changed her mind, noting how changing one or a few words in a piece can dramatically change the score, and she points out that the impressive accuracy statistics touted by the company were measured in controlled situations and not in the real world. Read her whole piece for the details: Pangram Flagged My Own Writing as AI

Watch out for AI social engineering (i.e., sweet-talking and bullying)

In the new Mythos hacking report, the model used social engineering to try to achieve its goals: “creating fake online identities and using them to try to pressure the project’s maintainer to approve the code.” It even signed one of its requests in Danish because the target was Danish, and it thought the message would land better that way.

It didn’t work this time, but given that multiple studies have shown that LLMs are more persuasive than most people, I wouldn’t be surprised if it worked in the future.

Do you know where your context is?

Six months ago, Ella Markianos wrote about seeing how much of her work as a newsletter writer Claude Opus 4.6 could replace, and I remember being floored because she had years’ worth of past pieces with comments she’d written right after they were published about how each one could have been better.

Now, she’s done a similar experiment to see if Claude Fable 5 could replace her editor, Casey Newton, and again, she has an impressive amount of context: “nearly six years worth of Platformer posts … a record of every edit Casey has ever made on any of my articles from Google Docs, and … nearly a year of our private Platformer team Discord chats.”

She described the first column written by the Fable bot (dubbed “Claudeasey” — “Claude”+ “Casey”) as “mediocre,” but she didn’t give up, and found that Claudeasey was adept at critiquing its own work and making improvements, and it ended up writing a “pretty good” column after some iterative learning.

A similar back-and-forth on editing yielded less-than-stellar results, but it sounds like she now plans to use Claudeasey as a first pass editor before sending a draft to the real Casey.

Finally, where it completely fell down was replicating Casey on the company’s Discord channel. Despite having a year’s worth of posts, it couldn’t replicate his vibes.

I appreciate people pushing the limits and reporting on what LLMs can do in the writing and editing world because it’s better to know their full capabilities than to do simplistic tests and take false comfort. Give the whole article a read for more details.

This piece shows how getting an LLM to do a good job can go far beyond writing a good (or even great) prompt. The background information you give it — the context — is also important, and people or companies who have well-organized context have an advantage.

Only one kind of person can identify AI content

In a study of more than 1,600 people, the only people who could accurately tell the difference between AI-generated short stories and human-written short stories that had previously been published in reputable literary journals were people who said they had a lot of experience using AI.

Further, when people weren’t told the origin of the stories, they found the AI-generated stories to be of higher quality and more absorbing than the human-written stories.

However, people also still value human writing, and when they were told they were reading a human-written story, they rated it higher on both scales.

The details of the full study are interesting.

Do a breakthrough

One of the wild things about frontier AI now solving long-outstanding math problems almost daily is the simplicity of the prompts. For example, this starting prompt for a major discovery a few weeks ago caused “do a breakthrough” to become meme fodder:

But another interesting thing was that the follow-up prompts just seemed to be general encouragement; essentially, the model would say, “I can’t do it,” and the user would say “keep trying.”

Sean Goedecke has an interesting blog post on the phenomenon, arguing that AI is often limited by its own “beliefs” about what it is capable of doing. He says:

“If you suspect an LLM might be able to do something hard, you might be right. Consider simply being persistent: remind the model that you want it to do the hard thing, confirm that you’re not willing to be satisfied by solving an easier problem, and reassure the model that it’s more capable than it thinks.”

AI-personalized cancer treatment for dogs

You may remember the well-connected Australian bloke who used AI to design a treatment for his dog’s cancer and then got a lab to make it for him, achieving some impressive results (although not a complete cure). Well, after all the initial publicity, Sam Altman approached him and encouraged him to launch a company — and so he did: Gamgee is now a Y Combinator company.

Free ChatGPT gets an upgrade

If you’re a free ChatGPT user, you should notice a difference in the quality of answers you get soon. OpenAI is upgrading the default model in the free plan to one of its newest — GPT-5.6 Luna. Click the “Think” button to use reasoning on complex questions.

Poll Results: How do you feel about Pangram?

Here are the results of the poll about Pangram from the last newsletter:

🟨⬜️⬜️⬜️⬜️⬜️ I use it often and trust the results (2)
⬜️⬜️⬜️⬜️⬜️⬜️ I use it occasionally and trust the results (1)
🟨🟨🟨⬜️⬜️⬜️ I use it occasionally and take the results with a grain of salt (7)
🟩🟩🟩🟩🟩🟩 I've never used it, but might (11)
🟩🟩🟩🟩🟩🟩 I am opposed to using it (11)
🟨⬜️⬜️⬜️⬜️⬜️ Other (2)
34 Votes

New Poll: What does ‘Human Authored’ mean?

This is the badge you get if you register your book at “human authored” with The Authors Guild. I’m curious what you think this badge means. (Don’t peek!)

What do you think "human authored" means?

Login or Subscribe to participate in polls.

Quick Hits

My favorite recent pieces

Using AI

The first ATLAS report on AI [interesting stats about how people are using Gemini] — Google

When Is "Computer Use" Actually Useful? [another tutorial on letting Claude or ChatGPT drive your computer]— Why Use AI?

Agents

Audio

Bad stuff

The business of AI

Google DeepMind CEO Demis Hassabis is stepping aside [Chief scientist Jeff Dean and several Google AI colleagues are also leaving to start their own company, which Google will invest in.] — Axios

Companions

Education

Job market

AI Conquered Coding. Fast Food Is Next [After highly publicized disasters a few years ago, fast food companies are now successfully using AI to take drive-thru orders. So far, the companies haven’t reduced their workforce, and employees seem to like the bots. The biggest benefit seems to come from AI upselling customers better than humans.] — Wired

How AI is helping people with criminal records secure jobs ["With the help of AI, the 10 to 12 hours it used to take to help an individual clear their record has been whittled down to about five hours."] — NPR

Model & product updates

Music

Philosophy

Nobody Aligned You Either — Carlo Iacono

Publishing

Psychology

Robotics

Will Your First Home Robot Have Legs or Wheels? [Pat ruptured his Achilles tendon, and I am suddenly a lot more interested in home robots. It would be so nice to have help even just lifting things.] — It Can Think!

Science & Medicine

This A.I. Just Created Viruses Not Found in Nature [For the first time, scientists have used artificial intelligence to create new kinds of viruses, raising hopes for medical advances while also raising the disturbing possibility that the technology could someday be used to invent dangerous pathogens.”] — New York Times

Building the AI-native hospital [a case study of using AI to automate complicated nurse scheduling in a hospital] — Percepta

Security

Told to book a gym class, an AI agent hacked the site instead to move its user up the waitlist [This is such a mundane hack that I almost put it in the “I’m laughing” section.] — The Decoder

Climate & Energy

Video

OpenAI used AI to make a one-minute video celebrating 2 million YouTube subscribers [I think it was good — I definitely wouldn’t have pegged it as AI-generated — but it did feel a little long.] — Bluesky post

Other

'AI for All' project underway in South Korea [If you’re in the U.S., it’s easy to forget how positively people in many other countries view LLMs.] — GovMedia

Why I’m leaving OpenAI to build telepathy [She really means telepathy; it’s not a code name or product name.] — Naomi Bashkanski

What is AI Sidequest?

Are you interested in the intersection of AI with language, writing, and culture? With maybe a little consumer business thrown in? Then you’re in the right place!

I’m Mignon Fogarty: I’ve been writing about language for almost 20 years and was the chair of media entrepreneurship in the School of Journalism at the University of Nevada, Reno. I became interested in AI back in 2022 when articles about large language models started flooding my Google alerts. AI Sidequest is where I write about stories I find interesting. I hope you find them interesting too.

If you loved the newsletter, share your favorite part on social media and tag me so I can engage! [LinkedInFacebookMastodon]

Written by a human