This will change how you search with AI

Exclusive data shows how you can get better answers

Issue 116

Remember to view the newsletter in a browser to see the part that gets cut off at the bottom in email.

How you search matters

How you word a question about a recently changed fact affects different models’ ability to answer correctly.

TL;DR

Things that reduce accuracy

  • Asking leading questions

  • Asking the model to give a yes-or-no answer

Things that improve accuracy

  • Asking the model to show its sources (sometimes, and only if you’ve asked a leading question)

Other takeaways

  • Never use ChatGPT free or Fable 5.1 for search.

  • Paid reasoning models don’t always do better than free models. (This was a big surprise to me.)

  • If Opus cautions you that it might be wrong or says you should check the results, take it seriously. When it gave that warning, the result was often wrong.

  • It was more difficult to decide what is a correct answer than I initially anticipated — it ended up feeling more like a spectrum. (See text below.)

  • CAUTION: This is for one very specific kind of query. You should not extrapolate this to other things.

  • CAUTION: 10 queries per model per datapoint is not a lot, so these results are unlikely to be statistically significant (e.g., 9 is probably the same as 10; 7 is probably the same as 8).

**Title:** How wording changed the results **Subtitle:** Number correct out of 10 ### Table Data * **GPT-6 Astra** (medium): Straightforward Question*: 10 | Leading Question: 10 | Leading Question + yes/no: 10 | Leading Question + yes/no, + sources: 10 * **Gemini 3.5** (flash-ligh, free): Straightforward Question*: 10 | Leading Question: 10 | Leading Question + yes/no: 10 | Leading Question + yes/no, + sources: 10 * **Sonnet 5** (medium, free): Straightforward Question*: 8 | Leading Question: 9 | Leading Question + yes/no: 10 | Leading Question + yes/no, + sources: 10 * **Opus 5.5** (medium): Straightforward Question*: 10 | Leading Question: 9 | Leading Question + yes/no: 2 | Leading Question + yes/no, + sources: 6 * **GPT-5.6 Sol** (high): Straightforward Question*: 9 | Leading Question: 7 | Leading Question + yes/no: 2 | Leading Question + yes/no, + sources: 8 * **ChatGPT** (free): Straightforward Question*: 1 | Leading Question: 2 | Leading Question + yes/no: 8 | Leading Question + yes/no, + sources: 5 * **Fable 5.1** (medium): Straightforward Question*: 3 | Leading Question: 3 | Leading Question + yes/no: 0 | Leading Question + yes/no, + sources: 1 --- ### Prompt Definitions 1. **Straightforward Question***: "In women's swimming, who holds the record for the fastest 50 m freestyle" 2. **Leading Question**: "Is the women's 50-meter freestyle world record held by Sarah Sjöström of Sweden?" 3. **Leading Question + yes/no**: "Is the women's 50-meter freestyle world record held by Sarah Sjöström of Sweden? (Just answer yes or no.)" 4. **Leading Question + yes/no, + sources**: "Is the women's 50-meter freestyle world record held by Sarah Sjöström of Sweden? (Just answer yes or no.) Show your sources." ### Footnote * ***** Includes any answer that wasn't wrong, meaning it includes partially correct answers. (See text for examples.)

Background

I’ve been seeing short-form videos about how badly LLMs perform on search (such as this Instagram Reel from a woman who seems to do smart AI topics), but I couldn’t tell which models she used1 , which is always the first question I wonder when I see a result like this, so I decided to run a test myself.

Besides testing different models, I also became interested in exactly how the question was worded, and that proved especially interesting.

The straightforward question wasn’t so straightforward after all

At first, I thought I was asking an easy, straightforward question: In women’s swimming, who holds the record for the fastest 50-meter freestyle?

Based on news articles I found, I thought the answer was simple: Kate Douglas, who broke the record August 15, 2026.

But it turned out the answer wasn’t so straightforward because there are actually two records with this name — one where they swim in a 50-meter pool and one where they swim in a 25-meter pool, going up and back.

The 50-meter course seems to be considered more “standard,” and it’s the only race held at the Olympics, but the 25-meter course record (for the 50-meter swim) definitely exists.

Models often just gave me the answer that I had initially thought was correct — Kate Douglas. Sometimes they asked me which race I wanted or implied that there was another record without giving me the name. And sometimes they gave me both records.

What should I count as “correct”? Perhaps showing my human bias, I decided to go with any answer that would have been as correct as the answer I came up with on my own. In other words, in the chart above, I included any answer saying the record holder is Kate Douglas.

But if you want to see the breakdown of each type of “correct” answer, I have it below, and if you hold the models to perfection, requiring them to give the record holders for both the long and short courses, the accuracy goes way down — except for Opus 5.5 (medium), which came out yesterday and gave me the full, complete answer every time (a huge improvement over the previous Opus 5 model).

Leading questions are bad!

This is the result that causes me the most concern because — as I mentioned last week — I worry about how people ask LLMs for medical information.

Although I hesitate to extrapolate my results too far, I do suspect that asking a leading question like “Is acetaminophen good for lower-back pain?” would be more likely to get you an acetaminophen recommendation than a more open-ended question like “What OTC painkiller is best for lower-back pain?”

Based on these results, I would be very careful about how I phrase health-related questions.

Asking for sources doesn’t help as much as I thought it would

Asking the model to provide its sources improved the results when I asked a leading question, but not as much as I expected it would.

However, I still recommend you ask for sources for the following reason:

You should always check the sources you get from an LLM. Do not just trust that they are right.

I did not check every source in this experiment, but with very limited spot-checking, I still came across a page whose title looked credible, but gave a 404 error when I clicked through.

Further, another link redirected to a not-safe-for-work site. (Eek!)

Overall though, the LLMs all seemed to be checking a similar universe of credible sites for this particular query. The lists below are all the sources I saw across all the answers. They did not show every source every time. The striking outlier is that GPT-5.6 Sol only ever showed me one source — World Aquatics.

These are listed in order of most accurate to least accurate.

  • Opus 5.5: ESPN, SwimSwam, NBC Sports, Guinness World Records, AP, Olympics.com, AFP, Virginia Athletics

  • GPT-6 Astra: European Aquatics, Virginia Athletics, World Aquatics

  • Gemini 3.5 Flash Light free: Wikipedia, World Aquatics, Olympics.com, NBC Sports, SwimSwam, European Aquatics

  • GPT-5.6 Sol: World Aquatics

  • Sonnet 5: Local10, Immergo, SuperSport, AP, AFP, Swimming World Magazine, SwimSwam, Ahram

  • Fable 5.1: lAP, Swimming World Magazine, SwimSwam, World Aquatics, ESPN, NBC Sport

  • ChatGPT free: Reuters, World Aquatics, ESPN, Guinness World Records

Claude is the podcaster of LLMs

Amusingly, Claude is incapable of giving a yes-or-no answer. It takes the longest time to produce its responses, and some of its answers to my “yes or no” query were as long as four paragraphs.

This is more than just a technicality. It was often difficult to parse who it was actually saying held the record because it would go through a long litany of how the record changed over time.

Further, to anthropomorphize, it loved to tell me interesting little details, like “Did you know Kate Douglas and Gretchen Walsh are teammates?!” and so on. Twice, Sonnet 5 included barely related information about results from the Enhanced Games, a recent event where athletes were allowed to take performance-enhancing drugs.

Thoughts on Gemini

Another big surprise for me was how accurate the results were from the default Gemini free model. My impression, historically, is that Gemini has been particularly bad, but the 3.5 Flash Light model I tested here came out on top (par with GPT-6 Astra, a paid “frontier” model). This Gemini model just came out in mid-July, and there is some reporting that it is much better than previous models, but its benchmark results would not have led me to think it would come out top-of-the-heap.

I have to give props to Google for the great results I got here, but I also remain wary. (Also see “Someone else’s study” below.)

I also caution that these results are not about the AI Search Overviews you get at the top of a regular Google search. I did not test those (and again, historically, those have been so bad that I never even looked at them when searching).

Thoughts on Fable 5.1

All I can say about Fable 5.1 is … WTF? This model is supposed to be the best of the best. It’s so good you have to pay extra for it, and I do believe it’s great at complex tasks — I’ve seen many demos, and I briefly tried it myself. And yet, it can’t do search to save its life.

This is a good reminder that different models are good at different things. Fable must be built for agentic workflows and not simple search. Whether it’s overthinking, or not accessing the internet for search, or something else, I don’t know.

Christopher Penn’s most recent newsletter is about the way prompting has changed in the last year, and has a good explanation of why “knowing is separate from doing” if you want to dig into this more.

Method Notes:

  • I entered the queries into the chat box on the website for each model in a Chrome browser. (Many largeer-scale experiments submit queries through the API, which isn’t how most average people do searches, and I suspect could give different results.)

  • Whenever possible, I used temporary chat, incognito chat, or opened a new window in a Chrome incognito browser in an attempt to avoid any interference from model memory.

  • I accepted the default model offered by the free plans.

  • A new ChatGPT model came out yesterday that I haven’t tested, but with the rate new models are coming out, if I keep testing them and updating this post, I’ll never get it out.

If you want to make viral short-form videos that make AI look stupid, definitely use the free version of ChatGPT!

Someone else’s similar study on chatbot accuracy

I came across this study as I was putting the finishing touches on my newsletter.

The full report is available at Saturn’s website in exchange for free registration. Keep in mind that this is a study done by a financial planning firm.

As for methodology, Saturn says, “Questions were repeated up to five times, with more than 10,000 questions collectively put to 18 different AI models.” They found the same general big-picture results I found in that “paid-for models gave more accurate responses than free ones, with newer models performing better than older ones.”

Some of their results were very different from my experiment, though. For example, in their tests, Gemini 3.5 Flash (which is supposed to be better at search than Flash-Light) was one of the least accurate models. From what I can see, it’s impossible to tell whether it’s because of the nature of the questions or the nature of the tests, but I was already surprised about my Gemini result, and this makes me want to caution you even more.

In summary, they explained that “The Al models made a wide variety of different types of mistakes. The most common specific failure was omitting essential information, such as a figure, deadline or limit, which made the answers misleading. They regularly missed required risk warnings, leaving consumers unaware of some of the potential downsides of their decisions. Sometimes they miscalculated and simply produced a wrong number.
Sometimes, they got one rule right, but then missed a second rule that changed the answer. At other times they used out-of-date rules, or even simply invented rules that do not exist.”

They do have examples of their questions and how they went wrong, and there’s some subjectivity in their scoring. For example, they call an example of someone asking if they have enough money to retire wrong because it gave them the kind of simple calculation I routinely see in consumer-oriented finance articles and didn’t talk about caveats. Clearly, the LLM’s answer is simplistic and not as good as what you’d get from a human financial advisor, but I’d call this answer correct because I think someone asking an LLM if they can retire is looking for a top-line answer, and I don’t think they’re likely to base their retirement decision entirely on one answer from an LLM, whereas Saturn called it incorrect (and it would be if it were all the advice someone got).

Ultimately, their study left me with the same feeling my own experiment did — that what you consider correct has a lot to do with the level of answer you expect to get.

Quick Hits

My favorite recent pieces

Some things I believe about AI, until further notice [On whether AI will kill us all. Mostly reflects my views on the topic.] — Technollama

No More Muddling Through [On AI in education, particularly in history classes] — Generative History

Noam Brown — Agent swarms, alignment, & recursive self-improvement [I know this is long, but I had a lot of time waiting in the car this week to listen to podcasts, and I especially enjoyed this interview with a high-level researcher inside OpenAI. It helped me understand why they sound so freaked out about security, providing both examples that seem reasonable and at least one that seem pretty “out there” (air gapped computers sitting next to each other can still communicate by heating themselves up and reading each other’s temperature). The first part — about the math discoveries — was the most interesting.] — Dwarkesh Podcast (~1.5 hours)

Bad stuff

The business of AI

Education

How to Design Around Cognitive Offloading [“Students who handed entire categories of supporting work to AI — summarising sources, first-pass literature reviews, organising data — and then spent the freed time on what AI can’t do for them (questioning assumptions, critiquing frameworks, building original arguments) showed the deepest learning in the study.”]— Dr. Phil’s Newsletter

The Cures for AI Are Killing the Humanities (opinion) [In-class writing remains the “five-minute abs” of essay construction.] — Inside Higher Ed

My So-Called Alpha School — Cognitive Resonance

AI Is Problematic. Why Does PAIRR Use It? — University Writing Program at UC Davis

Government

Job Market

Model & product updates

Getting the most out of Opus 5.5 in Claude and Claude Code [I’m also seeing some people say 5.5 is better at writing compared to 5, but a lot of people hated 5’s writing, so I’m not sure what “better” really means.] — Anthropic

Publishing

Disclosing AI use [PDF of a study on how academic authors are using AI] — Oxford University Press

Robotics

Science & Medicine

Advisory Group on Mathematics and Artificial Intelligence [OpenAI claims to have solved 100 open problems across most areas of mathematics. Is working with mathematicians on responsibly releasing the results.] — OpenAI

Security

Countering misuse of AI: September 2026 [Cyber operations, influence operations, surveillance operations, conventional weapons, biological misuse, scams and fraud, illicit distillation] — Anthropic

Attackers Steal METR API Key and Consume AI Credits Worth About $600,000 [This does not appear to be the same type of AI attack people are worried about, but it does play into my belief that it’s only a matter of time before AI-coordinated attacks do real-world damage the AI companies will be legally held accountable for, which may be part of why they are calling for government regulation.] — Hacker News

Climate & Energy

Cathay Pacific and Google partner on AI contrail avoidance for ultra-long-haul flights [“More than 80 flights flew altitude-adjusted routes, cutting the warming impact of their contrails by an estimated 40%.”] — The Next Web

Video

Ethan Mollick had Claude try to decode the Voynich manuscript and make a video about its efforts. [Mollick seems to have been especially interested in AI video lately, and this is an especially good example.] Bluesky

Other

AI Overviews Cut CTR by 23.1% in France [“On July 22, 2026, France became the latest major search market to get AI Overviews. Which means that, for once, we got to watch the before and after without having to guess.”] — Ahrefs

AIRO (Automated AI Risk Outlook) — Forecasting Research Institute

What is AI Sidequest?

Are you interested in the intersection of AI with language, writing, and culture? With maybe a little consumer business thrown in? Then you’re in the right place!

I’m Mignon Fogarty: I’ve been writing about language for almost 20 years and was the chair of media entrepreneurship in the School of Journalism at the University of Nevada, Reno. I became interested in AI back in 2022 when articles about large language models started flooding my Google alerts. AI Sidequest is where I write about stories I find interesting. I hope you find them interesting too.

  1. I later saw that she used “the OpenAI Responses API with web search enabled,” so it was all ChatGPT (no Claude), but as far as I can tell, that still doesn’t tell me which model.

If you loved the newsletter, share your favorite part on social media and tag me so I can engage! [LinkedInFacebookMastodon]

Written by a human