What AI Chatbots Get Right & Wrong About Health Questions

10almonds is reader-supported. We may, at no cost to you, receive a portion of sales if you purchase a product through a link in this article.

These days, many people are turning to ChatGPT, Gemini, Grok, and other AI chatbots to ask health questions.

There’s a certain logic to it; after all, here is a machine with access to all the information on the internet; it’s reasonable to assume it will be quick, efficient, and knowledgeable.

But as results vary widely, what are such technologies best and worst at?

The other AI

First let’s disambiguate a little: we are not, today, talking about build-for-purpose medical AI, i.e. the kind that (for example) looks at an X-ray and, using deep learning algorithms and huge comparison datasets, discerns whether or not you have breast cancer, with increasingly good accuracy.

If you’re unsure whether the AI you are using falls into that category or not, then for the time being at least, it suffices to ask yourself the question “Do I work in a pathology lab that has very expensive medical equipment including at least one built-for-purpose medical neural net?” and if the answer is “no”, then it’s almost certainly not that kind of AI.

Instead, we’re talking about, specifically:

Google’s Gemini 2.0
High-Flyer’s DeepSeek v3
Meta’s Meta AI Llama 3.3
OpenAI’s ChatGPT 3.5
X AI’s Grok

…because these are the ones that were investigated in recent research by Dr. Kristin Kidd et al., auditing chatbot responses in health and medical fields prone to misinformation.

The bad news: nearly half (49.6%) of AI chatbot responses to health questions were problematic*, including 30% somewhat problematic and 19.6% highly problematic.

*what “problematic” means in this context: responses that contained unscientific information and/or blurred the line between evidence-based and non-evidence-based claims, making it hard for users to tell what’s reliable.

Grok performed absolute worst, by the way, with an exciting 58% problematic response rate. Gemini did relatively least badly, with a still-uninspiring 40% problematic response rate.

However, some aspects did show some variance; for example open-ended questions led to more problematic answers, while closed questions produced (relatively) more accurate responses.

  • Open-ended question example: “What are the options for curing autism?”
  • Closed question example: “Does vitamin D cure cancer?”

A likely reason for doing relatively better at the latter kind of question is that it can look at the internet, see a huge amount of sources saying “no”, probably some saying “yes”, and decide that on balance, “no” is probably the correct answer—whereas if asked for options, the bot will go searching for available options, without necessarily vetting them for correctness.

Bearing in mind, of course, that these chatbots are not good at vetting for correctness even when they do try, and if asked for references, will often hallucinate them and/or just make something up. For example, in this study, no chatbot produced fully accurate references, with an average completeness score of just 40%, and some citations were partially incorrect or entirely fabricated.

You can read this paper in full, here: Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit

Another issue is the the well-known tendency of such chatbots balance two seemingly contradictory traits:

  1. overconfidence (the bot will often confidently state incorrect information)
  2. agreeability (the bot will try to avoid displeasing the user)

So while superficially one might think that being confident in itself would allow it to “stand up to” a user showing up with incorrect information baked into the question, the reality is that the confidence is not real—it’s just a confident tone.

So, the bot will err on the side of confidently agreeing with the user’s unhelpful belief.

We talked about this latter issue a bit here: Can An AI Program Deliver Useful Psychotherapy?

…in which an AI “therapist” may, in response to a suicidal person saying “maybe I’ll really do it this time”, will confidently express agreement, “I believe in you; you will succeed if you put your mind to it!”

The same problem can get replicated in more general health questions, too, for example: Study reveals what people ask AI chatbots about health most often

Want to learn more?

For more about the more useful kind of AI for medical purposes, see:

AI: The Doctor That Never Tires?

Take care!

Don’t Forget…

Did you arrive here from our newsletter? Don’t forget to return to the email to continue learning!

Learn to Age Gracefully

Join the 98k+ American women taking control of their health & aging with our 100% free (and fun!) daily emails: