Study: Nearly Half of Medical Advice Given by AI Is Misleading

Bloomberg reported on Wednesday (April 15) that researchers from the United States, Canada, and the United Kingdom evaluated five leading artificial-intelligence (AI) platforms—ChatGPT, Gemini, Meta AI, Grok, and DeepSeek; each had to answer 10 questions across five health-care categories.

The researchers reported this week in BMJ Open that about 50% of the answers given by these AI chatbots were rated "problematic," and nearly 20% of those problems were "very serious."

The results highlight concerns about growing public reliance on generative-AI platforms, which hold neither the licensure required to give medical advice nor the clinical judgment needed for diagnosis.

The safety risks cannot be ignored. OpenAI data show that at least 200 million people ask ChatGPT health-related questions every week—a huge number. If so many answers are problematic, AI medical advice poses real safety risks to many people.

Over the past two or three years I have now and then used AI to look things up, and I have found its error rate high. But in my own medical queries the error rate was not high, perhaps because, having some medical knowledge, I choose keywords very precisely.

In fact, some friends who used the medical AIs I recommended got very different-quality results from mine. The reason is that most people do not know how to query professional medical knowledge; they lack even the ability to ask the right questions, and so in the end do not get high-quality replies.

The AI wave is sweeping every industry; medicine has been affected, and ordinary people's lifestyles are quietly changing. But AI exhibits the trait of "meeting the strong strongly and the weak weakly": in every field only professionals can interact with AI in depth, and only professionals can judge whether AI's output is right or wrong.

Whether AI can replace professionals is a matter of debate. I lean toward the view that AI cannot replace truly valuable expert talent, but it can indeed replace many mid-level professionals, because their work was not very good to begin with. AI's output is riddled with errors, but we know from real work that humans are no different. I recall a survey of American oncologists a few years ago that found more than half were not up to standard overall.

High-level medical research cannot be done by AI alone; medical practice requires many innovative attempts, and every advance in medicine has been forged through vast clinical experience. So in the future, human doctors with creative thinking and hands-on ability should remain scarce talents.