Can Artificial Intelligence Replace Doctors? Exploring the Future of AI in Medicine

Can Artificial Intelligence Replace Doctors? Exploring the Future of AI in Medicine

By Akhilesh Kuppili·
AI & Data SciencePublic HealthNew Innovations

Original: Will AI Replace Physicians in the Near Future? AI Adoption Barriers in Medicine

Rafał Obuchowicz, Adam Piórkowski , Karolina Nurzyńska, Barbara Obuchowicz, Michał Strzelecki , Marzena Bielecka

Introduction

Artificial intelligence is becoming increasingly common in medicine. AI can analyze medical images, organize patient notes, write notes, answer queries, and advise physicians about diseases. In some tasks, algorithms were found to be comparable to physicians. However, the success in one task does not imply that AI can replace physicians in all activities related to patient care. This study evaluates how close AI is to replacing physicians in general and what obstacles prevent the adoption of medical AI.

Two flavors of AI dominate the conversation in the paper – convolutional neural networks (CNN) and large language models. The former excel at analyzing medical images, such as X-rays, CT scans, MRI, and mammography. The latter understands and generates natural language and can be applied to extract information, answer questions, or write notes.

AI has the promise to make healthcare efficient, but performing a specific task is not the same as tending to a complete patient. The researchers posed the question of whether medical algorithms have the required clinical and decision-making power to substitute physicians and what challenges prevent such replacement from happening.

Methods

The paper is a narrative review, which means that the authors synthesized the information from prior studies rather than designing and implementing one study.

The researchers evaluated the performance of AI in medical image analysis, large language models (LLMs), diagnostic and treatment advice compared to physicians’ performance, patient-physician interactions, physical exams, medical procedures, addressed biases, reliability, ethical implications, legal accountability, and other factors.

The authors compared the relative ease of automation of specific tasks, such as scanning films and writing notes, and the ones requiring nuanced human judgment.

Results and Limitations

AI Performance in Medical Image Analysis

The narrative review identified that AI can perform remarkably well if the task is well-defined and falls within the limited scope. For instance, in radiology, convolutional neural networks (CNN) can assist clinicians by highlighting possible bleeding, blood clots, tumors, or other objects. Moreover, algorithms can measure the size of lesions, track the changes between scans, and prioritize cases based on the degree of urgency.

For instance, in some studies, AI for identifying intracranial hemorrhage demonstrated a sensitivity between 90% and 100%. Similarly, algorithms for reading mammography and detecting breast cancer showed outstanding performance.

AI can be particularly beneficial for repetitive tasks because it rarely makes errors due to fatigue. For instance, the same algorithm can analyze thousands of x-rays without getting tired. Such consistency can make clinicians more efficient by standardizing the measurements and reducing intra-rater variability.

Performance of Large Language Models in Healthcare

LLMs similarly demonstrate excellent language understanding and can be employed in answering medical questions, summarizing patient notes, preparing clinicians’ notes, and other natural language processing tasks. In particular, such algorithms can allow clinicians to delegate note-writing and documentation tasks to AI.

However, current LLMs suffer from “hallucination,” meaning that they might make things up. In particular, in medical communication, wrong information can be detrimental, therefore erroneous answers in algorithms presenting as certain might lead to severe consequences.

Challenges Specific to Medical AI

A significant challenge for medical AI, in particular to CNNs, lies in the out-of-distribution generalization, meaning that the performance of AI decreases if a case is presented, which was not seen during training. This is particularly concerning for medical applications, where patient demographics, hospital settings, and scanners can vary between institutions.

Out-of-distribution cases can lead to AI providing incorrect answers with a high degree of confidence, which might be particularly concerning if such algorithms operate independently from physicians.

Moreover, medical CNNs can be “adversarial,” meaning that specific inputs can trick the algorithm into giving incorrect responses. In particular, a minor change in pixels, which would be imperceptible to a human, can cause the AI to identify a healthy brain as glioma. However, the researchers note that such attacks usually require extensive testing and do not occur naturally. Similar to the challenges with generalization, such vulnerabilities are particularly concerning if clinicians rely on unproven algorithms.

Other challenges include the inability of current AI to handle many tasks related to physical exams, human judgment, and legal and ethical considerations.

Physical Examinations

Examinations currently present a significant limitation for medical AI. For instance, physicians can assess the patient’s posture, breathing pattern, expressions, movements, tenderness, temperature, and many other factors, which are currently beyond the capabilities of AI.

In addition, many medical specialties, such as surgery, dentistry, or emergency medicine, require a physical approach to patients. Therefore, the ability to examine a patient is central to medical practice, while current algorithms cannot perform all the tasks of physicians. Moreover, some aspects of the physical exam, such as fine motor skills or fine-touch sensation, cannot be handled by robots.

Human Judgment, Legal, and Ethical Considerations

Most of the tasks related to legal and ethical considerations in medicine currently belong to clinicians. Moreover, despite advances in AI, such tasks as diagnosing patients still require human judgment. The researchers note that even if AI handles most of the diagnostic and treatment planning tasks, clinicians remain accountable for patient care. Thus, in cases of erroneous recommendations, it might be hard to attribute responsibility to a specific entity, such as a physician, hospital, or an software developer. Consequently, ethical and legal issues present significant challenges to adopting medical AI.

Limitations

Since the study was a narrative review, the authors did not perform one experiment to evaluate the performance of AI algorithms. Instead, they analyzed several studies in the field. Therefore, the performance of AI might vary depending on the task, algorithms, patients, and other factors. Moreover, the field is evolving, and future developments might alter the findings of the review.

Conclusion

This narrative review suggests that the prospect of replacing physicians with AI is limited. Substituting doctors with algorithms can be particularly challenging due to clinicians’ need for judgment, accountability, and the physical side of medical practice. However, in specific tasks, such as analysis of medical images, lesion detection and measurement, documentation, and other similar activities, AI can be remarkably efficient. The review highlights the need to continue exploring the role of medical algorithms and improving generalization, training better safety measures, and other changes. Instead of acting as a replacement for physicians, AI can serve as an assistant by handling routine tasks, allowing clinicians to focus on more complex cases and provide better care for patients.

Akhilesh Kuppili

Akhilesh Kuppili

Founder & Co-President