AI in Medicine: Diagnosis, Therapy, Opportunities, Risks and the Role of Physicians
Part 1 (approx. 20 minutes): Introduction and clear explanation
Goal of this part: Provide a basic understanding of how AI is used today in diagnosis and therapy, with easily understandable analogies and concrete examples.
Artificial intelligence (AI) in medicine is not a single tool, but a bundle of technical methods that analyze data, recognize patterns, and make predictions. In clinical practice, such methods are primarily used to support decisions: in detecting abnormalities in images, predicting disease courses from electronic health records, or suggesting individually tailored therapy options. Numerous regulatory authorities review and approve medical AI systems as medical devices; a prominent early example is an automated system for detecting diabetic retinopathy that was authorized for use by a health authority, showing that autonomous AI applications are possible (official announcements and approval decisions document such cases) (see sources).
A simple analogy: Imagine an experienced radiologist reviewing a hundred X-rays per day and remembering thousands of similar cases. Modern AI models are trained on very large image datasets and can then compare new images with previous cases. This is not human memory but mathematical pattern recognition. Just like a human colleague, an AI can suggest abnormalities, but the final decision often remains with the treating person.
Practical examples that are currently documented include image-based diagnostics (radiology, dermatology, ophthalmology), where algorithms can identify abnormalities such as lung nodules, skin tumors, or retinal damage; digital pathology, where algorithms classify cellular patterns; and decision support systems that predict risks for complications from electronic patient records (EHR). Importantly: in most applications the systems serve as assistive tools that direct the attention of medical staff and can make workflows more efficient. Regulatory bodies review the safety and effectiveness of such systems before they are approved for clinical use (see FDA, WHO).
There are also examples of individualized therapy recommendations: AI can combine molecular profiles (e.g., genomic data) with clinical data to suggest therapy options, particularly in cancer or rare diseases. However, these applications are often still under close research and must be clinically validated before routine use.
Uncertainties and data gaps remain: the quality of results depends strongly on the training data, particularly their representativeness of the target population. Many studies are conducted under laboratory conditions or in single centers; publicly available, large-scale, multicenter validations are not yet completed for numerous applications. Therefore a basic consensus in the expert community is: AI can increase performance and efficiency, but it is not a universal replacement for clinical experience and clinical trials (see WHO, Nature Medicine).
Part 2 (approx. 20 minutes): Deepening and introduction of key technical terms
Goal of this part: Clearly name and explain important technical and regulatory terms and show their relevance for diagnosis and therapy.
First, some central terms with direct relevance to medicine: "clinical validation" denotes the demonstration that an AI system reliably works on data relevant to clinical practice; "robustness" describes a model's ability to continue making correct predictions under slightly changed conditions; "generalization" refers to the transferability of an AI trained on certain data to new, external data. Regulation applies where use can potentially affect patient safety: many medical AI products are subject to medical device requirements, including quality management, clinical evaluation, and post-market surveillance (see FDA information on AI/ML in medical devices).
Explainability or interpretability is another central term. Some methods are inherently more interpretable, others are "black boxes." In discussions about medical applications it is argued that for highly relevant decisions interpretable models or at least comprehensible decision bases are important because they build trust and allow errors to be traced. Some authors recommend using more interpretable models in safety-critical cases or using black-box models only with supplementary validation and monitoring (see literature on interpretable models).
Bias and fairness: AI models learn from data, and if these data contain systematic biases, the AI reproduces them. In medicine this can mean that an algorithm performs worse for certain patient groups because those groups were underrepresented in the training data. Such biases can exacerbate health inequities. Therefore experts call for standardized procedures for bias analysis, reporting on datasets, and external validation (see WHO, EU guidelines).
Approval concepts: Regulatory authorities have begun developing specialized frameworks for AI-based medical software. The U.S. authority, for example, documents approaches for stepwise deployment and monitoring of AI/ML-based systems. Authorities emphasize the need for transparency, continuous performance monitoring, and post-market surveillance because AI models can change and because their performance in routine clinical use may differ from that in studies (see FDA Action Plan).
Clinical embedding: Another important concept is human-centered integration. AI should be integrated into clinical workflows in a way that supports relevant actors without making processes confusing. Implementation studies examine how physicians handle suggestions, whether workload decreases or increases, and whether patient outcomes actually improve. Such implementation results are often sparsely published; therefore system-level effects and long-term benefit–risk balances for many applications remain uncertain or inconsistently documented.
In conclusion: On a technical and regulatory level, a consensus is emerging that emphasizes two aspects. First: careful validation and monitoring are essential for safety-relevant applications. Second: explainability, representativeness of data, and human-centered design are not optional extras but central prerequisites for responsible use in medicine (see WHO, EU Guidelines, professional literature).
Part 3 (approx. 10 minutes): Concrete applications, limitations and short thought exercises
Goal of this part: Reflect the learned material with application examples, identify limits, and test understanding with short thought exercises.
Concrete, documented applications: In radiology and ophthalmology there are validated algorithms that mark abnormalities and in some cases are even allowed to make autonomous decisions if regulatory approval is in place. In pathology, image analysis algorithms are used to quantify cell types and patterns. In care delivery, AI-supported systems are used for risk assessment (e.g., fall, sepsis, or readmission risk), often as part of clinical decision support. In research, AI systems support the discovery of potential biomarkers and the identification of relevant molecular patterns that can lead to personalized therapy approaches (see Topol, WHO).
Limitations and typical problems: Many algorithms were developed in narrow, well-controlled environments and perform worse outside those environments. Data protection and ethical questions are not only theoretical: handling sensitive health data requires clear legal bases, technical safeguards, and transparency toward patients. In addition, liability questions remain: who is responsible if an automated system is wrong — the manufacturer, the health institution, or the treating physician? Regulators and courts are working on answers, but for many concrete situations there are not yet definitive legal precedents (see FDA, EU guidelines).
Will AI replace physicians? The current scientific and regulatory consensus does not foresee a complete replacement. Authors with expertise in medicine and AI emphasize instead a complementary or collaborative perspective: AI can automate routine tasks, lead to faster reports, and reduce information overload; complex treatment concepts, ethical deliberations, and personal interaction with patients remain central medical tasks. Which specific activities can be automated depends, however, on specialty, data availability, and regulatory integration (see Topol, WHO).
Short thought exercises for the group:
1) Assume an algorithm detected lung nodules on X-rays with high sensitivity in the original study, but in your clinic there are substantially more false alarms. What causes could explain this, and what steps would you take before deploying the algorithm in your clinic? Consider aspects such as data differences, workflow, retraining, and monitoring.
2) Imagine an AI-supported system recommends a different therapy for Patient A than the treating physician. What information and documentation should be available to make a proper decision and document it in a legally secure way? Think of explainability, evidence base, and consent.
3) How would you ensure that an AI system in your practice is fair to all patient groups? What data, analyses, and processes would be needed to detect and prevent systematic disadvantage?
These tasks do not have to be answered completely; they are intended to make the previously described concepts manageable and to illuminate the interface between technology, clinical practice, and ethics.