Why Is AI still Struggling to Gain a Foothold in Healthcare?
This article is part 1 of a 4-part series on the barriers and solutions to building reliable AI in healthcare.
Artificial intelligence is already shaping our daily lives. From virtual assistants to productivity tools, it is gradually being embedded in both our personal and professional practices. ChatGPT, Copilot, and other consumer solutions are democratizing access to analytical capabilities that were once reserved for experts.
Paradoxically, in the most critical sectors (healthcare, pharmaceuticals, medical research), AI is present, but its use often remains limited to isolated cases or disconnected from core business processes.
Pharmaceutical professionals, faced with massive volumes of complex information, have access to powerful tools but still struggle to fully integrate them into strategic, regulatory, or decision-making workflows.
Why such a gap? The answer does not lie in resistance to change, but in the very nature of health and scientific challenges. In an environment where mistakes are not an option, AI must meet strict requirements for reliability, traceability, and regulatory compliance.
To understand how to overcome these barriers, it is essential to clearly identify the challenges and obstacles slowing down AI adoption in healthcare environments.
The Four Major Barriers to AI in the Pharmaceutical Industry
Barrier #1: Data Fragmentation in Healthcare
The problem: a siloed data ecosystem
One of the main obstacles to medical AI lies in the extreme fragmentation of healthcare data. Today, pharmaceutical companies have access to a growing mass of public health data: reimbursement databases, regulatory repositories, clinical trials, epidemiological data, pricing, ATC or ICD codes, scientific publications, and more.
But this data comes with three major challenges:
- Dispersion – spread across dozens of heterogeneous sources (institutional websites, PDFs, HTML tables, unstructured documents).
- Heterogeneity – formats, granularity, nomenclatures, and semantic structures are not harmonized.
- Inconsistency – the same drug or pathology may be referenced differently depending on the database. Product names, codes, indications, or classifications vary across the BDPM (French Public Drug Database), Orphanet, ATC, ICD-10, or HAS databases.
The case of rare diseases
In rare diseases, this challenge is particularly evident. Teams often need to cross-reference data from:
- Orphanet, which uses a specific ontology (OrphaCode, detailed clinical concepts),
- ICD-10, much broader and more exhaustive but often too generic to accurately describe rare diseases,
- Various file formats (PDF, Excel, etc.), containing key information but not automatically exploitable.
The result: analyzing the epidemiology of a rare disease or building a reimbursement dossier requires hours of manual data cleaning, reprocessing, and validation across multiple teams. In this context, an AI without a concrete database to rely on may struggle to distinguish critical nuances or interpret data within the correct regulatory framework.
Take Huntington’s disease, for example. A generalist AI might not distinguish between the classic and juvenile forms, even though this nuance has major implications: genomically, in-patient management, and in determining eligibility for therapeutic protocols.
Without a rigorous knowledge base to guide its reasoning, AI risks merging non-comparable data or producing biased analyses. This issue is compounded by informational noise: an overload of poorly structured data from multiple formats and ontologies. The AI can no longer distinguish what is essential from what is secondary, increasing the risk of errors in an environment where precision is non-negotiable.
Impact on expert teams’ efficiency
For healthcare teams, this fragmentation results in:
- Significant time spent manually reconciling data across sources.
- Increased risk of errors or inconsistencies due to varied formats.
- Limited visibility over a therapeutic area.
Barrier #2: Lack of Traceability in AI-Generated Responses
The problem: information without source or method
Large language models such as ChatGPT or other generative AI tools can produce convincing answers. But in a regulated environment like healthcare, traceability is not a luxury, it is an obligation. Yet these AI systems do not provide:
- Precise traceability of the data source (scientific publication, regulatory database, unverified website, etc.),
- Transparency about the reasoning process (synthesis, inference, extrapolation),
- Guarantees that the information is up to date, compliant with the latest regulations, or not already refuted.
This lack of transparency makes these tools difficult to use for strategic deliverables such as pricing and reimbursement dossiers, budget impact studies, or multi-country benchmarks.
The hallucination problem: when AI invents without warning
This opacity creates a well-known risk: hallucination. An AI “hallucinates” when it provides false or partial answers, but in a plausible and convincing way, without indicating that it is extrapolating.
In a scientific context, this may lead to:
- Fabricated bibliographic sources that appear authentic.
- Invented drug indications presented as established facts.
- Erroneous product profiles, resulting from a blend of similar molecules.
- Misinterpretation of causal links (e.g., molecule - gene, symptoms - disease).
Such errors can be especially problematic when they go unnoticed, since the user has limited ways to verify whether the information is scientifically sound or partly fabricated.
The risk is compounded by the difficulty of structuring sources: when overwhelmed with unranked, non-hierarchical information, AI struggles to separate the essential from the secondary, producing erroneous answers without warning the user.
Business requirement: only justifiable information is actionable
In science, information has value only if it can be traced back to a precise, validated source, contextualized (by therapeutic area, population, regulatory status), and justified by a clear business or scientific logic.
Without these elements, no AI, no matter how sophisticated it appears, can be integrated into the tools and processes used to defend pricing, structure economic models, or meet health authority requirements.
This gap creates the illusion that AI “understands” the issues, when in fact it only reproduces likely patterns without considering the sector’s specific constraints. In a context where reproducibility, traceability, and interpretability are essential, this approach limits the operational value of AI-generated deliverables.