Authors:
- Andreas Dimakakos, Director of Scientific Insights, Intelligencia AI
- Kate Smietana, Director of Knowledge and Thought Leadership, Intelligencia AI
- Sofia Mastoraki, Senior Customer Experience and Insights Associate, Intelligencia AI
- Panos Karelis, VP of Commercial, Intelligencia AI
- Dimitrios Skaltsas, CEO and Co-founder, Intelligencia AI
The Information Paradox
Artificial intelligence is rapidly reshaping pharmaceutical R&D, enabling organizations to retrieve, summarize and interpret scientific information at unprecedented speed. Yet despite remarkable advances in model capability, high-value portfolio, licensing and clinical development decisions remain difficult. Our analysis points to an increasingly important reason why: for complex pharmaceutical questions, model capability is no longer the principal constraint. Without information structured around explicit relationships across assets, trials, indications and outcomes—and domain expertise to supplement incomplete evidence and the deep ontological nuances of clinical development—even highly capable LLMs struggle and produce flawed and biased representations of the industry's innovation landscape.
The challenge is not a shortage of information, but how pharmaceutical knowledge is organized. The evidence needed to support complex decisions is fragmented across clinical trials, regulatory decisions, publications, development programs and competitive intelligence—sources that use different identifiers, structures and levels of granularity. High-value questions therefore require information to be connected across scientific, clinical and regulatory domains before meaningful reasoning can begin.
The industry's response has largely focused on model capability. Each new generation promises stronger reasoning and fewer hallucinations, while retrieval-augmented generation, proprietary knowledge bases and web search have expanded access to information. These advances matter, but analytical power and access alone do not resolve fragmentation.
This brings the information layer into focus. Structured, expert-curated datasets and semantic relationships between assets, diseases, targets, trials and outcomes can provide LLMs with context that fragmented sources cannot consistently offer. As models improve, the question may therefore be shifting from which model performs best to how best to combine analytical capability, information breadth and domain-specific knowledge to support reliable decisions.
Truth and Trust: A Simple Framework to Evaluate AI Performance
Most public AI benchmarks focus on accuracy, reasoning or domain-specific knowledge. These measures remain important, but they capture only part of what matters in pharmaceutical R&D, where decisions often depend on evidence synthesized across multiple independent sources. An answer can be factually correct while omitting critical evidence, or appear convincing without providing a sufficiently complete basis for decision-making.
Pharmaceutical organizations can effectively evaluate AI through two complementary dimensions:
Truth measures whether an answer is factually correct and sufficiently complete to support a business or scientific decision.
Trust reflects whether users can rely on an answer, including consistency across repeated analyses and traceability to credible supporting evidence.
To ensure reliability, both dimensions need to be assessed and scored by human experts, not AI models. Together, Truth and Trust provide a practical framework for assessing AI systems on complex, evidence-driven pharmaceutical questions.
Applying the Framework
To explore how these dimensions perform in practice, we evaluated three frontier LLMs—ChatGPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8—using twenty oncology-focused questions representative of the analyses routinely performed across pharmaceutical R&D. The evaluation progressed from straightforward information requests to questions requiring historical reconstruction, competitive landscape assessment and evidence synthesis across multiple public sources.
Responses were compared against a human-verified reference dataset and assessed using the Truth and Trust framework. The objective was not to identify the "winning" model, but to understand where today's frontier LLMs succeed, where they struggle, and what drives those limitations.
When Better AI Doesn't Produce Better Decisions
The results revealed a remarkably consistent pattern. All three models performed well on straightforward retrieval or summarization. When questions required evidence to be assembled across disconnected sources, performance declined substantially. Yet the decline was not accompanied by equivalent warning signs: outputs remained reproducible and supported by credible references, keeping Trust high.
Figure 1. As question complexity increases, Truth score declines while Trust score remains comparatively stable, creating a widening gap between the quality of the underlying answer and its apparent reliability.
Truth declined steeply as question complexity increased. Answers remained largely accurate, but became progressively less complete, omitting historical programs, supporting evidence and competitive context. For pharmaceutical decision-making, incomplete representations of reality can be as misleading as incorrect facts.
The Common Failure Pattern
The consistency of this behavior suggests that the primary limitation is not model-specific. Despite differences in architecture and training, all three frontier models exhibited a similar failure pattern once questions depended on reconstructing fragmented pharmaceutical knowledge.
Figure 2. On a maximum score of 8, Truth remains low and declines with increasing information complexity across all three LLMs. Trust holds up better, but remains well below the maximum, leaving substantial room for improvement across both dimensions.
The most reliable tasks were those supported by a single authoritative source, such as identifying an FDA approval or retrieving a clinical trial. Performance declined when questions required information to be integrated across multiple repositories and historical records. Models frequently missed or misclassified discontinued programs, struggled to distinguish an active trial from an asset still in active development, and produced unreliable counts when assets, trials and indications had to be reconciled across sources.
These failures share an important characteristic: the required answer does not exist explicitly within any single source. Reconstructing development histories or competitive landscapes requires relationships across drugs, indications, trials and outcomes to be explicitly mapped, alongside expert-defined rules that resolve ambiguity across the underlying evidence. For example, an active ClinicalTrials.gov record cannot by itself establish that an asset remains in active development; program status must incorporate subsequent sponsor disclosures, regulatory events and other evidence.
This distinction helps explain why greater analytical power alone may not resolve the problem. To examine whether model evolution itself could overcome these limitations, we applied the same question set to three successive Claude Opus generations released over approximately six months. Across versions 4.6, 4.8 and 5, Truth improved by only about 7% on average per generation, while performance on the most demanding evidence-synthesis questions remained largely unchanged. Future models may make substantial advances, but our results suggest that waiting for the next generation is not, on its own, a reliable strategy for addressing these failures.
Improving the information available to the model had a larger effect. Direct access to authoritative public databases improved accuracy and completeness, with Claude Science increasing average Truth by 41% relative to the base configuration. Yet substantial gaps remained, with gains concentrated in specific questions and little improvement where answers required complete historical reconstruction or resolution of ambiguous data. Broader access clearly matters, but access alone cannot create relationships or resolve ambiguities that are not explicit in the underlying sources.
Together, these results point beyond both model capability and information access. For complex pharmaceutical questions, performance increasingly depends on whether evidence has been harmonized, connected through meaningful relationships and interpreted using domain-specific rules before it reaches the model.
The Next Competitive Advantage
The implications extend beyond choosing the right model or giving it access to more information. Decision-grade pharmaceutical AI requires a third layer: domain-specific knowledge that connects the evidence, resolves ambiguity and encodes the rules and ontological relationships that define drug development. Without it, even powerful models working across broad information sources remain exposed to the same gaps our analysis identified.
This is the philosophy behind our approach at Intelligencia AI. Frontier models provide analytical power, continuously maintained datasets provide breadth and context, and years of expert-led curation, ontology development and rules-based processes transform that evidence into decision-ready knowledge. The competitive advantage lies in bringing all three together—and particularly in the expertise required to structure pharmaceutical knowledge before AI is asked to reason over it.
For pharmaceutical organizations, this creates an opportunity today. The next advantage in AI-enabled R&D will not come from waiting for the next model release. It will come from building the knowledge foundation that allows increasingly capable models to deliver on their potential. In an industry where a single portfolio or development decision can create—or destroy—billions in value, that distinction matters.
1Methodology: The evaluation was conducted in June, 2026, using ChatGPT Plus v5.5, Gemini 3.1 Pro and Claude Max (Opus v4.8). Twenty oncology-focused questions spanning Historical Data (8), Competitive Landscape (7), and Trial Design & Outcomes (5) were tested, with questions progressing from direct-source retrieval to multi-source retrieval, source selection and evidence synthesis. Each question was run 10 times per model, generating 600 outputs in total. Truth and Trust are each scored on a 0–8 scale. Truth combines Accuracy and Completeness (0–4 each), while Trust combines Repeatability and Traceability (0–4 each). The same evaluation team was used throughout, with each output independently assessed by two reviewers against Intelligencia AI’s proprietary, human-verified reference dataset. Outputs were constrained to publicly available information and oncology, FDA-track, interventional, industry-led development. All reported question-level results represent averages across the three evaluated models. Follow-up analyses applied the same question set and evaluation framework across successive Claude Opus generations (v4.6, v4.8, v5.0) and configurations with enhanced information access, including direct connection to ClinicalTrials.gov, combined access to ClinicalTrials.gov, PubMed and FDA sources, and Claude Science (beta). Claude was selected for longitudinal analysis as the most balanced performer in the initial cross-model evaluation.