European Pharmacopoeia 5.38: a new chapter on data quality
Since 1 July 2026, the new European Pharmacopoeia (Ph. Eur.) general chapter 5.38, Quality of data, has been in force as part of Issue 12.3.
Published for information, this new chapter addresses the management of digital data quality throughout the data life cycle. Its publication comes at a time when the way pharmaceutical data are generated and used is evolving rapidly, with advanced analytics, machine learning and AI progressively finding their place in pharmaceutical development and quality control.
In this context, the quality of the data used to feed these approaches deserves increasing attention.
Looking at quality throughout the data life cycle
Chapter 5.38 describes several dimensions that contribute to data quality, including accuracy, completeness, consistency, timeliness and veracity, as well as availability and interoperability.
These considerations apply throughout the data life cycle, from data generation and acquisition to processing, transfer and storage. The chapter also specifically addresses the Extract-Transform-Load (ETL) process used to collect data from different sources, transform them and make them available in another environment.

This is particularly relevant in pharmaceutical laboratories, such as Quality Assistance, where analytical data may come from multiple instruments and systems and may subsequently need to be consolidated with other datasets.
The quality and compatibility of the different sources therefore need to be considered when data are aggregated. Harmonisation plays an important role here, both when data enter laboratory systems and when they are prepared for transfer or further use.
This also connects with a broader challenge in pharmaceutical development: maintaining continuity between analytical data generated across different methods, systems and stages of a product's lifecycle. We recently explored why this analytical continuity matters, and how a more integrated analytical environment can help preserve it.
FAIR and ALCOA+ principles
The chapter places these considerations within principles that are already familiar to the pharmaceutical industry.
Data reproducibility and stewardship are addressed through the ALCOA+ framework and the FAIR principles: Findable, Accessible, Interoperable and Reusable. Open and well-documented file formats, together with controlled data flows, contribute to maintaining these characteristics over time.
Interoperability is becoming particularly relevant as analytical data increasingly need to move between organisations and systems.
For a sponsor working with an external analytical laboratory, the usefulness and value of a dataset also depends on its ability to be transferred, integrated and reused within the sponsor's own data environment. Consistent structures and formats can facilitate this process and reduce the need for additional data transformation after delivery.
Scientific expertise remains part of the process
Chapter 5.38 also addresses the use of data in automated decision-making systems and highlights the continued involvement of subject matter experts.

As machine learning and AI-based tools are progressively introduced into pharmaceutical applications, the datasets used by these systems need to remain scientifically relevant and appropriate for their intended use. The quality of an automated output will inevitably depend on the quality and context of the data provided to the system.
Keeping scientific expertise involved in these processes therefore remains an important part of data quality management.
How we approach data quality at Quality Assistance
At Quality Assistance, analytical data management already aligns with these principles. Using our LIMS and ELN/LES (LabVantage) environments, we harmonise data flows and apply automated quality checks to ensure datasets are consistent, traceable, interoperable and aligned with FAIR and ALCOA+ principles. As data play a growing role in AQbD and advanced analytics, Ph. Eur. chapter 5.38 provides a valuable framework for managing data quality across the entire lifecycle and for preparing data for effective exchange and future use.