Data Integrity and Compliance

Data Integrity

Last updated:
October 3, 2026

TL;DR

Data integrity is the extent to which lab data stay complete, consistent and accurate for their whole life cycle, from the moment a value is recorded to the end of its retention period. Regulators judge it against the ALCOA+ attributes and against the controls around each record, such as audit trails and second-person review.

What is data integrity?

Data integrity is the extent to which data stay complete, consistent and accurate across their whole life cycle, from the first recording of a value through processing, review, reporting and archiving. The FDA's 2018 guidance on data integrity in drug CGMP defines it as "the completeness, consistency, and accuracy of data," and the MHRA's 2018 GxP guidance adds that those characteristics have to be maintained throughout the data life cycle.

It became a named regulatory priority in the 2010s. Inspections kept turning up deleted chromatography runs and samples retested until one passed, so the FDA, the MHRA and the WHO each published dedicated data integrity guidance, all of it built on the ALCOA+ attributes.

Data security gets mixed up with it a lot, and the two do overlap. Security keeps data away from the wrong people and safe from loss, through access control and backups. Integrity asks whether you can trust the record itself, including how it was corrected and who reviewed it. So an encrypted, backed-up spreadsheet in which someone typed over a failing value still has a data integrity problem.

What do the FDA and MHRA expect to see?

They expect every reported result to trace back to its raw data, with each change visible and explained. The regulations spell some of this out. For studies conducted under GLP, 21 CFR 58.130(e) requires that a change to an entry doesn't obscure the original, gives a reason, and is dated and signed. For drug testing under CGMP, 21 CFR 211.194(a)(8) requires the initials or signature of a second person showing that the original records were reviewed for accuracy, completeness and compliance with established standards.

On paper, labs meet these expectations with bound notebooks and single-line strike-throughs. On electronic systems it comes down to individual logins and an audit trail users can't switch off. And the person who generated a result shouldn't have permission to delete it.

Data integrity failures turn up regularly in the FDA warning letters published on the agency's website. One finding can put every result produced on the same system in doubt, which is why a QA lead will worry about a single shared login on one instrument PC.

What does a trustworthy result need behind it?

A chain of small records, each one keeping the link between the number you report and the work behind it. Take a stability study where a technician pulls the 6-month vial from a -80 °C freezer, runs it on the HPLC and reports a purity value. For that value to hold up, the record has to show which vial was pulled, by sample ID and box position, and who pulled it. The raw chromatogram should be there exactly as the instrument produced it, with any reprocessing kept next to the original integration, and every injection belongs in the record, including the one aborted when the column pressure spiked.

If the reported value changes, the original stays visible with a reason and a timestamp. Then a second person reviews the raw data and the calculation before the result goes anywhere.

Outside regulated work the cost shows up as lost time. A postdoc leaves, the raw files behind a key figure turn out to live on a personal laptop, and the experiment gets repeated before the paper can be revised.

To find the weak links in your own lab, map where the data actually live. For each instrument and record type, write down where raw files are saved, who can edit or delete them, how corrections are made and who reviews the results. The exercise usually turns up at least one instrument PC with a shared login and raw files that exist only on its local drive. Fix those first. Then add a review step for results that feed a decision, like a batch release or a go/no-go on a lead compound, and explain the reason for each control to the people who follow it, because a rule that looks arbitrary gets worked around at 7 pm on a Friday.

IGOR records every inventory action in Sample History with a timestamp and the user who performed it. Notebook entries have a full audit trail, and access is managed through role-based access control, enforced in the application and at the database level through row-level security.

Frequently asked questions

What is the difference between data integrity and data quality?

Data integrity is about whether a record is trustworthy and unaltered, and data quality is about whether the data are fit for the purpose you're using them for. An assay result can be fully attributable and untouched and still be scientifically weak because the wrong control was used. A result needs both before it can support a decision.

What are the most common data integrity problems in research labs?

Shared logins and raw data that live only on an instrument PC come up most often. Hand transcription from a printout into a spreadsheet is another, and so are failed runs that never make it into the notebook.

Is data integrity a legal requirement for academic research?

Only for regulated work, such as studies conducted under GLP or testing under GMP, and most academic research falls outside those rules. Academic labs are still bound by institutional research integrity policies and funder requirements. The NIH, for example, has required data management and sharing plans since January 2023. Poor records also make a misconduct allegation much harder to answer.

How does an audit trail support data integrity?

An audit trail records who created or changed a record, when, and what the previous value was, so every change is attributable and the original stays visible. That only works if users can't switch it off and someone actually reviews it. In IGOR, the review history and full audit trail of a notebook entry can be viewed at any time.

What happens during a data integrity audit?

The auditor picks reported results and traces them back to the raw data, checking each step against ALCOA+. Typically that means asking for the original instrument file, comparing it with the reported value, and reading the audit trail for edits, deletions, aborted runs or reprocessing around the time of the test. Expect questions about who holds administrator rights on each system.

Related terms

  • ALCOA+: the nine attributes regulators use to assess data integrity.
  • Audit Trail: the timestamped log of who changed a record and when.
  • 21 CFR Part 11: the FDA regulation for electronic records and electronic signatures.
  • Chain of Custody: the documented record of who handled a sample, when and where.
  • Electronic Signature: how a signer is linked to an electronic record.
  • ELN and LIMS glossary: the full list of lab software and compliance terms.

References

  1. U.S. Food and Drug Administration. Data Integrity and Compliance With Drug CGMP: Questions and Answers. Guidance for Industry, December 2018.
  2. Medicines and Healthcare products Regulatory Agency (MHRA). 'GXP' Data Integrity Guidance and Definitions, Revision 1, March 2018.
  3. World Health Organization. Guideline on data integrity. WHO Technical Report Series No. 1033, Annex 4, 2021.
  4. 21 CFR 58.130, Conduct of a nonclinical laboratory study (Good Laboratory Practice), eCFR.
  5. 21 CFR 211.194, Laboratory records (Current Good Manufacturing Practice for finished pharmaceuticals), eCFR.

This page is a general overview of data integrity for lab scientists. It is not formal compliance or legal advice, and requirements for a specific study or product depend on the regulations that apply to it and on your organization's own procedures.