Insight
Open Data: The Promise and the Limits
Open publication makes verification and reuse possible — but openness alone guarantees neither quality nor privacy nor use.
Summary
Open data — data anyone can access, use and share — enables verification, secondary research and public tools, and the FAIR principles (findable, accessible, interoperable, reusable) describe what makes it genuinely usable. But openness is not sufficiency: open datasets can be low-quality or unusable in practice, privacy limits what should be opened at record level, and publication without documentation or capacity produces transparency in name only.
The strongest argument for open data is epistemic: claims based on closed data must be taken on trust, while claims based on open data can be checked, re-analysed and corrected. Open publication is how errors get found — and how findings earn durable credibility.
#What ‘good’ open data means
The widely adopted FAIR principles (Wilkinson et al., 2016) hold that data should be findable (indexed, with persistent identifiers), accessible (retrievable by standard protocols), interoperable (standard formats and vocabularies) and reusable (documented, with clear licences). The principles matter because raw openness fails without them: an undocumented spreadsheet at an unstable URL is technically open and practically useless.
#Three limits to keep in view
- Privacy. Record-level data about people can often be re-identified even after names are removed, especially when combined with other sources. Aggregation, access controls for sensitive microdata, and formal disclosure-control methods exist precisely because ‘open everything’ is not a responsible default for personal data.
- Quality. Openness and accuracy are independent properties. Open publication makes errors findable; it does not prevent them. Documentation, versioning and known provenance are what allow users to judge fitness for purpose.
- Capacity and ‘open-washing’. Publishing data is cheap; making it usable — cleaned, documented, maintained, answerable — is not. Portals full of stale, undocumented files deliver the appearance of transparency without its substance.
The practical conclusion: open data is infrastructure, and like all infrastructure it works when maintained. For research organisations, the same logic applies to publications themselves — openly licensed, documented and stably addressed work is work that others can verify and build on, a theme we develop in What Makes Research Citable.
Terms used in this document
How to cite this article
Helsen Institute for Public Research (2026). “Open Data: The Promise and the Limits.” Helsen Institute Insight, published July 17, 2026. https://helsen-institute.vercel.app/insights/open-data-promise-and-limits
This work is licensed under CC BY 4.0. You may republish, translate, and adapt it — including for commercial purposes — with attribution to the Helsen Institute for Public Research and a link to https://helsen-institute.vercel.app/insights/open-data-promise-and-limits.
Related from the Institute
Insight
What Makes Research Citable — by People and by Machines
Stable references, explicit dates, open licences and plain summaries determine whether research gets cited accurately, by readers and increasingly by AI systems.
Insight
Why Definitions Decide Debates in Public Statistics
Unemployment, poverty, migration: in each case, the definitional choice — not the data collection — often determines the headline.