SciLifeLab PULSE Future Leaders Perspective

STAY UP TO DATE

Open Data: SciLifeLab PULSE Future Leaders’ Perspective

Future Leaders’ Perspectives share the thoughts of the SciLifeLab PULSE Postdoc cohort on leadership, science communication, and key competencies. Each installment is part of the transferrable skills training program, where postdocs reflect on expert interviews and practical assignments within these topics.

Open Data: Challenges and opportunities in making data truly reusable

by Nina Trubanová, Álvaro Serrano, and Patricia Sosa  

Editor: Suné Joubert

Science thrives on shared knowledge1,2. The principle sounds straightforward: deposit your data, let others build on it. Yet for early-career researchers working across the life sciences, the daily reality of data reuse is far messier than any policy document suggests. As SciLifeLab PULSE postdoctoral students, we recently examined our own experiences with reusing public datasets and producing data intended for others. What emerged was a picture of shared frustration, unexpected consensus, and cautious optimism.

We all hit the same wall

Across the groups, one experience was universal: the gap between data being available and data being (re)usable. Public repositories, such as NCBI GEO3 for sequencing data, PRIDE4 for proteomics, and the PDB5 for structural data, hold an extraordinary wealth of information. Datasets are findable, downloadable, and technically open6. On paper, this is a triumph of the FAIR principles.

In practice, finding the data is only the beginning. Incomplete metadata was a recurring frustration across every discipline we represented. Sequencing datasets sometimes lack sample-level condition labels. Proteomics submissions may omit instrument settings or sample preparation details critical for reproducing a quantification workflow. Affinity or interaction datasets can be deposited without the binding conditions or controls needed to interpret them. As one colleague put it during our discussion: “(Because of this lack of metadata) the inability even to perform quality control on a dataset renders it effectively useless.” If you cannot assess data quality, you cannot decide whether to use it. And if that basic step is impossible, you are stuck.

Batch effects and protocol heterogeneity compound the problem, regardless of data type. Different mass spectrometry platforms, library preparation kits, sequencing depths, or prediction pipelines all introduce systematic variation that demands careful normalisation before any cross-study comparison can be attempted. Each of these barriers demands time, expertise, and often bespoke computational solutions, resources that early-career researchers frequently lack7.

Where we disagree: simplicity versus completeness

Perhaps the most lively part of our discussion revealed a genuine tension within the group. On one side, some argued passionately for richer, more detailed metadata standards, such as sharing models, documenting analytical choices, and making every step reproducible. “Please share more details about your models and how you did it,” one participant urged, “so we can reproduce them, or at least confirm how the data was really acquired.”

On the other side, an equally compelling voice cut through: “Sometimes I just want the spreadsheet.”

This is not a trivial disagreement. It reflects a real trade-off in open science. Overcomplicating data deposition can discourage researchers from sharing at all. If the activation energy for sharing is too high, data stays on personal hard drives. Yet oversimplified deposits, in the form of raw files dumped without context, can be nearly as inaccessible as no data at all. The sweet spot lies somewhere in between: standardised, machine-readable metadata that captures the essentials without demanding a technical manifest alongside every upload. What “the essentials” look like will differ between a proteomics experiment and a genomics time course, but the underlying principle is the same: enough context for a stranger to judge quality and applicability8.

The sensitive data dilemma: anonymisation and access

For those working with clinical or human-derived samples, or patentable data, the push for open science introduces a unique set of hurdles. The tension between maximising data utility and protecting patient privacy was a key point of reflection in our group. Proper data anonymisation is non-negotiable for sharing sensitive information, yet scrubbing a dataset of identifiable markers often risks removing the demographic or clinical metadata that makes it scientifically valuable9.

We argued that “open” cannot always mean “publicly downloadable.” Navigating this requires a nuanced approach to sharing sensitive data. The community needs better, more standardised frameworks for managed access, like secure data enclaves and tiered access models, allowing sensitive datasets to remain FAIR without compromising ethical obligations10,11. In the meantime, we can only adhere to the standard open-access mantra, “as open as possible, as closed as necessary”12, at a very broad level.

What we want stakeholders to know

If we could send a single message upwards, it would be this: incentivise quality data sharing, not just data deposition. Current funder mandates, whether from SciLifeLab, the MSCA, the ERC, or others, increasingly require open data deposition, and that is welcome progress. But mandates alone do not guarantee usability. We would like to see funders support dedicated data curation roles within research groups, recognise data contributions in hiring and promotion criteria, and invest in the long-term maintenance of repositories. Data infrastructure is scientific infrastructure; it deserves sustained funding, not one-off grants.

For principal investigators, we ask for a culture shift. Early-career researchers who spend weeks curating a dataset for public deposition are doing essential scientific work. That effort should be visible in performance evaluations, not treated as an administrative afterthought. And for future early-career researchers, we would urge them to treat data management as a core research skill from day one. The time invested pays dividends: for your future self trying to reproduce an analysis two years later, and for every researcher who follows.

The sustainability argument we rarely make

There is an underappreciated dimension to data reuse that deserves a louder voice: sustainability. Reusing existing datasets reduces reagent consumption, instrument time, and, in fields involving animal or clinical samples, the ethical cost of generating new material. A well-designed meta-analysis or computational screen can extract new biological insight from already-funded experiments without a single additional sample being collected. Responsible reuse is not just efficient science; it is more ethical science13,14.

But sustainability has another dimension that the open science community has been slower to confront: our digital footprint. Storing massive amounts of high-throughput data is both financially and environmentally expensive, contributing significantly to global greenhouse gas emissions15,16. Data centers consume immense amounts of electricity, accounting for over 1% of total global electricity demand17, and the academic community cannot afford to treat cloud storage as infinite. To make open data truly sustainable, we must address the entire lifecycle of what we store, expanding traditional open science practices to actively minimise this digital environmental footprint15. Implementing robust data versioning and end-to-end traceability is essential here18. As datasets are updated, re-processed, or corrected, proper version control allows researchers to track analytical changes and maintain strict reproducibility without needlessly duplicating petabytes of raw files18,19. Sustainable open science means being as thoughtful about how we store and version our data as we are about generating it.

Education, equity, and community engagement

Just as we learn to walk before we run, researchers should encounter FAIR principles long before their PhD. Building high-quality and trustworthy data should start early in a scientist’s career20. Training initiatives and dedicated courses focused on navigating repositories, identifying reliable datasets, and depositing high-quality data are essential to strengthening open science practices.

The production of high-quality data has been widely discussed and refined over the years, with approaches that are highly specific to each research area. Of course, data quality can and should remain an open topic for discussion. However, this issue goes beyond data repositories. 

In clinical or field research settings, data collection occurs on a daily basis, and this represents a bottleneck in some countries. Clinicians working in overstretched healthcare systems, without full staffing or adequate IT infrastructure, face difficulties in producing high-quality data21: handwritten diagnoses, lost papers, and incomplete medical records are a common reality. These infrastructural constraints are not evenly distributed, they disproportionately affect researchers and clinicians in lower-resourced institutions and countries. As a result, the burden of producing FAIR-compliant data falls unevenly across the global research community, risking a widening of existing gaps in scientific capacity and participation22. Broadening this conversation beyond our institutions and actively supporting other countries in addressing the challenges of producing FAIR-compliant data will be critical to ensuring that the benefits of FAIR data are equitably shared across the global scientific community and that it accelerates research development.

A shared commitment

Despite the frustrations, our group was unanimous on one point. When asked whether they intended to make their own data reusable, the answer was immediate and emphatic: yes. We want to publish, we want to share everything, and we hope people will reuse it. That consensus is significant. It suggests that the next generation of researchers is not disillusioned by the imperfections of open science; rather, we are motivated to do it better. We recognise that our datasets will only be as useful as the metadata we attach to them, and we are willing to put in the work.

The foundations for a more open, reusable scientific data ecosystem are already in place. FAIR principles provide the framework. Public repositories provide the infrastructure. Funder mandates provide the push. What is still needed is the connective tissue: better metadata standards that balance thoroughness with practicality, institutional recognition of data stewardship, sustained funding for repositories, and a research culture that values the quiet, unglamorous work of making data truly usable.

We are optimistic. The datasets accumulating across life science repositories represent an extraordinary collective resource. With continued community investment in annotation, curation, and open sharing, the next generation of researchers will inherit not just raw data, but genuinely reusable scientific commons. That is a future worth building toward.

References

1.         Ellemers N. Science as collaborative knowledge generation. British J Social Psychol. 2021 Jan;60(1):1–28. doi:10.1111/bjso.12430 

2.         Vicente-Saez R, Martinez-Fuentes C. Open Science now: A systematic literature review for an integrated definition. Journal of Business Research. 2018 Jul;88:428–36. doi:10.1016/j.jbusres.2017.12.043 

3.         Clough E, Barrett T, Wilhite SE, Ledoux P, Evangelista C, Kim IF, et al. NCBI GEO: archive for gene expression and epigenomics data sets: 23-year update. Nucleic Acids Research. 2024 Jan 5;52(D1):D138–44. doi:10.1093/nar/gkad965 

4.         Perez-Riverol Y, Bandla C, Kundu DJ, Kamatchinathan S, Bai J, Hewapathirana S, et al. The PRIDE database at 20 years: 2025 update. Nucleic Acids Research. 2025 Jan 6;53(D1):D543–53. doi:10.1093/nar/gkae1011 

5.         Berman H, Henrick K, Nakamura H, Markley JL. The worldwide Protein Data Bank (wwPDB): ensuring a single, uniform archive of PDB data. Nucleic Acids Research. 2007 Jan 3;35(Database):D301–3. doi:10.1093/nar/gkl971 

6.         Wilkinson MD, Dumontier M, Aalbersberg IjJ, Appleton G, Axton M, Baak A, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016 Mar 15;3(1):160018. doi:10.1038/sdata.2016.18 

7.         Hughes LD, Tsueng G, DiGiovanna J, Horvath TD, Rasmussen LV, Savidge TC, et al. Addressing barriers in FAIR data practices for biomedical data. Sci Data. 2023 Feb 23;10(1):98. doi:10.1038/s41597-023-01969-8 

8.         Gomes DGE, Pottier P, Crystal-Ornelas R, Hudgins EJ, Foroughirad V, Sánchez-Reyes LL, et al. Why don’t we share data and code? Perceived barriers and benefits to public archiving practices. Proc R Soc B. 2022 Nov 30;289(1987):20221113. doi:10.1098/rspb.2022.1113 

9.         Arellano AM, Dai W, Wang S, Jiang X, Ohno-Machado L. Privacy Policy and Technology in Biomedical Data Science. Annu Rev Biomed Data Sci. 2018 Jul 20;1(1):115–29. doi:10.1146/annurev-biodatasci-080917-013416 

10.       Lvovs D, Creason AL, Levine SS, Noble M, Mahurkar A, White O, et al. Balancing ethical data sharing and open science for reproducible research in biomedical data science. Cell Reports Medicine. 2025 Apr;6(4):102080. doi:10.1016/j.xcrm.2025.102080 

11.       Wu M, Löffler F, Mathiak B, Psomopoulos F, Schindler U, Aryani A, et al. Bridging the Data Discovery Gap: User-Centric Recommendations for Research Data Repositories. Data Science Journal. 2026 Feb 12;25:6. doi:10.5334/dsj-2026-006 

12.       Recommendations for FAIR and open research data [text] [Internet]. 2020 [cited 2026 May 29]. Available from: https://www.vr.se/english/mandates/open-science/open-access-to-research-data/recommendations-for-fair-och-open-research-data.html

13.       Holub P, Kohlmayer F, Prasser F, Mayrhofer MTh, Schlünder I, Martin GM, et al. Enhancing Reuse of Data and Biological Material in Medical Research: From FAIR to FAIR-Health. Biopreservation and Biobanking. 2018 Apr 1;16(2):97–105. doi:10.1089/bio.2017.0110 

14.       Sielemann K, Hafner A, Pucker B. The reuse of public datasets in the life sciences: potential risks and rewards. PeerJ. 2020 Sep 22;8:e9954. doi:10.7717/peerj.9954 

15.       Labadie M, Desconnets JC, Sabot F. Digital data and sustainability. In: Dangles O, Fréour C, editors. Sustainability Science – Volume 1 [Internet]. Marseille: IRD Éditions; 2023 [cited 2026 May 29]. p. 150–3. Available from: https://books.openedition.org/irdeditions/65622 doi:10.4000/15pkn 

16.       Samuel G, Lucassen AM. The environmental sustainability of data-driven health research: A scoping review. DIGITAL HEALTH. 2022 Jan;8:205520762211112. doi:10.1177/20552076221111297 

17.       Mersico L, Abroshan H, Sanchez-Velazquez E, Saheer LB, Simandjuntak S, Dhar-Bhattacharjee S, et al. Challenges and Solutions for Sustainable ICT: The Role of File Storage. Sustainability. 2024 Sep 14;16(18):8043. doi:10.3390/su16188043 

18.       Cejudo A, Tellechea Y, Calvo A, Almeida A, Martín C, Beristain A. Scalable Big Data Platform With End-to-End Traceability for Health Data Monitoring in Older Adults: Development and Performance Evaluation. JMIR Med Inform. 2025 Dec 22;13:e81701–e81701. doi:10.2196/81701 

19.       Antunes B, Hill DRC. Reproducibility, Replicability and Repeatability: A survey of reproducible research with a focus on high performance computing. Computer Science Review. 2024 Aug;53:100655. doi:10.1016/j.cosrev.2024.100655 

20.       FAIRsFAIR [Internet]. 2020 [cited 2026 May 29]. FAIR in European Higher Education. Available from: https://www.fairsfair.eu/fair-european-higher-education 

21.       Bernardi FA, Alves D, Crepaldi N, Yamada DB, Lima VC, Rijo R. Data Quality in Health Research: Integrative Literature Review. J Med Internet Res. 2023 Oct 31;25:e41446. doi:10.2196/41446 

22.       Rangel Teixeira A, Kleinlein R, Kalema NL, Agyemang GO, Senteio C, Celi LA, et al. Reimagining open science for global health: Epistemic power and the pursuit of health equity. Meudec M, editor. PLOS Glob Public Health. 2025 Nov 6;5(11):e0005300. doi:10.1371/journal.pgph.0005300

This perspective was developed as part of the SciLifeLab PULSE Transferable Skills Training programme, 2026. Co-funded by the European Union. Views and opinions expressed are those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency.


STAY UP TO DATE

Last updated: 2026-08-28

Content Responsible: Kristen Schroeder(kristen.schroeder@scilifelab.se)