A total of 193 respondents took part in the survey on research data storage, collectively reporting 290 repositories, databases and other storage solutions. After consolidating different name variants and verifying the links, these responses correspond to 70 distinct services.
The dataset containing the results of this round of the survey is publicly available in the National Data Repository.
More than a quarter of responses point to Zenodo
Respondents mentioned Zenodo a total of 78 times, meaning it accounts for almost 27% of all reported storage solutions. It is followed at some distance by databases operated by the European Bioinformatics Institute (EMBL-EBI), with 23 mentions, and databases of the US-based NCBI, with 18. ASEP, the institutional repository of the Czech Academy of Sciences, was mentioned by 17 respondents. CESNET and IT4Innovations storage services appeared in the responses 11 times.
Specialised services used in molecular biology, genomics, proteomics and other areas of the life sciences are also strongly represented. These include EMBL-EBI and NCBI databases, the Protein Data Bank and ProteomeXchange. These disciplines are characterised by long-established infrastructures for storing and sharing data, as well as the use of discipline-specific persistent identifiers (PXD ID, UniProt ID, PDB ID), which make individual datasets unambiguously findable and citable.
Specialised databases also illustrate why the 70 reported items cannot simply be understood as 70 independent repositories. Individual services may function as repositories, databases, catalogues or aggregators, and their roles may overlap. For example, the ProteomeXchange infrastructure brings together several repositories for proteomics data, including PRIDE, MassIVE and jPOST, while ProteomeCentral serves as their central catalogue. PRIDE is also part of the EMBL-EBI infrastructure. In practice, a response such as “EMBL-EBI” may therefore refer to data stored in one of the many specialised resources covered by this infrastructure.
The way stored data are identified also varies. General-purpose repositories such as Zenodo or Figshare primarily use DOIs, while disciplinary databases often rely on their own accession numbers. In the life sciences, datasets may therefore have, for example, a PXD identifier in ProteomeXchange or specific identifiers assigned by NCBI databases. In addition to showing the number of services being used, the survey results therefore reveal a highly heterogeneous infrastructure that differs considerably across research disciplines.
The survey results will also be used to explore the possibility of harvesting metadata on Czech research outputs stored in external, including international, repositories into the National Metadata Directory (NMA). The survey may therefore help to better map the sources from which metadata on Czech research data could potentially be harvested into the NMA.
Alongside repositories and disciplinary databases, respondents also reported storage methods that unfortunately do not support the effective long-term management and accessibility of research data. Nine respondents use internal or local institutional storage, eight reported OneDrive, SharePoint or Microsoft Teams, and seven use flash drives or external USB drives. Other frequently mentioned services include GitHub, Google Drive, Figshare and the Open Science Framework.
Data Most Often Accompany Publications
The second part of the survey focused on the stage at which researchers deposit their data. Most commonly, these are data used as supporting material for publications. This option was selected by 127 respondents, representing approximately two thirds of all participants. Raw data are deposited by 88 respondents, while 84 deposit data for long-term archiving. Final processed and validated data were reported by 79 respondents, or 41%, while 55 respondents deposit partially cleaned or calibrated data.
Specifically modified datasets are deposited less frequently. Aggregated data were reported by 31 respondents, anonymised data by 26, and pseudonymised data by only ten. In total, 18% of respondents stated that they do not deposit data in a repository at all.
The responses therefore do not point to a single standard approach to data storage. Respondents reported a total of 67 different combinations of stages at which they deposit their data, highlighting considerable variation in practice across the research environment. Even the most common responses do not dominate significantly: 24 respondents stated only that they do not deposit data in a repository, while another 17 deposit exclusively supporting data for publications. According to the survey results, data deposition is therefore often closely linked to the publication process. At the same time, a wide range of other approaches is also represented, from depositing raw data and final processed versions to long-term archiving. This diversity reflects not only differences between disciplines, but also varying practices across individual research teams and institutions.
The Survey Continues
The survey remains open, and researchers can still share their experiences with storing research data. Newly collected responses will be evaluated as part of the next round of data collection.
Note: The composition of the survey sample provides important context for interpreting the results. Researchers from the University of Chemistry and Technology Prague and Masaryk University were the most strongly represented, while a significant proportion of respondents also came from the Institute of Molecular Genetics of the Czech Academy of Sciences, the Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, and the Institute of Physics of the Czech Academy of Sciences. Due to the uneven representation of individual institutions, the data cannot be generalised to the entire Czech research environment. The results should therefore be interpreted as a snapshot of the current practices of the respondents involved, capturing different approaches to research data storage and the use of related services.