Zenodo Leads the Way, but Czech Researchers Use Dozens of Different Repositories for Their Data

Where do Czech researchers store their data, and at what stage of the research process are data most often deposited in repositories? The results collected so far from 193 respondents identified a total of 70 different repositories, databases and other storage solutions. Zenodo is the clear leader, while specialised databases used in molecular biology, genomics and proteomics are also strongly represented. Data are most often deposited as supporting material for publications, whereas almost one fifth of respondents do not deposit their data in repositories at all. Alongside established repositories, more traditional and riskier storage practices also remain common, including cloud storage, local servers and external drives.

16 Sep 2026 Martin Dvořák Karolína Smetanová

No description

A total of 193 respondents took part in the survey on research data storage, collectively reporting 290 repositories, databases and other storage solutions. After consolidating different name variants and verifying the links, these responses correspond to 70 distinct services.

The dataset containing the results of this round of the survey is publicly available in the National Data Repository.


More than a quarter of responses point to Zenodo

Respondents mentioned Zenodo a total of 78 times, meaning it accounts for almost 27% of all reported storage solutions. It is followed at some distance by databases operated by the European Bioinformatics Institute (EMBL-EBI), with 23 mentions, and databases of the US-based NCBI, with 18. ASEP, the institutional repository of the Czech Academy of Sciences, was mentioned by 17 respondents. CESNET and IT4Innovations storage services appeared in the responses 11 times.

Specialised services used in molecular biology, genomics, proteomics and other areas of the life sciences are also strongly represented. These include EMBL-EBI and NCBI databases, the Protein Data Bank and ProteomeXchange. These disciplines are characterised by long-established infrastructures for storing and sharing data, as well as the use of discipline-specific persistent identifiers (PXD ID, UniProt ID, PDB ID), which make individual datasets unambiguously findable and citable.

Specialised databases also illustrate why the 70 reported items cannot simply be understood as 70 independent repositories. Individual services may function as repositories, databases, catalogues or aggregators, and their roles may overlap. For example, the ProteomeXchange infrastructure brings together several repositories for proteomics data, including PRIDE, MassIVE and jPOST, while ProteomeCentral serves as their central catalogue. PRIDE is also part of the EMBL-EBI infrastructure. In practice, a response such as “EMBL-EBI” may therefore refer to data stored in one of the many specialised resources covered by this infrastructure.

The way stored data are identified also varies. General-purpose repositories such as Zenodo or Figshare primarily use DOIs, while disciplinary databases often rely on their own accession numbers. In the life sciences, datasets may therefore have, for example, a PXD identifier in ProteomeXchange or specific identifiers assigned by NCBI databases. In addition to showing the number of services being used, the survey results therefore reveal a highly heterogeneous infrastructure that differs considerably across research disciplines.

The survey results will also be used to explore the possibility of harvesting metadata on Czech research outputs stored in external, including international, repositories into the National Metadata Directory (NMA). The survey may therefore help to better map the sources from which metadata on Czech research data could potentially be harvested into the NMA.

Alongside repositories and disciplinary databases, respondents also reported storage methods that unfortunately do not support the effective long-term management and accessibility of research data. Nine respondents use internal or local institutional storage, eight reported OneDrive, SharePoint or Microsoft Teams, and seven use flash drives or external USB drives. Other frequently mentioned services include GitHub, Google Drive, Figshare and the Open Science Framework.


Data Most Often Accompany Publications

The second part of the survey focused on the stage at which researchers deposit their data. Most commonly, these are data used as supporting material for publications. This option was selected by 127 respondents, representing approximately two thirds of all participants. Raw data are deposited by 88 respondents, while 84 deposit data for long-term archiving. Final processed and validated data were reported by 79 respondents, or 41%, while 55 respondents deposit partially cleaned or calibrated data.

Specifically modified datasets are deposited less frequently. Aggregated data were reported by 31 respondents, anonymised data by 26, and pseudonymised data by only ten. In total, 18% of respondents stated that they do not deposit data in a repository at all.

The responses therefore do not point to a single standard approach to data storage. Respondents reported a total of 67 different combinations of stages at which they deposit their data, highlighting considerable variation in practice across the research environment. Even the most common responses do not dominate significantly: 24 respondents stated only that they do not deposit data in a repository, while another 17 deposit exclusively supporting data for publications. According to the survey results, data deposition is therefore often closely linked to the publication process. At the same time, a wide range of other approaches is also represented, from depositing raw data and final processed versions to long-term archiving. This diversity reflects not only differences between disciplines, but also varying practices across individual research teams and institutions.


The Survey Continues

The survey remains open, and researchers can still share their experiences with storing research data. Newly collected responses will be evaluated as part of the next round of data collection.

Czech Version of the Survey English Version of the Survey


Note: The composition of the survey sample provides important context for interpreting the results. Researchers from the University of Chemistry and Technology Prague and Masaryk University were the most strongly represented, while a significant proportion of respondents also came from the Institute of Molecular Genetics of the Czech Academy of Sciences, the Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, and the Institute of Physics of the Czech Academy of Sciences. Due to the uneven representation of individual institutions, the data cannot be generalised to the entire Czech research environment. The results should therefore be interpreted as a snapshot of the current practices of the respondents involved, capturing different approaches to research data storage and the use of related services.


More articles

All articles

You are running an old browser version. We recommend updating your browser to its latest version.