“Poorly described data are like a book shelved in the wrong place. They practically don’t exist,” explains Petra Černohlávková.

A researcher may spend years collecting valuable data. But if those data remain on a personal drive without proper description, they effectively do not exist for anyone else. Metadata determine whether data can be found, understood and reused, and their importance is growing further with the rise of artificial intelligence. “For AI tools, well-structured and rich metadata are a gold mine,” says Petra Černohlávková, Head of the Repositories and Metadata Management Department at the National Library of Technology. In the interview, she explains why simply storing data is not enough, how the right repository can make a difference, and where AI can already reliably take over part of the work.

18 Sep 2026 Karolína Smetanová

No description

Metadata are one of the most frequently used terms in research data management. Yet many people are still unsure what exactly the term refers to. How would you explain metadata in simple terms, and why are they so important for research?

Metadata are all around us. In grocery stores, they can describe a product’s name, price or code. In transport apps, they may include information about a passenger and their travel history, while in healthcare systems they can refer to information about patients and their medical history. Put simply, metadata are information about objects, whether physical or digital.

In our context, research data are the objects being described, and if we want to work with them effectively, it is essential to describe them properly using well-structured metadata. It is like a book that has been shelved in the wrong place in a library. If it is placed on the wrong shelf, it is practically lost. No one would look for Čapek’s Dášeňka in the organic chemistry section. The same applies to data. If they are not described and identified at all, they practically do not exist. If they are described poorly, the chances of someone finding them decrease significantly.


Researchers often pay close attention to the data themselves, but less to how they are described. What consequences can poor-quality or missing metadata have? Could you give an example of a situation where this made it difficult or impossible to use or share the data effectively?

The consequences can be significant. Let me illustrate this with the example of a department that had measurement data stored in cabinets for several years. Apart from the researcher concerned, no one knew that the data existed, how they were structured, or what exactly they contained. As a result, the potential of data generated through publicly funded research remained untapped. Collecting such data is often costly, and storing them on individuals’ local drives can in practice mean that no one else will ever be able to access them. Poor metadata management can also make data reuse more difficult, for example when information about licences or conditions of use is missing or inaccurate. This may not necessarily prevent further use altogether, but it can make it considerably more complicated and expensive, whether in terms of time or, for example, legal costs.


Where do you think researchers most often make mistakes when working with metadata? Are there a few key principles that everyone should follow if they want their data to remain usable in the long term?

“Poor metadata management can also make data reuse more difficult, for example when information about licences or conditions of use is missing or inaccurate. This may not prevent reuse altogether, but it can make it significantly more complicated and costly, whether in terms of time or, for example, legal expenses.”

I would not say that researchers necessarily make mistakes, but rather that they often underestimate the overall role of metadata and the related persistent identifiers. Together, they form an infrastructure that enables metadata to be shared effectively across the global research, development and innovation ecosystem. For example, using the ORCID iD researcher identifier can reduce administrative burden throughout the research process – when applying for a grant, depositing research data, publishing an article and reporting research outputs. The real value lies in the integration of different types of systems and identifiers. A repository with well-designed integrations can make depositing research data much easier. Through the ORCID machine interface, it can retrieve information about authors, use a DOI to obtain information about related outputs and grants, and so on.

There are many principles for working with metadata, but the foundation is choosing the right repository. By selecting a repository, researchers are also choosing which metadata they are required and able to provide, as well as which identifiers can be used. General-purpose repositories such as Zenodo will never be able to offer the same level of metadata richness as disciplinary repositories and therefore cannot support the same level of advanced data analysis directly within the repository. Researchers can seek support with these issues from a data steward at their institution or research team, as well as from the various support nodes of the National Data Infrastructure (NDI).


Supporting Components of the NDI

identifikatory.cz ccmm.cz

Metadata are also closely linked to the FAIR principles. How important are metadata in putting these principles into practice? Can we say that without high-quality metadata, data cannot truly be FAIR?

Yes. Without high-quality metadata, we cannot really speak of FAIR data at all. Metadata and persistent identifiers are fundamental building blocks of FAIR. If data are not surrounded by a layer of machine-actionable metadata and all key objects related to the dataset are not properly identified, important context is missing that is essential for smooth and unambiguous further use. As a result, the data will be neither findable, accessible, interoperable nor reusable.


The National Library of Technology is one of the institutions that has long supported research data and metadata management in Czechia. What questions or challenges do researchers and research institutions most often approach you with? Have you noticed any changes in their attitudes towards metadata in recent years?

The range of questions is broad, from theoretical aspects of metadata management and metadata interoperability between different systems to the practical implementation of metadata in various repository systems. Approaches to metadata are changing significantly, particularly thanks to the EOSC CZ initiative, which brings together key stakeholders from the Czech research, development and innovation ecosystem to build the National Data Infrastructure. The discussion is gradually moving from theory towards practical questions and the implementation of concrete solutions in the domain-specific repositories currently being developed.


As artificial intelligence continues to develop, the issue of quality is being discussed more and more often. What role do you think metadata will play at a time when AI tools are increasingly working with scientific data?

For AI tools, well-structured and rich metadata are a gold mine. If the metadata are high-quality and accurate, the results produced by AI can also be more precise. Additional value comes from metadata structured according to established Semantic Web practices, which enable machines to determine the context of the entity being described. Take the concept of migration, for example. Data on bird migration will not necessarily be relevant to a researcher studying human migration. To prevent data that appear to concern the same topic from being mixed together, high-quality machine-readable metadata descriptions at the level of linked data are essential.


Can artificial intelligence or other automated tools already help researchers create metadata in practice? And where does the human role remain indispensable?

“Approaches to metadata are changing significantly, particularly thanks to the EOSC CZ initiative, which brings together key stakeholders from the Czech research, development and innovation ecosystem to build the National Data Infrastructure. The focus is gradually shifting from theoretical discussions towards practical questions and the implementation of concrete solutions in the domain-specific repositories currently being developed.”

Yes, they can. The use of persistent identifiers mentioned above already helps to automate metadata workflows significantly, even without the use of artificial intelligence. AI can then assist by generating draft descriptions of datasets or validating information that has been entered. For now, however, the final decision should still be made by a person who understands the context and can assess more accurately whether something is actually an error. The human role is also indispensable when setting up metadata practices within a particular ecosystem, whether that is a repository, an institution or a research team. AI is a valuable support tool, but in my view, people are still better able to understand the full complexity of the context.


How do you think Czech research institutions’ approach to metadata management has changed in recent years? Is this area moving in the right direction, or is there still a long way to go?

We are moving from a decentralised approach fragmented across disciplines towards deeper cooperation. Many institutions are becoming increasingly aware of the importance of high-quality metadata descriptions. Although we are moving in the right direction, FAIR research data management is a complex area, and there is still a long way to go before we can say that rich, high-quality metadata are being produced as standard, with their creation automated to the greatest possible extent.


If you had to convince a researcher who sees metadata as just another administrative burden, what would you say to them? And what would you like to see change in metadata management in the near future?

I would tell them that they have probably chosen the wrong repository and suggest that they contact their data steward to look for ways to optimise and automate the process. I would also use a practical example to demonstrate the consequences of poor-quality or missing metadata. I would like to see more metadata specialists in this field. It is a highly specialised area at the intersection of computer science and library and information science, and there are currently not enough such experts on the Czech labour market.

“The final decision should be made by a person who understands the context and can assess more accurately whether something is actually an error. The human role is also indispensable when setting up metadata practices within a particular ecosystem, whether that is a repository, an institution or a research team. AI is a valuable support tool, but in my view, people are still better able to understand the full complexity of the context.”

No description

Petra Černohlávková


is Head of the Centre for Repositories and Metadata Management at the National Library of Technology, expert lead for the key activity Methodological Support for Research Data Metadata and PIDs within the IPs CARDS project, and lead for the National Repository within the NCIP project. She has been working in metadata and data curation for more than ten years, including five years in a leadership role. Her experience spans librarianship, institutional repositories, national services, and the coordination of research data metadata harmonisation in Czechia. She studied Library and Information Science at Charles University. She also brings this expertise to the Metadata Working Group, which she leads with the aim of opening it up to the wider research community and focusing on the further development of the CCMM metadata model, the creation of crosswalks, and the use of controlled vocabularies within the National Data Infrastructure (NDI). Thanks to the Working Group’s close connection with the CARDS project, combining these roles will support a smooth flow of information and a consistent direction in the area of metadata.


More articles

All articles

You are running an old browser version. We recommend updating your browser to its latest version.