Other models are available - MetaROR review

By Bianca Kramer

September 15, 2026

Response to Market dynamics, governance and open research metadata in the AI era - Daniel Hook ( https://doi.org/10.48550/arXiv.2604.19507)

Originally published at MetaROR: https://doi.org/10.70744/MetaROR.427.1.rv1

For additional perspectives, see the two other MetaROR reviews by Cameron Neylon and Neil Jacobs.

Review

The paper ‘Market dynamics, governance and open research metadata in the AI era’ introduces and discusses the ‘innovation annulus’ as a zone of closed structured metadata that separates a core of fully open metadata and an advancing frontier of refined knowledge products, and argues that the annulus exists because the cost of producing and refining structured knowledge data is real and persistent, shaped by production frictions that technology reduces but cannot eliminate.

In this review, I focus on a number of counterarguments to the premise of the article. The review does not go into detail on the mathematical details of the model presented, but instead, hopes to contribute to a discussion on the assumptions underlying the model as a whole.

From zero-sum game to win-win scenario?

The author argues that the debate about scholarly knowledge infrastructure has traditionally been framed as a zero-sum game between openness and commercial enclosure, with every advance in openness a retreat for commercial interests and vice versa.

The proposed model is presented as a positive sum game (a win-win scenario) where a base layer of structured metadata is openly available, and development of new or enriched metadata is in the hands of commercial providers. The paper further sees a role for both economics and community governance in determining where the dividing line between the two classes of metadata should lie (recognizing that this can differ for different types of metadata).

This model keeps commercial interests at its center by postulating that innovation can or will only take place in a commercial setting. In effect, this perpetuates a dependency on commercial systems, not only as a locus of innovation, but also as source of structured and refined metadata that kept close (for now) to recoup financial investments and make a profit.

Other models are available

In the paper, the existence of a zone of closed structured metadata is justified by stating that the cost of producing and refining structured knowledge data is real and persistent. In our view, the latter is a given, but the conclusions derived from that in the paper are not.

Provision of structured metadata at source

First, producers of scholarly metadata play an important role in providing structured metadata at source. An obvious example is publishers depositing publication metadata through Crossref, but this also involves institutional and subject repositories that expose metadata for publications, as well as data repositories and software repositories.

The author acknowledges that provision of better structured metadata at source (helped by technical advances, standardization and community norms) does reduce third-party efforts for metadata structuring and enrichment. This increases the proportion of metadata (of a given type) that is openly available, reducing the width of zone of closed structured metadata.

There are, however, potential additional dynamics at play that could influence this provision of open metadata at source. When publishers provide access to full text or JATS XML access to bibliographic databases to use in the extraction and structuring of metadata, there may be less incentive to provide those same metadata openly at source. A similar development has been observed with publishers requesting (open) bibliographic databases to take down abstracts at a time where abstracts are increasingly valuable as training. material for LLMs12. In these cases, the legal and contractual frictions described by the author may contribute to a non-level playing field for metadata structuring and enrichment, and the existence of a closed zone of structured metadata in itself may limit the provision of structured metadata at source.

Alternative financial models for structuring and enriching metadata

Second, where third-party efforts are required to harmonize, structure and enhance scholarly metadata, there are multiple examples of this work being taken on not as commercial activity (with the resulting metadata, at least initially, being kept in the closed zone), but by organizations that operate under different financial models and, make the resulting metadata immediately available as open metadata as part of their ethos and practice.

While the author discusses a limited role for ‘state investment’, the financial models used by infrastructures that provide the metadata they enrich and structure directly as open metadata are much more varied than this term suggests.

For example, OpenAlex receives project funding from charitable funders for innovation, but also financial support from research performing organizations and funders through their institutional membership route, and direct revenue for services provided on top of their database of openly available metadata. OpenAIRE, originally a direct recipient of European Commission funding through consecutive Framework Programmes, has diversified revenue streams through an institutional membership programme, participation in funded projects and direct collaborations with e.g. national governments and library consortia. As a third example, EuropePMC works on longer-term operational funding from a group of both national and charitable funders in the medical domain to link and enrich metadata in a specialized domain, going far beyond publication metadata only.

Crucially, all these models involve decisions by research institutions and funders to financially contribute to the generation of structured enriched metadata that is then made openly available rather than kept closed to be licensed to other users.

A different look at the frontier

The paper argues that more specialized demands for metadata, e.g. for contextualised analytical products, data on research capabilities beyond publication counts and (often domain-specific) data to support corporate R&D activities, require development and innovation by commercial actors, as they “require a quality level that raw open metadata does not yet consistently deliver”.

A counterargument to this would be that, like above, there is no inherent reason why such development and innovation cannot be supported by other financial models, if downstream users would elect to pay for these. The difference is not raw open data versus closed high-quality data, but high-quality data directly released as open data or restricted as closed data. An underlying question could be whether corporate research organizations would consider it in their (competitive) interest to not just pay for access to data, but for the creation of high-quality data that would also be publicly available. Here, it should be noted that any competitive advantage from paying for closed data would be mitigated by the fact that the same data would be available to any other party willing and able to pay for them.

There is an additional argument against considering work on ‘new’ metadata types as naturally in the purview of commercial development because these types of data only serve specialized usage. This is that limited, closed availability of these data types in itself slows down uptake and usage, and by definition excludes lesser resources actors, not because they have no use for these data types, but because they cannot afford access to them. Of the examples given in the paper, funding flow analysis stands out as an area where there is a lot of interest from research funders in the Global South34. Here it should also be noted that access to otherwise closed data for specific users (e.g. in the context of a research project), especially without the right to share the data, does not represent the same value and benefits as true open availability of such data does.

Finally, when considering the ‘frontier’ of specialized metadata and metadata usage, a distinction can be made between the data itself and (analytical or other) applications and services built on top of these data. It could be argued that a financial model that charges for services while having the underlying data openly available would be in line with the Principles of Open Scholarly Infrastructure, at least for this aspect. In addition, it would keep the field open for (competitive) innovation and development to take place on top of open metadata.

The role of community

The paper sets out a role for the scholarly community, and explicitly the Barcelona Declaration, to develop a normative shared understanding of which data types should be inside the ‘open core’, what quality and provenance standards should apply, and what responsibilities producers, consumers and aggregators of metadata should have towards each other.

While there certainly is value in such collaborative discussions, they should not be positioned to implicitly endorse a model where the existence of a zone of closed structured metadata, produced and restricted by commercial actors, is considered both inevitable and inherently beneficial.

As the closing sentence of the paper reads: “The question is whether we govern it wisely: ensuring that (…) the scholarly data ecosystem serves all its users—including those who cannot afford to pay for frontier refinement— as equitably and efficiently as the state of the art allows.”

This ambition deserves the consideration of multiple models for enabling the structuring and enrichment of metadata that truly benefit all users, as well as the role that individual research performing and funding organizations have in deciding where to allocate both financial resources and in kind efforts (e.g. in participating in governance bodies and the integration of data sources in institutional processes).

This is not to discount the potential value of commercial actors in this space (especially when they operate on a service- not data-based revenue model), but to challenge their ‘natural’ role in the provision of high-quality metadata. The argument is not about whether producing metadata is somehow easy or cost-free (it isn’t), but about the different ways this production can be organized and financed.

Final remark

The paper offers a valuable contribution in theorizing and modelling the forces at play in shaping the way scholarly metadata are created, structured, enriched and made available for use by the scholarly community. Further explorations of the model under different assumptions, including those outlined in this response, could contribute to the discussion of the role of different types of actors (including commercial actors) in this space.

About the author:

The author of this response is independent advisort and research analyst at Sesame Open Science (working in the areas of open science, open metadata and open infrastructure) as well as Executive Director of the Barcelona Declaration on Open Research Information. This response is written in a personal capacity, representing my personal views and opinions.


  1. Kramer, B. (2024). More open abstracts? Sesame Open Science. https://doi.org/10.59350/vy0mr-wcc38 ↩︎

  2. Tay, A. (2025) The Petrol Tank for AI Discovery Might be Running Dry as Publishers close access to scholarly content such as abstracts due to AI incentives. Aaron Tay’s Musings about Librarianship. https://aarontay.substack.com/p/the-petrol-tank-for-ai-discovery ↩︎

  3. https://www.clacso.org/fundingflows/; ↩︎

  4. https://idrc-crdi.ca/en/what-we-do/projects-we-support/project/state-science-technology-and-innovation-africa-science ↩︎

Posted on:
September 15, 2026
Length:
9 minute read, 1750 words
See Also: