When people want to know whether a piece of research is any good, they usually have two options: ask an expert, or count citations. Both have well-known problems. Expert review is slow, expensive, subjective, and does not scale to portfolios of thousands of papers. Citations take years to accumulate, reward the already well-cited, and say more about whether a paper was used than whether it said anything new.
That second gap is the interesting one. Creative work, including scientific discovery, is generally understood to be both new and useful. Citations are a reasonable proxy for usefulness. We have had no widely trusted, scalable measure for new. Funders say they want novel research, and then evaluate it with tools that cannot see novelty.
Earlier this year an international challenge set out to change that, and the winning approach now exists as a working prototype. We are considering adding a novelty score to OpenAlex. Before we write a spec or scope resources to implement it, we want to know whether you would use it, what for, and what would make you trust it.
The challenge
The Metascience Novelty Indicators Challenge was launched in September 2025 by the UK Metascience Unit at UK Research and Innovation; designed by the Science Policy Research Unit (SPRU) at the University of Sussex and RAND Europe; and operated by Challenge Works. Teams were asked to build a tool that could assess the novelty of a research paper at the time of publication.
To evaluate the entries, SPRU built a ground-truth dataset the hard way: a random sample of recent publications across all disciplines (drawn from OpenAlex), each rated for novelty by human experts in the relevant field who were provided with the full text. Challenge entrants never saw those ratings. Their indicators were scored on how closely they matched the experts’, and the evaluators were blinded to who had submitted what.
The winner, announced just a few months ago in June 2026, was a team from Forschungszentrum Jülich in Germany: Jan Göpfert, Samuel Kieling, Jann Weinand, Titan Hartono, and Patrick Kuckertz. Their approach came top on most evaluation criteria (in a “winner of winners approach”), and they received a £300,000 prize, provided by Coefficient Giving, to develop it further.
Coverage: https://www.nature.com/articles/d41586-025-01882-7 and https://www.science.org/content/article/how-novel-research-paper-competition-quantify-concept-crowns-winner.
How the Jülich indicator works
Most existing novelty indicators work from metadata: unusual combinations of references, keywords, or journals. The Jülich indicator reads the paper instead.
It uses a large language model to analyse the study and a selection of the works it cites, and from those, it reconstructs the state of knowledge at the time of publication, including the open questions in the field. It then asks what the focal study contributes anew: an original method, a surprising result, or a solution to a problem that was previously unsolved. It deliberately collects arguments for and against the paper’s novelty and weighs them. The output is a score from 0 to 100, an interval showing how confident the model is, and a written justification that a reader can check.
Jülich press release: https://www.fz-juelich.de/en/news/archive/press-release/2026/ai-identifies-scientific-novelty-julich-team-wins-international-challenge
The team is clear that this is a support for human judgement rather than a replacement for it, and that a score which could be gamed, or which quietly disadvantaged some fields or languages, would be worse than no score at all. Those are our concerns too.
Why OpenAlex is interested, and what it would take
OpenAlex already carries citation counts, topics, and a growing set of work-level signals (e.g., field-weighted citation impact). A novelty score that is transparent, reproducible, and available for every work, free, in the same place, would be a genuinely new thing for funders, institutions, journals, and researchers to work with. The UK Metascience Unit has told us that it would be of interest to a large proportion of the roughly 100 research teams it funds, as well as to UKRI’s own analysts. And we would rather that such a novelty indicator be built in the open than watch it appear behind a paywall.
It is not a small undertaking, and the open questions are the reason for this post:
- How far back? Scoring every work in OpenAlex is very different from scoring everything published since, say, 2019.
- Which works? Journal articles only, or preprints, conference papers, book chapters, and dissertations as well? The indicator was validated so far only on journal articles. Maybe that’s enough (after all, FWCI only applies to a few work types), but maybe it’s worth the effort to evaluate the indicator in a broader set of work types first.
- One-off or continuous? A static file scored once is much cheaper than a live score that updates as new works arrive and as records change. A live production database like OpenAlex really wants the latter, but maybe the community only needs static scores, updated yearly.
- Open-weight or commercial models? Open-weight LLMs would make the scores reproducible over the long term and could run on public research computing. Jülich is testing whether they perform as well.
- Cost and compute. Running a language model over hundreds of millions of works is a significant compute bill, however it is done. That’s why answers to the questions above are especially important before we fully spec this out.
- Transparency versus gaming. Publishing the full method makes the score checkable and makes it easier to write to. Where should the line sit and is that an argument for static historic application vs. on-going scoring of the indicator against new works?
- Use in evaluation. A score built to describe a literature can be misused to judge individuals. What safeguards would you want before it existed at all?
What we are asking
This is a vibe check with a light-touch questionnaire, not a full consultation. We want enough signal on demand and scope to write a spec and a resourcing case, and we will go deeper later if the response warrants it.
Fill in the form. About six questions, five minutes, any language. It stays open until Friday 13 November 2026.
Join the webinar. Wednesday 7 October 2026, 8:00 to 9:00am Pacific (11am Eastern, 4pm UK, 5pm Central Europe). Ben Steyn (UK Metascience Unit) on why the challenge was run, Jan Göpfert (Jülich) on how the indicator works, and Kyle Demes (OpenAlex) on what integrating it into OpenAlex would look like, followed by half an hour of questions. Sarah Otner (SPRU) will be on hand for questions about the evaluation. Recording and slides posted afterwards.
https://zoom.us/webinar/register/WN_EetlA-l3QgK2FvQVxE0GLg
Pass it on. Funders, research offices, metascience groups, tool builders, and anyone who has ever wished that citations told them something they do not.
What happens next
We will share a synthesis of the responses with the UK Metascience Unit and the Jülich team in late November, and publish a summary here. If there is clear demand and a workable scope, the next step is a spec and a scoped first run. If there is not, we will say so and keep the door open.
Thanks in advance for any input you share. If you have any other questions about this initiative, let me know at: kyle@openalex.org