We’ve made a major, systematic improvement to how OpenAlex finds and assigns corresponding authors, and the corresponding institutions tied to them. On a hand-checked gold standard, precision rose from 0.60 to 0.92 while recall held nearly steady (0.91 to 0.88), lifting our F1 score from 0.72 to 0.90. We also added real corresponding authors to […]
The post A big improvement to our corresponding-author data appeared first on OpenAlex blog.
We’ve made a major, systematic improvement to how OpenAlex finds and assigns corresponding authors, and the corresponding institutions tied to them. On a hand-checked gold standard, precision rose from 0.60 to 0.92 while recall held nearly steady (0.91 to 0.88), lifting our F1 score from 0.72 to 0.90. We also added real corresponding authors to roughly 7 million works that were missing them. If you use OpenAlex to track transformative agreements, negotiate with publishers, or attribute output to the institution that led a paper, this one matters.
Here’s what we did, how we measured it, and what it means for your work.
Why corresponding authors matterThe corresponding author is usually the person who submitted the paper and who handles publishing decisions, including who pays any article-processing charge. For libraries and consortia, that makes corresponding-author and corresponding-institution data the backbone of transformative-agreement (read-and-publish) tracking: eligibility for most of these deals is determined by the corresponding author’s affiliation. If that field is wrong or missing, the analysis underneath a negotiation is wrong or missing too.
This work grew directly out of a collaboration with the University of California’s California Digital Library (CDL), who resourced the project. CDL flagged that gaps and errors in corresponding-author coverage were constraining their publishing analyses and negotiation strategy, and partnered with us to fix it at the source rather than work around it. We’re grateful for their support and for pushing us toward a problem whose solution benefits the whole community.
Building a gold standard, then improving against itThe foundation of this project was measurement. We built a hand-checked gold standard of corresponding author assignments and used it as ground truth, both to see how we were actually doing and to drive systematic improvements.
OpenAlex marks corresponding authors with the is_corresponding flag on each authorship, and exposes them on a work through corresponding_author_ids and corresponding_institution_ids. Against the gold standard, our starting point was a precision of just 0.60 (paired with a recall of 0.91), for an F1 of 0.72. In plain terms: while we correctly found 91% of true corresponding authors, when we labeled an author as corresponding it was only correct 60% of the time. That gave us a clear baseline and a way to test every change we made.
With that baseline in hand, we focused on the core problem: getting much better at recognizing and attributing corresponding-author information from the unstructured text of the records and documents we ingest. Corresponding-author details are usually present somewhere, but they’re expressed in messy, inconsistent, free-text ways rather than in a clean structured field. We substantially improved our ability to recognize and extract that information, which let us assign genuine corresponding authors to millions of works where we previously had none, and correct many where we had the wrong one.
What the gold standard taught us about an old assumptionThe gold standard also let us test an assumption baked into our older system. When we couldn’t find explicit corresponding-author information for a paper, we used to fall back on treating the first author as corresponding. We knew that this assumption wasn’t perfect and was becoming less reliable in recent years (particularly after transformative agreements linked APC fees and payments to corresponding authors), but this exercise helped show how unreliable that assumption had become. The first author does turn out to be the corresponding author more than half the time, but that fallback was wrong in nearly half of cases. Leaning on that assumption propped up our previous recall (0.91) of corresponding authors, but is why our previous precision was so low (0.6).
Through this work, we were able to remove the first-author fallback entirely, and our recall barely moved: it went from 0.91 to 0.88. Normally, throwing away a blanket assumption like that would tank recall. It didn’t, because our improved text recognition now finds the real corresponding author in nearly all the cases the assumption used to paper over. Precision, meanwhile, climbed from 0.60 to 0.92. We kept almost all of our coverage because we no longer needed to rely on the old assumption. Where we now have genuine evidence we use it; where we still don’t, we say so (i.e., null values over assumed ones). The result is more accurate and reliable data.
The resultsMeasured against the hand-checked gold standard:
| Metric | Before | After |
| Precision | 0.60 | 0.92 |
| Recall | 0.91 | 0.88 |
| F1 | 0.72 | 0.90 |
We want to be straight about the limits. The cases we still struggle with are mostly very small publishers, where the corresponding author is mentioned only inside an email address or tucked into an obscure corner of the record, with no consistent pattern to recognize. We’ll keep working on the long tail, and as always, you can help by curating records where you spot a problem (send error reports to support@openalex.org).
What this means for youIf you rely on corresponding_author_ids or corresponding_institution_ids, your results should be both more accurate and more complete than they were.
If you’ve run these analyses recently, you’ll want to repeat them. The nice part is that all of the queries you’ve already built will work exactly the same way; they’ll just return more accurate data now.
A couple of things to keep in mind:
For transformative-agreement tracking specifically, the combination of more correct corresponding authors and more correct corresponding institutions should give a cleaner picture of which articles belong to which institution, and a firmer footing for the analyses that support your negotiations.
Questions, or a case where it still looks wrong? We’d love to hear it: support@openalex.org.
Finally, our thanks again to the University of California and the California Digital Library for resourcing this work and for collaborating with us on a problem that helps the entire OpenAlex community. If your institution is keen to explore a similar collaboration with us, we’d love to hear from you: reach out to kyle@openalex.org.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | A major cleanup of journal records in OpenAlex | 0 | 7.07 | 29-07-2026 |
| 2 | An Overhaul of Type Classification | 0 | 9.82 | 15-07-2026 |
| 3 | Opening your research funding data: a practical guide for funders | 0 | 7.89 | 13-04-2026 |
| 4 | When affiliation errors become a research security problem | 0 | 8.19 | 27-07-2026 |
| 5 | Linking the world’s research to the code it runs on | 0 | 6.45 | 04-08-2026 |
| 6 | Who funded this dataset? Let’s ask DataCite. | 0 | 6.52 | 16-07-2026 |
| 7 | Recommitting to the Principles of Open Scholarly Infrastructure (POSI) | 0 | 5.81 | 30-03-2026 |
| 8 | The Hakai Institute, as seen by OpenAlex | 0 | 6.61 | 15-06-2026 |
| 9 | Q2 2026 Town Hall: What We Shipped and What’s Next | 0 | 11.86 | 25-04-2026 |
| 10 | Медицинский ИИ OpenEvidence оказался надежнее универсальных моделей в работе с источниками | 0 | 10.37 | 13-08-2026 |