*This post was written with Claude. Every number and example was verified by a human against the live data. Unlike most scholarly databases, OpenAlex doesn’t start from journals. Other databases pick a journal, index it, and attach its articles — so their journal metadata is clean by construction. OpenAlex works the other way: we index […]
The post A major cleanup of journal records in OpenAlex appeared first on OpenAlex blog.
*This post was written with Claude. Every number and example was verified by a human against the live data.
Unlike most scholarly databases, OpenAlex doesn’t start from journals. Other databases pick a journal, index it, and attach its articles — so their journal metadata is clean by construction. OpenAlex works the other way: we index scholarly works first, then connect each one to the rest of the research graph — authors, institutions, and publication venues. Venue information arrives from many different sources, unstandardized and often without persistent identifiers. That’s how one journal can quietly become two or three records. And because source metadata is critical for many of our users — especially librarians — it was time for a comprehensive cleanup.
Over the past two weeks we merged 27,728 duplicate source records — the same journal listed under a translated name, an old name, or a small spelling difference, its papers split among the copies — correcting the venue records of more than 6 million works. After the cleanup, those papers live under one journal: one complete works count, one citation profile, one page. The active catalog went from 282,924 source records to 255,434; every removed record was a duplicate, and no real venue lost its page.
Three examplesChemischer Informationsdienst → ChemInform. Wiley renamed this chemistry alerting service decades ago — but the catalog carried both names as separate journals, splitting fifty years of chemistry down the middle: 307,000 works under the German name, the rest under the English one. It’s one record now, 794,000 works, with the German identity preserved as an alternate title and searchable in both languages.
Journal of the American Medical Association → JAMA. Even the world’s best-known journals weren’t immune: 85,611 papers were filed under the spelled-out name as if it were a separate journal from JAMA. Roughly a sixth of JAMA’s output was invisible from its own page. One record now — 354,735 works — and the spelled-out name remains searchable as an alternate title.
镇江医学院学报 → Journal of Zhenjiang Medical College. This journal existed twice in the catalog — once under its Chinese name, once under its English one. The telling detail: the duplicate’s publication activity stops in 2001 — the exact year Zhenjiang Medical College merged into Jiangsu University. The bibliographic data echoed institutional history; the two records are now one.
How it worksThe hard part of this work isn’t finding lookalike records, it’s that academic publishing is full of journals with similar names that must not be merged. Every merge decision ran through four independent layers:
More than 158,000 examined pairs were confirmed distinct — disjoint ISSN families, different publishers, independent publication histories — and recorded, so future passes build on settled ground instead of starting over. This was a first comprehensive pass; as metadata improves, more merges will follow.
What you’ll noticeNew records arrive every day, and the pipeline that did this — find candidates, judge on full evidence, verify independently — is now a repeatable tool for keeping the catalog clean. Next up for source metadata: cleaning up the publisher and host-organization lists, and improving how works get matched to their journals in the first place.
| # | Наименование новости | Тональность | Информативность | Дата публикации |
|---|---|---|---|---|
| 1 | A big improvement to our corresponding-author data | 0 | 14.9 | 23-06-2026 |
| 2 | An Overhaul of Type Classification | 0 | 9.82 | 15-07-2026 |
| 3 | Opening your research funding data: a practical guide for funders | 0 | 7.89 | 13-04-2026 |
| 4 | Linking the world’s research to the code it runs on | 0 | 6.45 | 04-08-2026 |
| 5 | Who funded this dataset? Let’s ask DataCite. | 0 | 6.52 | 16-07-2026 |
| 6 | When affiliation errors become a research security problem | 0 | 8.19 | 27-07-2026 |
| 7 | Recommitting to the Principles of Open Scholarly Infrastructure (POSI) | 0 | 5.81 | 30-03-2026 |
| 8 | The Hakai Institute, as seen by OpenAlex | 0 | 6.61 | 15-06-2026 |
| 9 | Q2 2026 Town Hall: What We Shipped and What’s Next | 0 | 11.86 | 25-04-2026 |
| 10 | Découvrez le modèle OpenAlex pour Lodex | 0 | 11.33 | 03-06-2026 |