The Original Paper and Its Unsettling Proliferation
The 2017 paper Attention Is All You Need, authored by Vaswani et al., is a cornerstone of modern AI, introducing the Transformer architecture that now underpins everything from large language models to advanced machine translation. Its impact is undeniable, evidenced by over 140,000 citations. However, a recent query on Crossref, the digital object identifier (DOI) registration agency, revealed a disturbing anomaly: five distinct records identical to the original paper have appeared, all bearing the same title, the same eight authors, and crucially, a publication year of 2025.
These five duplicate records are not mere misprints. Each possesses a unique DOI prefixing 10.65215, a prefix registered not with a major computer science publisher, but with an entity identified as the Shenzhen Medical Academy of Research and Translation. The papers themselves are hosted via a Chinese preprint server. At the time of this writing, each of these DOIs resolves to a live, indexed web page, presenting a clear case of academic identity theft or, at the very least, significant metadata manipulation.
Unpacking the Anomaly: DOIs, Dates, and Discrepancies
The implications of this discovery are far-reaching. The original Attention Is All You Need paper was published in 2017 and has been a stable, verifiable academic artifact ever since. The sudden appearance of five byte-for-byte identical copies, all dated 2025, raises immediate questions about academic integrity, the robustness of metadata systems, and the potential for malicious actors to inject fabricated research into the scholarly record. The DOIs themselves, such as 10.65215/r5bs2d54, 10.65215/ysbyhc05, 10.65215/mdcm8z23, 10.65215/nxvz2v36, and 10.65215/2q58a426, are all registered under the same suspicious prefix and point to the same erroneous publication year.
The Shenzhen Medical Academy of Research and Translation's involvement, via its chosen preprint host, is particularly perplexing. While medical research often intersects with AI, the replication of a foundational computer science paper under this umbrella, and with a future publication date, suggests a deliberate act. It is not uncommon for preprints to be hosted on various platforms, but replicating a landmark paper with identical author lists and future dates is unprecedented and deeply concerning. This isn't just about a paper being available on multiple sites; it's about the systematic creation of false metadata linked to a critical piece of AI literature.
Broader Implications for Academic Publishing and AI Research
This incident is more than just a digital curiosity; it strikes at the heart of how scholarly work is cataloged, discovered, and trusted. For developers and researchers building upon the Transformer architecture, the existence of these fakes could lead to confusion, misattribution, or even the citation of non-existent or misrepresented work. Imagine a researcher citing one of these 2025 papers, only to find that the underlying work is not yet published or is a fraudulent copy. The integrity of citation counts, academic performance metrics, and the very foundation of the scientific record are at stake.
The fact that Crossref, an organization dedicated to maintaining the integrity of scholarly metadata, is involved is particularly troubling. While Crossref itself does not host content, it serves as a central registry for DOIs. The issue here lies with the registration agency (Shenzhen Medical Academy of Research and Translation) and the preprint host. This situation highlights a potential vulnerability in the system: how can registration agencies ensure the legitimacy of the content they are associating DOIs with, especially when a paper is so well-known and its metadata could be easily mimicked?
Furthermore, the choice of a future publication date (2025) for identical copies of a 2017 paper is a bizarre and potentially strategic choice. It might be an attempt to obscure the fraudulent nature of the papers by placing them in the future, or perhaps to create a false sense of legitimacy by implying an upcoming, updated version. Regardless of the intent, it adds another layer of manipulation to an already suspicious scenario.
What nobody has fully addressed yet is the potential for these fabricated records to influence AI models that are trained on scholarly metadata. If future AI systems ingest these fraudulent entries as legitimate, it could subtly skew our understanding of the development timeline and origins of key AI technologies. This incident serves as a stark reminder that as AI becomes more integrated into research workflows, the integrity of the data sources it relies upon becomes paramount.
This situation demands a thorough investigation by Crossref and the Digital Object Identifier Foundation. The Shenzhen Medical Academy of Research and Translation and the involved preprint host must provide an explanation for these duplicate records. Until then, researchers and developers should exercise extreme caution when encountering records for Attention Is All You Need that deviate from the original 2017 publication details.
