What a DOI Is, and Why Our AI RPG Benchmark Has One
A DOI is a permanent serial number for a published thing, backed by a promise from somebody other than the publisher that whatever sits at the end of it will never quietly change.
The Arcanum AI RPG Benchmark has one. It is 10.5281/zenodo.21721591, it has had one since 29 July 2026, and in the weeks since we started printing it on the benchmark page, the number of readers who have asked what it means is roughly zero — not because everyone understood it, but because an unfamiliar string of digits next to a familiar word is easy to scroll past.
That is a fair reaction to an acronym nobody explained. So here is the explanation, in ordinary words, along with the more useful thing: a way to use it to check whether we are lying to you.
A DOI Is a Serial Number That Somebody Else Guarantees
Everything on the internet has an address. Addresses are unreliable. The site moves, the company folds, someone reorganises the URLs, the page is quietly edited at two in the morning and the sentence you remembered reading is not there any more. If you have ever followed a link in a four-year-old forum post, you already know the failure.
A DOI — Digital Object Identifier — is the fix the research world settled on. Instead of pointing at a place, you point at a registered record. The identifier is handed out by a registration agency, the agency keeps a lookup table, and you resolve it by sticking doi.org/ in front:
The 10.5281 at the front is the prefix belonging to the repository that issued it. Everything after the slash identifies our specific deposit. If that repository ever reorganised its site, the DOI would still resolve, because the registry — not the repository’s URL structure — is what the identifier is attached to.
The important part is not the permanence of the link. It is the permanence of the contents. A DOI is issued against a fixed deposit. We cannot swap the files underneath it. If we want to publish different numbers, we have to publish a new version, which gets its own DOI, and the old one stays exactly where it is, still resolving, still saying what it always said.
That is the whole trick. It converts “trust us” into “check us.”
Zenodo Is Where the Copy Lives, and It Isn’t Ours
Zenodo is a free, open repository run by CERN — the European particle physics laboratory, the one with the Large Hadron Collider — and built under an EU open-science programme. Researchers use it to deposit datasets, code and papers so other people can cite them permanently. It is one of the standard places to put a thing you want to survive you.
Three properties matter here, and none of them are about prestige:
- It is not our server. If arcanumrpgs.com goes down tomorrow, is sold, or gets rewritten by someone with different opinions, the deposited benchmark is still sitting on CERN infrastructure, unaffected.
- We cannot delete it. Once a Zenodo record is published, the depositor cannot take it back. We can supersede it with a newer version. We cannot make the July numbers stop existing.
- We cannot edit it. The files are frozen at deposit. There is no “just fix that one score” button, and that is the point.
Put those together and you get something a review site does not normally have: a copy of our own claims that we have no power over.
Why This Matters If You Just Want to Pick a Game Tonight
Fair question. You are trying to decide whether to spend fifteen dollars a month on an AI roleplaying platform, and a particle physics laboratory has nothing to do with that.
Here is the connection. Every review site on the internet has the same structural weakness: the scores live on a page the reviewer controls. A rating can be nudged up after a developer sends a friendly email. A harsh line can vanish when a company becomes an advertiser. A prediction that aged badly can quietly disappear so the site looks more prescient than it was. Almost none of this leaves a trace, because the only copy of the old version was on the same server as the new one.
You cannot detect that from the outside. You can only not be able to detect it, which feels the same as everything being fine.
An external archive changes the shape of the problem. We are not asking you to believe we would never edit a score under pressure. We are removing our ability to do it invisibly. If a number moves, the frozen copy is still out there disagreeing with us, and anyone who cares can put the two side by side.
This is also why we publish the rubric itself and why every rated entry states which test tier it got. A score you cannot interrogate is a verdict wearing a number’s clothes. We wrote about that at length when we re-scored every platform in the directory and thirteen of fifteen went down — an uncomfortable thing to publish, and much harder to publish if the old numbers had not been written down somewhere we could be held to.
The Part You Can Check Yourself, Right Now
Enough theory. Here is the archived snapshot doing its job.
The deposited edition is dated 29 July 2026 and covers 15 platforms. The live board today covers 25. Compare them:
| Archived July snapshot | This site today | |
|---|---|---|
| Platforms scored | 15 | 25 |
| Highest composite | 3.9 | 4.4 |
| Lowest composite | 1.4 | 1.4 |
| Strongest axis field-wide | NPC Fidelity, 3.20 | NPC Fidelity, 3.24 |
| Weakest axis field-wide | Mechanical Depth, 2.13 | Mechanical Depth, 2.52 |
| Axes where nothing has scored 5 | Memory, Agency, Longevity | Memory, Longevity |
Two things to notice.
The ceiling moved, and the archive proves we did not pretend otherwise. In July the best platform on our board scored 3.9. Today it is 4.4. If we had wanted to look consistent, the tempting move would be to quietly regrade July so the field looked like it had always had a 4-plus in it. We cannot. The frozen copy says 3.9, and it will say 3.9 forever.
A prediction from July survived six new platforms, and then broke on the seventh. The deposited README makes a specific, falsifiable claim: three of the seven axes — Memory & Continuity, Player Agency and Longevity — have no top score anywhere in the field, while every 5 we have ever awarded sits on NPC Fidelity, Mechanical Depth, Determinism & Fairness or Signature Design. The reasoning was that the claimed axes are the ones a competent team can solve with craft, and the unclaimed ones need the underlying technology to improve.
Six platforms were scored without denting it. The seventh did. On 4 September 2026 NOPOTIONS took Player Agency — 5, the first top score on any of the three, and the July claim is now two-thirds right instead of wholly right.
We would rather show you that than bury it, because it is the entire argument for depositing anything. The README is frozen: it still predicts three unclaimed axes, and anyone can pull the July file and read the sentence that turned out to be wrong. We cannot go back and soften it into “two axes” so the record looks prescient. What we can do is say which part broke and why it matters — and the why is the interesting half. Agency did not fall to better technology, which was the stated reasoning. It fell to restraint: a narrator that describes the scene and then stops, on a map open enough that there is nothing to railroad you back onto. That is craft, not capability, which means the axis was misfiled in July rather than merely unsolved.
The other two held. Memory & Continuity has crept from a field maximum of 4 to 4.25 — held alone by Voyage, and a quarter point is a tie-break, not a shade of opinion — and still nobody anywhere has a 5. Longevity still tops out at 4. Meanwhile Mechanical Depth is still the weakest axis in the category despite having a maximum of 5, which says the field is not uniformly shallow so much as split between a few real engines and a long tail of chat interfaces.
That is what a dated archive buys a reader. Not a nicer number — a claim with a timestamp on it that either survives contact with new evidence or doesn’t, in public, where you can watch. This one didn’t, and you can check.
Access disclosure: Latitude provided beta access to Voyage, and the developers of Craft provided beta access and a subscription. No conditions were attached to any of our coverage. Arcanum is independent and unpaid — we do not accept payment for reviews or ratings.
What’s Actually Inside the Deposit
Three files, and it is worth knowing what they are, because “we archived it” can mean anything from a full dataset to a screenshot.
benchmark.csv— the board. One row per platform, one column per axis, plus the composite, the test tier it received, and a link to its page.benchmark.json— the same scores plus the axis definitions, the field averages, the scale semantics, and the list of platforms we have deliberately not measured.README.md— a data dictionary explaining every column, the method, the inclusion rules, the findings, the disclosures, and the limitations.
The limitations section is the part we would point at if you only read one page of it. It states, in our own words and permanently:
- One rater. Every score comes from the same person applying the same rubric, so there is no inter-rater reliability statistic for this edition, because there is no second rater. We call this the single largest limitation of the dataset.
- Ordinal judgements, arithmetic composite. We average seven ordinal ratings. That is conventional, but the gap between 2 and 3 is not guaranteed to equal the gap between 4 and 5, so composites are ranking information rather than measurements on a ratio scale.
- Card-driven platforms carry extra variance. Where you bring a community-made character card and the platform runs it, a share of the experience is the card, not the product.
- Limited reproducibility, on purpose. We publish the method and withhold the specific probe items, because a benchmark that publishes its test items stops measuring quality and starts measuring who read the test. That is a real cost and we record it as one rather than pretending it away.
- A non-random sample with a short shelf life. These are the platforms one person has played, scored on a specific date, in a category that ships fast.
Writing your own weak points into a file you can never edit is a different act from writing them into a page you can revise. That asymmetry is most of why we did it.
Two DOIs, and Which One You Want
Zenodo issues two identifiers, and they are not interchangeable.
- The version DOI —
10.5281/zenodo.21721591— pins this edition’s numbers, permanently. Use it when you are quoting a specific figure, because a later edition will move that figure and you want your citation to keep meaning what it meant when you wrote it. - The concept DOI —
10.5281/zenodo.21721590— always resolves to the newest edition. Use it when you mean “the Arcanum benchmark” as an ongoing thing rather than one snapshot of it.
The rule of thumb: a number gets the version DOI, the project gets the concept DOI. If you are writing “Arcanum scored it 3.1,” you want the version. If you are writing “Arcanum maintains a seven-axis benchmark,” you want the concept.
The citation line, for anyone who needs to paste one:
The Arcanum AI RPG Benchmark, 2026 H2 edition (archived snapshot, 2026-07-29). Arcanum RPGs. Zenodo. https://doi.org/10.5281/zenodo.21721591 (CC BY 4.0).
The Archive Is Deliberately Behind This Site
You may have noticed the gap already: the deposit covers 15 platforms and the live board has 25.
That is not neglect. The two run on different clocks on purpose. A platform joins the live board the moment it earns a rating — there is no waiting list, entry into the benchmark is the rating — so that page moves several times a quarter. The archive is refreshed quarterly, because an archive that changed as often as the page would defeat its own purpose. A snapshot that keeps re-snapping is not a snapshot.
So between deposits, the site is legitimately ahead of the DOI, and the benchmark page states both dates side by side rather than hiding the discrepancy. Cite the DOI when you need a number that cannot move under you. Cite the page when you want the field as it stands today. Both are correct answers to different questions.
At the next edition the new numbers go up as a new version of the same record, never as a separate deposit — which is what keeps the citation chain intact and lets every earlier edition stay reachable at its own DOI.
What It Costs Us
It would be dishonest to present this as pure upside, so here is the bill.
We can never quietly fix a mistake. If a score in the deposited edition is wrong, that error is permanent and public. The best we can do is publish a corrected version and let both stand. We correct a published edition only for verifiable factual error — never because someone disagreed with a verdict — and the wrong version stays reachable either way.
What the broken prediction implies for the rest of the field — that more of what looks capability-bound may be craft-bound — is argued in the next AI RPG war won’t be about better AI.
Every prediction is on the record. The predictions in our current state-of-the-field edition are made the same way, for the same reason. The three-unclaimed-axes claim held this time. Next edition it might not, and there is no version of that outcome where we get to have not said it.
Anyone can reprint the whole thing. The dataset is CC BY 4.0, which means a competitor can lift the entire table, chart it, and publish it commercially, needing nothing from us but attribution and a link. We chose that deliberately. A licence that forbids reproduction would defeat the purpose of publishing the data at all, and being checked and named beats being tidy and ignored.
It takes real work per edition. Assembling the files, writing the data dictionary, keeping the site and the deposit honest about their differences — none of it is automatic, and none of it makes a single reader’s evening more fun.
We think it is worth it anyway, because the alternative is asking you to take a number on faith in a category where every incentive runs the other way.
What a DOI Does Not Mean
This is the section that matters most, and it is the one an over-eager version of this article would leave out.
A DOI is not peer review. Nobody checked our scores before they were deposited. No reviewer, no editor, no committee.
A DOI is not accreditation or a quality stamp. Depositing on Zenodo requires a login, not an approval. Anyone can deposit anything. If you ever see a site imply that having a DOI makes its claims scientifically validated, that site is misusing the word, and you should discount it accordingly.
A DOI does not make the numbers right. Everything in the limitations section above is still true. One rater is one rater whether or not CERN is storing the file.
What a DOI actually does is narrower, and it is the only thing we claim for it: it fixes the claim, dates it, attaches a name to it, and puts a copy somewhere we cannot reach. It turns an opinion into a record you can hold us to later.
That is a small promise. It is also one we can actually keep, which is more than can be said for most of the trust signals in this category.
How to Audit Us in Sixty Seconds
You do not have to take any of the above on faith either. Here is the whole procedure:
- Open doi.org/10.5281/zenodo.21721591. Download
benchmark.csvfrom the Zenodo record. - Download the current
benchmark.csvfrom this site. - Open both. For any platform in both files, the axis scores should match.
- Where they don’t match, the platform’s review on this site should say the score changed and why. Re-scored entries carry a dated note explaining what they used to score.
- If you find a score that moved with no explanation anywhere, that is a real finding, and we would like to hear about it: [email protected].
Step 5 is not a rhetorical flourish. Disagreeing with a verdict is not a correction — we will cheerfully keep a score you think is too harsh. But a number that moved without a stated reason is exactly the failure this whole apparatus exists to make visible, and if the apparatus ever catches us, publishing that is the point of having built it.
Frequently Asked Questions
What is a DOI?
A DOI, or Digital Object Identifier, is a permanent serial number for a published thing. It looks like 10.5281/zenodo.21721591 and you open it by putting doi.org/ in front of it. The difference between a DOI and an ordinary link is who guarantees it. A link points at a place, and whoever owns that place can change what sits there or delete it. A DOI points at a registered record held by a third party, and the registration is the promise that the thing at the other end will still be there and will not have been silently edited.
What is Zenodo?
Zenodo is a free open research repository run by CERN, the European particle physics laboratory, and built under an EU open-science programme. Researchers use it to deposit datasets, papers and code so that other people can cite them permanently. It issues a DOI for every deposit and it keeps the files whether or not the person who uploaded them still has a website. Once a record is published there, the depositor cannot delete it. A newer version can supersede it, but the old version keeps its own DOI and stays reachable.
Does a DOI mean the benchmark is peer reviewed?
No, and we would rather say so than let the acronym do work it has not earned. Nobody checked our scores before they were deposited. Depositing on Zenodo requires a login, not an approval, and a DOI is not accreditation, peer review, or a quality stamp of any kind. What it does is fix the claim in place, date it, attach a name to it, and put a copy somewhere we cannot reach. It does not make the numbers correct. It makes them checkable, which is a smaller promise that we can actually keep.
Why is the archived snapshot behind the live benchmark page?
Because they run on different clocks on purpose. A platform joins the live board the moment it earns a rating, so that page changes several times a quarter. The Zenodo deposit is refreshed quarterly, so between deposits the site is legitimately ahead of the archive. The archived edition is dated 29 July 2026 and covers 15 platforms; the live board has 24. The gap is stated on the benchmark page rather than papered over. Cite the DOI when you need a number that cannot move under you, and cite the page when you want the field as it stands today.
How can I check that Arcanum has not quietly changed a score?
Open the DOI, download benchmark.csv from the Zenodo record, then download the current benchmark.csv from arcanumrpgs.com and compare the two. Any platform present in both files should hold the same axis scores unless a review on the site says the score was changed and why. That comparison takes about a minute and it does not require our cooperation, which is the entire reason the archived copy exists.
Can I reuse the benchmark data?
Yes. The dataset is published under Creative Commons Attribution 4.0, which lets you reprint, chart, redistribute or build on the whole table, commercially or not, as long as you credit Arcanum with a link. That licence covers the dataset specifically; the articles on this site are governed separately by our terms. We would rather be cited and argued with than kept tidy and ignored.
The short version, for anyone who scrolled: every rating on this site runs through the same published rubric, and a frozen copy of the benchmark sits somewhere we cannot edit it. If we ever moved a score to make ourselves look better, the old copy is still there, and anyone can pull it up and catch us.
That is the whole claim. Not that we are right — that we can be checked. Nothing here is just vibes.
The board is here, the rubric is here, and the frozen copy is here.