Fieldnote

Why semantic equivalence is not identity equality

Two records that look alike are still two records

Federation brings duplicates into view. Two Semantic Spaces describe a grain mill in nearly the same words; an external catalogue asserts sameAs between two records; a migration finds two legacy entries built from the same source file. The obvious move is to merge them. RSM ID deliberately does not let that move happen on its own, and the reason says a lot about what identity is for.

Semantic equivalence is a judgement about descriptions: these two records seem to talk about the same thing. Identity equality is a fact about referents: these two identifiers denote one resource. The first is cheap, probabilistic, and often right. The second is meant to be permanent, because other systems will build relationships, citations, and permissions on top of it. A merge that turns out to be wrong cannot be quietly undone once a thousand references have followed it.

Where similarity misleads

The fictional Valley Grain Mill 451.GP6FR9ZG54T51AZF has a cleaning and milling line. A second network's record describes "the Valley mill's cleaning line" with the same address and capacity. It might be the same line — or the decommissioned original line, which has its own retired identity, or a new line installed in the same building. Descriptions are similar precisely in the cases where getting the referent right matters most. Migration multiplies the risk: Document 08 warns that similar titles, shared subjects, equivalent descriptions, or common source files do not establish identity equivalence.

What an equivalence decision needs

So equivalence is possible, but it is a governed act rather than an inference: it needs authority, evidence, and a record. External identifiers follow the same discipline. A DOI or ORCID can be associated with an RSM resource through relationships such as sameAs, identifiedBy, or derivedFrom, "subject to evidence and equivalence rules", and RSM never silently treats an external identifier as an RSM-issued RID.

Neither document yet specifies the procedure, the evidence format, or what becomes of the identity that is reconciled away; those remain open questions, and this note does not fill them in. What is fixed is the boundary: until an authorized decision exists, two identities stay two.

The design in one sentence

Let discovery be generous and identity be strict. Atlas may rank, cluster, and suggest likely duplicates as freely as it likes, because that is relevance; it may deduplicate only by RID or by verified equivalence, because that is identity. Keeping the two apart is what allows search to improve without ever corrupting the references that depend on it — and it is why RSM ID can promise that an RID means tomorrow exactly what it meant today.

A note for migrations

Migration is where this boundary is tested hardest, because legacy systems were built on paths, titles, and local keys. Document 08 lets existing identifiers survive as aliases and asks migration to register explicit relationships between legacy IDs, RRNs, and newly issued RIDs. The temptation is to collapse every apparent duplicate on the way in. Resisting it costs some tidiness at first: two RIDs may exist for a while where one referent eventually emerges. That untidiness is recoverable through governed reconciliation; a wrong merge, once references have built on it, is not.