Catalogue dataset: method and limits
This dataset is a projection of public OYOTTA catalogue records, book metadata, reading collections, and published Stream edition metadata. The hosting worker rebuilds the snapshot when normalized records or relationships change. A timer alone does not create a new edition.
What each field means
- id
- The source object identifier, retained for joins. A second identifier can point to the same URL, so record count and distinct URL count are reported separately.
- title and description
- Published catalogue wording, not an independent review of a work. Descriptions are bounded to 1,800 characters.
- object_class and medium
- Source classifications. Missing classes are retained as OBJECT; a classification is not proof of professional or institutional standing.
- url
- An allowed first-party public destination. Private edition paths, administrative paths, credentials in URLs, fragments, and query aliases are excluded.
- source
- The public file used for this record. The JSON snapshot includes source-file hashes for reproducibility.
Joining the relationship files
Use subject and object as foreign keys to id, and preserve the predicate as the source’s stated relationship. Rows whose endpoints are absent from this snapshot are omitted. Exact duplicate relationships are removed. This provides an internally joinable extract, not a claim that every relationship in the estate is represented.
A reproducible comparison
- Save a dated copy of catalogue.json and relationships.json.
- Join a later snapshot by id, then compare title, URL, description, and class.
- Keep additions separate from changes to existing records. A changed website record does not establish the original creation date of its underlying work.
- For a relationship change, inspect the linked sources before treating the change as a new creative or production fact.
Boundaries
The snapshot does not contain full paid manuscripts, private audience information, account data, inferred nationality, or inferred birthdays. Listing a public URL does not prove it is indexed by a search engine. Search-engine submission outcomes are recorded separately by the hosting worker and are never labeled indexed.
CSV exports prefix formula-like values with an apostrophe for safer spreadsheet use. JSON retains the original displayed strings. Missing records or a source disagreement should be investigated through the source paths rather than filled with invented data.