Dataset Overview
The OYOTTA Public Source Library captures public HTML source pages published across OYOTTA domains, providing full cleaned text, canonical URL targets, SHA-256 cryptographic provenance for raw HTML and extracted text, and explicit source limitations.
Selection & Provenance Boundaries
- Public Pages Only: Selected from OYOTTA public catalogs and verified public endpoints.
- Exclusions: Paid manuscript editions (
/edition/), cron tasks, administration endpoints, and private previews are strictly excluded. - Redirect Resolution: HTTP redirects and HTML client-side meta-refresh destinations are resolved to canonical target bodies.
- No Speculation: All dates, hashes, and text lengths reflect direct network observations without invented citations.