Research corpus
- status
- active · mixed
- when
- 2026-07 → now
- came from
- shellm
- led to
- Marilyn, Paper ingestion & arXiv weekly
About 4.3 billion tokens of public text, code and de-LaTeX’d papers, assembled, deduplicated and gated on invariants.
Why it started
shellm had a 56 MB corpus. Marilyn needed ten tokens per parameter for a 500M model. That is a data-engineering project, not a download.
What it is
Fourteen tiers from public datasets (educational web, math, code, papers), plus a tier of about 22,000 arXiv e-prints converted from LaTeX source to prose, plus a smaller OCR tier for PDF-only papers that is kept apart because OCR misreads math. Composition runs as a Temporal pipeline: near-duplicate detection between overlapping sources, licence filter, quality gate. The backup of the one irreplaceable part, the derived arXiv prose, was proven by restore.1
Where it stands
Active; v5 corpus finalised 2026-08-13 and pinned. The provenance gap is open: no tier records the dataset revision it was fetched from, so tier integrity cannot be re-verified today. The fix is scoped at about an hour and not yet done.
The hardest week of the project produced the site’s most-cited lesson: gate on invariants, not symptoms.
Lineage
- Came from: shellm
- Led to: Marilyn (it is her food), paper ingestion (the arXiv tooling grew into a nightly pipeline)
Sub-pages
- Wins · Losses · Rabbit holes
Sources
Numbers on this page are quoted from the lab's own records. The records are private; each note gives the record's date.
- lab record ↩