stather (v.)

Research corpus

status
active · mixed
when
2026-07 → now
came from
shellm
led to
Marilyn, Paper ingestion & arXiv weekly

About 4.3 billion tokens of public text, code and de-LaTeX’d papers, assembled, deduplicated and gated on invariants.

Why it started

shellm had a 56 MB corpus. Marilyn needed ten tokens per parameter for a 500M model. That is a data-engineering project, not a download.

What it is

Fourteen tiers from public datasets (educational web, math, code, papers), plus a tier of about 22,000 arXiv e-prints converted from LaTeX source to prose, plus a smaller OCR tier for PDF-only papers that is kept apart because OCR misreads math. Composition runs as a Temporal pipeline: near-duplicate detection between overlapping sources, licence filter, quality gate. The backup of the one irreplaceable part, the derived arXiv prose, was proven by restore.1

Where it stands

Active; v5 corpus finalised 2026-08-13 and pinned. The provenance gap is open: no tier records the dataset revision it was fetched from, so tier integrity cannot be re-verified today. The fix is scoped at about an hour and not yet done.

The hardest week of the project produced the site’s most-cited lesson: gate on invariants, not symptoms.

Lineage

Sub-pages

Sources

Numbers on this page are quoted from the lab's own records. The records are private; each note gives the record's date.

  1. lab record