Memento 3 preprint: frozen LLM rewrites its rulebook on the fly
An October 11, 2026 preprint, Memento 3, pairs a frozen language model with a human-readable rulebook it edits in an observe-hypothesize-test-validate loop, with no gradient updates. On the PAST-Bench benchmark, task success rose from 62% to 89% over 30 days while compute costs fell about 70% versus weekly fine-tuning.
- PAST-Bench task success rose from 62% to 89% in 30 days
- Safety violations dropped from 12 to 0 versus the static baseline
- Compute spend about 70% lower than weekly fine-tuning
- Companion reports: SAHOO with a Goal-Drift Index and PAST-Bench
Read next
AI