Test-time Memorization
Test-time memorization 是模型在推論期間把新資訊寫入可用記憶的能力,不等同於單純拉長 context window。Google Research 的 Titans + MIRAS 文章把它定義成:模型運行時根據 surprise metric 判斷哪些輸入值得進入長期記憶,並在不重新離線訓練的情況下更新記憶模組。
Titans + MIRAS 的分工
Titans 是具體架構:用 attention 處理精準短期脈絡,用神經式 long-term memory 壓縮並保存大量過去資訊,再用 persistent memory 保留固定知識。MIRAS 是理論框架:描述 memory size、surprise metric、decay rule 與 memory update algorithm 如何決定模型是否該記住新事件。
這和 reasoningbank-agent-memory 的層級不同:ReasoningBank 是 agent 執行後把 trajectory 萃取成策略記憶;test-time memorization 則是模型內部在生成過程中即時更新記憶狀態。它也補強 reasoning-for-factual-recall:reasoning trace 能提升 factual recall,但若模型架構本身能把高 surprise 資訊寫入長期記憶,長上下文任務就不必只依賴更大的 attention window。nested-learning-continual-learning 則把 Titans / MIRAS 往上抽象成多時間尺度的 nested optimization:不只問「要不要記住」,也問 architecture、optimizer、memory module 各自多久更新。
為什麼重要
長 context 不是免費午餐:把上下文塞得更長會增加成本,也不保證模型知道哪些細節值得保留。Titans / MIRAS 的重點是用 surprise 作為寫入門檻,讓模型在極長文件、串流資料或邊緣推論中保留真正改變預期的資訊。
Ponytail 判準:產品層不要先自建複雜 long-term memory system;先判斷問題是 retrieval、agent memory,還是模型層長上下文瓶頸。只有當任務需要跨超長序列保留新 facts,才需要關注 test-time memorization 類架構。