Time-series Foundation Models

Time-series foundation models 是把跨領域時間序列資料預訓練成可在新資料集上 zero-shot 或低調整成本預測的模型。與每個資料集各自訓練的 forecasting pipeline 相比,這個方向把模型、資料與 inference contract 預先泛化,再依任務提供 target、歷史 covariates 與已知的未來訊號。TimesFM-3 是 Google Research 在 2026-08-31 公布的 multivariate 案例;其模型規模、訓練資料量與 benchmark 結果均是來源研究的報告,不是本頁的獨立重現。1

TimesFM-3 的輸入與輸出契約

  • Multiple targets:同時預測彼此共同演化的多條 target series,而不是把每條序列完全隔離。
  • Past covariates:加入只在歷史可觀測的輔助特徵,例如過去流量或活動量。
  • Past-future dynamic covariates:加入預測 horizon 內已知的事件,例如促銷排程或天氣預報。
  • Point + quantile forecasts:除了單一點預測,也輸出多個 quantiles,讓工作流保留不確定性資訊。

TimesFM-3 的實作資料顯示,這些介面被放在同一個原生 multivariate model 中;官方文章稱模型有 330M parameters,並以超過 1 trillion time points 的 real-world/synthetic corpus 預訓練。這些數字應保留 Google Research attribution,不應直接解讀成對任意企業資料的泛化保證。1

架構與推論

來源描述 TimesFM-3 延續 decoder-only transformer,並以 per-series normalization 處理不同量綱;target、past covariate 與 past-future covariate 會以不同方式構造 multivariate tokens。對未來已知的 covariates,token 可採 lookahead 方式攜帶後續訊號。

模型使用 Contiguous Patch Masking(CPM)把未知的 forecast horizon 放入 masked placeholder,讓多個未來 patch 在 single forward pass 中一起產生,而不是逐步 autoregressive loop。Google Research 文章稱每個 target、每個 horizon step 會產生第 10 至第 90 percentile 的 9 個 quantiles。這使 forecast workflow 可以把預測值與不確定性一起交給後續的決策或人工審查;是否因此降低特定任務延遲或成本,仍需實測。1

Hugging Face model card 另列出 20 層 transformer、model dimension 1280、16 heads、context patch length 32 與 forecast horizon patch length 64,並將架構標為 Stacked Mixing Transformer with Variate Attention and CPM Iterative RevIN。這些是目前 model card 的 release metadata,不等於完整研究論文或獨立實作審查。1

評測不能只看一個平均分數

Google Research 報告 TimesFM-3 在 Gift-Eval、FEV-Bench 與 TIME 三個公開 benchmark 的 point 與 probabilistic forecasting average rank 都居預訓練 foundation models 之首;官方 GitHub README 進一步列出 FEV-Bench 100 個 real-world tasks、TIME 50 個 domain datasets/98 個 evaluation tasks,以及各 benchmark 的 rank #1 說法。這些是官方來源所報告的結果,沒有在本次 ingest 中重跑,不能升格為跨資料集、跨時間外推或 production SOTA。1

評估 time-series foundation model 時,至少要固定:

  1. target 與 covariate 在 forecast origin 的實際可得時間;
  2. forecast horizon、資料頻率、缺值與多序列的對齊方式;
  3. point metrics 與 probabilistic metrics 是否同時符合任務成本;
  4. random split、rolling split、unseen time period 或 unseen entity 的泛化邊界;
  5. inference latency、重試/重算成本、quantile calibration 與人工 review;
  6. 模型權重、商業使用、production deployment 與資料外送的授權條件。

這個檢核表可接到 eval-is-spec 的 task-level eval,也可接到 data-science-agents 的「資料檢查 → 可執行分析 → verifier → 報告」流程;benchmark rank 只是初篩訊號,不是任務完成或商業價值的替代品。

釋出與部署邊界

TimesFM-3 已由官方 GitHub repository 與 Hugging Face model card 提供;GitHub README 同時說明 source code 採 Apache-2.0,但 3.0 預訓練權重目前採 timesfm-non-commercial-license-v1.0,限制 non-commercial、non-production use。Google Research 文章所稱的 BigQuery integration 是「接下來幾週」的計畫,本次沒有把它寫成已上線能力。實際導入前應重新讀取版本與 license 文件,而不是只依賴模型名稱。1

這個授權與 runtime 邊界也說明了 model-harness fit 的另一面:模型的輸入格式、quantile output、硬體/推論環境、資料治理與商業授權必須一起評估。可參考 model-harness-fit、agent-ready-data-governance 與 open-weight-model-strategy。

與既有知識的關係

  • tabular-foundation-models:TabFM 將 zero-shot / in-context learning 用於表格分類與回歸;TimesFM-3 則將相近的 foundation-model pattern 用於有順序與共變量的時間序列。
  • wearable-health-foundation-models:SensorFM 與 GlucoFM 處理 wearable/CGM 的生理時間序列,但其 health cohort、label、device 與臨床風險邊界不同,不能直接合併 benchmark。
  • planetary-prediction-engine:PPE 將資料選擇、資料集整理與預測編排成 geospatial workflow;TimesFM-3 是其中可被選用的預測模型類型,而不是完整的資料治理或 agent pipeline。
  • data-science-agents:預測模型應被放進可回查的資料與 verifier loop,而不是把一次 forecast 當作結論。
  • eval-is-spec:模型選型要由任務資料分布、錯誤成本、時間外推與不確定性需求定義 eval。

未解事項

  • TimesFM-3 三個 benchmark 的完整資料切分、版本、baseline 與 calibration 細節,仍需回到 benchmark/repo 的可重現設定核對。
  • 本次未找到 TimesFM-3 專屬研究論文的完整內容;GitHub README 連結的是 TimesFM 原始 ICML 2024 paper,不能自動視為 3.0 的完整方法論文件。
  • BigQuery integration 是否已在特定 region 或產品介面上線,需要以當下 Google Cloud 官方文件確認。
  • 官方 benchmark 排名如何轉化到企業資料品質、長期漂移、缺值、成本、ROI 與決策安全,尚無本筆來源可支持的結論。

Footnotes

  1. google-research-timesfm-3-multivariate-forecasting-2026-08-31。 ↩ ↩2 ↩3 ↩4 ↩5 ↩6