Agent Trace Observability

Agent Trace Observability 是用 trace 觀察 agent 每一步決策、工具呼叫、錯誤、回饋與修正的工程方法。aihao-blog 的 agent trace 分析文章把 trace 視為持續改進 agent 產品的核心資料,而不是只在出錯時才看的 debug log。

2026-02-17 的 AIHAO 文章進一步把 trace 定義成 AI agent 時代的「source of truth」:source code 只能描述靜態邏輯,真正能解釋 agent 行為的是 messages、tool calls、observations 與 state transitions 連成的 execution trajectory。這讓 trace 同時成為 eval-is-spec 的資料來源、self-improving-harness 的改進燃料,以及 agent-post-training-workflow 中可被蒸餾或標註的經驗資料。

Google Research 的 Chain-of-Evidence(CoE)補上一個研究 agent 的 artifact-level trace 視角:citation graph、solver branch、原始 evaluator output、workspace artifact 與 claim-evidence binding 都應保留,讓論文中的 claim 能回到真正使用的證據。CoE Audit 再用獨立重跑、reference lookup 與 method-code comparison 檢查這條鏈是否完整且正確;這使 chain-of-evidence-autonomous-research 成為 trace → eval 閉環在 scientific research 的具體案例。

核心用途

trace 讓團隊能回答三個問題:

  1. agent 在哪一步偏離使用者目標?
  2. 哪些工具呼叫、上下文或 prompt 造成失敗?
  3. 哪些成功路徑可以轉成 eval、workflow 或 product rule?

這與 replit-agent-eval-scale 的 Telescope / trace clustering 思路相近:只看最終成功率不夠,還要知道成功或失敗是怎麼發生的。

對 loop engineering 的意義

在 aiark-loop-engineering 中,queue record、raw source、wiki diff、verify result 都可以視為 trace 的一部分。這讓每輪 watch / ingest / verify 不只是執行結果,而是後續校準 score、review gate 與來源品質的觀測資料。

最小做法不需要獨立 observability stack:

  • data/watch-queue.jsonl 保留分數、理由與狀態。
  • log.md 保留每輪動作與結果。
  • verify script 輸出 broken links、sha drift、index marker 等檢查結果。

這符合 ponytail-loop-review-gate:先把現有 artifacts 當 trace,用到痛了再抽系統。

Google Research 的 privacy-preserving chatbot analytics 文章補上一個限制:production conversations 和 agent traces 不應預設被原樣讀取或聚類。若 trace 來自真實使用者,應先用 privacy-preserving-chatbot-analytics 這類 aggregate / DP 方法降低單一使用者洩漏風險,再把趨勢回饋成 eval 或 workflow 修正。

實務判準

Agent trace 值得進入 wiki 或 synthesis 的條件:

  • 能解釋 agent 失敗原因,而不只是記錄錯誤訊息。
  • 能轉成 eval case、workflow 修正或 prompt / tool policy。
  • 能回饋到 source scoring、review gate 或 loop 停損線。

Trace → eval 的最小閉環

最小可行做法不是先買 observability 平台,而是讓每筆重要 trace 都能被轉成一個 eval case:

  1. 從 production trace 或失敗回報定位到出錯步驟。
  2. 保留當時的 context、tool result 與期望 outcome。
  3. 把它寫成 component / trajectory / full-turn eval。
  4. 修 harness、prompt 或 tool contract 後重跑。

這個閉環補上 harness-engineering-for-ai-coding 的觀測面:tool feedback、Goal 驗收與 outer loop 都需要 trace 才知道哪裡真的改善,否則只是把 slop 自動化。

Hex 的 Context Studio 提供一個資料 agent 版本:不要求資料團隊讀完所有 production 對話,而是用 LLM 標記可能出錯、agent 搞混或與 semantic model 衝突的案例,再把問題按失敗類型分群,回寫 guide、語意模型與 warehouse context。這把 trace 的用途從單次 debug 擴展成 context repair;個人 memory、組織治理 context 與可疑答案應保留來源層級,避免把錯誤回饋直接升格成真實規則。

從 production trace 選擇 eval 投資

AIHAO 對 Braintrust 與 howtoeval 的比較補上一條實作順序:小流量產品先人工讀 log、追蹤重複問題並把能重現的錯誤寫成高訊號 eval;流量與 harness 複雜度上升後,再把 trace 分層到 component、trajectory、replay/shadow 與線上抽樣。這讓 trace 不只是觀測資料,也成為決定測試預算與 suite 修剪的依據。

相關頁面