User Simulator Evaluation
User simulator evaluation 是評估「用 LLM 扮演使用者」是否真的像真人的測試方法。Google Research 的 ConvApparel 文章指出,若 conversational agent 只拿過度有耐心、知識過完整、語氣太規整的 synthetic users 做訓練,產品可能在真實使用者面前失效。
ConvApparel 的方法
ConvApparel 用 apparel shopping domain 建立 4,000+ human-AI multi-turn conversations,並把真人請求隨機導到兩種 recommender:一個 helpful 的 Good agent,以及刻意誤解、檢索較差、容易讓人挫折的 Bad agent。這個 dual-agent protocol 讓資料同時包含滿意與挫折情境,並要求使用者逐回合標註 satisfaction、frustration、purchase likelihood 等內部狀態。
它的評估框架有三層:
- Population-level statistical alignment:比對真人與模擬對話在長度、每回合字數、dialog acts 等分布是否接近。
- Human-likeness score:訓練 discriminator 判斷對話像真人還是 synthetic conversation。
- Counterfactual validation:只用 Good agent 對話訓練 simulator,再測它面對 Bad agent 時是否會像真人一樣變得挫折、拒絕或降低滿意度。
對 eval 的意義
這個概念補強 eval-is-spec:AI 產品的 eval 不只要有 final answer correctness,也要測試互動對象、情境分布與 out-of-distribution 行為。它也連到 agent-trace-observability,因為 user simulator 的可信度需要從 turn-by-turn trace 與使用者內部狀態回推,而不是只看整段對話是否「看起來合理」。
Ponytail 判準:不要先打造完整 synthetic user platform;先用一組真實失敗對話、少量情境標註與 counterfactual bad-agent 測試,確認 simulator 不會只是在模仿表面語氣。
DialogLab:多方人機對話模擬
Google Research 的 DialogLab 把 user simulator 從「單一使用者是否像真人」延伸到多方人機互動:會議、課堂、社交場合與訓練情境都有多人、角色、子群組、插話與 backchanneling。它把 group dynamics(group、parties、elements)和 conversation flow dynamics(snippets、turn order、interaction style)分開建模,讓設計者用 author-test-verify workflow 快速迭代。
這補強 eval-is-spec 的邊界:對話 agent 的規格不只是一組 final answers,也包含 turn-taking 分布、情緒流、誰能插話、何時需要人類控制等互動條件。DialogLab 的 14 人評估顯示,human control mode 比 fully autonomous 或 reactive mode 更讓參與者覺得自然、可控與有沉浸感;這提醒 agent-trace-observability 要保留可診斷的多方互動 trace,而不是只存單線 transcript。
風險
- Prompt-only simulator 容易保持不自然的禮貌與耐心。
- SFT / ICL simulator 在統計分布上更接近真人,但仍可能被 discriminator 看出 synthetic artifacts。
- 用不真實 simulator 最佳化 agent,可能提高測試分數但傷害真實使用者體驗。