愛好 AI 工程 Blog (aihao.tw)
基本資訊
- 名稱: 愛好 AI 工程 Blog
- 網址: https://blog.aihao.tw
- 類型: 繁體中文科技部落格
- 地區: 台灣(繁體中文)
- 語言: 中文(zh-TW)
- 描述: 以 AI 工程實務為主題的個人部落格,專注於 AI Agent、Coding Agent、規模化評測等深度技術分析
內容定位
- AI Agent 產品與技術深度分析
- Coding Agent (Vibe Coding) 實戰與評測方法論
- 產業技術會議重點整理(如 Anthropic Code with Claude 等)
- 以「小編」視角進行技術報導與評論
相關領域
- AI Agent 規模化評測 (Eval at Scale)
- Vibe Coding 與 AI 輔助程式開發
- AI 產品工程與 A/B Testing
監測筆記
- 文章發布頻率:不定期(深度文為主)
- 付費牆:無
- 特色:長文深度整理,注重技術脈絡與實戰架構
- 2026-07-06 新文:老師傅知識萃取成 Agent Skill,補強 skill source / extraction route 的方法論
- 2026-07-06 新文:SemiAnalysis 記憶體短缺與 context window 上限,補強 context engineering 與 HBM 供給限制
- 2026-07-18 監看:兩篇 2026-07-08 草稿頁被選入 raw capture;一篇補上 AI cognitive/belief offloading 與 anti-sycophancy,另一篇補上 code generation 成本與 production software 驗證成本的拆分。兩篇均保留 Draft 狀態,研究與產業數字需回查原始引用。
- 2026-07-25 新文:以四篇「Rethinking Agent Harness」導讀整理 function call 的 Decision/Serialization/Guarantee 分層、task-level Skill、filesystem 檢索的完整性訊號,以及 LLM × Harness × Data × Task 的適用邊界;更新既有 harness、agent skill、agentic search 與 LLM Wiki 概念頁。
- 2026-07-25 新文:整理 Matt Pocock 的 Agent Skill 設計哲學;以 user/model invocation、branch-based progressive disclosure、leading words、completion criteria 與 no-op pruning 更新既有 Agent Skills、harness 與 instruction-file concepts。
- 2026-07-25 草稿:整理 Simon Willison《Agentic Engineering Patterns》;以「程式碼變便宜、好程式碼沒有」、red/green TDD、手動測試、認知債、subagent context 邊界與 PR evidence 更新既有 coding-agent workflow cluster。
- 2026-07-26 新文:導讀 LlamaIndex《Beyond RAG》工作坊,補強文件解析、spatial text、結構化 chunking、agentic retrieval 與 parser/RAG eval 的上游 failure taxonomy。
- 2026-08-05 新文:整理「很難 eval」是產品設計警訊,補強可驗證 UX、provenance、progressive disclosure、atomic review 與使用者信任校準;更新 eval-is-spec 與 agent-experience,Hamel Husain 原文與案例屬來源歸屬,不視為獨立產品 benchmark。
- 2026-08-05 新文:整理 Hex 資料 agent engineering,補強資料分析的驗證缺口、capability modules、tool search、semantic model、Context Studio 回饋迴路與長時程 Metric City eval;更新 data-science-agents、eval-is-spec、agent-experience、agent-ready-data-governance、context-engineering、agent-skills、loop-engineering 與 agent-trace-observability,受訪者/媒體數字保留來源歸屬,不視為獨立 benchmark。
- 2026-08-05 草稿:比較 Braintrust 的架構分層 eval 與 howtoeval 的 production error analysis;更新 eval-is-spec、agent-trace-observability,並建立 agent-eval-methodologies,保留平台方法與數字的來源歸屬。
- 2026-08-22 canonical revision:同一主題由 draft 發布為正式 canonical page;補上 2025 論戰脈絡、floor raising、code-aware eval、offline/online 轉折、六代 agent 架構與 harness 塌縮/變厚的分歧,更新 agent-eval-methodologies,保留 draft raw 作為 source family。
- 2026-08-22 新文:以 Excel 週報比較 AI-enabled、AI First 與 AI Native,補強從 task augmentation 到 workflow/operating model redesign、zero-basing 與企業治理的判斷框架;更新 ai-native-enterprise-governance 與 ai-fitness-and-enterprise-ai-maturity。
- 2026-08-24 新文:整理 Shreya Shankar 的 Evals 自動化演講,補強 Analyze/Measure/Improve、error-discovery、criteria drift 與「人定義、AI 規模化」邊界;更新 eval-is-spec,演講中的 skill 與 benchmark 數字保留來源歸屬。
- 2026-08-24 擷取新文:整理《The New Frontier of AI Search》,補強 agentic-search-tool-curation 的 hybrid search 基本盤、agentic search 補救迴圈、行為訊號、presentation bias、LLM judge calibration 與 retrieval 成本取捨;講者成熟度與數字保留來源歸屬。
- 2026-08-24 canonical revision:同一 AI Search URL 更新標題並補上「現在就做/值得投入/先知道」成熟度表,明確把 sparse/dense 的漸進部署、signals boosting、quantization 與場景化採用順序接回 agentic-search-tool-curation;保留舊 raw 作為 source family。
相關連結
- 網站: https://blog.aihao.tw
- 標籤: AIAgent、Eval、Coding
關聯頁面
-
replit-agent-eval-scale — 該站報導的核心概念:Replit 規模化評測體系
-
test-time-compute-evaluation — 推論時算力對模型評測、成本與安全評估的影響
-
dynamic-agent-workflows — Code Act、tool use 與 Claude Code Dynamic Workflows 的 coding agent 工作流概念
-
agent-trace-observability — 用 trace 分析 agent 決策、工具呼叫與失敗路徑的觀測方法
-
agentic-ai-cost-management — GitHub Copilot scale 文章延伸出的 prompt caching、advisor 與多模型調度成本控制概念
-
coding-agent-as-optimizer — coding agent 作為外層優化器的迭代評測模式
-
agentic-search-tool-curation — 以高訊號字串、混合訊號與語意檢索為主的 agent 檢索工具策展概念
-
document-parsing-first-rag — 先以文件結構、parser 與 metadata 判斷 RAG failure,再選檢索與 agentic loop
-
agent-event-streaming-format — Token 串流到 Agent 事件串流的語意事件、namespace 與 projection 設計
-
eval-is-spec — eval 資料集作為 AI 產品規格書、把品質定義成可測量分布的概念
-
rpi-crispy-workflow — RPI 修正版:由研究/計畫/實作轉為 Questions → Research → Design → Structure Outline → Plan → Worktree → Implement → Pull Request
-
harness-engineering-for-ai-coding — 駕馭工程系列把 Deep Agent 六項能力、guides / sensors / tool feedback / mid-run injection / Goal 驗收轉成 AI coding agent 的可執行回饋迴路
-
loop-engineering — 駕馭工程外層 loop 文章補上 Ralph、Symphony、Cron 與跨 context 狀態管理的邊界
-
self-improving-harness — 駕馭工程第 7 篇補上 production trace、eval、regression gate 與 rollback 驅動的 harness 自我改進
-
model-harness-fit — 駕馭工程第 8 篇補上模型與 harness 配對、過期與評估問題
-
agent-framework-selection — 駕馭工程第 9 篇整理自建 agent framework 的 Deep Agent / 基礎構建選型
-
cognitive-load-agent-orchestration — 用認知負荷理論分析多 coding agent 協調、spec 與 memory 設計
-
microsoft-mai-model-platform — Microsoft MAI / Phi 的模型線分工、Foundry / Copilot / Office 分發與模型選型策略
-
ai-native-engineering-organization — Anthropic AI-native engineering org 演講與反對聲音整理
-
agent-trace-observability — trace 作為 agent 行為 source of truth,連接除錯、eval、production analytics 與 harness 改進
-
eval-is-spec — 通用 AI 指標不能替代 application-specific eval,產品級評測要從真實失敗案例與 domain criteria 建立
-
ai-collaboration-compounding-workflow — Eugene Yan 五層 AI 協作複利工作法,連接 context engineering、workflow 與團隊基礎設施
-
agent-instruction-files — AGENTS.md / CLAUDE.md 該寫什麼、不該寫什麼的最小化判準
-
agent-skills — Agent Skills 從建立到評估的工作流封裝方法
-
agent-experience — Agent Experience(AX)把產品、CLI、API 與文件設計成 agent 可穩定操作的介面
-
cloud-agent-vs-localhost-agent — cloud delegation 與 localhost pairing 的 coding agent 部署邊界
-
agent-sandbox-architecture — agent sandbox 的整合模式、隔離方式與 execution safety 取捨
-
ai-cognitive-offloading-and-agency — AI cognitive/belief offloading 與 agency gate
-
ai-code-validation-bottleneck — AI 時代 code review 成為 PR / 驗證瓶頸,需用規格、AI reviewer、CI feedback 與工作證據重設審查流程
-
code-abundance-product-taste — AI 時代 PM 角色被 prototype、eval、模型時機與工程產能暴增重新定義
-
agentic-search-tool-curation — RAG 不應只靠 vector search;先做 query understanding 與資料 affordance 建模,再把 embedding 放在 fallback / rerank 位置
-
skills-vs-mcp — 後 MCP 時代的 Skills、CLI / script 與 MCP server 工具設計邊界
-
unix-style-agent-architecture — Marc Andreessen / Latent Space 訪談整理出的 LLM + shell + filesystem + Markdown + cron 最小 agent 架構
-
model-specific-system-prompts — Fable 5 prompting 文章補上模型換代時刪除過時 prompt / skill scaffold 的判準
-
model-harness-fit — Fable 5 prompting 文章補強 effort、refusal/fallback、長時程 agent 驗證與 memory notes 的模型貼合問題
-
agent-skills — knowledge-to-skill framework 補強專家知識萃取成 skill 的來源與容器設計
-
agentic-ai-cost-management — context window / HBM 短缺補強成本管理與 context engineering 的上限判準