Wearable Health Foundation Models
Wearable Health Foundation Model 是把長時間、多模態 wearable sensor stream 預訓練成可跨任務重用的生理表徵,再透過輕量 prediction head 或 agent workflow 適配新健康任務的模式。Google Research 的 SensorFM 是一個具體案例:模型從超過一兆分鐘、五百萬名經同意參與者的未標註感測資料學習,並以 35 個 health prediction tasks 檢驗跨領域轉移能力。這裡記錄的是可重用的架構與評估模式,不把單一研究結果視為臨床驗證。
SensorFM 的方法
- Population-scale self-supervision:使用 PPG、加速度、EDA、皮膚溫度與高度計等五種 sensor modality 的每分鐘 aggregate features,透過 Adaptive and Inherited Masking(AIM)做 missing-aware masked reconstruction。真實感測缺口不被粗暴補值或丟棄,而是被視為資料本身的一部分。
- 共同擴展資料與模型:文章比較約 2 million 到 2 billion sensor-hours、100K 到 100M parameters 的組合;報告指出資料量與模型容量同步增加時,下游分類與回歸表現也持續改善,最大版本在 35 個任務中勝出 33 個。
- Frozen encoder + light adaptation:固定 SensorFM encoder,只訓練輕量 linear head,讓同一個 representation 跨 cardiovascular、metabolic、sleep、mental health、lifestyle 與 demographics 任務使用;文章報告 linear probe 在 35 個任務中的 34 個優於 engineered-feature supervised baseline,並具 label efficiency。
Agentic classroom:把 foundation model 適配變成搜尋問題
SensorFM 不只提供 embedding,也把 prediction-head adaptation 交給一組協作與競爭的 LLM agents。這個 agentic classroom 反覆產生、執行、測試與修正可執行程式,在文章的實驗中探索超過 30,000 個候選解;agent-designed heads 在 20 個 classification tasks 中有 16 個、15 個 regression tasks 中有 12 個勝過簡單 linear probe。
這個模式可接到 data-science-agents:agent 的工作不是憑空生成結論,而是在固定 representation 上搜尋可執行的 adaptation code,並以任務指標回饋下一輪。文章也指出解品質會隨底層 LLM 能力提升,較弱模型可透過 agent collaboration 縮小差距;這是 loop-engineering 中「產生 → 執行 → 驗證 → 迭代」的 domain-specific 例子。
Grounding Personal Health Agent
研究把 SensorFM predictions 當作工具接入 Personal Health Agent,與「daily wearable metrics + ground-truth measurements」及只用 demographics / daily metrics 的 baseline 比較。臨床專家以 context、relevance、justifiability、personalization、potential for harm 五項 rubric 評分 31 個 participant profiles,共 1,860 次評分;文章報告 SensorFM predictions 相對 baseline 在每一項都改善,且與 ground-truth 條件的差異未達統計顯著。
這補強 personal-general-ai-assistant 的一個重要邊界:個人助理的個人化不必全部塞進通用 LLM,而可以由領域 foundation model 先把原始資料轉成可檢查的工具輸出,再讓 agent 以該輸出生成摘要或建議。評估也應接到 eval-is-spec,把專家 rubric、可辯護性與 potential for harm 納入產品級 eval,而不是只看模型的單一 benchmark。
PhotoScan:從 wearable representation 到 optical phenotyping
Google Research 的 PhotoScan 把 SensorFM 的「長期生理訊號表徵」路線延伸到另一個輸入面:用一般 smartphone 的 2D 正面與側面影像估計 body fat percentage、Android-to-Gynoid fat ratio(A/G)與 Visceral-to-Subcutaneous fat area ratio(V/S),再把這些 body-composition features 與 demographics 交給 gradient-boosting classifier 預測 insulin resistance。方法分成三段:以 UK Biobank 的 MRI/DXA ground truth 預訓練 ResNet-50、用 677 人的 PhotoBIA 真實手機照片做 5-fold fine-tuning,再以 132 人的 MetabolicMosaic longitudinal cohort 做獨立驗證。
這個案例補強 eval-is-spec 的 multimodal health eval:不能只比較 body-fat MAE,還要固定資料來源、影像姿勢、外部 cohort、leak-free split、BMI/insulin-resistance 平衡、DXA 對照,以及從 baseline demographics 到 BIA、PhotoScan、DXA 的 feature ablation。來源報告 PhotoScan 在其研究 cohort 上讓 insulin-resistance AUROC 從 0.692 提升到 0.760,接近 DXA 的 0.773;這是 Google Research 的研究結果,不等於臨床等價或可直接部署。
與 SensorFM 的 Personal Health Agent grounding 相比,PhotoScan 展示的是「影像 → 可檢查的身體組成指標 → 風險分類」的工具鏈,而不是讓通用 agent 直接解讀原始影像。這讓 personal-general-ai-assistant 的個人化邊界更清楚:先由領域模型輸出可驗證的中間 representation,再交給 agent 整理;但醫療資料的 consent、privacy、bias、外部族群驗證與人類責任仍是 deployment gate。
Biomarker Discovery Framework:把 wearable research 變成 adversarial loop
Google Research 的 Biomarker Discovery Framework 不是另一個 wearable foundation model,而是疊在 wearable time series 與 clinical lab data 上的 multi-agent research harness。它把候選 biomarker prioritization 組成一個由人類監督的閉環:Orchestrator 將自然語言研究指令拆成計畫,再由 Scout、Literature、Hypotheses、Statistical、ML、Critic、Defender、Mechanism、Novelty、Strategy 與 Report 等專責 agent 協作。
這個案例補強 data-science-agents 與 loop-engineering 的健康研究版本:
- 先固定資料邊界:Scout 先檢查 schema、missingness、時間結構與 endpoint,並把 target label 與 feature construction 分離,降低 leakage。
- 把生成式推理包在 deterministic analysis 外:agent 可提出 hypothesis、文獻機制與 composite feature,但關聯估計、multiple-testing adjustment、model training 與數值驗證由可重現程式執行。
- 把 Critic / Defender 變成 gate:11 項 adversarial battery 檢查 leakage、overfitting、confounding、construct overlap、instability 與生理合理性,最後以 screened、conditional、exploratory、rejected、unstable 等 label 回報,而不是把所有候選都寫成結論。
- 共享 state 讓研究可追溯:shared memory、structured fact sheet、common tools 與 report verification 把假設、數值、圖表、文獻與限制連回同一條 evidence chain。
來源在三個 cohort、共 9,279 participant-observations 上找出 41 個 mental-health 與 25 個 metabolic candidate digital biomarkers;例如以睡眠時間變異與 depression severity 的關聯形成 circadian-instability hypothesis,並以 held-out、subgroup、leakage 與 alternative-explanation 檢查限制解讀。來源同時報告 15 位 domain experts 的 blinded manuscript evaluation;這屬於該研究 protocol 下的 research-system 與報告品質結果,不等於 causal discovery、clinical validation 或醫療部署安全性。這使 eval-is-spec 的要求更具體:健康 agent 的 eval 不只看 prediction,而要檢查 statistical validity、literature grounding、human review、potential harm 與是否把 hypothesis 誤報成 diagnosis。
與 SensorFM 的 foundation representation、agentic classroom 適配,以及 PhotoScan 的「影像 → 身體組成指標 → 風險分類」工具鏈相比,Biomarker Discovery Framework 的可重用模式是 domain data + deterministic computation + adversarial multi-agent review + human sign-off。它顯示 agent 的價值不在於取代統計或專家,而在於把研究假設、執行、反駁、文獻查核與報告組裝成可重複的 hypothesis-to-validation loop。
GlucoFM:把 CGM 的多時間尺度拆成可轉移表徵
Google Research 的 GlucoFM 把 wearable health foundation model 路線聚焦到 continuous glucose monitoring(CGM):先把每筆紀錄對齊到 24 小時、每 5 分鐘的網格,保留 observation mask,再用 dual-stream encoder 分開較慢的 glycemic trend 與較快的 residual event。後者可能來自飲食、活動、生理變化或 sensor artifact;模型因此不把 CGM 當成單一、同質的序列,而是保留 time-of-day、missingness 與兩種時間尺度的互補訊號。
GlucoFM 在 109,066 小時的 unlabeled CGM 資料上預訓練,涵蓋 Wear-CGM 與四個公開資料集、共 477 筆 participant/session records。它不直接重建可能受噪聲影響的原始 glucose readings,而是用 latent contextual prediction 預測被遮罩區段的 representation,再用 temporal-dynamics objective 預測 baseline 與短期偏差從一小時到下一小時的變化;baseline drift、sparser sampling 與短暫斷線等 augmentations 則把真實缺失與 sensor variation 放進訓練。
Transfer、few-shot 與跨裝置評估
來源在 CGMacros、Stanford、Hall、ShanghaiT2DM 四個 cohort 上評估 diabetes risk、insulin resistance、beta-cell dysfunction、hyperlipidemia、hypoglycemia、obesity 與 glucotype 七項任務,共 14 個 cohort–task evaluations。subject-disjoint linear probing 中,GlucoFM 的平均 PR-AUC 為 58.8,相較同一 corpus 重訓的最佳 CGM-specific baseline 54.7,高 4.1 個 absolute points;文章摘要段落另以最佳 GluFormer variant 報告平均優勢 5.8 個 percentage points,兩者是不同的 baseline 比較。GlucoFM 在所有 diabetes-risk 與 beta-cell-dysfunction 評估,以及四個 insulin-resistance 評估中的三個取得最高 PR-AUC。
在 postprandial glycemic response(PPGR)任務,研究使用 34 位參與者的 874 個 paired meal events,分別以 Dexcom 與 Libre、相同 subject-disjoint cross-validation 預測餐後兩小時完整 glucose-change trajectory。加入 pre-meal CGM、餐點營養、fasting glucose、BMI 與 diabetes status 後,GlucoFM 的平均 MAE 為 21.88 mg/dL,低於最佳 baseline 的 22.90 與 train-fold mean 的 27.69。把每天 representation 平均到最多七天,多數設定的 PR-AUC 也提升;跨 cohort transfer 則在 12 個 diabetes-risk/insulin-resistance 評估中贏 11 次。
Few-shot 結果顯示,在每類只有一名標註參與者、或只使用每位參與者 1% observation 的最低資料預算下,GlucoFM 仍在來源比較中維持最高的 task-averaged PR-AUC。這補強 eval-is-spec 的一個實務要求:health foundation model 不能只報一個平均 benchmark,還應固定 participant split、sensor/device、cohort transfer、multi-day aggregation 與 labeled-data budget,才能知道表徵是否真的跨人、跨裝置、跨資料集轉移。
在既有 health agent cluster 的位置
相較 SensorFM 的多模態 wearable representation、PhotoScan 的「影像 → body-composition 指標 → insulin-resistance 分類」,以及 Biomarker Discovery Framework 的「假設 → deterministic analysis → adversarial review」研究 loop,GlucoFM 的新增判準是 CGM multiscale decomposition + missing-aware latent prediction + frozen representation transfer。它可作為 data-science-agents 的 domain representation 來源,也讓 personal-general-ai-assistant 中「先由領域模型產生可檢查中間輸出,再交給通用 agent 整理」的邊界更具體;但本文沒有提出 Personal Health Agent、臨床診斷產品或部署安全證據,不應把 GlucoFM 的 prediction result 直接解讀成醫療建議。
限制與使用邊界
- SensorFM 的數據、效能與 clinician ratings 都來自來源文章的研究報告;本頁不把它們解讀成獨立重現或臨床安全性證明。
- GlucoFM 的數據、效能與跨 cohort 結果同樣是來源研究 protocol 下的 attribution;其 pre-training population 仍 modest,且目前以獨立 24 小時窗口處理,不能視為臨床等價或即時醫療部署驗證。
- 健康資料具有高度敏感性;即使來源描述資料為 de-identified 且參與者同意研究用途,實際部署仍需處理 consent、access control、資料最小化與醫療責任邊界。
- 「模型預測與 ground truth 條件無統計顯著差異」不等於兩者臨床等價;面向高風險決策仍需要專業審核與外部驗證。