GlucoFM: Foundation model for continuous glucose monitoring

Article summary

GlucoFM is a lightweight, self-supervised continuous glucose monitoring (CGM) foundation model. It models slower glucose trends and short-term deviations in separate streams, preserving time-of-day and missingness, and produces transferable representations for metabolic prediction tasks including diabetes risk, insulin resistance, beta-cell dysfunction, hyperlipidemia, hypoglycemia, obesity, and glucotype.

Captured article text

Consumer wearables use motion and physiological sensors to estimate activity and sleep, but these signals provide only an indirect view of glucose regulation. Continuous glucose monitors complement these measurements by tracking interstitial glucose every few minutes through a small sensor inserted under the skin, capturing fasting, overnight, and post-meal patterns. Making sense of these traces remains challenging because high-quality clinical labels are sparse and costly to obtain.

Many existing CGM foundation models, including CGMformer, GluFormer, and CGM-JEPA, process glucose through a single representation stream rather than explicitly separating slow baseline and transient event dynamics. CGM contains relatively slow baseline patterns punctuated by short-term deviations that may reflect meals, activity, or sensor artifacts. GlucoFM is designed to use daily CGM data to estimate diabetes risk, insulin resistance, beta-cell dysfunction, and postprandial glycemic response with limited labeled data.

GlucoFM is a self-supervised foundation model with a dual-stream design. It separates slower glycemic trends from short-term deviations while preserving time-of-day and missingness. Latent-prediction objectives learn daily context and temporal evolution. The model was evaluated across four diverse cohorts on seven clinical prediction tasks, comprising 14 cohort-task evaluations. The article reports that GlucoFM’s PR-AUC was 5.8 percentage points higher on average than the best-performing GluFormer variant evaluated when both were pre-trained on the same corpus. In the detailed linear-probe comparison, average PR-AUC increased from 54.7 for the strongest CGM-specific baseline retrained on the same data to 58.8 for GlucoFM, an absolute gain of 4.1 points.

Training GlucoFM to understand metabolism

GlucoFM was pre-trained on 109,066 hours of unlabeled CGM data from Wear-CGM and four published datasets, totaling 477 participant/session records.

CGM recordings can contain gaps, different sampling intervals, and sensor artifacts. GlucoFM aligns each recording to a 24-hour, five-minute grid and retains an observation mask, keeping measured and unobserved positions distinct. Its dual-stream encoder separates a lower-frequency state component, representing slower glycemic trends, from a residual event component capturing short-term deviations that may arise from physiology, behavior, or sensing artifacts.

Rather than reconstructing exact raw glucose readings, which can be affected by measurement noise and sensor artifacts, GlucoFM uses latent predictive pre-training with two complementary tasks:

  • Contextual prediction: parts of a daily glucose sequence are masked and the model predicts their latent representations from surrounding context. Predicting in latent space aims to capture broader daily glucose patterns without reconstructing every sensor reading.
  • Temporal dynamics: the model predicts how a person’s steady baseline and short-term deviations will shift from one hour to the next, encouraging it to capture the continuous nature of glucose dynamics rather than treating readings as isolated snapshots.

CGM-aware augmentations introduce baseline drift, compression-like drops, sparser sampling, and short disconnections, exposing the model to variation and missingness encountered in real CGM recordings.

What GlucoFM can do

Google Research evaluated GlucoFM across four cohorts — CGMacros, Stanford, Hall, and ShanghaiT2DM — and seven clinical prediction tasks, alongside a separate assessment of two-hour postprandial glycemic response prediction. The study tested whether frozen representations were informative for individual 24-hour windows from unseen participants, whether they provided historical context for postprandial prediction, whether multiple days improved subject-level prediction, how representations transferred to new cohorts, and how effectively they adapted when labeled data were limited.

Accuracy across metabolic tasks

The study used subject-disjoint window-level linear probing. Each model encoder was frozen, a linear classifier was trained on individual 24-hour representations, and no participant appeared in both training and test folds.

Across 14 cohort-task evaluations, GlucoFM achieved the strongest task-averaged PR-AUC among the evaluated methods. It achieved the highest PR-AUC in all diabetes-risk and beta-cell-dysfunction evaluations and in three of four insulin-resistance evaluations.

Predicting postprandial glycemic responses

For a dynamic prediction task, information available before each logged meal was used to predict the complete two-hour glucose-change trajectory relative to the meal-start value. The evaluation covered 874 paired meal events from 34 participants, with subject-disjoint cross-validation and separate modeling of Dexcom and Libre under identical splits.

The study progressively combined each frozen model representation with one hour of pre-meal CGM, meal nutrition — energy, carbohydrate, fat, protein, and dietary fiber — and participant-level information such as fasting glucose, BMI, and diabetes status. With the full context, GlucoFM achieved the lowest mean absolute error among the evaluated models: 21.88 mg/dL, compared with 22.90 mg/dL for the best baseline and 27.69 mg/dL for the train-fold mean baseline. These results suggest that GlucoFM provides complementary historical context for predicting postprandial glucose changes.

Looking beyond a single day

A single 24-hour trace may not fully capture a person’s glucose patterns. GlucoFM encoded each day separately and averaged representations across up to seven days, with each participant contributing equally.

Additional days improved PR-AUC in most settings across most datasets, including gains of 9.6 points for Stanford beta-cell dysfunction and 14.0 points for Hall diabetes prediction. CGMacros also showed mostly positive gains across Dexcom, Libre, and fused sensor data. ShanghaiT2DM insulin resistance was the main exception under simple averaging, indicating that the best aggregation strategy can vary by task. Overall, frozen daily representations could be combined to strengthen subject-level prediction without retraining the encoder.

Crossing the cohort divide

The cross-dataset transfer evaluation asked whether a diabetes-risk classifier trained on one clinical cohort would work on patients from a different study. For diabetes risk and insulin resistance, GlucoFM led in 11 of 12 evaluations by 0.5–8.6 PR-AUC points and trailed once by 0.6 points. Its absolute PR-AUC ranged from 61.6% for Stanford-to-Hall tasks to 90.0% for Hall-to-CGMacros insulin-resistance prediction.

Learning with less data

GlucoFM was evaluated under two few-shot settings: varying the number of labeled participants per class and varying the fraction of observations available from each participant. The article reports that GlucoFM markers were highest at every evaluated data budget, including the most limited settings of one participant per class and 1% of observations. The advantage was especially clear when labeled subjects were scarce.

Modeling glucose dynamics at two timescales

The full dual-stream design was compared with simpler alternatives: raw glucose processed directly, a design emphasizing slower trends, and a design emphasizing faster short-term deviations. The event-only version was the weakest, while raw-input and state-only versions were competitive. The full dual-stream model consistently performed best, supporting the organization of slower and faster glucose dynamics as complementary streams before combination.

Conclusion

The results suggest that CGM models can benefit from explicitly accounting for the multiscale structure of glucose dynamics, including slower trends, short-term deviations, daily timing, and sensor missingness. By learning reusable patterns from unlabeled CGM, GlucoFM produced representations that performed strongly across prediction, transfer, and few-shot settings, offering a way to make better use of limited labeled clinical data.

The current pre-training population remains modest. The authors’ next steps are larger and more diverse populations, native multi-day modeling rather than independently processed 24-hour windows, and evaluation of real-time changes.

Source attribution and limits

This capture preserves the Google Research Blog article’s claims, methods, reported metrics, cohorts, and stated limitations. The results are from the described research datasets and evaluation protocols; they are not independently reproduced benchmarks, a clinical diagnostic approval, or evidence of clinical equivalence. The article notes that CGM data include missingness and sensor artifacts, and that broader populations and longer time horizons remain future work.

Labels

  • Health & Bioscience
  • Machine Intelligence