How mobility gives language models a deeper understanding of place

Article summary

We introduce a dynamic, mobility-informed framework that allows AI models to understand the temporal activity rhythms of places over time, and, in doing so, significantly improve predictions about real-world attributes like opening hours, price levels, and busyness.

Captured article text

Artificial intelligence has made incredible progress in understanding the world through text. However, to build AI models that truly understand the physical world, they must comprehend more than just words: they need to capture the dynamic, real-world functionality of the built environment. Every place has two distinct signatures: its identity on paper, and its actual functional rhythm.

Traditional language models typically build representations of places (commonly referred to as “points of interest” or POIs), whether it’s a business or a place like a park or landmark, by relying heavily on this static metadata. They successfully analyze addresses, business categories, and text descriptions. While world-class language models like Gemini are incredibly proficient at processing text data, their geospatial representations can be significantly enriched by incorporating the real-world functional dynamics of the urban environment. Complementing semantic labels with mobility data can enable these models to effectively capture the unique temporal activity rhythms of POIs in a city.

To demonstrate this complementary capability, we introduce Mobility-Embedded POIs (ME-POIs), a novel framework that improves text-based place representations derived by language models. Using publicly available benchmark datasets, ME-POIs incorporates aggregated and anonymized mobility patterns, such as arrival times, stay durations, and surrounding movement patterns. Rather than treating a place as a frozen set of words, ME-POIs use a self-supervised approach to blend text descriptions with large-scale, anonymized mobility patterns from public benchmarks, capturing the aggregate spatial activity footprints of the environment throughout the day. The model constructs a numerical vector representation that encodes both the identity of a place and its dynamic functionality. Integrating ME-POIs with advanced text models delivered a context advantage that yielded up to an 81.9% relative gain in predicting visit intent, a 75.1% improvement in price level classification, and a 24.7% increase in busyness estimation accuracy across unseen places.

How the ME-POIs framework works

By providing a pre-enriched representation of a place, the ME-POIs framework makes it easier for AI models to draw inferences about attributes such as operating hours, target price levels, and current business status without calculating those attributes from scratch each time.

To build a model with a deeper understanding of places, the framework distinguishes between a place’s identity (its name and category) and its function (its aggregated visit footprint). In prior geospatial AI research, mobility patterns were almost exclusively applied to predicting the next POI a user will visit. ME-POIs shifts mobility from an output prediction task to an input feature that defines the place itself.

The framework uses a three-step pipeline: visit alignment, spatial multiscale visit propagation, and text-mobility synergy.

Visit alignment

The model treats aggregate visits to a specific POI as data points. It analyzes temporal arrival windows, departure trends, and typical stay durations. Rather than calculating simple averages, a temporal encoder maps these sequences into a dense vector space. This establishes a “functional centroid”: a multidimensional signature that maps the aggregate anonymized mobility patterns associated with a place over a one-year cycle and across different days of the week.

Solving the long tail of data sparsity

Famous landmarks, large shopping malls, and popular downtown chains generate abundant visit data, while most local businesses suffer from severe data sparsity. When a model encounters a place with few or no recorded visits, it may incorrectly assume the place has zero activity, leading to broken predictions.

ME-POIs addresses this through spatial multiscale visit propagation. It looks at adjacent places across multiple spatial scales — the immediate street, block, and wider neighborhood — and statistically transfers aggregated visit patterns from busy, data-rich neighbors to nearby sparse places. By learning a regional rhythm from active areas, the framework gives smaller shops a geographical prior even when they have few or no appearances in the data.

Text-mobility synergy

The framework enriches textual descriptions by aligning high-level language embeddings with the generated mobility vectors by maximizing cosine similarity. The mobility signal is layered directly on top of the language representation, creating a more holistic view of a place or business.

This hybrid approach preserves structural semantics, such as knowing that a place sells food from its text description, while absorbing operational context, such as whether it functions as a lunch spot or a late-night diner from its mobility signature.

Experiments

The researchers evaluated ME-POIs across two large, culturally distinct metropolitan areas: Los Angeles and Houston. The framework was trained on observed places and tested on entirely unseen places to assess whether it generalized beyond memorizing local patterns.

The five downstream tasks were:

  • Opening/closing hours prediction: infer the exact schedule of a business.
  • Price-level classification: distinguish high-end businesses from lower-cost businesses using mobility context.
  • Permanent closure detection: flag businesses that have gone dark before online profiles or reports are updated.
  • Visit intent classification: estimate aggregate search and navigation interest as a proxy for popularity.
  • Busyness forecasting: predict future crowd densities and peak-hour dynamics.

The comparison included standard text-only embedding models such as Gemini embeddings, trajectory-based geospatial models such as TrajGPT, and hybrid variations to isolate the value of mobility patterns.

Results

Across unseen test places, integrating ME-POIs improved performance over both purely text-based and mobility-based baselines across all predictive tasks. In some cases, including price-level classification, a model trained exclusively on mobility data surpassed text-only language models. The result suggests that collective activity at a physical place can be more descriptive of its function than the formal words used to label it.

Conclusion and limitations

ME-POIs demonstrates a way to move beyond static digital labels and enable AI to model the physical world. The framework focuses on aggregate properties of places and cannot draw conclusions about individual users or provide individual personalization. Its value is in creating rich aggregate numerical signatures of places that reduce the computational burden on downstream systems when they infer place attributes.

The authors position ME-POIs as part of Google Earth AI’s broader effort to create geospatial models and datasets that turn planetary data into actionable intelligence. Acknowledged co-authors include Neha Arora of Google Research, Prof. Cyrus Shahabi, and Shang Ling Hsu of the University of Southern California.

Labels

  • Algorithms & Theory
  • Earth AI
  • Machine Intelligence

Source attribution

This capture preserves the Google Research Blog article’s claims, method description, reported metrics, and stated limitations. The reported gains and benchmark comparisons remain Google Research attribution; they are not treated as independently reproduced benchmarks.