
Humans are remarkably good at not walking into each other.
We rarely think about it, but every busy footpath, station and pedestrian crossing is a continuous exercise in prediction.
We read speed, direction, hesitation and commitment almost unconsciously. We adjust our own movement in response, often before another person has made their intention explicit.
Recently, we took that observation to the Melbourne Foundational Models Meetup at Melbourne Connect, hosted by SMEC AI.
The talk introduced an early version of SpatioTemporal’s Motion Intelligence thesis: machines are becoming very good at seeing the world, but seeing movement is not the same as understanding what that movement means.
We also demonstrated the idea experimentally.
In our toughest NVIDIA Cosmos crossing scenario, the baseline robot recorded near-collisions in 24% of runs. With SpatioTemporal’s motion-derived intent signal added, that fell to 2%.
The robot had not gained better perception. It had gained another signal between perception and planning: an interpretation of how the humans around it were moving, and what they were likely to do next. That distinction has become central to our work.
Perception tells a machine what is there. Planning tells it what to do. Motion Intelligence helps it understand what is about to happen. Our approach is to learn the grammar of movement, compressing space and time into motion tokens that capture the patterns humans recognise instinctively.
The technology has moved considerably since that first public demonstration. The thesis has become simpler.
If Physical AI is going to move safely and naturally through our homes, workplaces, roads and public spaces, machines will need more than eyesight and a plan.
They will need to learn how to read us.









Thank you to SMEC AI, Andrew Lai and Desmond John for bringing the evening together, and to everyone who stayed around afterwards to keep pulling the idea apart with us.