
SpatioTemporal founder Andrew Ballard recently joined Andrew Lai and Amir Nissen from SMEC AI for a conversation about Motion Intelligence, foundation models and a different approach to building AI for the physical world.
The discussion explored a simple idea: Physical AI does not necessarily need larger models. It needs the right models.
Large general-purpose models have demonstrated what can be achieved by learning from enormous volumes of data and compute. Physical AI introduces a different set of constraints. Robots, autonomous vehicles and other machines operating in the real world need intelligence that can run quickly, efficiently and close to where decisions are being made.
SpatioTemporal approaches that problem by focusing on movement itself.
Rather than learning primarily from pixels, our Motion Intelligence work compresses space and time into motion tokens, creating a compact representation of how people, vehicles and other objects move. The objective is to help machines interpret not only what is around them, but what that movement may indicate about intent and what is likely to happen next.
A pedestrian can be clearly detected by a perception system, for example, while the more difficult question remains unanswered: are they waiting, hesitating or about to cross?
We see that as part of the missing middle layer between perception and planning.
The interview also explored the efficiency that comes from choosing a more focused representation. SpatioTemporal’s models are designed to be small enough to operate on edge hardware, with current models capable of running on devices as modest as a smartphone.
For Australian AI companies, there is a broader opportunity here. Competing in foundation models does not have to mean competing with the world’s largest laboratories on model size, training budgets or compute infrastructure. There remain important physical problems where the advantage may come from finding the right abstraction and building a specialised model around it.
For SpatioTemporal, that abstraction is movement.
Compressing space and time into motion tokens, so machines can begin to understand the grammar of movement.
Robots that read the room. Cars that read the road.
Thanks to Andrew, Amir and the SMEC AI team for the conversation and for featuring our work on ROI from AI.
Watch the original interview and post