GOOGLE'S GENIE 2 | OPEN-WORLD GAMES OUT OF NOTHING 🚀
TODAY IT’S ABOUT A REVOLUTION IN THE WORLD OF VIDEO GAMES,
THANKS TO ARTIFICIAL INTELLIGENCE
Imagine you could generate a complete 3D open-world video game from just a single image or sentence. Sounds impossible? That is exactly what Google DeepMind has made possible with Genie 2 – a true revolution in the world of video games!
We’re looking at the FOUNDATION WORLD MODEL from Google’s DeepMind .
Genie 2 is a groundbreaking technology that lets us create interactive 3D worlds out of nothing – and in real time! It is not a classic game engine but a diffusion model that generates images as we change the perspective.
👉 Learn more about Genie 2: https://deepmind.google/discover/blog/genie-2-a-large-scale-foundation-world-model/
Many may still remember the ‘AI Doom’ project.
The “AI Doom” project is an experiment in which an artificial intelligence creates a Doom-like game in real time. Instead of all the graphics and game mechanics being programmed beforehand, a neural network (more precisely: a diffusion model supported by reinforcement learning) generates every single image dynamically. The system, however, first had to be laboriously fed with Doom to create a roughly comparable gameplay experience.
Genie 2 goes much further, though. It is a true ‘World Model’ with a much greater understanding of the world. It can understand and implement concepts such as jumping, running, swimming or ramming a wall at high speed.
The big breakthrough is the ability to generalize. Genie 2 can infer and understand ideas about an environment.
In contrast to Genie 1, which still offered very simple environments with blurry characters and a gameplay duration of only 2 seconds, Genie 2 is a huge step forward.
Environments in Genie 2 can be controlled with keyboard and mouse by humans and AI agents.
HOW DOES GENIE 2 WORK?
Imagine you give Genie 2 a simple image or a description and it creates a complete world from it that keeps evolving. The model also remembers things that lie outside your field of view – important for a consistent and immersive experience.
Genie 2 responds to actions and knows that arrow keys move a character and not the environment.
The system can generate different perspectives, including first-person, isometric and third-person views.
APPLICATIONS AND POSSIBILITIES
Genie 2 enables rapid prototyping of interactive experiences. Researchers can thus train and test AI agents in new environments.
Concept drawings can be turned into interactive environments.
AI agents, such as the SIMA agent, can complete tasks and follow instructions in the worlds generated by Genie 2.
Real images can also serve as a basis, for example to simulate grass in the wind.
EXAMPLES AND CAPABILITIES IN DETAIL
Genie 2 masters:
- Object interactions, such as bursting balloons.
- Character animations.
- NPCs (non-player characters).
- Physical effects such as water, smoke and gravity.
LIMITATIONS
Of course Genie 2 still has its quirks – for example when a ghost suddenly sneaks through your garden or a snowboarder decides on an unexpected parkour adventure. But hey, that’s part of experimenting!
Currently Genie 2 can generate a consistent world for only about a minute. After that come the typical AI video hallucinations.
CONCLUSION
Genie 2 is a milestone and shows the potential of world models. It paves the way for more advanced AI systems.
Even though it is still a research project, there are already similar projects, e.g. from Tencent or World Labs.
Combining Genie 2 with an AI agent like SIMA and an LLM like Gemini opens up incredible possibilities.
Real-time interaction and autonomous agents:
By integrating with an AI agent like SIMA, agents can independently perform tasks within the worlds generated by Genie 2 – such as navigating, interacting with objects or executing commands like “Open the blue door”. This leads to a new kind of agent-based simulation in which AI doesn’t just act in static, pre-programmed scenarios but in dynamically changing environments.
Infinitely varied training environments:
From a single image prompt, Genie 2 can generate an almost unlimited variety of playable 3D worlds. These environments can serve as simulated training grounds for AI agents, allowing them to be trained in ever new, challenging scenarios – a crucial step in improving their robustness and generalization ability.
Natural language control and complex planning:
A powerful LLM like Gemini makes it possible to control these interactive worlds with natural language. Users can formulate complex commands or requests that Gemini then interprets – from simple instructions to multi-step, planning-based tasks. This opens up the possibility of developing interactive experiences or games that adapt flexibly to users’ inputs and preferences.
Revolution in game design and prototyping:
The entire pipeline – from the spontaneous generation of a 3D world (Genie 2) through autonomous control by an agent (SIMA) to natural interaction via Gemini – makes it possible to develop games or interactive experiences almost “on the fly”. Developers can create prototypes very quickly and test experimental concepts without relying on traditional, time-consuming modeling and programming.
And if you spin that further, it also becomes interesting for video productions. From a single frame I let a world emerge and then place my characters in it. Give them backstories, political views and sexual preferences. And then let them interact with each other. Watch and place my cameras wherever I want. Or I program a VR world right away and the viewer stands in the middle of the film. I already made a video about the edit suite of the future – feel free to check it out here.
Which world would you like to have created, which open world would you like to play?
Write it in the comments below.
