Delen
New tools for benchmarking AI help bridge the gap between laboratory success and real-world performance

Putting AI to the test before it reaches the real world

15 juni 2026

PhD researcher Tristan Tomilin developed new virtual ‘driving tests’ to reveal how artificial intelligence really behaves when faced with unexpected real-world challenges.

/

In a branch of artificial intelligence (AI) called reinforcement learning, digital AI systems learn by trial and error, in the process earning ‘rewards’ when they make the right decision. As training physical robots in the real world is too slow, expensive, and dangerous, scientists train them inside virtual, video-game-like simulations first.

However, most standard virtual tests only evaluate an AI on a single, unchanging task. In reality, practical applications require AI systems that can adapt to a chaotic world, cooperate with human teammates, and avoid dangerous mistakes.

To address this, PhD researcher Tristan Tomilin developed five new, open-source testing grounds that push AI out of its comfort zone. He defended his PhD thesis at the Department of Mathematics and Computer Science at ºÚÁϸ£ÀûÍø on Friday, June 12.

Designing richer and more realistic benchmarks

By building upon well-known virtual environments (including the 3D game Doom and the cooperative kitchen simulator Overcooked), translated his research goals into five distinct simulation benchmarks.

These flexible, high-efficiency testbeds allow researchers to systematically analyze how AI behaves when forced to think beyond a single, isolated task.

Rather than testing AI systems on a single, isolated task under fixed conditions, these environments introduce a series of realistic challenges that better reflect the complexity of the real world. These include visual stress tests that change the visual surroundings, memory tests that track how well the AI retains skills over time, teamwork setups that require coordinating with unfamiliar partners, and strict safety limits that force the AI to actively balance chasing high scores with avoiding dangerous risks.

Exposing limitations of current reinforcement learning methods

Experiments across all of these ‘driving tests’ reveal a consistent and sobering pattern: today’s leading AI methods are far less reliable than their famous successes suggest.

In the visual LevDoom test (a benchmark that challenges AI systems with changing lighting, textures, and object appearances) popular AI systems proved to be extremely fragile. When faced with simple, simultaneous changes in lighting, wall textures, or object sizes, the AI completely broke down. At the highest difficulty level, performance dropped dramatically, approaching the level of random guessing.

Furthermore, a long-term learning benchmark called COOM (Continual Reinforcement Learning Benchmark) exposed a critical bottleneck known as the ‘loss of network plasticity’. As the AI was trained on a continuous sequence of new tasks over time, its digital brain gradually became less flexible, eventually losing much of its ability to learn new tasks. It completely lost the ability to learn new skills, even though the computer system technically still had the capacity to do so.

Safety also proved to be a major roadblock. In a 3D safety test (HASARD), standard ‘Safe AI’ methods failed drastically when forced to navigate dangers based purely on what they could see through a virtual camera. Instead of avoiding hazards, the AI routinely ignored safety rules and risked major ‘accidents’ just to chase higher digital scores.

Finally, in a chaotic kitchen simulator (MEAL), the AI struggled to coordinate when paired with new, unfamiliar partners, especially when it had limited vision and could not see the whole room at once. In every scenario studied, human players easily outperformed the automated methods, particularly when the situations became complex.

Alongside these reality checks, Tomilin’s research delivers a significant practical breakthrough by vastly speeding up the underlying technology.

By utilizing the power of modern graphics cards through a new system called CRAX, Tomilin’s tests run up to 100 times faster than existing alternatives. This means complex experiments that used to take days or weeks can now be finished in just minutes, making large-scale AI research more accessible and cutting down on waiting times.

/

Broader implications

The results highlight that a flawless score on a single virtual test does not mean an AI is ready for the real world. It offers no guarantee that the system can handle complex settings, retain knowledge over time, cooperate effectively, or behave safely under pressure. By making these hidden limitations visible and measurable, Tomilin’s work takes a crucial step toward developing truly reliable AI.

These findings also carry a vital message beyond the lab, offering a timely reality check for marketing, communications, and administrative teams. While AI is often presented through its most impressive headlines, Tomilin’s PhD research shows that these systems can be surprisingly fragile when conditions change. In high-stakes areas like robotics, self-driving cars, and healthcare assistants, his research helps measure exactly when and how such systems can be safely trusted.

Finally, the high speed and efficiency of these new tests directly support university sustainability goals.

As they require a fraction of the computer power used by traditional methods, they lower electricity costs and the carbon footprint of large-scale AI testing. This can also have global implications, allowing smaller research groups with tighter budgets to participate in building the safe, adaptable AI of tomorrow.

Tomilin’s testing tools are now openly available to support the global community in creating AI that is truly ready for the real world. All benchmarks and software tools developed during this PhD have been released as open-source software and can be accessed through


PhD researcher Tristan Tomilin. 

  • Supervisors

    Mykola Pechenizkiy, Meng Fang

Written by

Bouri, Danai
(Communications Advisor M&CS)

More on AI and Data Science

Latest news

keep following us