Perplexity chief gave nine AI models a task to breach sandbox platform: Here are the results – CNBC TV18

Perplexity chief gave nine AI models a task to breach sandbox platform: Here are the results – CNBC TV18

By CNBCTV18.com
Publication Date: 2026-09-29 06:09:00

Amid the ongoing debate over AI safety regulations, Perplexity has announced that it is collaborating with Nvidia and over 100 industry partners to create infrastructure that contains rogue artificial intelligence agents.

CEO and co-founder Aravind Srinivas announced the results of a month-long test the company’s security team did to breach SPACE, the sandbox platform behind Perplexity Computer.

A sandbox is a secure, isolated environment where AI coding agents and the generated code can run safely without damaging host machines in any manner.


The term became associated with AI safety after the Hugging Face attack earlier this year, where a combination of OpenAI models gained internet access during internal testing and inadvertently breached the AI platform Hugging Face.

In a blog post on Perplexity’s website, researchers explained that a useful sandbox must contain accidental failure and also prevent AI agents that actively attempt to escape.

Perplexity CEO on his team’s new research

Srinivas said that his team gave nine models, including OpenAI’s GPT-5.6 and Anthropic’s Opus 5, root access inside the virtual machine (VM). He explained in a post on X that in some runs, the models were also given access to the full sandbox code.

The task assigned to the models was to either escape their host or reach a URL that SPACE’s network policy blocked.

“None of them got out of the VM in 108 runs,” Srinivas said.He added that once the researchers allowed some network access…