Some big AI companies seem to think that the best way to keep models from becoming too chaotic or too mischievous is to keep them locked inside of labs. If only a chosen few can access them, the thinking goes, they can do less damage in the real world while researchers try to understand what they’re capable of.
Nathan Lambert and Tom Zick, two industry scientists, believe the opposite. The pair founded a nonprofit, Trillium Labs, that will work on various areas of AI research—including potentially problematic areas like recursive self-improvement (RSI) and agents—in a more transparent way. In practice, this will mean publishing the details of experiments so that outside scientists can study and replicate them.
Lambert says the way frontier AI labs keep their work secret reduces the community’s ability to scrutinize ideas and contribute new approaches. He believes that letting outside experts see how models are built and tuned could be crucial to mitigating risks.
“Over the past few millennia, humanity has had the scientific method in our toolbox as a way to mitigate harms and build better futures,” Lambert tells WIRED. “The current closed trajectory of frontier AI development is taking us a step backwards.”
The world’s most powerful models, like those from OpenAI and Anthropic, can only be accessed through an app or an application programming interface (API). Often, this comes at the cost of transparency about how the model is built and how it behaves.
Other companies, especially those in China, offer relatively powerful models that can be downloaded and run on a user’s own hardware. The Chinese company Xiaomi, for example, recently published live details of a major training run involving one of its models. And researchers at Stanford are pretraining the AI model Marin in the open.
The industry is currently locked in a battle over which strategy is best, mostly because of how powerful frontier models now are. They can automate the discovery of new software vulnerabilities and automatically probe and hack into systems, and recent high-profile hacking sprees have prompted even greater scrutiny.
Proponents of a limited-access system say it’s crucial to keep that power in the hands of a trusted few, while those in Lambert and Zick’s camp believe that a shared understanding of the risks means we’re all better off.
Lambert previously worked at Ai2, a research lab that has taken an unusually open approach to AI, including publishing details of the data and the training methods used to build models alongside the models themselves. He previously worked at Hugging Face, runs a popular technical blog, and founded the American Truly Open Models, an initiative aimed at encouraging US companies to release more open models. Zick worked at Harvard University and helped Charles Schwab devise policies around “responsible AI.”
The two met over Zoom during the COVID-19 pandemic, when both were graduate students at UC Berkeley working on AI. They got the idea for the new nonprofit after seeing how disconnected industry AI research has become from academic work; Lambert says professors and students are often unable to replicate the work going on inside big company labs because they lack the resources required.
Zick says Trillium Labs, which launched today, will initially focus on post-training—fine-tuning large models after they’ve been built. Another key area will be RSI, a process for developing new models by having AI contribute research. The prospect that ongoing progress could continue indefinitely, leading to a loss of human control, has alarmed many AI researchers. The issue gained mainstream attention earlier this month when an Anthropic researcher left the company and warned that RSI could pose an existential threat to humankind.
The nonprofit will also look at how reinforcement learning, which rewards a model for good results and punishes it for bad outcomes, can improve its capabilities. That approach has made agents far more capable, but also more inclined to do unexpected things. They’ll study how reinforcement learning shapes the character and behavior of AI models, a method that can pose problems when a model becomes overly sycophantic, for example.
“To understand something like how reinforcement learning scales in post-training, you need significant compute and a lot of careful experimentation,” Zick says. She says that publishing details of how reinforcement training runs work could yield surprising insights as outside researchers scrutinize the work.
The lab has raised an undisclosed sum from Schmidt Sciences, Halcyon Futures, and others. The founders say they aim to raise $40 to $100 million in total and plan to spend $30 million on training over the next 18 months.
“I’m a massive fan of much more transparency than we currently have in R&D,” Tim Fist, director of emerging technology policy at the Institute for Progress, a policy thinktank, tells WIRED.
Lambert and Zick ultimately hope that Trillim Labs will contribute some much-needed nuance to the wider discussion about how best to build AI.
“We’re in an era of AI discourse dominated by a few world views,” Lambert says. “We believe that the scientific method and careful measurement of recent events is the best way to understand new behaviors of AI models.”
