Forget chatbot training. AI's next big data grab is about learning how humans work.
Google's Mechanize talks and Meta's moves reveal the next AI race: training agents to do real jobs, not just answer questions.
Courtesy Tamay Besiroglu
- Mechanize aims to automate jobs using RL environments, drawing interest from Google.
- AI development shifts from chatbot training to RL environments, enabling complex task learning.
- Meta tried to track employees' workplace actions and workflows, for use in AI development.
The AI race is shifting away from teaching chatbots how to answer questions and toward teaching agents how to perform entire jobs.
After years spent training models on internet text and paying contractors to rate chatbot responses, tech companies are increasingly focused on something new: creating realistic digital workplaces where AI can practice coding, using business software, making decisions, and completing long-running tasks much like a human employee.
These digital training realms are called reinforcement learning environments, or RL environments for short. Two recent developments show how this is catching on across the industry.
Google is in talks to invest more than $1.5 billion in Mechanize, a startup that builds virtual work environments to train AI agents, Business Insider exclusively reported last week. The potential deal would bring Mechanize talent to Google while the tech giant licenses some of its technology, with the team working on model evaluation and development.
Meanwhile, Meta has been trying to collect employees' keystrokes, mouse movements, clicks, and other screen activity, saying the goal is to help AI learn how people actually use computers — from keyboard shortcuts to navigating workplace software.
From chats to workflow
Taken together, these moves point to a broader shift in how frontier AI systems are being developed.
The first generation of large language models was built on two main data foundations: enormous amounts of text scraped from books and the internet, followed by human feedback from contractors who judged whether chatbot answers were good or bad.
That approach created remarkably capable conversational AI. But it has struggled to produce AI agents that can independently complete complex, multi-step work over hours or days. Increasingly, the industry's answer is RL environments, realistic simulations where AI agents learn by doing rather than simply predicting the next word.
Scale's pivot
Scale AI, one of the biggest suppliers of training data, says top AI companies are moving away from relying solely on static datasets and human preference feedback toward simulated environments where agents can safely learn through trial and error.
"The way models are trained is evolving," Chetan Rane, head of product for agents & RL environments at Scale AI, wrote in a recent blog post. "Frontier models increasingly need to learn through trial and error in realistic simulated environments rather than relying solely on static datasets or human preference feedback."
Nearly half of the company's new AI training projects now involve RL environments, which model realistic coding, computer use, and enterprise workflows, he noted.
Instead of asking whether a chatbot produced the right answer, these environments allow an AI to attempt an entire workflow, make mistakes, recover, and improve based on whether it successfully completed the task.
"Full automation"
Mechanize is betting this will become one of the most important parts of AI development.
The startup was launched last year by AI researcher Tamay Besiroglu with an unusually ambitious goal: "full automation of the economy." At the time, he was criticized for such a bold mission.
"We will achieve this by creating simulated environments and evaluations that capture the full scope of what people do at their jobs," Besiroglu and his cofounders wrote in a blog announcing Mechanize. "The market potential here is absurdly large: workers in the US are paid around $18 trillion per year in aggregate. For the entire world, the number is over three times greater, around $60 trillion per year."
Reward signals
The startup says today's AI systems remain unreliable at long-running work because they lack realistic environments in which to learn.
Its first target is software engineering. Mechanize argues that future coding agents will first learn from examples of professional programmers before improving through reinforcement learning inside increasingly realistic software environments that capture the complexity of real engineering projects.
As those environments improve, the company believes the same approach can expand into all kinds of white-collar work.
Reward signals are a key part of these RL environments, he told Business Insider in an interview last year.
"You want to be able to tell the model you did the task correctly versus incorrectly," Besiroglu explained. "Then you want to leverage that to reinforce the kind of patterns of behavior that resulted in it correctly performing the task."
Meta's data grab
That thinking also offers a possible explanation for why large technology companies are suddenly interested in collecting data on how their own employees actually work.
Meta's internal announcement said its tracking software would help AI understand how people complete everyday computer tasks because agents need to be trained on real examples. The software captures inputs such as mouse movements, clicks, keystrokes, and screen context across approved workplace applications.
Uber's "Agentic Pods"
Other companies are pursuing similar ideas.
Uber recently said it has begun embedding top AI engineers inside departments including finance, legal, HR, marketing, procurement, and customer support to observe how employees work before redesigning those workflows around AI. The company calls the initiative "Agentic Pods."
The result is that the AI industry's newest race may no longer be about building smarter chatbots.
Instead, companies are competing to build detailed digital versions of real workplaces, where AI agents can practice the thousands of decisions, actions, and workflows that make up modern jobs.
If the industry's biggest players are right, these virtual workplaces will become the training grounds for the next generation of AI.
Sign up for BI's Tech Memo newsletter here. Reach out to me via email at [email protected].
Read the original article on Business Insider