Crazy Wisdom

Crazy Wisdom

Episode #520: Training Super Intelligence One Simulated Workflow at a Time

Episode #520 · January 5, 2026 · 50 min

MP3 · Apple Podcasts · Spotify

About this episode

In this episode of the Crazy Wisdom podcast, host Stewart Alsop sits down with Josh Halliday, who works on training super intelligence with frontier data at Turing. The conversation explores the fascinating world of reinforcement learning (RL) environments, synthetic data generation, and the crucial role of high-quality human expertise in AI training. Josh shares insights from his years working at Unity Technologies building simulated environments for everything from oil and gas safety scenarios to space debris detection, and discusses how the field has evolved from quantity-focused data collection to specialized, expert-verified training data that's becoming the key bottleneck in AI development. They also touch on the philosophical implications of our increasing dependence on AI technology and the emerging job market around AI training and data acquisition.
Timestamps
  • 00:00Introduction to AI and Reinforcement Learning
  • 03:12The Evolution of AI Training Data
  • 05:59Gaming Engines and AI Development
  • 08:51Virtual Reality and Robotics Training
  • 11:52The Future of Robotics and AI Collaboration
  • 14:55Building Applications with AI Tools
  • 17:57The Philosophical Implications of AI
  • 20:49Real-World Workflows and RL Environments
  • 26:35The Impact of Technology on Human Cognition
  • 28:36Cultural Resistance to AI and Data Collection
  • 31:12The Bottleneck of High-Quality Data in AI
  • 32:57Philosophical Perspectives on Data
  • 35:43The Future of AI Training and Human Collaboration
  • 39:09The Role of Subject Matter Experts in Data Quality
  • 43:20The Evolution of Work in the Age of AI
  • 46:48Convergence of AI and Human Experience
Key Insights
  1. Reinforcement Learning environments are sophisticated simulations that replicate real-world enterprise workflows and applications. These environments serve as training grounds for AI agents by creating detailed replicas of tools like Salesforce, complete with specific tasks and verification systems. The agent attempts tasks, receives feedback on failures, and iterates until achieving consistent success rates, effectively learning through trial and error in a controlled digital environment.
  2. Gaming engines like Unity have evolved into powerful platforms for generating synthetic training data across diverse industries. From oil and gas companies needing hazardous scenario data to space intelligence firms tracking orbital debris, these real-time 3D engines with advanced physics can create high-fidelity simulations that capture edge cases too dangerous or expensive to collect in reality, bridging the gap where real-world data falls short.
  3. The bottleneck in AI development has fundamentally shifted from data quantity to data quality. The industry has completely reversed course from the previous "scale at all costs" approach to focusing intensively on smaller, higher-quality datasets curated by subject matter experts. This represents a philosophical pivot toward precision over volume in training next-generation AI systems.
  4. Remote teleoperation through VR is creating a new global workforce for robotics training. Workers wearing VR headsets can remotely control humanoid robots across the globe, teaching them tasks through direct demonstration. This creates opportunities for distributed talent while generating the nuanced human behavioral data needed to train autonomous systems.
  5. Human expertise remains irreplaceable in the AI training pipeline despite advancing automation. Subject matter experts provide crucial qualitative insights that go beyond binary evaluations, offering the contextual "why" and "how" that transforms raw data into meaningful training material. The challenge lies in identifying, retaining, and properly incentivizing these specialists as demand intensifies.
  6. First-person perspective data collection represents the frontier of human-like AI training. Companies are now paying people to life-log their daily experiences, capturing petabytes of egocentric data to train models more similarly to how human children learn through constant environmental observation, rather than traditional batch-processing approaches.
  7. The convergence of simulation, robotics, and AI is creating unprecedented philosophical and practical challenges. As synthetic worlds become indistinguishable from reality and AI agents gain autonomy, we're entering a phase where the boundaries between digital and physical, human and artificial intelligence, become increasingly blurred, requiring careful consideration of dependency, agency, and the preservation of human capabilities.
Episode transcript
Stewart Alsop III00:00

Welcome to the Crazy Wisdom Podcast. This podcast is for you. If you have an insane drive to find the truth of things, it's not the good answers that we seek, but the good questions. I interview a range of different guests from many different fields, all with the intention to uncover the simple truths that are hidden in plain sight. Most people don't want to go there. I go there, my guests go there, and you benefit. Please let me know if you enjoy these episodes and as always, subscribe on itunes, Spotify or wherever you listen to the podcasts.

Stewart Alsop III00:35

Welcome to the Crazy Wisdom Podcast. I've got Josh Holiday here and he is training superintelligence with Frontier Data at Turing. Welcome to the show.

Josh Holiday00:45

Thanks very much, Stuart. Great to be here. Excited to have a conversation with you today.

Stewart Alsop III00:49

So we were just talking about RL environments and I realized that I don't really know what an RL environment is. I know it stands for Reinforcement Learning Environment, and you've been doing this for a really long time. Can you share more what it is and how you got involved in this?

Josh Holiday01:04

Yeah, yeah, of course. It kind of ties nicely, actually. So I kind of first got into AI training data around five, six years ago. I was working at a company called Unity Technologies. You might be very familiar with them as like one of the gaming engines for 2D 3D games, but I was actually working on the industrial side. So we were working with enterprise organizations and we were building out a function for them essentially to generate simulated data where real world data couldn't fill, let's say, edge cases. This was kind of early on. So we were using this real time 3D gaming engine with physics embedded in it. And let's say we were working with an oil and gas company and they would come to us and they say, look, we've got this real world data set. There's certain situations or scenarios where it might be potentially hazardous for us to actually collect the data. But we need that data in case there's an accident. And we need our vision model to be able to predict what the outcome of that might be. So what we would do then is based on a scope, we would generate a 3D environment, so let's say like a construction site. And then we would have assets as well. So an asset could be like a person or an industrial vehicle. And we would create many, many different types of scenarios with many different variables in it to try and capture as much data as we could. And that was like the core idea behind it. And that was a very simple use case. You know, it actually it went so much deeper. We were talking with space intelligence companies that were building solutions like autonomous rendezvous and docking for satellites in space. And you know, how, how do you collect data for that? Well, you, you actually have to have a launch and you have to go up and collect the data. But what they're kind of looking for out there is space debris, which is essentially when you're looking at an image only really represented by a pixel, it's a tiny little thing. You can have a launch, you can collect some data, but it might not necessarily even have the full breadth of what you need. We were helping move from Blender pipelines into the Unity engine, which was real time 3D. What's funny about the other side of what I was doing was also supporting machine learning agents into Unity engine. So they go on a Python wrapper and go into the Unity engine and they would be trained on like a 2D game, let's say. So it's reinforcement. So every time that the agent gets the game wrong, it goes back to the beginning. It learns from its previous mistake with reward modeling and eventually it becomes very, very good at the game. This was like still very early days and it was more just experimental than anything. But it, for me it was just super interesting to be into, involved in that world. And it felt like very forward thinking at the time. Felt like we were really onto something.

Your first full transcript is free. After that, an email opens every transcript in the index — a list of readers we can write to, not a guest book.

Subscribe