Motivation
What evolves together fits together
Our goal at Extropic is to build a new substrate for intelligence, which is fundamentally a full-stack undertaking involving hardware-algorithmic co-design. As algorithms on our hardware can be between 100x and 10,000x more efficient, there is urgency to port as much inference to this novel hardware as soon as possible.
Modern deep learning grew out of roughly fifteen years of researchers honing the architectures, objectives, optimizers, and training recipes in a way that is best suited to matrix multiplication accelerators such as GPUs and TPUs. One can view this decades-long process as a distributed evolutionary search over the hyperparameter space of deep learning models and methods. In order to successfully migrate most of AI inference to our hardware, we need thermo ML algorithms to be as mature as classic deep learning.
As thermodynamic computing represents a fundamentally different species of hardware, we need a corresponding new species of algorithms that are optimized for this substrate. We recently open sourced our thermodynamic programming stack to accelerate the search over thermodynamic algorithms and Thermo Model hyperparameters. From high-level programming with Torx, to compiling stochastic differentiable programs to our thermo hardware fabric with Thermalizers, to our low-level sampling framework THRML, we have created this stack to accelerate thermodynamic algorithm discovery.
Building the stack was only the first step. Most top programmers today program in natural language using LLM models trained on coding tasks. Unfortunately, as Thermo ML and stochastic differentiable programming are fairly nascent subfields, there is not as much literature and data in the pre-training runs of most frontier LLMs.
In order to both greatly accelerate the speed of research in thermo ML, and to make the migration for new users as smooth as possible, we are post-training our own language models on thermo ML tasks so as to build synthetic agents that are specialized Thermo AI researchers. This will in turn allow us to simply throw more compute towards rapid algorithmic maturation for this new substrate.
Mission
Dreams of Thermo RSI
As we seek to accelerate the search over architectures and hyperparameters for Thermo ML, we would like to emulate the algorithmic maturation of a distributed research process carried out by an academic community by instead using LLM-based AI agents. This will effectively allow us to “speedrun” decades of algorithmic progress in a few short years, ensuring algorithmic maturity as thermodynamic hardware begins to scale.
We are therefore post-training models to carry out Thermodynamic ML research: proposing methods, implementing them, running experiments, and revising their approach in light of results. Our aim is to bring this process to THRML, our framework for building probabilistic models, and to our simulator API, so agents can search for algorithms that compile and run well on our hardware.
This post introduces the first step in this mission: training an open-source model to reproduce classic experiments from the Hinton tradition. After 100 steps of RL, we roughly triple its base score on held-out tasks and get to results competitive with much larger frontier models.
Background
Why start with classic experiments
Before GPUs made backpropagation through large feedforward networks the standard, a branch of Machine Learning research treated learning as a problem of probability and energy. These were the early days of what we call Thermodynamic ML. Methods like Boltzmann machines and the wake-sleep algorithm, among others, learn by drawing samples from a probabilistic model, and on deterministic digital hardware that sampling is expensive, which is part of why the field moved on.
Our thermodynamic sampling units (TSUs) change that trade-off: TSUs are built to sample from probabilistic energy-based models natively, so the operation that made these methods costly becomes the one our hardware does best.
Reproducing these experiments also teaches an agent the right habits: translating a paper's description into working code, setting up sampling procedures correctly, and judging whether its results match the original. These are skills we will need to reach thermo RSI.
Post-training
Learning to reproduce foundational experiments
We ran post-training on Prime Intellect, who provided optimized inference and training tools, reproducible code sandboxes, and GPU resources to train and benchmark open-source models for this project.
We begin with roughly 50 coding tasks adapted from open-source reproductions of classic connectionist experiments in the Hinton tradition, including Boltzmann machines and wake-sleep.
We post-train Qwen3.6-35B-A3B on these tasks using a GRPO objective, with Nemotron 3 Super 120B as an LLM judge for reward scoring. The reward combines two signals, an execution score and a rubric score:
The execution score is deterministic, so no LLM judge is needed. It rewards an implementation first for running successfully (0.15 of the score), then for reproducing the reference task's core metrics, such as FID or loss (the remaining 0.85). Together, these tests measure whether the model writes code that both runs and recovers the behavior of the original experiment.
The rubric score uses Nemotron 3 Super 120B as an LLM judge. It checks the implementation against a task-specific set of criteria derived in advance from the original work. These criteria capture what the metrics can miss: whether the architecture matches the original, whether the experimental procedure was followed, and whether the implementation is reproducible.
We chose to weight execution more heavily because it is the harder signal to game: code either runs and hits the target metrics we set or it doesn't.
After 100 training steps, reward on held-out problems nearly triples, rising from 0.127 to 0.361.
Evaluation
Comparing against frontier models
We next compare our post-trained model against larger general-purpose models run zero-shot on the same suite of tasks. Post-training takes a small open-source model, Qwen3.6 (35B total, 3B active), from 0.127 to 0.361 on held-out reproduction tasks, about 3x its base score.
Our model is highly competitive with much larger frontier models on the research workflows we're optimizing for.
Outlook
Big things have small beginnings
These results are the first sparks of a broader effort to post-train relatively small models with useful research capabilities for thermodynamic computing. The natural next step is to move beyond reproducing established experiments toward discovering new algorithms.
Reproducing established experiments lays the groundwork for agents to devise new learning rules for our thermodynamic sampling units (TSUs), new algorithms, and new architectures. We plan to test those ideas in simulation, then, as our first large-scale chips come online in 2027, use measurements from real thermodynamic chips as feedback to inform subsequent experiments, thus creating the first automated Thermo ML research loop with real TSUs.
Deep learning's recipes were found by thousands of researchers over more than a decade. Thermodynamic computing doesn't and shouldn't have to wait that long. Agents that can read an idea, build it, test it, and learn from the result can explore the space of thermodynamic and stochastic algorithms far faster than we ever could in the past.
Our ambition is to enable recursive discovery for our thermodynamic platform: agents uncover thermo algorithms that expand what the hardware can do, and the hardware, in turn, can be co-evolved with the newly discovered algorithms. Ultimately, our novel hardware can also accelerate the inference of the agentic rollout, and thus make more discoveries per joule than ever before.
This is only the beginning. We plan to create far more ambitious post-training runs in the future, offer TSU-based sandboxes as environments, and offer inference on post-trained models.
We believe this will kick off a race for recursively self-improving hardware substrates optimizing for power efficiency.
Resources
Get involved
If the above sounds exciting to you, consider joining our team, or explore our software tools to post-train your own models to accelerate thermo AI research.
Note that everything in this post runs on tools you can use today. You can build probabilistic models with THRML, try your own algorithms on the simulator API, or explore the task data behind these results.