Marin: What an Open Lab Actually Solves
Most open models come with weights and a report. Marin also publishes the lab notebook, failed experiments included, while the training is still running.
I'm teaching myself how LLMs are trained, beginning with derivatives. I gave myself one end goal: understand Marin's current 535B run well enough to redesign it on paper. So I've spent a lot of time in their reports, GitHub issues and blog posts.
What Marin is
Marin is an open lab for foundation models. It started at Stanford's Center for Research on Foundation Models and was announced in May 2025, together with Marin 8B.
Each experiment is a GitHub issue that doubles as a small preregistration, saying what they will try and why. The code goes in a pull request that anyone can review. Training logs are public on Weights & Biases, and the failures stay up next to the wins.
So far there is Marin 8B (May 2025), Marin 32B (late 2025), and a 535B mixture-of-experts model that is training right now.
Problem 1: "Open" usually means open weights
Their announcement says it plainly: "We have open-weight models (e.g., Llama, DeepSeek, Gemma), sometimes mistakenly called open-source models, but the code and data (recipe) used to produce the model are not released."
Weights let you use a model. They don't teach you how to make one. Some projects, like AI2's OLMo and EleutherAI's Pythia, already release the code and data along with the final model. Marin also makes the process public while it happens, with the decisions and dead ends included.
Problem 2: Loss spikes the usual fixes can't fix
Their 32B retrospective is the best training war story I have read. Roughly what happened:
- For the first 70K steps, training looked fine but had many more loss spikes than the 8B run. Some people told them they were doomed.
- They tightened gradient clipping. Normal gradient norms sat around 0.2, and big jumps usually came right before a spike.
- At about 72K steps they added "update norm" clipping. It tracks a rolling mean and standard deviation of the optimizer's step size and clips anything more than two standard deviations above it.
- They added a rule to skip bad steps.
- They accidentally turned the update-norm filter off for a few thousand steps, between 74K and 80K.
- At 80K the spikes became impossible to avoid. A "necromancy" restart, which rebuilt the optimizer state offline, held for a few thousand steps and then relapsed. A run with the Muon optimizer diverged.

The fix was in the architecture. They switched to the Qwen3 32B layout, which adds QK-norm to attention, and warm-started it from the Llama weights at step 80K. The loss took a one-time hit and recovered in about 10B tokens. After that, the spikes were gone completely.

Why not use QK-norm from the start? Their honest answer: "we had some hubris from our 8B experience." The 8B model had trained fine without it, and so had an earlier 70B trial.
Their own summary of the episode is the line I'd tape to my monitor: "instrumentation alone could not replace an architectural fix at this scale."
A normal technical report would turn all of this into one sentence: "we use QK-norm." The retrospective shows what that sentence cost.
Problem 3: The data bugs nobody writes about
During the cooldown, GSM8K test data leaked into training through a cached data bundle. Separately, a permutation built on a linear congruential generator ordered the data badly, and they replaced it with a Feistel-based shuffle.
Both bugs change the numbers you would report, and bugs like these rarely make it into a technical report.
Problem 4: Predicting a big run from small ones
You only get one shot at a 535B run, so you want to know the outcome before you start. Marin built a scaling suite called Delphi for this, and the first version failed in a useful way. Fitted on small runs, it missed the 1e22-FLOP run by 2.5%, and the 1e23 run diverged completely.
The causes were ordinary. The learning-rate rule gave rates that were too high for long, data-heavy runs with large batches. Weight decay was fixed at 0.1 for every run, so it did not change with training length or model width.
They moved to an optimizer with no weight decay to tune (AdamH), added a correction that lowers the learning rate for longer runs, and retuned the reference model. The second version was fitted only on runs up to 3e20 FLOPs, and it predicted the 1e21, 1e22 and 1e23 runs to within +0.5%, +0.2% and +0.2%.
Before launching the 535B run, they wrote down its predicted final loss: about 2.04 on the Paloma evaluation set. The prediction comes from a ladder of smaller runs that used about 1% of the total compute. When the run ends, anyone can check it.
Problem 5: Frontier-scale training from an academic lab
The current run is Marin 535B-A23B: 535B parameters in total, with about 23B active for each token. The main numbers, from the run's public spec:
- 48 layers, 384 routed experts with 8 picked per token, plus 2 shared experts that every token uses
- 18T training tokens, about 2.7e24 FLOPs in total
- 11 GB200 NVL72 racks on CoreWeave, funded by a grant from the Jen-Hsun and Lori Huang Foundation. Their earlier models trained on TPUs from Google's TPU Research Cloud.
Even the parallelism choice is written up. The full training state takes 16 bytes per parameter, close to 8 TiB in total. Sharding the experts with FSDP would move about 20 GiB of weights per GPU per layer, while expert parallelism moves about 6 GiB of activations. In their small-scale ladder, that gave an estimated 18 to 35% better compute efficiency. Hardware utilization went from 15% in early tests to about 23% in production.
You can also just ask them things. In the run's GitHub issue, someone asked why they project tokens into a smaller "latent" before sending them to experts. A Marin engineer answered with numbers. On its own, the latent projection halves the communication cost. Without a learnable norm on it, quality got 30% worse. Moving from 4-of-192 experts to 8-of-384 smaller experts improved quality by 15% but doubled communication, and the latent trick gets that gain without paying for it in communication or compute.
As of October 7, the run is 56% done (10.1T of 18T tokens), and the projected end is November 14. The status page lists every stop. The latest was on October 6, when Kubernetes marked one node bad and deleted its training task, which stopped the other 175 tasks too.
What I'm taking from it
- Their MoE experiment digest covers 80 ideas: 16 worked and 32 did not. Publishing the 32 failures saves everyone else from paying for them again.
- Writing the predicted loss down before the run turns "it went well" into a claim anyone can check.
- Most of the hard work is unglamorous: loss spikes, a bad shuffle, contaminated data, a dead node. The architecture is one decision. Keeping the run alive takes hundreds more.
- It is the best textbook I have found. The issues read like lab notebooks. If you want to learn how training really works, read the 8B and 32B retrospectives before any survey paper.
Sources: Marin, 8B retrospective, 32B retrospective, Delphi, Expert parallelism at hero scale, 535B launch note, run tracker, GitHub issue #8435.
Cover: Joseph Wright of Derby, The Alchemist Discovering Phosphorus, 1771, detail. Public domain. In Chinese ML slang, training a model is 炼丹, "refining an elixir." Marin is trying to turn the alchemy into chemistry.