DeepSeek's move to put training on AMD was fully completed yesterday
DeepSeek moved its training onto AMD, and yesterday the whole thing was fully completed. Yesterday in the AI circle, there was something that looked purely technical but was actually quite significant. The LMSYS team released an engineering report yesterday, stating that the reinforcement learning training for DeepSeek's latest-generation V4 Flash model—this step is the critical training phase where a large model goes from "being able to chat" to "being able to solve problems"—the entire pipeline has already been
💡 What You Will Learn
DeepSeek moved its training onto AMD, and yesterday the whole thing was fully completed. Yesterday in the AI circle, there was something that looked purely technical but was actually quite significant
DeepSeek's Move to Put Training on AMD — Fully Completed Yesterday
There was something in the AI world yesterday that looked purely engineering-focused, but was actually pretty significant.
The LMSYS team published an engineering report yesterday. It said that DeepSeek's latest-generation V4 Flash model's reinforcement learning training — the critical training phase where a model goes from "being able to chat" to "being able to solve problems" — has been run end-to-end on AMD's MI355X GPUs.
The report has 9 figures total, from training pipeline diagrams and architecture diagrams to precision comparison charts and reward curves, breaking down every step for you.
Why is this worth talking about? Not because AMD suddenly rose to prominence, but because domestic large-model training has been almost entirely dependent on a single thread — NVIDIA — for the past few years. GPU models, VRAM, and the surrounding software ecosystem all revolve around one company.
This time, DeepSeek has paved that path.
A Few Key Numbers
The model has 284 billion total parameters, but only activates 13 billion per response. Training used 4 nodes, with 8 AMD MI355X GPUs per node — 32 GPUs in total.
Over 100+ training steps, the "response probability" and "scoring probability" aligned, with the gap staying stable around 0.09 — no drift.
The final result on the AIME-2024 math problem set: single-attempt accuracy improved from 0.39 to 0.49, and 8-sample coverage improved from 0.53 to 0.67. The rate of truncated responses dropped from 60% to 55% — meaning the model's ability to "finish writing" also improved.
These numbers aren't mind-blowing on their own.
But they prove one thing: AMD's hardware path has gone from "it can run" to "it can reliably train models that produce results."
What's Actually Hard About This
The hard part isn't "getting the model to run" — it's "getting two different engines to produce nearly identical answers to the same problem."
You can think of reinforcement learning as "letting the model self-grade while solving problems." Get it right, and it rewards itself. Get it wrong, and it penalizes itself. Then it adjusts its reasoning based on the scores. But there's a prerequisite: the "answer generator" responsible for producing candidate responses and the "scoring engine" responsible for grading and tuning must output nearly identical probabilities for the same problem.
If there's even a slight mismatch, the model goes off track.
Why has NVIDIA's CUDA ecosystem dominated large-model training for years? Because this probability-alignment engineering between the "answer generator" and "scoring engine" has been polished to near-perfection on CUDA. Switching to AMD's ROCm stack means recalibrating every precision detail, every weight synchronization, and every round of multi-node communication from scratch.
Plenty of teams domestically and internationally have wanted to do this. But whenever it touches engineering details like reinforcement learning, mixed precision, or multi-node coordination, they almost all fall back to NVIDIA.
DeepSeek is the first to fully complete this pipeline.
What It Means for the Domestic AI Stack
It's a "hardware insurance policy."
Over the past few years, the training pace of domestic large models has been too tightly coupled to NVIDIA's supply rhythm. GPU production capacity, export controls, VRAM upgrade cycles — if any single variable hiccups, the entire training chain stalls.
Now that AMD's path has been proven to produce results, it means the next time you're buying new GPUs, negotiating prices, or building domestic inference services, you have another card in hand.
The impact on the consumer side will come more slowly, but the direction is already set. When the training side is no longer locked to a single vendor, hardware choices, pricing, and API costs on the inference side will loosen up. The next time you open a domestic AI assistant, the GPUs running behind it might no longer be a uniform model.
This isn't about money. It's about choice.
What's worth watching next is full-pipeline FP8 training, performance tuning, and validation at larger scale. Now that DeepSeek has paved this path, teams like Xiaohongshu, Kuaishou, Meituan, and Baidu — which are already using domestic GPUs for inference — will likely follow this route and migrate their training side to the AMD stack as well.
That day probably isn't too far off.
[Image suggestions] figure-1-pipeline.png (training pipeline breakdown) goes before the "What's Actually Hard About This" section; figure-2-metrics.png (data comparison) goes before "A Few Key Numbers"; figure-3-ecosystem.png (ecosystem diffusion) goes before "What It Means for the Domestic AI Stack." The three figures tell the story in sequence: how hard it is → what results were achieved → what happens next.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
