Multimodal large model training: fast computation blows up VRAM, saving VRAM crawls like a snail? Xiaohongshu open-sources BigMac, no more choosing between the two
Training a multimodal large model — either it's fast but blows up your GPU memory, or it saves memory but crawls like a snail? Xiaohongshu has open-sourced BigMac, so you no longer have to make that choice. How hard is it to train a multimodal large model? You can write the code, you can buy the compute, but the training framework has long been stuck in a "pick one" deadlock — either it runs fast but eats up memory, or it saves memory but moves at a painfully slow pace. The Dots Infra team at Xiaohongshu has cut through that knot.
💡 What You Will Learn
Training a multimodal large model — either it's fast but blows up your GPU memory, or it saves memory but crawls like a snail? Xiaohongshu has open-sourced BigMac, so you no longer have to make that c
📜 Table of Contents
Multimodal LLM Training: Fast Means Blowing Up VRAM, Saving VRAM Means Snail-Paced? Xiaohongshu Open-Sources BigMac, No More Trade-Offs
How hard is it to train a multimodal LLM? You can write the code, you can buy the compute, but the training framework has long been stuck in a "pick one" deadlock—either it runs fast but blows up your VRAM, or it saves VRAM but crawls at a pace that tests your patience. Xiaohongshu's Dots Infra team just cut through that knot.
On July 22, Xiaohongshu officially open-sourced BigMac, a new framework designed specifically for pipeline-parallel training of multimodal LLMs. The paper is up on arXiv, the code is publicly available on GitHub under the MIT license, and it's already been battle-tested in production.
What Does a Multimodal Model Look Like?
Let's crack one open and see what's inside.
Three components: the encoder converts images and audio into vectors, the LLM backbone handles reasoning, and the generator outputs results as images or speech.
These three pieces look very different from each other, and when you put them into the same training pipeline, that's when the trouble starts.
The Old Solutions: Both Roads Lead Nowhere
The industry has had two options.
✅ Compute-first approach: Pull the encoder and generator out of the LLM pipeline and run them separately, in parallel. The LLM doesn't get slowed down, but VRAM usage grows linearly with batch size—once the scale gets big, it's an instant OOM, and no amount of GPUs can save you.
❌ VRAM-first approach: Stuff all three components into the same pipeline and run them sequentially. VRAM usage drops, but if the encoder or generator lags even one step, the entire LLM has to wait. Imagine one person on an assembly line getting stuck while everyone behind them just stares—that's the "tail bubble." The larger the scale, the worse the idle time.
Two paths: one can't stay fast, the other can't get fast.
BigMac's Solution: Don't Split, Don't Stuff—Hitch a Ride
BigMac's approach is actually pretty straightforward.
The LLM pipeline itself is already highly optimized and runs reliably in production—no need to reinvent the wheel. So what about the encoder and generator?
Nest them into the natural gaps in the LLM pipeline.
How does that work? The LLM has its own execution rhythm—when one batch finishes computing and it's waiting for the next, there are idle gaps in the pipeline. BigMac inserts encoder and generator computations into those gaps—when input is ready, it slots in and computes; when done, it releases the VRAM immediately, without disrupting the LLM backbone's pace.
The team calls this "quasi-dependency-safe nested pipeline." In plain terms: the LLM is the main road, and the encoder and generator hitch a ride, squeezing into the gaps, finishing their work, and leaving no trace.
How Does It Perform?
Understanding tasks (with Qwen3-30B-A3B as the backbone and a 1.3B ViT encoder): - 1.08–1.1× faster than the compute-first approach - 1.6–1.9× faster than the VRAM-first approach - VRAM usage stays stable as batch size grows, while the compute-first approach blows up at larger batches
Generation tasks (with a 20B MMDiT generator added): - The compute-first approach OOMs at every batch size - BigMac not only runs it successfully but is also 1.5–1.9× faster than the VRAM-first approach
Why BigMac Is More Than a Paper—It's a Toolbox
BigMac open-sources three things:
- Scheduler: Generates a global operator table covering all pipeline stages, micro-batches, and module types
- Executor: Interfaces with backends like Megatron-Core, breaking the scheduler's plan into per-GPU local execution sequences
- Simulator: Estimates throughput and bubble rates for different configurations without actually running training
Algorithm engineers only need to describe what each module produces and consumes—everything else, from stage partitioning to activation passing to cross-device communication, is fully handled by BigMac. A multimodal experiment that runs on a single GPU can now scale more naturally to pipeline parallelism.
A Trend Worth Noting
Xiaohongshu's investment in AI infrastructure has been increasingly visible over the past couple of years. From dots.llm1 to dots.ocr to dots.vlm1, this company known for "grass-planting" (product recommendations) is taking a different path in underlying technology—not just building models, but building the tools that make models actually run.
The open-sourcing of BigMac means it's not just Xiaohongshu that can use this approach—any team working on multimodal LLM training can pick it up. For teams caught between VRAM limits and speed constraints, this could save months of reinventing the wheel.
Have you recently run into VRAM shortages or speed bottlenecks while training multimodal models? Drop a comment below and see if BigMac can save you some trouble.
[Image suggestion] figure-1-bigmac-comparison.png (comparison chart of compute-first vs. VRAM-first vs. BigMac) — insert before the "The Old Solutions: Both Roads Lead Nowhere" heading; figure-2-bigmac-data.png (benchmark comparison table for understanding + generation tasks) — insert before the "How Does It Perform?" heading.
Written by our editorial team; tools listed here are tested or verified against public sources. Links point to official sites or GitHub repos for reference only — no paid placements.
