build1 publisher
Dropping the critic network moves an RL fine-tune's cost from VRAM to rollout tokens
A dev.to walkthrough of DeepSeek's GRPO puts a 70B PPO training loop over 600GB of VRAM before any sharding. The critic-free version deletes the most expensive network and pays for it in sampled completions.
Publishers:dev.to
Reality
- Evidence32
- Adoption
- Insufficient
- Hype gap+30
- Incentives42
- Confidence45