EMNLP 2026 · BabyLM WorkshopAcceptedArchival

Halved CLM Exposure Mitigates Late-Training Degradation in Small Recurrent Language Models

Yash Maheshwari

BabyLM Workshop at EMNLP 2026, 2026

BLiMP accuracy over training for vanilla RWKV-7, vanilla-skip and BiRWKV-7, and the peak-to-final drop for each across three seeds.

TL;DR

Halving CLM exposure mitigates late-training BLiMP degradation in small RWKV-7 models; auxiliary objectives are not required for the observed stability, while learning-rate trajectory remains coupled to update count.

Abstract

Small recurrent language models can degrade late in training under standard causal language modeling (CLM). On the BabyLM Strict-Small corpus, vanilla RWKV-7 peaks near 69% BLiMP early, then loses 2.57 to 3.85 pp by step 18k across all three seeds we ran. Halving how often CLM updates the weights mitigates this decline, via two mechanisms: alternating a backward MLM objective every other step (BiRWKV-7), or skipping every other gradient step (vanilla-skip). Vanilla-skip carries no auxiliary objectives yet beats BiRWKV-7 on peak BLiMP in every seed, so the auxiliaries are not necessary for stability. What halving changes is coupled: it reduces the CLM update count and doubles the learning-rate decay in update space. A constant-rate control still degrades at the two higher rates we tested, so a low terminal rate is not necessary for the decline; but at fixed update count, flooring the rate earlier recovers most of it. The update-space rate trajectory therefore also modulates the decline; it stays coupled to update count, and our design does not separate the two. On the official BabyLM 2026 pipeline our submission scores 68.70% BLiMP, above the GPT-2 baseline (65.23%, 98.4M) at about 28% as many parameters.

Keywords

BabyLM · RWKV-7 · recurrent language models · sample-efficient pretraining · data-constrained pretraining · causal language modeling · late-training degradation · BLiMP