NeurIPS 2026 · ODI WorkshopAcceptedPoster
Below One Bit: Training Route and Budget Shape Robustness Under On-Device Storage Constraints
NeurIPS 2026 Workshop on On-Device Intelligence: Foundation Models under Real-World Constraints (ODI), 2026
TL;DR
Under a fixed in-format token budget, how a sub-1-bit model was trained into its format changes how much it degrades on corrupted input, especially on CPython, but the contrast is strongly budget-dependent, so bit width alone does not describe a compressed on-device model.
Abstract
Sub-1-bit weight quantization fits language models under stringent on-device storage budgets and is judged almost entirely on clean benchmarks. We ask what the route to those bits costs under corrupted input. Byte-level models trained from scratch across three corpora, four scales (0.86M to 10.03M parameters), 0.41 to 4.19 effective bits per weight, and six seeds per cell compare two routes to the same format at matched in-format token budgets: native quantization-aware training (QAT) and calibrated post-training quantization plus in-format fine-tuning (PTQ+FT). Under nine synthetic character-corruption families, the native route degrades less on CPython at every tested scale, width, and seed, and again on an independent Python corpus at 4.87M; on natural language the pooled advantage, normalized by each domain's own degradation scale, is 8 to 20 times smaller at comparable widths and not generally seed-verified. The contrast tracks the recipe, not the format: quadrupling the native token budget reverses the sign, and matching total lifetime tokens leaves no seed-verified difference at either code scale tested. On an Apple M5 Max the packed checkpoint is 8.8 to 22.9 times smaller than 16 bits per parameter, growing with scale, but with no fused sub-1-bit kernel, unpacking to execute costs 6.0 to 16.1 times keeping the weights resident at batch 1. Bit width alone does not describe a compressed on-device model; route and token budget belong on the spec sheet next to it.
Keywords
on-device inference · quantization · quantization-aware training · post-training quantization · sub-1-bit compression · robustness · input corruption · byte-level language models · model compression