# Yash Maheshwari: Publications > Accepted research papers by Yash Maheshwari, research intern at the Shah Lab (Stanford Medicine) and the Lemons Lab (Stanford Graduate School of Education), and an independent researcher. The work covers small language models: sample-efficient pretraining, sub-1-bit quantization for on-device inference, and fine-tuned models as math tutors. This page lists accepted work only. ## Publications - [Halved CLM Exposure Mitigates Late-Training Degradation in Small Recurrent Language Models](https://research.yash-maheshwari.com/papers/babylm-2026). Yash Maheshwari. *BabyLM Workshop at EMNLP 2026*, 2026. Status: Accepted, Archival. OpenReview: https://openreview.net/forum?id=Dt0kB5nTNu PDF: https://research.yash-maheshwari.com/assets/papers/maheshwari-2026-halved-clm-exposure-babylm.pdf Halving CLM exposure mitigates late-training BLiMP degradation in small RWKV-7 models; auxiliary objectives are not required for the observed stability, while learning-rate trajectory remains coupled to update count. - [Below One Bit: Training Route and Budget Shape Robustness Under On-Device Storage Constraints](https://research.yash-maheshwari.com/papers/odi-2026). Yash Maheshwari. *NeurIPS 2026 Workshop on On-Device Intelligence: Foundation Models under Real-World Constraints (ODI)*, 2026. Status: Accepted, Poster. OpenReview: https://openreview.net/forum?id=uGyEuoSW6N Under a fixed in-format token budget, how a sub-1-bit model was trained into its format changes how much it degrades on corrupted input, especially on CPython, but the contrast is strongly budget-dependent, so bit width alone does not describe a compressed on-device model. - [Training and comparison of fine-tuned small language models as math tutoring assistants](https://research.yash-maheshwari.com/papers/jei-2026). Yash Maheshwari, Rinky Gupta. *Journal of Emerging Investigators (JEI)*, 2026. Status: Accepted, In press. Across nine fine-tuned Llama math tutors (1B, 3B and 8B parameters, 1 to 5 training epochs), parameter count mattered more than training time; the 8B model trained for five epochs gave a correct first step 93.65% of the time. ## Links - Publications page: https://research.yash-maheshwari.com/ - Main website: https://www.yash-maheshwari.com/ - Main website summary for LLMs: https://www.yash-maheshwari.com/llms.txt - GitHub: https://github.com/yashmahe2020 - LinkedIn: https://www.linkedin.com/in/yashmaheshwari2009/