Recent advances in language modeling have led to computationally intensive\nand resource-demanding state-of-the-art models. In an effort towards\nsustainable practices, we study the impact of pre-training data volume on\ncompact language models. Multiple BERT-based models are trained on gradually\nincreasing amounts of French text. Through fine-tuning on the French Question\nAnswering Dataset (FQuAD), we observe that well-performing models are obtained\nwith as little as 100 MB of text. In addition, we show that past critically low\namounts of pre-training data, an intermediate pre-training step on the\ntask-specific corpus does not yield substantial improvements.\n