Mixture weights (how much code vs conversation vs multilingual web) shape what the base model is good at before any instruction tuning. Crawls include licenses, PII, and quality cliffs, so labs filter aggressively and still miss things. The corpus is usually not fully public even when weights are.