SLMs trade world knowledge and multi-step reasoning for lower memory, faster decode, and the option to run in a browser or on a phone. MicroLLM Lab (September 2026) demonstrates Q4 models from about 26 million to 360 million parameters in WebGPU; they are commonly used for classification, routing, and privacy-sensitive triage rather than as general assistants.