It is a stricter case of edge inference: the target is a battery-powered sensor, wearable, or industrial controller with milliwatt power budgets and often no network connection at all. Models are aggressively quantized (commonly int8 or lower) and pruned to fit in on-chip SRAM/flash, and the whole pipeline — capture, inference, action — typically runs in microseconds to milliseconds on a single core.