It loads or compiles operations, manages tensor memory, selects kernels, and schedules work across devices. Graph optimizations and hardware-specific implementations can improve performance while preserving the model's numerical behavior within tolerance.