FLOP/s measures how many floating-point math operations a processor can perform each second. AI workloads primarily use matrix multiplications, so GPU specs emphasize tensor FLOP/s at different precisions — FP16, BF16, FP8, and INT8. NVIDIA's Blackwell B200 achieves roughly 9 petaFLOP/s at FP4 precision. Training compute is often reported in total FLOPs (not per second), with frontier models requiring 10^25+ FLOPs.