It's a fast, cheap way to sanity-check a prompt or model change before investing in formal benchmarks or evals. It's useful for catching obvious regressions, but it doesn't substitute for measured evaluation on a representative test set when the stakes are higher.