An agent observes environment state, chooses actions, and receives feedback over a sequence of steps. Task success is measured by environment-specific outcomes rather than response style alone. Tool interfaces and prompting policies must remain consistent for meaningful comparisons.