Chapter 35 of 36 · ~1 min

Measure It

Evaluate the system against a defined benchmark rather than by the few impressive examples you remember. Build the evaluation set from chapter 18, define success criteria, run the prototype against it repeatedly, and record the results. Then change something and run it again.

This closes the loop on one of the central lessons of applied AI: if you cannot measure it, you do not really know whether you have improved it. The benchmark is also what lets you defend a decision to deploy, or not to deploy, to someone who was not in the room.

Exercise

Do it yourself

Score your prototype on your benchmark. Make one change. Score it again. Report both numbers and what changed.

Big question

What number would have to move for you to call this done?