Chapter 35 of 36 · ~1 min
Measure It
Evaluate the system against a defined benchmark rather than by the few impressive examples you remember. Build the evaluation set from chapter 18, define success criteria, run the prototype against it repeatedly, and record the results. Then change something and run it again.
This closes the loop on one of the central lessons of applied AI: if you cannot measure it, you do not really know whether you have improved it. The benchmark is also what lets you defend a decision to deploy, or not to deploy, to someone who was not in the room.
Exercise
Do it yourselfScore your prototype on your benchmark. Make one change. Score it again. Report both numbers and what changed.
Big question
What number would have to move for you to call this done?
