Model Evaluation Should Be a First-Class Engineering Discipline

https://hackernoon.imgix.net/images/y4RwzlKAVleRegqPbstUBzukKU32-ci83b4t.png

Ask an ML engineer what they work on and you will hear about models, training runs, and serving infrastructure. Almost nobody says evaluation. Yet for nearly two years I led the design of an evaluation framework for the perception system of an autonomous vehicle, and I came away convinced that measurement is the one of the hardest engineering problems in production machine learning, and often the least discussed.

Here is the uncomfortable claim at the center of this piece: your accuracy number may be perfectly computed – and still answer the wrong question. And when the system in question can fail in the physical world, that gap is not academic.

The aggregate metric is structurally inadequate

Every ML team lives by a handful of aggregate numbers. Accuracy, precision and recall, mean average precision. These numbers feel objective, and on a fixed benchmark they are. The problem is what they silently...

Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE

Read more