Most AI projects are evaluated on how good the demo looks. A model reads a messy PDF perfectly, or answers a question fluently, and the project is approved. Months later, nobody can say whether the business is better off. The fix is to decide what will be measured before the pilot starts, and to measure the current process first.
The core operating metrics
Processing time: the elapsed time from work arriving to work completed, such as from purchase order email to order entered, or RFQ received to quote sent. This is usually the metric customers feel.
Manual touches: how many times a person has to handle each item. Reducing touches is often more valuable than reducing time, because every touch is a chance for error and an interruption.
Exception rate: the share of items the workflow could not complete without help, and why. A healthy workflow has a stable, well-understood exception rate and clear routing for the exceptions.
Cost per transaction: the fully loaded cost of handling one order, quote, or request. This is the number that connects the workflow to the business case.
Set the baseline first
None of these metrics mean anything without a baseline. Before a pilot, sample the current process for a few weeks: time a representative set of items, count the touches, and log the errors that are corrected downstream. It is tedious, but it is the only way to know later whether the workflow made a difference.
Report from real use
Once the workflow is live, report against the same metrics on a regular cadence, using the system's own logs and the approval data. Include what did not improve. A report that only shows good news is a sales document, not an operating tool, and the teams who rely on the workflow will notice the difference.
See how this applies in practice: how an engagement works.

