A contract analysis deployment can look successful on a dashboard while quietly failing in reality. The trap is measuring activity, documents processed, clauses extracted, when what matters is outcome, decisions made faster and risks caught earlier. This piece defines the KPIs that distinguish the two, explains how to instrument them without fooling yourself, and shows how to read the signal once the numbers come in.
The framing throughout is that a metric is only useful if a bad value would change what you do. A number you would never act on is decoration. Every KPI below comes with the action a poor reading should trigger, because that is what separates measurement from theater.
Instrument these from day one, not after rollout. The most valuable comparison is before-and-after, and you cannot reconstruct a baseline you never captured.
Accuracy Metrics
Accuracy is the foundation, because every downstream benefit assumes the tool's output is trustworthy.
Recall, or the miss rate
Of the clauses that should have been flagged, how many did the tool catch? Misses are the dangerous failure, because a missed liability cap or auto-renewal can cost real money. Instrument this by having reviewers log clauses the tool failed to surface during the verify stage described in Triage, Extract, Verify: A Reusable Model for Reviewing Agreements.
Precision, or the false-alarm rate
Of the clauses the tool flagged, how many were genuinely worth a human's attention? Too many false alarms train reviewers to ignore flags, which silently destroys the tool's value. Track the share of flags a reviewer dismisses as noise.
Reading the accuracy signal
High recall with tolerable precision is the goal. If precision collapses, tune the playbook before reviewers stop trusting the tool. If recall is low, narrow the tool's scope to the documents it handles well.
The trade-off you cannot escape
Recall and precision pull against each other. Make the tool flag more aggressively to catch every important clause and you raise false alarms; tighten it to reduce noise and you risk missing things. There is no setting that maximizes both, so the real decision is where to sit on that curve for each document type. For high-stakes contracts, lean toward recall and accept more noise, because a missed liability cap costs more than a few wasted reviewer minutes. For low-stakes, high-volume work, lean toward precision so reviewers are not buried. Picking that balance deliberately, rather than accepting the vendor's default, is one of the highest-leverage configuration choices you will make.
Throughput and Time Metrics
Once accuracy clears a bar, time savings become the headline benefit.
Cycle time per document type
Measure turnaround separately for each document class, because the tool helps templated work far more than negotiated work. A blended average hides where the value actually is.
Reviewer hours reallocated
The point is rarely fewer hours overall; it is hours shifting from boilerplate to judgment. Track where reviewer time goes before and after, not just the total.
Queue depth and aging
Beyond average cycle time, watch how many documents sit in the review queue and how long the oldest ones have waited. Averages can look healthy while a handful of documents age dangerously, and an aging contract is exactly where a renewal window slips past unnoticed. A rising count of stale items signals a verify-stage bottleneck, often a missing owner, that no accuracy improvement will fix. This is an operational metric, not a model metric, and it frequently explains why a technically accurate tool still fails to deliver.
Risk and Outcome Metrics
These are the hardest to measure and the most persuasive to a decision-maker.
Caught obligations
Count the at-risk items the tool surfaced that would otherwise have been missed: renewals inside their window, missing required clauses, out-of-policy terms. Each caught item is a concrete, defensible result, the kind that anchors the Cost, Payback, and the Business Case for Review Automation.
Escaped misses
Track the items the tool missed that surfaced later through other means. This is uncomfortable to measure and the single most honest signal of real coverage.
Instrumenting Without Fooling Yourself
The metrics are only as good as the discipline behind them.
Capture a baseline first
Record current cycle times and known miss rates before the tool goes live. Without a baseline, every post-rollout number is unanchored and easy to spin.
Log disagreements, not just outputs
The richest signal is the gap between what the tool flagged and what the reviewer decided. A running disagreement log turns vague confidence into an error picture you can act on, exactly the practice that made the Inside One Legal Team's Move to Automated Redline Review rollout credible.
Reading the Whole Picture
No single metric decides the verdict. The skill is reading them together.
Combining the signals
A healthy deployment shows high recall, reviewers who still trust the flags, cycle time down on templated work, and a steady stream of caught obligations. If throughput is up but escaped misses are rising, you have bought speed at the cost of safety, which is a worse position than where you started. Read the metrics as a system, and let the weakest one set your next action.
Avoiding the vanity-metric trap
The numbers most often reported are the least meaningful: documents processed, clauses extracted, hours of usage. They rise simply because the tool is running and tell you nothing about whether it is helping. When you build a dashboard, resist filling it with these counts. A reviewer can process a thousand documents through a tool that catches nothing important, and the activity metrics will look triumphant. Anchor every reported number to a decision: would a bad value here change what we do? If not, it does not belong on the dashboard, no matter how good it looks in a status update.
Reviewing the metrics on a cadence
Metrics are not a launch checklist; they are a standing instrument. Set a regular review, monthly is common, where the team looks at recall, precision, queue aging, and caught versus escaped obligations together. The cadence matters because deployments decay quietly: document mix drifts, a model update shifts behavior, reviewers gradually start trusting flags they should question. A scheduled review catches the drift while it is still cheap to correct, turning measurement from a one-time justification into an ongoing control on quality.
Building a Dashboard That Tells the Truth
Metrics scattered across spreadsheets get ignored. A small, honest dashboard keeps the signal visible.
What belongs on it
A truthful dashboard shows recall and precision by document type, cycle time against the captured baseline, queue depth and the age of the oldest item, and the running count of caught versus escaped obligations. Each of these maps to an action, which is the test for inclusion. Notably absent are the activity counts, documents processed and hours used, that fill most vendor dashboards and prove nothing.
Who reads it and when
A dashboard with no audience decays. Assign the metric review to the same owner who runs the verify queue, and put it on a regular cadence so trends surface before they become problems. The most valuable view is change over time, not a snapshot, because a single good week tells you little while a steady drift in precision or queue age tells you everything. Pairing the dashboard with the disagreement log gives you both the what and the why: the dashboard shows that flags are getting noisier, and the log shows which playbook rule is responsible.
Frequently Asked Questions
Which metric matters most?
Recall, the miss rate, because a missed clause can carry real financial or legal consequence. Speed and volume mean nothing if the tool is silently letting important terms slip through unflagged.
Why is precision worth tracking if misses are the dangerous failure?
Because too many false alarms train reviewers to ignore flags entirely. Once people stop trusting the output, even accurate flags get dismissed, and the tool's value collapses regardless of its recall.
How do I measure something the tool missed?
Through escaped misses, items that surface later by other means, and through reviewer logs during the verify stage. It is uncomfortable to track, but it is the most honest signal of true coverage.
Why capture a baseline before rollout?
Because the most persuasive evidence is before-and-after, and you cannot reconstruct a baseline you never recorded. Without it, every post-rollout number is unanchored and easy to dismiss or inflate.
Should I report a single blended cycle-time number?
No. Blended averages hide where the value is, since the tool helps templated documents far more than negotiated ones. Report cycle time by document type so the real impact is visible.
Key Takeaways
- Measure outcomes, decisions made faster and risks caught, not activity like documents processed.
- A metric is only useful if a bad value would change what you do.
- Recall is the foundational metric, but precision protects reviewer trust in the flags.
- Capture a baseline and log tool-versus-reviewer disagreements from day one.
- Read the KPIs as a system, and treat rising throughput alongside rising escaped misses as a warning, not a win.