Human review cost in AI workflows is treated as a safety margin. It is not. It is a line item with a price per minute, and like every line item it earns its place or it does not. A reviewer checking an AI-extracted lease abstract is buying down expected loss, one field at a time. Where the expected loss caught exceeds the minutes spent, review pays. Where it does not, review is overhead dressed as diligence, and it crowds out the checking that matters.
Key Takeaways
The value of reviewing one field equals the field's error rate, times the reviewer's catch rate, times the cost of the error going uncaught.
The cost of reviewing one field equals the minutes spent on it, times the reviewer's fully loaded hourly cost.
Reviewers do not catch everything. Research on spreadsheet inspection summarized by Raymond Panko found individual inspectors detect roughly 50% to 60% of seeded errors.
Review priced per document hides the economics. Review priced per field shows that a few fields carry nearly all the value and most fields carry almost none.
Review stops paying when it is spread evenly across fields whose errors are cheap, and it keeps paying on fields whose errors move rent, dates, or options.
What does human review of an AI extraction cost per document?
Human review costs the minutes spent per document multiplied by the reviewer's fully loaded hourly rate. Using public wage data, a paralegal-level reviewer costs roughly $42 per hour loaded, or about 70 cents per minute. A 40-minute full review of one lease abstract therefore costs about $28 before any error is found.
The U.S. Bureau of Labor Statistics Occupational Outlook Handbook reports a median annual wage of $61,010 for paralegals and legal assistants in May 2024. Divided by 2,080 working hours, that is about $29.33 per hour in wages. The BLS Employer Costs for Employee Compensation release for June 2025 found wages and salaries were 70.2% of total compensation cost for private industry workers. Grossing up the wage by that share gives about $41.78 per hour, or $0.70 per minute.
Review scope per lease | Minutes | Cost at $0.70 per minute |
|---|---|---|
Full abstract, every field | 40 | about $28 |
Money and date fields only | 10 | about $7 |
One high-consequence field | 3 | about $2 |
One descriptive field | 1 | about $0.70 |
The minute counts are illustrative inputs; every firm should time its own reviewers. The point survives any reasonable inputs: review is cheap per field and expensive in aggregate, which is why it gets applied everywhere and questioned nowhere.
How do you decide whether reviewing a field is worth it?
Review a field when its expected loss caught exceeds its review cost. Expected loss caught is the probability the extraction is wrong, times the probability the reviewer notices, times the dollar cost if the error survives. If that product exceeds the minutes times the loaded rate, review pays. If not, the firm is buying assurance it does not need.
Review pays when error rate x catch rate x cost of error > minutes x loaded rate per minute.
Rearranged, it gives a break-even error rate for any field: the minutes times the rate, divided by the catch rate times the cost of the error. Below that error rate, the field should not be reviewed by default. Most review policies in CRE are set by habit, not calculation.
The catch rate is the input firms most often assume is 100%. It is not. Panko's summary of human error research, What We Know About Spreadsheet Errors, reports that individual inspectors in controlled experiments found roughly 50% to 60% of seeded spreadsheet errors. The lesson transfers to lease review: a single reviewer is a filter, not a guarantee.
Worked example: which fields on a lease abstract deserve review?
Review pays heavily on fields whose errors cost thousands and not at all on fields whose errors cost a few dollars. Running three fields through the rule, with the same assumed error rate and catch rate, shows value per lease ranging from $1,800 down to 30 cents. The consequence of the error decides everything.
Stated inputs: loaded reviewer cost of $0.70 per minute (derived above), a catch rate of 60% (the upper end of Panko's range for individual inspectors), and illustrative error rates and loss amounts chosen for a mid-size office lease.
Field | Assumed error rate | Assumed cost if uncaught | Expected loss caught (rate x 0.60 x cost) | Review minutes | Review cost | Verdict |
|---|---|---|---|---|---|---|
Renewal notice deadline | 2.0% | $150,000 | $1,800.00 | 3 | $2.09 | Review always |
Tenant notice address | 2.0% | $200 | $2.40 | 2 | $1.39 | Marginal |
Suite number | 0.5% | $100 | $0.30 | 1 | $0.70 | Do not review by default |
The renewal notice deadline pays for its review roughly 860 times over. The loss figure reflects a mis-dated option on a below-market renewal, as described in the real cost of a missed lease option. The suite number costs more to review than it saves. The notice address sits near break-even, so small input changes flip the answer.
Apply the rule to the whole abstract. If a lease has 60 fields and 8 are money and date fields, a 40-minute full review spends most of its $28 on fields where review returns less than it costs. A 10-minute targeted review of the 8 captures nearly all of the expected loss caught for a quarter of the price.
When does human review stop paying?
Review stops paying in three situations: when it is spread across fields whose errors are cheap, when a second or third pass adds little beyond the first, and when reviewers mostly confirm correct values and stop looking hard. Each converts a control into a cost nobody chose.
Failure mode | What happens | What the rule says |
|---|---|---|
Uniform review | Every field gets the same minutes regardless of consequence | Reallocate minutes from descriptive fields to money and date fields |
Stacked passes | A second reviewer re-checks the whole abstract | Each pass catches a share of what is left; value falls with every pass |
Confirmation drift | Reviewers see mostly correct values and approve by default | The catch rate in the formula falls, and so does the value of every minute |
Stacked passes deserve a number. If each independent pass catches 60% of remaining errors, one pass leaves 40% uncaught, two leave 16%, and three leave 6.4%. The second pass catches 24 points of the original errors and the third catches fewer than 10. On a renewal deadline worth $150,000, a second pass still pays. On a suite number, even the first pass did not.
Confirmation drift is the least visible. A workflow that sends a reviewer 60 fields per lease, 59 of them correct, trains the reviewer to confirm. The fix is not more review. It is less review, aimed better, so each field a human sees carries a real chance of being wrong. That is the logic behind routing by exception, covered in designing the extraction workflow around the exceptions, and behind reading a model's confidence score as a prediction, not a grade.
The line worth keeping: human review should be priced per field and spent where errors are expensive, because review applied evenly is paid for evenly and pays off almost nowhere.
How should a firm set its review budget?
Set the review budget field by field, from the bottom up. For each field, estimate the error rate from the firm's own testing, the cost of an uncaught error from the asset team, and the review minutes from timing real reviewers. Review the fields that clear break-even. Sample the rest periodically to confirm the error rate has not moved.
Error rates come from a held-out test of the firm's own documents, the same test described in the AI acceptance test. Loss amounts come from the people who reconcile rent rolls and track options. Review minutes come from a stopwatch.
Two rules keep the budget honest. First, never set a field's review to zero permanently: sample it at a low rate so a drift in error rate shows up before it compounds. Second, rerun the calculation whenever the extraction system changes, because a new model changes the error rate, and the error rate is the input that moves the verdict fastest.
Firms that price review per field compound an advantage: every reviewer minute goes where it buys down the most loss. Firms that review uniformly pay full price for assurance on suite numbers and still miss renewal-date errors.
Frequently Asked Questions
How much does human review of an AI-extracted lease abstract cost?
At a loaded cost of about $42 per hour derived from BLS data, review costs roughly 70 cents per minute. A 40-minute full review of one abstract costs about $28, and a 10-minute review of money and date fields costs about $7.
Should every AI-extracted field get human review?
No. A field deserves default review only when its error rate, times the reviewer's catch rate, times the cost of an uncaught error, exceeds the cost of the minutes spent on it. Low-consequence fields should be sampled, not reviewed on every document.
Does a second reviewer make the abstract safe?
A second reviewer reduces uncaught errors but does not eliminate them. If each pass catches 60% of remaining errors, two passes still leave 16% of the original errors uncaught, so a second pass is worth it only on fields where errors are expensive.
Where do the inputs for the review calculation come from?
Error rates come from testing the extraction system on the firm's own documents. Loss amounts come from the asset and accounting teams who absorb the errors. Review minutes come from timing actual reviewers on actual documents.
Conclusion
Human review is a cost line and should be managed like one. Run the per-field numbers and the answer is rarely "review everything." It is review the handful of fields that move rent, dates, and options, sample the rest, and rerun the math whenever the system changes. A firm that cannot say what its review costs per field cannot say whether its review is working.