Menu

  1. Jul 20, 2026

    Field Extraction vs Full-Text Summary: What Your Underwriters Actually Need

Field extraction and full-text summary are sold as two flavors of the same thing. They are not. A summary is prose an underwriter reads once and forgets. Field extraction is structured data that flows into the model, gets validated against a rule, and traces back to a page. When a document AI tool hands your team a fluent paragraph instead of typed, sourced fields, it has moved the reading burden, not the modeling burden. The underwriter still keys every number by hand. The product you needed was the fields.

Key Takeaways

  • Field extraction returns named data points mapped to model inputs. A full-text summary returns prose a human still has to read and re-key. Only one of them ends the manual work.

  • A summary cannot be validated field by field, cannot be checked against a range rule, and cannot be reconciled to a rent roll. Structured fields can be all three.

  • IDC estimates roughly 80 percent of the world's data is unstructured. In CRE that unstructured mass is the offering memorandum, the lease, and the T-12. Turning it into fields is the actual job.

  • Manual single-entry keying carries a pooled error rate near 0.29 percent per field in controlled studies. On a 300-field deal that is roughly one silent error per model, and a summary hides it rather than surfacing it.

  • Buy the layer that produces auditable fields. Treat the readable summary as a convenience on top, not the deliverable.

What is the difference between field extraction and a full-text summary?

Field extraction pulls specific, named values out of a document and maps each one to a structured field: base rent, lease commencement, renewal option, cap rate, going-in NOI. A full-text summary compresses the document into readable prose. The first is data an underwriter models with. The second is a narrative an underwriter reads, then re-enters by hand.

The distinction is not cosmetic. A field has a name, a value, a type, and a location in the source. "Base rent" equals "$32.50 per square foot" found on "page 14, Section 4.1." That structure is what lets software validate it, sort it, and reconcile it. A summary has none of that. "The property offers strong in-place cash flow with modest escalations" is pleasant and useless to a model. You cannot subtract it from expenses. You cannot flag it when it falls outside a plausible range.

Why can't underwriters model from a summary?

A model runs on typed cells, not sentences. Underwriting requires every input to occupy a defined position: this number is the going-in cap rate, that number is the exit assumption, this one is year-one NOI. A summary delivers all of it as undifferentiated prose, so the underwriter has to read, locate, interpret, and transcribe each value before the model can move. The summary added a reading step and removed nothing.

Consider what the underwriter does with an offering memorandum. They are not trying to understand the property in general. They are trying to populate maybe 40 to 80 fields that drive the return: unit mix, in-place rents, market rents, operating expenses by line, debt terms, capital plan. A summary that says the property "has upside through a value-add renovation program" does not tell them the per-unit renovation budget, the rent premium assumed, or the number of units already turned. Those are fields. Only fields underwrite.

Prose also launders uncertainty. When an AI writes "the property generates approximately $2.4 million in net operating income," the underwriter cannot see whether the model read that off the T-12, computed it from line items, or inferred it. A structured field carries its own provenance and its own confidence. A sentence carries neither. This is the same reason source citations matter more than raw accuracy in extraction tools: a number you cannot trace is a number you have to redo.

What does field extraction give you that a summary cannot?

Field extraction gives you three things a summary structurally cannot: validation, reconciliation, and provenance. Each named field can be checked against a rule, compared against another document, and traced to the exact page it came from. Prose fails all three tests because it has no fields to check, nothing to reconcile against, and no per-value source.

Capability

Field extraction

Full-text summary

Flows directly into a model

Yes, mapped to input cells

No, must be re-keyed

Validates against a range rule

Yes, per field

No

Reconciles OM to rent roll to T-12

Yes, field by field

No

Carries page-level provenance

Yes, per value

No, at best a document-level note

Per-value confidence score

Yes

No

Surfaces the outlier for review

Yes, flags the field

No, buries it in prose

Validation is the sharpest example. A structured "year-one cap rate" field of 3.1 percent can be flagged automatically as implausible for the asset and market. The same number inside a paragraph passes unread. Data validation rules only work on fields. They have nothing to grip in prose. The moment your output is structured, the machine can defend the number. The moment it is prose, a human is the only line of defense, and humans skim.

When does a full-text summary earn its place?

A summary earns its place as a top layer, not the deliverable. It is useful for orientation: a broker's framing, the business plan narrative, unusual clauses that do not fit a field, the "why" behind a deal. Read it first to decide whether to underwrite at all. Never confuse reading the summary with having the data.

The honest case for summaries is real. Some content genuinely resists fielding. A co-tenancy clause with cascading conditions, an unusual earnout, a ground lease with a reset mechanism: these are narrative by nature, and a good summary that flags them for a human is worth more than a mangled attempt to force them into cells. The failure is not summaries. The failure is selling a summary as the finished underwriting input when the model needs fields. The best systems do both and keep them separate: fields for the model, a short narrative for context, and every field linked to its source so the two never drift.

How should you evaluate a document AI tool on this axis?

Test it on your own documents and ask one question: after the tool runs, how many numbers does the underwriter still type? If the answer is "most of them," you bought a summary with extra steps. A real field-extraction tool hands back a structured object, one value per field, each with a page citation and a confidence score, ready to load into the model.

Run a worked example to see the stakes. Assume a deal with 300 extractable fields across the OM, rent roll, and T-12. Manual single-entry keying carries a pooled error rate near 0.29 percent per field in controlled research, so 300 fields yields roughly 0.87 errors, call it about one silent error per deal on a good day. That is the controlled floor. Under real deadline pressure the same literature reports error rates climbing far higher. Now screen 40 deals a month. Even at the floor, that is roughly 35 undetected errors a month entering live models, each one buried inside a number no one re-derived. Field extraction with precision and recall you can measure and per-field confidence surfaces those outliers for review. A summary cannot surface what it never separated into fields.

The audit dimension compounds it. When an investment committee asks where a number came from, "the AI summarized the OM" is not an answer. "Base rent of $32.50 per square foot, page 14, Section 4.1, confidence 0.98" is. Structured extraction with a defensible audit trail is the only version that survives the question. As one way to put it: a summary tells you what a document says, a field tells you what to model and where you got it.

Frequently Asked Questions

Is a full-text summary ever the right output for underwriting?

As a top layer, yes. Use it for orientation, deal narrative, and clauses that resist fielding. Never use it as the input the model runs on, because prose cannot be validated, reconciled, or traced field by field.

What is field extraction in commercial real estate document AI?

Field extraction is the process of pulling named values, such as base rent, lease term, and going-in cap rate, out of a document and mapping each to a structured field with its source location. It produces model-ready data rather than readable prose.

Why does field-level provenance matter more than a good summary?

Because an investment committee asks where each number came from. A field carries its own page citation and confidence score. A summary carries neither, so any number inside it has to be re-verified against the source, which erases the time the AI was supposed to save.

How do I know if a tool does real field extraction?

Count how many numbers your underwriter still types after the tool runs. Real extraction returns a structured object, one value per field, each with a page citation and a confidence score. If the team is re-keying from prose, the tool produced a summary.

Conclusion

The question is not whether AI can read a document. It can. The question is what it hands back. A full-text summary moves the reading burden and leaves the modeling burden untouched, so the underwriter still keys every field by hand and still owns every silent error. Field extraction returns typed, sourced, validated data the model can consume and the committee can audit. For an underwriter, the fields are the product. The summary is a convenience printed on top. Buy the tool that knows the difference.

Related Reading

Get Started

Upload your lease documents. Rets does the rest.

Get Started

Upload your lease documents. Rets does the rest.