The CRE AI data moat does not sit inside the model. It sits in the data. A firm can license the same foundation model as everyone else, and most firms do, because the frontier is converging fast. What no competitor can license is a firm's own corpus of corrected, labeled commercial real estate documents: the leases, offering memoranda, and rent rolls a team has read, tagged, and fixed over years. The model is a rented commodity. The proprietary data behind it is the only asset that compounds and the only one a rival cannot buy.
Key Takeaways
Foundation models are converging. Stanford HAI's 2025 AI Index found the score gap between the top and tenth-ranked model fell from 11.9% to 5.4% in a single year, and the top two models were separated by 0.7%. A model that any competitor can license is not a moat.
The defensible layer in CRE AI is proprietary labeled data, not the base model. Anyone can call the same model. No one else holds your corrected corpus of leases, offering memoranda, and rent rolls.
General-purpose models fail on verifiable domain facts. The Stanford study "Large Legal Fictions" found leading models hallucinated between 58% and 88% of the time on specific, checkable legal questions. Domain labels are what close that gap.
Data moats are conditional, not automatic. Andreessen Horowitz argues raw data advantage suffers diminishing returns. The moat holds only where the data is proprietary, labeled, and drawn from a domain with no public corpus, which describes CRE documents.
The same base model yields different accuracy for different firms. The difference is the labeled data used to adapt it, and that data is the asset an operator should be accumulating.
Why isn't the model itself a moat in CRE AI?
The model is not a moat because every firm can rent the same one, and the best models now perform within a few points of each other. Stanford HAI's 2025 AI Index reported that the score gap between the top and tenth-ranked model narrowed from 11.9% to 5.4% in a year. Convergence, not divergence, defines the frontier.
When the top ten systems cluster within roughly five points, the choice of model stops being a source of advantage. It becomes a procurement decision, like choosing a cloud region. The same report found that nearly 90% of notable models in 2024 came from industry, which means the supply of capable base models is expanding, not concentrating. Abundant supply is the opposite of scarcity, and moats are built on scarcity.
There is a quotable version of this: a capability every competitor can buy at list price is a feature, not a moat. The base model is the feature. The question that decides outcomes is what each firm feeds that model and how well it has taught the model to read documents that no general corpus contains.
What actually makes proprietary data a defensible moat?
Proprietary data is defensible when it is labeled, domain-specific, and impossible to reconstruct from public sources. A commercial lease with a firm's own corrections, field tags, and audited outcomes cannot be scraped, licensed, or synthesized by a rival. The moat is not data volume. It is verified, structured knowledge of documents that live behind NDAs and deal rooms.
The distinction becomes clear when you separate the model layer from the data layer and ask a single question of each: who can copy it.
Layer | What it is | Who can copy it | Defensibility |
|---|---|---|---|
Base model | A general foundation model licensed by API | Any competitor, same day, same price | None. It is a commodity input |
Prompts and workflow | The instructions and steps wrapped around the model | A competent competitor, within weeks | Low. Observable behavior can be reverse-engineered |
Public data | Filings, listings, published comps | Anyone who buys the same feed | None. Shared supply confers no edge |
Proprietary labeled data | Your corrected leases, OMs, and rent rolls with field-level tags | No one. It does not exist outside your systems | High. It compounds with every corrected document |
The bottom row is the moat. It is the row a competitor cannot shortcut with capital, because the asset is not for sale. Every lease your team abstracts and corrects, every offering memorandum your model misreads and a human fixes, becomes a labeled example. That labeled example is training data no one else holds. Over years, the corpus becomes a private map of how CRE documents actually behave, including the messy amendments and side letters a general model has never seen. This is the substance of vertical AI: a system adapted to one domain's documents rather than a general model prompted to guess at them.
Why does the same base model produce different accuracy on proprietary labeled CRE data?
The same base model produces different accuracy because two firms adapt it with different data. One firm prompts the raw model. The other fine-tunes it on thousands of corrected examples from its own deals. Same weights at the start, different accuracy at the end, because the second firm taught the model what its documents mean. The gap is the data, not the model.
Consider a worked example. Assume two firms license the identical base model to extract fields from offering memoranda. The numbers below are illustrative, derived from the stated inputs to show the mechanism, not measured benchmarks.
Firm A prompts the base model with no domain adaptation. On a single field, say in-place net operating income, assume it reads the value correctly 90% of the time. That sounds strong.
A completed extraction is not one field. Assume an OM summary requires 12 fields read correctly to be usable without rework: NOI, occupancy, in-place rent, market rent, expense ratio, and so on.
The probability that all 12 independent fields are correct at 90% each is 0.90 to the 12th power, which is about 0.28. So roughly 72% of documents need a human to catch and fix an error.
Now Firm B. It has abstracted and corrected 40,000 OM fields across three years of deals. It fine-tunes the same base model on that labeled set, teaching it the vocabulary, layout quirks, and edge cases of its own document flow.
Assume that adaptation lifts per-field accuracy from 90% to 98%, a plausible gain when a model stops guessing at domain conventions and starts recognizing them.
Now the probability that all 12 fields are correct is 0.98 to the 12th power, which is about 0.78. Roughly 22% of documents need a human touch, down from 72%.
Same base model. The only variable that changed was the labeled data Firm B owned and Firm A did not. This is why fine-tuning CRE models beats generic LLM extraction: the ceiling on a general model is set by what it has never read, and proprietary labels raise that ceiling. The compounding is the point. Firm B's 40,000 labels become 60,000 next year, and the accuracy gap widens while Firm A stays flat.
When is a data moat empty rather than defensible?
A data moat is empty when the data is generic, abundant, or subject to diminishing returns. Andreessen Horowitz argued in "The Empty Promise of Data Moats" that raw data advantage often saturates: past a threshold, more of the same data stops improving the model, and a smaller high-quality set closes the gap. The warning is real, and it sharpens the thesis rather than refuting it.
The moat holds under specific conditions. The data must be proprietary, meaning no competitor can license the same feed. It must be labeled, meaning it carries verified field-level ground truth, not raw text. And it must come from a domain with no public corpus, so that scale cannot be bought. Commercial real estate documents meet all three. Executed leases, offering memoranda, and rent rolls are private, they require expert judgment to label, and no open dataset of corrected CRE abstractions exists. That is why the diminishing-returns critique applies to consumer clickstream data far more than to a corrected corpus of two-hundred-page leases.
Data quality is the multiplier that separates a real moat from a warehouse of noise. A million unlabeled documents teach a model nothing it can verify. Ten thousand corrected, enriched documents teach it exactly how a percentage rent clause or a co-tenancy provision reads in practice. This is also why the extraction product is the structured output, not the prose: the value compounds only when reads are captured as structured data an AI system can reuse and audit, not left as free text a human rereads each time.
Frequently Asked Questions
Is the CRE AI data moat about having more data than competitors?
No. It is about having proprietary, labeled data competitors cannot obtain. Volume alone suffers diminishing returns. A moat comes from verified field-level labels on private documents such as leases and offering memoranda, which no rival can license or scrape at any price.
If foundation models keep improving, will the data moat disappear?
Improving base models raise the floor for everyone equally, which erodes model-based advantage rather than data-based advantage. As the tenth-ranked model approaches the first, the differentiator shifts further toward proprietary data. Better commodity models make the data layer more decisive, not less.
Can synthetic data replace a proprietary CRE corpus?
Synthetic data can supplement training, but it cannot reproduce the real distribution of executed CRE documents: the actual amendments, side letters, and drafting quirks a portfolio contains. Ground truth for those documents comes from expert correction of real files, which is the asset a firm accumulates and a competitor cannot fabricate.
Does this mean building a custom model is required?
No. Most firms should license a commodity base model and invest in the data layer around it. The strategic work is capturing corrected outputs as structured, labeled data over time. The model can be rented. The labeled corpus is the thing worth owning.
Conclusion
The strategic error in CRE AI is treating the model as the asset. The model is a rented commodity, and the frontier is converging fast enough that the specific one a firm picks will not decide outcomes. What decides outcomes is the proprietary, labeled data a firm accumulates: the corrected leases, offering memoranda, and rent rolls that no competitor can license, scrape, or synthesize. The CRE AI data moat is the data, not the model. An operator's job is to run the commodity model well and, more importantly, to capture every corrected document as a labeled asset. The model depreciates the moment a better one ships. The data compounds every time a human fixes a field.