PropTech Report

AI Valuation Models versus Traditional Appraisals

AVMs excel with volume data but struggle where appraisers shine most.

Staff Writer · · 11 min read · Updated
Cover illustration for “AI Valuation Models versus Traditional Appraisals”
AI in Real Estate · August 11, 2026 · 11 min read · 2,528 words

AVMs are data-hungry machines, and they perform exactly as well as their inputs allow. In dense urban and suburban residential markets with frequent sales turnover, standardized property types, and robust MLS data sharing, they operate near peak reliability. Tract homes in high-volume suburbs, condos in large uniform buildings, multifamily assets with consistent unit layouts: these give a model the training volume it needs to produce a tight estimate. That is the environment AVMs were built for.

The picture changes fast in thin markets. Rural and exurban areas have sparse transaction histories and long physical distances between comparables. Custom and historic homes present renovation and architectural combinations no model has encountered before. Commercial and specialty property, office buildings, hospitality assets, land, sits largely outside the data environments that residential AVMs were designed to exploit. When inputs are sparse and idiosyncratic, confidence intervals widen even when the model is still willing to produce a number. The model does not always tell you when it is guessing.

This is where traditional appraisal holds a structural advantage that does not get enough credit. An experienced appraiser can work with two or three comparables, apply reasoned judgment about condition and local market dynamics, and still produce a defensible opinion of value. That skill does not require volume. It is precisely the advantage that matters where AVMs struggle most.

Traditional appraisal is not immune to data problems, though. Comp selection is subjective. In thin markets, the adjustments applied to the few available sales require assumptions that are genuinely hard to validate, and appraiser quality varies more than the profession typically acknowledges. The real reliability question is never which method is more accurate in the abstract. It is which method has better inputs for this specific property, in this specific market, at this specific moment.

What accuracy figures actually measure (and what they obscure)

The headline accuracy figures for modern AVMs are genuinely impressive. Median error rates for standard residential properties have fallen sharply over the past decade, and an FHFA 2024 study found that top-tier cascade AVMs came within ten percent of subsequent appraisal value on more than 87 percent of transactions. For anyone who worked with early-generation AVM outputs, that is a real and substantial improvement. The technology got better. That is not in dispute.

What those figures actually measure, though, is agreement with transaction prices or subsequent appraisals, not agreement with some observable "true" value. Transaction prices embed buyer urgency, financing constraints, and negotiation dynamics, none of which are intrinsic property characteristics. A model that closely predicts sale price is not necessarily capturing underlying value. It is capturing the specific market conditions under which that sale occurred.

Commercial real estate shows how wide the spread can be. In data-dense multifamily environments, AVM accuracy reaches the mid-to-high nineties. Office properties, particularly older suburban assets, land considerably lower, with meaningful variance around the average. A 2025 peer-reviewed study in Advances in Consumer Research benchmarked Zestimate valuations against New York City Department of Finance assessments across 294 properties and found a median absolute percentage error of 17.5 percent, a considerably wider band than industry benchmarks suggest for well-served markets. The gap between tight residential accuracy and that number is not a contradiction. It reflects market type, model architecture, and what the benchmark is measuring in the first place.

Here is the part that rarely gets acknowledged on either side of this debate: traditional appraisals are almost never benchmarked in the same systematic way. Their error rate is largely assumed rather than empirically tracked. That assumption has served the profession for decades, but it is increasingly a problem for lenders and regulators who need defensible accuracy standards on both sides of the comparison, not just one.

How the appraisal workforce shortage is accelerating AVM adoption

The appraiser workforce has declined more than twenty percent from its mid-2000s peak, per the Appraisal Institute. The Bureau of Labor Statistics counted tens of thousands of actively employed appraisers as of May 2024, and a significant share of those holding credentials are no longer actively practicing.

The demographic profile of the remaining workforce is the deeper problem. The profession is aging fast, retirements are accelerating, and new entrants are not arriving in sufficient numbers to close the gap. Part of the reason is structural: the Appraisal Foundation has reported thousands of aspiring appraisers unable to find the supervisory relationships required to complete their licensing. The apprenticeship model that governs entry into the profession is itself a supply constraint, and it does not respond quickly to demand signals. You cannot solve a workforce shortage with a credentialing system that bottlenecks on mentor availability.

The geographic consequences show up at the transaction level. In rural and underserved markets, scheduling a traditional appraisal can add two to three weeks to a closing timeline. Fees have risen with scarcity; complex or remote properties routinely command fees north of seven hundred dollars. In rate-sensitive markets, delayed closings carry real financial consequences for buyers locked into rate commitments and sellers managing contingent transactions on both sides.

The point is that AVM adoption is not purely a preference for speed. For many lenders, particularly those serving markets where appraiser availability is genuinely constrained, it is partly a structural necessity. The workforce dynamics are not reversing. Every year that passes without a meaningful change to how new appraisers enter the profession is another year the industry tilts further toward automated alternatives by default.

How regulators have responded (and what they now permit)

Diagram: The GSE Three-Tier Valuation Framework. Visualizes: Show the three distinct tiers now formalized in the GSE appraisal framework as a stepped or layered diagram, moving from full appraisal to hybrid appraisal to appraisal waiver (Value…

COVID-19 was the inflection point. Fannie Mae and Freddie Mac introduced emergency desktop appraisal and waiver flexibilities in 2020 out of operational necessity. What proved effective at reducing cost and timeline without materially increasing risk did not get rolled back when the emergency ended. It became permanent, then it expanded.

Fannie Mae's Q1 2025 eligibility changes extended Value Acceptance, the formal name for appraisal waivers, from 80 percent to 90 percent loan-to-value for purchase loans on primary residences and second homes. Value Acceptance with Property Data was expanded to program limits. As of October 2024, 97.9 percent of GSE direct sellers had delivered loans with appraisal waivers. Fannie Mae estimates appraisal alternatives have saved mortgage borrowers more than $2.5 billion since early 2020, per the company's published reporting.

The hybrid appraisal was formalized as a third tier in early 2025, with both Desktop Underwriter and Loan Product Advisor expanding eligibility. A property data collector conducts the physical inspection; a licensed appraiser reviews the data and issues an opinion remotely. The GSE framework now has three distinct tiers: full appraisal, hybrid appraisal, and appraisal waiver, each mapped to loan type, property type, LTV, and risk profile.

On the quality control side, the FHFA's proposed AVM rule pushes vendors toward documented accuracy standards and explicit bias mitigation programs. In July 2024, six federal agencies jointly issued guidance targeting deficient valuations, specifically flagging errors, omissions, and discriminatory conclusions as areas requiring lender response. In Europe, the revised CRR3 criteria for advanced statistical models signal that regulatory acceptance is conditional and technically specific, not blanket.

The trajectory across jurisdictions is not replacement of appraisals by AVMs. It is stratification, with method calibrated to risk profile rather than assigned by default. The regulatory structure is catching up to where lender practice already was.

The bias question (and why neither method is a clean solution)

The traditional appraisal's bias problem is long-documented and not seriously contested in the research literature. Lower valuations in majority-Black and majority-Hispanic neighborhoods, independent of property characteristics, trace a pattern connecting to historical redlining through comp selection and neighborhood adjustment methodology. The 2024 interagency guidance from six federal agencies was a direct regulatory acknowledgment that this problem is real, systemic, and requiring active intervention, not future study.

AVM proponents make a legitimate argument in response. Remove human judgment, remove one vector for conscious or unconscious discrimination. A model does not know the race of the buyer or the seller. It is not applying a neighborhood adjustment based on demographic intuition.

The counterargument is equally legitimate, and it does not get dismissed easily. The model does know zip code, census tract, school district, and flood zone, all of which correlate with demographic composition and with the historical pricing patterns that emerged from discriminatory practices. A model trained on historical transaction prices inherits the discrimination embedded in those prices. Accurately predicting a biased market does not correct the bias; it perpetuates it at scale, with the added credibility of algorithmic authority. That is the part that makes this problem harder than it looks. An algorithm replicating historical discrimination is often invisible to the people relying on its outputs.

The FHFA's AVM quality control rule names bias mitigation as a specific regulatory requirement. The International Valuation Standards Council and the Royal Institution of Chartered Surveyors have both emphasized active identification of bias and transparency about model inputs in their 2024 and 2025 standards updates. Neither body has declared the problem solved. Both have declared it an area of active obligation, not an aspirational goal.

The practical question is whether model bias is more auditable and correctable than human bias. An algorithm's inputs can be examined in ways that a human appraiser's reasoning often cannot. But auditability is only useful if someone is doing the auditing, and the regulatory framework is still defining what that standard actually requires in practice.

A decision framework. Which method fits which situation

Table: Which Valuation Method Fits Which Situation. Compares Best Property Type, Market Conditions, Transaction Context, Key Advantage, and 2 more by AVM, Hybrid Appraisal and Traditional Appraisal.

The choice is not philosophical. It is situational, and the variables are identifiable.

When an AVM is the right tool

AVMs perform best on standard residential properties in data-dense markets with frequent sales turnover. Single-family homes and condos in high-volume suburbs, where the model has processed thousands of comparable transactions, are its native environment. For transactions within GSE waiver eligibility thresholds, where LTV is relatively low and risk exposure is contained, the speed and cost advantage of an AVM is real and the accuracy trade-off is modest.

Refinances, portfolio monitoring, and rapid-close competitive purchase scenarios all favor AVM use. Lenders processing large transaction volumes face a practical reality: staffing a traditional appraisal for every loan in the pipeline is not operationally viable, and for standard-property, low-LTV transactions, it is not warranted by the risk profile. That said, concentration in a single model creates correlated risk across a portfolio. Cascade architectures and model diversity matter for exactly this reason.

When traditional appraisal remains irreplaceable

Unique, custom, historic, or significantly renovated properties are poor AVM candidates. The model lacks adequate training data, and the confidence interval widens to the point of limited utility. Rural and exurban markets present the same problem from a different direction: statistical inference requires volume, and the volume is not there.

High-value or high-LTV transactions, where the cost of a valuation error substantially exceeds the cost of an appraisal, warrant the more thorough method. Legal, estate, and dispute contexts require a signed opinion of value from a licensed professional; a model output does not carry the same legal standing. Commercial and specialty property, where income approach or cost approach analysis is central to the valuation, still requires the kind of judgment that no current model reliably replicates.

The hybrid as a practical middle path

The hybrid appraisal, now formally recognized within the GSE framework, captures meaningful portions of both methods' advantages. A property data collector handles the physical inspection. A licensed appraiser reviews the collected data and issues a remote opinion. The physical observation that a pure AVM cannot perform gets done, while some of the time and cost premium of a full appraisal is recovered. For transactions sitting at the margin between full appraisal and waiver eligibility, the hybrid is increasingly the most sensible default.

Framing it by stakeholder

Venn diagram: AVMs vs. Traditional Appraisals. Compares AVMs and Traditional Appraisal; overlap: Shared Concerns.

For buyers: understanding that an appraisal waiver is not a no-valuation decision matters. The lender is relying on a model. In markets where homes are selling above list price, a model trained on recent transaction data will lag actual current conditions. Buyers accepting waivers in fast-moving markets should care about the model's confidence interval, not just its point estimate.

For investors: portfolio monitoring and acquisition screening are legitimate high-volume AVM use cases. Individual asset underwriting for complex properties still warrants a human opinion, particularly in post-pandemic office and retail sectors where model training data predates structural demand shifts that have not yet fully resolved.

Risk tolerance is the deciding variable throughout. AVMs trade speed and cost savings for a wider confidence interval. Traditional appraisals trade time and money for tighter certainty. Hybrid methods sit between. That is precisely why they are growing.

The direction the industry is moving and what remains unsettled

The AI property valuation market was valued at $2.85 billion in 2025 and is projected to reach $11.60 billion by 2034 at a compound annual growth rate of 16.9 percent. That differential relative to the traditional appraisal services market is structural, not cyclical. Appraiser workforce demographics, GSE policy direction, and lender digitization are all pointing the same direction, and none of those forces reverse on a short timeline.

Three problems remain technically unsolved, and they are not minor.

Interior condition is unobservable without physical inspection. No amount of satellite imagery, street-level photography, or permit history tells you what is behind the finished basement wall. That is a fundamental epistemic limit, not an engineering problem awaiting a better algorithm. The model has never been in the building.

Market discontinuities invalidate training data faster than models can retrain. The pandemic-era demand shifts, the rate spike of 2022 and 2023, local economic shocks: all of these created conditions where historical transaction data was not a reliable guide to current value. Models learn from history. They are structurally exposed when history stops predicting the present.

Bias embedded in historical transaction prices cannot be corrected by a model that learns from those prices without explicit intervention. Naming it as a regulatory requirement is necessary. It is not the same as solving it.

The FHFA quality control rule, when finalized, will raise the floor on AVM standards and consolidate the market around vendors with documented accuracy and bias mitigation programs. The GSE hybrid appraisal expansion is the most significant near-term structural change; it creates a credible third tier likely to absorb a substantial share of transactions currently sitting at the margin between full appraisal and full waiver.

The likeliest outcome across the next decade is not replacement but stratification. AVMs dominate high-volume, standard-property, low-LTV transactions. Traditional appraisals concentrate in complex, high-stakes, legally consequential situations. Hybrid methods expand the middle. Meanwhile, the skills involved in each method are quietly converging: appraisers increasingly use AVM outputs as a starting point; AVM vendors increasingly embed appraiser-derived adjustment logic into their models. The boundary between the two is becoming more permeable every year.

The most useful question was never which method wins. It is which method has the right inputs for this property, this transaction, and this level of acceptable risk. Ask which model, what confidence score, what the last comparable transaction date was in this market. The practitioners who ask those specific questions make better decisions than those treating the whole debate as a generational preference. That is not a prediction. That is just what I have watched happen.

Sources

  1. workingre.com
  2. bankingjournal.aba.com
  3. rmahq.org
  4. appraisalbuzz.com
  5. bisnow.com
  6. appraisersblogs.com

More in AI in Real Estate