September 8, 2026 By Yodaplus
AI is measurably more reliable than standard industry-code screening at selecting truly comparable companies, but it still requires analyst validation before feeding a valuation, since wrong peer selection alone can swing an implied valuation by 50% or more regardless of how accurate the multiples applied to that peer set are. A 2023 study published in the Journal of Accounting Research found machine learning models using gradient boosting substantially outperformed traditional relative valuation models in out-of-sample accuracy, with the resulting valuations behaving like genuine fundamental values rather than noise-driven estimates.
That outperformance is real and well-documented in academic research. It is also not the same thing as saying AI gets peer selection right without oversight. Here is what the evidence actually shows.
Selecting the right comparable companies is widely regarded as the single most consequential step in relative valuation, more important than the sophistication of the multiple analysis applied afterward. No amount of careful EV/EBITDA or P/E calculation compensates for a peer group built on the wrong companies. This is precisely why the reliability of the selection method itself, not just the accuracy of the resulting multiples, deserves scrutiny.
The Journal of Accounting Research study tested gradient boosting machine models against traditional peer-based valuation approaches and found the machine learning models produced valuation multiples that behaved consistently with fundamental value: stocks the model flagged as overvalued tended to decline in price the following month, and undervalued stocks tended to rise. This out-of-sample performance persisted across different types of firms and held over time, rather than reflecting a one-off result specific to a single market period.
Separately, research applying named entity recognition and natural language processing to identify comparable companies has moved peer identification beyond keyword or classification matching toward genuinely understanding business descriptions, products, and services, closer to how an experienced analyst would reason about similarity.
The Global Industry Classification Standard remains the dominant system used by investment professionals, organizing companies into an eight-digit code across sectors, industry groups, industries, and sub-industries. It is a useful starting point, but academic research going back decades has shown industry grouping alone explains only a small share of stock return variance, and most classification systems, with GICS performing somewhat better than older alternatives, do not fully explain the cross-sectional share price movement within a given industry or sub-industry code.
AI-driven peer selection tools address this gap directly by screening on deeper operational characteristics, business model, growth profile, and margin structure, rather than defaulting to a classification code as the primary filter. This matters concretely: a company whose business model has shifted meaningfully since it was first assigned a classification code can retain a stale code for years, quietly distorting any comps analysis built purely on that classification.
The gradient boosting approach used in the Journal of Accounting Research study offers a concrete window into the mechanism. The model builds a series of decision trees that allocate firms into different groupings based on financial fundamentals, and firms frequently allocated to the same grouping can be treated as close peers, while those rarely grouped together are flagged as dissimilar, even if they share the same industry code. This produces a peer weight for every candidate company, giving analysts a quantified, scrutinizable measure of comparability rather than a binary include-or-exclude decision.
Separately, GPT-based and natural language processing approaches identify comparable companies by parsing business descriptions and named entities directly, capturing similarity in products, services, and market positioning that a purely numerical or classification-based model can miss entirely.
Despite this documented outperformance, AI-driven peer selection is not infallible, and researchers building these models have been explicit that the resulting peer weights are meant to be used to scrutinize comp choices, not replace that scrutiny entirely. Earlier research found that analysts and equity valuation experts weigh factors like size, asset turnover, industry classification, and trading volume when selecting peers, some of which are genuinely predictive and some of which carry less relevance than commonly assumed, a nuance a purely automated model can still misjudge without a human sense-check.
AI models are also trained on historical data and existing relationships between fundamentals and valuation multiples, which means a company undergoing a genuine structural shift, a new competitor entering from an adjacent business model, or a market entering unprecedented conditions can all fall outside what a trained model reliably captures.
Given this evidence, the most defensible current practice treats AI-generated comparability scores or peer weights as a strong, quantified starting point rather than a final answer. Analysts using tools that generate peer weights can specifically use those weights to question which companies belong in a set and which should be dropped, applying the same judgment that separates professional relative valuation from a mechanical exercise, just informed by a much richer signal than a classification code alone provides.
Overconfidence in a single quantitative score A high peer weight from a model reflects historical statistical similarity, not necessarily forward-looking business comparability, particularly for companies undergoing rapid change.
Training data limitations Models trained on historical fundamentals and pricing relationships can underperform when a target or its industry faces conditions genuinely outside that historical pattern.
Loss of qualitative context Numerical similarity models can miss qualitative factors, such as management quality or an emerging competitive threat, that a well-informed analyst would weigh heavily.
Verification burden shifting rather than disappearing Analysts still need to verify why a model flagged a company as comparable, which requires understanding the model’s logic well enough to trust or challenge its output, a genuinely new skill requirement.
Academic evidence increasingly supports machine learning-based peer selection as a genuine improvement over classification-only screening, and continued research combining gradient boosting approaches with natural language business description matching points toward peer selection tools that reason more like an experienced analyst over time. Full reliability without human validation remains a distant goal, though, since even the strongest published results treat model output as a tool for scrutinizing comp choices rather than a replacement for that scrutiny.
AI is genuinely more reliable than classification-code screening alone at selecting comparable companies, backed by peer-reviewed research showing real out-of-sample valuation accuracy gains. It is not yet reliable enough to skip analyst review, particularly for companies whose fundamentals or competitive position have shifted in ways a historically trained model may not fully capture.
Yodaplus helps equity research and financial services teams deploy this kind of AI-assisted peer selection responsibly. Our enterprise AI solutions combine multi-agent AI with intelligent document processing to surface comparability signals beyond standard classification codes, while keeping analyst review and audit trails built into every step, within a governance-first AI architecture designed for the accuracy standards equity valuation work demands.
Wrong peer selection alone can swing an implied valuation by 50% or more, making it more consequential than errors in the multiple calculations applied to a peer set, since no amount of precise multiple analysis compensates for the wrong companies.
Yes. A study published in the Journal of Accounting Research found gradient boosting machine learning models substantially outperformed traditional relative valuation approaches in out-of-sample accuracy, with results holding across firm types and over time.
Research dating back decades has shown industry classification alone explains only a small portion of stock return variance, and most systems do not fully capture the cross-sectional price movement within a given industry or sub-industry code, missing companies whose business models have shifted.
This is where AI reliability weakens most, since models trained on historical data and fundamentals-to-multiple relationships can underperform for companies facing conditions outside that historical training pattern.
No. Even the strongest published research treats AI-generated peer weights or comparability scores as a tool for scrutinising comp choices, not a replacement for analyst judgment on qualitative fit and business context.