Predicting post-merger operating performance with interpretable ML
This thesis tests whether public pre-deal firm and transaction information can rank later operating outcomes without confusing fit for foresight. Five outcome definitions, chronological validation, and interpretable attribution keep the modest predictive signal in view alongside its limits.
Optimised for desktop. Complete on smaller screens.
Answer first
Pre-deal information carries a weak but real rank-ordering signal, concentrated in operational and financial channels rather than classic revenue-synergy stories.
Boundary. Predictive performance is weak and depends on the outcome measure: reported test-set R² ranges from 0.0221 for the Healy CFROA baseline to 0.1167 for the asset-turnover alternative.
Framing
One plus one should exceed two
M&A is sold on the claim that one plus one can exceed two through cost, revenue, asset-use, and financing synergies.
Framing
The observability problem
The integration behaviours that decide success happen after the deal and are not recorded in pre-deal information.
Public data can describe the firms before close. The integration choices that determine realized value happen later.
Observable before close
- Public annual reports
- Firm and transaction characteristics
- Pre-deal operating and financial structure
Unobservable before close
- Integration execution
- Management decisions after close
- Realized synergy
Framing
The bounded question
The thesis asks how much of a post-merger outcome is visible before close, and which part remains out of reach.
Can pre-deal firm and transaction information predict post-merger operating performance?
Predict
Test whether visible pre-deal information carries signal about later outcomes.
Do not infer
Explain why integration succeeds or fails, or identify a causal mechanism.
Prediction is not causal explanation. This is a signal detector, not a crystal ball.
46 sec
Project film
The thesis in motion
A short visual overview of the research question, model, and screening application before the full empirical story continues.
Film summary
The film asks whether pre-deal public information can provide an early signal of post-merger operating performance. It shows the financial and operational inputs, the machine-learning step, the expected direction of change, and a demonstrator for comparing potential synergy signals across deals.
Method
The public-data design
A public-report, target-family design uses 29 pre-deal features and chronological validation without future information.
The design follows one audit trail: define which deals can be measured, keep the outcome family intact, restrict the model to pre-deal evidence, then let later time periods grade it.
1995–2022
13.7%of the universe has every outcome-label ingredient
What the label requires
- Pre-deal cash flow and assets for both firms at year t
- Combined cash flow and assets for the merged entity at t+3
- SIC-year median benchmarks at t and t+3
The evidence covers large, listed, well-documented deals; it does not cover the full M&A universe or private SMEs.
5 readings of operating performance
Healy CFROA
Cash-flow efficiency
CFO / assetsAsset turnover
Asset productivity
revenue / assetsOperating ROA
Profitability
EBIT / assetsOperating margin
Margin discipline
EBIT / salesCAPEX intensity
Investment behaviour
CAPEX / assetsNo single target is a universal merger-success score. Realized synergy is not directly observable in the pre-deal data, but supervised machine learning needs a target to predict. These five later operating outcomes are measurable proxies for synergy, not synergy itself, and the family tests whether learnability changes with the proxy definition.
Is the family cherry-picking? The opposite. All five are estimated with the identical pipeline and reported together, strong and weak alike. Publishing the whole family is precisely what makes selective reporting impossible.
Supervised learning needs a number to predict. Getting that number wrong is how merger studies accidentally measure the industry cycle, or the accounting, instead of the deal.
Accounting outliers are clipped rather than dropped, so a handful of restatements cannot drive the fit.
Healy-style CFROA is the anchor target, not the only one. Four further operating outcomes are estimated with the identical pipeline.
Text equivalent
The Healy-style outcome is built in five steps: cash flow from operations over total assets; pooled across acquirer and target before the deal; adjusted by the SIC2 industry-year median; differenced between the deal year and three years later; and winsorised at the 1st and 99th percentiles.
Acquirer and target public annual reports at FY−1 only
The model tests whether visible pre-deal information carries signal; it does not claim to know how integration will go.
Hover a node for its formula · select a channel to open its list
Five channels, one economic mechanism each. The grouping is what makes the SHAP attribution in chapter 09 readable: in a 29-column model, correlated single-feature rankings are easy to over-read, so channel aggregation reduces brittleness without claiming stability across repeated fits.
Every feature is engineered from the last annual report filed before the deal was announced. Nothing dated after the announcement enters the model.
Text equivalent
The 29 pre-deal features are grouped into five synergy channels: operational (6), financial (9), cost (7), revenue (5), and macro (2). Grouping is what makes the later SHAP attribution readable, because single-feature stories in a 29-column model are brittle.
Knowledge cutoff · end 2018
Older deals train the model, intermediate deals tune it, and the 2019–2022 cohort grades it once. Every prediction uses only information available before the deal.
Results
How to read the results
Spearman ρ asks whether deals are ordered well; held-out R² asks whether outcome magnitudes are predicted accurately.
R² asks whether the predicted outcome magnitude is close to what happened.
Match every pair · negative R² is worse than using the meanSpearman ρ asks whether the model puts deals in a useful order.
The loop compares true order with model order · rank, not distanceText equivalent
R² measures how closely predicted outcome magnitudes match held-out outcomes; it can be negative when predictions are worse than using the average. The paired bars are illustrative rather than thesis observations. Spearman ρ measures whether higher predicted outcomes are ranked above lower ones, regardless of the exact distance between predictions. Its fixed visible domain is Spearman ρ 0.00 to 0.30.
Source: Published thesis and Public thesis defense presentation.
Text equivalent
Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.
Source: Published thesis and Public thesis defense presentation.
R² of 0.022 sounds like nothing. Is the model broken?
Not necessarily. Post-acquisition performance is unusually difficult to explain from pre-deal filings, but unlike-for-unlike studies cannot supply a numeric pass mark. The defensible reading comes from the held-out metrics, their uncertainty, and the rank-order evidence together.
Context, not a scoring ruler
- M&A performance literatureKing et al. (2004)
Four heavily studied pre-deal moderators explain near-zero variance across the meta-analytic evidence. This establishes difficulty, not a numeric benchmark band.
- Financial-ML precedentGu, Kelly & Xiu (2020)
Strict out-of-sample validation and nonlinear ensembles are method precedents. Equity-return prediction is a different target and is not used as a numerical comparator.
- Why other finance tasks can look easierAmini et al. (2021)
Capital structure has an equilibrium forcing mechanism; one-time integration outcomes do not. Its much higher R² therefore demonstrates a different problem structure, not a threshold this study should clear.
- This thesis · Healy CFROA baselineHeld-out 2019–20220.0220
R² is positive but its bootstrap interval includes zero. The load-bearing result is modest rank ordering: Spearman ρ = 0.161.
- This thesis · asset turnoverHeld-out 2019–20220.1167
The larger R² is target-specific. Removing operating-level features reduces it to 0.0254, making mean reversion part of the interpretation.
Low R² is contextualised, not celebrated. The model is a weak rank-order screen, never a point predictor.
Text equivalent
The literature explains why post-merger prediction is difficult and why strict out-of-sample machine-learning validation is appropriate; it does not provide numerical benchmark bands comparable across different targets. This thesis therefore reports its held-out values directly: the Healy CFROA baseline has R² 0.022 with a confidence interval that includes zero and Spearman rho 0.161, while asset turnover has R² 0.1167 before an ablation reduces it to 0.0254.
Results
What is learnable
Asset productivity is more learnable than the cash-flow anchor, showing that target definition changes the available signal.
Pre-deal levels describe where a firm starts.
Mean reversion can make later movement partly learnable without observing synergy.
Most of the signal is consistent with mean-reversion risk. Residual signal remains, so the claim shrinks rather than vanishes.
597 held-out deals · hollow markers show the levels-removed specification.
Hollow markers remove pre-deal operating levels.
Text equivalent
Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.
Source: Published thesis and Public thesis defense presentation.
Results
Profitability: order without calibration
Operating ROA and operating margin retain modest rank information while failing to predict the magnitude of profitability changes.
The model can place some stronger deals above weaker ones. It cannot reliably say how much profitability will change.
Operating ROA
Profit earned from the combined asset base
Operating margin
Profit retained from each unit of revenue
Text equivalent
Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.
Source: Published thesis and Public thesis defense presentation.
Useful for ranking and triage, not for forecasting the size of a profitability gain.
Text equivalent
Operating ROA has held-out R² −0.0550 and Spearman ρ 0.2390. Operating margin has held-out R² −0.0070 and Spearman ρ 0.1940. The model therefore retains modest ranking information for both outcomes while failing to predict their magnitudes better than the held-out mean.
Source: Published thesis and Public thesis defense presentation.
Diagnostic
CAPEX as a diagnostic
CAPEX intensity tracks investment behaviour but does not provide a standalone synergy interpretation.
Text equivalent
Healy CFROA: R² 0.0220, Spearman ρ 0.1610; Asset turnover: R² 0.1167, Spearman ρ 0.2761; Operating ROA: R² -0.0550, Spearman ρ 0.2390; Operating margin: R² -0.0070, Spearman ρ 0.1940; CAPEX intensity: R² -0.0290, Spearman ρ 0.1360.
Source: Published thesis and Public thesis defense presentation.
Efficiency · disciplineremoving duplicated assets
Growth investmentcapacity for the combined firm
Underinvestmentstarving the asset base
Overinvestmentempire-building
CAPEX direction is not a universal success label.
Text equivalent
CAPEX can fall through efficiency or discipline, or through underinvestment. It can rise through growth investment, or through overinvestment. CAPEX intensity, measured as CAPEX divided by assets, has held-out R² −0.029 and Spearman ρ 0.136. It is retained as a diagnostic of investment behaviour.
Source: Published thesis and Public thesis defense presentation.
Interpretability
What SHAP actually does
Every prediction splits into a baseline plus one exact contribution per feature, computable on trees but only ever about the model.
If a model predicts Netherlands 2–1 Argentina, SHAP asks how much each player contributed to those predicted goals across possible line-ups.
baseline 0.67+player credits 1.33=2 predicted goals
- Player
- Model feature
- Line-up
- Feature coalition
- Predicted goals
- Predicted post-merger outcome
SHAP explains the prediction, not who caused the real match. In the thesis it shows what the model relied on, never what caused the post-merger outcome.
Text equivalent
The model predicts Netherlands 2 and Argentina 1. For Netherlands, a baseline of 0.67 plus 1.33 of illustrative player credits equals 2 predicted goals. For Argentina, the same baseline plus 0.33 equals 1. Players stand for model features, possible line-ups stand for feature coalitions, and predicted goals stand for the predicted post-merger outcome. This allocates prediction credit and does not establish real-world causality.
f(x) = φ₀ + Σᵢ φᵢ
Every prediction decomposes exactly into a baseline plus one contribution per feature. The parts always sum back to the whole.
Illustrative decomposition · the arithmetic, not a thesis observation
The permutation definition averages a feature’s marginal contribution over every possible ordering. Enumerating those orderings directly is not a practical implementation at 29 features.
TreeSHAP uses tree-path dynamic programming across the trained ensemble to return exact φ values in polynomial time. It avoids brute-force permutation or coalition enumeration without pretending the computation is a single pass.
Exact, not approximate. But exact about the model, never about the world.
Feature relevance is a property of the model, not of the world.
- Janzing, Minorics & Blöbaum (2020)
Attribution is a causal problem; SHAP without a causal model does not deliver causal effects.
- Sundararajan & Najmi (2020)
Shapley variants disagree; conditional-expectation SHAP can violate the axioms it is sold on.
say “the model weights X”
never “X causes Y”
Method anchors: Shapley (1953) · Lundberg & Lee (2017) · Lundberg, Erion & Lee (2018)
Text equivalent
SHAP decomposes a single prediction into additive per-feature contributions that sum to the prediction. The permutation definition considers all 29-factorial conceptual feature orderings; TreeSHAP avoids direct enumeration through exact polynomial-time dynamic programming over the trained tree paths. The exactness is computational: SHAP describes model behaviour, not the economic data-generating process.
Interpretability
Inside the model
SHAP shows that the baseline model relies most on operational and financial structure, not a causal story.
Allocation analogy: give each channel the share of the prediction it contributes at the margin. That marginal contribution is a model reading, not a claim about cause.
Single-feature SHAP stories are brittle and easy to over-read: correlated ratios trade rank between runs. Summing absolute attribution into five economic channels is more stable and far harder to cherry-pick after the fact.
Channel shares are relative shares of mean absolute SHAP. Aggregation reduces brittleness; it does not convert attribution into causation.SHAP explains the prediction, never the causal mechanism in the world.
Text equivalent
For the baseline Healy model, operational features account for 34.2% of mean absolute SHAP attribution, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. Pre-deal accounting reads structure but cannot observe post-deal execution.
Source: Published thesis and Public thesis defense presentation.
Interpretability
How to read the SHAP figures
Spread, colour, slope, and scatter each answer a different question about what the model is leaning on.
Horizontal position is the SHAP push on a single deal. Colour is that deal’s own feature value.
- Wide spread
Moves predictions a lot, and variably. High but heterogeneous reliance.
- Tight at zero
Barely moves any prediction. The model effectively ignores it.
- Red one side, blue the other
Consistent direction: high values push one way, low values the other.
- Red and blue interleaved
Non-monotonic or interaction-driven. The push depends on something else.
x is the raw feature value, y is that feature’s SHAP push for that deal. Colour is a second feature SHAP picks automatically.
A clean marginal response. Target cash-flow margin slopes down: high values push the prediction down.
Interaction-heavy. At one x value the push splits top-to-bottom, and colour explains the split.
Colour splitting top-to-bottom at a single x value means interaction. Colour merely following the line means the two features are correlated, not interacting.
Spread and slope describe the model’s response surface, never a causal dose–response curve.
Text equivalent
In a SHAP beeswarm each row is a feature and horizontal spread is the size and variability of that feature’s push on the prediction, with colour showing the feature’s own value. In a dependence plot each dot is one deal, the slope shows the direction the model moves, and vertical scatter at a fixed feature value indicates interaction with a second feature.
Limits
Where the model gives up
The tails pull inward: shrinkage in a low-signal fit, with winsorising dampening how extreme the actuals even look.
The dots reconstruct the tail-compression mechanism from the reported held-out metrics; they are not the 78 held-out observations.
perfect calibrationillustrative compressed slope
Shrinkage toward the conditional mean
The ordering survives: the lowest prediction really is a bad outcome and the highest really is a good one. What compresses is the magnitude. With little signal in a high-variance target, the loss-minimising predictor pulls inward, because confident extreme calls would inflate squared error. That is exactly what ρ = 0.161 alongside R² = 0.022 looks like.
ρ 0.161the order survivesR² 0.022the magnitude does not
Winsorising at the 1st / 99th percentile
The target is capped before fitting, so the plotted “extremes” are already clipped values. Winsorising dampens how extreme the actuals look; it does not create the gap. Without it the underestimation would look larger, not smaller.
The claim is modest rank ordering, never calibrated magnitudes.
Text equivalent
Extreme outcomes are underestimated for two reasons. The primary cause is shrinkage: in a low-R² fit on a high-variance target, the error-minimising prediction is pulled toward the conditional mean. The secondary factor is that the target is winsorised at the 1st and 99th percentiles, so the plotted extremes are already capped; removing winsorisation would widen the gap rather than close it.
Robustness
Robustness checks
Distress, culture, and rolling-window checks add nuance while retaining a weak, bounded conclusion.
- Supports bounded conclusion
Main asset-turnover model
Does the conclusion survive this change?
A modest held-out signal makes pre-deal asset-productivity drift inspectable. - Weakens claim
Levels-removed ablation
Does the conclusion survive this change?
R² 0.1167 → 0.0254; much of the result is consistent with mean-reversion risk. - Boundary remains
Altman Z comparison
Does the conclusion survive this change?
One compact financial-health score: many problems, one needle. It is a diagnostic comparison, not a replacement conclusion. - Boundary remains
Pompe–Bilderbeek comparison
Does the conclusion survive this change?
R² 0.022 → 0.030; the richer distress panel improves diagnostics modestly. - Boundary remains
Culture and alternative-target checks
Does the conclusion survive this change?
Country scores are not firm culture; integration happens between firms, not flags. Alternative targets: Healy CFROA R² 0.0220, Spearman ρ 0.1610; Operating ROA R² -0.0550, Spearman ρ 0.2390; Operating margin R² -0.0070, Spearman ρ 0.1940; CAPEX intensity R² -0.0290, Spearman ρ 0.1360.
Text equivalent
The main asset-turnover model supports a bounded reading. Removing levels weakens that reading, while Altman Z, Pompe–Bilderbeek, culture, rolling-window, and alternative-target checks retain its boundary rather than establishing a stronger claim. Alternative-target evidence: Healy CFROA R² 0.0220, Spearman ρ 0.1610; Operating ROA R² -0.0550, Spearman ρ 0.2390; Operating margin R² -0.0070, Spearman ρ 0.1940; CAPEX intensity R² -0.0290, Spearman ρ 0.1360. The rolling-window rank correlations are 0.184 for 2013–2015, 0.082 for 2016–2018, and 0.161 for 2019–2022. The check supports partial persistence, not full temporal robustness.
Source: Published thesis and Public thesis defense presentation.
Learn from the past, score only on the next unseen years, then roll the window forward. Across 2013–2022 the rank signal stays positive but breathes with the deal cycle.
- train 1995–2012test 2013–2015ρ 0.184Positive
- train 1995–2015test 2016–2018ρ 0.082The dip
- train 1995–2018test 2019–2022ρ 0.161Recovers · COVID inside
Expanding-window forward test · chronological split · no look-ahead
This supports partial persistence, nothing more. Robustness across market regimes is not claimed, the final window contains COVID and the post-COVID deal surge, and this test rolls the CFROA baseline, so asset-turnover stability over time remains future work.
Text equivalent
Expanding-window forward tests give Spearman ρ of 0.184 for 2013–2015, 0.082 for 2016–2018, and 0.161 for 2019–2022 on the Healy CFROA baseline. The rank signal stays positive across all three windows but varies with the deal cycle, supporting partial persistence rather than full temporal robustness.
Economics
Why the model leans on cash
Free-cash-flow theory explains the financial channel, and predicts that the sign reverses outside listed firms.
- 01Cash-rich acquirer
Free cash flow beyond what the firm’s own project set can absorb.
- 02Agency cost
Managers with more money than good ideas face weaker discipline from capital markets.
- 03Overinvestment
Acquisitions become the place the surplus goes.
- 04Value-destroying deals
Cash reserves are empirically associated with worse acquisition outcomes.
- Jensen (1986)Free cash flow → agency cost → overinvestment.
- Harford (1999)Cash reserves → more value-destroying acquisitions.
- Altman (1968)Distress score as a compact financial-health proxy.
Listed firms: idle cash reads negative, as Jensen-style agency slack.
Family-owned SMEs: the same cash reads positive, as a Harford-style precautionary buffer that keeps integration alive.
Text equivalent
Free-cash-flow theory predicts that cash-rich acquirers over-invest, and the evidence associates cash reserves with value-destroying acquisitions. The model picks up this financial-health structure. In SME settings the sign may reverse, because the same cash balance can be a precautionary buffer rather than agency slack.
Limits
What the finding does not claim
After the target, attribution, and robustness checks, the surviving claim is predictive and bounded, not causal.
What is learnable
Learnable: some pre-deal asset-productivity drift.
The available signal changes with the target definition, and a small residual remains after levels are removed.
What it does not prove
Not proven: stronger synergy forecasting.
It does not observe integration execution, establish a causal mechanism, or turn asset turnover into a universal merger-success score.
Transfer
Why SME transfer is not copy-paste
The inputs can be computed for private-company records; the labelled outcome cannot, so the chapter stops at feasibility.
Computable inputs
Two of three answers are encouraging. The third decides the chapter: the inputs compute, and there is still nothing to train against.
94% of listed-firm gain attribution is computable
0% of it is validated · the gate above is still shut
- DDirectly transferable14 features · ~47% of gain
An ORBIS-equivalent field exists and the logic is unchanged: asset-turnover and ROA gaps, revenue size, SIC overlap, market factors.
- MTransferable with adaptation8 features · ~30% of gain
Valid, but the proxy is redefined: Altman Z′ substituting book equity for market cap, quick ratios, deal value.
- RProxy replaced5 features · ~17% of gain
The concept holds but no ORBIS proxy exists today: inventory turnover, R&D, CAPEX, intangibles, cash ratios.
- NNot transferable2 features · ~6% of gain
No SME analogue, so they are dropped: tender-offer and stock-payment flags, since SME deals are cash-financed.
- 01Acquirer cash flips signDirection can reverse
The listed-firm model reads cash as negative (Jensen overinvestment). For owner-managed SMEs expect it positive (Harford cash slack). The feature transfers; the sign does not.
- 02A constrained distress bundleAltman Z′ is a proxy
Private-firm Z′ substitutes book equity for market cap in X₄. Retained earnings are unavailable and total liabilities were not exported, so it is a bundle proxy rather than a reproduced score.
- 03EBITDA is not cash flowCFO is thin below €10m
Operating cash flow has low ORBIS coverage for small firms, so EBITDA margin substitutes, but EBITDA ignores working-capital movement. That makes it a CFROA proxy, not the Healy target.
Field-level computability under workable ORBIS conditions is not uniform availability, and construct-level portability is not validated predictive portability. Without a labelled SME outcome there is still no training target and no test set.
Text equivalent
For 3,110 revenue-reporting Western European SMEs, asset-turnover inputs are computable for 3,090 firms (99.4%) and EBIT/assets inputs for 2,991 firms (96.2%), while zero labelled post-merger SME outcomes are available. Grading all 29 features for ORBIS portability leaves roughly 94% of listed-firm gain attribution in the directly-transferable, adaptable, or proxy-replaced classes, but computability is not validation.
Transfer
The practitioner boundary
The practical output is a compass for inspecting pre-deal signals, not a validated SME prediction model or crystal ball.
Bounded pre-deal signalEvidence, not a verdict
Use for ranking and triage
Compare the visible pre-deal profile across a pipeline, then identify cases that merit deeper diligence.
Use for sensitivity analysis
Test whether a provisional view changes when observable assumptions, outcomes, or scenarios move.
Do not use as causal proof
A model reliance pattern does not establish why an operating outcome changed after a deal.
Do not use as an automatic deal decision
The score cannot see integration quality, management judgement, or the commercial case that remains to be tested.
What this thesis adds
- Methodological
A chronological, leak-free, target-family ML design for post-merger operating performance.
- Empirical
Predictability is target-dependent; asset-productivity drift is more learnable than CFROA-style synergy, with the mean-reversion boundary explicit.
- Practical
An interpretable, feasibility-gated advisory framework: inspectable pre-deal signals with explicit limits.
Evidence
What the claims stand on
Ten load-bearing claims, each tied to the papers that anchor it and the chapter that defends it.
The most-used anchors carry 2 of the 10 claims: Ghosh (2001), Healy, Palepu & Ruback (1992), King et al. (2004). 3 of 19 works support more than one claim; the rest carry one claim apiece.
These references support bounded claims. None of them makes the model causal, and none of them turns a feasibility map into a validated SME result.
Text equivalent
Ten load-bearing claims in the thesis are each tied to their anchor papers: target construction (Healy, Palepu & Ruback 1992; Ghosh 2001), target family (Healy, King, Devos), low predictability (King 2004; Gu, Kelly & Xiu 2020; Amini 2021), mean reversion (Ghosh 2001), SHAP method (Shapley 1953; Lundberg 2017, 2018), SHAP limits (Janzing 2020; Sundararajan & Najmi 2020), distress (Altman 1968; Pompe & Bilderbeek 2005), culture (Hofstede), the cash channel (Jensen 1986; Harford 1999), and SME feasibility (Arvanitis & Stucki; Golubov & Xiong).
Evidence
Evidence archive
The public presentation, repository, and published thesis preserve the evidence and its limits for inspection.
These are the inspectable artefacts behind the results, limitations, and practitioner boundary above. Captions link every visual back to its public evidence.
Dot plot comparing published out-of-sample R-squared and Spearman rank correlation across six post-merger outcome models. Public thesis extension summary.
Text equivalent
Test R² / Spearman ρ: Healy CFROA 0.0221 / 0.1609; Pompe 0.0296 / 0.1975; culture 0.0120 / 0.1318; Pompe plus culture 0.0304 / 0.2035; asset turnover 0.1167 / 0.2761; asset turnover without level features 0.0254 / 0.2327. Asset turnover is more predictable, but mean-reversion concerns bound the conclusion.
Chronological model validation timeline with training from 1995 to 2015, validation from 2016 to 2018, and held-out testing from 2019 to 2022. Public thesis defense deck.
Text equivalent
Chronological split: train on 1995–2015 deals, validate on 2016–2018 deals, and test once on the held-out 2019–2022 cohort. The knowledge cutoff is the end of 2018, so the model never trains on future deals.
Horizontal bars showing operational and financial features account for most mean absolute SHAP attribution in the baseline post-merger model. Public thesis defense deck.
Text equivalent
Baseline Healy-model shares of mean absolute SHAP attribution: operational 34.2%, financial 30.9%, cost 20.3%, revenue 9.7%, and macro 4.8%. SHAP describes what the model relied on, not a causal mechanism.