Roughly 150 to 200 U.S. credit unions disappear into mergers every year, out of a population that has fallen from over 10,000 in the 1990s to around 4,300 today. If you sell to credit unions, advise them, or buy them, knowing which ones are on that path twelve months early is worth real money — a vendor does not want to implement a core system at an institution that will be absorbed next spring, and an acquirer wants to open the conversation before three other acquirers do.
So the question is not whether merger prediction is useful. It is whether any given model actually predicts, or merely describes.
What a merging credit union looks like before it merges
Consolidation in this industry is overwhelmingly a scale story, and it shows up on the balance sheet long before it shows up in a board vote. The pattern that recurs:
- Sub-scale assets. Fixed cost — compliance, cybersecurity, a digital banking stack members now expect — does not shrink with the institution. Below roughly $50M in assets, that cost is spread across too few members to carry.
- Membership going backwards. A shrinking member count is the clearest single predictor, because it removes the one thing that could eventually fix the scale problem.
- Earnings that cannot fund reinvestment. A return on assets near or below zero means the institution is not generating the capital to buy its way out.
- Capital thinning toward the regulatory floor. The 7% net worth ratio that separates "well capitalized" from everything below it is a real cliff, and boards behave differently as they approach it.
- An aging charter with an aging field of membership. A sixty-year-old single-employer charter whose sponsor has shrunk is a specific and common story.
- Credit deteriorating faster than the institution can absorb. Charge-offs climbing on a thin capital base compress the timeline.
None of these is decisive alone. A small credit union with strong earnings and growing membership is not a merger candidate; a large one shedding members with weak earnings may well be. The signal is in the combination, which is what a scorecard is for.
Scorecard, then calibration
There are two separable jobs here, and models that conflate them are hard to trust.
The scorecard takes each of those components, scales it to a 0–1 range against sensible bounds, and blends them at published weights into a 0–100 susceptibility score. Its virtue is that it is legible: every score decomposes back into weight × feature, so you can see that an institution scored 74 because of membership decline and charter age rather than because of a black box. A scorecard is an argument you can audit.
The calibration is a fitted model that converts the scorecard's inputs into an actual one-year probability, trained against confirmed historical merger events. Its virtue is that a probability means something a score does not: "18%" is a claim you can be wrong about, and being wrong about it is measurable.
You want both. A scorecard alone tells you the ranking but not the odds. A fitted probability alone gives you a number with no explanation attached, which is unusable in a conversation with a board.
The only validation that counts is out-of-time
Here is where most merger lists fall down, and the question worth asking any vendor.
A model trained and tested on the same period will look excellent. It has seen which institutions merged and can find the features that separated them — including features that were consequences of the merger process rather than causes of it. That is not prediction. That is description with extra steps.
Out-of-time validation trains on one period and tests on a later one the model has never seen, scoring institutions as of a date and then checking what actually happened in the following year. It is a harder test and it always produces worse-looking numbers, which is exactly why it is the one to insist on.
Three numbers to ask for:
| Metric | What it answers | What "good" looks like here |
|---|---|---|
| AUC, out-of-time | Given a merged and an unmerged institution, how often does the model rank the merged one higher? | 0.75+. Below 0.65 is barely better than sorting by asset size; above 0.90, ask what leaked. |
| Top-decile capture | What share of actual mergers land in the model's riskiest 10%? | 40%+. This is the number that translates to your call list. |
| Lift over random | How much better than a random list of the same length? | 4×+. |
Two more, without which the first three mean very little.
How many actual mergers were in the test window? An AUC of 0.90 measured on eleven events is a statement about eleven events. Mergers are rare, held-out windows are short, and a vendor who will quote you three decimal places but not a count is telling you which of the two they would rather you looked at.
And what counts as a positive? A model that predicts any merger is predicting a different, easier thing than one that predicts a distress merger, because most mergers are strategic combinations between two healthy charters. Widening the label roughly quadruples the positive class and flatters every metric in the table. Ask what the label was before you compare two AUCs.
For reference, the merger model behind this site publishes 0.86 AUC out-of-time against a distress label — mergers a regulator attributes to the credit union's own condition — with the top scored decile capturing roughly 56% of them, about 5.6× lift. That rests on 89 events in the held-out window, which puts the standard error on the AUC near 0.04 — and that estimate is itself optimistic, because the fit pools consecutive as-of quarters and rows across cohorts are not independent. Read it as mid-0.8s. It is not a number to defend to the third decimal, and we do not. The held-out window also contained no distress merger above $500M in assets, so above that threshold the model is untested rather than weak. Those are backtest results on historical data, not a guarantee about any future period, and the known limitations are published rather than buried.
An AUC in the mid-0.8s is genuinely useful and genuinely not clairvoyant. It means the ranking is strongly informative and any individual prediction is uncertain. Anyone quoting you a 0.95 on a problem like this has either leaked future information into the training set or is measuring something other than what you need.
Susceptibility is not the same as value
A second, subtler mistake: treating "likely to merge" as "worth acquiring."
They are close to opposites. The institutions scoring highest on susceptibility are small, shrinking and thinly capitalized — which is precisely why they merge, and precisely why acquiring one may absorb capital without buying much. What an acquirer actually wants is a franchise: a low-cost core deposit base, a real member relationship, clean credit, enough scale to matter.
So susceptibility and franchise value have to be scored separately, and the interesting set is the intersection — institutions that are plausibly available and worth having. That set is always much smaller than either list alone, and it is the only one worth a banker's time.
Four questions to ask before you buy a merger list
- Is the validation out-of-time, and what is the AUC? If the answer is a single accuracy percentage with no time split described, the model has not been tested.
- Does every score decompose? If you cannot see which factors produced a given institution's number, you cannot defend it in a meeting, and you cannot catch it when it is wrong.
- Is susceptibility scored separately from franchise value? One list conflating the two is a list of small institutions.
- How often does it refresh, and what does it tell you changed? A static ranking is worked through in a quarter. The recurring value is in transitions — who crossed a band this quarter and why — not in the level.
What this does not tell you
A merger is a board decision made by people, driven by succession, sponsor relationships, regulatory conversations and personalities that no balance sheet records. A model reading quarterly financials is reading the pressure, not the decision. It will rank the pressure well and miss the institution whose CEO simply decided to retire.
Used as a ranking that tells you where to spend attention first, that is fine. Used as a prediction about a specific named institution, it is not — and no score on this site, or anywhere else, is investment, credit or merger advice. See the disclaimer and disclosures.