Skip to content

Cluster analysis · ALB Equity

Two companies in the same sector can have nothing in common.

Sector limits are how most mandates describe risk. Grouped instead by what the model actually measures, 943 companies fall into 6 profiles that cut straight across the 9 GICS sectors — and the sector label turns out to account for only 12% of what makes these companies alike.

943
Companies grouped
6
Fundamental profiles
12%
Explained by GICS sector
4.8×
More explained by profile

Release 2026.06 · 30 June 2026 · recomputed from the model output on every release, not maintained by hand.

The finding

How much does a classification actually tell you?

Each bar is the share of variation in the model’s own readings that a grouping accounts for. The comparison only means something with all three present — see the note beneath it.

GICS sector · 9 groups12.0%

The classification most mandates and risk reports are written in.

GICS sub-industry · 59 groups27.0%

Far more groups, hand-built — the fairest test of classification alone.

Fundamental clusters · 6 groups57.5%

Fitted to maximise exactly this. Shown for scale, not as a fair contest.

Why this matters to a mandate

Sector limits are the instrument most portfolios use to control concentration. If the sector label captures a modest share of what makes companies behave alike, a book can satisfy every sector limit it has and still be a single bet — long cash generation and short momentum, say, with nothing in the sector report to show it.

The point is not that sectors should be dropped. They are how mandates, benchmarks and reporting are written, and that will not change. It is that a second lens is needed alongside them, and that the second lens should come from the data rather than from another committee.

The profiles

6 groups, each defined by what it is strong and weak at.

Named from their own two strongest analytical families — generated from the data, not written by an analyst, so they stay accurate when the next release moves the boundaries.

C1206 · 21.9%

Weak expectations & implied growth, strong relative performance

Strongest
Relative performance, Market indicators
Weakest
Expectations & implied growth, Quality
Median size
$18.4bn
Median engine rank
8.01of 10, 0 best

Most represented

  • Utilities29% of cluster(84% of sector)
  • Industrials18% of cluster(16% of sector)
  • Materials14% of cluster(36% of sector)
C2182 · 19.3%

Strong free cash flow yield, weak relative performance

Strongest
Free cash flow yield, Traditional valuation
Weakest
Relative performance, Market indicators
Median size
$14.5bn
Median engine rank
5.08of 10, 0 best

Most represented

  • Industrials25% of cluster(19% of sector)
  • Consumer Discretionary20% of cluster(33% of sector)
  • Consumer Staples15% of cluster(37% of sector)
C3173 · 18.4%

Strong quality, weak relative performance

Strongest
Quality, Growth
Weakest
Relative performance, Market indicators
Median size
$34.1bn
Median engine rank
2.36of 10, 0 best

Most represented

  • Industrials31% of cluster(22% of sector)
  • Health Care23% of cluster(33% of sector)
  • Information Technology22% of cluster(26% of sector)
C4141 · 14.9%

Weak traditional valuation, strong relative performance

Strongest
Relative performance, Leverage & management behaviour
Weakest
Traditional valuation, Free cash flow yield
Median size
$41.0bn
Median engine rank
2.96of 10, 0 best

Most represented

  • Industrials37% of cluster(22% of sector)
  • Information Technology34% of cluster(33% of sector)
  • Materials9% of cluster(16% of sector)
C5130 · 13.8%

Strong free cash flow yield, weak growth

Strongest
Free cash flow yield, Traditional valuation
Weakest
Growth, Leverage & management behaviour
Median size
$35.4bn
Median engine rank
2.64of 10, 0 best

Most represented

  • Energy22% of cluster(48% of sector)
  • Industrials18% of cluster(10% of sector)
  • Consumer Staples15% of cluster(25% of sector)
C6111 · 11.8%

Weak free cash flow yield, strong growth

Strongest
Growth, Leverage & management behaviour
Weakest
Free cash flow yield, Expectations & implied growth
Median size
$16.5bn
Median engine rank
8.08of 10, 0 best

Most represented

  • Information Technology25% of cluster(19% of sector)
  • Industrials22% of cluster(10% of sector)
  • Consumer Discretionary16% of cluster(16% of sector)

Profile detail

The same groups, across all nine analytical families.

Analytical familyC1n=206C2n=182C3n=173C4n=141C5n=130C6n=111
Growth
-0.69
-0.27
+0.84
+0.45
-0.93
+0.92
Free cash flow yield
-0.45
+0.96
+0.26
-0.75
+0.96
-1.32
Expectations & implied growth
-0.81
+0.68
+0.72
-0.21
+0.57
-1.12
Traditional valuation
+0.13
+0.90
-0.32
-1.15
+0.75
-0.64
Margins & profitability
-0.12
-0.32
+0.67
+0.27
+0.06
-0.71
Leverage & management behaviour
-0.74
-0.25
+0.67
+0.76
-0.25
+0.07
Market indicators
+0.16
-0.42
-0.38
+0.58
+0.36
-0.20
Quality
-0.78
-0.09
+0.88
+0.58
+0.45
-1.04
Relative performance
+0.36
-0.81
-0.67
+0.93
+0.49
-0.05

Standard deviations from the universe mean, on the model’s own analytical families. Positive is the favourable direction throughout — the engine scores 0 as best, and these readings are inverted once, here, so a green cell always means strength. These are positions within a ranking, not accounting quantities: there is no cash-flow yield or margin figure on this page, because those are inputs the model does not publish.

Sectors

Some sectors are one thing. Most are several.

Utilities is genuinely homogeneous — 84% of it sits in a single profile. Industrials spreads across every one, its largest group holding only 22%. For a sector like that, the label says where a company is listed in a taxonomy, not how its business works.

C1 n=206C2 n=182C3 n=173C4 n=141C5 n=130C6 n=111
  • Utilities70 companies · largest cluster 84%
  • Energy58 companies · largest cluster 48%
  • Consumer Staples75 companies · largest cluster 37%
  • Information Technology144 companies · largest cluster 33%
  • Materials77 companies · largest cluster 36%
  • Communication Services49 companies · largest cluster 31%
  • Consumer Discretionary111 companies · largest cluster 33%
  • Health Care123 companies · largest cluster 33%
  • Industrials236 companies · largest cluster 22%

Method

How it is built, and where it is weak.

What it is grouping on

Every company in the release is placed by its 14 published readings — the nine analytical families, the four style scores and its size. Those are converted to van der Waerden rank-normal scores, reduced to 4 components by horn parallel analysis, 20 permutations (74% of the variance), and partitioned by k-means. The same release in always produces the same groups out: every random draw runs from a fixed seed.

It is not the study it reconstructs

This follows a research study that clustered the same universe on the model’s raw inputs. Those inputs are not published — which ones they are, and in what direction they act, is the part of the model that stays closed — so this works one level up, on the families the engine does publish. The structural findings hold: sector explains a small share of the variance, sub-industry more, the clusters more again, and the same sectors turn out homogeneous. The per-cluster cash-flow and margin figures of the original do not appear here, and no substitute for them has been invented.

Why 6 groups, when the data prefers 3

Resampling 70% of the universe 20 times and refitting, the partition into 6 groups reproduces itself at an adjusted Rand index of 0.67. A coarser split into 3 is steadier still. 6 is used because it is the resolution the original study established and its groups do correspond — but the steadier figure belongs on the page next to it, not in a footnote. Read the six as a map of similarity, not as a classification with hard edges.

Real Estate is excluded, not missing

6 of the 949 companies in the release sit in Real Estate. Too few to form a group of their own, they would be either stranded as a cluster of their own or absorbed into one they have nothing in common with, so they are left out entirely — the same exclusion the original study made. 943 companies are grouped.

Why the cluster share is not the headline

The clusters account for 57.5% of the variance against 12.0% for the 9 GICS sectors, but the clusters were fitted to maximise exactly that quantity and the sectors were not — on its own the comparison proves nothing. The 59 sub-industries are the real test: a hand-built taxonomy with 10 times as many groups as there are clusters, reaching 27.0%. That gap is the finding.

Names are generated, not chosen

Each group is named from its own two strongest analytical families. Nobody writes these labels, which is why they read plainly — and why they stay correct after a release moves the boundaries. Group numbers run largest-first and are stable within a release, but not across releases: C3 next month need not be C3 today.

Which profile is each company in?

This page shows the structure. Inside ALB Equity, every one of the 943 companies carries its group, its distance from the centre of it, and the strongest names within each profile by the engine’s own ranking.