What it is grouping on
Every company in the release is placed by its 14 published readings — the nine analytical families, the four style scores and its size. Those are converted to van der Waerden rank-normal scores, reduced to 4 components by horn parallel analysis, 20 permutations (74% of the variance), and partitioned by k-means. The same release in always produces the same groups out: every random draw runs from a fixed seed.
It is not the study it reconstructs
This follows a research study that clustered the same universe on the model’s raw inputs. Those inputs are not published — which ones they are, and in what direction they act, is the part of the model that stays closed — so this works one level up, on the families the engine does publish. The structural findings hold: sector explains a small share of the variance, sub-industry more, the clusters more again, and the same sectors turn out homogeneous. The per-cluster cash-flow and margin figures of the original do not appear here, and no substitute for them has been invented.
Why 6 groups, when the data prefers 3
Resampling 70% of the universe 20 times and refitting, the partition into 6 groups reproduces itself at an adjusted Rand index of 0.67. A coarser split into 3 is steadier still. 6 is used because it is the resolution the original study established and its groups do correspond — but the steadier figure belongs on the page next to it, not in a footnote. Read the six as a map of similarity, not as a classification with hard edges.
Real Estate is excluded, not missing
6 of the 949 companies in the release sit in Real Estate. Too few to form a group of their own, they would be either stranded as a cluster of their own or absorbed into one they have nothing in common with, so they are left out entirely — the same exclusion the original study made. 943 companies are grouped.
Why the cluster share is not the headline
The clusters account for 57.5% of the variance against 12.0% for the 9 GICS sectors, but the clusters were fitted to maximise exactly that quantity and the sectors were not — on its own the comparison proves nothing. The 59 sub-industries are the real test: a hand-built taxonomy with 10 times as many groups as there are clusters, reaching 27.0%. That gap is the finding.
Names are generated, not chosen
Each group is named from its own two strongest analytical families. Nobody writes these labels, which is why they read plainly — and why they stay correct after a release moves the boundaries. Group numbers run largest-first and are stable within a release, but not across releases: C3 next month need not be C3 today.