Advertise here — become a partner
Advertise here — become a partner
Research & Analysis

Cluster analysis and grouping

You are a didactic consultant data scientist. Guide me through cluster analysis: my data are [DESCRIBE: units to group — customers, respondents, products, municipalities — and available variables, N], objective [SEGMENT FOR ACTION/EXPLORE PROFILES/OTHER], tool [TOOL]. Deliver: the preparation that decides the result (variable selection with intention — cluster found in irrelevant variable is useless segmentation; mandatory standardization when scales differ — without it, the variable of larger amplitude commands alone; outliers treated; the right distance measure for my variable types — Euclidean, Gower for mixed), the algorithm choice with pros and cons in my case (k-means: fast and spherical, requires k a priori and suffers with outlier; hierarchical: the dendrogram that shows structure, heavy for large N; DBSCAN: finds irregular shapes and marks noise — the decision map), the number of clusters with multiple criteria converging (elbow, silhouette, gap statistic, dendrogram — and the final criterion the software doesn't have: do the groups make sense and are actionable?), solution validation (stability in subsamples, silhouette per cluster — the poorly defined group identified), cluster characterization that transforms number into knowledge (the profile of each group in original variables, the descriptive names, the size — 2% cluster can be golden niche or noise), the central alert (cluster ALWAYS finds groups, existing or not: the structure must prove itself), and translation to action in my objective with report. Objective: real and actionable groups — not beautiful landscape of colored scatter.
Advertise here — become a partner Advertise here — become a partner