> For the complete documentation index, see [llms.txt](https://documentation.connexica.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.connexica.com/reporting/clustering.md).

# 12. Clustering

## 03fiii. Clustering

With a set of data points, clustering algorithms can be used to classify each data point into a specific group.

For unlabeled data sets, clustering is used to determine the possible groups that lie within the records. In theory, data points that are in the same group should have similar properties or features, while data points in different groups should have highly dissimilar properties or features.

{% stepper %}
{% step %}

### Prepare Data

Using the checkboxes for each column, select the fields that will be included in the model output.

Use this step to omit fields that are not required for the reporting output and to reduce processing time.

Click **Save Preparation** to proceed to the next step.
{% endstep %}

{% step %}

### Create New Model

With the data set that will be used for the model now selected, the appropriate **Model Type** can now be selected to scan and cluster the data. Depending on the Model Type, these clusters will be generated using a different algorithm.

The goal of the **K-Means** algorithm is to discover groups in the data, with the number of groups represented by the variable ‘k’.

Specify the ‘k’ value using the **Number of Clusters** textbox. The algorithm then works iteratively to assign each data point to one of the ‘k’ groups based on the features in the data set, with each data point clustered based on similarity.

The **K-Means** algorithm is effective when segmenting large data sets due to its efficiency and relatively low computational cost. It is also able to adapt to changes making it the preferred choice when a data set is being frequently updated. However, the **K-Means** algorithm cannot detect the optimal number of clusters, requiring it to be specified by the user before running the model. This algorithm is also only compatible with numerical attributes and spherical data sets.

The **Expectation Maximisation (EM)** algorithm comprises of two steps.

The estimation step estimates a value for latent variables for each data point. Latent variables are variables that cannot be objectively observed, such as ‘happiness’, ‘morale’ and ‘confidence’. The maximisation step then optimises the parameters of the probability distributions to best capture the density of the data. These steps are repeated until a good set of latent values and maximum likelihood is achieved that best fits the data.

When selected, the **Let Model Decide** option for the **Number of Clusters** option becomes available, as the optimal number can be detected from the loaded data set using this algorithm. The **EM** algorithm does not assume clusters to be of any geometry, and works well for non-linear geometric distributions. Cluster sizes are also not biased to have specific structures. However, The **EM** algorithm is highly complex and must utilise all the features it has access to when running the model. This may result in longer processing time.

Enter a name for the model in the **Name** textbox, and select the required model from the Model Type drop-down list.

Click the **Advanced Settings** button to change the parameters of the selected model. It is recommended that the default values are used when first building the model, as they can be tweaked to suit requirements once the model has built successfully.

The following advanced settings are available:

The **Elbow Method** graph at the bottom of the page plots the number of clusters against the sum of squared differences to visually represent the optimal number of clusters for the data set. The goal is to choose a small ‘k’ value that still has a low SSE. The chart represents where the diminishing returns occur when increasing the ‘k’ value.
{% endstep %}

{% step %}

### Model Results

With the **Model Type** selected, the **Create Models** screen displays all currently configured models along with performance scores and a number of tabs to further visualise the clusters that have been generated.

Use these screens to compare and contrast the created models before generating the reporting output.

Once a model has been selected, click Build Report from Selected Model from any tab to navigate to the report that has been generated based on the findings of the model.
{% endstep %}
{% endstepper %}

### Performance Scores

As clustering is an unsupervised task with no ‘correct’ answer, all areas of the generated report should be carefully reviewed to make informed decisions when selecting the best model for the corresponding problem.

The following metrics are mathematical attempts to measure the clustering performance:

| Metric               | Description                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Davies-Bouldin Index | Commonly used to evaluate the performance of clustering algorithms, with a lower value indicating better clustering. However, a low value does not imply that the best information has been retrieved.                                                                                                                                                                                                                                                  |
| Silhouette Score     | <p>A metric used to calculate the performance of a clustering technique by calculating the separation between clusters. Its value ranges from -1 to 1, with a higher number implying better clustering result.<br><br>A score of 1 means the clusters are far apart from each other and clearly distinguished, while a score of 0 means that the distance is not significant. If a score of -1 is displayed, the clusters are assigned incorrectly.</p> |

### Table

Created models are displayed along with their current configuration and performance metrics. Use the **Sort By** drop-down list to order the created models as required.

Click **New Model** to create a new model based on the currently loaded configuration.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

### Scatter

The scatter graph plots each cluster on a set scale to allow for direct comparisons in the data. To change the colour of a cluster, click the **Edit** icon in the key.

### Centroid Plot

This graph plots the centroid, the centre of a cluster. This is based on their average value for fields in the data set.

Each cluster should be as far away from other clusters as possible, indicating that each cluster contains a distinctive set of data points.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

### Centroid Table

In a similar analysis to the **Centroid Plot** tab, values here are instead listed rather than plotted on a graph. A centroid is the centre of a cluster.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

This graph plots the centroid, the centre of a cluster. This is based on their average value for fields in the data set.

### Distribution

Here, the relative size of each cluster is displayed on a pie chart. This can be used as a high-level metric for measuring the comparative size of the groupings across the data set.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

### Comparison

The **Comparison** tab displays all data for a cluster to allow for direct comparison against other clusters. Click the required cluster name in the **Key** to select a cluster of interest.

This provides a visual representation of the clustering distributions created for the data set, with the selected cluster displayed prominently against a backdrop of the other clusters on the same scale.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

Consistent values across the graph indicate more compact clusters containing values that are closer together. On the graph, clusters with a straight line can be considered compact without too much deviation compared to other items within the same cluster.

## View Report

### Summary

This tab displays the model summary along with a number of key metrics to represent the clusters that have been detected in the data set.

The values at the top of the screen communicate the settings used along with the appropriate performance scores.

### Cluster Summary

The **Cluster Summary** displays the cluster size distribution over the data set.

For each cluster, a summary displays the key statistical characteristics along with the percentage difference between a cluster average and the whole data set average.

### Clustered Data Table

This table displays the entire data set with a **Clusters** column to display the assigned cluster for each row.

Use the page numbers at the top of the screen to navigate through the entire data set, and sort the values using the column headings and by clicking values of interest.

### Data Plot

The **Data Plot** tab displays all data for a cluster to allow for direct comparison against other clusters. Click the required cluster name in the **Key** to select a cluster of interest.

This provides a visual representation of the clustering distributions created for the data set, with the selected cluster displayed prominently against a backdrop of the other clusters on the same scale.

Click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

Consistent values across the graph indicate more compact clusters containing values that are closer together. On the graph, clusters with a straight line can be considered compact without too much deviation compared to other items within the same cluster.
