> For the complete documentation index, see [llms.txt](https://documentation.connexica.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.connexica.com/reporting/association.md).

# 11. Association

## Association

Using association rules analysis, associations and correlations can be determined to better understand the connections between itemsets across the data.

## Prepare Data

Creating an Association model requires the data to be in a specific shape, with each row representing a single transactional-level event and each column containing a specific outcome. For example, if using retail data, transactional point-of-sale data must be transposed so that each row represents a single transaction and column names are used for each product, with fields containing a binary flag to indicate whether this item was included in the sale.

### Select Data

Using the checkboxes for each column, select the fields that will be included in the model output.

Use this step to omit fields that are not required for the reporting output and to reduce processing time.

### Create Model

With the data preparation saved, the predictive models can now be built based on this configuration.

#### Select Model

To identify items that often occur together, itemsets must first be recognised. Use the **Model Type** drop-down list to select the between the **Apriori Algorithm** and **FP Growth** options that will be used to generate the association rules based on the discovered relationships.

The **Apriori Algorithm** is an iterative process that first finds the frequent itemsets by generating candidates before returning the appropriate association rules that determine how strongly items are related to each other. The entire data set is scanned multiple time to discover these rules.

**FP Growth** addresses a key shortcoming that the **Apriori Algorithm** poses by not requiring any candidate generation prior to association rule generation. Instead, data is stored more efficiently by mapping itemsets in condensed paths to discover frequency.

Use the **Minimum Support** textbox to specify the threshold for frequency in itemsets. Increasing this value will result in the model only including itemsets with higher frequency counts, therefore creating less association rules. It is recommended that this is left at the default value and then increased should too many association rules be created in the next stage. If **FP Growth** is selected from the **Model Type** drop-down list, an **Automatic** checkbox is available to set this value based on the findings of the model.

Click the **Advanced Settings** button to change the parameters of the selected model. It is recommended that the default values are used when first building the model, as they can be tweaked to suit requirements once the model has built successfully. The following settings are available:

Click **Model Results** to save the current configuration and view the output of the model.

### Model Results

The Model Results page displays the association rules that have been generated from the data set using the specified algorithm.

For each row, a value is specified for the **Antecedent** and **Consequent**. The value in the **Consequent** column represents a condition that occurs when the value in the **Antecedent** column is true. For example, if the value in the **Antecedent** column is ‘Bread’, the value in the **Consequent** column may be ‘Butter’.

If too many association rules have been created, it is recommended that the **Minimum Support** value is increased to filter out less frequent itemsets. To change this value, click the **Settings** button and make the required change using the **Minimum Support** textbox before clicking **Build Model** to rebuild the model with the change applied.

As a starting point, it is recommended that less than 100 association rules should be generated from the model. This ensures that only the most frequent itemsets are included in the output.

When the required number of association rules have been created, click **Create Report** to navigate to the reporting output of the model.

#### Performance

The following metrics are available when evaluating model performance:

| Metric     | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                   |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Support    | <p>Measures how frequently the itemset appears in the data set. A higher <strong>Support</strong> value for a rule indicates that the combination of <strong>Antecedent</strong> and <strong>Consequent</strong> is more frequent in a data set.<br><br>In a retail example, a high <strong>Support</strong> value would indicate that the combination of items is popular for many customers.</p>                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                            |
| Confidence | <p>Measures how likely <strong>Consequent</strong> value appears when the <strong>Antecedent</strong> value is present.<br><br>It is important to remember that this value may misrepresent correlations for some items that are very popular within the data set. In a retail example, ‘Onion’ and ‘Milk’ may not have a direct connection from the perspective of the customer, but as they often appear in baskets together due to their popularity, it may result in a high <strong>Confidence</strong> value.<br><br>To account for the shortcomings of this value, the <strong>Lift</strong> value can instead be used.</p>                                                                                                                                                                                                                                                                                                                                                                                                                             |
| Conviction | <p>The <strong>Conviction</strong> value compares the probability that <strong>Antecedent</strong> value appears without the <strong>Consequent</strong> value.<br><br>A higher <strong>Conviction</strong> value indicates that the <strong>Consequent</strong> value is more dependent on the <strong>Antecedent</strong> value.<br><br>In a retail example, correlations between ‘Birthday Cake’ and ‘Birthday Candles’ would be common, as the <strong>Consequent</strong> (‘Birthday Candles’) are rarely purchased without the <strong>Antecedent</strong> (‘Birthday Cake’).</p>                                                                                                                                                                                                                                                                                                                                                                                                                                                                       |
| Lift       | <p>Measures the independence level between the <strong>Antecedent</strong> and <strong>Consequent</strong> values, with a value of 1 indicating that there is no association.<br><br>A <strong>Lift</strong> value above 1 indicates there is a positive correlation between <strong>Antecedent</strong> and <strong>Consequent</strong>, and that when the <strong>Antecedent</strong> value is present, it is likely that the <strong>Consequent</strong> value is also present.<br><br>A <strong>Lift</strong> value below 1 indicates there is a negative correlation between <strong>Antecedent</strong> and <strong>Consequent</strong>, and that when the <strong>Antecedent</strong> value is present, it is likely that the <strong>Consequent</strong> value is not present. In a retail example, this can be used to interpret substitution products in basket analysis. For example, if a customer purchases washing detergent in a bottle, they are unlikely to also purchase washing detergent capsules, as they can substitute each other.</p> |

## View Report

With a set of association rules generated from the model, the reporting output can now be viewed. This uses the previously displayed performance metrics to graphically represent the associations discovered in the data set.

Use the tabs on the left side of the report to navigate to the different outputs.
