> For the complete documentation index, see [llms.txt](https://documentation.connexica.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://documentation.connexica.com/reporting/prediction.md).

# 10. Prediction

## Prediction

Classification is one of the main methods that can be applied in various scenarios to generate business insights. Prediction is a machine learning task that refers to using predictive modelling to predict a class label. The possible application ranges widely from loan default prediction, customer churn analysis, market subscription prediction and medical diagnosis.

## Data Preparation

When preparing data for a prediction model, a number of key stages are required to first train then apply the model to a data set.

{% stepper %}
{% step %}

### Select Training Data

Before selecting the data that will have an outcome predicted, a sample of historical records must first be selected for the model to use as **Training Data**.

The **Training Data** will contain records that, once analysed, will provide a framework that the model will use to formulate its predictions.

Click the **Filter by Saved Query** button to load a pre-defined set of records, or click the field names to manually select records that will be used. Page numbers are also available to view the entire data set. It is recommended that the training data should at least match or exceed the number of records that will be predicted.

The quality of the output model will depend on the validity of the subset used in the training process, with valid data providing a platform on which repeatable predictions can be made.

Please note that a high record count with a large number of fields will increase the processing time. While more data may provide a more valid overview, model quality will not scale linearly with data quantity.
{% endstep %}

{% step %}

### Select Model Data

Now that the data that will be used to train the model has been selected, the **Model Data** must now be specified. The **Model Data** will contain the records that will be processed by the trained model and will contain the rows that will have an outcome predicted.

To ensure valid results, this cohort must maintain the same shape as the **Training Data**.

Click the **Filter by Saved Query** button to load a pre-defined set of records, or click the field names to manually select records that will be used. Page numbers are also available to view the entire data set.

Please note that a high record count with a large number of fields will increase the processing time.
{% endstep %}

{% step %}

### Select Target Field

With the **Model Data** selected, the **Target Field** must now be established.

The **Target Field** is the field that will be predicted. This may be a field that determines a specific status, or a key identifier used to define an outcome.

Once selected, the **Class of Higher Interest** can be set. This is used to identify a value from the **Target Field** to be predicted.

This drop-down list will populate with every unique value from the **Target Field**, where the outcome of interest can then be selected.
{% endstep %}

{% step %}

### Select Input Fields

Select the fields that will be included in the model output.

Use this step to omit fields that are not required for the reporting output and to reduce processing time.

Click **Save Data Preparation** to save the current configuration and proceed to the model building screens.
{% endstep %}
{% endstepper %}

## Model Creation

With the data preparation saved, the predictive models can now be built based on this configuration.

### Create New Model

Enter a name for the model in the **Name** textbox, and select a model from the **Model Type** drop-down list.

The **Bayesian Network** model type denotes probabilities of the possible outcomes of events and how they are conditionally linked. The structure of a Bayesian network consists of individual nodes that represent a belief about a particular event and arrows to denote the nodes that are probabilistically linked. This is to provide an understanding of the system that is being modelled, allowing predictions to be made about how the system will behave in the future.

The **Decision Tree** model type uses the C4.5 algorithm to calculate the nodes that are most influential on a field of interest to predict the value of the chosen field. The model is displayed with decisions and their subsequent consequences as a graph that will display outcomes as variables are selected, allowing users to pinpoint the variables that have a specific effect on an outcome.

Click the **Advanced Settings** button to change the parameters of the selected model. The following settings are available:

| Setting                                                                                                 | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| ------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Learning Algorithm                                                                                      | The algorithm used for learning the structure of the network.                                                                                                                                                                                                                                                                                                                                                                                                           |
| Use Local Score Metric                                                                                  | <p>Determines how the network is scored at each stage of the algorithm. If selected, the score of the network is determined by the sum of the scores of the individual nodes.<br><br>If unselected, a Global Score Metric is used. This repeatedly splits the data into training and validation sets and assesses the quality of the network on how well it predicts the validation data. This greatly increases the time it takes for the network to be generated.</p> |
| Markov Blanket Classifier                                                                               | Ensures that all nodes are either a parent, child or any other parent of the classifier node.                                                                                                                                                                                                                                                                                                                                                                           |
| Score Type                                                                                              | <p>When Local Score Metric is used, this option determines how the network is scored:<br><br><strong>BDe</strong> (Bayesian Dirichlet likelihood-equivalence)<br><br>Seeks to maximise the joint probability of the data and the network. Evaluates the model and the data given the model.<br><br><strong>MDL</strong> (Minimum Description Length)<br><br>Biases the algorithm to favour a simpler network when choosing between multiple potential outputs.</p>      |
| Initialise as Naïve Bayes                                                                               | Forces the initial structure of the network as a Naïve Bayes before the learning algorithm is applied, where the classifier node has an arrow pointing to every other node.                                                                                                                                                                                                                                                                                             |
| Max Number of Parents                                                                                   | The max number of parents for each node in the network.                                                                                                                                                                                                                                                                                                                                                                                                                 |
| Use Arc Reversal                                                                                        | Considers reversing the direction of any arrows at each step of the generation process.                                                                                                                                                                                                                                                                                                                                                                                 |
| Sample Type                                                                                             | <p>Instead of all the data being used to train the network, a sample can be used:<br><br><strong>First</strong><br>Uses the first portion of the data for training.<br><br><strong>Last</strong><br>Uses the last portion of the data for training.<br><br><strong>Interval</strong><br>Used data at even intervals for training.<br><br><strong>Random</strong><br>Uses a random subset of data for training.</p>                                                      |
| Sample Percentage                                                                                       | The percentage of the data that will be used to create the model.                                                                                                                                                                                                                                                                                                                                                                                                       |
| <p>Random Layout<br><em>Applicable to K2</em></p>                                                       | Randomises the initial ordering of the nodes. Otherwise the order the nodes will be ordered alphabetically.                                                                                                                                                                                                                                                                                                                                                             |
| <p>Number of Look Ahead Steps<br><em>Applicable to Look Ahead Hill Climber</em></p>                     | How many steps ahead the algorithm looks.                                                                                                                                                                                                                                                                                                                                                                                                                               |
| <p>Number of Good Operations<br><em>Applicable to Look Ahead Hill Climber</em></p>                      | How many operations are stored at each step.                                                                                                                                                                                                                                                                                                                                                                                                                            |
| <p>Runs<br><em>Applicable to Look Ahead Hill Climber, Tabu Search, Simulated Annealing and ICS</em></p> | The amount of times an algorithm will run. The algorithm will start with a random network for each run and chooses the best scoring network overall.                                                                                                                                                                                                                                                                                                                    |
| <p>Tabu-List Length<br><em>Applicable to Tabu Search</em></p>                                           | The list of generation steps that the algorithm will not repeat.                                                                                                                                                                                                                                                                                                                                                                                                        |
| Balance Data                                                                                            | Reweights the data so that each class has the same total weight.                                                                                                                                                                                                                                                                                                                                                                                                        |
| Run in Background                                                                                       | Generates the network in the background and saves the report once completed.                                                                                                                                                                                                                                                                                                                                                                                            |

### Model Dashboard

Created models are displayed along with their current configuration and performance metrics. Use the **Sort By** drop-down list to order the created models as required.

In the **Table** tab, click **New Model** to create a new model based on the currently loaded configuration.

Click the **View** button to display a created model, or click the **Edit** button to modify the configuration of a model. To remove a model, click the **Delete** button.

#### Performance

The following metrics are available when evaluating model performance:

| Metric    | Description                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                  |
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Recall    | The measure of how effectively the model detects true positives. A higher **Recall** figure ensures true positives are included, potentially resulting in a higher number of false negatives. This can act as a safety net to ensure true positives are not missed.                                                                                                                                                                                                                                                                                                                                                                                                                                                                                          |
| Precision | The ratio between every positive and all true positives. A model that has a higher **Precision** figure is stricter at declaring a positive value, resulting in a higher number of false negative values. This is desirable when curating a smaller, more accurate list of outcomes at the expense of missing potentially true targets.                                                                                                                                                                                                                                                                                                                                                                                                                      |
| F-Measure | <p>When there is no preference between <strong>Recall</strong> and <strong>Precision</strong>, the <strong>F-Measure</strong> score can be used. This value is a measure of accuracy derived from both the <strong>Recall</strong> and <strong>Precision</strong> scores, with a high score denoting a low number of false positives and false negatives. This means that true positives are being correctly identified and are not being masked by false positives.<br><br>Reviewing both <strong>Precision</strong> and <strong>Recall</strong> is useful in cases where there is an imbalance in the observations between the two classes. Specifically, there are many examples of no event (class 0) and only a few examples of an event (class 1).</p> |

**Charting**

In the **Chart** tab, the **ROC Curve** and **RP Curve** are also plotted to graphically represent the model performance.

The **ROC Curve** measures the ability of the model to classify both true positives and true negatives. A good model will not only be able to accurately predict positive values, but can also differentiate between, and predict, negative values. An accurate test will plot a curve that follows the left-hand border and then the top border.

The **PR Curve** plots the **Precision** value on the y-axis and the **Recall** value on the x-axis, with the area under the curve representing high **Recall** and high **Precision** scores. When both score highly, the model will have greater accuracy.

### Model Report

Now that the model has finished building, a report can be generated from its findings. Two tabs are available in the report: **Business Insights** and **Model Performance**.

In the **Model Prediction** table, the **Prediction** column denotes the predicted outcome for the row based on the currently loaded model.

The **Confidence** column displays a value between 0-1 to score the confidence of the **Prediction** outcome. The closer the value to 1, the more confident the model is about the prediction result.

For example, when identifying individuals who cancel a subscription, a ‘Subscription Status’ field would be selected as the **Target Field**. A **Confidence** score of 1 indicates the model is very confident about the prediction outcome, while a score closer to 0.5 indicates that it only 50% confident.
