Skip to content

Normalization

The normalization engine is designed to standardize commercial parameter data, making different parameters comparable by adjusting them to fit within a range of [0->1]. The engine also groups each normalized value into a partition, based on the parameter type.

To enhance response times, the system periodically pre-calculates and stores these normalized values, eliminating the need for on-the-fly calculations. Any changes made to the commercial parameter data will automatically prompt a new normalization calculation.

Hold "Ctrl" to enable pan & zoom
flowchart LR
    subgraph parameter["Parameter"]
        direction TB
            A[Partition 1]---B
            B[Partition 2]-.-N["Partition N"];
    end
    ingested[Ingested \n Parameter] --> parameter
    normalized["Normalized \n Value"]
    parameter -- Partition into a \n Normalized Value --->  normalized

Commercial properties

A normalized value will be created for each partition, which means that is possible to see how well each item fits within any given partition. All the normalized values for an item is called the commercial properties of that item:

Item Parameter Partition Value
WINE-42 STOCK FEW 1
STOCK SOLD OUT 0
STOCK PLENTY 0
NEW PRODUCT 7 DAYS 0
NEW PRODUCT 14 DAYS 0
NEW PRODUCT 21 DAYS 1
PRODUCT CLICK RANK 0.75

The table above shows an example of how the normalized values could be calculated for a single item.

Parameter normalization

The following section describes how each parameter is normalized, depending on the parameter type.

Term

A term normalization will look at the input value and match the input against the terms defined on partitions to find a fitting partition.

Warning

The matching is case sensitive e.g. "FEW" and "few" are two distinct values.

An input value matching the partition would get the normalized value 1, and the value 0 for all other partitions of the parameter.

Example of Term Normalization

Given the following SKU with a STOCK parameter:

Item Parameter Ingested Value
WINE-42 STOCK FEW

Partitions

Parameter Partition Term
STOCK FEW FEW
STOCK SOLD OUT SOLDOUT
STOCK PLENTY PLENTY

Result of Normalization

Item Parameter Partition Value
WINE-42 STOCK FEW 1
STOCK SOLD OUT 0
STOCK PLENTY 0

Range

A range normalization matches the input value against the ranges defined on the partitions to find a fitting partition.

An input value within the range of a partition gets the normalized value 1, and all other partitions of the parameter get 0 as their normalized value.

Example of Range Normalization

Given the following SKU with a NEW PRODUCT parameter:

Item Parameter Ingested Value
WINE-42 NEW PRODUCT 12

Partitions

Parameter Partition From To
NEW PRODUCT 7 DAYS 0 8
NEW PRODUCT 14 DAYS 8 15
NEW PRODUCT 21 DAYS 15 22

Result of Normalization

Item Parameter Partition Value
WINE-42 NEW PRODUCT 7 DAYS 1
NEW PRODUCT 14 DAYS 0
NEW PRODUCT 21 DAYS 0

Rank

A rank normalization uses the input value to numerically sort and rank the items.

The system ranks a sample of up to 5,000 distinct values and groups them into at most 100 rank clusters. The normalized value of an item is interpolated linearly between the two clusters closest to its value. Items with the same value get the same normalized value. The lowest and highest clusters lie slightly beyond the lowest and highest ingested values, so the items at the extremes get values close to, but not exactly, 0 and 1.

Example of Rank Normalization

Given the following SKU's with a CLICKS parameter:

Item Parameter Ingested Value
WINE-42 CLICKS 12
WHISKEY-22 CLICKS 50
GIN-10 CLICKS 26
RUM-14 CLICKS 26
BEER-4 CLICKS 30

Partitions

Parameter Partition
CLICKS Rank

Result of Normalization

Item Parameter Partition Value
WINE-42 CLICKS Rank 0.05
WHISKEY-22 CLICKS Rank 0.87
GIN-10 CLICKS Rank 0.33
RUM-14 CLICKS Rank 0.33
BEER-4 CLICKS Rank 0.67

Cluster rank

A cluster rank normalization uses the same rank clusters as a rank normalization, but with a key difference. Instead of interpolating between clusters, each item gets the normalized value of the cluster closest to its value.

Clustering groups items based on similarity, resulting in clusters that vary in size. All items within a cluster get the same normalized value, which makes the normalized values step-wise rather than continuous.

Proxy

A proxy normalization directly uses the input value as the normalized value.

While any value is accepted, it's recommended that the value falls within the range of [0->1]. This ensures consistency with the normalized values generated for other parameters.