Association Rules, and What Confidence Leaves Out
What you will be able to do
The learner can compute support, confidence and lift for a rule over a small transaction set, explain why the anti-monotone property lets Apriori skip most of the candidate lattice, and decide whether a rule with high confidence describes an association at all, recognising that confidence above a threshold is compatible with the antecedent making the consequent less likely.
Orientation
A rule with high confidence and lift below one
Ten transactions from a small shop. Bread appears in eight of them, eggs in six, and both together in four.
Mine the rule
That would be acting against the data.
Eggs appear in
below one. The rule is perfectly true as a frequency, and the association it describes runs in the opposite direction from the one the confidence suggests.
No threshold on confidence can catch this, because confidence never looks at how common the consequent already is. The correction is one division, by the consequent's own support, and it changes the question from "how often does
A second thing about rule mining is worth knowing before starting. The search space is every subset of every item: for five items that is
This unit covers the three measures and what each leaves out, the pruning property and the saving it produces, and the judgement the algorithm cannot make: which of the rules it returns are worth anyone's attention.
Definition
What each measure is a proportion of
The canonical statement above gives the three formulas. What follows is what each one divides by, since that is what decides which question it answers and which it cannot.
What each measure divides by.
| Measure | Divides pair count by | Answers |
|---|---|---|
| support | all transactions | how often does this combination occur at all |
| confidence | transactions containing | given |
| lift | expected count under independence | is |
Support and confidence share a numerator and differ only in what they are proportions of. That is why a rule can be strong by one and negligible by the other: a combination occurring in two transactions out of ten has support
Why lift is the one that answers the question people mean. Writing it out,
the denominator is the support the pair would have if
Symmetry follows from the same expression. Swapping
The anti-monotone property, and why it is exact. Adding an item to a set can only narrow the set of transactions containing it, so
What the thresholds decide before any mining happens. The support threshold fixes what is generated at all. A rule below it never appears, whatever its lift would have been. So the choice of threshold is a decision about which patterns are visible, made before the data is examined, and a mining run reports nothing about the patterns its own threshold excluded.
What the output is. Every rule meeting the thresholds, exhaustively. Not the interesting ones, not the actionable ones, all of them. The measures rank; they do not select.
Intuition
Why the popular item wins every rule
Suppose an item appears in ninety percent of transactions. Then almost any rule ending in that item has high confidence, not because the antecedent has anything to do with it, but because the item is nearly always there. Confidence measures how often
This is the whole difficulty in one observation. The measure people reach for rewards popularity, and popularity is exactly what is not interesting: nobody needs a rule to discover that most baskets contain bread.
Lift asks the question that was meant. Divide the confidence by how often
Consider the second case.
Why the search is affordable. The candidates are every subset of the catalogue, which doubles with each item. That sounds hopeless, and it would be but for one fact: adding an item to a set can only shrink the set of transactions containing it. Support never rises as an itemset grows.
So once an item is too rare to pass the threshold, every set containing it is too rare, and the search can skip that entire branch without counting anything in it. On this unit's data butter has support
The saving is real:
What the algorithm hands back, and what it does not. Every rule above the thresholds, which on a real catalogue is thousands. Each is a true statement about the transactions. Whether any of them means anything, whether the co-occurrence reflects a shared cause, an artefact of how the data was collected, or a pattern worth acting on, is untouched by the computation. The mining is the easy part.
Example
Every two-item rule, ranked by confidence
All twenty directed rules between pairs of items in the unit's ten transactions, ordered by confidence, which is how a mining tool would present them by default.
| rule | support | confidence | lift |
|---|---|---|---|
| butter → milk | |||
| butter → bread | |||
| jam → milk | |||
| jam → bread | |||
| milk → bread | |||
| bread → milk | |||
| eggs → milk | |||
| eggs → bread | |||
| milk → eggs | |||
| bread → eggs | |||
| butter → eggs | |||
| milk → butter | |||
| bread → butter | |||
| eggs → butter | |||
| milk → jam | |||
| bread → jam |
(The four rules pairing jam with eggs or butter have support
The top four rules have perfect confidence and rest on two transactions and one transaction respectively. The most trustworthy rules in the table, bread and milk, on seven transactions, sit fifth and sixth. Sorting by confidence puts the flimsiest evidence at the top.
Five rules pass a
The symmetry is visible. Each lift value appears in both directions of its pair:
Only two distinct lifts appear,
---
Confidence ranks rules by how common the consequent is, mostly. Lift separates them into those occurring more often than chance and those occurring less. Support says how much data stands behind either verdict. Three columns, three different questions, and a tool presenting only the second column, sorted, would mislead on all of them.
Procedure
Mining rules, and deciding which to keep
To compute the measures for one rule.
- Count transactions, not items. A transaction containing an item twice contributes one.
- Count three things: transactions containing
, transactions containing , and transactions containing both. - Divide each by
to get , and . - Confidence is
, and lift is that divided by. - Report all three. Any one alone is misleading in a way the other two would have exposed.
To run Apriori by hand.
- Fix the support threshold first, and record it. It determines what can be found.
- Level 1: count every single item and discard those below threshold.
- Level
: form candidates only from frequent -itemsets, and discard any candidate having an infrequent subset before counting it, that check is cheaper than a pass over the data. - Count the surviving candidates and keep those meeting the threshold.
- Stop when a level produces no frequent itemsets.
- Report how many candidates were counted against how many exist,
foritems. The ratio is what makes the pruning visible.
To generate rules from a frequent itemset.
- For each frequent itemset, consider each way of splitting it into antecedent and consequent.
- Compute confidence for each split. Support is the same for all of them, since it depends only on the whole set.
- Compute lift for each, and note that splits which are reverses of one another share it.
- Discard nothing yet. Filtering belongs in the next stage, where judgement applies.
To decide which rules to report.
- Check lift against one first. Below one, the rule describes a negative association regardless of its confidence, and reporting it as a recommendation inverts its meaning.
- Check the transaction count behind the rule. A confidence of
resting on two transactions is one contrary purchase away from . Multiply support by to get the count, and state it. - Ask whether the lift is large enough to act on. A lift of
is above one and describes a nine percent enhancement, which may or may not justify a change. - Ask what would explain the co-occurrence other than a relationship between the items. A shared cause, a promotion running during collection, the way transactions were defined.
- Say what the rule would be used for, since the same figures justify reporting in one setting and discarding in another.
Checks. Confirm support never rises as an itemset grows, if a computed superset has higher support than its subset, the counting is wrong. Confirm that a rule and its converse share a lift; differing lifts mean an arithmetic error. And before reporting any rule, compute what its confidence would be if the antecedent were irrelevant: that number is
Worked example
Ten transactions, counted through
The data.
| # | items |
|---|---|
| 1 | bread, milk |
| 2 | bread, milk, eggs |
| 3 | bread, milk, butter |
| 4 | bread, milk, eggs, butter |
| 5 | bread, milk |
| 6 | bread, milk, jam |
| 7 | bread, eggs |
| 8 | milk, eggs |
| 9 | bread, milk, eggs |
| 10 | eggs |
Step 1: single-item support. Count transactions containing each item and divide by
| item | count | support |
|---|---|---|
| bread | ||
| milk | ||
| eggs | ||
| butter | ||
| jam |
Step 2: the rule that looks fine. Bread and eggs occur together in transactions
Half of bread buyers also buy eggs. On confidence alone this is a reportable rule.
Step 3: the division that changes the verdict.
Below one. Eggs appear in
Nothing in step 2 could have revealed this, because
Step 4: a rule that survives the check. Bread and milk occur together in transactions
Above one, so this pair really does occur more often than independence predicts, about nine percent more. Note how much smaller that margin is than the confidence suggests:
Step 5: direction. Reverse the first rule.
The confidence changed,
Step 6: the low-support end. Butter appears twice, both times with milk and bread.
Perfect confidence and lift above one, and the entire rule rests on two transactions. A single additional butter-without-milk purchase would drop the confidence to
Step 7: Apriori on the same data, at a support threshold of
- Level 1. Count all
single items. Bread, milk and eggs pass; butter and jam fail. - Level 2. Form pairs from the three survivors only,
candidates, not the pairs available. All three pass: bread-milk , bread-eggs , eggs-milk . - Level 3. One candidate, bread-eggs-milk, whose subsets are all frequent. It passes at
. - Level 4. No candidates remain.
The measures rank rules and do not select them; confidence without lift can point the wrong way; lift without support can rest on almost nothing; and the search that produces all of this is cheap and exact. What remains, deciding which of the surviving rules is worth acting on, is not in any of these numbers.
Contrast
Pairs that differ in one respect
Two rules with comparable confidence and opposite verdicts.
| bread → milk | bread → eggs | |
|---|---|---|
| support | ||
| confidence | ||
| consequent's own support | ||
| lift | ||
| verdict | occurs more than chance | occurs less than chance |
Both confidences are respectable. The column that decides is the third, and neither confidence contains it. This is the single comparison worth taking from the unit: a confidence is only interpretable beside the consequent's base rate.
A rule against its converse.
So the question "are bread and eggs associated" has one answer, and the question "given bread, expect eggs" has a different answer from "given eggs, expect bread". Choosing the measure means choosing which question is being asked.
Perfect confidence on two transactions against modest confidence on seven.
Support is the column that separates them, which is why it is reported rather than used only as a filter.
A pruned branch against an examined one.
Butter has support
One measurement settles a whole branch; another opens one. The asymmetry is what makes the search affordable, and it rests on support never rising as a set grows.
Exhaustive search against a selective one.
Apriori returns every rule above the thresholds, all of them, guaranteed complete. That completeness is a strength for the search and a burden afterwards: on real data it produces thousands of true statements, and nothing in the procedure ranks them by interest. A method returning fewer, better rules would be making judgements the algorithm has no basis for making.
Lift above one against a cause.
A lift of
Warning
Rules that pass every filter and say nothing
Filtering on confidence alone. Confidence never refers to how common the consequent is, so an item appearing in most transactions produces high-confidence rules with every antecedent. On this unit's data, three rules clear a
Sorting by confidence and reading from the top. The four highest-confidence rules here rest on two transactions and one transaction. The best-supported rule in the dataset sits fifth. Default orderings put the flimsiest evidence first, and a report following that order has been shaped by the tool rather than by the data.
Treating a lift above one as a discovery. A lift of
Ignoring the transaction count.
Reading a rule as causal. Co-occurrence above chance has at least four explanations: one item drives the other, a third factor drives both, a promotion ran during collection, or the transaction boundary was defined in a way that groups them. Lift distinguishes none of these. It compares against independence, which is a much weaker baseline than causation.
Forgetting that the threshold decided what could be found. Rules below the support threshold are never generated, whatever their lift would have been. A mining run therefore reports nothing about the patterns its own threshold excluded, and a rare but genuine association is invisible rather than absent.
---
Two that follow from the search being exhaustive.
Treating the output as a set of findings. Apriori returns every rule meeting the thresholds, on a real catalogue, thousands of true statements. They are raw material. Presenting them as discoveries mistakes completeness for selectivity, and the selection is the analyst's work rather than the algorithm's.
Multiplicity. Mining thousands of rules and reporting the most extreme is testing thousands of hypotheses and reporting the largest effect. Some rules will look impressive by chance alone, and the usual correction, validating a rule on transactions not used to mine it, is rarely applied, though nothing prevents it.
---
And one about what the measures are for. They rank; they do not decide. A rule with acceptable support, lift comfortably above one, and a plausible mechanism is worth considering. The same figures without the mechanism are worth investigating. The figures alone, whatever their size, are a description of ten or ten million transactions, and the question of what to do about them is not in the arithmetic.
Application
Acting on mined rules, and the assumptions involved
Store layout and promotion. The original use, and the one that exposes the confidence trap most sharply: moving two items together assumes they are bought together more often than chance, which is a statement about lift rather than confidence. Acting on
A subtler point: a rule found in historical transactions describes behaviour under the old layout. Rearranging the shelves changes the process that generated the data, so the rule's own success undermines the evidence for it.
Recommendation systems. "Customers who bought this also bought that" is confidence, computed per antecedent, and it is genuinely what is wanted here, because the recommendation is conditional on a specific item the customer has chosen. The failure mode is recommending whatever is most popular, since popular items have high confidence after everything. Systems that correct for base rates, which is lift, produce less obvious and more useful recommendations.
Medical co-occurrence and pharmacovigilance. Diagnoses or drugs appearing together across patient records are mined the same way, and the stakes change the thresholds. A rule resting on few patients is precisely the signal sought for a rare adverse reaction, so the support threshold that makes retail mining tractable would discard the signal being looked for. The multiplicity problem also bites hardest here: mining thousands of drug pairs guarantees some extreme co-occurrences by chance, and confirmation on separate records is not optional.
Web usage and process mining. Pages visited in a session, or activities appearing in a case, are transactions. What differs is that order often matters, which plain itemset mining discards entirely, sequence mining exists for this reason, and applying association rules to ordered data silently throws the ordering away.
Fraud and anomaly screening. Here the useful rules are the ones that fail: a combination with lift far below one is a combination that essentially does not occur, so an instance of it is worth examining. The same arithmetic serves the opposite purpose, and the rules a retail analyst would discard are the ones a fraud analyst keeps.
---
The arithmetic is identical across all of them; what changes is which measure matches the decision, and what the support threshold makes invisible. A rule is a statement about transactions that were collected under particular conditions, and every application above depends on those conditions still holding when the rule is acted on.