Association Rules, and What Confidence Leaves Out
How support, confidence and lift are computed from transaction counts, why the anti-monotone property lets a level-wise search examine a fraction of the candidate lattice, and the case that matters most: a rule whose confidence passes any conventional threshold while its lift falls below one, so the consequent is less frequent among transactions containing the antecedent than it is overall, which is the opposite of what the confidence figure suggests.
Definition
Transaction data is a collection of sets. Each transaction lists the items that occurred together, with no order and no quantities.
An itemset is any set of items. Its support is the fraction of transactions containing it:
An association rule
Support says how often the combination occurs at all. Confidence says how often
The anti-monotone property. If
Apriori uses this to search level by level: count the support of single items, discard the infrequent, form candidate pairs only from survivors, and continue. Whole branches of the lattice are never examined, because one infrequent subset rules out everything above it.
What the procedure does not supply. The support threshold is an input, not a finding; rules below it are never generated, whatever their lift. And the search is exhaustive within the threshold, so it returns every qualifying rule rather than the interesting ones, selecting among them is the analyst's work.
Assumptions and scope
Support and confidence are computed from counts of transactions, not of items. A transaction containing an item twice counts once, and quantities are outside the model.
The support threshold determines what is generated. A rule below it is never produced, however strong its lift would have been, so a rare but genuine pattern is invisible to a search tuned for common ones.
Lift equals one under independence and is symmetric in the two itemsets. Confidence is not symmetric, so a rule and its converse share a lift and generally differ in confidence.
Lift is a comparison against independence within this transaction set. It supports no causal claim: two items may co-occur because one drives the other, because both follow a third, or because of how the transactions were collected.
A high confidence computed from very few transactions is weak evidence. A rule holding in
of transactions and one holding in of have the same support, , and the same confidence; what differs is the number of transactions behind the estimate and therefore its sampling uncertainty. Support is a proportion and does not record it. Reporting the raw counts, or an interval for the rule measure, is what separates the two cases. The figures in this unit come from exact rational arithmetic on the stated ten-transaction set, with the Apriori candidate counts obtained by running the level-wise search and comparing against the full
lattice.
Worked material
Example
Every two-item rule, ranked by confidence
All twenty directed rules between pairs of items in the unit's ten transactions, ordered by confidence, which is how a mining tool would present them by default.
| rule | support | confidence | lift |
|---|---|---|---|
| butter → milk | |||
| butter → bread | |||
| jam → milk | |||
| jam → bread | |||
| milk → bread | |||
| bread → milk | |||
| eggs → milk | |||
| eggs → bread | |||
| milk → eggs | |||
| bread → eggs | |||
| butter → eggs | |||
| milk → butter | |||
| bread → butter | |||
| eggs → butter | |||
| milk → jam | |||
| bread → jam |
(The four rules pairing jam with eggs or butter have support
The top four rules have perfect confidence and rest on two transactions and one transaction respectively. The most trustworthy rules in the table, bread and milk, on seven transactions, sit fifth and sixth. Sorting by confidence puts the flimsiest evidence at the top.
Five rules pass a
The symmetry is visible. Each lift value appears in both directions of its pair:
Only two distinct lifts appear,
---
Confidence ranks rules by how common the consequent is, mostly. Lift separates them into those occurring more often than chance and those occurring less. Support says how much data stands behind either verdict. Three columns, three different questions, and a tool presenting only the second column, sorted, would mislead on all of them.
Contrast
Pairs that differ in one respect
Two rules with comparable confidence and opposite verdicts.
| bread → milk | bread → eggs | |
|---|---|---|
| support | ||
| confidence | ||
| consequent's own support | ||
| lift | ||
| verdict | occurs more than chance | occurs less than chance |
Both confidences are respectable. The column that decides is the third, and neither confidence contains it. This is the single comparison worth taking from the unit: a confidence is only interpretable beside the consequent's base rate.
A rule against its converse.
So the question "are bread and eggs associated" has one answer, and the question "given bread, expect eggs" has a different answer from "given eggs, expect bread". Choosing the measure means choosing which question is being asked.
Perfect confidence on two transactions against modest confidence on seven.
Support is the column that separates them, which is why it is reported rather than used only as a filter.
A pruned branch against an examined one.
Butter has support
One measurement settles a whole branch; another opens one. The asymmetry is what makes the search affordable, and it rests on support never rising as a set grows.
Exhaustive search against a selective one.
Apriori returns every rule above the thresholds, all of them, guaranteed complete. That completeness is a strength for the search and a burden afterwards: on real data it produces thousands of true statements, and nothing in the procedure ranks them by interest. A method returning fewer, better rules would be making judgements the algorithm has no basis for making.
Lift above one against a cause.
A lift of
Common errors
Common misconception
That a rule with high confidence has found an association, so a threshold on confidence is enough to filter mined rules. Confidence is the proportion of transactions containing the antecedent that also contain the consequent, and it takes no account of how common the consequent already is. On the ten-transaction set of this unit,
Common misconception
That
Related units
Requires
Connected
Learn this topic
Used in
Sources
- Introduction to Data Mining (2019)