Association Rules, and What Confidence Leaves Out

What you will be able to do

The learner can compute support, confidence and lift for a rule over a small transaction set, explain why the anti-monotone property lets Apriori skip most of the candidate lattice, and decide whether a rule with high confidence describes an association at all, recognising that confidence above a threshold is compatible with the antecedent making the consequent less likely.

Orientation

A rule with high confidence and lift below one

Ten transactions from a small shop. Bread appears in eight of them, eggs in six, and both together in four.

Mine the rule bread → eggs . Its confidence is 4 / 8 = 0.5000 , half of all bread buyers also buy eggs, which passes most conventional thresholds and reads like a finding. A shop acting on it would move the eggs next to the bread.

That would be acting against the data.

Eggs appear in 6 / 10 = 0.6000 of transactions overall. Among bread buyers they appear in 0.5000 . Bread buyers are less likely to buy eggs than customers in general, and the lift states it precisely:

0.5000 0.6000 = 5 6 = 0.833333 ,

below one. The rule is perfectly true as a frequency, and the association it describes runs in the opposite direction from the one the confidence suggests.

No threshold on confidence can catch this, because confidence never looks at how common the consequent already is. The correction is one division, by the consequent's own support, and it changes the question from "how often does B follow A " to "does A make B more likely than it already was". Only the second question bears on whether to act.

A second thing about rule mining is worth knowing before starting. The search space is every subset of every item: for five items that is 31 candidate itemsets, and it doubles with each item added. Yet a level-wise search settles this data by examining 9 of the 31 , under thirty percent, while being guaranteed to miss nothing frequent. The property making that safe is simple enough to state in one line, and it is what makes mining feasible on a real catalogue.

This unit covers the three measures and what each leaves out, the pruning property and the saving it produces, and the judgement the algorithm cannot make: which of the rules it returns are worth anyone's attention.

Definition

What each measure is a proportion of

The canonical statement above gives the three formulas. What follows is what each one divides by, since that is what decides which question it answers and which it cannot.

What each measure divides by.

MeasureDivides pair count byAnswers
supportall transactionshow often does this combination occur at all
confidencetransactions containing A given A , how often does B follow
liftexpected count under independenceis B more frequent given A than marginally

Support and confidence share a numerator and differ only in what they are proportions of. That is why a rule can be strong by one and negligible by the other: a combination occurring in two transactions out of ten has support 0.2000 , and if both of those transactions contain the antecedent, confidence 1.0000 . Note that support is a proportion, so the same figures arise from 200 of 1000 transactions; what separates the two is the count behind the estimate and the uncertainty that follows from it, which neither measure reports.

Why lift is the one that answers the question people mean. Writing it out,

lift ⁡ ( A → B ) = supp ⁡ ( A ∪ B ) supp ⁡ ( A ) supp ⁡ ( B ) ,

the denominator is the support the pair would have if A and B were independent. So lift is an observed-to-expected ratio, and comparing it against 1 is comparing the data against independence. Neither support nor confidence contains that comparison.

Symmetry follows from the same expression. Swapping A and B leaves that formula unchanged, so lift is symmetric. Confidence divides by supp ⁡ ( A ) alone, so reversing the rule changes the denominator and generally the value. A rule and its converse therefore share one lift and have two confidences, which is not an inconsistency but a consequence of their answering different questions.

The anti-monotone property, and why it is exact. Adding an item to a set can only narrow the set of transactions containing it, so supp never increases as an itemset grows. A level-wise search exploits this to skip candidates, and the skipping is lossless: nothing frequent is discarded, because anything frequent has only frequent subsets. This is a guarantee rather than a heuristic, which is what distinguishes Apriori from an approximation.

What the thresholds decide before any mining happens. The support threshold fixes what is generated at all. A rule below it never appears, whatever its lift would have been. So the choice of threshold is a decision about which patterns are visible, made before the data is examined, and a mining run reports nothing about the patterns its own threshold excluded.

What the output is. Every rule meeting the thresholds, exhaustively. Not the interesting ones, not the actionable ones, all of them. The measures rank; they do not select.

Intuition

Why the popular item wins every rule

Suppose an item appears in ninety percent of transactions. Then almost any rule ending in that item has high confidence, not because the antecedent has anything to do with it, but because the item is nearly always there. Confidence measures how often B follows A , and an item that follows everything follows A too.

This is the whole difficulty in one observation. The measure people reach for rewards popularity, and popularity is exactly what is not interesting: nobody needs a rule to discover that most baskets contain bread.

Lift asks the question that was meant. Divide the confidence by how often B occurs anyway, and a common consequent stops earning credit for being common. What survives is the difference the antecedent makes. A lift of 1.093750 says the pair occurs about nine percent more often than independence predicts; a lift of 0.833333 says about seventeen percent less.

Consider the second case. bread → eggs has confidence 0.5000 , a figure most analysts would accept, and lift 0.833333 . The rule is true as a frequency, and the lift says the association is negative: among bread buyers eggs are rarer than in the catalogue as a whole. That is a statement about co-occurrence in this data, not a prediction about what shelving them together would do.

Why the search is affordable. The candidates are every subset of the catalogue, which doubles with each item. That sounds hopeless, and it would be but for one fact: adding an item to a set can only shrink the set of transactions containing it. Support never rises as an itemset grows.

So once an item is too rare to pass the threshold, every set containing it is too rare, and the search can skip that entire branch without counting anything in it. On this unit's data butter has support 0.2000 against a threshold of 0.3 , and with that one measurement every pair, triple and quadruple involving butter is settled, unexamined, and provably not frequent.

The saving is real: 9 candidates counted out of 31 , with nothing frequent missed. The guarantee matters as much as the speed, since an approximation that might drop a genuine pattern would need justifying and this does not.

What the algorithm hands back, and what it does not. Every rule above the thresholds, which on a real catalogue is thousands. Each is a true statement about the transactions. Whether any of them means anything, whether the co-occurrence reflects a shared cause, an artefact of how the data was collected, or a pattern worth acting on, is untouched by the computation. The mining is the easy part.

Example

Every two-item rule, ranked by confidence

All twenty directed rules between pairs of items in the unit's ten transactions, ordered by confidence, which is how a mining tool would present them by default.

rulesupportconfidencelift
butter → milk 0.2000 1.0000 1.250000
butter → bread 0.2000 1.0000 1.250000
jam → milk 0.1000 1.0000 1.250000
jam → bread 0.1000 1.0000 1.250000
milk → bread 0.7000 0.8750 1.093750
bread → milk 0.7000 0.8750 1.093750
eggs → milk 0.4000 0.6667 0.833333
eggs → bread 0.4000 0.6667 0.833333
milk → eggs 0.4000 0.5000 0.833333
bread → eggs 0.4000 0.5000 0.833333
butter → eggs 0.1000 0.5000 0.833333
milk → butter 0.2000 0.2500 1.250000
bread → butter 0.2000 0.2500 1.250000
eggs → butter 0.1000 0.1667 0.833333
milk → jam 0.1000 0.1250 1.250000
bread → jam 0.1000 0.1250 1.250000

(The four rules pairing jam with eggs or butter have support 0 and are omitted.)

The top four rules have perfect confidence and rest on two transactions and one transaction respectively. The most trustworthy rules in the table, bread and milk, on seven transactions, sit fifth and sixth. Sorting by confidence puts the flimsiest evidence at the top.

Five rules pass a 0.5 confidence threshold with lift below one. Every rule involving eggs with bread or milk has lift 0.833333 , and three of them clear a conventional confidence bar. A pipeline filtering on confidence alone would surface all three as findings, each describing an association that runs the opposite way.

The symmetry is visible. Each lift value appears in both directions of its pair: bread → eggs and eggs → bread both show 0.833333 while their confidences are 0.5000 and 0.6667 . Milk and bread both show 1.093750 and, because the two items have identical support, identical confidences of 0.8750 . A coincidence of this data, not a rule.

Only two distinct lifts appear, 1.250000 and 0.833333 , plus 1.093750 for the bread–milk pair. That is an artefact of a tiny dataset with few distinct support values, and it should be read as such rather than as structure.

---

Confidence ranks rules by how common the consequent is, mostly. Lift separates them into those occurring more often than chance and those occurring less. Support says how much data stands behind either verdict. Three columns, three different questions, and a tool presenting only the second column, sorted, would mislead on all of them.

Procedure

Mining rules, and deciding which to keep

To compute the measures for one rule.

  1. Count transactions, not items. A transaction containing an item twice contributes one.
  2. Count three things: transactions containing A , transactions containing B , and transactions containing both.
  3. Divide each by n to get supp ⁡ ( A ) , supp ⁡ ( B ) and supp ⁡ ( A ∪ B ) .
  4. Confidence is supp ⁡ ( A ∪ B ) / supp ⁡ ( A ) , and lift is that divided by supp ⁡ ( B ) .
  5. Report all three. Any one alone is misleading in a way the other two would have exposed.

To run Apriori by hand.

  1. Fix the support threshold first, and record it. It determines what can be found.
  2. Level 1: count every single item and discard those below threshold.
  3. Level k : form candidates only from frequent ( k − 1 ) -itemsets, and discard any candidate having an infrequent subset before counting it, that check is cheaper than a pass over the data.
  4. Count the surviving candidates and keep those meeting the threshold.
  5. Stop when a level produces no frequent itemsets.
  6. Report how many candidates were counted against how many exist, 2 m − 1 for m items. The ratio is what makes the pruning visible.

To generate rules from a frequent itemset.

  1. For each frequent itemset, consider each way of splitting it into antecedent and consequent.
  2. Compute confidence for each split. Support is the same for all of them, since it depends only on the whole set.
  3. Compute lift for each, and note that splits which are reverses of one another share it.
  4. Discard nothing yet. Filtering belongs in the next stage, where judgement applies.

To decide which rules to report.

  1. Check lift against one first. Below one, the rule describes a negative association regardless of its confidence, and reporting it as a recommendation inverts its meaning.
  2. Check the transaction count behind the rule. A confidence of 1.0000 resting on two transactions is one contrary purchase away from 0.6667 . Multiply support by n to get the count, and state it.
  3. Ask whether the lift is large enough to act on. A lift of 1.093750 is above one and describes a nine percent enhancement, which may or may not justify a change.
  4. Ask what would explain the co-occurrence other than a relationship between the items. A shared cause, a promotion running during collection, the way transactions were defined.
  5. Say what the rule would be used for, since the same figures justify reporting in one setting and discarding in another.

Checks. Confirm support never rises as an itemset grows, if a computed superset has higher support than its subset, the counting is wrong. Confirm that a rule and its converse share a lift; differing lifts mean an arithmetic error. And before reporting any rule, compute what its confidence would be if the antecedent were irrelevant: that number is supp ⁡ ( B ) , and a confidence near it is a rule saying nothing.

Worked example

Ten transactions, counted through

The data.

#items
1bread, milk
2bread, milk, eggs
3bread, milk, butter
4bread, milk, eggs, butter
5bread, milk
6bread, milk, jam
7bread, eggs
8milk, eggs
9bread, milk, eggs
10eggs

Step 1: single-item support. Count transactions containing each item and divide by 10 .

itemcountsupport
bread 8 4 / 5 = 0.8000
milk 8 4 / 5 = 0.8000
eggs 6 3 / 5 = 0.6000
butter 2 1 / 5 = 0.2000
jam 1 1 / 10 = 0.1000

Step 2: the rule that looks fine. Bread and eggs occur together in transactions 2 , 4 , 7 and 9 , four of them, so supp = 2 / 5 = 0.4000 .

conf ⁡ ( bread → eggs ) = 0.4000 0.8000 = 1 2 = 0.5000 .

Half of bread buyers also buy eggs. On confidence alone this is a reportable rule.

Step 3: the division that changes the verdict.

lift ⁡ ( bread → eggs ) = 0.5000 0.6000 = 5 6 = 0.833333 .

Below one. Eggs appear in 60 % of all transactions but only 50 % of bread transactions, so buying bread is associated with buying eggs less often. The rule describes a negative association, and any action premised on pairing the two runs against the evidence.

Nothing in step 2 could have revealed this, because 0.5000 was computed without ever referring to how common eggs are.

Step 4: a rule that survives the check. Bread and milk occur together in transactions 1 – 6 and 9 , seven of them, supp = 0.7000 .

conf = 0.7000 0.8000 = 7 8 = 0.8750 , lift = 0.8750 0.8000 = 35 32 = 1.093750 .

Above one, so this pair really does occur more often than independence predicts, about nine percent more. Note how much smaller that margin is than the confidence suggests: 0.8750 sounds emphatic, and the association it reflects is modest, because both items are common.

Step 5: direction. Reverse the first rule.

conf ⁡ ( eggs → bread ) = 0.4000 0.6000 = 2 3 = 0.6667 , lift = 0.6667 0.8000 = 0.833333 .

The confidence changed, 0.5000 became 0.6667 , because the denominator is now the rarer item, while the lift is identical. One co-occurrence, two confidences, one lift.

Step 6: the low-support end. Butter appears twice, both times with milk and bread.

conf ⁡ ( butter → milk ) = 0.2000 0.2000 = 1.0000 , lift = 1.0000 0.8000 = 1.250000 .

Perfect confidence and lift above one, and the entire rule rests on two transactions. A single additional butter-without-milk purchase would drop the confidence to 0.6667 . The support figure 0.2000 is what exposes the fragility, which is why support is reported alongside and not treated as a mere filter.

Step 7: Apriori on the same data, at a support threshold of 0.3 .

  • Level 1. Count all 5 single items. Bread, milk and eggs pass; butter ( 0.2000 ) and jam ( 0.1000 ) fail.
  • Level 2. Form pairs from the three survivors only, 3 candidates, not the 10 pairs available. All three pass: bread-milk 0.7000 , bread-eggs 0.4000 , eggs-milk 0.4000 .
  • Level 3. One candidate, bread-eggs-milk, whose subsets are all frequent. It passes at 0.3000 .
  • Level 4. No candidates remain.
9  candidates counted , 31  in the full lattice , 9 / 31 = 0.290323 .

22 support computations were never performed, and 7 frequent itemsets were found. Every pair, triple and quadruple containing butter or jam was eliminated by two measurements at level 1, not sampled, not estimated, but proved infrequent by the anti-monotone property.

The measures rank rules and do not select them; confidence without lift can point the wrong way; lift without support can rest on almost nothing; and the search that produces all of this is cheap and exact. What remains, deciding which of the surviving rules is worth acting on, is not in any of these numbers.

Contrast

Pairs that differ in one respect

Two rules with comparable confidence and opposite verdicts.

bread → milkbread → eggs
support 0.7000 0.4000
confidence 0.8750 0.5000
consequent's own support 0.8000 0.6000
lift 1.093750 0.833333
verdictoccurs more than chanceoccurs less than chance

Both confidences are respectable. The column that decides is the third, and neither confidence contains it. This is the single comparison worth taking from the unit: a confidence is only interpretable beside the consequent's base rate.

A rule against its converse.

bread → eggs has confidence 0.5000 ; eggs → bread has 0.6667 . Same four transactions, same co-occurrence, different denominators, 0.8000 against 0.6000 . The lift is 0.833333 for both.

So the question "are bread and eggs associated" has one answer, and the question "given bread, expect eggs" has a different answer from "given eggs, expect bread". Choosing the measure means choosing which question is being asked.

Perfect confidence on two transactions against modest confidence on seven.

butter → milk scores confidence 1.0000 and lift 1.250000 , resting on two transactions. bread → milk scores 0.8750 and 1.093750 , resting on seven. The first ranks higher on both measures and is far weaker evidence: one butter purchase without milk would take its confidence to 0.6667 , while one bread purchase without milk barely moves the second.

Support is the column that separates them, which is why it is reported rather than used only as a filter.

A pruned branch against an examined one.

Butter has support 0.2000 , below the 0.3 threshold, so every itemset containing butter is eliminated by that one measurement, provably, not probably. Bread has support 0.8000 , so every pair containing bread is a live candidate and must be counted.

One measurement settles a whole branch; another opens one. The asymmetry is what makes the search affordable, and it rests on support never rising as a set grows.

Exhaustive search against a selective one.

Apriori returns every rule above the thresholds, all of them, guaranteed complete. That completeness is a strength for the search and a burden afterwards: on real data it produces thousands of true statements, and nothing in the procedure ranks them by interest. A method returning fewer, better rules would be making judgements the algorithm has no basis for making.

Lift above one against a cause.

A lift of 1.093750 says two items co-occur about nine percent more often than independence predicts, in this transaction set. It does not say that buying one leads to buying the other, that a shared cause drives both, or that the pattern will hold next month. Those are separate claims, and every one of them needs evidence the transaction counts do not contain.

Warning

Rules that pass every filter and say nothing

Filtering on confidence alone. Confidence never refers to how common the consequent is, so an item appearing in most transactions produces high-confidence rules with every antecedent. On this unit's data, three rules clear a 0.5 confidence bar while having lift 0.833333 . Each describing an association that runs the opposite way from what the rule suggests. A confidence threshold cannot detect this, because the quantity being thresholded does not contain the comparison.

Sorting by confidence and reading from the top. The four highest-confidence rules here rest on two transactions and one transaction. The best-supported rule in the dataset sits fifth. Default orderings put the flimsiest evidence first, and a report following that order has been shaped by the tool rather than by the data.

Treating a lift above one as a discovery. A lift of 1.093750 means about nine percent more co-occurrence than independence predicts. A real but modest enhancement, and a much weaker statement than the confidence of 0.8750 beside it suggests. "Above one" and "worth acting on" are different thresholds, and only the first is computed.

Ignoring the transaction count. supp = 0.2000 on ten transactions is two purchases. A confidence of 1.0000 built on two observations is one contrary case away from 0.6667 . Multiply support by n and state the count, because a proportion conceals how little may be behind it.

Reading a rule as causal. Co-occurrence above chance has at least four explanations: one item drives the other, a third factor drives both, a promotion ran during collection, or the transaction boundary was defined in a way that groups them. Lift distinguishes none of these. It compares against independence, which is a much weaker baseline than causation.

Forgetting that the threshold decided what could be found. Rules below the support threshold are never generated, whatever their lift would have been. A mining run therefore reports nothing about the patterns its own threshold excluded, and a rare but genuine association is invisible rather than absent.

---

Two that follow from the search being exhaustive.

Treating the output as a set of findings. Apriori returns every rule meeting the thresholds, on a real catalogue, thousands of true statements. They are raw material. Presenting them as discoveries mistakes completeness for selectivity, and the selection is the analyst's work rather than the algorithm's.

Multiplicity. Mining thousands of rules and reporting the most extreme is testing thousands of hypotheses and reporting the largest effect. Some rules will look impressive by chance alone, and the usual correction, validating a rule on transactions not used to mine it, is rarely applied, though nothing prevents it.

---

And one about what the measures are for. They rank; they do not decide. A rule with acceptable support, lift comfortably above one, and a plausible mechanism is worth considering. The same figures without the mechanism are worth investigating. The figures alone, whatever their size, are a description of ten or ten million transactions, and the question of what to do about them is not in the arithmetic.

Application

Acting on mined rules, and the assumptions involved

Store layout and promotion. The original use, and the one that exposes the confidence trap most sharply: moving two items together assumes they are bought together more often than chance, which is a statement about lift rather than confidence. Acting on bread → eggs at confidence 0.5000 and lift 0.833333 would pair items that co-occur less than independence predicts. The measure that matches the intervention is the one comparing against the base rate.

A subtler point: a rule found in historical transactions describes behaviour under the old layout. Rearranging the shelves changes the process that generated the data, so the rule's own success undermines the evidence for it.

Recommendation systems. "Customers who bought this also bought that" is confidence, computed per antecedent, and it is genuinely what is wanted here, because the recommendation is conditional on a specific item the customer has chosen. The failure mode is recommending whatever is most popular, since popular items have high confidence after everything. Systems that correct for base rates, which is lift, produce less obvious and more useful recommendations.

Medical co-occurrence and pharmacovigilance. Diagnoses or drugs appearing together across patient records are mined the same way, and the stakes change the thresholds. A rule resting on few patients is precisely the signal sought for a rare adverse reaction, so the support threshold that makes retail mining tractable would discard the signal being looked for. The multiplicity problem also bites hardest here: mining thousands of drug pairs guarantees some extreme co-occurrences by chance, and confirmation on separate records is not optional.

Web usage and process mining. Pages visited in a session, or activities appearing in a case, are transactions. What differs is that order often matters, which plain itemset mining discards entirely, sequence mining exists for this reason, and applying association rules to ordered data silently throws the ordering away.

Fraud and anomaly screening. Here the useful rules are the ones that fail: a combination with lift far below one is a combination that essentially does not occur, so an instance of it is worth examining. The same arithmetic serves the opposite purpose, and the rules a retail analyst would discard are the ones a fraud analyst keeps.

---

The arithmetic is identical across all of them; what changes is which measure matches the decision, and what the support threshold makes invisible. A rule is a statement about transactions that were collected under particular conditions, and every application above depends on those conditions still holding when the rule is acted on.

Next step

Practice Association Rules, and What Confidence Leaves Out

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.