Practice: Association Rules, and What Confidence Leaves Out

Direct application

In a set of ten transactions, bread appears in eight, milk appears in eight, and both appear together in seven.

Compute the confidence of the rule bread → milk , to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Confidence is a proportion of the transactions containing the antecedent, not of all transactions.

Hint 2: Next step

Of the eight transactions containing bread, how many also contain milk?

Direct application

In ten transactions, bread appears in eight, eggs appear in six, and both appear together in four. The rule bread → eggs therefore has confidence 0.5000 .

Compute its lift, to six decimal places.

(Then note, for yourself, which side of 1 the answer falls on and what that says about the rule.)

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Lift compares the confidence against how often the consequent occurs anyway.

Hint 2: Next step

The consequent is eggs, with support 6 / 10 .

Direct application

In ten transactions, bread appears in eight, eggs appear in six, and both appear together in four. The rule bread → eggs has confidence 0.5000 .

Compute the confidence of the reversed rule eggs → bread , to four decimal places.

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Confidence divides by the support of whichever itemset is the antecedent.

Hint 2: Next step

Now the antecedent is eggs, which appear in six of the ten transactions.

Direct application

Five items appear across ten transactions, with supports: bread 0.8000 , milk 0.8000 , eggs 0.6000 , butter 0.2000 , jam 0.1000 . The support threshold is 0.3 .

Apriori proceeds level by level, forming candidates at each level only from the frequent itemsets of the level below:

  • Level 1 counts all five single items.
  • Level 2 forms pairs from the level-1 survivors only; all pass.
  • Level 3 forms one candidate from the level-2 survivors; it passes.
  • Level 4 has no candidates.

How many candidate itemsets does Apriori count in total across all levels?

Enter the value. It is checked against the answer and the precision this task asks for.

2 hints available, least help first.

Hint 1: Retrieval cue

Only items passing the threshold at level 1 can appear in level-2 candidates.

Hint 2: Next step

Three items survive level 1. How many pairs can be formed from three items?

Recognition · Error diagnosis

In ten transactions, bread appears in eight, eggs in six, and both together in four.

A learner computes conf ⁡ ( bread → eggs ) = 0.4000 / 0.8000 = 0.5000 and reports: "so the rule between bread and eggs has confidence 0.5000 in either direction, since the four transactions containing both are the same four either way."

Which response identifies the error?

2 hints available, least help first.

Hint 1: Retrieval cue

In the definition of confidence, whose support is the denominator?

Hint 2: Concept cue

Bread appears in eight transactions and eggs in six. Reversing the rule changes which of those two is divided by.

Direct application · Explanation

Five items appear across ten transactions. Butter has support 0.2000 ; the support threshold is 0.3 .

(a) State the anti-monotone property and explain why it holds, from the definition of support.

(b) List the itemsets containing butter that a level-wise search never counts, and say how many there are.

(c) A colleague suggests counting a few of the butter pairs anyway, in case one turns out frequent. Explain why this cannot happen, and distinguish the guarantee from a heuristic that usually works.

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

If a transaction contains { a , b } , does it contain { a } ?

Hint 2: Concept cue

For (c), ask whether the conclusion follows from the definition or from observing many datasets.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The property.

If I ⊆ J then supp ⁡ ( J ) ≤ supp ⁡ ( I ) .

The reason is immediate from the definition. Support counts transactions containing the whole itemset. Every transaction containing J must contain each of its subsets, including I , so the transactions counted for J are a subset of those counted for I , and a subset cannot be larger. Adding an item to a set can only narrow the collection of transactions containing it.

The contrapositive is what the search uses: if I is infrequent, every superset J ⊇ I is infrequent too.

(b) What is skipped.

With five items, the itemsets containing butter are butter itself plus every combination of butter with a subset of the other four items. That is 2 4 = 16 itemsets in total: butter alone, four pairs, six triples, four quadruples, and the full five-item set.

Once butter is measured at 0.2000 and found infrequent, the other fifteen are settled without being counted:

  • pairs: butter-bread, butter-milk, butter-eggs, butter-jam
  • triples: the six combinations of butter with two others
  • quadruples: the four combinations of butter with three others
  • the five-item set

One measurement disposes of fifteen candidates.

(c) Why counting them is wasted effort.

It cannot happen, and the reason is the property in (a) rather than an empirical observation. Any itemset containing butter has support at most supp ⁡ ( butter ) = 0.2000 , which is below the threshold of 0.3 . So each of the fifteen is provably infrequent before any transaction is examined.

The distinction from a heuristic matters. A heuristic that usually works would leave open the possibility of a missed pattern, and using one would oblige an analyst to say what was risked, perhaps to spot-check, as the colleague suggests. Here the pruning is exact: the set of frequent itemsets found by Apriori is identical to the set found by examining all 31 candidates. Nothing is traded away for the speed.

This is unusual. Most ways of making a search cheaper give up a guarantee, and this one does not, which is why the property is the foundation of the algorithm rather than an optimisation bolted onto it.

A complete answer does each of these:

  • applies antimonotone pruning

Interpretation · Evaluation

A mining tool returns these rules from ten transactions, sorted by confidence as it does by default:

rulesupportconfidencelift
butter → milk 0.2000 1.0000 1.250000
jam → bread 0.1000 1.0000 1.250000
bread → milk 0.7000 0.8750 1.093750
eggs → bread 0.4000 0.6667 0.833333
bread → eggs 0.4000 0.5000 0.833333

An analyst reports the top two as the headline findings and recommends stocking milk beside butter and bread beside jam.

(a) Say what is wrong with taking the top of this list, referring to the support column.

(b) Two rules here have lift below one. Say what that means and whether either should be reported as a finding.

(c) Which rule in the table is best supported, and what would you say about it, including what it does not establish?

Write your answer, then compare it with the worked solution.

2 hints available, least help first.

Hint 1: Retrieval cue

Support times the number of transactions gives the count behind the rule.

Hint 2: Concept cue

Compare each confidence against the consequent's own frequency before judging it.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The top of the list is the weakest evidence. Support 0.2000 on ten transactions means two transactions; support 0.1000 means one. So butter → milk has perfect confidence because both butter purchases happened to include milk, and jam → bread because the single jam purchase included bread. These are fragile to a degree the confidence column conceals. One butter purchase without milk would take that rule from 1.0000 to 0.6667 ; one more jam transaction could halve the second. A rule computed from one observation is not a rule. The ordering is what misleads here. Sorting by confidence systematically promotes rules with few supporting transactions, because a small denominator reaches 1.0000 easily. The best-supported rule in this table sits third. (b) Lift below one means a negative association. eggs → bread and bread → eggs both have lift 0.833333 . The same value, as lift is symmetric. The two items co-occur about seventeen percent less often than independence would predict. Neither should be reported as a finding in the sense the analyst intends. A confidence of 0.6667 looks like a pattern worth acting on, and acting on it, pairing bread and eggs to encourage the combination, would work against what the data shows. If anything is reportable here it is the negative association itself, stated as such. Note that eggs → bread at confidence 0.6667 outranks nothing in this table by lift, yet sits fourth by confidence, above a rule with a genuinely positive association. The two columns disagree about the ordering, and the lift column is the one bearing on whether the items go together. (c) The best-supported rule is bread → milk . Its support of 0.7000 is seven transactions, by far the most evidence in the table, with confidence 0.8750 and lift 1.093750 . It is the only rule here combining substantial support with a lift above one. What I would say about it: bread and milk co-occur about nine percent more often than independence predicts, on seven of ten transactions. That the lift is 1.093750 rather than something dramatic matters because the confidence of 0.8750 sounds much stronger than the association actually is. Both items are common, and most of that 0.8750 is explained by milk appearing in eighty percent of transactions regardless. What it does not establish: that buying bread causes anyone to buy milk. Co-occurrence above chance is consistent with a shared cause (a weekly staples shop), with how the transaction boundary was drawn, or with a promotion running while the data was collected. Lift compares against independence, which is a far weaker baseline than causation. It also does not establish that the pattern will persist. Ten transactions are few, and a rule mined from historical data describes behaviour under the conditions that generated it, including the store layout, which an intervention based on this rule would change.

A complete answer does each of these:

  • judges rule worth
  • reads lift against one

Transfer · Evaluation · Explanation

A health system mines 400,000 prescription records for drug pairs that co-occur with an adverse-event code. Each record is one patient's medication list plus any recorded events.

The team reports: "We set a minimum support of 0.01 and mined all rules, then ranked by confidence and reviewed the top 50. The rule { drug A } → adverse event has confidence 0.42 , well above our 0.30 threshold, and appears in 6,200 records. We conclude drug A elevates the risk of this event and recommend a prescribing warning. Rules below our support threshold were not examined."

The adverse event's overall rate in the dataset is 0.55 .

Write a review covering:

(a) What the reported confidence of 0.42 establishes, given the event's base rate, with the relevant figure computed.

(b) Whether the support threshold was appropriate for this question, and what it may have excluded.

(c) What ranking by confidence over a large rule set does, and what statistical problem the "top 50" review introduces.

(d) Whether the causal conclusion follows, and what else could produce this co-occurrence.

(e) What you would do instead, and what you would report.

Write your answer, then compare it with the worked solution.

3 hints available, least help first.

Hint 1: Retrieval cue

The event occurs in 0.55 of all records. What does that make the lift of a rule with confidence 0.42 ?

Hint 2: Concept cue

Minimum support 0.01 on 400,000 records is a floor of 4,000 records. How common is a reaction that would matter clinically?

Hint 3: Strategy cue

Separate four claims: the measure, the threshold, the ranking, and the causal step. Each fails for a different reason.

Compare with the worked solution

Comparing does not record a result. Judging your own written answer cannot show that you can do this without help.

(a) The confidence of 0.42 points the opposite way from their conclusion. The adverse event occurs in 0.55 of all records. Among patients on drug A it occurs in 0.42 . So

lift = 0.42 0.55 = 0.763636 ,

below one. Patients taking drug A experience this event less often than patients generally, about twenty-four percent less. Their own figures, correctly read, are evidence against the association they are reporting. A threshold of 0.30 on confidence cannot detect this, because confidence never refers to the consequent's base rate. Any event occurring in over half of all records will produce confidences above 0.30 with almost every antecedent, purely by being common. The team has set a bar that a null result clears easily. This is the central failure. A prescribing warning issued on this basis would restrict a drug whose recorded association with the event is protective in this dataset. (b) The support threshold was wrong for this question. Minimum support 0.01 on 400,000 records requires a pattern in at least 4,000 of them. Adverse drug reactions of clinical concern are frequently much rarer. A reaction in 1 in 5,000 patients would appear in about 80 records and be invisible to this search. So the threshold was chosen as though this were a retail-scale mining exercise, where it makes the computation tractable and discards only trivia. Here it discards exactly the class of signal the study exists to find. The rules excluded are not reported as excluded, so the write-up contains no trace of what was made invisible, and "we found no rare signals" would be an unsupported claim, since the search was configured not to look. The anti-monotone property is why the threshold has this reach: every itemset containing a rare drug is at most as frequent as that drug, so the whole branch is pruned without examination. That pruning is exact and lossless relative to the threshold, and the threshold is the assumption doing the damage. (c) Ranking and reviewing the top 50. Sorting by confidence promotes rules whose consequent is common, since a common consequent follows almost any antecedent at a high rate. With the event at 0.55 , the ranking is substantially a ranking by how common each drug is, not by how much each drug changes the risk. The genuinely interesting rules, large lift, adequate support, will not concentrate at the top. The review of the top 50 introduces a multiplicity problem. Mining 400,000 records across many drugs generates a very large number of rules; selecting the 50 most extreme by one measure and interpreting them is testing many hypotheses and reporting only the largest effects. Some will look extreme by chance alone, and the usual safeguards, a correction for multiple comparisons, or confirmation of each candidate rule on records not used to mine it, are absent here and would be straightforward to apply. (d) The causal conclusion does not follow, and would not even if the lift were above one. Lift compares the observed co-occurrence against independence. Independence is a much weaker baseline than causation, and several explanations remain open: - Confounding by indication. Drug A is prescribed for a condition that itself raises or lowers the event rate. This is the dominant concern in prescription data and is entirely invisible to co-occurrence counting.
- A third factor. Age, comorbidity, or care setting drives both the prescription and the event.
- Recording artefacts. Patients on drug A may be monitored more or less closely, so events are detected at different rates rather than occurring at different rates.
- How a record was defined. The medication list's time window determines what counts as co-occurring at all. The direction of the rule is also not settled by the data. Confidence is directional, { drug A } → event and event → { drug A } have different confidences because the denominators differ, while lift is symmetric and identical both ways. So the symmetric measure cannot support a directional claim, and the directional measure does not establish causation. Neither licenses "drug A elevates the risk". (e) What to do instead. Report lift against one, with the base rate beside it. Every rule should carry support, confidence, lift and the consequent's base rate. On the flagged rule that is 6,200 records, 0.42 , 0.763636 , 0.55 , and stated that way the rule is plainly not a safety signal. Set the support threshold from the clinical question. If reactions occurring in 1 in 5,000 patients matter, the threshold must permit patterns at that frequency, accepting the extra computation. State the threshold and what it excludes, so the write-up records what could not have been found. Rank by lift among rules with adequate support, rather than by confidence, and apply a multiplicity correction or hold out a portion of the records to confirm candidate rules found in the rest. Treat the output as candidates. Mining generates hypotheses; it does not test them. Each surviving candidate needs review against the indication, whether the drug is prescribed for a condition already associated with the event, and confirmation in an independent dataset. What I would report on this rule. That drug A was flagged by the confidence-based screen, that its lift is 0.763636 against a base rate of 0.55 , that the association in these records is therefore negative rather than positive, and that no prescribing warning is supported. I would also report that the screen as configured could not have detected reactions rarer than about 4,000 records, which is a limitation of the study rather than a finding about the drugs.

A complete answer does each of these:

  • computes rule measures
  • reads lift against one
  • distinguishes direction
  • applies antimonotone pruning
  • judges rule worth
Practice data

Your practice record is stored in this browser only. Clearing it removes every answer and every scheduled review, and cannot be undone.

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.