Choosing What the Reader Will Judge

What you will be able to do

The learner can choose the graphical encoding a stated question requires, justify the choice by what the eye judges accurately, set the scale and baseline so that the visual comparison matches the numerical one, and state which readings the chosen display supports and which it forecloses.

Orientation

A display assigns the reader a task

Choosing how to show data is usually described as choosing a picture. It is more useful to describe it as assigning the reader a job: judge which dot is higher, judge how many times longer this bar is, judge what fraction of the circle this wedge covers.

Those jobs are not equally easy, and the difference has been measured rather than asserted. People compare positions on a shared axis almost exactly and compare angles poorly, so the same numbers shown two ways produce readings of different quality. The data did not change; the task did.

Two further decisions carry as much weight and are often filed as formatting.

The scale decides which comparison the eye performs. A linear axis makes equal differences equal distances; a logarithmic axis makes equal ratios equal distances. A quantity growing at a steady rate bends upward on the first and lies straight on the second, so one display answers "how much was added" and the other answers "is the rate steady".

The baseline decides what the lengths mean. Where a bar's length carries the value, the reader compares lengths measured from the baseline, and moving it rescales every comparison in the picture while every number stays where it was.

This unit covers the perceptual ordering, those two decisions, and one more stated in advance: every display conceals something, and the useful question is whether it conceals what this reader needs.

Definition

What the ordering is a claim about

The canonical statements above give the perceptual ordering and the two scale decisions. What follows is the scope of each, since that is what decides an unfamiliar case.

The ordering ranks tasks, not chart types. A bar chart used to compare heights against a common axis is a position judgement and sits at the top; the same bar chart used to compare the lengths of segments stacked above different baselines is a length judgement and sits lower. The chart name is not the encoding, and one chart can present several tasks at once.

It is a claim about average accuracy on comparison tasks. It does not say a lower-ranked encoding is never right. Colour ranks poorly for magnitude and is the natural encoding for unordered categories; area ranks poorly for precise comparison and is appropriate on a map, where position is already spoken for by geography. The ordering applies when the reader must judge how much.

TaskTypical useWhat it costs
position, common scaledot plot, scatter, lineneeds an axis both series share
position, non-aligned scalessmall multiplescomparison across panels is weaker
lengthbar from a zero baselinerequires the zero
angle, slopepie, line steepnessslope depends on aspect ratio
areabubble, treemapsystematically underestimated
colour saturationheatmapcoarse, and unreliable in greyscale

The logarithmic scale has a precondition and a cost. It requires strictly positive values, so a series touching zero cannot use it. Its cost is that a reader unused to the convention reads the flattening of a curve as a slowdown in absolute terms, when what has been flattened is a constant proportional rate.

The zero-baseline rule is narrower than it is usually quoted. It binds where length or area encodes the value, because those are read from the baseline. It does not bind for position encodings: a line chart of body temperature or a dot plot of test scores is not improved by forcing zero into view, and doing so compresses the variation being examined into an unreadable band.

Banking is an optimisation, not an instruction. Centring segment orientations near 45 degrees maximises how well slope differences are discriminated. Where the display exists to compare levels rather than rates, a different aspect ratio may serve better, and the choice should be made deliberately rather than inherited from a default canvas size.

Intuition

Why the same numbers support different conclusions

A display is an instrument, and like any instrument it has a resolution that depends on how it was built rather than on what it is pointed at.

The resolution comes from the task. Position on a shared axis is read almost exactly because the eye is comparing two points against the same reference. An angle has no such reference: judging that one wedge is 27% and another 31% means estimating two rotations with nothing to align them against, and people are reliably poor at it. Both displays contain the same four numbers. Only one lets the reader recover them.

The scale decides what the eye is comparing. Consider a quantity rising 7% per period for twelve periods. On a linear axis the first increment is 7.000 and the last is 14.734 , a factor of 2.1049 , so the curve steepens and a reader concludes the growth is accelerating. On a logarithmic axis every step is log 10 ⁡ 1.07 = 0.02938378 , constant to within 4 × 10 − 16 , and the same data form a straight line whose slope is the growth rate. Neither picture lies. They answer different questions, and the axis chose which.

The baseline decides what the lengths mean. Two bars for 102 and 108 drawn from zero stand in a height ratio of 1.0588 , which is the true ratio. Move the baseline to 100 and the visible lengths become 2 and 8 , a ratio of 4.0000 . A difference of 5.9% now looks like a fourfold one. Nothing about the data changed; the reader is simply measuring from somewhere else.

The aspect ratio decides judged slope. Two segments rising 0.2 and 1.0 over the same horizontal distance subtend 2.9 ° and 14.0 ° in a wide, flat frame, barely distinguishable, and 38.7 ° and 76.0 ° in a tall one. The rates never changed. Banking centres the orientations near 45 degrees because that is where a difference in slope is easiest to see.

And a summary is not the data. Four datasets can agree on the mean of x at 9.000 , the mean of y at 7.501 , a fitted slope near 0.500 and a correlation near 0.816 , and look nothing alike when plotted: one linear, one curved, one linear with a single outlier, one almost entirely vertical. The statistics are correct and they are not sufficient. That is the strongest argument for plotting, and it cuts both ways. A plot chosen to summarise will conceal in exactly the same manner.

Example

Five questions and the display each one implies

Which of eight regions had the highest revenue? A ranking question, answered by position along a common scale. A dot plot with regions sorted by value answers it in one glance; a pie chart asks the reader to compare eight angles and answers it poorly. Sorting is part of the encoding here, not decoration: an alphabetical ordering forces a search where a sorted one gives the answer by position.

Did the eight regions grow at the same rate? A ratio question. Revenue over time on a logarithmic axis makes equal growth rates parallel lines, so the comparison becomes "are these parallel" rather than "is this gap widening". The same series on a linear axis separates the large regions from the small ones and says almost nothing about rates.

How is response time distributed? Neither a mean nor a bar. The question is about shape, so it needs a display that shows shape: a histogram, or the points themselves when there are few. Reporting a mean of 240 ms conceals whether that is a tight cluster or a bimodal split between a fast path and a slow one, which is usually the thing worth knowing.

Do these two variables move together? A scatter plot, which is a position judgement on two axes at once. A correlation coefficient answers a narrower question, how well a straight line fits, and four datasets sharing r = 0.816 can be linear, curved, linear-with-an-outlier, or determined by a single point. The coefficient is not wrong; it is not sufficient.

What proportion of total cost is each category? The case where a part-to-whole encoding is genuinely apt, and it still competes with a sorted bar chart from zero. A pie is readable when there are two or three slices whose sizes differ by enough that the angle comparison is not what limits the reading; with seven similar slices the angle judgement fails and the labels do the work the picture was supposed to do.

---

The pattern across the five. The question names a comparison, the comparison names a perceptual task, and the task names the encoding. Working the other way, choosing a chart and then deciding what it shows, produces displays that are accurate and unable to answer anything in particular.

Procedure

Designing the display, and checking it before it ships

To choose an encoding.

  1. Write the question as a sentence the reader should be able to answer. "Which region grew fastest" and "how much did each region contribute" lead to different displays from the same table.
  2. Name the comparison that answers it, which is larger, how many times larger, is the rate constant, how are these two related, how is this distributed.
  3. Pick the highest-ranked encoding that supports that comparison. Position along a common scale where values must be compared precisely; length from a zero baseline where magnitudes are the subject; colour and area for categories and for maps, where they are not being read for magnitude.
  4. Check whether one display can carry the question. Several comparisons usually need several panels rather than one crowded chart, and small multiples keep each comparison a position judgement.

To set the scale.

  1. Ask whether the claim is about differences or ratios. Differences take a linear axis; ratios and growth rates take a logarithmic one.
  2. For a logarithmic axis, confirm every value is strictly positive and label the axis so the convention is visible, decade gridlines, or explicit tick values.
  3. State in the caption what the scale makes visible, because a reader who misses the convention will read a flattening curve as a slowdown.

To set the baseline.

  1. If length or area carries the value, the baseline is zero. There is no version of this that depends on taste.
  2. If the meaningful variation is far from zero, do not truncate the bars; change the encoding. A dot plot or a line chart on a restricted axis is a position judgement and is honest on a range that excludes zero.
  3. If an axis is restricted, say so plainly in the axis label rather than relying on the tick values to be noticed.

To set the aspect ratio. Where slopes are the subject, bank so that segment orientations centre near 45 degrees. Where levels are the subject, choose the ratio that gives the vertical range room, and do not inherit whatever the default canvas was.

Checks before it ships.

  • Read the display as the intended audience and write down the conclusion it produces. If that conclusion is not the one the data support, the display is wrong however correct its numbers.
  • Compute the ratio the picture implies and compare it with the ratio in the data. On a truncated bar chart these differ, often by a lot.
  • Look at it in greyscale. Anything carried by colour alone disappears, which tells you whether colour was decoration or load-bearing.
  • Name one thing the display hides. Every display hides something; being unable to name it means not having looked.

Worked example

Four decisions, each measured

(a) The scale, on a steadily growing series. A quantity starts at 100 and rises 7% per period for twelve periods: 100.00 , 107.00 , 114.49 , 122.50 , 131.08 , and so on to 225.22 .

ReadingLinear axisLogarithmic axis
first step + 7.000 + 0.02938378
last step + 14.734 + 0.02938378
ratio of last to first 2.1049 1.0000
shapecurve steepeningstraight line

On the logarithmic axis the twelve step sizes differ by at most 4.44 × 10 − 16 , which is floating-point noise: they are identical. The slope of that line is log 10 ⁡ 1.07 = 0.02938378 per period, and a constant slope is a constant growth rate.

What each display answers. The linear plot answers "how much is being added each period", and the answer is that the additions are growing. The logarithmic plot answers "is the rate steady", and the answer is yes, exactly. A reader shown only the first and asked about the rate will say it is accelerating, which is false.

(b) The baseline, on two close values. Two categories measure 102 and 108 .

BaselineBar heightsHeight ratioTrue ratio
0 102 , 108 1.0588 1.0588
100 2 , 8 4.0000 1.0588

The true difference is 5.9 % . Truncating the axis at 100 presents the reader with bars in a ratio of four to one. The numbers printed beside the bars can be perfectly correct while the comparison the reader performs is wrong by a factor of nearly four.

The repair is not always a zero baseline. If the interesting variation genuinely sits between 100 and 110 , forcing zero into view compresses it to nothing. The answer is then to stop encoding with length: a dot plot on an axis running 100 to 110 is a position judgement, where the reader reads values off the scale rather than comparing lengths from a baseline.

(c) The aspect ratio, on two successive slopes. Two segments rise 0.2 and 1.0 over equal horizontal distances.

Height-to-widthFirst segmentSecond segmentDifference
0.25 2.9 ° 14.0 ° 11.2 °
1.00 11.3 ° 45.0 ° 33.7 °
4.00 38.7 ° 76.0 ° 37.3 °

In the flat frame the two rates are nearly indistinguishable. At a height-to-width ratio of 1 the steeper segment banks to 45 ° and the angular separation is three times larger. The underlying rates are identical in all three panels.

(d) What the summary cannot carry. Four datasets of eleven points each:

Setmean x mean y interceptslope r
I 9.000 7.501 3.0001 0.5001 0.8164
II 9.000 7.501 3.0009 0.5000 0.8162
III 9.000 7.500 3.0025 0.4997 0.8163
IV 9.000 7.501 3.0017 0.4999 0.8165

Every column agrees to three decimals. Plotted, set I is a noisy straight line, set II is a clean parabola, set III is a perfect line with one point far off it, and set IV has all its x values equal except one, so a single observation determines the entire fit.

A regression table reporting these four would present them as the same finding. Only the plot separates them, and a box plot of y would not, because it summarises too.

---

In each case the numbers are untouched and the reader's conclusion moves. That is what makes these decisions part of the analysis rather than part of the presentation.

Contrast

Pairs that differ in one decision

A pie chart against a sorted dot plot, on the same shares.

piedot plot
perceptual taskangleposition, common scale
rank in measured accuracy4th1st
answers "which is largest"poorly when shares are closeimmediately
answers "do these sum to a whole"yes, visiblyonly if stated

The pie carries one thing the dot plot does not: the visual fact that the parts compose a whole. Where that is the point and the slices are few and unequal, it earns its place. Where the reader must rank or compare, it assigns a task the eye performs badly.

Linear against logarithmic, on the same growth series.

linearlogarithmic
first step + 7.000 + 0.02938378
last step + 14.734 + 0.02938378
shapecurve steepeningstraight
reader concludesgrowth is acceleratingrate is constant

Both are faithful renderings of the same twelve numbers. The first is the right choice when the quantity added each period is the subject. A budget, a headcount. The second is right when the rate is the subject.

Zero baseline against truncated, on values of 102 and 108.

baseline 0baseline 100
bar heights 102 , 108 2 , 8
height ratio 1.0588 4.0000
matches the datayesno

The truncated version is not a stronger presentation of a small difference; it is a different comparison. If 5.9 % matters and is hard to see, the answer is a display where small differences are legible without being magnified. A dot plot on a restricted axis, where the reader reads values rather than comparing lengths.

Aspect ratios 0.25 and 1.0, on the same two slopes.

Segments rising 0.2 and 1.0 subtend 2.9 ° and 14.0 ° in the flat frame, a separation of 11.2 ° ; at a ratio of 1 they subtend 11.3 ° and 45.0 ° , a separation of 33.7 ° . The rates are identical in both. Banking is the deliberate version of a choice that is otherwise made by whatever canvas size the tool defaulted to.

A summary statistic against the points.

Four datasets agreeing on means of 9.000 and 7.501 , slopes near 0.500 and correlations near 0.816 look nothing alike. The table treats them as one finding; the scatter plots separate them at a glance. The general lesson is not "always plot" but that every aggregation, including a plot that aggregates, discards the structure it summarises, and the author chooses which structure to discard.

Warning

Distortions that survive correct numbers

Every failure below is compatible with a display whose printed values are exactly right. That is what makes them hard to catch in review: checking the numbers does not check the picture.

A truncated length encoding. Bars for 102 and 108 drawn from a baseline of 100 present a height ratio of 4.0000 where the data give 1.0588 . The axis is labelled, the values may even be printed on the bars, and the comparison the reader performs is still wrong by nearly fourfold. If a small difference genuinely matters, change the encoding rather than the baseline.

An aspect ratio inherited from the tool. The same two rates separate by 11.2 ° in a flat frame and 33.7 ° at a height-to-width ratio of 1. Nobody decided this; the canvas did. A trend described as "flattening" is often a trend drawn wide.

A logarithmic axis without a signpost. The convention is invisible to a reader who does not notice the tick spacing, and the characteristic flattening of a constant-rate curve then reads as a slowdown. The scale needs to be announced, not merely used.

Area for magnitude. Doubling a circle's radius quadruples its area, and readers underestimate area differences even when the mapping is done correctly. Where the mapping is done to radius rather than area, the exaggeration compounds the misjudgement.

Colour as the only channel. Anything carried by hue alone vanishes in greyscale, in projection, and for a substantial fraction of readers. Checking a display in greyscale takes seconds and reveals immediately whether colour was decoration or load-bearing.

---

And two failures of omission.

A summary presented as the data. Four datasets agreeing on every reported statistic to three decimals can be a line, a parabola, a line with one outlier, and a vertical stack. A table of coefficients would report them as the same result. Nothing in the summary signals that it is concealing the difference.

A plot presented as complete. The same applies one level up: a display that aggregates hides the distribution behind each point, a time series hides composition, a two-variable scatter hides the third variable. This is unavoidable and is not a defect. The defect is failing to say which omission was made, because the reader cannot infer it from the picture.

---

The check that catches most of these. Read the finished display as the intended audience, write down the conclusion it produces, and compare that sentence with what the data support. A distortion that survives correct numbers will not survive that comparison, and nothing else in a normal review process performs it.

Application

Where the encoding decision has consequences

Epidemic reporting. Case counts are published on logarithmic axes because the question during growth is whether the rate is changing, and a constant rate is a straight line there. The same series on a linear axis produces headlines about acceleration while the growth rate is flat. Public dashboards that offer a linear-logarithmic toggle are letting the reader choose which question to ask, and most readers do not know that is what the control does.

Clinical trial reporting. Guidelines discourage bar charts of group means because they encode a summary with length and hide the distribution behind it. Showing the individual points, or a box plot with the points overlaid, answers the question a bar cannot: whether the difference in means reflects a shift in the whole distribution or a few extreme observations.

Financial disclosure. Truncated axes on revenue charts are common enough that some regulators and exchanges comment on them. The numbers in the filing are audited; the baseline is not, and a 5.9 % change drawn from a baseline near the values reads as a fourfold one.

Model diagnostics. A residual plot exists because the summary statistics of a fit do not reveal curvature, heteroscedasticity or an influential point. This is the same argument as the four-dataset example, applied inside the analysis rather than to its presentation: the plot is chosen so that a specific failure becomes visible.

Scientific figures under review. Journals increasingly ask for the underlying points alongside any summary display, and for colour schemes that survive greyscale reproduction and common forms of colour vision deficiency. Both requirements are perceptual rather than statistical, and neither is checked by verifying the numbers.

---

The common thread. In each case a correct calculation reaches a reader through a display, and the display decides what they conclude. The decisions that do that work, the encoding, the scale, the baseline, the aspect ratio, what is shown rather than summarised, are part of the analysis and are best defended in the same way: by stating the question, and saying why this display answers it and what it leaves out.

Next step

Practice Choosing What the Reader Will Judge

Practice this

Results update as you type. Use the up and down arrow keys to move between results, Enter to open one, and Escape to close.

Type to search.

Settings

Appearance

Interface density

Your record

Your progress is stored in this browser and nowhere else: an identifier, the answers you have given, the mastery states and review schedule derived from them, and the lesson you last opened. Clearing it makes you a new learner on this device. It cannot be undone, and it will not affect your appearance or density settings.

Focus timer

Focus--minutes remaining

Phase

Kept in this browser only, and used to label the session in your own history.

Today

Nothing recorded yet. Finish a focus session and it will appear here.

Settings

Focus sessions between long breaks.

Sessions you are aiming for in a day.

Notifications

Your history

Sessions are stored in this browser and nowhere else. They are not evidence and never reach your mastery record.