Data Visualisation: Choosing a Chart That Does Not Mislead
A chart earns its place when the reader needs a shape and a table wins when the reader needs a value. Choose the chart from the shape: a movement over time, a comparison across categories, a part against a whole, or a relationship between two measures. Correct numbers mislead through the axis, the units and the caption, so check all three.
A chart is a claim drawn rather than written, and a reader takes in a drawn claim faster than a written one because the picture lands before the sentence does. The eye has already settled which bar is tallest while the heading above it is still being read. The speed is the entire value of a chart and the entire danger of one. Keeping the drawn claim and the true claim the same thing is the whole discipline.
One assignment runs underneath this guide from beginning to end. The Kavery research desk, four people inside an invented firm called Kavery Capital Services Private Limited, prepared a research note on Meenakshi Tubes Private Limited. Sharada Iyer wrote it, Prakash Nadar reviewed it, and Latha Menon commissioned it and took the decision at the end. The analysis is already finished and every number in it has already been agreed. Some of those numbers now have to become pictures.
The product Meenakshi Tubes Private Limited actually manufactures never matters to a single rule that follows. A chart is chosen from the shape of the data and never from what the data means. The rules concern lengths, areas, axes and labels rather than businesses. Substitute a bond, a fund, a warehouse or a single machine and not one of them needs rewriting. Chart choice as a named practice belongs to Cole Nussbaumer Knaflic, whose Storytelling with Data, 2015, is where most working analysts first meet the idea that a chart is a decision somebody made rather than a default the software handed over.
Which question comes before choosing a chart at all?
Almost everybody starts in the wrong place. Numbers arrive, the numbers feel important, and the next move is to pick a chart type. The first question is not which chart, it is whether a chart at all, and skipping that question is where most poor exhibits are born. A chart is not a neutral container for numbers. Drawing something makes a claim about it that the numbers alone did not make, and a claim nobody intended is an error added to a correct set of figures.
Two numbers show what happens. Stating that operating profit was Rs 2,30,00,000 in the year just reported against Rs 2,21,00,000 in the prior year states two values. Drawing those same two values as a line says something extra and unbidden: that there is a slope here, that it continues, and that the movement between the two points is the story. Two points always make a perfectly straight line, so a two point line chart looks like a trend whatever the underlying figures do. Nobody typed a false number. The picture added the claim by itself.
Here is the everyday version. A household writes down what it spent on vegetables in January and in February. Two numbers on a slip of paper is a record. Draw them as a rising line taped to the fridge and the household has quietly told itself that vegetable spending is climbing. Two months cannot possibly establish that. The slip of paper was honest and the line on the fridge is an argument.
Data Visualisation vs Data Table: what does each one do that the other cannot?
Both are ways of putting numbers in front of a reader and they fail at completely different jobs, so each one is worth defining fully before they are set against each other.
A data visualisationA claim about numbers made as a picture rather than a sentence. turns quantities into visual properties. A number becomes a length, an area, a position along a scale, an angle or a shade, and the reader compares those properties instead of comparing digits. The substitution is what makes a chart fast. Comparing two lengths takes no conscious effort at all. Comparing Rs 4,80,00,000 with Rs 5,20,00,000 takes a second or two of genuine reading. The price of the substitution is precision. Nobody reads a value off a bar to the rupee, and nobody is meant to.
A data tableThe same numbers laid out so each value can be read exactly. keeps the numbers as numbers and arranges them so that any one of them can be found and read exactly. A table makes no claim about shape at all. A table will not announce that one row towers over another; a reader has to work that out by reading. In exchange it gives the figure to the last digit, it lets a reader recompute the arithmetic, and it survives being copied into somebody else's work without losing anything.
The two are not ranked against each other, they are aimed at different readers doing different things, and the whole choice turns on which reader the exhibit actually has. A reader who is going to quote the figure, key it into something else, or check whether it matches a filing needs a table and is actively hindered by a chart. A grid of digits gives up its shape slowly or not at all, so a reader who is going to look once, take away an impression and move on needs a chart and is actively hindered by a table.
| Question | A data visualisation | A data table |
|---|---|---|
| What does it hand the reader | A shape, taken in at a glance | A value, readable to the last digit |
| What is it poor at | Precision. Nobody reads a bar to the rupee | Shape. A grid gives up its pattern slowly |
| What does it add that nobody typed | A claim about relative size, trend or proportion | Nothing. It states the figures and stops |
| How many items does it handle well | Many. A hundred points is still one shape | Few. Twenty rows is already hard work |
| What can a reviewer do with it | Challenge the encoding, the axis and the caption | Recompute every figure in it |
How is the table-or-chart test run on a real exhibit?
The test is one question and it is about the reader rather than about the data: what is this person going to do with the exhibit, read a value off it or see a shape in it? The amount of data decides nothing and the reader decides everything, so the question is what the reader will do with the exhibit rather than how much data is in hand. Four numbers can absolutely be a chart if the reader only needs the shape, and two hundred numbers stay a table if somebody has to reconcile them line by line.
In practice a rule of thumb does most of the work. The shape of four numbers can be stated in one sentence and does not need a picture at all, so four numbers that somebody will want to read precisely is a table almost every time. When four numbers are about to be drawn, the sentence is worth writing first, and then checking whether the sentence was all that was needed. Usually it was.
The Kavery research desk held operating margin and revenue for two years: two measures across two periods. Chart or table?
Which four shapes cover almost everything an analyst has to draw?
Once a chart is warranted, the chart type is not a matter of taste. The shape of the claim, rather than the shape of the spreadsheet, is what decides, and the chart then picks itself. Four shapes cover almost everything a working analyst produces: a movement over time, a comparison across categories, a part against a whole, and a relationship between two measures. Habit means whichever chart the last exhibit used, so picking by habit rather than by shape is where most poor charts start.
Shape one, a movement over time: which chart does that shape call for?
A movement over time means one measure observed at points along a continuous timeline, where the claim concerns the path between the points rather than any single point on it. A movement over time calls for a line, and the reason is exact. The horizontal scale is continuous, so the space between two points genuinely means something, and joining them therefore means something too. A line says the measure passed through the space between the observations, and along a real timeline it did.
Two conditions have to hold before one is drawn. Unequal gaps drawn at equal widths distort the slope without touching a single value, so the observations should be evenly spaced in time. And there have to be enough observations for a path to exist at all. Three points is not a path, it is three points with two straight segments the software invented. The scenario curve below carries twenty one observations, and twenty one is a shape.
Shape two, a comparison across categories: when do the bars turn horizontal?
A comparison across categories means several separate things measured the same way, where the claim concerns relative size. Bars, always, and from a zero baseline. Length is what carries the number here: a bar twice as long has to mean twice as much, and length only stays proportional if the measurement starts at nothing.
Time can run along the bottom of the chart while the shape is still a comparison rather than a movement, and that combination trips people more than any other. Four quarters looks like time, so the reflex is a line. Ask which claim the exhibit is making. If the claim is that one quarter towers over the other three, that is a statement about relative size and it is bars, with time merely deciding the left to right order. If the claim is that the measure has been climbing quarter after quarter, that is a path and it is a line. The Kavery research desk wanted the first of those, so the quarterly exhibit is a bar chart.
Past about six categories the bars turn on their side, and the reason is prosaic rather than statistical. Category labels are words, words run horizontally, and a vertical bar gives its label only the width of one bar to sit in. Once the label is wider than the bar pitch the software tilts it, and a tilted label is read more slowly than a flat one because the reader has to work at an angle. Turn the bars horizontal and every label gets a full line of reading width, sitting flat where the eye already is. Horizontal bars are not a stylistic preference past six categories, they are what happens when the labels stop fitting.
Shape three, a part against a whole: why does this shape get drawn wrongly most often?
A part against a whole means one total broken into pieces, where the claim concerns how much of the total each piece takes. A part against a whole is the shape most often drawn wrongly, and it is wrong before the drawing starts. Three conditions have to hold, and nobody checks them. The pieces must sum exactly to the whole. The pieces must be mutually exclusive, so nothing is counted in two of them. And every piece must be present, including the small ones nobody cares about.
Break any of the three and the picture lies while every number in it stays correct. Pieces that do not sum to the whole make each piece look larger than it is. Pieces that overlap double count. Pieces quietly dropped for being small make the survivors look bigger. The missing category set below is the same device.
Given the three conditions hold, a single stacked bar read against one baseline beats a circle cut into wedges, for a reason taken up again below: a reader compares lengths accurately and compares angles and areas badly. A stacked bar encodes each piece as a length along one common line. A circle asks the reader to compare wedges that share only a centre point.
Shape four, a relationship between two measures: when is a second axis warranted?
A relationship between two measures means each observation carries two numbers, and the claim concerns how one moves with the other. Each point is positioned by both measures at once, and position along a scale is the most accurately read of all the visual properties. Where the observations have no natural order, they stay unjoined and the cloud makes the claim. Where one of the two measures runs in order, joining them gives a curve rather than a cloud. The scenario exhibit in this assignment is the second kind: the horizontal measure runs from zero upward in steps, so the twenty one points join into a path.
A dual axisTwo different scales on one chart, one on each side. is warranted in one narrow case: when the reader has to see two measures in different units against the same horizontal scale, and the claim concerns the timing of their turns rather than their relative size. The moment a second vertical scale appears, the two lines stop sharing a ruler, and the reader loses the ability to compare the two heights. Anybody drawing one has to know that they have traded away comparison to buy timing, and most people who draw one have not made that trade deliberately.
Sharada Iyer wants to show that Q3 at Rs 6,40,00,000 is much larger than the other three quarters. Which shape is that, and which chart does it call for?
How does a chart built from correct numbers mislead its reader?
Choosing the chart is only half of it. The work after the choice is where a careful analyst separates from a quick one. A chart can be built from figures that are correct to the rupee, checked twice and traceable to a filing, and still leave the reader holding something false. Four devices do it, and not one of them requires a wrong number anywhere on the exhibit.
The first is the axis startThe value the vertical scale begins at, which sets how large a change looks.. Begin the vertical scale somewhere other than nothing and every bar keeps its correct value while losing its correct proportion. The second is the truncated category setShowing some of the categories and not saying which are missing., where the bars drawn are all correct and some of the bars that exist are simply not there. The third is the dual axis, where two scales in one frame can be slid against each other until the two lines cross wherever the author would like. The fourth is encodingThe visual property, length, area or position, that carries the number., where the quantity is handed to the reader as an area while the drawing was scaled by width.
The category set with members missing
Take the second device on its own. The missing category feels least like a trick while doing the most damage. A reader looking at a chart of categories makes one assumption without ever noticing they made it: that the categories shown are all the categories there are. Show three quarters out of four and the reader does not think a quarter is missing, the reader thinks the year had three quarters in it and reads all the proportions accordingly.
Including everything is sometimes genuinely impossible. The fix is to say on the chart what is not there and why. A line under the exhibit reading that the fourth quarter is excluded, and the reason, restores the reader to the position of knowing what they are looking at. Everyday version: a shop that shows three quotations and does not mention the fourth is not lying about any of the three, and the customer still walks out with a false idea of the range.
An exhibit shows the three largest quarters of the year and leaves out the fourth, and every bar drawn on it is correct to the rupee. What has to happen before it goes out?
The encoding, which is the one the reader cannot see happening
The fourth device is different in kind from the other three and deserves its own treatment. Encoding is the visual property that carries the number: a length, an area, a position along a scale, an angle, a shade. Every chart picks one, most people never notice a choice was made, and the choice decides how accurately the reader can read the picture at all. Position along a common scale is read most accurately, length against a common baseline next, and area, angle and shade a long way behind.
Here is the whole problem in one line. The eye takes in area, and area grows as the square of width. A shape drawn a certain number of times wider therefore does not look that many times bigger, it looks the square of that many times bigger. A circle scaled to twice the width to show a quantity that doubled covers four times the area. The reader reads four. Nothing on the chart is wrong and nothing on the chart says so.
Which of the four devices has the reader no way of noticing, even while checking every figure printed on the chart?
The panel below puts both encodings on screen at once so the gap can be watched opening. At a ratio of 2.00 the bar reads 2.00 times and the circle reads 4.00 times, so a doubling is shown to the reader as a quadrupling. At a ratio of 1.50 the bar reads 1.50 and the circle reads 2.25. At 3.00 the bar reads 3.00 and the circle reads 9.00. The overstatement is never fixed, it is equal to the ratio itself, so the bigger the true difference the bigger the exaggeration on top of it.
A quantity doubled, and somebody drew it as a circle twice as wide. Before the control below is moved: what multiple does the reader see?
Move the ratio between two quantities, and watch one encoding tell the truth while the other one does not.
One control moves: how many times larger the second quantity is than the first, from 1.00 up to 3.00. Both panels are then handed exactly the same pair of numbers. On the left the ratio becomes a bar length. On the right it becomes a circle scaled by width, exactly what happens when somebody drags a corner. The two readings sit on top of each other at 1.00 and separate the moment the control moves away from it.
Set the control to 2.00 and the panel returns the worked case exactly: the bar reads 2.00 times and the circle reads 4.00 times. Now walk it up to 3.00. The bar reads 3.00, still the truth, and the circle reads 9.00. The dashed outline is the width the circle should have had for its area to carry the ratio honestly, and the distance between the dashed outline and the solid edge is the whole of the deception, drawn. Length is the safe default for almost every quantity precisely because the reader cannot be wrong about it, and area is a choice that has to be earned.
Where should a vertical axis start, and when is starting elsewhere honest?
The rule has two halves and almost everybody remembers only the first one. On a chart whose claim is the size of a change, the vertical axis starts at zero, and there is no exception to that half. The reason is not etiquette. Length is the encoding, the reader compares lengths, and a length only stays proportional to a quantity if the measurement began at nothing. Cut the bottom off and the bars stop being lengths of anything at all; they become lengths of the leftover after an amount the author chose to hide.
Here is the half people forget. When the claim concerns the level rather than the size of the movement, starting somewhere other than zero is honest and often the only way to see anything. A measure that sits between 10.7 and 10.9 all year, plotted from zero, is a flat line telling the reader nothing. Plotted from 10.5 it shows what it actually did. The exhibit is then making a claim about position rather than about magnitude, the two are different claims, and the caption has to say which one it is making and where the scale begins.
So the working rule is a pair. The first move is to ask which claim the picture is making. If it claims that something moved by a lot, the axis starts at zero. If it claims that something is sitting at a particular level, the axis may start anywhere, provided the starting value is printed in the caption where the reader will find it without hunting.
The chart that was correct and misled anyway
The Kavery research desk drew operating margin for two years, and every number on it was right. The margin fell from 12.0 per cent to 10.8 per cent, a fall of 1.2 points. The frame stopped at 13 per cent. Drawn from zero, that 1.2 point fall takes up 9.2 per cent of the height of the frame, a visible decline that looks like one. Drawn from 9 per cent, the same 1.2 points takes up 30.0 per cent of the frame: three and a quarter times as much, and it reads as a collapse.
Nobody typed a wrong figure and nobody intended anything. The margins were fine. Prakash Nadar caught it in review by looking at where the scale started rather than by recomputing anything. The person who drew it was being honest, so being honest is not what protects the axis. A rule is. The rule is the one above: when the claim is the size of a change, the axis starts at zero.
When is starting a vertical axis somewhere other than zero honest rather than misleading?
What does a caption owe the reader, in exactly two lines?
A captionThe line under a chart stating in words what it shows. is not a title and it is not a label. A title names the exhibit and a caption states its content, and the difference matters more than it sounds because of what happens to an exhibit after it leaves the analyst's hands. Somebody screenshots it into an email. Somebody quotes it in a memo. Somebody reads the deck as a document six weeks after the meeting, with nobody in the room to explain it. In all three cases the picture arrives without its author.
Two lines, and the shape of them is fixed: line one says what is drawn, and line two says the one thing to notice in it. Line one carries the measure, the period and, where it matters, where the vertical scale starts. Line two carries the claim. A person quoting the exhibit in running text keeps the sentence and drops the picture, so line two is the line that survives. A caption reading Operating margin, two years is a label. The label says what is drawn, stops there, and leaves nothing behind at all when the chart is stripped away.
With the picture covered, the caption underneath reads: Operating margin, two years. Is that caption doing its job?
How is a chart tested before it goes out?
Three tests, and each one takes under a minute. Each of the three catches a fault the other two walk straight past, so run all three on every exhibit.
Test one: cover the picture and read the caption alone. If the caption still tells somebody something they did not know, the exhibit will survive being quoted. If it reads as a label, rewrite the second line before anything else.
Test two covers the caption and leaves only the picture, and what the picture appears to claim is said out loud. If that sentence is stronger than, or different from, what was meant, the picture is making a claim nobody wrote, and the usual culprit is the axis or the encoding.
Test three reads the furniture out loud. Where does the vertical scale start. What are the units and are they the same on both sides. How many categories exist and how many are drawn. A fault looked at forty times becomes invisible to the eye and stays audible to the ear, so reading it aloud matters more than it sounds.
The Kavery research desk built three charts for this assignment. Before reading on: how many of them survived review?
What happened to the three charts in this assignment?
Three charts were drawn and two went out.
| The exhibit | The shape it carries | What it became |
|---|---|---|
| Operating margin and revenue, two years | Four numbers a reader wants to read exactly | Cut, and typed as a table |
| Revenue for the four quarters | A comparison across categories | Bars from zero, and it survived |
| Margin against the share of revenue that does not repeat | A relationship between two measures, in order | A line of twenty one points, and it survived |
The one that was cut is worth a sentence on its own. The cut exhibit carried operating margin of 12.0 per cent and 10.8 per cent against revenue of Rs 18,40,00,000 and Rs 21,20,00,000. Two measures across two periods makes four numbers. Anybody reading it wants those values, not the slope between them, so it went into a table and the exhibit disappeared. Cutting a chart at review is an ordinary outcome rather than a sign that something went wrong.
The quarterly exhibit went out because the shape is the whole point and no reader needs Q2 to the rupee. The four quarters were Rs 4,80,00,000, Rs 5,20,00,000, Rs 6,40,00,000 and Rs 4,80,00,000, and they sum to Rs 21,20,00,000. The filed revenue for the year is the same figure to the rupee. A chart whose parts reconcile to a filed total is a chart a reviewer can trust without redrawing it, so the sum is worth printing under the chart.
| Quarter | Prior year | Year just reported |
|---|---|---|
| Q1, April to June | Rs 4,20,00,000 | Rs 4,80,00,000 |
| Q2, July to September | Rs 4,60,00,000 | Rs 5,20,00,000 |
| Q3, October to December | Rs 5,20,00,000 | Rs 6,40,00,000 |
| Q4, January to March | Rs 4,40,00,000 | Rs 4,80,00,000 |
| Sum of the four quarters | Rs 18,40,00,000 | Rs 21,20,00,000 |
| Revenue as filed | Rs 18,40,00,000 | Rs 21,20,00,000 |
| Difference | Nil | Nil |
The second surviving exhibit is a line, and it exists because twenty one points is a shape no table can carry. The line plots the operating margin as the share of revenue that does not repeat moves from 0 to 20 per cent. At zero it reads 10.8 per cent. At 14 per cent, the size of the single December order of Rs 2,96,80,000 sitting inside Q3, it reads 5.0 per cent. At 20 per cent it reads 1.9 per cent. The scenario line is a constructed case built to test a boundary rather than a forecast, and its caption says so in those words.
The two surviving exhibits are a pair rather than two separate items, and the pairing is the reason the quarterly chart is in the deliverable at all. The quarterly bars show where the concentration sits, and the scenario line shows what the concentration costs, so the first exhibit is the question the second one answers. Put the second one in front of a reader who has not seen the first and it looks like arithmetic about nothing.
How does a lender, an analyst or somebody working alone use any of this?
A lender reading a credit file does not admire the charts, a lender interrogates them. The first move is to find the vertical scale on every exhibit and read where it starts. Thirty seconds of work catches the most common distortion in circulation. The second is to count the categories and ask what is not drawn. A lender who does those two things before reading a single number has already filtered most of what a picture can do to them.
An analyst runs the same rules in the other direction, building rather than reading. The habit worth acquiring is to write the caption first, before drawing anything. When the second line, the one saying what to notice, can be written, the exhibit has a purpose and the chart type usually falls out of the sentence. When that line cannot be written, there is no exhibit yet, only some numbers and an intention, and drawing them will not fix it.
Somebody working alone, with no reviewer, gets more from this than a desk of four does rather than less. There is nobody to say the axis looks odd, so the three tests are the only review available, and they have to be run in earnest rather than hoped for. The household version is exactly the same discipline at a smaller scale. A household comparing quotes for a wedding hall can draw four bars from zero and see the ranking honestly, or can start the scale at the cheapest quote and turn a small difference into a chasm, and the underlying numbers are identical either way. The discipline is not about charts at all, it is about noticing that a picture is an argument and checking the argument before accepting it.
A length is a length and an area is an area wherever the exhibit is drawn, so none of these rules is specific to India or to any one market. The rules for choosing and checking a chart are universal. The conduct duties that govern how work is documented and reviewed are a separate matter, and those do vary by regime.
References
| Source | Document | Where |
|---|---|---|
| Securities and Exchange Board of India | Conduct and disclosure duties applying to registered intermediaries, including a duty not to present information in a misleading form | sebi.gov.in |
| Institute of Chartered Accountants of India | Documentation and engagement review standards, including the requirement to document work and to have it reviewed before it is issued | icai.org |
| International Organization of Securities Commissions | Published conduct principles that several national regimes draw on, covering cross-border expectations around fair presentation | iosco.org |
| Cole Nussbaumer Knaflic | Storytelling with Data, 2015, the origin of chart choice as a taught practice | Wiley |
The Kavery research desk, Kavery Capital Services Private Limited, Meenakshi Tubes Private Limited, Sharada Iyer, Prakash Nadar and Latha Menon are invented.
Educational material. Not advice on any investment, tax, budget or market position.
