Bounded Rationality: Deciding Well Enough With Limited Capacity
Bounded rationality is deciding as well as possible with limited attention, information and time, and it is not a weaker grade of rationality. Herbert Simon argued that the costs of deciding are real costs, so stopping once an option is good enough is the correct response to them rather than a lapse. The shortcuts this forces are the subject of everything that follows.
The clearest starting point has nothing to do with money. A shopper is at a vegetable market with about forty stalls and wants tomatoes. In principle the whole row could be walked, every stall priced, quality weighed against price at each one, and the best stall chosen. Nobody has ever done this. The shopper checks two or three stalls, finds one that is fine at a price that is fine, and buys. The shopper who priced all forty went home with slightly better tomatoes and no afternoon left. The arithmetic that matters includes what the searching itself costs, and the forty-stall shopper left that term out. Bounded rationality is the claim that the shopper who stopped at three was the one doing the arithmetic properly.
The shopper who stopped at three is the whole idea. The phrase gets used carelessly, so its content is worth stating precisely. Bounded rationality is not saying people are stupid, or lazy, or emotional, or in need of correction. Bounded rationality says that the standard account of a good decision was written without any accounting for what reaching it costs, and that once that cost is put back in, a large amount of behaviour that looked like carelessness turns out to be arithmetic.
What did Simon actually claim?
Herbert Simon set the claim out in a paper called A Behavioral Model of Rational Choice, published in the Quarterly Journal of Economics in 1955. The account he was arguing with had been given its cleanest statement by John von Neumann and Oskar Morgenstern in Theory of Games and Economic Behavior in 1944: a decider who can hold every option in view, evaluate each one against a consistent set of preferences, and select the best available.
Simon did not say that people fall short of this. He said something more awkward. He said the description had been written as though deciding were free, and that it is not. Three things are always finite. Information has to be gone and got, so the decider knows only part of what the options hold. Working out the consequences of a choice takes effort that runs out, so the decider can compute only so much. And the time available to decide ends, usually before the options do. Simon's objection was not that people are worse than the ideal decider; it was that the ideal decider was described without any account of what reaching that standard costs.
Put that way, the correction is small and the consequences are enormous. If deciding is free, then more deciding is always at least as good as less, and the only sensible amount of search is all of it. If deciding costs something, the sensible amount of search is finite, and finding it becomes a calculation in its own right. Bounded rationality is what follows from moving that one term from zero to something above zero.
The phrase is often used loosely to mean that human reasoning is a bit dim. The loose use is close to the opposite of what the phrase says. Bounded rationalityDeciding as well as possible given limited attention, information and time. describes a decider who is optimising, but optimising against a problem that includes the cost of the optimising. A person who spends four hours choosing between two schemes that differ by a small amount has not been more rational than a person who spent ten minutes. The four-hour chooser has paid four hours for a difference worth less than four hours.
Is bounded rationality a kind of irrationality?
How does satisficing differ from maximising?
Two people can face the same list of options, want exactly the same thing, and reason without a single error, and still behave completely differently. The only thing that separates them is the rule they use to decide when to stop looking.
MaximisingContinuing the search until the best available option has been identified. stops when every option has been examined. Only then can the best one be known with certainty. SatisficingStopping the search at the first option that is good enough, by a standard fixed in advance. stops at the first option that clears a standard fixed before the search began. Simon built the word from satisfy and suffice, and it is not a synonym for compromising. A satisficer has a standard and refuses everything below it. A satisficer will not keep looking after the standard has been met.
Notice what these two rules have in common. The shared ground is more than the difference. Both deciders want the highest value they can get. Both are consistent. Neither is confused about what they are looking for. The satisficer and the maximiser differ in one place only: the instruction that tells them when to stop. Every difference in what happens to them comes out of that single line.
Here is the part that surprises people. On the arithmetic worked out below, the maximiser genuinely finds the better option and still ends up worse off by a wide margin. The extra looking cost more than the better option was worth. The reversal is not a trick, and it does not depend on the maximiser being slow or clumsy. Setting a cost everybody knows exists beside the benefit it buys is all it takes.
Where does an aspiration level come from?
A stopping rule of the satisficing kind needs a number to compare against, and Simon called that number the aspiration levelThe standard an option has to clear for the search to stop.. The aspiration level is the standard an option has to clear for the looking to end. Everything about how satisficing behaves depends on where this number comes from, and the answer is that it comes from outside the current search.
The level comes from what was obtained last time. The level comes from what people nearby appear to have obtained. The level comes from what the situation has taught the decider is available. A household that has always paid about Rs 12,000/- a month in rent has an aspiration level for rent, and it did not arrive by reasoning; it arrived by living in that market. The level also moves. Cleared easily three times, it drifts up. Left uncleared for a month, it drifts down. A person looking for work eventually accepts something they would have refused at the start. The drift is not weakness. The standard is being corrected by evidence about what the environment actually contains.
A standard that updates to whatever was just looked at can never be cleared, and the search then has no stopping point at all. The one thing an aspiration level cannot be, therefore, is a running comparison with the best option seen so far. This is the difference between a standard and a preference, and it is where satisficing quietly succeeds or fails in practice. Set the number first and the search terminates. Discover the number while looking and every option becomes a comparison with the last one. The shopper who does that never leaves the market.
A standard implies a length of search whether or not anybody computes it. The level is where the arithmetic below meets ordinary life. If options really are spread evenly between 0 and 100, then a standard of 80 is cleared by one option in five, so the search runs five options long on average before it stops. Five options is an average, not a promise about any single afternoon, and it sits one step below the optimum computed below. A person who has never seen the arithmetic can land near its answer by holding a sensible standard.
Where does an aspiration level have to be set for satisficing to work at all?
Why does the second look add less than the first one did?
Searching costs something, and the cost per look does not fall. Reading one more scheme document takes about as long as reading the first one. Visiting one more stall takes about as long as visiting the first. Phoning one more lender, checking one more quotation, sitting through one more meeting: the charge is roughly flat. Call this the search costWhat it costs in time and attention to examine one more option. and hold it constant. In most real searching it very nearly is.
The return on a look is a different matter entirely. Before the first look there was nothing, so the first option examined improves the position enormously. The second can only help if it beats the first, and it does so about half the time. The tenth can only help if it beats the best of nine, and it does so about one time in ten. The gain from one more look shrinks as the best so far rises. The cost of that look stays exactly where it was, and a shrinking gain set against a flat charge produces a peak every single time. This is diminishing returnsEach extra unit of effort adding less than the one before it did. meeting a constant charge, and the meeting point is the whole of Simon's argument in one line.
Work the shrinking out. If the options are spread evenly between 0 and 100, examining a number of them and keeping the best returns on average 100 multiplied by that number, divided by that number plus one. One option returns 50.0 on average. Two return 66.7. Three return 75.0. So the second look added 16.7 and the third added 8.3. Carry it on and the additions get small very quickly. The charge of 2 points a look never moves.
| Which look | What it adds on average | Adds | Costs | Worth taking? |
|---|---|---|---|---|
| the fourth | only helps if it beats the best of three | 5.00 | 2.00 | yes |
| the fifth | only helps if it beats the best of four | 3.33 | 2.00 | yes |
| the sixth | only helps if it beats the best of five | 2.38 | 2.00 | yes, only just |
| the seventh | only helps if it beats the best of six | 1.79 | 2.00 | no |
| the eighth | only helps if it beats the best of seven | 1.39 | 2.00 | no |
Read the table as a stopping rule and it answers the question on its own. Take a look while the look adds more than it charges, and stop when it does not. The sixth look adds 2.38 against a charge of 2.00 and is worth taking. The seventh adds 1.79 against the same 2.00 and is not. The rule is optimal stoppingThe point at which one more examination costs more than it is expected to add. in one paragraph, and it lands on six.
Why does the best-found curve flatten out while the cost of searching does not?
How is the point at which searching should stop computed?
The two halves put together give the answer directly. The assumption is that the options are worth something between 0 and 100 and that which is which cannot be told without looking. Examining a number of them and keeping the best leaves, on average, 100 multiplied by that number and divided by that number plus one. Every examination costs 2 points. The net is the first figure less the second, and the net is what the searcher actually walks away with. No other figure matters.
| Options examined | Best found, out of 100 | Paid for the searching | Net value |
|---|---|---|---|
| one | 50.0 | 2.0 | 48.0 |
| three | 75.0 | 6.0 | 69.0 |
| five | 83.3 | 10.0 | 73.3 |
| six, the highest net available | 85.7 | 12.0 | 73.7 |
| seven | 87.5 | 14.0 | 73.5 |
| ten | 90.9 | 20.0 | 70.9 |
| thirty | 96.8 | 60.0 | 36.8 |
Read the last column downwards. The net climbs, turns over at six, then falls, and it keeps falling for the rest of the list. Six is the whole answer, and six is a startlingly small number. Almost nobody guesses it before they see the arithmetic. Examining thirty options leaves a net of 36.8, which is worse than examining three, and very much worse than examining six. The top is also nearly flat: five gives 73.3 and seven gives 73.5, so being one out either way costs almost nothing. Being twenty out costs half of everything.
The assumptions are obviously simplifications, and none of the argument depends on them. Options are not spread evenly between 0 and 100 in real life, and looking at one does not cost exactly 2 points. Change either and the numbers move. The shape does not move: a gain that shrinks with each look, set against a charge that does not, always climbs to a peak and then falls away. Make searching cheaper and the peak moves right. Make it dearer and the peak moves left. The peak never disappears.
Examining one option costs 2 points and options are worth up to 100. Before the control below is touched, how many options should be examined?
Move the search longer and watch the net value turn over
One variable moves: how many options get examined, from 1 to 30. One consequence: the net value of the search, meaning the best option found less what all that looking cost. Examining one option leaves a best found of 50.0 on average, for a net of 48.0. Examining six leaves a best found of 85.7 for a net of 73.7, the highest net value anywhere on the scale. Examining all thirty leaves a best found of 96.8, the highest such figure shown anywhere, for a net of 36.8, the lowest. The peak sits at six options.
Examining 6 options, the best one found is worth 85.7 and the searching has cost 12.0, so the net is 73.7. That is the highest net value available anywhere on this scale, because the sixth look added 2.38 against a charge of 2.00 and a seventh would add only 1.79.
What happens when a long list is put in front of real people?
The arithmetic above is worked from assumed numbers, and assumed numbers prove nothing about how anybody behaves. So set a measurement beside it. In the invented Palash decision log, thirty readers were shown a list of 214 options and asked to pick one. Eleven of them picked anything at all, or 36.7 per cent. A second group of thirty was shown a shortlist of 7 and asked the same question. Twenty one picked, or 70.0 per cent. The options were the same options and the people were drawn the same way. The only thing that moved was how long the list was.
The measurement shows that searching costs something real enough that people stop paying it, and the cost shows up not as a worse choice but as no choice at all. That is the observation the argument needs. Nineteen of the first thirty walked away with nothing, not because they were confused about what they wanted, but because working through 214 options at any honest rate of attention is a job, and the job was worth less to them than the afternoon it would have taken.
The limits of those two numbers matter more than the finding. Nothing in them says the choices made from the shortlist were better ones. Completing a choice and making a good choice are different measurements, and only the first one was taken. Nothing in them says a shorter list is a better list either, or that anybody should be offered seven of anything. How many options ought to be put in front of a person is a separate question with a separate literature. The measurement establishes one thing: that the cost of searching is not a theoretical device invented to make the arithmetic tidy.
One coincidence is worth noticing, and then worth refusing to make anything of. The arithmetic put the optimum at six options. The shortlist that nearly doubled the completion rate had seven. The match is a coincidence, one measurement on one invented cohort, and two numbers landing next to each other is not evidence of anything whatsoever. A reader who spots the match should hear it named as a coincidence rather than quietly file it as a finding.
Shown 214 options, 11 of 30 chose. Shown 7, 21 of 30 chose. What does that measure?
The error that gets made, and what it costs
The error is hearing satisficing and thinking settling. The maximiser sounds more serious. Examining everything sounds like diligence, and stopping at the first acceptable option sounds like not being bothered, so the two get ranked by how much effort they display rather than by what they return.
Rank them by what they found and the maximiser wins: 96.8 against 85.7, and the maximiser genuinely did find the better option. Nothing in the setup is rigged, and that is what makes the case interesting. Rank them by what each one had left after paying for the search and the order reverses completely: the satisficer keeps 73.7 and the maximiser keeps 36.8, less than half. The entire gap is search cost.
The searching cost more than the improvement it bought. Calling the satisficer lazy requires ignoring the only line item that separates the two of them, so that one sentence is both the whole of Simon's argument and the whole of the error. The error costs a reversed judgment: the person being praised has done worse, the person being criticised has done better, and nobody involved has made a mistake in reasoning.
The maximiser found 96.8 and the satisficer found 85.7. Who ended better off?
What are the two blades of Simon's scissors?
Simon gave the idea its best image in a second paper, Rational Choice and the Structure of the Environment, published in Psychological Review in 1956. Behaviour, he said, is shaped like something cut by a pair of scissors. One blade is what a mind can do: how much it can attend to, how much it can hold, how long it has. The other blade is how the world is arranged: how many options there are, how they are laid out, how much time the situation allows. One blade alone cannot explain the cut.
A shortcut is never good or bad on its own; it is good or bad in a setting, and naming the setting is half of every honest explanation of behaviour. A rule that works beautifully when options are few and roughly comparable can produce a bad result when options are many and dressed up to look alike. Nothing about the mind changed between those two cases. The other blade moved.
Take it out of finance again. A person crossing a quiet lane judges the gap by eye and is right every time for forty years. The same eye, the same rule, on a road where vehicles arrive at four times the speed, is wrong. The person did not become a worse judge of gaps. The environment changed, and the rule was tuned for the old one. The scissors is why bounded rationality is not a list of the ways people are deficient.
What are the two blades of Simon's scissors?
Is any of this a criticism of the person deciding?
No, and the reason is worth stating bluntly. A bounded decider is not a defective unbounded decider. The bounded problem includes the cost of solving it and the unbounded problem quietly assumed that cost away, so the bounded decider is solving the harder problem of the two.
Run the comparison honestly. The unbounded decider examines everything and pays nothing for doing it. The unbounded decider is not a higher standard that people fail to reach, but a description of somebody living under different physics. Judged by what they walk away with, a person who stops at six on the arithmetic above beats a person who examines all thirty by almost double. The stopping is the intelligent part, not the lapse, and any account that reads bounded rationality as a shortfall has inverted the finding it reports.
The shortcuts and the errors they produce are set out under heuristics and biases, and a list of named errors is very easy to read as a catalogue of human failure. A catalogue of failure is the wrong reading. Every shortcut named there exists because deciding is expensive, and every one of them was worth having in the setting it was tuned for. The errors are what happens when the second blade moves and the rule does not.
What does bounded rationality predict that the unbounded account does not?
A relaxed assumption earns its place by predicting something the original could not. Bounded rationality earns its place several times over, and its predictions are the kind that can be checked rather than admired.
| What is observed | The unbounded account says | Bounded rationality says |
|---|---|---|
| Search stops while options remain | it should not, since examining more is free | it should, at a point that can be computed |
| Two people, same options, different lengths of search | nothing, since both should examine everything | their standards differ, so their stopping points do |
| Making a decision easier lengthens the search | nothing, since ease does not enter | a cheaper look moves the peak to the right |
| People given a very long list often choose nothing | nothing, since more options can only help | the search is worth less than it costs, so it is not begun |
| People develop rules of thumb and keep using them | nothing, since exact evaluation is free | a cheap rule that is usually right is valuable |
The last row of that table is the bridge from bounded rationality to everything else in behavioural finance. Once deciding is expensive, a rule that is quick and usually right becomes genuinely valuable, and people will find such rules and keep them. In a world where evaluating is free there is nothing for a shortcut to save, so the unbounded account cannot predict that at all.
How this gets used, by an adviser and by a person deciding alone
Devika Rao, the adviser at the invented Palash Advisory Services Private Limited, does not hand anybody a list of everything available. A complete list would be useless, and the completion figures above say what would happen to it. Her method has two steps, and the two steps are the two halves of the argument. First, fix the standard before anything is shown: what does this money have to do, by when, and what would count as good enough. Second, present a short set that already clears the standard, and say so out loud. The person is then choosing among acceptable options rather than searching for an acceptable one.
A person deciding alone, with no adviser and no committee, runs the identical two steps unaided. Meera Sundaram puts Rs 25,000/- a month in by standing instruction, and the instruction is doing exactly this work: the standard was set once and the searching does not have to be redone every month. Written down before the looking starts, a standard converts an endless comparison into a decision with an ending, and that is the practical whole of satisficing. The failure mode to watch for is the one in the second panel of the flow above: a standard that quietly becomes whatever was last looked at.
If capacity is limited and searching is costly, what should people be expected to develop?
Where does the argument stop, and what picks it up?
The argument stops at the exact point where limited capacity forces a shortcut. Deciding costs something, the cost makes the sensible amount of searching finite, the finite amount can be computed and is smaller than anybody guesses, and stopping there is the correct answer rather than a lapse.
The shortcuts themselves are a subject of their own, set out under heuristics and biases. The programme that named and measured them was opened by Amos Tversky and Daniel Kahneman in Judgment under Uncertainty: Heuristics and Biases, published in Science in 1974, and everything in it rests on the argument set out above: shortcuts exist because deciding is expensive, and they produce errors because the environment they were tuned for is not always the one they get used in. Read in that order, the errors that follow are the price of a sensible economy rather than evidence of a deficient mind.
Sources
| Source | Document | Site |
|---|---|---|
| Herbert Simon | A Behavioral Model of Rational Choice, Quarterly Journal of Economics, 1955, where bounded rationality is first set out | ssrn.com |
| Herbert Simon | Rational Choice and the Structure of the Environment, Psychological Review, 1956, where the scissors image is introduced | ssrn.com |
| John von Neumann and Oskar Morgenstern | Theory of Games and Economic Behavior, 1944, the cleanest statement of the unbounded account of choice | Princeton University Press |
| Amos Tversky and Daniel Kahneman | Judgment under Uncertainty: Heuristics and Biases, Science, 1974, which opened the programme on named shortcuts | ssrn.com |
Meera Sundaram, Devika Rao, Palash Advisory Services Private Limited and the Palash decision log are invented.
Educational material. Not advice on any investment, tax, budget or market position.
