Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
001Four people must cross a narrow bridge at night with one torch. At most two can be on the bridge at once, anyone crossing must carry the torch, and a pair walks at the slower person's pace. They take 1, 2, 5 and 10 minutes. What is the shortest total time to get everyone across?Belvedere TradingChicago · 2021
Try it first
Before you plan it: what is the fastest time?
Show the worked solution
17 minutes. Send 1 and 2 over (2 minutes), 1 comes back (1), 5 and 10 cross together (10), 2 comes back (2), and 1 and 2 cross again (2). The obvious plan, where the fastest person escorts each of the others, takes 19. The saving comes from putting the two slowest walkers on the bridge at the same time.
Why does the obvious plan lose two minutes?
Think of two slow parcels going by the same courier. If each travels on its own trip you pay for both trips; if they share a van, you pay for the slower one only. Every crossing costs the slower walker's time, so a slow person paired with a fast one wastes the fast one, and two slow people paired together waste nothing. Escorting with the fastest person pays 10 and then 5 as separate crossings. Pairing 5 with 10 pays 10 once and the 5 minutes are free.
Pairing the 5 and 10 minute walkers on one crossing finishes in 17 minutes, while letting the 1 minute walker escort everyone pays for the 10 and the 5 separately and finishes in 19. What is the price of pairing the slow two?
Somebody has to bring the torch back after the slow pair crosses, and it must not be one of them. So the plan first ferries two fast people over, leaves one on the far side to carry the torch back later, and spends the 2 minute walker's return trip to buy the 5 minute saving. The trade is 1 + 2 extra minutes of shuttling against 5 minutes saved on the slow side, a net gain of 2. With different speeds the trade can flip, which is the real content of the puzzle.
The relationshipa, b the two fastest times, here 1 and 2 c, d the two slowest times, here 5 and 10 What it says in wordsPairing the slow two is better exactly when twice the second fastest time is less than the fastest plus the second slowest.How do you convince the interviewer 17 cannot be beaten?
There must be at least five crossings, three over and two back, because each trip over moves at most two people and someone must return the torch. The 10 minute walker costs 10 on whatever crossing carries them. If 5 and 10 cross separately you already spend 15 on those two trips, and the three remaining crossings cost at least 1 + 1 + 2, which is 19; if they cross together, the best you can do with the other four crossings is 2 + 1 + 2 + 2. A brute force over every schedule gives the same minimum, 17 minutes.
Where candidates lose it
The strong candidate's trap is a fast answer of 19. Letting the quickest person run every errand feels efficient, and it is the right instinct for returning the torch, but it is the wrong instinct for the slow walkers.
The second loss is getting 17 by trial and error and then being unable to say why. State the principle, that the slow pair shares one crossing, and give the rule for when it wins: when twice the second fastest time is below the fastest plus the second slowest.
What the interviewer asks next
- What if the times are 1, 4, 5 and 10?
- Six people with times 1, 2, 5, 10, 20 and 25: what is the plan?
- Write the general algorithm for n people and say its running time.
Asked at Belvedere Trading, Trading, Chicago, 2021 (Wall Street Oasis):
crossing the bridge in the shortest amount of time with one flashlight brainteaser
002You are flying to a city where it rains on 25% of days. You phone three friends who live there. Each tells the truth with probability 2/3, independently of the others, and all three say it is raining. What is the probability that it is actually raining?Jane StreetNew York · 2025
Try it first
Pick your answer before working it.
Show the worked solution
8/11, about 72.7%. If it is raining, all three say yes with probability (2/3)^3 = 8/27. If it is dry, all three must be lying, (1/3)^3 = 1/27. Weight each by how often it happens: 1/4 x 8/27 against 3/4 x 1/27, which is 8 parts to 3. Three agreeing witnesses move a 25% prior a long way, but not to certainty.
Why is the answer not simply 8/9?
Picture a clinic where a test is quite reliable but the illness is uncommon. A positive result makes the illness more likely, but how much more depends on how rare it was to begin with. The friends' agreement tells you how much more likely rain makes their answer than dry does, eight times, but it does not erase the fact that dry days are three times as common. 8/9 is the answer you get if rain and dry start level. Here they do not.
Rain covers a quarter of days and all three friends say yes on 8/27 of those, while dry days cover three quarters and all three lie on only 1/27 of them, so the shaded areas stand 8 to 3 and the chance of rain given three yeses is 8/11, about 72.7%. How do you set it up so the arithmetic stays small?
Use odds rather than probabilities. Posterior odds are prior odds times the likelihood ratioHow many times more likely the evidence is if the hypothesis is true than if it is false., and both are easy numbers here. Prior odds of rain are 1 to 3. The likelihood ratio of three yeses is (2/3)^3 over (1/3)^3, which is 2 cubed, 8. So the posterior odds are 8 to 3, and the probability is 8 over 8 plus 3, 8/11. Each additional agreeing friend would double the odds again.
The relationshipR, D rain and dry YYY all three friends say yes (2/3)^3 and (1/3)^3 the chance of three yeses when it rains, and when it is dry What it says in wordsMultiply the prior odds by how much more likely the evidence is under rain, then turn the odds back into a probability.What assumption is doing the work, and should you say it?
The calculation needs the friends to lie independently. If they could be coordinating a joke, three yeses are really one piece of evidence, and the answer falls back towards the one-friend figure of 2/5. Say the independence assumption out loud, then give 8/11. Interviewers often follow up by making one friend unreliable or by letting them talk to each other.
Where candidates lose it
The most common wrong answer is 8/9: the candidate compares the chance of three truths with the chance of three lies and forgets the weather's own odds. The rain prior is a quarter, and leaving it out quietly assumes it is a coin flip.
The second loss is writing out a full Bayes formula with 27ths and 108ths and losing the thread under time pressure. Odds times likelihood ratio gets 8 to 3 in two lines and is easier to check out loud.
What the interviewer asks next
- What if only two of the three friends say yes?
- How many agreeing friends would you need before you were 95% sure it is raining?
- What changes if the friends can talk to each other before answering?
Asked at Jane Street, Generalist, New York, 2025 (Wall Street Oasis):
There was a question about the probability of rain the next day that relied on a very in depth understanding of bayes theorem
007You roll a fair six-sided die six times. What is the expected number of different faces that appear?Quant tradingProp trading firms
Try it first
Your estimate before any working?
Show the worked solution
About 3.99. Give each face an indicator that equals 1 if that face appears at least once. A given face is missed on all six rolls with probability (5/6) to the 6th, about 0.335, so it appears with probability 0.665. The expected count of distinct faces is the sum of the six indicators' expectations: 6 x 0.665 = 3.99. No case listing is needed.
Why not list the cases?
You could work out the chance of exactly one, two, up to six distinct faces and average them. It works, but it needs Stirling numbersCounts of the ways to split a set of items into a given number of non-empty groups; they appear when counting surjections. or a lot of careful counting, and it is easy to slip. Linearity of expectation lets you ignore how the faces interact: the expected total of several indicators is the sum of their expectations, whether or not they are independent. Here the six indicators are clearly dependent, since seeing many faces leaves fewer rolls for the others, and it does not matter at all.
Each of the six faces appears at least once with probability 66.5%, so the expected number of distinct faces is six times that, 3.99; the exact distribution peaks at four distinct faces and has all six only 1.5% of the time. How does the indicator trick work step by step?
Think of a teacher counting how many of six friends turn up to a party. Instead of listing every guest list, she asks for each friend separately how likely that friend is to come, then adds. Write the count as I1 + I2 + ... + I6, where I_k is 1 if face k appears; take expectations; each E[I_k] is just the probability face k appears. Face k is missed on one roll with probability 5/6, on all six with (5/6) to the 6th, 0.335. So each indicator averages 0.665 and the total averages 3.99.
The relationshipD the number of distinct faces seen in six rolls (5/6)^6 the chance a given face never appears in six rolls What it says in wordsThe expected number of distinct faces is six times the chance that any one face appears.Where does this pattern reappear?
The same shape answers how many distinct birthdays a group of n people has, how many of n hash buckets get used, and how many different stocks a random sample of trades touches. For n faces and n rolls, the expected share of faces seen is 1 - (1 - 1/n) to the n, which tends to 1 - 1/e, about 63.2%, as n grows. Six faces give 66.5%, already close. The exact enumeration of all 46,656 rolls gives a mean of 3.9906, the same number.
Where candidates lose it
The instinctive answer is 6, or something close to it, because six rolls over six faces feels like one of each. In reality repeats are the norm, and all six different faces happen in under 2% of runs.
The costlier trap is starting to enumerate cases under time pressure. Candidates who try to list exactly four distinct faces lose minutes. Say indicator variables and linearity in the first sentence.
What the interviewer asks next
- What is the expected number of faces that appear exactly once?
- How many rolls do you need, on average, to see all six faces?
- What is the variance of the number of distinct faces?
008Z1 and Z2 are independent standard normal random variables. What are the mean and variance of Z1 squared + Z2 squared, and what is the probability that it exceeds 2?Quant researchQuant trading
Try it first
What is P(Z1 squared + Z2 squared > 2)?
Show the worked solution
Mean 2, variance 4, and the probability of exceeding 2 is e to the -1, about 36.8%. Each Z squared has mean 1 and variance 2, so the sum has mean 2 and variance 4. The sum is a chi-squared with two degrees of freedom, which happens to be exactly an exponential with mean 2. Its tail beyond t is e to the -t/2, so beyond 2 it is e to the -1.
How do you get the mean and variance without the distribution?
E[Z squared] is the variance of Z, which is 1. For the variance of Z squared you need the fourth moment: E[Z to the 4] is 3 for a standard normal. So Var(Z squared) = 3 - 1 = 2, and because Z1 and Z2 are independent the variances add, giving a mean of 2 and a variance of 4. Say the fourth moment of 3 aloud; it is the number interviewers check you know, and it is why normal kurtosis is quoted as 3.
The sum of two squared standard normals has density one half times e to the minus x over 2, an exponential with mean 2 and variance 4, and the area beyond 2 is exactly e to the minus 1, about 36.8%. Why is this particular sum exponential?
Think of a dart thrown at a board where both the horizontal and vertical errors are independent standard normals. The joint density depends only on the distance from the centre, so the dart's direction is uniform and all the information is in the radiusThe distance of the point (Z1, Z2) from the origin, the square root of Z1 squared plus Z2 squared.. Switching to polar coordinates, the chance that the squared distance exceeds t is e to the -t/2, which is the tail of an exponential with mean 2. At t = 2 the answer is e to the -1.
The relationshipr the distance of (Z1, Z2) from the origin \frac{1}{2\pi} e^{-r^2/2} the joint density of two independent standard normals e^{-t/2} the tail of an exponential distribution with mean 2 What it says in wordsIn polar coordinates the angle integrates out and the radius gives an exponential tail.Why is 50% the tempting wrong answer?
Because 2 is the mean and people read the mean as the middle. For a right-skewed distribution the mean sits above the median, so less than half the mass lies beyond it; here the median is 2 ln 2, about 1.39. A simulation of 200,000 pairs gives a mean of 1.997, a variance of 3.96 and a tail share of 0.368, matching the exact results. The same polar trick is what powers the Box-Muller method for generating normal random numbers.
Where candidates lose it
The common wrong answer is 50%, from treating the mean as the median. Chi-squared variables are skewed to the right, and the skew is largest with few degrees of freedom.
The second trap is the variance. Candidates who say the variance of Z squared is 1 have confused it with the variance of Z. The fourth moment of 3 is the step, and missing it gives a variance of 2 for the sum instead of 4.
What the interviewer asks next
- What is the distribution of the square root of Z1 squared + Z2 squared?
- How would you use this to generate normal random numbers from uniforms?
- What are the mean and variance of a chi-squared with k degrees of freedom?
009A stock trades at 100 and in one period will be either 120 or 80. Interest rates are zero. Price a call option struck at 100 by building a portfolio of shares and borrowing that copies it, and explain why the real-world probability of the up move does not appear in the price.Options market makingQuant trading
Try it first
If you believe the stock goes up with probability 90%, what is the call worth?
Show the worked solution
The call is worth 10. It pays 20 if the stock goes to 120 and 0 at 80. Half a share pays 60 or 40, so half a share with a loan of 40 pays 20 or 0, exactly the call. That portfolio costs 50 - 40 = 10 today. If the call traded at any other price, you could buy the cheap one and sell the dear one for a riskless profit, so no probability is needed.
How do you build the copy?
Match the swing first. The call's payoff moves by 20 between the two states while the stock moves by 40, so the copy needs 20/40 = 0.5 of a share: that ratio is the option's deltaHow much an option's value changes for a one-unit change in the underlying price; here, the number of shares that copies the option.. Half a share is worth 60 or 40 at the end, which is 40 more than the call in both states. Borrow 40 today, repay 40 at the end with zero interest, and the copy pays exactly 20 or 0.
A call struck at 100 on a stock that moves to 120 or 80 is copied by half a share and a loan of 40, which pays 20 or 0 exactly as the call does and costs 10 today, so the call is worth 10 with no probability used. Why does the chance of the up move not matter?
Think of a shop selling a bundle of two items that you can also buy separately. The bundle's price is pinned by the parts, whatever you think about how useful the items are. The call is a bundle of half a share and a loan; the share price already reflects everyone's views about the up move, so the option inherits them and adds none of its own. A 90% view is a reason to hold the stock itself, not a reason to pay more for the call than its parts cost.
The relationship\Delta shares held in the copy B the amount borrowed, 0.5 x 80 - 0 = 40 q the risk-neutral weight on the up state, fixed by the prices, not by beliefs What it says in wordsThe copy's cost gives the price, and the same price is an average of the payoffs using weights set by today's stock price.What would you do if the call traded at 12?
Sell the dear thing and buy the cheap one. Sell the call for 12, buy half a share for 50 and borrow 40, a net cash inflow of 2 today; at the end the portfolio pays exactly what you owe on the call in either state. The 2 is kept whatever happens. The weight q = 0.5 that reproduces the price is called the risk-neutral probability, but it is a pricing weight backed out of the stock price, not a forecast. Say that distinction; interviewers listen for it.
Where candidates lose it
The trap is pricing the call as an expected payoff under your own view: 90% of 20 is 18. That price can be arbitraged against the stock, so nobody could trade it for long, and the interviewer wants to hear that the copying portfolio pins the price.
The second loss is getting 10 by assuming a 50% chance. The number is right by coincidence of the symmetric tree; ask yourself what happens with an up move to 130, and the risk-neutral weight changes to 1/2.5 = 0.4.
What the interviewer asks next
- Price the put struck at 100 and check put-call parity.
- What changes if interest rates are 5% for the period?
- The stock can go to 130 or 80 instead. Price the call again.
011An urn starts with one red ball and one blue ball. Each turn you draw a ball at random, put it back, and add another ball of the same colour. After ten draws, what is the probability that exactly five of the draws were red?Quant researchQuant trading
Try it first
Your first instinct for P(exactly five reds)?
Show the worked solution
1/11, about 9.1%. Any particular sequence with five reds and five blues has probability 5! x 5! / 11!, whatever the order, because the numerators just count up the reds and blues separately. There are 10 choose 5 = 252 such sequences, and 252 x 5! x 5! / 11! = 1/11. The same algebra gives 1/11 for every count from 0 to 10.
Why is this not ten coin flips?
Think of a new café and its first customers. If the first few visitors like it and bring friends, it fills up; if they do not, it stays empty. Early luck compounds. In this urn each red draw adds a red ball, so it raises the chance of the next red: the draws reinforce each other, and runs that start lopsided tend to stay lopsided. That spreads the count of reds far more widely than independent coin flips, which bunch around five.
After ten draws from an urn that starts with one red and one blue ball and adds a ball of the drawn colour each time, every count of reds from 0 to 10 has probability 1/11, a flat line, while ten fair coin flips would pile up at five with 24.6%. Why does the order of the draws not matter?
Write out one sequence, say five reds then five blues. The chances are 1/2, 2/3, 3/4, 4/5, 5/6 for the reds, then 1/7, 2/8, 3/9, 4/10, 5/11 for the blues. Shuffle the order and the denominators are still 2 up to 11, while the red numerators still run 1 to 5 and the blue ones 1 to 5, so every arrangement of five reds and five blues has the same probability, 5! x 5! / 11!. Draws whose joint probability ignores order are called exchangeableA sequence of random variables whose joint distribution does not change when you reorder them, even though they need not be independent., and that is the property doing the work.
The relationship\binom{n}{k} the number of orders in which k reds can appear k!(n-k)!/(n+1)! the probability of any one such order What it says in wordsThe number of orders times the probability of each order is always one over n plus one.Is there a picture that makes 1/11 obvious?
Yes. This urn behaves exactly as if nature first picked a hidden red probability p uniformly between 0 and 1, and then flipped ten independent coins with that p. Averaging a binomial over a uniform p gives every count the same weight, 1/(n + 1). The urn is also a small model of momentum and of market share: an early lead makes further gains more likely, so the final split is highly uncertain, even though each step looks like a fair draw.
Where candidates lose it
The fast wrong answer is 252/1024, about 24.6%, from treating the draws as independent coin flips. The reinforcement is the whole question, and it spreads the outcomes out rather than pulling them to the middle.
The second trap is trying to add up paths through a tree of ten levels. Spot that every order of a given mix has the same probability, and the problem collapses to one line.
What the interviewer asks next
- What is the probability the eleventh draw is red, given that five of the first ten were red?
- What if the urn starts with two red and one blue ball?
- As the number of draws grows, what does the fraction of red balls converge to?
013Users join a server at times 1, 2, 4, 5 and 7 and leave at times 7, 3, 8, 9 and 10 respectively. A leave at the same moment as a join is processed first. What is the maximum number of users online at once, and how would you compute it efficiently for a million users?Two SigmaNew York · 2025
Try it first
What is the peak number of users online together?
Show the worked solution
The peak is 3 users. Turn every join into a +1 event and every leave into a -1 event, sort all ten events by time with leaves before joins at equal times, and keep a running total. It goes 1, 2, 1, 2, 3, then at time 7 down to 2 and back to 3, then 2, 1, 0. Sorting costs n log n, and the sweep itself is linear.
Why not check every moment in time?
A shopkeeper who wants to know the busiest moment of the day does not count heads every second; they note each time the door opens in or out and keep a tally. The count of users can change only at a join or a leave, so the maximum must occur just after some join, and you only need to look at the 2n event times. Checking every time step costs time proportional to the length of the day, and comparing every pair of users costs n squared; both are far too slow at a million users.
Each join adds one user and each leave removes one, so a running total over the sorted events finds the peak of 3; processing the join at time 7 before the leave would produce a false peak of 4. Why does the tie rule matter so much?
At time 7 user 1 leaves and user 5 joins. If a user's session is taken to end just before the moment they leave, then a leave and a join at the same instant never overlap, and the leave must be sorted first; sorting the other way invents a user who was never there. In code this is one comparison in the sort key, and it is exactly the detail interviewers use to separate a working answer from a nearly working one. Ask which convention applies before writing any code.
The relationshipd_i +1 for a join and -1 for a leave (t_i, d_i) the sort key: time first, then leaves (-1) before joins (+1) What it says in wordsSort the events so leaves come first at equal times, then the peak is the largest running total.Is there a version that avoids building the event list?
Sort the join times and the leave times separately and walk two pointers through them. At each step take the earlier of the next join and the next leave, taking the leave on a tie, and adjust the count; this is the same sweep without allocating 2n tuples. If times are small integers you can go further: add +1 and -1 into an array indexed by time and take a running sum, which is linear. Mention both and say which you would use for a million users with timestamps in milliseconds.
Where candidates lose it
The trap is the tie. Many candidates write a correct sweep, sort by time alone and get 4, because the join at time 7 is processed before the leave. The question states the convention precisely to see whether you use it.
The second loss is proposing a double loop that checks every pair of sessions. It gives the right answer on five users and fails the question, which asked how you would do it efficiently.
What the interviewer asks next
- Return the time interval during which the peak occurs, not just the count.
- Users arrive as a stream and you must report the current count at any moment. What data structure do you use?
- How many servers are needed if each can hold at most two users at once?
Asked at Two Sigma, Equity Hedge, New York, 2025 (Wall Street Oasis):
Given arrays (start & end) of the times users join and leave a server, find the max number of concurrent users on the server
014Five per cent of fund managers are skilled and beat the market in any given year with probability 60%; the rest are unskilled and beat it with probability 50%. Years are independent. A manager has beaten the market in exactly 8 of the last 10 years. What is the probability the manager is skilled?Citadel SecuritiesMiami · 2022
Try it first
Roughly how likely is it that this manager is skilled?
Show the worked solution
About 12.7%. A skilled manager wins exactly 8 of 10 with probability 0.1209; an unskilled one with 0.0439, a likelihood ratio of about 2.75. Prior odds of skill are 5 to 95, 1 to 19. Posterior odds are 2.75 to 19, so the probability is 0.05 x 0.1209 / (0.05 x 0.1209 + 0.95 x 0.0439) = 0.127. The record helps, but luck has far more players.
Why does an impressive record move the needle so little?
Imagine a thousand people each tossing a coin ten times. About 55 of them will get eight heads or better with a fair coin. If a handful of the thousand had slightly biased coins, you still could not pick them out from the lucky crowd by one run of ten. Evidence moves a belief in proportion to how much more likely it is under one explanation than the other, and 8 wins in 10 is not much more likely from a 60% manager than from a 50% one. The ratio is about 2.75.
Eight wins in ten years has probability 12.1% for a 60% manager and 4.4% for a 50% manager, but after weighting by how common each type is, 5% against 95%, the chance that an eight-win manager is skilled is only 12.7%. How do you set it up quickly?
Use odds. Prior odds of skill are 1 to 19; the likelihood ratioHow many times more likely the evidence is under one hypothesis than under the other. of the record is (0.6/0.5) to the 8 times (0.4/0.5) squared, which is 1.2 to the 8 times 0.64, about 2.75; multiply to get posterior odds of about 0.145. Converting, 2.75 over 2.75 + 19 is 12.7%. The binomial coefficient, 45, is the same in both likelihoods and cancels, so you never need it.
The relationshipS, U skilled and unskilled 0.6^8 0.4^2 the chance of one particular sequence of 8 wins and 2 losses for a skilled manager 0.5^{10} the same for an unskilled manager What it says in wordsPrior odds of 1 to 19, times a likelihood ratio of 2.75, give a posterior of about 12.7%.What does this say about picking managers?
When skill is rare and its edge is small, even a long, strong track record leaves luck as the likelier explanation. Using 8 or more wins instead of exactly 8 barely changes things: the answer becomes 13.9%. The honest limitation is that the model is stylised: real skill is not a fixed 60%, and survivorship means the managers you hear about were already filtered for good records, which pushes the true figure lower still.
Where candidates lose it
The common answer is around 80%, reading the record's win rate as the chance of skill. That skips the prior entirely, and with only 5% of managers skilled, the prior dominates.
The quieter trap is computing the full binomial probabilities, 45 x 0.6 to the 8 x 0.4 squared and so on, and getting lost in decimals. The coefficient cancels. Say odds and likelihood ratio and the arithmetic stays on one line.
What the interviewer asks next
- How many years of 80% wins would you need before the manager is more likely skilled than not?
- What if 20% of managers were skilled?
- How does survivorship bias change the answer if you only ever see managers with good records?
Asked at Citadel Securities, Sales and Trading, Miami, 2022 (Wall Street Oasis):
I got a question about Bayes' theorem applied to a practical scenario, which I handled decently
015You roll three fair dice. Is a total of 9 or a total of 10 more likely, given that each can be written as exactly six unordered combinations of three faces?Quant tradingProp trading firms
Try it first
Which total is more likely?
Show the worked solution
10 is more likely: 27 ways out of 216 against 25. The equally likely outcomes are the 216 ordered rolls, not the unordered combinations. A combination of three different faces covers six ordered rolls, one with a pair covers three, and a triple covers one. Total 9 includes 3 + 3 + 3, which counts once, and one fewer all-different combination, so it loses two orderings to 10.
Why are combinations the wrong thing to count?
Think of dealing two cards and asking whether a pair of kings or a king with a queen is more likely. There is one combination of each, but king and queen can arrive in either order while two kings are just two kings. Probability comes from counting outcomes that are equally likely, and with dice those are the ordered rolls: first die, second die, third die, 6 x 6 x 6 = 216 of them. Unordered combinations bundle different numbers of those outcomes, so counting combinations gives the wrong weights.
Totals of 9 and 10 each have six unordered combinations, but weighting each combination by its number of orderings gives 27 of 216 rolls for 10 against 25 for 9, mostly because 9 includes the triple 3, 3, 3, which happens only one way. How do you count the orderings fast?
Classify each combination by its repeats. Three different faces give 3! = 6 orders, a pair gives 3 orders, one for each position of the odd die, and a triple gives 1. For 9: 1 2 6, 1 3 5 and 2 3 4 are all different, 18; 1 4 4 and 2 2 5 are pairs, 6; 3 3 3 is a triple, 1; total 25. For 10: 1 3 6, 1 4 5 and 2 3 5 give 18; 2 2 6, 2 4 4 and 3 3 4 give 9; total 27. So 10 comes up 12.5% of the time and 9 only 11.6%.
The relationship6, 3, 1 the orderings of an all-different, a pair and a triple combination 216 the ordered outcomes of three dice What it says in wordsWeight each combination by its orderings and 10 beats 9 by two rolls in 216.Is there a shortcut that avoids listing?
Yes: symmetry. Replacing each face x by 7 - x maps a total of t to 21 - t, so the distribution of three dice is symmetric about 10.5, and 10 and 11 are the two most likely totals, each 27/216. Anything further from 10.5, including 9, must be less likely or equal; a quick count confirms it is 25. Historically this is the question gamblers put to Galileo, who answered it by counting ordered outcomes, which is still the method.
Where candidates lose it
The trap is the question's own framing: six combinations each invites the answer that the totals are equally likely. The interviewer wants you to reject the framing, not accept it.
The second loss is listing all 216 rolls, or writing out every ordering. Classifying by repeats, six, three or one, gets both totals in under a minute.
What the interviewer asks next
- What is the most likely total with four dice, and its probability?
- What is P(total is 9) with two dice, and why does 9 behave differently?
- How many ordered outcomes of three dice sum to 7?
018I will draw a card from a shuffled deck. You may pay Rs 6 to play a bet that pays Rs 10 if the card is red. Before deciding, you may pay to be told the card's colour. What is the most you should pay for that information?OptiverChicago · 2025
Try it first
What is the information worth?
Show the worked solution
Rs 2. Without information the bet is worth 0.5 x 10 - 6 = -1, so you decline and your value is 0. With the colour known, you play on red and make 4, and skip black and make 0, which averages 2. Information is worth the improvement in your best decision: 2 - 0 = 2. If it would not change what you do, it is worth nothing.
How do you value a piece of information?
Suppose a weather forecast costs money and you are deciding whether to carry an umbrella. If you would carry it anyway, the forecast is worthless to you; it is valuable only if some answer would change what you do. The value of information is the expected value of your best decision with it, minus the expected value of your best decision without it. Work out both decision trees separately and subtract. Never value information by the size of the payout it relates to.
Blind, the bet has an expected value of minus 1 so you decline and get 0; told the colour first, you play only on red and make 4 half the time, an average of 2, so the information is worth Rs 2. Why is it not worth Rs 4 or Rs 5?
Rs 4 is what you make when the card is red, but it is red only half the time. Rs 5 is half the payout, which ignores the Rs 6 you pay to play. The information saves you from the losing half of the bet and lets you keep the winning half, and that is worth half of Rs 4, which is Rs 2. Pay more than Rs 2 and you would do better declining the offer of information and declining the bet.
The relationshippayoff 10 - 6 = 4 on red, -6 on black E[max(payoff, 0)] your value when you can choose after seeing the colour max(E[payoff], 0) your value when you must choose blind What it says in wordsInformation is worth the gap between deciding after you know and deciding before.When is information worth the most?
Vary the price of the bet. At a price of 5 you are exactly indifferent blind, and the information is worth 2.50, its maximum; at a price of 0 you would always play, and it is worth 0. Information is valuable when you are close to indifferent and the decision could go either way. The formula also has the shape of an option payoff: knowing first lets you exercise only when it pays, which is why traders talk about paying for optionality and paying for information in the same breath.
Where candidates lose it
The trap is answering with the size of the win, Rs 4, or half the payout, Rs 5. Both value the information by the bet it is about, not by the decision it improves.
The second loss is forgetting that without information you would decline. Candidates who compare with playing blind, at -1, get 3. The comparison is always with your best action without the information, which here is to walk away.
What the interviewer asks next
- What is the information worth if the bet costs Rs 3?
- What if the information is only 80% reliable?
- You can pay to see one card of a two-card hand before betting. How do you decide what that is worth?
Asked at Optiver, Quantitative Research, Chicago, 2025 (Wall Street Oasis):
Valuing information, taking directional bets when not plus EV.

