Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
001How many golf balls fit inside a 100-storey office tower? Talk me through your thought process; I care more about the structure than the final number.Tower Research CapitalNew York · 2013
Try it first
Before any arithmetic: which single assumption will move your answer the most, once the building's size is fixed?
Show the worked solution
About 12 billion, with an honest range of 10 to 14 billion. A floor plate of 4,000 square metres at 3.8 metres a floor gives 1.52 million cubic metres over 100 floors; take half as usable, 0.76 million. A golf ball is 4.27 centimetres across, about 41 cubic centimetres, so 24,500 would fit in a solid cubic metre. Spheres leave gaps: at a packing factor of 0.64 that is 15,700 per cubic metre, and 0.76 million times 15,700 is 11.9 billion.
Why build a chain instead of guessing a number?
If someone asks how much rice a storeroom holds, you do not guess a tonnage. You measure the room, decide how much of it is shelving, and know how much a sack holds. An estimation question is marked on whether each link in the chain is stated, sized and defensible, not on the final figure. Here the chain is floors, floor plate, floor height, usable fraction, ball volume, packing factor. Say the six links before you fill any of them in, so the interviewer can follow and correct a single link rather than the whole answer.
The chain runs from 100 floors, a 4,000 square metre plate and 3.8 metre floors to 1.52 million gross cubic metres and 0.76 million usable, then from a 40.8 cubic centimetre ball to 24,531 balls per solid cubic metre, and the packing factor turns that into 10.4, 11.9 or 13.8 billion balls, a spread of about a third from the packing assumption alone. Where does the packing factor come from, and why does it matter so much?
Pour marbles into a jar and shake it: they settle at roughly 64% of the jar's volume filled, with the rest air. Stack them by hand in the tightest possible pattern and you reach about 74%. Tip them in gently without shaking and you can be as low as 56%. The floor plate and height are things you can look up or pace out, but the packing factor is a physical assumption you must own, and it moves the answer from 10.4 to 13.8 billion on its own. Say which one you are using and why: balls poured into a building are a shaken random pile, so 0.64 is the defensible middle.
The relationship100 x 4,000 x 3.8 floors times floor plate times floor height, the gross volume in cubic metres 0.5 the usable fraction after the core, structure and services 40.8 x 10^-6 one golf ball in cubic metres 0.64 the packing factor for a shaken random pile of spheres What it says in wordsUsable volume divided by the volume of one ball, scaled down for the gaps between balls.What do you say when the interviewer pushes on the usable fraction?
Lifts, stairs, structural columns, ducts and ceiling voids take a large slice of a tower, and nobody knows it to the percent. Give a range and show which way it pushes the answer: 40% usable takes you to about 9.5 billion, 60% to about 14 billion, so the final figure is somewhere between 10 and 15 billion whatever you assume. A candidate who reports a single number to three figures has missed the point of the question. A candidate who says 12 billion, give or take a third, with the reasons, has answered it.
Where candidates lose it
The common failure is to fix on one number early, such as the number of balls in a room, and then multiply by a guessed room count without ever stating the building's volume. The interviewer cannot follow it, cannot correct it, and marks it as a guess.
The second loss is forgetting that spheres do not fill space. Dividing the volume by the ball's volume overstates the answer by about half. Saying the words packing factor, and a number for it, is what separates a trader's estimate from a schoolchild's.
What the interviewer asks next
- Now the balls are tennis balls, 6.7 centimetres across. Roughly how does your answer change?
- If I told you the real answer was 20 billion, which of your assumptions would you revisit first?
- Make me a market on it: give me a bid and an offer in billions, and tell me how wide and why.
Asked at Tower Research Capital, Assistant Trader, New York, 2013 (Wall Street Oasis):
How many golf balls fit in the empire state building? Explain thought process and detailed solution
002Every market day is either trending or choppy. A trending day is followed by another trending day 70% of the time, and a choppy day by another choppy day 60% of the time. In the long run, what fraction of days are trending? And when a trend starts, how many days does it last on average?Quant tradingHedge funds
Try it first
Gut call before the algebra: which is larger, the share of trending days or the share of choppy days?
Show the worked solution
4/7 of days trend, about 57.1%, and a trend lasts 3.3 days on average. In the long run the flow out of trending must equal the flow in: 0.3 times the trending share equals 0.4 times the choppy share, so the shares stand 4 to 3. A trend ends on any given day with probability 0.3, so its expected length is 1/0.3, and a choppy spell lasts 1/0.4 = 2.5 days.
Why must the two flows balance?
Picture two rooms at a party with a door between them. Each minute, 30% of the people in room A wander into B and 40% of those in B wander into A. The crowd settles when the two queues through the door carry the same number of people; otherwise one room keeps filling. A stationary split is one where the number of days leaving each state equals the number entering it, and that single equation fixes the split. Days leaving trending: 0.3 times the trending share. Days entering it from choppy: 0.4 times the choppy share. Set them equal and the ratio is 4 to 3.
Trending days keep 0.7 of their successors and lose 0.3 to choppy, while choppy days keep 0.6 and lose 0.4 back, so the long-run shares stand 4 to 3, 57.1% trending and 42.9% choppy, with trends lasting 3.3 days and choppy spells 2.5 days on average. The relationshippi_T, pi_C the long-run shares of trending and choppy days 0.3, 0.4 the chance a trending day flips to choppy, and a choppy day flips to trending 1/0.3 expected length of a run that ends with probability 0.3 each day What it says in wordsEach state's share is the other state's flip rate divided by the sum of the two flip rates, and a run's length is one over its own flip rate.Why is the average streak 1/0.3 and not something longer?
A trend that has lasted five days is no more likely to end tomorrow than one that started today: the chain has no memory beyond yesterday. Each trending day ends the run with probability 0.3, independent of its age, so the run length is a geometric count with mean 1/0.3 = 3.33 days. That is the same reason the expected number of rolls to a six is 6. The two answers also check each other: the share of a state equals how often a run of it starts, times how long it lasts, and 4/7 against 3/7 is exactly 3.33 against 2.5 scaled by the same start rate.
What does the chain say about tomorrow, given today?
This is where the puzzle connects to trading. Today's state carries real information: after a trending day, tomorrow trends with probability 0.7, well above the unconditional 57%. After a choppy day it is only 0.4. The long-run split tells you nothing about tomorrow; the transition row for today's state does. Say that distinction out loud, because an interviewer who hears 57% quoted as a one-day forecast knows you have confused the stationary distribution with a conditional one. The limitation to add: a two-state chain with fixed probabilities is a toy, and real regime persistence drifts over time.
Where candidates lose it
The first wrong answer is 50%, on the grounds that each state has one way in and one way out. The flows are not equal in rate: trending leaks at 0.3, choppy at 0.4, and the slower leak wins more of the time.
The second loss is writing out eigenvectors of a two-by-two matrix under time pressure. The balance equation, flow out equals flow in, takes one line and is what the interviewer wants to hear.
What the interviewer asks next
- Starting from a choppy day, what is the chance that the day after tomorrow is trending?
- Add a third state, a crash day, that follows a choppy day 5% of the time. How does the method change?
- How would you estimate these transition probabilities from a year of daily data, and how noisy would they be?
005A stock trades at 100. Each day for three days it moves up 10 or down 10, each with probability one half, and interest rates are zero. What is a call struck at 100 and expiring after the third day worth?
Try it first
Before drawing the tree: how many distinct end prices are there, and how many equally likely paths?
Show the worked solution
7.50. Three moves of plus or minus 10 end at 130, 110, 90 or 70, reached by 1, 3, 3 and 1 of the eight equally likely paths. The 100 call pays 30 at 130, 10 at 110 and nothing below. Its value is the average payoff: (1 x 30 + 3 x 10) over 8, which is 60 over 8, or 7.50. With zero rates and symmetric moves, the real-world probabilities are already the pricing probabilities, so no discounting and no adjustment is needed.
Why is counting paths all the tree needs?
Think of three coin tosses where you get a sweet for each head. The chance of exactly two heads is not one in four; it is three in eight, because there are three orders in which two heads can arrive. The tree recombines, so the value of an end node is its payoff weighted by how many of the eight paths reach it, and the path counts are the binomial coefficients 1, 3, 3, 1. Nothing else in the problem carries information: the step size fixes the end prices and the counts fix the weights.
From 100, three moves of plus or minus 10 reach 130, 110, 90 or 70 by 1, 3, 3 and 1 of the eight equally likely paths, the call pays 30 and 10 at the two upper nodes and nothing below, and the average payoff (1 x 30 + 3 x 10) over 8 gives a price of 7.50. The relationshipk the number of up days out of three C(3, k) the number of paths with k up days: 1, 3, 3, 1 100 + 10(2k - 3) the end price after k ups and 3 minus k downs What it says in wordsThe call is the payoff at each end price, weighted by the share of paths that reach it.Where does the risk-neutral machinery go?
In a general tree you would replace the real probabilities with the risk-neutral ones, chosen so that the stock's expected growth equals the interest rate. Here rates are zero and the moves are symmetric, so the stock already has zero expected drift and the risk-neutral probability is the same one half you were given. Say that out loud: it shows you know the shortcut is a coincidence of the setup, not a rule. If the up move were 10 and the down move 5, one half would no longer price the stock and you would have to solve for the probability that does.
What sanity checks do you say before the number?
Two quick ones. The call cannot be worth more than the expected value of the stock above the strike ignoring the max, which is zero here, so the call is worth exactly the expected positive part, and that is what 7.50 is. And put-call parity with zero rates says the 100 put must also be 7.50, which you can confirm from the lower nodes: (3 x 10 + 1 x 30) over 8. Giving the put price unprompted, and showing it matches, is the cheapest way to prove the tree was right.
Where candidates lose it
The common error is to treat the four end prices as equally likely, which gives (30 + 10) over 4 = 10. The outer nodes are reached by one path each and the inner ones by three; the weights are 1, 3, 3, 1, not 1, 1, 1, 1.
The second loss is reaching for a risk-neutral formula and getting lost in it. With zero rates and symmetric moves, the given probabilities already price the stock. Say why, then count.
What the interviewer asks next
- Now the up move is 10 and the down move 5. What probability prices the stock, and what is the call worth?
- Price the 110 call and the 90 put on the same tree.
- Four days instead of three: what is the 100 call worth, and why does it rise?
007A stock goes up 10% one day and down 10% the next, and keeps alternating for 250 trading days. Where does it end relative to its start? And where does a fund that delivers three times the stock's daily move end up?Volatility tradingWealth management
Try it first
Before multiplying: after one up day and one down day, is the stock back where it started?
Show the worked solution
The stock ends at about 28% of its start; the three-times fund at roughly 8 millionths of its start, effectively zero. Each up-and-down pair multiplies the stock by 1.1 x 0.9 = 0.99, and 125 pairs give 0.99^125 = 0.285. The leveraged fund moves 30% each way, so each pair is 1.3 x 0.7 = 0.91, and 0.91^125 is about 7.6e-06. The arithmetic average return is zero in both cases; the compounded return is not.
Why does a zero average return lose money?
Take a 100 rupee note to a shop that marks everything up 10% in the morning and discounts 10% in the afternoon. The afternoon discount is taken off a bigger number, so the price ends at 99, not 100. A gain and a loss of the same percentage do not cancel, because the loss acts on the larger base; the pair costs the square of the move, 1% for a 10% swing. Repeat that 125 times and the 1% losses compound to a 72% fall. This is volatility drag: the gap between the average return and the compounded return.
On a log scale both paths step down in straight lines: the stock loses 1% per up-and-down pair and ends at 28% of its start after 250 days, while the three-times fund loses 9% per pair and ends at about 8 millionths of its start, so leverage multiplies the drag by far more than three. The relationshipm the daily move, 0.10 for the stock and 0.30 for the three-times fund 1 - m^2 what one up-and-down pair leaves of the value 125 the number of pairs in 250 days What it says in wordsEach pair loses the square of the move, and the leveraged fund's loss per pair is nine times the stock's because 0.3 squared is nine times 0.1 squared.Why is three times the move so much worse than three times the loss?
The drag per pair is the square of the move. Tripling the move multiplies the drag by nine, not three: the stock loses 1% per pair, the fund loses 9%. That is why a daily-rebalanced leveraged fund in a choppy, sideways market bleeds even when the underlying ends flat. The general rule you can quote: over many periods the compounded growth rate is roughly the average return minus half the variance, and leverage multiplies the variance by the square of the leverage.
What would you say to a client who holds the three-times fund?
That the product tracks three times the daily move, exactly as promised, and that this is not the same as three times the return over a year. A leveraged fund is a tool for a view on the next day or week; held through a sideways year it loses to its own rebalancing. The limitation to state: the alternating path is the worst case for drag, and a strongly trending market can make a leveraged fund return more than three times the underlying. The drag is about path, not just direction.
Where candidates lose it
The common answer is that the stock ends flat, because plus 10 and minus 10 seem to cancel. They cancel in arithmetic and not in compounding; the second move acts on a different base.
The second loss is saying the leveraged fund ends at three times the stock's loss, or at 28% cubed. The right route is per pair: 1.3 x 0.7 = 0.91, then raise to the 125th power. The drag scales with the square of the leverage.
What the interviewer asks next
- Make the daily move 1% instead of 10%. Where does the stock end after 250 days?
- Over a year with 16% annual volatility and zero average daily return, roughly what is the compounded return?
- Why do leveraged funds rebalance daily, and what would change if they rebalanced monthly?
010In some stock, the 99-strike call trades at 5.60 and the 101-strike call at 4.70, same expiry. Estimate the price of a digital option that pays 1 if the stock finishes above 100 at that expiry.Exotics tradingStructured products
Try it first
Before any arithmetic: which combination of the two calls has a payoff that looks most like a step at 100?
Show the worked solution
About 0.45. Buying the 99 call and selling the 101 call pays 0 below 99, 2 above 101 and a straight ramp between. Divide that by the width of 2 and the payoff is 0 below 99, 1 above 101 and a ramp through 100: a digital with its edge smoothed over two points. Its cost is (5.60 minus 4.70) over 2, which is 0.45. The narrower the spread, the closer the ramp sits to the step, and the price converges to the digital.
Why does a call spread stand in for a digital?
A light switch is a step: off or on. A dimmer that goes from fully off to fully on over a tiny turn of the knob is, for every practical purpose, the same switch. A call spread over its width is a dimmer: it ramps from 0 to 1 across the two strikes, and as the strikes close in on 100 the ramp becomes the step. So the digital is the limit of a scaled call spread, and a traded call spread gives you a price for it without any model.
The 99 to 101 call spread divided by its width of 2 pays 0 below 99, ramps to 1 at 101 and crosses the digital's step exactly at 100, so the two payoffs differ only inside the narrow band between the strikes and the spread's price, (5.60 minus 4.70) over 2, gives a digital value of 0.45. The relationshipC(K) the price of a call struck at K D(K) the price of a digital paying 1 above K (C(99) - C(101)) / 2 the slope of the call price in strike, estimated across 100 What it says in wordsThe digital is minus the slope of the call price with respect to strike, and a centred call spread measures that slope.Is 0.45 the digital's price or an approximation, and which way is it off?
It is an approximation to the slope at 100 taken from two points either side. Because the spread is centred on 100, the first-order error cancels and what remains is small, of the order of the curvature of the call price between 99 and 101. If the digital were struck at 99 instead, the same spread would overstate it, because the call price is convex in strike and the ramp sits above the step on that side. On a desk you would quote the digital from the tightest spread the market will show you, and hedge it with that spread, so the approximation is also the hedge.
What does 0.45 say about the market, and what is the limitation?
A digital paying 1 above 100 at 0.45, with rates near zero, means the pricing probability of finishing above 100 is about 45%, slightly below one half. That is a risk-neutral probability, not a forecast, and it is pulled down by the skew: with a steeper put skew, out-of-the-money calls are cheaper in volatility terms and the slope in strike is steeper, which moves the digital. Say that a flat-volatility formula would miss this, and that the call spread picks the skew up automatically because it uses the two traded prices.
Where candidates lose it
The common error is to take the difference of the two call prices, 0.90, and present it as the digital. That is the price of a spread that pays 2 above 101, not 1. Divide by the width.
The second loss is reaching for a lognormal formula with a guessed volatility. The question gives you two traded prices precisely so you can price the digital without a model; use them.
What the interviewer asks next
- The 99.5 and 100.5 calls are 5.37 and 4.93. What does that pair say about the digital, and why might it differ from 0.45?
- How would you hedge a short digital you sold at 0.45, and what goes wrong near expiry?
- Price a digital that pays 1 if the stock finishes below 100.
013How many people work in a large bank's 45-storey London headquarters tower? Give me a number and the assumptions behind it.HSBCCentral · 2026
Try it first
Before building anything: which route gives an estimate the interviewer can check link by link?
Show the worked solution
About 8,400 people, with an honest range of roughly 5,558 to 12,994. Of 45 floors, take 40 as ordinary office floors after plant rooms, lobby and trading floors. A floor plate of 3,000 square metres with 70% usable gives 2,100 square metres of desk space per floor. At 10 square metres per person that is 210 people a floor, and 40 floors give 8,400. The widest assumption is the space per person, 8 to 12 square metres, which alone moves the answer by half.
Why build from the floor rather than from the bank?
If you wanted to know how many people a wedding hall holds, you would not guess from the size of the family; you would pace the hall and think about chairs per row. An estimate is only as good as the link the listener can check, and anyone who has worked on an office floor has a feel for how many desks it holds. So the chain is office floors, floor plate, usable share, square metres per person. Say the four links before any number, so the interviewer can argue with one of them rather than with the whole answer.
Forty office floors of 3,000 square metres at 70% usable and 10 square metres a person give 8,400 people, while taking every assumption at its low end gives 5,558 and at its high end 12,994, and the space per person is the link with the widest range, a ratio of 1.5 between its ends. The relationshipfloors office floors out of 45, after plant, lobby and other uses plate gross floor area in square metres usable the share of a floor that holds desks rather than lifts, cores and meeting rooms sq m per person the desk plus a share of corridors and kitchens What it says in wordsPeople equals total desk area divided by the area each person uses.Which assumption should you spend your time on?
The one with the widest range. Floors run 38 to 42, a ratio of 1.11; the plate 2,700 to 3,300, a ratio of 1.22; the usable share 65% to 75%, 1.15; but space per person runs 8 to 12 square metres, a ratio of 1.5, so it moves the answer most. Check it against a floor you know: a trading floor packs people at 6 to 8 square metres, a floor of meeting rooms and offices spreads them at 15 or more. If the interviewer gives you one fact, ask for that one.
What would you add about occupancy?
That desks and people are different counts. With hot-desking, a floor of 210 desks might be home to 250 or 300 people who are not all in on the same day, so a question about who works in the building can give a larger answer than a question about who is in it. Say which one you are answering. The limitation to state: the usable share and the space per person are guesses from general experience, not measurements, and a real number would come from the building's floor plans and the badge-in data.
Where candidates lose it
The common failure is to answer with a bare number, often a round 10,000, and then be unable to defend any part of it. The number is not what is marked; the chain is.
The second loss is spending the time on floors and plate, which are tight, and waving at the space per person, which is loose. Put the effort where the range is.
What the interviewer asks next
- Now estimate how many lifts the tower needs to get everyone in between 8 and 9 in the morning.
- How many taxis operate in a city's central business district on a weekday morning? Build the chain.
- If the building's badge data showed 6,000 entries a day, which assumption would you revisit first?
Asked at HSBC, Sales and Trading, Central, 2026 (Wall Street Oasis):
How many employees in London hsbc building How many taxis are in HK central How many beds in the nyc hotel
014The sample variance with n minus 1 in the denominator is an unbiased estimate of the true variance. Is its square root an unbiased estimate of the standard deviation? If not, which way is it off, and does it matter when you estimate volatility from 20 daily returns?Squarepoint CapitalLondon · 2026
Try it first
Pick before you reason: the sample standard deviation on average is
Show the worked solution
No. The sample standard deviation is biased low, by about 1.3% at n = 20, and for a volatility estimate that is small next to the sampling noise. The square root is concave, so by Jensen's inequality the expected value of the root is below the root of the expected value. For normal returns the exact factor is c4(n): the expected sample standard deviation is 0.9869 times the true one at n = 20. On the same 20 points the estimate's own standard error is around 16%, so the bias is not what you should worry about.
Why does an unbiased variance give a biased standard deviation?
If you average the areas of several square plots and then take the square root, you do not get the average of their side lengths; the big plots pull the area average up more than they pull the side average. Unbiasedness is a statement about averages, and averages do not pass through a curved function: for a concave function like the square root, the average of the outputs is below the output of the average. That is Jensen's inequality, and it is the whole answer. The variance is unbiased; its root is not, and the direction is down.
Two sample variances of 0.1 and 1.9 times the truth average to 1.0, whose root is 1.00, but their roots average only 0.85, and for 20 normal returns the exact expected shortfall is c4(20) = 0.9869, so the sample standard deviation runs about 1.3% low on average. The relationships the sample standard deviation, the root of the n minus 1 sample variance sigma the true standard deviation c4(n) the exact correction factor for normal data, always below 1 and rising to 1 as n grows What it says in wordsFor normal data the sample standard deviation underestimates the true one by a known factor that depends only on the sample size.n c4(n) shortfall 5 0.9400 6.0% 10 0.9727 2.7% 20 0.9869 1.3% 50 0.9949 0.5% 100 0.9975 0.2% The shortfall of the sample standard deviation falls quickly with the sample size, from about 6% at five observations to about 1.3% at twenty and half a percent at fifty. Does it matter for a 20-day volatility estimate?
Not much, and saying why is the second half of the marks. The bias is 1.3%, but the standard error of a standard deviation from 20 observations is roughly 1 over the root of 2 times 19, about 16% of the true value, so the noise is more than ten times the bias. You can multiply by 1/c4 to remove the bias if you like, but you cannot remove the noise without more data, and 20 daily returns is simply a short window. On a volatility desk the honest answer is that a 20-day estimate of 16% could easily have been 13% or 19%.
What assumption does the exact factor need?
Normal returns. The direction of the bias, downward, holds for any distribution, because it comes from concavity alone; the size of the factor c4 depends on the distribution, and fat-tailed daily returns change it. Say the general result first, Jensen, then the normal-case number, then the limitation. That order shows you know what is a theorem and what is a model.
Where candidates lose it
The fast wrong answer is yes, on the grounds that the variance is unbiased and the root is just a relabelling. The root is a curved function, and expectations do not commute with curved functions; the sign of the curvature gives the direction.
The second loss is stopping at biased low. The interviewer asked whether it matters, and the answer is a comparison of 1.3% of bias against roughly 16% of sampling noise. Numbers, not adjectives.
What the interviewer asks next
- Derive the direction of the bias from Jensen's inequality in one line.
- What is the standard error of the sample standard deviation for normal data with n observations, roughly?
- Would you use the n or the n minus 1 version for a volatility estimate, and does the choice matter at n = 20?
Asked at Squarepoint Capital, Quantitative Research, London, 2026 (Wall Street Oasis):
Is the square root of the unbiased estimator for sample variance unbiased for standard deviation?
015You have three six-sided dice. Red has the faces 2, 6, 7; green has 1, 5, 12; blue has 3, 4, 8, with each number on two faces. Two players each pick a die and roll; the higher number wins. Which die would you choose to play with?Belvedere TradingChicago · 2022
Try it first
Before working the pairs: green has the highest average face, 6 against 5 for red and blue. Does that make green the best die?
Show the worked solution
Let the other player choose first, then take the die that beats theirs. Red beats green 5 times in 9, green beats blue 5 in 9, and blue beats red 5 in 9. The three dice form a cycle like rock, paper, scissors, so no die is best on its own; the advantage belongs to whoever picks second. If you must pick first, no choice does better than 4 in 9 against a wise opponent, and you should say so rather than pretend one die is stronger.
How can three dice with the same average not have a best one?
Three cricket teams can each beat one of the others and lose to the third; a league table would show them level, and still no team is the best. Winning a roll depends only on which die shows the higher face, pair by pair, and pairwise comparisons do not have to line up in a single order the way averages do. Red and blue average 5 and green averages 6, and yet green loses to red 5 times in 9: every head-to-head is lopsided, 5 to 4, in a circle, and the highest average sits inside it. Green's 12 wins by a mile and its 1 loses by a mile, and a roll pays nothing for the margin. The question tests whether you check the comparison that matters instead of the summary that does not.
Red beats green in 5 of the 9 equally likely face pairs, green beats blue in 5 of 9 and blue beats red in 5 of 9, so the three dice form a cycle with no best die and the player who picks second always holds a 5 in 9 edge. How do you check a pair quickly in the room?
Write one die's faces across and the other's down and count the cells where the first is higher. Red against green: 2 beats only the 1; 6 beats 1 and 5; 7 beats 1 and 5; that is 1 + 2 + 2 = 5 of 9. Green against blue: 1 beats nothing, 5 beats 3 and 4, 12 beats everything, again 5 of 9. Blue against red: 3 and 4 each beat the 2, and 8 beats 2, 6 and 7, again 5 of 9. Three counts, under a minute, and the cycle appears.
The relationshipR, G, B the face shown by the red, green and blue die 9 the number of equally likely face pairs, three distinct faces on each die 5/9 each die's edge over the next one around the cycle What it says in wordsCount the winning face pairs out of nine for each ordered pair, and the three results form a cycle.What is the trading lesson the interviewer is after?
That the order of moves can be worth more than the thing being chosen. The second mover has a guaranteed 5 in 9; the first mover, against someone who knows the cycle, has at best 4 in 9, so you should pay to move second and never volunteer to move first. That is the same instinct as quoting after you have seen the other side's interest rather than before. The limitation to state: the edge is only 5 to 4, so over a few rolls luck dominates, and a one-roll bet on it is a small edge with a large variance.
Where candidates lose it
The common answer is green, because it has the biggest face and the highest average. Both facts are true and both are irrelevant: a roll pays for being higher, not for being higher by a lot, and the pairwise count is the only thing that decides it.
The second loss is finding the cycle and still naming a die. The answer to which die is a question back: which one is the other player taking? Say that you want to choose second, and why.
What the interviewer asks next
- Each player rolls their die twice and the totals are compared. Does the cycle survive, and does it change direction?
- Design a fourth die that beats all three of these more often than not, or show that none exists.
- Where on a trading desk does moving second carry an edge, and where does it cost you?
Asked at Belvedere Trading, Prop Trading, Chicago, 2022 (Wall Street Oasis):
You have 3 dice: red has 2, 6, 7; green has 1, 5, 12; blue has 3, 4, 8. Highest number wins the game. Which one would you choose to play with?
016You walk into a casino with Rs 63,000 and bet Rs 1,000 on red at even money, where red comes up 48% of the time. Every time you lose, you double the bet. You stop at the first win, or when you cannot cover the next bet. What is the chance you lose everything, and what is your expected result?Risk managementProp trading firms
Try it first
Before any arithmetic: the plan ends a session up Rs 1,000 about 98 times in 100. What is its expected result per session?
Show the worked solution
You lose everything about 2.0% of the time, 0.52 to the sixth power, and the expected result is about minus Rs 265. Rs 63,000 covers exactly six bets: 1, 2, 4, 8, 16 and 32 thousand. A win at any of them recovers every earlier loss and nets Rs 1,000, which happens 98.0% of the time. Six losses in a row cost all Rs 63,000. Weighted, 980 of expected winnings against 1,246 of expected loss leaves minus Rs 265.
Why does a plan that wins 98 times in 100 still lose money?
Picture a friend who sells phone insurance to classmates for Rs 50 a month. Month after month nobody drops a phone, and the Rs 50 notes pile up; it feels like free money until the month three phones go into a pond. A win rate tells you how often you are paid, not how much you are paid against how much you can lose, and the expected value needs both. Doubling after every loss builds exactly that shape: Rs 1,000 collected almost every time, and Rs 63,000 handed back rarely. The rare branch is 63 times the size of the common one, so a 2% chance of it more than cancels a 98% chance of the small win.
The doubling plan ends a session up Rs 1,000 with probability 98.0% and down Rs 63,000 with probability 2.0%, and weighting the two gives plus 980 against minus 1,246, an expected result of minus Rs 265, which is also 4% of the Rs 6,633 the plan expects to stake. How do you lay out the six bets in the room?
Write the ladder down before computing anything. The stakes are 1, 2, 4, 8, 16 and 32 thousand, which add to 63 thousand exactly, so the seventh bet of 64 thousand can never be placed. If the first win comes at bet k, it pays 2 to the power k minus 1 thousand, and the losses before it add to one thousand less than that, so every winning session nets exactly plus Rs 1,000. There are only two outcomes, and the table shows how quickly the chance of reaching each rung falls: by the sixth bet you are staking Rs 32,000 to recover Rs 31,000 of losses and win one more thousand.
Bet Stake (Rs) Lost before it (Rs) Chance of reaching it 1 1,000 0 100.0% 2 2,000 1,000 52.0% 3 4,000 3,000 27.0% 4 8,000 7,000 14.1% 5 16,000 15,000 7.3% 6 32,000 31,000 3.8% Each rung doubles the stake while the chance of reaching it falls by a factor of 0.52, and the chance of losing the sixth bet as well is 1.98%, the probability of ruin. The relationship0.52^6 the chance of six losses in a row, about 2% +1,000 the net result of any session that wins before the money runs out -63,000 the whole bankroll, lost when all six bets lose What it says in wordsThe expected result is the frequent small win times its probability plus the rare total loss times its probability, and the second term is larger.Is there a faster way to see the sign without the ladder?
Yes, and it is the one a trader reaches for first. Every rupee placed on red loses 4 paise on average, whatever happened on the previous spin, because the wheel has no memory. The expected result of any staking plan is the edge per rupee times the expected total amount staked: here minus 4% of Rs 6,633, which is minus Rs 265, the same figure as the ladder. Doubling raises the amount you put down when you are losing; it cannot change the sign of the edge. On a fair 50/50 wheel the same plan has an expected value of exactly zero, with the same lopsided shape.
Why does a desk interviewer care about a roulette plan?
Because the shape is the shape of selling far out-of-the-money options, or of adding to a losing position to get back to flat. Both produce a long run of small gains and a rare large loss, and a good-looking track record says almost nothing about the tail. Repetition makes the rare branch common: play 50 sessions and the chance of at least one ruin is 1 minus 0.98 to the 50th, about 63%. The limitation to state is that the plan assumes no table limit; a casino maximum bet cuts the ladder short and makes ruin more likely, not less.
Where candidates lose it
The common answer is that the plan wins, because it almost always wins. Candidates quote the 98% and stop, never weighing it against the size of the 2% branch. A probability without a payoff is half an expected value.
The second loss is the opposite slip: computing minus 4% of the Rs 63,000 bankroll, about minus Rs 2,520. The edge applies to rupees actually staked, and most sessions stake only Rs 1,000 or Rs 3,000 before the first win. Expected stake, Rs 6,633, is the base.
What the interviewer asks next
- The wheel is fair, 50/50. What is the expected result now, and what is the chance of ruin?
- You have unlimited money but the table caps any single bet at Rs 16,000. How does the picture change?
- Name a trading strategy with the same payoff shape, and say how you would size it.
017Two friends agree to meet at a cafe between 1 pm and 2 pm. Each arrives at a uniformly random time within that hour, independently of the other, and waits 20 minutes for the other before leaving (or until 2 pm, whichever is sooner). What is the probability they meet?Jane StreetNew York · 2026
Try it first
Pick before you draw anything: the chance the two friends meet is
Show the worked solution
5/9, about 55.6%. Put A's arrival time across and B's up a 60 by 60 square; every pair of times is a point, all equally likely, so probability is area. They meet when the times are within 20 minutes, the band either side of the diagonal. They miss in two corner triangles, each with legs of 40 minutes and area 800 of 3,600, which is 2/9. So the meeting chance is 1 minus 4/9 = 5/9.
Why turn two arrival times into a square?
Throw a dart at a square board without aiming and the chance it lands in any patch is just that patch's share of the board. Two independent arrival times, each spread evenly over the hour, behave exactly like that dart: A's time picks a position across, B's time picks a position up. With two independent uniform times, every pair of arrivals is a point in a 60 by 60 square, equally likely anywhere in it, so a probability becomes an area you can see. The event they meet is the set of points where the two times differ by less than 20 minutes, a band hugging the diagonal.
In the 60 by 60 square of arrival times the friends meet in the band within 20 minutes of the diagonal and miss in two corner triangles with legs of 40 minutes, each 2/9 of the area, so the meeting probability is 1 minus 4/9, which is 5/9 or about 55.6%. Why is it easier to compute where they miss?
The band is an awkward six-sided shape; the regions outside it are two clean triangles. In the top-left triangle B arrives more than 20 minutes after A, so A has gone; in the bottom-right one, A is the late one. Each triangle has legs of 60 minus 20 = 40 minutes, so its area is 40 x 40 / 2 = 800 square minutes out of 3,600, which is 2/9, and the two together are 4/9. The clause about leaving at 2 pm changes nothing, because no one can arrive after 2 pm anyway; it only stops the question from looking ambiguous.
The relationshipw the waiting time in minutes, here 20 60 the length of the window in minutes (1 - w/60)^2 the two miss triangles together, which fit into one square of side 1 - w/60 What it says in wordsThe chance of meeting is one minus the square of the share of the hour that falls outside the waiting time.How does the answer move with the waiting time?
Not in a straight line, and that is a common follow-up. Doubling the wait from 10 to 20 minutes takes the meeting chance from about 31% to about 56%, not from one third to two thirds, because the miss region shrinks as a square. The table runs the formula for four waits. The trading version is two orders that must arrive within a latency window to match: halving the gap you can tolerate does more than halve the matches, and a picture of the square is the fastest way to see by how much. The limitation to state is the uniform assumption; real arrivals bunch near the hour, which raises the meeting chance.
Wait (minutes) Miss region Meet probability 10 25/36 11/36, 30.6% 20 4/9 5/9, 55.6% 30 1/4 3/4, 75.0% 40 1/9 8/9, 88.9% The meeting probability rises faster than the waiting time at first and then flattens, because the miss region is the square of the share of the hour outside the wait. Where candidates lose it
The fast wrong answer is one third, from reading the 20 minutes as a share of the hour. It forgets that either friend can be the late one and that the window is cut off at both ends of the hour. Without a picture, people also land on two thirds by doubling the window.
The second loss is drawing the square and then computing the band directly, with a hexagon and several pieces. The interviewer is watching for the complement: two identical triangles, one line of arithmetic, done in under a minute.
What the interviewer asks next
- Each friend now waits 20 minutes but B always arrives in the second half hour. What is the probability they meet?
- Three friends, each waiting 20 minutes. What is the chance all three are there at once?
- What waiting time gives a meeting chance of exactly one half?
Asked at Jane Street, Technology, New York, 2026 (Wall Street Oasis):
two people arrive at a location uniform random time within an hour, each wait 20min, what's the prob they meet

