Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
041Five observations come from a uniform distribution on 0 to theta: 3.1, 7.4, 5.2, 9.0 and 1.8. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
What is the maximum likelihood estimate of theta?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and multiplying by (n + 1)/n = 6/5 gives an unbiased 10.8. Each observation has density 1/theta when theta covers it, so the likelihood is theta to the minus 5 for theta at least 9.0 and zero below. That peaks at 9.0. But the sample maximum averages 5/6 of theta, never above it, so scale it up by 6/5.
Why does the likelihood peak at the largest observation?
Suppose raffle tickets are numbered 1 to N and you see five, the highest being 90. N cannot be below 90, and the smaller N is, the more likely it was to produce those particular five tickets. For a uniform on 0 to theta, each observation has density 1/theta, so the likelihood is theta to the minus 5, which only falls as theta grows, but it is zero for any theta below an observation. The best allowed value is the smallest theta that covers all the data: the maximum, 9.0. Calculus does not help here, because the peak sits at the edge where the likelihood jumps from zero.
The likelihood is zero for theta below 9.0, peaks at 9.0 and then falls as theta to the minus 5, down to 40% of the peak at 10.8; on the data line, the MLE of 9.0 sits at the largest observation, the method of moments gives 10.6 and the bias-corrected estimate is 10.8. Why is 9.0 biased, and what is the right correction?
The sample maximum can never exceed theta, so it can only err on the low side. Five points drop into 0 to theta and cut it into six gaps of the same average size, so the largest point sits on average one gap short of theta: at 5/6 of theta. Scaling the maximum by (n + 1)/n removes that bias: 9.0 x 6/5 = 10.8. The same logic underlies the classic serial-number estimation problem from wartime production counts.
The relationshiptheta the unknown upper end of the uniform n = 5 number of observations max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the sample maximum, which on average falls short of theta by a factor n/(n + 1), so scale it up.An interviewer may ask why not use twice the mean, 2 x 5.3 = 10.6, which is also unbiased. The corrected maximum is far more precise: its variance is theta squared over n(n + 2), against theta squared over 3n for twice the mean, so twice the mean is 2.3 times as variable with five points. Twice the mean can even land below the largest observation, an estimate the data have already ruled out. The limitation of the correction is that unbiased is not the only goal: the multiple of the maximum with the smallest mean squared error is (n + 2)/(n + 1), which gives 10.5 here, and saying you would choose by the loss that matters shows you know the trade.
Where candidates lose it
The common loss is setting the derivative of the log-likelihood to zero, getting -5/theta = 0, and concluding there is no maximum. The maximum is at a boundary, where the indicator switches on, and that is the point of the question.
The second is answering 9.0 and stopping. The follow-up is always the bias; say that the maximum sits below theta on average and give the (n + 1)/n correction with its one-line reason.
What the interviewer asks next
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the variance of the corrected estimator.
- The observations come from a uniform on theta minus 1 to theta plus 1. What is the MLE now?
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
042A queue holds between 0 and 3 orders. Each tick at most one thing happens: with probability 0.3 a new order arrives (if there is room), with probability 0.5 one order is filled (if the queue is not empty), and otherwise nothing changes. In the long run, what fraction of ticks is the queue full?DRWNew York · 2026
Try it first
Roughly what share of ticks is the queue full?
Show the worked solution
27/272, about 9.9% of ticks. In a birth-death chain the long-run flow up across each cut equals the flow down, so share(k) x 0.3 = share(k + 1) x 0.5. Each state's share is 0.6 times the one below: weights 1, 0.6, 0.36 and 0.216, summing to 2.176. The full state gets 0.216/2.176, about 9.9%, and the queue is empty about 46% of the time.
Why can you skip solving the full set of equations?
Stand at a doorway between two rooms at a party that has settled down. Over an evening, the number of people walking through one way must match the number walking back, or one room would keep filling. In a chain that only steps up or down by one, the long-run flow across the boundary between neighbouring states must balance, which gives one simple equation per cut. Here flow up from state k is its share times 0.3, and flow down from state k + 1 is its share times 0.5.
Arrivals push the queue up with probability 0.3 and fills pull it down with probability 0.5, so each state's long-run share is 0.6 times the one below: 46.0% empty, 27.6% with one order, 16.5% with two and 9.9% full. How do the cut equations give the answer?
Write each share relative to the empty state. Each cut gives share(k + 1) = share(k) x 0.3/0.5 = 0.6 x share(k), so the weights are 1, 0.6, 0.36 and 0.216. They sum to 2.176, so the full queue holds 0.216/2.176 = 27/272 of the time, about 9.9%. The staying probabilities, 0.2 in the middle states and 0.5 when full, never enter; a chain that pauses on a state does not change the balance across cuts.
The relationshippi_k long-run share of ticks with k orders in the queue 0.3 chance of an arrival when there is room 0.5 chance of a fill when the queue is not empty What it says in wordsEach state is visited 0.6 times as often as the one below it; normalise the four weights to add to one.Check with conservation. Orders accepted per tick are 0.3 x (1 - 0.099) = 0.2702, and orders filled per tick are 0.5 x (1 - 0.460) = 0.2702: the same, as they must be. That gives a useful business number: arrivals turned away because the queue is full run at 0.3 x 0.099, about 0.030 per tick, or one arrival in ten. The model assumes one event per tick; if an arrival and a fill could happen in the same tick, the chain changes and so do the numbers.
Where candidates lose it
The common loss is assuming the four states are equally likely, or writing out all four balance equations with the self-loops and solving a 4 by 4 system under time pressure. The cut method needs three one-line ratios.
The second is inverting the ratio, using 0.5/0.3, which makes the full state the most common. Fills are faster than arrivals, so the queue must lean toward empty; check the direction before you normalise.
What the interviewer asks next
- What is the average queue length?
- What arrival probability would make the queue full 25% of the time?
- How does the answer change if the queue can hold unlimited orders?
Asked at DRW, Quantitative Research, New York, 2026 (Wall Street Oasis):
There was a problem on Chi-squared distributions which was difficult and also one on birth death chains.
043A 3 x 3 x 3 cube is painted on the outside and cut into 27 small cubes. How many small cubes have 3, 2, 1 and 0 painted faces? You pick a small cube at random and roll it like a die: what is the probability the top face is painted?Jane StreetNew York · 2026
Try it first
What is the probability the top face is painted?
Show the worked solution
8 cubes have 3 painted faces, 12 have 2, 6 have 1 and 1 has none; the chance the top face is painted is exactly 1/3. Corners carry three, edge middles two, face centres one, and the core none. Picking a random cube and rolling it picks a random small face out of 27 x 6 = 162. The painted ones are the big cube's surface, 6 x 9 = 54, so the probability is 54/162 = 1/3.
Where do the 8, 12, 6 and 1 come from?
Think of a Rubik's cube: its pieces are corners, edges and centres, plus a hidden core. A small cube's painted faces equal the number of outer walls it touches: a corner touches three, an edge middle two, a face centre one, the core none. A cube has 8 corners, 12 edges with one middle piece each, and 6 faces with one centre each. That accounts for 8 + 12 + 6 = 26 cubes; the 27th is the core.
Slicing the cube into three layers shows 8 corner cubes with three painted faces, 12 edge cubes with two, 6 face centres with one and a single unpainted core, which together carry 54 painted faces out of 162, exactly one third. Why is the roll probability exactly one third?
Do it the long way first: weight each cube type by its share of cubes and its share of painted faces. 8/27 x 3/6 + 12/27 x 2/6 + 6/27 x 1/6 + 1/27 x 0 = (24 + 24 + 6)/162 = 54/162. Then notice the shortcut: a random cube with a random face up is a uniform pick from all 162 small faces, and the painted small faces are exactly the big cube's surface, 6 x 9 = 54. That gives 1/3 without any case split, and it works for any size: an n x n x n cube gives 6n squared over 6n cubed, which is 1/n.
The relationship8, 12, 6, 1 numbers of corner, edge, face-centre and core cubes 3, 2, 1, 0 painted faces on each type 27 x 6 all small faces, each equally likely to land on top What it says in wordsCount painted small faces over all small faces, because the roll makes every small face equally likely.Interviewers use the second part to see whether you look for the structure before the arithmetic. Counting faces instead of cubes turns a four-case weighted average into one division. Say both routes: the case split proves you can count, the face count proves you can see.
Where candidates lose it
The common loss is answering about cubes when the question is about faces: 26 of 27 cubes have paint, so candidates say 26/27, forgetting that a painted cube still shows an unpainted face most of the time.
The second is miscounting edges, using 8 or 24 instead of 12. Say the cube's shape out loud, 8 corners, 12 edges, 6 faces, and check 8 + 12 + 6 + 1 = 27.
What the interviewer asks next
- For a 4 x 4 x 4 cube, how many small cubes have exactly two painted faces?
- You roll a random small cube and see a painted top. What is the chance it is a corner cube?
- For which n does an n x n x n cube have more unpainted small cubes than painted ones?
Asked at Jane Street, Engineering, New York, 2026 (Wall Street Oasis):
How you got to the answer matters even if you got the question right. Strawberry question + 3x3 cube question
044A stick of length 1 is broken at three independent uniform points into four pieces. What is the expected length of the longest piece?Hudson River TradingNew York · 2024
Try it first
What is the expected length of the longest piece?
Show the worked solution
25/48, about 0.521. For a stick broken into n pieces, the expected k-th smallest piece is (1/n)(1/n + 1/(n - 1) + ... ) with k terms. For n = 4 the sorted pieces average 3/48, 7/48, 13/48 and 25/48, which add to 1. The longest is (1/4)(1 + 1/2 + 1/3 + 1/4) = 25/48, more than twice the average piece of 1/4.
Why is the longest piece so much longer than a quarter?
Cut a sheet of dough at three random spots and the pieces are rarely even: one is usually a big slab and one a sliver. Random breaks produce uneven pieces, and the longest piece collects the unevenness, so its average sits far above the average piece. The average piece is always 1/4; the question asks about the largest of four correlated lengths, which is an {term('order statistic', 'The k-th smallest value in a sample, for example the minimum, the median or the maximum.')}.
Sorted by length, the four pieces of a randomly broken stick average 3/48, 7/48, 13/48 and 25/48 of its length, so the longest piece averages about 0.52, more than twice the naive quarter, and a 100,000-stick simulation agrees to three decimals. Where does the harmonic formula come from?
Start with the shortest piece. The chance that all four pieces exceed x is (1 - 4x) cubed: take x off every piece and the three breaks must fit into the remaining length 1 - 4x. Integrating that from 0 to 1/4 gives an expected shortest piece of 1/16. Then the key fact: the step from each sorted piece to the next adds on average (1/n) times 1 over the number of pieces still longer. After the shortest, three pieces remain longer, so the next piece averages 1/16 + (1/4)(1/3); then add (1/4)(1/2), then (1/4)(1). The longest piece is (1/4)(1/4 + 1/3 + 1/2 + 1) = 25/48.
The relationshipL_(4) the longest of the four pieces 1/4 one over the number of pieces 1 + 1/2 + 1/3 + 1/4 the harmonic sum up to the number of pieces What it says in wordsThe longest piece averages one quarter of the fourth harmonic number.The step rule comes from the fact that the pieces behave like independent exponential lengths scaled to total 1, and the gap between successive minima of exponentials is memoryless. You can name that in the room rather than prove it. The check that the formula is right: the four sorted averages add to exactly 1, and a simulation of 100,000 sticks gives 0.521 for the longest. For n pieces in general, the longest averages (1/n) times the n-th harmonic number, which grows like (ln n)/n.
Where candidates lose it
The common loss is answering 1/4, the average piece. The question asks for the average of the largest piece, and the largest of four uneven pieces is usually more than half the stick.
The second is trying to integrate the maximum directly over the three break points, which gets messy fast. Start from the minimum, use the step rule, and check that the four averages add to 1.
What the interviewer asks next
- What is the expected length of the shortest piece for n pieces?
- What is the probability the four pieces can form a quadrilateral?
- Break the stick at two points instead. What is the expected longest piece?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
I was asked an expected value question involving order statistics.
045Simplified poker with three cards, A, K and Q: each player antes 1, you are dealt one card and I am dealt another. You may bet 1 or check; if you bet, I call or fold, and if you check the higher card wins the antes. How often should you bluff with the Q, and how often should I call with the K, in equilibrium?Old Mission CapitalNew York · 2022
Try it first
How often should I call a bet when I hold the K?
Show the worked solution
Bluff with the Q one time in three, and call with the K one time in three. You always bet the A and always check the K; I always call with the A and fold the Q. A Q bluff risks 1 more to win the 2 antes, so I must defend two thirds of hands facing it: the A gives half, the K calling one time in three gives the rest. The game is worth 1/18 per hand to you.
Which hands are easy, and where is the real decision?
Clear away the obvious hands first. With the A you always bet, because you win whether I call or fold. With the K, betting only gets called by the A and folds out the Q, which you beat anyway, so you check. As the caller, I always call with the A and always fold the Q. The whole game comes down to two mixed choices: how often you bluff with the Q, and how often I call with the K.
In equilibrium the bettor always bets the A, always checks the K and bluffs the Q one time in three, while the caller always calls the A, folds the Q and calls with the K one time in three; together the A and the K calls defend two thirds of hands facing a bluff. How do the indifference conditions fix both frequencies?
A teacher who spot-checks homework faces the same logic: check every paper and time is wasted, never check and everyone copies, check at the right rate and copying stops paying. In equilibrium each player mixes at the rate that makes the other indifferent between their two options. Your Q loses 1 by checking. Bluffing loses 2 against my A, and against my K loses 2 if I call and wins 1 if I fold. The two are equal only when I call with the K one time in three. My K loses 1 by folding; calling loses 2 against your A and wins 2 against a bluff, which is worth -1 only when you bluff one time in three.
The relationshipc how often the caller calls with the K b how often the bettor bluffs with the Q -1 the payoff of the alternative: checking the Q or folding the K, losing the ante What it says in wordsEach frequency is set so the opponent's two choices are worth the same.Check against the pot-odds rule. A bluff risks 1 extra to win the 2 antes, so the caller must defend 2/(2 + 1) = 2/3 of the hands facing it; the A already covers half, and the K calling one time in three covers the other sixth. That two thirds is the number people misremember as the K's calling rate. Put the strategies together and the bettor, who acts first with more information about their own hand, earns 1/18 of a chip per hand. The limitation: with a bigger bet or more cards the ratios change, but the method, indifference on both sides, carries over.
Where candidates lose it
The common loss is setting the K's calling rate to two thirds. Two thirds is the total defence the caller needs against a bluff, and the A already provides half of it, so the K calls only one time in three.
The second is never bluffing because the Q cannot win a showdown. A player who never bluffs lets the caller fold every K to a bet, and the A's bets stop earning. The bluff is what gets the A paid.
What the interviewer asks next
- What is the value of the game to each player?
- The bet size doubles to 2. How do the bluffing and calling frequencies change?
- Now the caller may also bet after a check. What changes?
Asked at Old Mission Capital, Quantitative Research, New York, 2022 (Wall Street Oasis):
Asking to find the game theory optimal strategy in a simplified poker game
046I offer you an even-money bet on either of two events. Which is more likely: at least one six in four rolls of one die, or at least one double six in 24 rolls of a pair of dice? Which side do you take, and what is your edge per rupee?Quant tradingProp trading firms
Try it first
Which event is more likely?
Show the worked solution
At least one six in four rolls is more likely, 51.8% against 49.1%, so take that side. The chance of no six in four rolls is (5/6)^4 = 625/1296, and the chance of no double six in 24 rolls is (35/36)^24. At even money the six earns about 3.5 paise per rupee staked; the double six loses about 1.7 paise and needs 25 rolls to become favourable.
Why does the obvious 2/3 come out the same for both bets?
The old gamblers' shortcut, sometimes called the rule of proportion, says that if one try succeeds with chance p, then n tries succeed with chance n times p. Four rolls at 1/6 gives 4/6, and 24 rolls at 1/36 gives 24/36, the same 2/3. This puzzle is the one the Chevalier de Mere is said to have brought to Pascal, because his winnings disagreed with the shortcut. Adding chances only works for events that cannot happen together, and two sixes in four rolls can. The shortcut counts a hand with two sixes twice, and that is why it breaks down completely after seven rolls, where it would claim 7/6.
A family checking the weather forecast makes the same slip: three days at a 40% chance of rain do not make a 120% chance of a wet weekend. The right question is the chance that it stays dry on all three days, 0.6 cubed, about 22%, so rain on at least one day is about 78%.
At least one six in four rolls of a die happens 51.8% of the time and at least one double six in 24 rolls of two dice happens 49.1% of the time, so the first bet clears even money and the second does not, although the rule of proportion gives two thirds for both. How do you compute each chance in your head?
Go through the complement. At least one success is one minus the chance that every try fails, and independent failures multiply. No six in four rolls is (5/6)^4 = 625/1296, which is just under one half, so the six wins 671/1296, or 0.5177. For the double six, use the approximation that (1 - 1/36)^24 is close to e to the power -24/36, which is e^(-2/3), about 0.513; the exact figure is 0.5086, so the double six comes up only 0.4914 of the time.
The relationship5/6 the chance one roll is not a six 35/36 the chance one roll of two dice is not a double six e^{-24/36} the approximation (1 - 1/n)^m close to e^{-m/n} when 1/n is small What it says in wordsThe chance of at least one success is one minus the chance of failing every time.What edge does each side carry at even money?
At even money you win a rupee with probability p and lose a rupee otherwise, so the expected gain per rupee is p - (1 - p) = 2p - 1. The six gives 2 x 0.5177 - 1 = +0.0355, about 3.5 paise per rupee, while the double six gives -0.0172. That is a thin edge: over 100 bets of Rs 1 the expected profit is about Rs 3.55 with a standard deviation near Rs 10, so it takes thousands of bets before the edge shows reliably. The break-even for the double six is the smallest n with (35/36)^n below one half, which is 25 rolls, giving 0.5055. The limitation: the dice must be fair and independent; any bias in the dice swamps an edge this size.
Where candidates lose it
The common loss is multiplying tries by chance and declaring the two bets equal at 2/3. The interviewer is waiting to see whether you notice that the formula can exceed one, which proves it is wrong, and switch to the complement.
The second is getting the probabilities right and then picking a side without stating the edge. On a trading desk the question is not only which side but how much it is worth: about 3.5 paise per rupee for the six, and a loss of about 1.7 paise for the double six.
What the interviewer asks next
- How many rolls of three dice do you need before at least one triple six is better than even money?
- I offer you 11 to 10 on the double six in 24 rolls. Do you take it?
- How much of your bankroll would you stake per bet on the single six?
047Two assets' daily returns are negatively correlated within every month, yet their monthly returns are positively correlated across the year. How can that happen? Build a small numerical example.Squarepoint CapitalMontreal · 2024
Try it first
Both assets share a drift that changes from month to month, plus daily noise that is negatively correlated. What happens to the correlation as you sum more days into one return?
Show the worked solution
A drift shared by both assets for the whole month can outweigh daily noise that moves them in opposite directions. Within a month the drift is constant, so only the noise shows and the correlation is negative. Summed over 21 days the drift's covariance grows with 21 squared but the noise's only with 21. With noise correlation -0.5 and drift standard deviation 0.3% a day, monthly correlation is +0.48.
What does a three-month example look like?
Picture two shops in the same market street. On any one day, a customer who buys from one did not buy from the other, so their daily takings move against each other. But in festival months the whole street is busy and in the rains the whole street is quiet, so their monthly takings rise and fall together. Correlation at one horizon says nothing on its own about another, because a slow common factor and fast opposing noise can sit in the same data.
Make it numerical with five-day months. In month 1 both assets drift at -0.8% a day, in month 2 at +0.2%, in month 3 at +1.2%. On top, asset A gets daily noise of +0.6, -0.3, 0, +0.3, -0.6 and asset B gets -0.3, +0.3, 0, -0.3, +0.3, which move in opposite directions. Within each month the correlation is -0.95. Summed over each month, the noise nets to zero, so both assets return -4%, +1% and +6%: identical, a monthly correlation of +1. Even all 15 days pooled show +0.71, because the month-to-month swing in drift is larger than the noise.
Within each five-day month the daily returns of the two assets slope downward with a correlation of -0.95, but the monthly drifts of -0.8%, +0.2% and +1.2% a day are shared, so the three cluster centres rise together and the monthly returns correlate at +1. Why does summing more days push the correlation positive?
Write each daily return as the month's drift m plus noise. Over n days the drift adds up to n times m, while the noise adds up to a sum of n separate shocks. The drift's contribution to covariance scales with n squared, the noise's only with n, so the longer the horizon the more the shared drift wins. Take noise with standard deviation 1% a day and a within-month correlation of -0.5, and a drift whose standard deviation across months is 0.3% a day. For a 21-day month the drift adds 39.69 to the covariance and the noise subtracts 10.5, for a monthly correlation of 29.19/60.69 = 0.48.
The relationshipn number of days summed into one return \sigma_m standard deviation of the shared daily drift across months, 0.3% c daily noise covariance within a month, -0.5 \sigma daily noise standard deviation, 1% What it says in wordsShared drift covariance grows with the square of the horizon, independent noise covariance only in proportion to it.For a 21-day month the shared drift adds 39.69 to the covariance and the opposing daily noise takes away 10.5, so monthly returns correlate at +0.48, and the correlation of summed returns crosses from negative to positive at about 5.6 days. The crossover sits where n times 0.09 equals 0.5, about 5.6 days, so even weekly returns of five days would still show a slightly negative correlation, -0.03. A hedge sized on daily correlation can therefore fail at a monthly horizon, which is why a desk measures correlation at the frequency it actually holds risk. The other mechanisms worth naming are mean reversion in the spread between the two assets and stale prices that lag by a day; both also make correlation depend on frequency. The limitation of the example is the assumption that drift is constant inside a month and noise is independent from day to day.
Where candidates lose it
The common loss is saying it is impossible, or that it must be a data error, because correlation feels like a fixed property of two assets. It is a property of two assets at a horizon, and the interviewer wants the decomposition into a slow shared part and a fast opposing part.
The second is a hand-waved answer with no numbers. Build the five-day example in a minute, state that drift covariance scales with n squared and noise with n, and the explanation becomes checkable.
What the interviewer asks next
- What would make daily correlation positive but monthly correlation negative?
- How would you estimate the shared monthly drift from daily data?
- A pairs trader hedges at the daily beta and holds for a month. What goes wrong?
Asked at Squarepoint Capital, Hedge Fund, Montreal, 2024 (Wall Street Oasis):
correlation can be negative intra-month but positive across a year, how?
048Calls on the same stock and expiry are quoted: the 95 strike at 9.80 bid, 10.20 offered, and the 100 strike at 4.30 bid, 4.50 offered. Is there an arbitrage, and exactly how would you trade it?Wolverine Trading, Chicago, ILUSA · 2019
Try it first
What can you lock in, per spread, at these quotes?
Show the worked solution
Yes: sell the 95 call at 9.80 and buy the 100 call at 4.50, collecting 5.30 for a position that can never cost more than 5. A 95/100 call spread pays between 0 and the strike gap of 5 at expiry, so its price must sit between 0 and 5. The market lets you sell it for 5.30, which locks in at least 0.30 per spread, more if the stock ends below 100.
What is a 95/100 call spread worth at most?
Think of two coupons for the same shirt: one lets you buy it for Rs 950, the other for Rs 1,000. The first is worth more, but never by more than Rs 50, because the most it can save you over the second is the Rs 50 difference in price. Long the 95 call and short the 100 call pays the stock's rise above 95, capped once it reaches 100, so at expiry it is worth between 0 and the strike gap of 5. Anything that is certain to pay no more than 5 cannot be worth more than 5 today; with interest it is worth at most 5 discounted, slightly less.
Selling the 95 call at its 9.80 bid and buying the 100 call at its 4.50 offer collects 5.30 for a spread whose payoff is zero below 95 and capped at 5 above 100, so at least 0.30 is kept whatever the stock does at expiry. Which side of each quote do you trade at?
This is where the question is really won or lost. You sell at the bid and buy at the offer, so the spread you can sell is worth 9.80 - 4.50 = 5.30 to you, not the mid of 5.60. 5.30 is still above 5, so the bound is broken at prices you can actually deal at. Buying the spread would cost 10.20 - 4.30 = 5.90 for something worth at most 5, a certain loss, so only one direction works. Check the stock price cases: below 95 both calls expire worthless and you keep 5.30; at 97 you owe 2 on the short call and keep 3.30; at 100 or above you owe exactly 5 net and keep 0.30.
Stock at expiry Short 95 call pays Long 100 call receives Net owed You keep 90 0.00 0.00 0.00 5.30 95 0.00 0.00 0.00 5.30 97 -2.00 0.00 2.00 3.30 100 -5.00 0.00 5.00 0.30 110 -15.00 +10.00 5.00 0.30 The 5.30 collected less what the spread owes at expiry is never below 0.30, because the short 95 call and the long 100 call together never owe more than 5. The relationshipC(K) price of the call with strike K, same stock and expiry e^{-rT} discount factor to expiry; it makes the upper bound slightly below 5 What it says in wordsA call spread is worth between zero and the discounted strike gap; selling it for more than the gap is free money.What could stop the arbitrage from paying?
If the calls are American and the short 95 is exercised early, exercise the 100 call too: you pay the stock price minus 95 and receive the stock price minus 100, a net 5, and you already hold 5.30. The real frictions are fees, the margin the short call ties up, and the risk that the quote vanishes after you trade one leg. Trade both legs together as a spread order. On a real screen a 0.30 bound violation lasts seconds, which is why the interviewer is testing whether you can see it fast and name the side, not whether such quotes are common.
Where candidates lose it
The common loss is reasoning with mid prices: 10.00 - 4.40 = 5.60 and a claimed profit of 0.60. Nobody deals at mids; you sell at the bid and buy at the offer, and the honest edge is 0.30.
The second is getting the direction backwards and buying the spread because the 95 call looks cheap next to its payoff. Say the bound first, the spread is worth at most 5, and the direction follows: sell it.
What the interviewer asks next
- The 105 call is quoted 1.10 bid, 1.30 offered. Is there a butterfly arbitrage across 95, 100 and 105?
- What is the lower bound on the 95/100 call spread, and what quotes would break it?
- How does a dividend before expiry change the early-exercise argument?
Asked at Wolverine Trading, Prop Trading, Chicago, IL, USA, 2019 (Wall Street Oasis):
pricing options given an ask and a bid price for options with different strikes if you were to short one and long another
049Which is larger, e to the power pi or pi to the power e? Prove it without a calculator.Quant researchQuant trading
Try it first
Which way does it go?
Show the worked solution
e^pi is larger: about 23.14 against 22.46 for pi^e. Take logs of both and divide by e x pi, which turns the question into comparing ln e / e with ln pi / pi. The function ln x / x rises up to x = e and falls after it, so its value at e beats its value at any other number, pi included. Undo the steps and the order holds.
How do you turn two awkward powers into one comparison?
When two people race on different tracks, you compare them by converting to the same distance. Here the base and the exponent both differ, so convert each number into a common form. Take logarithms and divide by e x pi: e^pi against pi^e becomes ln e / e against ln pi / pi, the same function evaluated at two points. Logs and division by a positive number both preserve order, so whichever side wins the new comparison wins the original.
The relationship? the unknown direction of the inequality, the same at every step \ln x / x the function whose largest value settles the question What it says in wordsTaking logs and dividing by e times pi turns the question into one function compared at e and at pi.Why is ln x / x largest at e?
Differentiate: the derivative of ln x / x is (1 - ln x) / x squared. It is positive while ln x is below 1 and negative once ln x passes 1, so the function climbs until x = e and falls after. Its largest value, 1/e, occurs only at x = e, so ln pi / pi must be smaller, and therefore e^pi beats pi^e. Equivalently, x^(1/x) peaks at e with value 1.4447, while pi^(1/pi) is 1.4396.
The curve x to the power 1/x reaches its maximum of 1.4447 at x = e and has already fallen to 1.4396 at x = pi, so raising both values to the power e times pi gives e^pi = 23.14, larger than pi^e = 22.46. There is a second proof that needs no calculus beyond one inequality. For any x other than zero, e^x is greater than 1 + x, because the exponential curve lies above its tangent line at zero. Put x = pi/e - 1, about 0.156: then e^(pi/e - 1) is greater than pi/e, so e^(pi/e) is greater than pi, and raising both to the power e gives e^pi greater than pi^e. Give that one if the interviewer asks for a proof without derivatives.
Notice how close the race is: the two values of x^(1/x) differ by only 0.0050, because pi sits near the flat top of the curve. That is why rough estimation is risky here and a proof is needed: 23.14 and 22.46 differ by about 3%. The same argument settles a whole family: for any two numbers a and b with e at most a, and a less than b, a^b is greater than b^a, which is why 3^4 = 81 beats 4^3 = 64.
Where candidates lose it
The common loss is answering pi^e because pi is the bigger base, or trying to estimate both numbers to a decimal and getting lost in the arithmetic. The gap is only about 3%, so mental estimates can land either way.
The second is proving it backwards: assuming the answer and manipulating until something true appears, without checking each step preserves the inequality. Say out loud that taking logs and dividing by the positive e x pi keep the order.
What the interviewer asks next
- Which is larger, 2^3 or 3^2, and why does the argument not apply to 2 and 4?
- Find all pairs of distinct positive integers with a^b = b^a.
- Which is larger, 99^100 or 100^99?
050A bus leaves the depot with some passengers. At stop 1 half of them get off; by stop 2 the number on board has grown by a third; at stop 3 half get off; by stop 4 the number has grown by a third again. There are now 16 people on board. How many started?Jane StreetNew York · 2026
Try it first
How many passengers started?
Show the worked solution
36 passengers started. Work backwards from 16 and undo each step with its inverse. Growing by a third multiplies by 4/3, so undo it by multiplying by 3/4: 16 becomes 12. Undo half getting off by doubling: 24. Then 3/4 again: 18. Double again: 36. Forwards it checks: 36, 18, 24, 12, 16.
Why work backwards instead of setting up an equation?
Retracing your route to find a dropped wallet works because you know where you ended up. Here you know the final count and every step, so the cheapest route is to run the film in reverse. Each step is a multiplication, so each can be undone by multiplying by its reciprocal, starting from the 16 and moving toward the depot. An equation also works, 4x/9 = 16, but the backward chain shows every intermediate count, which lets you check that each one is a whole number of people.
Going forward the count is multiplied by 1/2, 4/3, 1/2 and 4/3 to reach 16; going backward from 16 the inverses 3/4, 2, 3/4 and 2 give 12, 24, 18 and finally 36 passengers at the depot. What is the inverse of growing by a third?
This is the step that catches people. Growing by a third means the new count is 4/3 of the old one. To undo a rise of a third you multiply by 3/4, which removes a quarter of the new number, not a third of it. Taking a third off 16 gives 10.67, which is not a whole person and is a red flag on its own. Multiplying by 3/4 gives 12, and 12 grown by a third is 16 again, so the step checks.
The relationshipx passengers leaving the depot 1/2 half get off 4/3 the count grows by a third 4/9 the net factor over all four stops What it says in wordsFour multiplications compound into one factor of 4/9, so the start is 16 divided by 4/9.The whole-number check also tells you what starting counts are possible at all. Every intermediate count, x/2, 2x/3, x/3 and 4x/9, must be a whole number, so x must be a multiple of 18. A free consistency check is worth saying aloud: 36 is a multiple of 18, and every count on the way, 18, 24 and 12, is whole. The same structure appears on a desk whenever a number passes through several percentage changes: a price up 10% and then down 10% ends at 99% of where it began, and undoing a change always means dividing by the factor, never subtracting the percentage.
Where candidates lose it
The common loss is undoing the growth by taking a third off the later number, which gives 10.67 and stalls. A third of the earlier count is a quarter of the later one, so the inverse is multiplying by 3/4.
The second is doing the steps in the wrong order when working backwards. The last thing that happened is the first thing to undo: start with the growth at stop 4, then the halving at stop 3.
What the interviewer asks next
- What is the smallest number of passengers the bus could have started with for every count to be whole?
- If the pattern repeats for eight stops and 64 people are left, how many started?
- A stock rises a third and then falls a quarter. Where does it end?
Asked at Jane Street, Technology, New York, 2026 (Wall Street Oasis):
x amount of people in the bus. 1/2 got off, 1/3 get in , and so on and so forth

