Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
041Five observations come from a uniform distribution on 0 to theta: 3.1, 7.4, 5.2, 9.0 and 1.8. What is the maximum likelihood estimate of theta, is it biased, and how would you correct it?AQR Capital ManagementTown of Greenwich · 2022
Try it first
What is the maximum likelihood estimate of theta?
Show the worked solution
The MLE is 9.0, the largest observation; it is biased low, and multiplying by (n + 1)/n = 6/5 gives an unbiased 10.8. Each observation has density 1/theta when theta covers it, so the likelihood is theta to the minus 5 for theta at least 9.0 and zero below. That peaks at 9.0. But the sample maximum averages 5/6 of theta, never above it, so scale it up by 6/5.
Why does the likelihood peak at the largest observation?
Suppose raffle tickets are numbered 1 to N and you see five, the highest being 90. N cannot be below 90, and the smaller N is, the more likely it was to produce those particular five tickets. For a uniform on 0 to theta, each observation has density 1/theta, so the likelihood is theta to the minus 5, which only falls as theta grows, but it is zero for any theta below an observation. The best allowed value is the smallest theta that covers all the data: the maximum, 9.0. Calculus does not help here, because the peak sits at the edge where the likelihood jumps from zero.
The likelihood is zero for theta below 9.0, peaks at 9.0 and then falls as theta to the minus 5, down to 40% of the peak at 10.8; on the data line, the MLE of 9.0 sits at the largest observation, the method of moments gives 10.6 and the bias-corrected estimate is 10.8. Why is 9.0 biased, and what is the right correction?
The sample maximum can never exceed theta, so it can only err on the low side. Five points drop into 0 to theta and cut it into six gaps of the same average size, so the largest point sits on average one gap short of theta: at 5/6 of theta. Scaling the maximum by (n + 1)/n removes that bias: 9.0 x 6/5 = 10.8. The same logic underlies the classic serial-number estimation problem from wartime production counts.
The relationshiptheta the unknown upper end of the uniform n = 5 number of observations max the largest observation, 9.0 What it says in wordsThe likelihood peaks at the sample maximum, which on average falls short of theta by a factor n/(n + 1), so scale it up.An interviewer may ask why not use twice the mean, 2 x 5.3 = 10.6, which is also unbiased. The corrected maximum is far more precise: its variance is theta squared over n(n + 2), against theta squared over 3n for twice the mean, so twice the mean is 2.3 times as variable with five points. Twice the mean can even land below the largest observation, an estimate the data have already ruled out. The limitation of the correction is that unbiased is not the only goal: the multiple of the maximum with the smallest mean squared error is (n + 2)/(n + 1), which gives 10.5 here, and saying you would choose by the loss that matters shows you know the trade.
Where candidates lose it
The common loss is setting the derivative of the log-likelihood to zero, getting -5/theta = 0, and concluding there is no maximum. The maximum is at a boundary, where the indicator switches on, and that is the point of the question.
The second is answering 9.0 and stopping. The follow-up is always the bias; say that the maximum sits below theta on average and give the (n + 1)/n correction with its one-line reason.
What the interviewer asks next
- What is the MLE if the distribution is uniform on theta to 2 theta?
- Derive the variance of the corrected estimator.
- The observations come from a uniform on theta minus 1 to theta plus 1. What is the MLE now?
Asked at AQR Capital Management, Trading, Town of Greenwich, 2022 (Wall Street Oasis):
Derive the mle for some given distribution. Explain linear regression intuitively and derive the ols estimate.
042A queue holds between 0 and 3 orders. Each tick at most one thing happens: with probability 0.3 a new order arrives (if there is room), with probability 0.5 one order is filled (if the queue is not empty), and otherwise nothing changes. In the long run, what fraction of ticks is the queue full?DRWNew York · 2026
Try it first
Roughly what share of ticks is the queue full?
Show the worked solution
27/272, about 9.9% of ticks. In a birth-death chain the long-run flow up across each cut equals the flow down, so share(k) x 0.3 = share(k + 1) x 0.5. Each state's share is 0.6 times the one below: weights 1, 0.6, 0.36 and 0.216, summing to 2.176. The full state gets 0.216/2.176, about 9.9%, and the queue is empty about 46% of the time.
Why can you skip solving the full set of equations?
Stand at a doorway between two rooms at a party that has settled down. Over an evening, the number of people walking through one way must match the number walking back, or one room would keep filling. In a chain that only steps up or down by one, the long-run flow across the boundary between neighbouring states must balance, which gives one simple equation per cut. Here flow up from state k is its share times 0.3, and flow down from state k + 1 is its share times 0.5.
Arrivals push the queue up with probability 0.3 and fills pull it down with probability 0.5, so each state's long-run share is 0.6 times the one below: 46.0% empty, 27.6% with one order, 16.5% with two and 9.9% full. How do the cut equations give the answer?
Write each share relative to the empty state. Each cut gives share(k + 1) = share(k) x 0.3/0.5 = 0.6 x share(k), so the weights are 1, 0.6, 0.36 and 0.216. They sum to 2.176, so the full queue holds 0.216/2.176 = 27/272 of the time, about 9.9%. The staying probabilities, 0.2 in the middle states and 0.5 when full, never enter; a chain that pauses on a state does not change the balance across cuts.
The relationshippi_k long-run share of ticks with k orders in the queue 0.3 chance of an arrival when there is room 0.5 chance of a fill when the queue is not empty What it says in wordsEach state is visited 0.6 times as often as the one below it; normalise the four weights to add to one.Check with conservation. Orders accepted per tick are 0.3 x (1 - 0.099) = 0.2702, and orders filled per tick are 0.5 x (1 - 0.460) = 0.2702: the same, as they must be. That gives a useful business number: arrivals turned away because the queue is full run at 0.3 x 0.099, about 0.030 per tick, or one arrival in ten. The model assumes one event per tick; if an arrival and a fill could happen in the same tick, the chain changes and so do the numbers.
Where candidates lose it
The common loss is assuming the four states are equally likely, or writing out all four balance equations with the self-loops and solving a 4 by 4 system under time pressure. The cut method needs three one-line ratios.
The second is inverting the ratio, using 0.5/0.3, which makes the full state the most common. Fills are faster than arrivals, so the queue must lean toward empty; check the direction before you normalise.
What the interviewer asks next
- What is the average queue length?
- What arrival probability would make the queue full 25% of the time?
- How does the answer change if the queue can hold unlimited orders?
Asked at DRW, Quantitative Research, New York, 2026 (Wall Street Oasis):
There was a problem on Chi-squared distributions which was difficult and also one on birth death chains.
043A 3 x 3 x 3 cube is painted on the outside and cut into 27 small cubes. How many small cubes have 3, 2, 1 and 0 painted faces? You pick a small cube at random and roll it like a die: what is the probability the top face is painted?Jane StreetNew York · 2026
Try it first
What is the probability the top face is painted?
Show the worked solution
8 cubes have 3 painted faces, 12 have 2, 6 have 1 and 1 has none; the chance the top face is painted is exactly 1/3. Corners carry three, edge middles two, face centres one, and the core none. Picking a random cube and rolling it picks a random small face out of 27 x 6 = 162. The painted ones are the big cube's surface, 6 x 9 = 54, so the probability is 54/162 = 1/3.
Where do the 8, 12, 6 and 1 come from?
Think of a Rubik's cube: its pieces are corners, edges and centres, plus a hidden core. A small cube's painted faces equal the number of outer walls it touches: a corner touches three, an edge middle two, a face centre one, the core none. A cube has 8 corners, 12 edges with one middle piece each, and 6 faces with one centre each. That accounts for 8 + 12 + 6 = 26 cubes; the 27th is the core.
Slicing the cube into three layers shows 8 corner cubes with three painted faces, 12 edge cubes with two, 6 face centres with one and a single unpainted core, which together carry 54 painted faces out of 162, exactly one third. Why is the roll probability exactly one third?
Do it the long way first: weight each cube type by its share of cubes and its share of painted faces. 8/27 x 3/6 + 12/27 x 2/6 + 6/27 x 1/6 + 1/27 x 0 = (24 + 24 + 6)/162 = 54/162. Then notice the shortcut: a random cube with a random face up is a uniform pick from all 162 small faces, and the painted small faces are exactly the big cube's surface, 6 x 9 = 54. That gives 1/3 without any case split, and it works for any size: an n x n x n cube gives 6n squared over 6n cubed, which is 1/n.
The relationship8, 12, 6, 1 numbers of corner, edge, face-centre and core cubes 3, 2, 1, 0 painted faces on each type 27 x 6 all small faces, each equally likely to land on top What it says in wordsCount painted small faces over all small faces, because the roll makes every small face equally likely.Interviewers use the second part to see whether you look for the structure before the arithmetic. Counting faces instead of cubes turns a four-case weighted average into one division. Say both routes: the case split proves you can count, the face count proves you can see.
Where candidates lose it
The common loss is answering about cubes when the question is about faces: 26 of 27 cubes have paint, so candidates say 26/27, forgetting that a painted cube still shows an unpainted face most of the time.
The second is miscounting edges, using 8 or 24 instead of 12. Say the cube's shape out loud, 8 corners, 12 edges, 6 faces, and check 8 + 12 + 6 + 1 = 27.
What the interviewer asks next
- For a 4 x 4 x 4 cube, how many small cubes have exactly two painted faces?
- You roll a random small cube and see a painted top. What is the chance it is a corner cube?
- For which n does an n x n x n cube have more unpainted small cubes than painted ones?
Asked at Jane Street, Engineering, New York, 2026 (Wall Street Oasis):
How you got to the answer matters even if you got the question right. Strawberry question + 3x3 cube question
046I offer you an even-money bet on either of two events. Which is more likely: at least one six in four rolls of one die, or at least one double six in 24 rolls of a pair of dice? Which side do you take, and what is your edge per rupee?Quant tradingProp trading firms
Try it first
Which event is more likely?
Show the worked solution
At least one six in four rolls is more likely, 51.8% against 49.1%, so take that side. The chance of no six in four rolls is (5/6)^4 = 625/1296, and the chance of no double six in 24 rolls is (35/36)^24. At even money the six earns about 3.5 paise per rupee staked; the double six loses about 1.7 paise and needs 25 rolls to become favourable.
Why does the obvious 2/3 come out the same for both bets?
The old gamblers' shortcut, sometimes called the rule of proportion, says that if one try succeeds with chance p, then n tries succeed with chance n times p. Four rolls at 1/6 gives 4/6, and 24 rolls at 1/36 gives 24/36, the same 2/3. This puzzle is the one the Chevalier de Mere is said to have brought to Pascal, because his winnings disagreed with the shortcut. Adding chances only works for events that cannot happen together, and two sixes in four rolls can. The shortcut counts a hand with two sixes twice, and that is why it breaks down completely after seven rolls, where it would claim 7/6.
A family checking the weather forecast makes the same slip: three days at a 40% chance of rain do not make a 120% chance of a wet weekend. The right question is the chance that it stays dry on all three days, 0.6 cubed, about 22%, so rain on at least one day is about 78%.
At least one six in four rolls of a die happens 51.8% of the time and at least one double six in 24 rolls of two dice happens 49.1% of the time, so the first bet clears even money and the second does not, although the rule of proportion gives two thirds for both. How do you compute each chance in your head?
Go through the complement. At least one success is one minus the chance that every try fails, and independent failures multiply. No six in four rolls is (5/6)^4 = 625/1296, which is just under one half, so the six wins 671/1296, or 0.5177. For the double six, use the approximation that (1 - 1/36)^24 is close to e to the power -24/36, which is e^(-2/3), about 0.513; the exact figure is 0.5086, so the double six comes up only 0.4914 of the time.
The relationship5/6 the chance one roll is not a six 35/36 the chance one roll of two dice is not a double six e^{-24/36} the approximation (1 - 1/n)^m close to e^{-m/n} when 1/n is small What it says in wordsThe chance of at least one success is one minus the chance of failing every time.What edge does each side carry at even money?
At even money you win a rupee with probability p and lose a rupee otherwise, so the expected gain per rupee is p - (1 - p) = 2p - 1. The six gives 2 x 0.5177 - 1 = +0.0355, about 3.5 paise per rupee, while the double six gives -0.0172. That is a thin edge: over 100 bets of Rs 1 the expected profit is about Rs 3.55 with a standard deviation near Rs 10, so it takes thousands of bets before the edge shows reliably. The break-even for the double six is the smallest n with (35/36)^n below one half, which is 25 rolls, giving 0.5055. The limitation: the dice must be fair and independent; any bias in the dice swamps an edge this size.
Where candidates lose it
The common loss is multiplying tries by chance and declaring the two bets equal at 2/3. The interviewer is waiting to see whether you notice that the formula can exceed one, which proves it is wrong, and switch to the complement.
The second is getting the probabilities right and then picking a side without stating the edge. On a trading desk the question is not only which side but how much it is worth: about 3.5 paise per rupee for the six, and a loss of about 1.7 paise for the double six.
What the interviewer asks next
- How many rolls of three dice do you need before at least one triple six is better than even money?
- I offer you 11 to 10 on the double six in 24 rolls. Do you take it?
- How much of your bankroll would you stake per bet on the single six?
048Calls on the same stock and expiry are quoted: the 95 strike at 9.80 bid, 10.20 offered, and the 100 strike at 4.30 bid, 4.50 offered. Is there an arbitrage, and exactly how would you trade it?Wolverine Trading, Chicago, ILUSA · 2019
Try it first
What can you lock in, per spread, at these quotes?
Show the worked solution
Yes: sell the 95 call at 9.80 and buy the 100 call at 4.50, collecting 5.30 for a position that can never cost more than 5. A 95/100 call spread pays between 0 and the strike gap of 5 at expiry, so its price must sit between 0 and 5. The market lets you sell it for 5.30, which locks in at least 0.30 per spread, more if the stock ends below 100.
What is a 95/100 call spread worth at most?
Think of two coupons for the same shirt: one lets you buy it for Rs 950, the other for Rs 1,000. The first is worth more, but never by more than Rs 50, because the most it can save you over the second is the Rs 50 difference in price. Long the 95 call and short the 100 call pays the stock's rise above 95, capped once it reaches 100, so at expiry it is worth between 0 and the strike gap of 5. Anything that is certain to pay no more than 5 cannot be worth more than 5 today; with interest it is worth at most 5 discounted, slightly less.
Selling the 95 call at its 9.80 bid and buying the 100 call at its 4.50 offer collects 5.30 for a spread whose payoff is zero below 95 and capped at 5 above 100, so at least 0.30 is kept whatever the stock does at expiry. Which side of each quote do you trade at?
This is where the question is really won or lost. You sell at the bid and buy at the offer, so the spread you can sell is worth 9.80 - 4.50 = 5.30 to you, not the mid of 5.60. 5.30 is still above 5, so the bound is broken at prices you can actually deal at. Buying the spread would cost 10.20 - 4.30 = 5.90 for something worth at most 5, a certain loss, so only one direction works. Check the stock price cases: below 95 both calls expire worthless and you keep 5.30; at 97 you owe 2 on the short call and keep 3.30; at 100 or above you owe exactly 5 net and keep 0.30.
Stock at expiry Short 95 call pays Long 100 call receives Net owed You keep 90 0.00 0.00 0.00 5.30 95 0.00 0.00 0.00 5.30 97 -2.00 0.00 2.00 3.30 100 -5.00 0.00 5.00 0.30 110 -15.00 +10.00 5.00 0.30 The 5.30 collected less what the spread owes at expiry is never below 0.30, because the short 95 call and the long 100 call together never owe more than 5. The relationshipC(K) price of the call with strike K, same stock and expiry e^{-rT} discount factor to expiry; it makes the upper bound slightly below 5 What it says in wordsA call spread is worth between zero and the discounted strike gap; selling it for more than the gap is free money.What could stop the arbitrage from paying?
If the calls are American and the short 95 is exercised early, exercise the 100 call too: you pay the stock price minus 95 and receive the stock price minus 100, a net 5, and you already hold 5.30. The real frictions are fees, the margin the short call ties up, and the risk that the quote vanishes after you trade one leg. Trade both legs together as a spread order. On a real screen a 0.30 bound violation lasts seconds, which is why the interviewer is testing whether you can see it fast and name the side, not whether such quotes are common.
Where candidates lose it
The common loss is reasoning with mid prices: 10.00 - 4.40 = 5.60 and a claimed profit of 0.60. Nobody deals at mids; you sell at the bid and buy at the offer, and the honest edge is 0.30.
The second is getting the direction backwards and buying the spread because the 95 call looks cheap next to its payoff. Say the bound first, the spread is worth at most 5, and the direction follows: sell it.
What the interviewer asks next
- The 105 call is quoted 1.10 bid, 1.30 offered. Is there a butterfly arbitrage across 95, 100 and 105?
- What is the lower bound on the 95/100 call spread, and what quotes would break it?
- How does a dividend before expiry change the early-exercise argument?
Asked at Wolverine Trading, Prop Trading, Chicago, IL, USA, 2019 (Wall Street Oasis):
pricing options given an ask and a bid price for options with different strikes if you were to short one and long another
049Which is larger, e to the power pi or pi to the power e? Prove it without a calculator.Quant researchQuant trading
Try it first
Which way does it go?
Show the worked solution
e^pi is larger: about 23.14 against 22.46 for pi^e. Take logs of both and divide by e x pi, which turns the question into comparing ln e / e with ln pi / pi. The function ln x / x rises up to x = e and falls after it, so its value at e beats its value at any other number, pi included. Undo the steps and the order holds.
How do you turn two awkward powers into one comparison?
When two people race on different tracks, you compare them by converting to the same distance. Here the base and the exponent both differ, so convert each number into a common form. Take logarithms and divide by e x pi: e^pi against pi^e becomes ln e / e against ln pi / pi, the same function evaluated at two points. Logs and division by a positive number both preserve order, so whichever side wins the new comparison wins the original.
The relationship? the unknown direction of the inequality, the same at every step \ln x / x the function whose largest value settles the question What it says in wordsTaking logs and dividing by e times pi turns the question into one function compared at e and at pi.Why is ln x / x largest at e?
Differentiate: the derivative of ln x / x is (1 - ln x) / x squared. It is positive while ln x is below 1 and negative once ln x passes 1, so the function climbs until x = e and falls after. Its largest value, 1/e, occurs only at x = e, so ln pi / pi must be smaller, and therefore e^pi beats pi^e. Equivalently, x^(1/x) peaks at e with value 1.4447, while pi^(1/pi) is 1.4396.
The curve x to the power 1/x reaches its maximum of 1.4447 at x = e and has already fallen to 1.4396 at x = pi, so raising both values to the power e times pi gives e^pi = 23.14, larger than pi^e = 22.46. There is a second proof that needs no calculus beyond one inequality. For any x other than zero, e^x is greater than 1 + x, because the exponential curve lies above its tangent line at zero. Put x = pi/e - 1, about 0.156: then e^(pi/e - 1) is greater than pi/e, so e^(pi/e) is greater than pi, and raising both to the power e gives e^pi greater than pi^e. Give that one if the interviewer asks for a proof without derivatives.
Notice how close the race is: the two values of x^(1/x) differ by only 0.0050, because pi sits near the flat top of the curve. That is why rough estimation is risky here and a proof is needed: 23.14 and 22.46 differ by about 3%. The same argument settles a whole family: for any two numbers a and b with e at most a, and a less than b, a^b is greater than b^a, which is why 3^4 = 81 beats 4^3 = 64.
Where candidates lose it
The common loss is answering pi^e because pi is the bigger base, or trying to estimate both numbers to a decimal and getting lost in the arithmetic. The gap is only about 3%, so mental estimates can land either way.
The second is proving it backwards: assuming the answer and manipulating until something true appears, without checking each step preserves the inequality. Say out loud that taking logs and dividing by the positive e x pi keep the order.
What the interviewer asks next
- Which is larger, 2^3 or 3^2, and why does the argument not apply to 2 and 4?
- Find all pairs of distinct positive integers with a^b = b^a.
- Which is larger, 99^100 or 100^99?
051You and a friend agree to meet at a spot some time between 5 and 6 pm. Each of you arrives at an independent, uniformly random time in that hour and waits 20 minutes for the other before leaving (or until 6 pm, whichever comes first). What is the probability you meet?Jane StreetNew York · 2026
Try it first
Before you draw anything: what is the chance you meet?
Show the worked solution
5/9, about 55.6%. Put your arrival time on one axis and your friend's on the other, so every outcome is a point in a unit square. You meet when the two times differ by at most a third of an hour, a band along the diagonal. The two corner triangles outside the band each have legs of 2/3, area 2/9, so the band is 1 - 4/9 = 5/9.
Why turn two arrival times into a square?
Think of two people trying to catch each other at a tea stall with no phones. Nothing about the answer depends on who is you and who is the friend; it depends only on the pair of times. With two independent uniform times, every pair is equally likely, so the pair is a point spread evenly over a square and any probability is simply an area. The event you meet becomes a region: the set of points where the two times are within 20 minutes of each other. That region is a diagonal band, because the line x = y is where you arrive together.
Plotting your arrival against your friend's, the meeting region is the diagonal band where the times differ by 20 minutes or less; the two white corner triangles, each 2/9 of the square, are the misses, so you meet with probability 5/9. How do you get the area without any integration?
Count the region you do not want. The two corners where one person arrives more than 20 minutes after the other are right triangles with both legs 40 minutes long, which is 2/3 of the side. Each has area (2/3) x (2/3) / 2 = 2/9, together 4/9, so the band is 5/9. Complements are the fastest route here, as they are for most geometric probability questions, because the leftover pieces are usually triangles.
The relationshipX, Y the two arrival times as fractions of the hour, independent and uniform on 0 to 1 w the waiting time as a fraction of the hour, here 20 of 60 minutes What it says in wordsThe chance of meeting is one minus the two corner triangles, whose legs are each one minus the waiting time.What does the general formula tell you that the number does not?
Read 2w - w squared term by term. The 2w is the naive answer of either person waiting, and the minus w squared removes the double count and the clipping at the edges of the hour. It also tells you how waiting time buys certainty: to meet half the time each person must wait about 17.6 minutes, and to be sure each must wait the whole hour. The same picture prices any tolerance between two random arrivals, such as two orders landing in the same matching window of an auction.
Where candidates lose it
The common answer is 1/3, from reading 20 minutes as a third of the hour. It ignores that either person can be the one who waits, and it has no way to handle the edges of the hour, where a person arriving at 5:55 can only wait five minutes.
The second trap is trying to integrate over one person's arrival time case by case near the edges. It works but wastes three minutes. Draw the square first and subtract the two triangles out loud.
What the interviewer asks next
- How long would each person need to wait for a 50% chance of meeting?
- You wait 10 minutes and your friend waits 30. What is the chance now?
- Three people arrive at random in the hour and each waits 20 minutes. What is the chance all three are together at some moment?
Asked at Jane Street, Technology, New York, 2026 (Wall Street Oasis):
1v1 math problems. bus stop. two people meeting probelm
052You roll a fair die repeatedly until the first six appears. What is the expected sum of all the rolls before the six, not counting the six itself?Quant tradingProp trading firms
Try it first
Pick your answer before working it.
Show the worked solution
15. The first six takes 6 rolls on average, so 5 rolls come before it. Each of those rolls is known not to be a six, so it is uniform on 1 to 5 and averages 3. Expected count times expected size gives 5 x 3 = 15. The check: all rolls including the six average 6 x 3.5 = 21, and taking off the final six leaves 15.
How many rolls come before the six?
Picture waiting at a stop where each minute a bus arrives with chance 1 in 6. On average you wait 6 minutes, and the sixth is the one where it comes. The number of rolls up to and including the first six is geometric with mean 1/p = 6, so the number strictly before it is 5. That is the first factor. Most candidates get this far; the loss comes in the second factor.
A typical game has five non-six rolls before the stopping six, and each of those rolls averages 3 because it is known not to be a six, so the expected sum is 5 x 3 = 15; the full-sum check of 6 x 3.5 less the final 6 also gives 15. Why is each of those rolls worth 3 and not 3.5?
Because you are told something about them. Every roll before the stopping six is, by definition, not a six, so its distribution is the die conditioned on 1 to 5, which averages exactly 3. Using 3.5 gives 17.5, the most common wrong answer. It is the same slip as averaging the income of people who did not win a prize with everyone's income, prize winners included.
The relationshipN the number of rolls before the first six, mean 5 X | X not 6 a roll known not to be a six, uniform on 1 to 5 E the expected sum from any fresh start What it says in wordsExpected count times the expected size of each piece gives 15, and the one-step recursion confirms it.How do you check 15 a second way in the room?
Two checks, both fast. The recursion: with chance 5/6 the next roll is not a six, adds 3 on average and you are back where you started, so E = (5/6)(3 + E), which solves to 15. The full sum: Wald's identityFor a stopping time N that does not look into the future, the expected sum of N independent identical draws equals E[N] times the mean of one draw. applied to every roll including the six gives 6 x 3.5 = 21, and the last roll is always exactly 6, so the rest must average 15. Say both; the second one shows you understand why the conditional mean is 3.
Where candidates lose it
The trap is 17.5: the right count, 5, multiplied by the unconditional mean of a die. The interviewer set the question up so that the rolls you sum are selected, not random, and wants to see whether you notice.
The second trap is multiplying 6 rolls by 3.5 and stopping at 21, which includes the six the question told you to exclude. Say what is counted before you multiply.
What the interviewer asks next
- What is the expected sum if you do count the six?
- What is the expected sum of the rolls before the first time you roll a 1 or a 2?
- What is the expected number of rolls until two sixes in a row?
057Z is a standard normal random variable. What is the expected value of max(Z, 0), and what does that number tell you about the price of an at-the-money option?Quant researchQuant trading
Try it first
Roughly what is E[max(Z, 0)]?
Show the worked solution
1/sqrt(2 pi), about 0.399. Only the positive half contributes, and there you integrate z times the normal density. Because the derivative of the density is minus z times the density, the integral is simply the density's height at zero. So an at-the-money call on a normally distributed move is worth about 0.4 standard deviations of that move: roughly 0.4 x S x volatility x the square root of time.
Why is the answer not zero, and not one half?
Think of a shop that keeps the profit on good days and closes, losing nothing, on bad days. Its average day is better than the average of all days, because the bad days have been floored. max(Z, 0) throws away every negative outcome and keeps every positive one at its full size, so its mean is the positive half's contribution alone: the integral of z times the density from zero to infinity. That is not one half, which is only the chance of being positive; the size of each positive draw matters too.
The shaded area under z times the normal density on the positive side is exactly 0.399, the same as the density's peak height, so an at-the-money call on a normal move of one standard deviation is worth about 0.4, which for a Rs 1,000 stock at 20% volatility over three months is about Rs 40. How do you do the integral in one line?
Notice what differentiating the density gives. The derivative of e to the minus z squared over 2 is minus z times itself, so z times the density is the negative derivative of the density, and its integral from 0 to infinity is the density at 0 minus the density at infinity. The density at infinity is zero, and at zero it is 1/sqrt(2 pi). No tables, no substitution: the answer is 0.3989. Doubling it gives E[|Z|], about 0.798, which is the at-the-money straddle.
The relationshipphi(z) the standard normal density, e^(-z^2/2) / sqrt(2 pi) phi(0) the height of the density at its peak What it says in wordsThe expected positive part of a standard normal equals the height of the bell at its centre.What does it say about an at-the-money option?
If the stock's move to expiry is roughly normal with standard deviation S x vol x sqrt(T), an at-the-money call pays the positive part of that move. So its value is about 0.4 x S x vol x sqrt(T), the rule of thumb option traders use to price at-the-money options in their heads. For a Rs 1,000 stock at 20% volatility and three months, sqrt(T) is 0.5 and the call is about 0.399 x 1,000 x 0.2 x 0.5 = Rs 39.9; the Black-Scholes value with zero rates is Rs 39.88. The rule loosens for long maturities and high volatilities, where the lognormal skew matters.
Where candidates lose it
The two fast wrong answers are 0, from averaging Z itself, and 0.5, from confusing the probability of a positive draw with its expected size. Both come from answering before writing down what is being averaged.
The second loss is getting 0.399 and not connecting it to options, which is why the question is asked on a trading desk. Say the 0.4 rule in the same breath.
What the interviewer asks next
- What is E[max(Z, 1)]?
- What is the variance of max(Z, 0)?
- Using the rule, what is an at-the-money straddle worth on a Rs 500 stock at 30% volatility for one month?
058A bet pays 2 to 1 and wins 40% of the time. What fraction of your bankroll does the Kelly criterion stake on each bet, what long-run growth rate does that give, and what happens if you bet twice that fraction?Quant tradingOptions market making
Try it first
At twice the Kelly stake, what happens to long-run growth?
Show the worked solution
Stake 10% of the bankroll; that grows wealth by about 0.97% a bet, and twice Kelly grows it by only about 0.07%. Kelly is edge over odds: (2 x 0.4 - 0.6)/2 = 0.1. The growth rate is 0.4 ln(1.2) + 0.6 ln(0.9). At 20% the losses compound away almost the whole edge, and above about 20.4% the bankroll shrinks in the long run despite a positive expected value.
Why not bet as much as possible on a good bet?
Think of a shopkeeper with a profitable weekly sale who puts the entire shop's stock on it every week. The average week is good, but one bad week ends the business. With repeated bets, wealth multiplies, so what matters is the average of the log of each outcome, not the average outcome, and a big loss costs more in log terms than an equal gain earns. This bet has a clear edge: each rupee staked returns 0.4 x 2 - 0.6 = Rs 0.20 on average. The question is how much of that edge survives compounding at each stake size.
How do you get the Kelly fraction and the growth rate?
Stake a fraction f. A win multiplies wealth by 1 + 2f, a loss by 1 - f, so the growth per bet is g(f) = 0.4 ln(1 + 2f) + 0.6 ln(1 - f). Set the derivative to zero: 0.8/(1 + 2f) = 0.6/(1 - f), giving f = 0.1. The Kelly stake is the edge divided by the odds, (bp - q)/b = 0.2/2 = 10%. Plugging in, g = 0.4 x 0.1823 - 0.6 x 0.1054, about 0.97% a bet, so the typical path doubles its wealth roughly every 71 bets.
Long-run growth peaks at 0.97% a bet at the Kelly stake of 10%; half Kelly keeps 76% of that growth, twice Kelly keeps almost none of it at 0.07%, and any stake above about 20.4% shrinks the bankroll over time. The relationshipb the net odds, 2 to 1 p, q the chances of winning and losing, 0.4 and 0.6 g(f) expected log growth of wealth per bet at stake f What it says in wordsKelly maximises the expected log of wealth, and its stake is the edge divided by the odds.Why is overbetting so much worse than underbetting?
Near the peak the growth curve is close to a parabola, so the cost of a sizing error grows with its square. Half Kelly gives up only about a quarter of the growth, while twice Kelly gives up nearly all of it, and three times Kelly shrinks wealth at about 2.6% a bet. Real edges are estimated, not known, so a trader who thinks the win rate is 40% but faces 35% is already overbetting at the full 10%. That asymmetry is why desks size at a fraction of Kelly.
Where candidates lose it
The first trap is stopping at the positive expected value and saying bet big. The interviewer is testing whether you know that repeated multiplicative bets are judged by log growth, where volatility itself costs money.
The second is misremembering the formula as p - q or as p/b. Derive it from the log growth in two lines; it is faster than recalling and it proves you know where it comes from.
What the interviewer asks next
- What is the Kelly fraction for an even-money bet that wins 55% of the time?
- Why might a trader deliberately stake half Kelly?
- How would you size two independent simultaneous bets like this one?

