Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
021You roll a fair die again and again and keep a running total. What is the probability that the running total is ever exactly 10? And what does the probability of hitting a given target settle to as the target grows large?Quant trading
Try it first
Before any recursion: for a very large target, the chance the running total lands on it exactly is closest to
Show the worked solution
For 10 the chance is 0.2893, and for large targets it settles at 2/7, about 0.286. The total lands on n only by landing on one of the six numbers before it and then rolling the exact gap, so p(n) is the average of the previous six values, starting from p(0) = 1. Running that recursion gives p(10) = 17,492,167/60,466,176. In the long run the total advances 3.5 per roll, so it lands on one number in 3.5.
How do you set up the recursion?
Think of climbing a staircase by jumping one to six steps at a time, each jump picked at random. To stand on step 10 you must at some point stand on one of steps 4 to 9 and then make exactly the right jump. The running total equals n only if it first equals one of n minus 1 down to n minus 6 and then the next roll is exactly the gap, and those six routes cannot both happen, so p(n) is the sum of p(n minus k) times 1/6 for k from 1 to 6. In words, each value is the average of the six before it, with p(0) = 1 because you start at zero and p of a negative number = 0.
The relationshipp(n) the probability that the running total ever equals n exactly p(n - k) the chance the total visits the number k below the target 1/6 the chance the next roll is exactly the gap k What it says in wordsThe chance of hitting a number is the average of the chances of hitting each of the six numbers just below it.n p(n) n p(n) 1 0.1667 6 0.3602 2 0.1944 7 0.2536 3 0.2269 8 0.2681 4 0.2647 9 0.2804 5 0.3088 10 0.2893 10 0.2893 The hit probability climbs to a peak of 0.3602 at six, falls back to 0.2536 at seven, and by ten is already within half a percentage point of its long-run level of 2/7. The chance the running total ever equals n climbs from 0.167 at one to a peak of 0.360 at six, drops at seven, and then wobbles in towards 2/7 = 0.286, with p(10) = 0.2893, because each bar is the average of the six bars before it. Why does it settle at 2/7?
If you walk down a long street taking steps that average 3.5 paving stones, then over a kilometre you will have stepped on about one stone in every 3.5. A long run of rolls moves the total forward 3.5 per roll on average, so the totals visited are a share 1/3.5 = 2/7 of all the numbers passed, and far from the start every number is equally likely to be one of them. That is the renewal argument, and it gives the limit without any recursion. It also explains why the answer is not 1/6: the total does not get one try at each number, it passes every number and either lands on it or steps over it.
Why the hump at 6, and what is the desk point?
Small totals have many routes compared with their distance from zero: you can reach 6 in one roll, or in two, three, up to six rolls. For n from 1 to 6, p(n) = (1/6)(7/6) to the power n minus 1, so it grows each step and peaks at 0.360 at six; after that the averaging takes over and damps the swings. The interview point, often set as a coding task, is dynamic programming: one pass, six additions per number, no enumeration of paths. The limitation to say out loud is that the 2/7 limit needs a fair die and nothing that depends on the total so far; a rule such as skip your turn above 50 breaks it.
Where candidates lose it
The common answer is 1/6, as though the total gets a single roll at landing on 10. It gets many chances, from 4, 5, 6, 7, 8 and 9, and the whole question is about adding those routes without double counting.
The second loss is trying to count sequences of rolls that sum to 10 and weight each by its length. It works in principle and collapses under the arithmetic in the room. The recursion on p(n) is the answer the interviewer is waiting for, followed by the 2/7 limit from the average step.
What the interviewer asks next
- Write the recursion as a loop and say how much work it takes to reach n = 1,000.
- The die is replaced by a coin that moves the total 1 or 2. What is the long-run hit chance, and what is p(n) exactly?
- What is the expected number of rolls until the running total first reaches 10 or more?
022Make me a market on the total number of Olympic gold medals the United States has ever won, Summer and Winter Games together. You cannot look anything up.Old Mission CapitalBoston · 2024
Try it first
Once the structure is written down, which input deserves most of your quote's width?
Show the worked solution
About 1,096 in the middle, so quote 1,000 bid, 1,200 offered, and say the share of golds is where you are least sure. Summer: about 29 Games attended, times about 190 events, times about 18% won by the United States, gives about 992. Winter: 24 Games, about 62 events, about 7%, adds about 104. The honest range is 756 to 1,491; the quote sits around the middle and moves as soon as someone trades on it.
What is the interviewer actually marking?
If a friend asks how much your monthly grocery bill is, you do not recall a number; you think of trips a week, spend a trip, four weeks, and say a figure with a range. A market on an unknown quantity is marked on the structure behind the middle, the honesty of the range, and what you do when someone trades with you, not on knowing the answer. The interviewer knows the real count. Treat every input below as an assumption to be checked against the published medal table, and say so; inventing precision is worse than a wide range.
Summer Games contribute about 29 x 190 x 18% = 992 golds and Winter Games about 24 x 62 x 7% = 104, a total of about 1,096 with an honest range of 756 to 1,491, and the share of golds is the widest link, so the quote of 1,000 at 1,200 sits around the middle. How do you build the middle?
Split it into Summer and Winter, because they differ by an order of magnitude. Modern Summer Games run every four years from 1896, less three wartime cancellations, about 30 Games, and take one off for the boycotted 1980 Games, 29. Events grew from a few dozen to over three hundred, so call the average about 190, with 170 to 210 as the bracket. The share of golds is the input you own: the United States wins more than any other country, so something like a sixth to a fifth of all golds, 18% in the middle, is defensible but soft. Winter adds 24 Games at about 62 events and a smaller share, about 7%.
The relationship29, 24 Summer and Winter Games the United States took part in, as estimated 190, 62 average events per Games, one gold each 0.18, 0.07 the assumed United States share of golds in each What it says in wordsGolds are the number of Games times events per Games times the share the United States wins, added across Summer and Winter.How wide should the quote be, and what happens when someone trades?
Running all lows together and all highs together gives 756 to 1,491, but every input is unlikely to sit at its extreme at once, so the quote can be much tighter than that range. Quote about 10% either side of the middle, 1,000 bid and 1,200 offered, and treat the first trade as information from someone who may know the answer. If the interviewer pays 1,200, move both sides up, perhaps to 1,150 at 1,350, and ask yourself which link you underestimated; the share is the first suspect. The limitation to admit is that a quote built from memory of rough facts can be confidently off-centre, which is exactly why the market moves when it is traded.
Where candidates lose it
The common failure is to blurt one number, or a market as wide as 500 at 2,000, with no structure behind it. A number with no structure cannot be defended, and a market that wide is a refusal to quote dressed up as caution.
The second loss is holding the quote after the interviewer lifts it twice. A counterparty who keeps buying is telling you the answer is higher. Move up, widen if you must, and say which assumption you are revising.
What the interviewer asks next
- I buy 10 at your offer. Where is your new market?
- Now make a market on the total golds won by all countries ever. Which of your inputs carries over?
- Your structure gave about 1,100. If the true answer were 1,500, which link would you suspect, and why?
Asked at Old Mission Capital, Equities, Boston, 2024 (Wall Street Oasis):
One was on making a market on the total number of Olympic gold medals the US has won.
023A ticket pays Rs 100 if a card drawn from a well-shuffled standard deck is a picture card (jack, queen or king), and nothing otherwise. I quote the ticket at 20 bid, 22 offered. Would you trade, and on which side?Akuna CapitalChicago · 2026
Try it first
Before you work out the fraction: the right action against a 20 bid, 22 offered quote is
Show the worked solution
Yes: buy at 22, because the ticket is worth 23.08. A standard deck has 12 picture cards out of 52, a 23.1% chance, so the fair value of a Rs 100 payout is 12/52 x 100 = 23.08. You buy at the offer, 22, which is below fair, and expect to make about Rs 1.08 a ticket. Selling at 20 would give away about Rs 3.08. The edge is small next to the outcome, so it pays only over many tickets.
How do you value the ticket?
If a friend offers you Rs 100 when a coin lands heads, you would pay up to Rs 50 and no more; the coin's odds set the price. A ticket that pays a fixed amount on an event is worth the payout times the probability of the event, before any talk of the quote. A standard deck has four suits, each with a jack, a queen and a king, so 12 picture cards in 52. The ticket is worth 12/52 x 100 = 3/13 x 100 = 23.08. Notice that the ace is not a picture card; counting it is the quickest way to get 16/52 and the wrong side.
The relationship12 picture cards: jack, queen and king in four suits 52 cards in the deck 100 the payout in rupees if a picture card is drawn What it says in wordsThe ticket is worth its payout times the chance of being paid.Which side of the quote do you compare with?
A quote has two prices, and you can only use one of each: you buy at the offer and sell at the bid. Trade only when fair value sits outside the quote: buy when it is above the offer, sell when it is below the bid, and pass when it falls between them. Here fair is 23.08 against a 22 offer, so you buy and expect 1.08. The figure runs two other quotes for contrast: at 22 bid, 24 offered, fair is inside and there is no trade; at 24 bid, 26 offered, you sell at 24 and expect 0.92.
Against a fair value of 23.08, the quote of 20 bid, 22 offered lets you buy below fair for an expected 1.08 a ticket, a quote of 22 at 24 straddles fair and gives no trade, and a quote of 24 at 26 lets you sell above fair for an expected 0.92. How much would you buy, and what could change your mind?
Each ticket pays 0 or 100, so its standard deviation is 100 x the square root of 0.231 x 0.769, about Rs 42, against an edge of Rs 1.08. The edge is real but small next to the noise, so it is worth taking in size over many independent draws and worth very little on a single ticket. Before trading, ask the two questions a desk would: is the deck standard and well shuffled, and does the person quoting know something you do not, such as a card already removed? The limitation of the clean answer is that it trusts the setup; in a game where the quoter controls the deck, a quote this generous is itself a warning.
Where candidates lose it
The common slip is to sell at 20 on the feeling that the ticket usually loses. That is true and already priced: losing most of the time is why the ticket is worth 23, not 50. Value it first, then compare.
The second loss is comparing fair value with the middle of the quote, 21, and saying buy without naming the price. You buy at the offer. Saying buy at 20 tells the interviewer you do not know which side of a quote you can trade on.
What the interviewer asks next
- The ticket now pays on a picture card or an ace. Where would you trade against the same quote?
- I draw a card, look at it without showing you, and then quote 20 at 22. What do you do now?
- Make me your own market on the original ticket, and say why you chose that width.
Asked at Akuna Capital, Junior Trader Interview, Chicago, 2026 (Wall Street Oasis):
if you win you get 1$. how much money would be a fair bet.
024I roll a fair die, then flip as many fair coins as the die shows. Make me a market on the number of heads.OptiverAustin · 2025
Try it first
Before the variance: what is the fair value, the centre of your market?
Show the worked solution
Centre on 1.75 and quote something like 1.6 bid, 1.9 offered. The die averages 3.5 coins and each coin gives half a head, so the mean is 1.75. The spread comes from two sources: the coin flips, E[N]/4 = 0.875, and the uncertain number of coins, Var(N)/4 = 0.729, a variance of 1.604 and a standard deviation of 1.27. The fair value is exact, so the quote can be tight; the spread tells you how hard to size it.
How do you get the centre?
Suppose a shop's daily customers vary and each spends Rs 200 on average. Average takings are average customers times Rs 200, whatever the day-to-day mix. When a random number of random things are added up, the mean is the expected count times the expected size of each, so here 3.5 coins times half a head = 1.75. That is the law of total expectation in one line. Note that 1.75 is not a possible outcome, and the single most likely outcome is one head, with a chance of 0.312; a market is centred on the mean because that is where neither side has an edge.
The number of heads has mean 1.75 and most of its mass on one and two heads, and its variance of 1.604 splits into 0.875 from the coin flips and 0.729 from not knowing how many coins the die will give, so the standard deviation is 1.27. Why is the spread bigger than the coins alone suggest?
Because you are uncertain about two things at once: how many coins, and how they land. The variance of a random sum is the average of the inner variance plus the variance of the inner mean: E[N] x 1/4 from the coins, plus Var(N) x 1/4 from the die, 0.875 + 0.729 = 1.604. If you knew the die would show 3.5 coins and ignored its own wobble, you would quote a standard deviation of 0.94 instead of 1.27 and size too large. The die's contribution is almost half the total, which is the step most candidates leave out.
The relationshipH the number of heads N the number of coins, the die roll, with mean 3.5 and variance 35/12 1/4 the variance of one fair coin, and the square of its half-head mean What it says in wordsThe mean is half the expected number of coins, and the variance adds the coin noise to the noise in the coin count.How tight should the market be, and what if the other side saw the die?
Width pays you for two risks: not knowing the fair value, and trading with someone who knows more. Here the fair value is exact and nobody has seen anything, so a tight quote around 1.75, such as 1.6 at 1.9, is right, and the standard deviation of 1.27 governs how many contracts you take, not where you centre. Change one fact and the answer changes: if the counterparty has seen the die, a buyer is telling you the die was high. A die of 6 implies 3 heads on average, and a die of 1 only 0.5, so widen sharply or ask to see the die before quoting. The limitation is that real games price in that information risk from the first quote.
Where candidates lose it
The common loss is centring the market on two, the most likely outcome of the coins, or on 1.5 from a guessed three coins. A market is centred on the expected value, and the expected value takes the die's average into account exactly.
The second loss is computing the spread as if the number of coins were fixed at 3.5. That leaves out the variance of the die, almost half the total, and makes you size the position as if it were safer than it is.
What the interviewer asks next
- I buy 5 from you at your offer. Where is your market now, and does it matter whether I saw the die?
- What is the probability of zero heads?
- Now the die decides the number of coins, and each head pays the die's value. What is the expected payout?
Asked at Optiver, Quantitative Research, Austin, 2025 (Wall Street Oasis):
Technical (Simulated EV Poker like game, with cards, coins and dice; Market Making and Taking)
025A stock trades at 500. The one-year 500-strike call is priced at 40 and the 500-strike put at 45, and interest rates are zero. What dividend is the options market pricing in, and what would you check before trading on it?Equity derivativesMarket making
Try it first
Before using any formula: the put costs more than the call at the same at-the-money strike. With zero rates, that tells you
Show the worked solution
A dividend of 5 per share. With zero rates, put-call parity says call minus put equals spot minus the dividend minus the strike. Here 40 minus 45 = 500 minus D minus 500, so D = 5, and the implied forward is 495. Before trading against it, check that the options are European or that early exercise is worth nothing, what it costs to borrow the stock, whether the ex-date falls before expiry, and whether all four prices are live and tradable.
Why does a put costing more than a call point to a dividend?
Suppose you agree today to buy a friend's scooter in a year for its fair price, but the friend will keep using it to deliver parcels and pocket the fees until then. You would pay less than today's price, by the fees they collect. Holding a call and selling a put at the same strike is an agreement to buy the stock at the strike, so call minus put must equal the forward price minus the strike, and the forward is spot minus whatever the holder collects before expiry. With zero rates there is no interest to account for, so a put dearer than the call by 5 says the forward is 5 below spot.
The relationshipC, P the call and put prices at the same strike and expiry S the spot price, 500 K the strike, 500 D the cash the stock pays out before expiry, here the implied dividend What it says in wordsWith zero rates, the call minus the put equals the spot less the dividend less the strike, so the dividend is what balances the equation.Call minus put is minus 5, so the options imply a forward of 495, and since the stock's forward is 500 minus the dividend, the market is pricing a dividend of 5; if the company pays only 3, the remaining 2 has to be explained by something else, such as the cost of borrowing the stock. What would you check before trading on it?
Say your own estimate of the dividend is 3. Then a reversal looks attractive: short the stock, buy the call, sell the put, collecting 500 + 45 minus 40 = 505; at expiry the options deliver the stock back at 500 and you owe the 3 dividend, 503 in all, a locked-in 2 per share. That 2 is only yours if you can borrow the stock for less than 0.4% of its price over the year; a hard-to-borrow stock shows up in parity exactly as an extra dividend. Then check the rest: American calls can be exercised just before a dividend, the ex-date must fall before expiry, and four bid-offer spreads can eat a gap of 2 on their own.
What is the general lesson the interviewer wants?
That parity is an accounting identity between three things you can trade, and any gap between what it implies and what you believe is a claim about something you have not yet priced. The options market does not quote a dividend; it quotes a forward, and the dividend, the borrow cost and the interest rate are the pieces you split it into. A good answer gives the 5, names the implied forward of 495, and then lists what could make the 5 something other than a dividend. The limitation to say plainly: with non-zero rates, the strike is discounted and the arithmetic shifts, so state the zero-rate assumption before quoting the number.
Where candidates lose it
The common answer is that the put is dearer because the market expects the stock to fall. At one strike and one expiry, direction cannot make the put dearer than parity allows; anyone could sell the put, buy the call and short the stock against it. The gap is a forward, not a view.
The second loss is stopping at 5. The question asks what you would check, and the borrow cost is the one interviewers wait for: an implied dividend above the announced one is often a stock that is expensive to short.
What the interviewer asks next
- Interest rates are now 6% a year. Redo the implied dividend, and say which way it moves.
- The company announces a dividend of 8. Which trade would you put on with these four prices, and what is the risk?
- Why might an American call on this stock be worth more than its European twin, and how would that distort the implied dividend?
026I will pay you if at least 60% of the flips of a fair coin come up heads. Do you want 10 flips or 100 flips?Hudson River TradingNew York · 2020
Try it first
Before you count anything: which do you take?
Show the worked solution
Take 10 flips. You are betting on luck, and luck averages out as the trials pile up. With 10 flips, 6 or more heads happens 386 times in 1,024, about 37.7%. With 100 flips, 60 or more heads happens about 2.8% of the time. The share of heads settles towards 50% at the rate of one over the square root of n, so the 60% line gets harder to reach with every extra flip.
Why does the number of flips change the odds at all?
Think of a school with two cricket teams, one of eleven and one of a hundred and ten. If someone offers a prize for a team whose average height is 10 cm above the national average, the small team is the one that can win: one or two tall players move its average, while the big team's average is pinned down by sheer numbers. The same coin, flipped more often, produces a share of heads that sits ever closer to a half, so a payout that needs an unusual share wants the fewest flips you can get. This is the law of large numbers working against you, and the question is testing whether you know which side of it you are on.
With 10 flips the bars at 6 heads and above hold 37.7% of the probability, but with 100 flips the bars at 60 heads and above hold only 2.8%, because the share of heads tightens around a half as the flips increase. How do you put a number on it without a table?
Count the small case exactly: 6 or more heads in 10 flips means adding the ways to get 6, 7, 8, 9 and 10 heads, which are 210, 120, 45, 10 and 1, a total of 386 out of 1,024, so 37.7%. For 100 flips use the normal approximation. The standard deviation of the share of heads is 0.5 over the square root of n: 15.8% for 10 flips, 5% for 100. The 60% line is 0.6 standard deviations out in the first case and 2.0 in the second, and two standard deviations in one tail is about 2.3%. The exact binomial answer is 2.8%; the approximation gets you to the right decision in one breath.
The relationshipsigma of X bar the standard deviation of the share of heads n the number of flips z how many standard deviations the 60% line sits from the mean of 50% What it says in wordsThe 60% line gets further from the centre, measured in standard deviations, as the flips increase, so it becomes rarer to cross.What is the interviewer listening for after the answer?
Say the general rule and then the exception. The rule: whenever a payout needs the sample to look unlike the population, choose the smallest sample. The exception: if the payout were for landing between 40% and 60%, you would want the most flips, for exactly the same reason. Read the sign of the bet before you choose the sample size: a bet on luck wants few trials and a bet on the average wants many. If the interviewer changes the wording to more than 60%, 7 or more heads in 10 is 17.2%, still far above the 100-flip figure. And if the choice is between 100 and 1,000 flips, 600 or more heads happens about 1.36e-10 of the time, which is as close to never as a desk needs.
Where candidates lose it
The fast wrong answer is 100 flips, because more flips feel like more chances. They are more chances for the average to assert itself, not for luck. Candidates who say it have the law of large numbers backwards, and the interviewer hears it immediately.
The second loss is saying 10 without a number. Have the 386 out of 1,024 ready, then the standard deviation argument for 100, so the answer sounds reasoned rather than remembered.
What the interviewer asks next
- I pay you if the share of heads is between 45% and 55%. Now which do you want?
- What if the coin has a 60% bias towards heads? Does the answer flip?
- Roughly how many flips make the 60% line a three standard deviation event?
Asked at Hudson River Trading, Prop Trading, New York, 2020 (Wall Street Oasis):
Questions on EV for coin tosses, law of large numbers, Bayes theorem
027You back out implied volatility from an option price with Newton's method. For an at-the-money call priced at 40 on a stock at 1,000 with three months to expiry and rates at zero, starting from 30%, how fast does it converge, when can it fail, and what starting guess do traders use?Akuna CapitalNew York · 2025
Try it first
Starting from 30%, how many Newton steps until the error is below one part in a million?
Show the worked solution
Two steps, because Newton converges quadratically near the root. From 30% the error goes 0.099, 8.2e-05, 4.3e-11, then machine precision: the correct digits roughly double each step. The implied volatility is 20.06%. Newton fails where vega is tiny, far out of the money or close to expiry, because dividing by a near-zero slope throws the next guess to nonsense, and it fails outright if the price sits outside the no-arbitrage bounds. Traders start from price over 0.4 x S x sqrt T, which is 20% here.
Why is Newton so fast on this option?
Picture walking towards a wall in the dark by stepping the full distance your outstretched hand estimates. If the floor is level the estimate is right and you arrive in one step; if it slopes gently you arrive in two. Newton does the same with the pricing function: it fits a straight line at the current guess and jumps to where that line hits the target price. The jump is as good as the line, and at the money the call price is almost a straight line in volatility, so the first jump lands within a hair of the answer. Here the price is roughly S x 0.4 x sigma x sqrt T, which is linear in sigma, with a small concave bend. Starting at 30% gives a price of 59.79 against a target of 40; one step takes sigma to 20.0532%, an error of 8.2e-05, and the next step clears ten digits.
The relationshipC(sigma) the model price at the current volatility guess C mkt the market price, 40 here dC/d sigma vega, the slope of price in volatility sigma star the implied volatility being solved for C'' over 2C' the curvature of price in volatility relative to its slope; small at the money, so the squaring bites hard What it says in wordsEach step divides the price gap by the slope, and once close the error is squared, so the correct digits double every step.For the at-the-money call the error in volatility falls from 0.10 to 8.2e-05 to 4.3e-11 and reaches machine precision by the third step, while for a far out-of-the-money call started where vega is tiny the first step overshoots to 114% and the method needs many more steps to crawl back. When does the method fail, and what does the failure look like?
Newton divides by vega, so it breaks where vega is close to zero: far out of the money, close to expiry, or at a very low starting volatility. Take an illustrative call struck at 1,300, 30% above spot, priced at 0.50. Its true implied volatility is 23.0%, but at a starting guess of 15% the model price is 0.005 and vega is only 0.50 per unit of volatility, so the first step jumps to 114% and the method needs 7 more steps to get within 7e-06. Start at 10% and vega is 2.4e-04, so the step divides by almost nothing and the next guess is a volatility of 2,095, which is garbage. The other failure is a price with no solution at all: a call priced below its intrinsic value or above the stock has no volatility that produces it, and Newton loops forever. Check the bounds before you iterate.
What starting guess do traders actually use?
Use the at-the-money approximation: an at-the-money call is worth about 0.4 x S x sigma x sqrt T, so invert it. Here that gives 40 / (0.4 x 1,000 x 0.5) = 20%, within 0.0006 of the true 20.06%; the version with the exact constant, sqrt(2 pi / T) x C / S, gives 20.05%. A guess that close means Newton is finishing a job that is already nearly done, which is why production code rarely needs more than three steps. For options away from the money, a guard is standard: a starting volatility of sqrt(2 |ln(S/K)| / T), 145% for the 1,300 strike, from which the method is known to converge, or a bracketed method such as bisection for the first few steps and Newton only to polish. Say the limitation too: all of this assumes a price that the model can reach, and real screens carry stale or crossed quotes that no solver can fix.
Where candidates lose it
The common loss is describing Newton as halving the error, which is bisection, or saying one step per digit, which is a linear method. The word the interviewer wants is quadratic, with the digits doubling, and the reason: near the root the error is squared.
The second is forgetting the failure cases. A candidate who only praises the speed has not run the method on a far out-of-the-money option, where a tiny vega sends the next guess negative. Name vega as the divisor and the failure explains itself.
What the interviewer asks next
- Why is the call price nearly linear in volatility at the money, and where does it stop being so?
- What goes wrong if you start Newton above the true volatility for a far out-of-the-money put?
- How would you make the solver robust enough for a live surface of ten thousand strikes?
- Price a call at 40 with the stock at 1,000: is any price between 0 and 1,000 reachable by some volatility?
Asked at Akuna Capital, Quantitative Research, New York, 2025 (Wall Street Oasis):
Convergence time of newton's method
028How many zeros are at the end of 1000! (1000 factorial)?Jump TradingChicago · 2013
Try it first
Pick the count before you work it out.
Show the worked solution
249 trailing zeros. A zero at the end is a factor of 10, and a 10 is a 2 times a 5. Among 1 to 1,000 there are 994 factors of 2 but only 249 factors of 5, so the 5s decide. Multiples of 5 contribute 200, multiples of 25 add a second 5 each for 40 more, multiples of 125 add 8, and 625 adds 1: 200 + 40 + 8 + 1 = 249.
Why count 5s and not 10s?
Imagine packing gift boxes that each need one lid and one base, from a pile of 994 lids and 249 bases. You can make 249 boxes, and the extra lids are useless. A trailing zero is a box: it needs one factor of 2 and one factor of 5, and in the product 1 x 2 x 3 x ... x 1,000 the 5s run out long before the 2s. Counting multiples of 10 misses every zero that comes from pairing a 5 in one number with a 2 in another; 4 x 5 is 20, which ends in a zero though neither factor is a multiple of 10. So the question is simply: how many times does 5 divide into 1,000 factorial?
Multiples of 5 up to 1,000 contribute 200 factors of 5, multiples of 25 a further 40, multiples of 125 another 8 and 625 one more, which add to 249, while the 994 factors of 2 are never the constraint. How do you count the 5s without missing the doubled ones?
Count in layers. Every multiple of 5 carries at least one 5: 1,000 divided by 5 is 200. Every multiple of 25 carries a second 5 that the first layer did not see: 1,000 over 25 is 40. Multiples of 125 carry a third, 8 of them, and 625 carries a fourth, once. Each layer counts only the extra 5 that the layer before it missed, which is why you add the plain quotients and never multiply. The sum is 249. The same method gives 24 zeros for 100 factorial, 20 plus 4, and the layering stops as soon as the power of 5 exceeds the number.
The relationshipZ(n) the number of trailing zeros of n factorial floor of n over 5^k how many numbers up to n carry at least k factors of 5 What it says in wordsAdd the whole-number quotients of n by each power of 5 until the power is bigger than n.What does the interviewer learn from how you say it?
The answer is a test of whether you decompose a question into its prime factors rather than reach for the obvious count. Say the pairing argument first, then the layered count, then check the shape: the answer is a little under n over 4, because 200 + 40 + 8 + 1 is a geometric series that sums towards 1,000 over 4. That check is useful on the follow-ups. For the number of factors of 2 the same layering gives 994; for 3s it gives 498, so in base 6, where a trailing zero needs a 2 and a 3, 1,000 factorial ends in 498 zeros. The limitation is worth a sentence: this counts zeros at the end, not zeros anywhere in the 2,568-digit number, which is a different and much harder question.
Where candidates lose it
The fast wrong answer is 100, one per multiple of 10, or 200, one per multiple of 5. The first misses every zero made by pairing a 5 with a 2 from a different number; the second misses the extra 5s inside 25, 125 and 625.
The other slip is counting 625 as four and 125 as three in the final sum, which double counts. Each layer adds only one new 5 per number, so the layers are added as plain quotients.
What the interviewer asks next
- How many trailing zeros does 100! have?
- How many trailing zeros does 1000! have when written in base 6? In base 12?
- What is the smallest n for which n! ends in exactly 100 zeros, and is there an n that ends in exactly 5 zeros?
Asked at Jump Trading, Trading, Chicago, 2013 (Wall Street Oasis):
How many 0's are in 1000! (factorial)?
029A point is dropped uniformly at random inside a unit square. What is the expected distance from the point to the nearest edge?Hudson River TradingNew York · 2024
Try it first
Pick the expected distance before you integrate.
Show the worked solution
One sixth. The distance to the nearest edge exceeds t only when the point lands inside the inner square of side 1 minus 2t, which has area (1 - 2t)^2. The expected value of a non-negative quantity is the integral of its tail probability, so E[D] is the integral of (1 - 2t)^2 from 0 to 1/2, which is 1/6, about 0.167. One axis alone would give 1/4; the second axis trims it.
Why work with the chance of being further than t, rather than the distance itself?
Suppose you want the average waiting time at a counter, and the only thing you can observe is, for each t, the fraction of people still waiting after t minutes. That is enough: add up the survival fractions over all t and you have the average wait. For any non-negative quantity, the expected value equals the area under its tail curve, and here the tail curve is easy to see: the point is further than t from every edge exactly when it lands inside the inner square of side 1 minus 2t. That inner square has area (1 - 2t)^2, and it shrinks to nothing at t equal one half, the centre of the square. Writing the density of the minimum of four dependent distances directly is far messier.
A random point is further than t from every edge only when it lands in the inner square of side 1 minus 2t, so the tail probability is (1 - 2t)^2, and the area under that curve from 0 to 1/2 is the expected distance of one sixth, confirmed by a 0.1667 Monte Carlo estimate. How does the integral come out to exactly 1/6?
Substitute u = 1 - 2t. As t runs from 0 to 1/2, u runs from 1 down to 0, and dt is minus du over 2. The integral becomes one half of the integral of u squared from 0 to 1, which is one half times one third. The answer is 1/6 because a squared tail integrates to a third and the half-width of the square halves it again. A sanity check: the median distance is where (1 - 2t)^2 equals a half, which is t = 0.146, a little below the mean, as you expect for a distribution with a long right tail reaching 0.5.
The relationshipD the distance from the point to the nearest edge (1 - 2t)^2 the area of the inner square where every edge is further than t away u the side of that inner square, 1 - 2t What it says in wordsThe expected distance is the area under the tail curve, and the tail curve is the area of a shrinking inner square.What is the fast cross-check that shows you understand the structure?
Do one axis first. The distance to the nearer of the left and right edges is min(x, 1 - x), a triangle-shaped quantity with mean 0.25. The nearest edge of the square is the smaller of two such independent quantities, one per axis, and taking the smaller of two pulls the mean down from 1/4 to 1/6. Say 1/4 for one axis, then 1/6 for two, and the interviewer hears that you see the minimum of independent pieces rather than a formula. The same structure gives 1/8 for a unit cube, where the tail is (1 - 2t)^3, and it shows the limitation of the method: it works because the inner region stays the same shape as it shrinks, which fails for a disc, where the nearest point on the boundary is not along an axis.
Where candidates lose it
The common loss is answering 1/4, the one-axis answer, forgetting that the nearest of four edges is smaller on average than the nearest of two. The second axis is independent and matters.
The other is trying to write the density of the minimum directly and running out of time. The tail curve is one line, (1 - 2t)^2, and the expected value is its integral. Learn that identity; it solves half the expected value questions a desk asks.
What the interviewer asks next
- What is the expected distance to the nearest face for a point dropped in a unit cube?
- What is the expected distance to the nearest edge in a 2 by 1 rectangle?
- What is the expected distance from the point to the centre of the square? Why is that a different kind of integral?
- Two points are dropped. What is the chance that both are further than 0.1 from every edge?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
expected value question involving the expected value among distance to an edge, with a randomly placed object
030A coin falls your way 60% of the time and pays even money. What fraction of your capital should you bet each time, and what growth rate does that give?OptiverAustin · 2025
Try it first
What fraction of your capital goes on each flip?
Show the worked solution
Bet 20% of your capital each time, for a growth rate of about 2.0% per bet. The Kelly fraction at even money is the edge, p minus q, which is 0.6 minus 0.4. The expected log growth is 0.6 x ln(1.2) + 0.4 x ln(0.8) = 0.0201, so capital typically grows by a factor of about 7.5 over 100 bets. Bet twice that, 40%, and the growth turns slightly negative, -0.24% a bet: anything beyond about 39% loses money despite the edge.
Why not bet everything on a coin that favours you?
A street vendor with a stall that makes money six days in seven does not spend the whole float on stock each morning, because the seventh day would end the business. Expected value is the right guide for a single bet with money you can replace, but when you must survive to play again, what matters is the growth of your capital over many bets, and that is governed by the average of the logarithm, not the average of the rupees. Betting all of it has the highest expected value and a certainty of ruin: one loss and there is nothing left to compound. The logarithm punishes that loss infinitely, which is the mathematical way of saying you cannot come back from zero.
The relationshipf the fraction of capital staked on each bet p, q the chance of winning and of losing, 0.6 and 0.4 f star the fraction that maximises long-run growth; at even money it is the edge g the expected log growth per bet, about 2% here What it says in wordsMaximise the average log of your wealth after one bet, and at even money the best stake is simply the edge.Expected log growth per bet rises from zero to a peak of 2.01% at a 20% stake and falls back through zero at about 39%, so half Kelly keeps three quarters of the growth while double Kelly, just past the crossing, loses a little on every bet. Where does the 20% come from, and what if the odds are not even?
Differentiate g with respect to f and set it to zero: p over (1 + f) equals q over (1 - f), which gives f = p - q. At even money that is the edge, 0.2. For a bet that pays b to 1 the formula becomes (bp - q) over b, edge over odds. Say the general form, because the interviewer will often move the payout: at 2 to 1 on the same coin the stake becomes (1.2 - 0.4) / 2 = 40%, and at odds where bp is below q the right stake is zero. The growth at 20% is 0.0201 per bet, which compounds to a factor of e to the 2.01, about 7.5 times, over 100 bets on the typical path. That is the median outcome, not the average: a few lucky paths do far better and pull the mean up.
Why do desks bet less than Kelly?
The curve is flat on the left and steep on the right. Half Kelly, 10%, still delivers 1.50% a bet, three quarters of the peak, with far smaller swings. Double Kelly, 40%, delivers -0.24%, already below zero, with wild swings: the crossing point is about 39%, and anything beyond it shrinks capital in the long run even though every single bet has positive expected value. Over-betting is punished much harder than under-betting, and in real trading the 60% is an estimate rather than a fact, so sizing at a fraction of Kelly is the standard discipline. The limitation to state: Kelly assumes you know p, that bets are independent and repeated, and that you can resize freely; on a desk, position limits and the uncertainty in your edge usually bind before the formula does.
Where candidates lose it
The common loss is answering with the expected value: 60% of the time you win, so bet big, or even bet everything. The interviewer wants to hear that growth is about the log, that ruin ends the game, and that the stake is the edge, 20%.
The second is giving 20% but no growth rate, or quoting the 20% expected return per rupee as the growth rate. The growth per bet is about 2%, far below the 20% edge, because the losses bite on a shrinking base.
What the interviewer asks next
- The coin pays 2 to 1 instead of even money. What is the stake now?
- Why does half Kelly give three quarters of the growth? What is the general relationship?
- You are only 60% sure the coin is 60% biased. How does that change the stake?
- What is the probability that a full-Kelly bettor halves their capital at some point before doubling it?
Asked at Optiver, Quantitative Research, Austin, 2025 (Wall Street Oasis):
Use Kelly Criterion
