Derivatives Foundation puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 66
- Topics
- 12
- Hard
- 29
002Every market day is either trending or choppy. A trending day is followed by another trending day 70% of the time, and a choppy day by another choppy day 60% of the time. In the long run, what fraction of days are trending? And when a trend starts, how many days does it last on average?Quant tradingHedge funds
Try it first
Gut call before the algebra: which is larger, the share of trending days or the share of choppy days?
Show the worked solution
4/7 of days trend, about 57.1%, and a trend lasts 3.3 days on average. In the long run the flow out of trending must equal the flow in: 0.3 times the trending share equals 0.4 times the choppy share, so the shares stand 4 to 3. A trend ends on any given day with probability 0.3, so its expected length is 1/0.3, and a choppy spell lasts 1/0.4 = 2.5 days.
Why must the two flows balance?
Picture two rooms at a party with a door between them. Each minute, 30% of the people in room A wander into B and 40% of those in B wander into A. The crowd settles when the two queues through the door carry the same number of people; otherwise one room keeps filling. A stationary split is one where the number of days leaving each state equals the number entering it, and that single equation fixes the split. Days leaving trending: 0.3 times the trending share. Days entering it from choppy: 0.4 times the choppy share. Set them equal and the ratio is 4 to 3.
Trending days keep 0.7 of their successors and lose 0.3 to choppy, while choppy days keep 0.6 and lose 0.4 back, so the long-run shares stand 4 to 3, 57.1% trending and 42.9% choppy, with trends lasting 3.3 days and choppy spells 2.5 days on average. The relationshippi_T, pi_C the long-run shares of trending and choppy days 0.3, 0.4 the chance a trending day flips to choppy, and a choppy day flips to trending 1/0.3 expected length of a run that ends with probability 0.3 each day What it says in wordsEach state's share is the other state's flip rate divided by the sum of the two flip rates, and a run's length is one over its own flip rate.Why is the average streak 1/0.3 and not something longer?
A trend that has lasted five days is no more likely to end tomorrow than one that started today: the chain has no memory beyond yesterday. Each trending day ends the run with probability 0.3, independent of its age, so the run length is a geometric count with mean 1/0.3 = 3.33 days. That is the same reason the expected number of rolls to a six is 6. The two answers also check each other: the share of a state equals how often a run of it starts, times how long it lasts, and 4/7 against 3/7 is exactly 3.33 against 2.5 scaled by the same start rate.
What does the chain say about tomorrow, given today?
This is where the puzzle connects to trading. Today's state carries real information: after a trending day, tomorrow trends with probability 0.7, well above the unconditional 57%. After a choppy day it is only 0.4. The long-run split tells you nothing about tomorrow; the transition row for today's state does. Say that distinction out loud, because an interviewer who hears 57% quoted as a one-day forecast knows you have confused the stationary distribution with a conditional one. The limitation to add: a two-state chain with fixed probabilities is a toy, and real regime persistence drifts over time.
Where candidates lose it
The first wrong answer is 50%, on the grounds that each state has one way in and one way out. The flows are not equal in rate: trending leaks at 0.3, choppy at 0.4, and the slower leak wins more of the time.
The second loss is writing out eigenvectors of a two-by-two matrix under time pressure. The balance equation, flow out equals flow in, takes one line and is what the interviewer wants to hear.
What the interviewer asks next
- Starting from a choppy day, what is the chance that the day after tomorrow is trending?
- Add a third state, a crash day, that follows a choppy day 5% of the time. How does the method change?
- How would you estimate these transition probabilities from a year of daily data, and how noisy would they be?
005A stock trades at 100. Each day for three days it moves up 10 or down 10, each with probability one half, and interest rates are zero. What is a call struck at 100 and expiring after the third day worth?
Try it first
Before drawing the tree: how many distinct end prices are there, and how many equally likely paths?
Show the worked solution
7.50. Three moves of plus or minus 10 end at 130, 110, 90 or 70, reached by 1, 3, 3 and 1 of the eight equally likely paths. The 100 call pays 30 at 130, 10 at 110 and nothing below. Its value is the average payoff: (1 x 30 + 3 x 10) over 8, which is 60 over 8, or 7.50. With zero rates and symmetric moves, the real-world probabilities are already the pricing probabilities, so no discounting and no adjustment is needed.
Why is counting paths all the tree needs?
Think of three coin tosses where you get a sweet for each head. The chance of exactly two heads is not one in four; it is three in eight, because there are three orders in which two heads can arrive. The tree recombines, so the value of an end node is its payoff weighted by how many of the eight paths reach it, and the path counts are the binomial coefficients 1, 3, 3, 1. Nothing else in the problem carries information: the step size fixes the end prices and the counts fix the weights.
From 100, three moves of plus or minus 10 reach 130, 110, 90 or 70 by 1, 3, 3 and 1 of the eight equally likely paths, the call pays 30 and 10 at the two upper nodes and nothing below, and the average payoff (1 x 30 + 3 x 10) over 8 gives a price of 7.50. The relationshipk the number of up days out of three C(3, k) the number of paths with k up days: 1, 3, 3, 1 100 + 10(2k - 3) the end price after k ups and 3 minus k downs What it says in wordsThe call is the payoff at each end price, weighted by the share of paths that reach it.Where does the risk-neutral machinery go?
In a general tree you would replace the real probabilities with the risk-neutral ones, chosen so that the stock's expected growth equals the interest rate. Here rates are zero and the moves are symmetric, so the stock already has zero expected drift and the risk-neutral probability is the same one half you were given. Say that out loud: it shows you know the shortcut is a coincidence of the setup, not a rule. If the up move were 10 and the down move 5, one half would no longer price the stock and you would have to solve for the probability that does.
What sanity checks do you say before the number?
Two quick ones. The call cannot be worth more than the expected value of the stock above the strike ignoring the max, which is zero here, so the call is worth exactly the expected positive part, and that is what 7.50 is. And put-call parity with zero rates says the 100 put must also be 7.50, which you can confirm from the lower nodes: (3 x 10 + 1 x 30) over 8. Giving the put price unprompted, and showing it matches, is the cheapest way to prove the tree was right.
Where candidates lose it
The common error is to treat the four end prices as equally likely, which gives (30 + 10) over 4 = 10. The outer nodes are reached by one path each and the inner ones by three; the weights are 1, 3, 3, 1, not 1, 1, 1, 1.
The second loss is reaching for a risk-neutral formula and getting lost in it. With zero rates and symmetric moves, the given probabilities already price the stock. Say why, then count.
What the interviewer asks next
- Now the up move is 10 and the down move 5. What probability prices the stock, and what is the call worth?
- Price the 110 call and the 90 put on the same tree.
- Four days instead of three: what is the 100 call worth, and why does it rise?
007A stock goes up 10% one day and down 10% the next, and keeps alternating for 250 trading days. Where does it end relative to its start? And where does a fund that delivers three times the stock's daily move end up?Volatility tradingWealth management
Try it first
Before multiplying: after one up day and one down day, is the stock back where it started?
Show the worked solution
The stock ends at about 28% of its start; the three-times fund at roughly 8 millionths of its start, effectively zero. Each up-and-down pair multiplies the stock by 1.1 x 0.9 = 0.99, and 125 pairs give 0.99^125 = 0.285. The leveraged fund moves 30% each way, so each pair is 1.3 x 0.7 = 0.91, and 0.91^125 is about 7.6e-06. The arithmetic average return is zero in both cases; the compounded return is not.
Why does a zero average return lose money?
Take a 100 rupee note to a shop that marks everything up 10% in the morning and discounts 10% in the afternoon. The afternoon discount is taken off a bigger number, so the price ends at 99, not 100. A gain and a loss of the same percentage do not cancel, because the loss acts on the larger base; the pair costs the square of the move, 1% for a 10% swing. Repeat that 125 times and the 1% losses compound to a 72% fall. This is volatility drag: the gap between the average return and the compounded return.
On a log scale both paths step down in straight lines: the stock loses 1% per up-and-down pair and ends at 28% of its start after 250 days, while the three-times fund loses 9% per pair and ends at about 8 millionths of its start, so leverage multiplies the drag by far more than three. The relationshipm the daily move, 0.10 for the stock and 0.30 for the three-times fund 1 - m^2 what one up-and-down pair leaves of the value 125 the number of pairs in 250 days What it says in wordsEach pair loses the square of the move, and the leveraged fund's loss per pair is nine times the stock's because 0.3 squared is nine times 0.1 squared.Why is three times the move so much worse than three times the loss?
The drag per pair is the square of the move. Tripling the move multiplies the drag by nine, not three: the stock loses 1% per pair, the fund loses 9%. That is why a daily-rebalanced leveraged fund in a choppy, sideways market bleeds even when the underlying ends flat. The general rule you can quote: over many periods the compounded growth rate is roughly the average return minus half the variance, and leverage multiplies the variance by the square of the leverage.
What would you say to a client who holds the three-times fund?
That the product tracks three times the daily move, exactly as promised, and that this is not the same as three times the return over a year. A leveraged fund is a tool for a view on the next day or week; held through a sideways year it loses to its own rebalancing. The limitation to state: the alternating path is the worst case for drag, and a strongly trending market can make a leveraged fund return more than three times the underlying. The drag is about path, not just direction.
Where candidates lose it
The common answer is that the stock ends flat, because plus 10 and minus 10 seem to cancel. They cancel in arithmetic and not in compounding; the second move acts on a different base.
The second loss is saying the leveraged fund ends at three times the stock's loss, or at 28% cubed. The right route is per pair: 1.3 x 0.7 = 0.91, then raise to the 125th power. The drag scales with the square of the leverage.
What the interviewer asks next
- Make the daily move 1% instead of 10%. Where does the stock end after 250 days?
- Over a year with 16% annual volatility and zero average daily return, roughly what is the compounded return?
- Why do leveraged funds rebalance daily, and what would change if they rebalanced monthly?
010In some stock, the 99-strike call trades at 5.60 and the 101-strike call at 4.70, same expiry. Estimate the price of a digital option that pays 1 if the stock finishes above 100 at that expiry.Exotics tradingStructured products
Try it first
Before any arithmetic: which combination of the two calls has a payoff that looks most like a step at 100?
Show the worked solution
About 0.45. Buying the 99 call and selling the 101 call pays 0 below 99, 2 above 101 and a straight ramp between. Divide that by the width of 2 and the payoff is 0 below 99, 1 above 101 and a ramp through 100: a digital with its edge smoothed over two points. Its cost is (5.60 minus 4.70) over 2, which is 0.45. The narrower the spread, the closer the ramp sits to the step, and the price converges to the digital.
Why does a call spread stand in for a digital?
A light switch is a step: off or on. A dimmer that goes from fully off to fully on over a tiny turn of the knob is, for every practical purpose, the same switch. A call spread over its width is a dimmer: it ramps from 0 to 1 across the two strikes, and as the strikes close in on 100 the ramp becomes the step. So the digital is the limit of a scaled call spread, and a traded call spread gives you a price for it without any model.
The 99 to 101 call spread divided by its width of 2 pays 0 below 99, ramps to 1 at 101 and crosses the digital's step exactly at 100, so the two payoffs differ only inside the narrow band between the strikes and the spread's price, (5.60 minus 4.70) over 2, gives a digital value of 0.45. The relationshipC(K) the price of a call struck at K D(K) the price of a digital paying 1 above K (C(99) - C(101)) / 2 the slope of the call price in strike, estimated across 100 What it says in wordsThe digital is minus the slope of the call price with respect to strike, and a centred call spread measures that slope.Is 0.45 the digital's price or an approximation, and which way is it off?
It is an approximation to the slope at 100 taken from two points either side. Because the spread is centred on 100, the first-order error cancels and what remains is small, of the order of the curvature of the call price between 99 and 101. If the digital were struck at 99 instead, the same spread would overstate it, because the call price is convex in strike and the ramp sits above the step on that side. On a desk you would quote the digital from the tightest spread the market will show you, and hedge it with that spread, so the approximation is also the hedge.
What does 0.45 say about the market, and what is the limitation?
A digital paying 1 above 100 at 0.45, with rates near zero, means the pricing probability of finishing above 100 is about 45%, slightly below one half. That is a risk-neutral probability, not a forecast, and it is pulled down by the skew: with a steeper put skew, out-of-the-money calls are cheaper in volatility terms and the slope in strike is steeper, which moves the digital. Say that a flat-volatility formula would miss this, and that the call spread picks the skew up automatically because it uses the two traded prices.
Where candidates lose it
The common error is to take the difference of the two call prices, 0.90, and present it as the digital. That is the price of a spread that pays 2 above 101, not 1. Divide by the width.
The second loss is reaching for a lognormal formula with a guessed volatility. The question gives you two traded prices precisely so you can price the digital without a model; use them.
What the interviewer asks next
- The 99.5 and 100.5 calls are 5.37 and 4.93. What does that pair say about the digital, and why might it differ from 0.45?
- How would you hedge a short digital you sold at 0.45, and what goes wrong near expiry?
- Price a digital that pays 1 if the stock finishes below 100.
016You walk into a casino with Rs 63,000 and bet Rs 1,000 on red at even money, where red comes up 48% of the time. Every time you lose, you double the bet. You stop at the first win, or when you cannot cover the next bet. What is the chance you lose everything, and what is your expected result?Risk managementProp trading firms
Try it first
Before any arithmetic: the plan ends a session up Rs 1,000 about 98 times in 100. What is its expected result per session?
Show the worked solution
You lose everything about 2.0% of the time, 0.52 to the sixth power, and the expected result is about minus Rs 265. Rs 63,000 covers exactly six bets: 1, 2, 4, 8, 16 and 32 thousand. A win at any of them recovers every earlier loss and nets Rs 1,000, which happens 98.0% of the time. Six losses in a row cost all Rs 63,000. Weighted, 980 of expected winnings against 1,246 of expected loss leaves minus Rs 265.
Why does a plan that wins 98 times in 100 still lose money?
Picture a friend who sells phone insurance to classmates for Rs 50 a month. Month after month nobody drops a phone, and the Rs 50 notes pile up; it feels like free money until the month three phones go into a pond. A win rate tells you how often you are paid, not how much you are paid against how much you can lose, and the expected value needs both. Doubling after every loss builds exactly that shape: Rs 1,000 collected almost every time, and Rs 63,000 handed back rarely. The rare branch is 63 times the size of the common one, so a 2% chance of it more than cancels a 98% chance of the small win.
The doubling plan ends a session up Rs 1,000 with probability 98.0% and down Rs 63,000 with probability 2.0%, and weighting the two gives plus 980 against minus 1,246, an expected result of minus Rs 265, which is also 4% of the Rs 6,633 the plan expects to stake. How do you lay out the six bets in the room?
Write the ladder down before computing anything. The stakes are 1, 2, 4, 8, 16 and 32 thousand, which add to 63 thousand exactly, so the seventh bet of 64 thousand can never be placed. If the first win comes at bet k, it pays 2 to the power k minus 1 thousand, and the losses before it add to one thousand less than that, so every winning session nets exactly plus Rs 1,000. There are only two outcomes, and the table shows how quickly the chance of reaching each rung falls: by the sixth bet you are staking Rs 32,000 to recover Rs 31,000 of losses and win one more thousand.
Bet Stake (Rs) Lost before it (Rs) Chance of reaching it 1 1,000 0 100.0% 2 2,000 1,000 52.0% 3 4,000 3,000 27.0% 4 8,000 7,000 14.1% 5 16,000 15,000 7.3% 6 32,000 31,000 3.8% Each rung doubles the stake while the chance of reaching it falls by a factor of 0.52, and the chance of losing the sixth bet as well is 1.98%, the probability of ruin. The relationship0.52^6 the chance of six losses in a row, about 2% +1,000 the net result of any session that wins before the money runs out -63,000 the whole bankroll, lost when all six bets lose What it says in wordsThe expected result is the frequent small win times its probability plus the rare total loss times its probability, and the second term is larger.Is there a faster way to see the sign without the ladder?
Yes, and it is the one a trader reaches for first. Every rupee placed on red loses 4 paise on average, whatever happened on the previous spin, because the wheel has no memory. The expected result of any staking plan is the edge per rupee times the expected total amount staked: here minus 4% of Rs 6,633, which is minus Rs 265, the same figure as the ladder. Doubling raises the amount you put down when you are losing; it cannot change the sign of the edge. On a fair 50/50 wheel the same plan has an expected value of exactly zero, with the same lopsided shape.
Why does a desk interviewer care about a roulette plan?
Because the shape is the shape of selling far out-of-the-money options, or of adding to a losing position to get back to flat. Both produce a long run of small gains and a rare large loss, and a good-looking track record says almost nothing about the tail. Repetition makes the rare branch common: play 50 sessions and the chance of at least one ruin is 1 minus 0.98 to the 50th, about 63%. The limitation to state is that the plan assumes no table limit; a casino maximum bet cuts the ladder short and makes ruin more likely, not less.
Where candidates lose it
The common answer is that the plan wins, because it almost always wins. Candidates quote the 98% and stop, never weighing it against the size of the 2% branch. A probability without a payoff is half an expected value.
The second loss is the opposite slip: computing minus 4% of the Rs 63,000 bankroll, about minus Rs 2,520. The edge applies to rupees actually staked, and most sessions stake only Rs 1,000 or Rs 3,000 before the first win. Expected stake, Rs 6,633, is the base.
What the interviewer asks next
- The wheel is fair, 50/50. What is the expected result now, and what is the chance of ruin?
- You have unlimited money but the table caps any single bet at Rs 16,000. How does the picture change?
- Name a trading strategy with the same payoff shape, and say how you would size it.
025A stock trades at 500. The one-year 500-strike call is priced at 40 and the 500-strike put at 45, and interest rates are zero. What dividend is the options market pricing in, and what would you check before trading on it?Equity derivativesMarket making
Try it first
Before using any formula: the put costs more than the call at the same at-the-money strike. With zero rates, that tells you
Show the worked solution
A dividend of 5 per share. With zero rates, put-call parity says call minus put equals spot minus the dividend minus the strike. Here 40 minus 45 = 500 minus D minus 500, so D = 5, and the implied forward is 495. Before trading against it, check that the options are European or that early exercise is worth nothing, what it costs to borrow the stock, whether the ex-date falls before expiry, and whether all four prices are live and tradable.
Why does a put costing more than a call point to a dividend?
Suppose you agree today to buy a friend's scooter in a year for its fair price, but the friend will keep using it to deliver parcels and pocket the fees until then. You would pay less than today's price, by the fees they collect. Holding a call and selling a put at the same strike is an agreement to buy the stock at the strike, so call minus put must equal the forward price minus the strike, and the forward is spot minus whatever the holder collects before expiry. With zero rates there is no interest to account for, so a put dearer than the call by 5 says the forward is 5 below spot.
The relationshipC, P the call and put prices at the same strike and expiry S the spot price, 500 K the strike, 500 D the cash the stock pays out before expiry, here the implied dividend What it says in wordsWith zero rates, the call minus the put equals the spot less the dividend less the strike, so the dividend is what balances the equation.Call minus put is minus 5, so the options imply a forward of 495, and since the stock's forward is 500 minus the dividend, the market is pricing a dividend of 5; if the company pays only 3, the remaining 2 has to be explained by something else, such as the cost of borrowing the stock. What would you check before trading on it?
Say your own estimate of the dividend is 3. Then a reversal looks attractive: short the stock, buy the call, sell the put, collecting 500 + 45 minus 40 = 505; at expiry the options deliver the stock back at 500 and you owe the 3 dividend, 503 in all, a locked-in 2 per share. That 2 is only yours if you can borrow the stock for less than 0.4% of its price over the year; a hard-to-borrow stock shows up in parity exactly as an extra dividend. Then check the rest: American calls can be exercised just before a dividend, the ex-date must fall before expiry, and four bid-offer spreads can eat a gap of 2 on their own.
What is the general lesson the interviewer wants?
That parity is an accounting identity between three things you can trade, and any gap between what it implies and what you believe is a claim about something you have not yet priced. The options market does not quote a dividend; it quotes a forward, and the dividend, the borrow cost and the interest rate are the pieces you split it into. A good answer gives the 5, names the implied forward of 495, and then lists what could make the 5 something other than a dividend. The limitation to say plainly: with non-zero rates, the strike is discounted and the arithmetic shifts, so state the zero-rate assumption before quoting the number.
Where candidates lose it
The common answer is that the put is dearer because the market expects the stock to fall. At one strike and one expiry, direction cannot make the put dearer than parity allows; anyone could sell the put, buy the call and short the stock against it. The gap is a forward, not a view.
The second loss is stopping at 5. The question asks what you would check, and the borrow cost is the one interviewers wait for: an implied dividend above the announced one is often a stock that is expensive to short.
What the interviewer asks next
- Interest rates are now 6% a year. Redo the implied dividend, and say which way it moves.
- The company announces a dividend of 8. Which trade would you put on with these four prices, and what is the risk?
- Why might an American call on this stock be worth more than its European twin, and how would that distort the implied dividend?
032One-month implied volatility is 20% and three-month implied volatility is 25%. What volatility is implied for the period from month one to month three?Volatility tradingEquity derivatives
Try it first
What is the forward volatility for months two and three?
Show the worked solution
About 27.2%. Variance is volatility squared times time, and variances add across periods. The three months carry 0.25 squared x 3 = 0.1875 of variance; the first month carries 0.20 squared x 1 = 0.04; so months two and three carry 0.1475 between them, 0.0737 per month on an annualised basis. The square root is 0.272. The forward volatility must sit above 25% because the quarter's average has to be pulled up from the 20% start.
Why do you add variances and not volatilities?
If you walk a random distance each hour, your spread after three hours is not three times the hourly spread, because some hours cancel others. What grows in a straight line is the variance, the spread squared. Implied volatility is quoted per year, so the variance an option carries is volatility squared times its life, and a longer option's variance is the sum of the variances of the pieces of time inside it. That is the whole mechanism. Twenty per cent for one month is 0.04 of variance; twenty-five per cent for three months is 0.1875; the difference belongs to the two months in between, and dividing by two months and taking the square root turns it back into a volatility.
Drawn as area, the first month's variance of 0.04 plus the two forward months' variance of 0.0737 each must fill the quarter's 0.1875, so the forward box has height 0.0737 and its square root is a forward volatility of 27.2%. The relationshipsigma 3, sigma 1 the three-month and one-month implied volatilities, 25% and 20% T the option life in months; the units cancel as long as both use the same one sigma 1,3 the forward volatility for the period between the two expiries What it says in wordsSubtract the short period's variance from the long period's and spread what is left over the time in between.What does the number tell a trader?
The 27.2% is the volatility you would lock in for months two and three by selling the one-month option and buying the three-month one in variance-weighted sizes, a calendar spread. If you believe realised volatility in those two months will be well below 27.2%, the three-month option is rich relative to the one-month, and the calendar is the trade that expresses it. Say the rounding: an upward sloping term structure, 20% then 25%, hides a forward that is steeper than either quote, 27.2%, and the trap of reading 25% as the forward is what the question is built to catch. The calculation uses calendar time in months; a desk would use trading days or variance-weighted business days, which shifts the number by a few tenths.
What is the limit of the calculation?
The subtraction has to leave something positive. If the three-month volatility were below 11.5% with the one-month still at 20%, the quarter would carry less variance than its first month alone, which is impossible without arbitrage: you could sell the one-month, buy the three-month and hold a position with negative forward variance. Forward variance can be small but never negative, so a term structure that inverts too sharply is a mispricing, not a forecast. The second limit is that both quotes must refer to the same strike in forward terms; mixing an at-the-money one-month with a three-month that has rolled away from the money adds skew to the comparison and the forward you compute is no longer clean.
Where candidates lose it
The common loss is averaging volatilities: 25% over three months with 20% in the first month gives 27.5% for the rest if you treat volatility as additive. It is close, which is why it survives, but it is the wrong quantity and on a steeper curve the gap is large.
The second is forgetting to weight by time: subtracting 0.04 from 0.0625 and taking the root. Write variance times time for each leg, every time.
What the interviewer asks next
- Six-month volatility is also 25%. What is the forward volatility for months four to six?
- The one-month is 30% and the three-month is 25%. What is the forward, and what does it say about the market?
- How would you actually lock in that forward volatility with listed options, and what would break the hedge?
- Why do desks compute this in trading days rather than calendar months?
038Box A holds 3 red and 1 black ball, box B holds 1 red and 3 black. You pick a box at random and draw a red ball. You then draw again from the same box without replacement. What is the probability the second ball is red?Prop trading firmsMarket making
Try it first
Pick the probability of a second red before you compute.
Show the worked solution
Exactly 1/2. The first red is evidence: box A produces a red with probability 3/4, box B with 1/4, so after seeing red the odds are 3 to 1 for A, a posterior of 3/4. Then update the contents. Box A has 2 red of 3 left, so a second red has probability 2/3; box B has 0 red of 3, probability 0. Average over the updated boxes: 3/4 x 2/3 + 1/4 x 0 = 1/2. Forgetting the update gives 1/3, which is wrong.
Why does the first red change which box you think you hold?
If a friend's cooking is spicy three times out of four and yours one time in four, and a dish picked at random from one of you turns out spicy, you would guess it was your friend's, and you would be right three times in four. The first red ball is the same clue: it comes from box A three times as easily as from box B, so the box you are holding is A with probability 3/4, not 1/2. The arithmetic is Bayes. The weight for A is 1/2 x 3/4 = 3/8, for B it is 1/2 x 1/4 = 1/8, and A's share of the total 1/2 is 3/4. The question is built so that candidates who skip this update get a clean wrong answer.
After a red first draw the box is A with posterior 3/4 and B with 1/4, box A has 2 red of 3 left and box B none, so the chance of a second red is 3/4 x 2/3 plus 1/4 x 0, which is exactly one half. What has to be updated besides the belief about the box?
The box itself. One red ball has left it. Box A now holds 2 red and 1 black, so a red is 2/3; box B now holds 0 red and 3 black, so a red is impossible. Both updates matter and they pull in opposite directions: believing more in A pushes the answer up, removing a red from A pushes it down, and here they land on exactly 1/2. Say that out loud, then check it by counting. Imagine 8 rounds, 4 with each box. Box A yields a first red in 3 of its 4 rounds and box B in 1 of its 4: four red first draws in all. In the three from A the second draw is red 2/3 of the time, 2 rounds; in the one from B, never. Two of four is 1/2.
The relationshipR1, R2 the first and second draws are red P(A | R1) the posterior that the box is A after one red, 3/4 P(R2 | A, R1) a second red from box A with one red already gone, 2/3 What it says in wordsWeight each box's chance of a second red by how likely that box is after the first red.How do the variants separate candidates who understand from those who memorised?
Put the first ball back and the box contents do not change but the belief still does: 3/4 x 3/4 + 1/4 x 1/4 = 5/8, which is above 1/2 because the posterior leans to A and A stays rich in red. Forget the belief update and keep the contents update, and you get 1/2 x 2/3 + 1/2 x 0 = 1/3. The gap between 1/3, 1/2 and 5/8 is the whole lesson: the first draw tells you about the box and changes the box, and you need both. If a second red does appear, the posterior for A becomes 1, since B has no reds left, and a third red is then 1/2. The limitation is the usual one for Bayes: the 1/2 prior for each box is given here, and in a trading setting the prior is the thing you are least sure of.
Where candidates lose it
The common loss is 1/3: keeping the 1/2 prior on each box and only updating the contents. The first red is evidence about the box, and leaving it out throws away half the information in the question.
The second is 2/3: updating to box A and then forgetting that B is still possible with probability 1/4, or forgetting that A has lost a red. Carry both boxes through to the end, then average.
What the interviewer asks next
- The first ball is put back before the second draw. What is the probability now?
- The second ball is also red. What is the probability the box is A, and that a third draw is red?
- Box B held 0 red and 4 black instead. What changes, and what is the answer?
- A trader sees one winning trade from a new signal. What is the analogue of the box update?
044Using only calls, build a payoff that is zero below 90, rises one for one to 10 at 100, falls back to zero at 110, and stays at zero above that.Equity derivativesStructured products
Try it first
Which call portfolio gives the tent?
Show the worked solution
Long one 90 call, short two 100 calls, long one 110 call: a call butterfly. Read the slope of the target from left to right: 0, +1, -1, 0. A long call adds +1 to the slope at its strike, so the changes of +1 at 90, -2 at 100 and +1 at 110 give the weights. Check the corners: 0 at 90, 10 at 100, 0 at 110 and above. At 20% volatility and three months it costs about 3.69.
How do you read a payoff picture as a list of calls?
A road that is flat, then climbs, then drops, then is flat again can be described by where the gradient changes and by how much. A payoff made of straight pieces is the same: each call adds one unit of slope from its strike onwards, so the number of calls at a strike is simply the change of slope at that strike. The tent has slope 0 below 90, +1 from 90 to 100, -1 from 100 to 110, and 0 above 110. The changes are +1 at 90, -2 at 100 and +1 at 110. So you buy one 90 call, sell two 100 calls and buy one 110 call. Above 110 the three legs pay (S - 90) - 2(S - 100) + (S - 110) = 0, which confirms the payoff returns to zero and stays there.
The long 90 call, the two short 100 calls and the long 110 call add up to a tent that is zero below 90, peaks at 10 at 100 and returns to zero from 110 onwards, because the slope changes by +1, -2 and +1 at the three strikes. The relationship(S - K)+ the payoff of a call struck at K, the larger of S - K and zero weight at K the number of calls to hold at that strike; negative means sell S the stock price at expiry What it says in wordsThe weight on each strike is the jump in slope there, which is how any straight-line payoff is built from calls.What does the butterfly cost, and what does its price tell you?
Take an illustrative stock at 100 with 20% volatility and three months to expiry. The calls cost 10.71, 3.99 and 0.95, so the butterfly costs 10.71 - 2 x 3.99 + 0.95 = 3.69. The payoff is a tent of height 10 and base 20, and its price is close to the chance of finishing near 100 times the peak. Divide the cost by the peak payoff of 10 and you get 0.369, close to the 0.383 risk-neutral chance that the stock ends between 95 and 105, which is why a butterfly is the market's way of pricing the probability of a narrow range. Shrink the strike gap towards zero and the scaled butterfly becomes the risk-neutral density itself, a result traders use to read the distribution off a strip of call prices.
Where does the no-arbitrage check come in?
The tent never pays less than zero, so it can never cost less than zero. That means C(90) - 2C(100) + C(110) must be at least 0: call prices must be convex in strike. Suppose a screen shows the 90 call at 11.00, the 100 call at 7.00 and the 110 call at 2.50. The butterfly costs 11.00 - 14.00 + 2.50 = -0.50: you are paid 0.50 to hold a payoff that is never below zero. A negative butterfly price is a free lunch, and spotting it on a quote sheet is the reason the question is asked. The limitation in practice is that each leg has a bid and an offer, so the check must use the prices you can actually trade at, buying at offers and selling at bids, and small apparent violations usually vanish once the spread is paid.
Where candidates lose it
The common loss is short one 100 call instead of two. One short call only cancels the slope of the 90 call, which flattens the payoff at 10 forever; it takes two to turn the slope down to -1 and bring it back to zero.
The second is building the tent and stopping. The interviewer usually follows with the price: a butterfly must cost more than zero because it never pays less than zero, and a candidate who connects that to convexity in strike has answered the real question.
What the interviewer asks next
- Build the same tent with puts only. Is the cost the same?
- Build a payoff that is 0 below 90, rises to 10 at 100 and stays at 10 above.
- The 90, 100 and 110 calls trade at 11.00, 7.00 and 2.50. What do you do?
- As the gap between the strikes shrinks, what does the scaled butterfly price approach?
046Give the next term in each sequence: 2, 6, 12, 20, 30, ... ; 1, 1, 2, 6, 24, ... ; 3, 5, 9, 17, 33, ...Prop trading firmsSell-side sales and trading
Try it first
What comes after 3, 5, 9, 17, 33?
Show the worked solution
42, 120 and 65. In the first, the differences 4, 6, 8, 10 rise by 2, so the next difference is 12 and the term is 42; the terms are n(n + 1). In the second, each term is the previous one times 1, 2, 3, 4, so the next multiplier is 5 and the term is 120; these are the factorials. In the third, the differences 2, 4, 8, 16 double, so add 32 to get 65; the terms are 2 to the n plus 1.
What order should you test patterns in?
A mechanic with an unknown rattle checks the cheap, common causes first and the exotic ones last. Sequence questions reward the same discipline: write the first differences, then the second differences, and only if neither settles try ratios and then rules that mix the two, such as double and subtract one. Each test takes a few seconds and most interview sequences give way to one of the first three. The order matters because the interviewer is timing you; the candidate who stares at the numbers hoping to recognise them is slower than the one who writes a row of differences under them without thinking.
Writing a row under each sequence exposes the rule: the first has constant second differences of 2 and continues to 42, the second has ratios climbing by one and continues to 120, and the third has doubling differences and continues to 65. How does each of the three give way?
For 2, 6, 12, 20, 30, the first differences are 4, 6, 8, 10 and the second differences are all 2. Constant second differences mean the sequence is a quadratic in n, so the next difference is 12 and the next term 42, and the closed form is n(n + 1): 1 x 2, 2 x 3, up to 6 x 7. For 1, 1, 2, 6, 24, the differences 0, 1, 4, 18 tell you nothing, so try ratios: 1, 2, 3, 4. The next ratio is 5 and the term is 120; the terms are 0!, 1!, 2!, 3!, 4!, and 5! is 120. For 3, 5, 9, 17, 33, the differences are 2, 4, 8, 16, a doubling, so the next difference is 32 and the term is 65. Spot the rule in one line too: each term is twice the previous minus 1, and every term is a power of two plus 1, so the next is 2 to the 6 plus 1.
The relationshipa n the n-th term of the first sequence, a quadratic, so its second differences are constant b n the factorials, each term the previous one times the next whole number c n a power of two plus one, so its differences are powers of two What it says in wordsEach sequence has a one-line rule, and the difference or ratio row is how you find it in seconds.Is the answer really unique, and what should you say if pushed?
Strictly, no. Any five numbers can be continued by any sixth, because a polynomial of degree five can be passed through all six points. For example n(n + 1) + (n - 1)(n - 2)(n - 3)(n - 4)(n - 5) matches 2, 6, 12, 20, 30 exactly and then gives 162. The expected answer is the simplest rule that fits, and the way to show you know that is to name the rule, not just the number. That habit carries over to the desk: a pattern in five data points is a hypothesis, and the trader who states the rule can test it on the sixth point, while the one who only extrapolates cannot tell when the pattern has broken. If the interviewer offers a sequence that resists all three tests, try alternating terms, or interleaved sequences, before guessing.
Where candidates lose it
The common loss is 66 for the third sequence: doubling 33 and forgetting that the rule doubles and then subtracts 1. Check the rule on an earlier pair, 17 to 33, before you say the answer.
The second is hunting the factorials through differences, which give 0, 1, 4, 18 and lead nowhere. When differences grow faster than the terms, switch to ratios immediately.
What the interviewer asks next
- What comes next: 1, 4, 9, 16, 25, 36, and what are its second differences?
- Next term: 1, 2, 6, 15, 31, ...
- Next term: 2, 3, 5, 7, 11, 13, ...
- Find a rule for 1, 3, 7, 15, 31 and give its closed form.
