Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
091Cards are turned over one at a time from a well-shuffled 52-card deck. On average, at what position does the first ace appear?Quant and systematic fundsProp and quant trading firms
Try it first
What is the expected position of the first ace?
Show the worked solution
Position 10.6. The four aces split the 48 non-aces into five gaps: before the first ace, between each pair of aces, and after the last. By symmetry, a given non-ace is equally likely to fall in any of the five gaps, so each gap holds 48/5 = 9.6 cards on average. The first ace comes straight after the first gap, at 9.6 + 1 = 10.6.
Why do the gaps have the same size on average?
Picture four red beads threaded at random among 48 white ones on a string. Pick any one white bead: it is as likely to sit before all four reds as between the first and second, or after the last, because a shuffle favours no order. Each non-ace lands in each of the five gaps with probability 1/5, so the expected size of every gap is 48/5 = 9.6 cards. Any single shuffle has uneven gaps; it is only the average that is equal.
One shuffle leaves uneven gaps of 4, 5, 14, 11, 14 non-aces around its four aces, but averaged over all shuffles every one of the five gaps holds 9.6 cards, so the aces sit at 10.6, 21.2, 31.8, 42.4. How do you turn gap sizes into a position?
The first ace sits immediately after the first gap. Expected position = expected size of the first gap + 1 = 9.6 + 1 = 10.6. The same logic puts the second ace at 21.2, the third at 31.8 and the last at 42.4, evenly spaced 10.6 apart. A quick check: the last ace at 42.4 plus a final gap of 9.6 reaches exactly 52.
The relationshipn cards in the deck, 52 k special cards, here the 4 aces (n - k)/(k + 1) the expected number of ordinary cards in each of the k + 1 gaps What it says in wordsWith n cards and k special ones, the first special card is expected at position (n + 1) divided by (k + 1).Why is 13 the wrong instinct?
Thirteen is the average wait when every draw has a fresh 1-in-13 chance, as if each card were put back and the deck reshuffled. Drawing without replacement makes the aces easier to reach, because every non-ace you turn over raises the share of aces in what is left. The general rule, (n + 1)/(k + 1), is worth keeping: the first spade is expected at 53/14, about card 3.79, and the same shape answers how many trades you expect to check before finding the first of a handful of booking errors in a batch.
Where candidates lose it
The fast wrong answer is 13, borrowed from a wait with replacement. The interviewer wants to see you notice that the deck is finite, so the first ace must appear by card 49 at the latest and is pulled earlier than a with-replacement wait would suggest.
The second loss is starting the direct sum, position times the chance the first ace is there, which is correct but slow and easy to get wrong at a whiteboard. The gap argument gets there in one line.
What the interviewer asks next
- What is the expected position of the last ace?
- What is the expected number of cards between the first and second ace?
- How many cards do you expect to turn before the first spade?
092A casino offers the St Petersburg game: a fair coin is tossed until the first tail, and if that takes n tosses you are paid Rs 2 to the power n. The casino can pay out at most Rs 1 crore. What is a fair price to play?Quant and systematic fundsProp and quant trading firms
Try it first
Roughly what is the capped game worth?
Show the worked solution
About Rs 24.19. Round n pays Rs 2^n with probability 1/2^n, so each round adds exactly Rs 1 of expected value. That holds up to n = 23, since 2^23 is about Rs 84 lakh and 2^24 is above the Rs 1 crore cap. From round 24 on, the payout is stuck at Rs 1 crore, with total probability 1/2^23, which adds about Rs 1.19. So 23 + 1.19, about Rs 24.
Why is the uncapped game worth an infinite amount?
Each round is a doubling bet: the payout doubles while the chance of reaching it halves. Think of a raffle where a ticket twice as valuable is half as likely to win: every prize tier is worth the same to you. Round n pays 2^n with probability 1/2^n, so every round contributes exactly Rs 1 to the expected value, and there are infinitely many rounds. That is the famous paradox: few people would pay even Rs 100, yet the expected value has no bound. The resolution that matters on a desk is not psychology; it is that nobody can pay out an unlimited amount.
Rounds 1 to 23 each add exactly Rs 1 to the expected value, and once the Rs 1 crore cap binds the later rounds add 0.60, 0.30, 0.15 and so on, Rs 1.19 in all, so the capped game is worth about Rs 24.19. How does the cap change the sum?
Find the round where the cap starts to bind. 2^23 is Rs 83,88,608, under Rs 1 crore; 2^24 is Rs 1,67,77,216, over it. So rounds 1 to 23 each add Rs 1, and every round from 24 on pays the capped Rs 1 crore, which together happen with probability 1/2^23 and add 1,00,00,000 / 83,88,608, about Rs 1.19. The fair price is about Rs 24.19. Nearly all of the textbook infinity lives in outcomes the casino cannot pay.
The relationship2^n the payout if the first tail arrives on toss n 1/2^n the chance the first tail arrives on toss n 10^7 / 2^23 the capped Rs 1 crore times the chance of reaching round 24 or later What it says in wordsEvery uncapped round is worth one rupee; the capped tail is worth the cap times the chance of getting that far.What does a bigger casino buy you?
Very little. The value grows only with the logarithm of the cap: each doubling of the casino's bankroll adds about Rs 1. A cap of Rs 1,000 crore, a thousand times larger, lifts the fair price only to about Rs 34. That is the lesson a risk manager takes away: a payoff whose expected value rests on rare, enormous outcomes is worth what the other side can actually pay, and any estimate built on the tail should be checked against who stands behind it.
Where candidates lose it
Answering infinity is the trap for anyone who knows the textbook game. The interviewer added the cap to see whether you can find where it binds and redo the sum, not recite the paradox.
The other loss is dropping the tail beyond the cap and saying Rs 23, or valuing it crudely at the full Rs 1 crore times a guessed chance. The tail is a clean sum: probability 1/2^23 of receiving Rs 1 crore.
What the interviewer asks next
- The cap rises to Rs 1,000 crore. What is the fair price now?
- With logarithmic utility and wealth of Rs 1 lakh, roughly what would you pay for the uncapped game?
- How is a book short deep out-of-the-money options like the casino in this game?
093Four people must cross a narrow bridge at night with one torch. They take 1, 2, 5 and 10 minutes to cross, at most two can cross at a time, a pair moves at the slower person's pace, and the torch must be carried on every crossing. What is the fastest time for all four to get across?Prop and quant trading firmsLong-short equity funds
Try it first
What is the fastest crossing?
Show the worked solution
17 minutes. 1 and 2 cross (2 minutes), 1 returns (1), 5 and 10 cross together (10), 2 returns (2), and 1 and 2 cross again (2). The total is 17. The obvious plan, with the fastest person escorting everyone, takes 19, because the 5 and the 10 each cost a separate crossing. Pairing the two slowest hides the 5 inside the 10.
Why is the obvious plan not the fastest?
The natural plan uses the quickest person as a shuttle: 1 walks each person over and comes back. That costs 2 + 1 + 5 + 1 + 10 = 19. Every slow person who crosses on a separate trip pays their own time in full, so the 5 and the 10 together cost 15 minutes of crossing. Think of sending two slow parcels in one courier van instead of two: the second rides along for free.
The shuttle plan spends 5 and 10 minutes on separate crossings and takes 19 minutes, while sending 5 and 10 together costs 10 minutes for both and brings the total to 17, even after the 2-minute person makes one extra return trip. How does the 17-minute plan pay for pairing the slow two?
Pairing 5 and 10 costs 10 instead of 15, a saving of 5. But someone fast must already be waiting on the far side to bring the torch back afterwards. That setup costs one extra return by the 2-minute person and an extra crossing of the pair 1 and 2, so the net saving is 2 minutes: 19 down to 17. The general rule: pair the two slowest when twice the second-fastest time is less than the fastest plus the second-slowest, here 4 against 6.
The relationship10 the one crossing that carries both the 5-minute and the 10-minute person 2 + 2 the extra cost of having the 2-minute person bring the torch back and cross again What it says in wordsPairing the two slowest saves 5 minutes of crossing at a cost of 3 extra minutes of torch returns and re-crossing.How do you show 17 is the minimum?
Five crossings are unavoidable: three over and two back. If the 5 and the 10 cross separately, the forward trips cost at least 10 + 5 + 2 and the two returns at least 1 each, which is 19; if they cross together, the person who returns next must already be across, which forces the returns to be the 1 and the 2, and the best total is 17. Talking through that bound is what separates a remembered answer from a reasoned one, and it is what lets you handle the follow-up with different times.
Where candidates lose it
19 minutes is the answer most people reach, and it comes so quickly that they stop there. The interviewer is waiting to see whether you ask what the 5-minute person costs and whether that cost can be hidden.
The second loss is finding 17 by trial and error without the reason. Name the saving, 5 minutes from pairing, and its price, the extra return and re-crossing; that is what makes the answer hold when the interviewer changes the times.
What the interviewer asks next
- The times are 1, 5, 6 and 10. Does pairing the two slowest still help?
- Add a fifth person who takes 20 minutes. What is the fastest crossing now?
- State the general rule for when to pair the two slowest walkers.
094You buy an at-the-money option at 20% implied volatility and delta-hedge it. Over its life the stock realises 30% volatility. Where does your P&L come from, and does the direction of the moves matter?Volatility and relative value fundsProp and quant trading firms
Try it first
What drives the hedged position's P&L?
Show the worked solution
From gamma: the hedged option earns on the size of moves, not their direction, and pays theta for the privilege. Each day the P&L is about half gamma times the squared move, minus theta. Theta is priced for 20% volatility, so a daily move of about 1.26% breaks even. At 30%, typical moves are 1.89%, and the squared gain is 2.25 times the decay. Up or down makes no difference; the path still does.
What is left after the delta hedge?
Think of a street vendor who sells umbrellas and sunglasses from the same cart: she no longer cares whether it rains or shines, only whether the weather changes enough to bring people out. A delta hedge sells enough stock to cancel the option's first-order bet on direction, so what remains is the curvature: the option gains a little more on the way up than the hedge loses, and loses a little less on the way down than the hedge gains. That curvature is gamma, and it pays on moves either way.
The option's value curves above the straight delta hedge line on both sides of Rs 1,000, so a move either way earns the shaded gap; a day at 30% realised volatility earns about Rs 0.71 of it per option against Rs 0.32 of time decay priced at 20%. How big is the daily P&L in numbers?
Take a stock at Rs 1,000 and a three-month at-the-money call priced at 20% volatility, with zero interest rates. Its gamma is about 0.0040 per rupee and its time decay about Rs 0.32 a day over 252 trading days. A day's move of 1.89%, typical of 30% volatility, is Rs 18.9, and half of 0.0040 times 18.9 squared is a gain of Rs 0.71, against Rs 0.32 of decay: about Rs 0.40 a day per option. The decay is exactly what a move of 1.26%, typical of 20% volatility, would pay back.
The relationshipGamma how fast the option's delta changes with the stock price Delta S the day's move in the stock price Theta the option's daily time decay sigma_r, sigma_i realised and implied volatility, 30% and 20% What it says in wordsEach day a hedged long option earns half its gamma times the squared move and pays its time decay; on average that is the gap between realised and implied variance.Why does direction not matter, but the path does?
The move enters squared, so plus Rs 18.9 and minus Rs 18.9 earn the same. The P&L depends on realised volatility against implied, not on where the stock ends up. But gamma is largest near the strike and shrinks as the stock drifts away, so 30% realised in a trend that carries the stock far from the strike early earns less than 30% realised while the stock chops around the strike. Hedging frequency matters too: hedge rarely and the P&L becomes noisy, even if its average barely changes.
Where candidates lose it
The common loss is saying the option made money because the stock went up. With a delta hedge in place, direction has been sold away; a candidate who talks about direction has not understood what the hedge does.
The second loss is saying a hedged position has no P&L. It has exactly one exposure left, realised against implied volatility, and naming it is the whole answer.
What the interviewer asks next
- Realised volatility comes in at 15% instead. What happens to the P&L, and why?
- Why is the same realised volatility worth more while the stock stays near the strike?
- You think implied volatility is too low but have no view on direction. What position expresses that?
095Daily returns have an autocorrelation of 0.1 at a one-day lag, decaying geometrically: 0.01 at two days, 0.001 at three, and so on. By how much is the true volatility over a 21-day month higher than daily volatility times the square root of 21?AQR Capital ManagementGreenwich · 2022
Try it first
How much higher is the true monthly volatility?
Show the worked solution
About 10% higher. The variance of a sum of 21 daily returns is 21 daily variances plus twice every covariance between pairs of days. With autocorrelation 0.1 at lag one, 0.01 at lag two and so on, the covariance terms add about 4.42 daily variances, lifting the total from 21 to about 25.4, a 21% rise. Volatility, the square root, rises by about 10%.
Why does a small autocorrelation matter over a month?
Think of a queue where each person who joins makes it slightly more likely that the next person joins too. Each nudge is small, but over a month of days they add up to longer queues than pure chance would give. The variance of a sum counts every pair of days, and there are 20 neighbouring pairs in a 21-day month, each counted twice, so a lag-one correlation of 0.1 alone adds 4 daily variances to the 21.
The square-root rule counts 21 daily variances for a month, but one-day lags add 4.0, two-day lags 0.38 and longer lags 0.04, for 25.42, so monthly volatility is 5.04 daily volatilities rather than 4.58, about 10% higher. How do you do the sum?
Write the monthly variance as the daily variance times 21, plus twice the sum over every lag. Lag one appears in 20 pairs with correlation 0.1, lag two in 19 pairs with 0.01, lag three in 18 with 0.001: twice (2.0 + 0.19 + 0.018 and so on) is about 4.42. The total is 25.42 daily variances instead of 21, a ratio of 1.210, and the volatility ratio is its square root, 1.100. Over a long horizon the ratio tends to (1 + 0.1)/(1 - 0.1) = 1.222, so a month is already close to the limit.
The relationshipsigma^2 the variance of one day's return n days in the month, 21 rho_k the autocorrelation at a lag of k days, 0.1 to the power k n - k how many pairs of days in the month are k days apart What it says in wordsA month's variance is the sum of the daily variances plus twice every covariance between pairs of days, weighted by how many such pairs the month holds.What does this mean for risk and performance numbers?
Scaling daily volatility by the square root of time silently assumes zero autocorrelation. Positive autocorrelation, which shows up in trend-following returns and in portfolios of thinly traded assets priced from stale quotes, makes the square-root rule understate longer-horizon risk. Negative autocorrelation, typical of mean-reverting strategies, does the reverse: at minus 0.1 the monthly variance is only 0.83 times the rule's. The same effect is why correlations between two assets measured on daily returns can differ from those measured on monthly returns: prices that react to the same news on different days look less related day by day than month by month.
Where candidates lose it
Candidates either dismiss 0.1 as too small to matter or report 21% as the answer. The first ignores that the correlation enters twenty times over; the second forgets that volatility is the square root of variance.
The other loss is refusing to estimate without doing every lag. The lag-one term alone gives 21 + 4.0, close to the full 25.42; say that the higher lags add about 0.42 and that you checked.
What the interviewer asks next
- What is the ratio if the lag-one autocorrelation is minus 0.1 instead?
- A fund holds thinly traded assets priced from stale quotes. Which way does its reported volatility err?
- Why can the correlation between two assets look higher on monthly returns than on daily returns?
Asked at AQR Capital Management, Quantitative Research, Greenwich, 2022 (Wall Street Oasis):
describe what covariance means what would be the difference between the correlation of daily vs monthly returns of a given year
096You roll a fair die repeatedly and keep a running total. What is the probability that the total is ever exactly 10? What does the answer approach for large targets?Quant and systematic fundsProp and quant trading firms
Try it first
Roughly what is the chance the running total ever hits exactly 10?
Show the worked solution
About 0.289, and it settles at 2/7, about 0.286, for large targets. Let p(n) be the chance the total ever equals n. To hit n, the total must first land on one of n - 1 to n - 6 and then roll exactly the gap, each with chance 1/6, so p(n) is the average of the six values before it, with p(0) = 1. Working up gives p(10) = 0.2893. Totals advance 3.5 a roll on average, so they land on 1 number in 3.5.
How do you set up the recursion?
Think about the last roll before the total reaches n. The total can land exactly on n only by first landing on one of n - 1 down to n - 6 and then rolling exactly the gap, each with chance 1/6, so p(n) = (1/6)[p(n - 1) + ... + p(n - 6)]. Start with p(0) = 1, because you begin at zero, and p of any negative number = 0. Then p(1) = 1/6, p(2) = 7/36, p(3) = 0.227, and so on up to p(10) = 0.2893.
The chance of ever landing on a total climbs from 1/6 at 1 to a peak of 0.360 at 6, then wobbles and settles onto 2/7, about 0.286; the target of 10 is hit with probability 0.289. Why does the answer settle at 2/7?
Picture stepping stones across a river, where each stride covers 1 to 6 stones with equal chance. Over a long walk you touch about one stone in every 3.5, because that is your average stride. The running total advances 3.5 per roll on average, so in the long run it lands on a fraction 1/3.5 = 2/7 of all numbers, and each far-off target is hit with probability close to 2/7. The early values wobble: p(6) is the highest, 0.360, because 6 is the last total a single roll from zero can reach directly, and the wobbles die out by about 20.
The relationshipp(n) the chance the running total ever equals n p(n - k) the chance of standing k below the target, one roll away E[roll] the average roll of a fair die, 3.5 What it says in wordsEach total's chance is the average of the six before it, and in the long run the totals land on one number in every 3.5.Why would a quant interviewer ask for a table or code here?
Because the recursion is dynamic programmingSolving a problem by building up answers to smaller versions of it and reusing them, instead of recomputing from scratch.: each value reuses the six before it, so a table of ten numbers is faster and safer than listing every sequence of rolls that sums to 10. Say the recursion, compute a few terms out loud, and give the limit with its reason; that is the complete answer. Do not try to enumerate paths: there are 492 ordered ways to reach 10 with rolls of 1 to 6, each with its own probability.
Where candidates lose it
The quick answers are 1/6, reasoning that some roll must land on 10 with one chance in six, and 2/7 stated as exact. The first ignores that most runs skip straight over 10; the second is close but is the long-run limit, and 10 is not yet far enough out for it to be exact.
The second loss is trying to count paths. Set up the recursion in one line instead and let it do the counting.
What the interviewer asks next
- What is the probability that the running total ever equals exactly 6?
- With a coin that adds 1 or 2 instead of a die, what does the hit probability approach?
- How would you write this as a dynamic programme in a few lines of code?
097A fund makes 50% in year one on Rs 100 crore, then takes in Rs 400 crore of new money and loses 20% in year two. What are its time-weighted and money-weighted returns, and did its investors make money?Fund of funds and allocatorsMulti-manager platforms
Try it first
The fund reports its two-year return. Which statement is right?
Show the worked solution
The time-weighted return is +20%; the money-weighted return is about -10.2% a year, and investors as a group lost Rs 60 crore. Rs 100 crore grows to Rs 150 crore, then Rs 400 crore arrives and the Rs 550 crore falls 20% to Rs 440 crore. Chaining 1.5 x 0.8 gives +20%, the manager's record. But Rs 500 crore went in and Rs 440 crore remains, because most of the money arrived just before the loss.
What does each measure answer?
Think of a restaurant that gets a glowing review after a good year, triples its tables, and then has a poor year. Its record over the two years looks fine, but most of the diners ate in the poor year. The time-weighted return measures the manager, chaining each period's return so the timing of money coming in or out does not count; the money-weighted return measures the investors, weighting each period by the money actually exposed to it.
The fund earns 50% on Rs 100 crore, takes in Rs 400 crore, then loses 20% of Rs 550 crore, so the manager's time-weighted record is +20% while investors put in Rs 500 crore, hold Rs 440 crore, and earned about -10.2% a year. How do you compute the money-weighted return?
Treat the investors' money as a set of cash flows and find the single annual rate that links them: Rs 100 crore in at the start, Rs 400 crore in after a year, Rs 440 crore out at the end. Solve 100(1 + r) squared + 400(1 + r) = 440; the root is 1 + r = 0.8983, so the money-weighted return is about -10.2% a year. It is negative because most of the money sat through the losing year: Rs 550 crore lost 20%, Rs 110 crore, while the good year earned only Rs 50 crore on Rs 100 crore.
The relationship100 Rs crore invested at the start 400 Rs crore added after one year 440 Rs crore the investors hold at the end r the money-weighted return, the internal rate of return on those flows What it says in wordsThe money-weighted return is the one rate that grows every rupee invested, from the day it arrived, into what the investors hold at the end.Which number should an allocator look at?
Both, for different questions. To judge the manager's skill, use the time-weighted +20%, because the manager did not choose when investors arrived. To judge whether the fund was good for the people in it, or whether money chased performance, use the money-weighted figure. A wide gap between the two, as here, is a warning about hot money: investors piling in after a strong year and bearing the next year's loss on a much bigger base. Had all Rs 500 crore arrived at the start, both measures would agree and the investors would hold Rs 600 crore.
Where candidates lose it
The common loss is reporting +20% and saying the investors made money. The chained return deliberately ignores that most rupees arrived just before the loss, and the interviewer built the numbers so the two measures point in opposite directions.
The second loss is averaging the two yearly returns, +50% and -20%, to get +15% a year. That is neither measure: the time-weighted annual figure is the square root of 1.2 minus 1, about 9.5% a year.
What the interviewer asks next
- What is the time-weighted return expressed as an annual rate?
- The Rs 400 crore had arrived at the start of year one instead. What are both returns now?
- Why do fund fact sheets report time-weighted returns?
098With orthonormal regressors, ordinary least squares gives coefficients of 0.6 and 0.15. What do ridge and lasso with a penalty of 0.2 give for each, and why does only lasso set a coefficient to zero?CitadelLondon · 2026
Try it first
Using half the residual sum of squares plus the penalty, what do the two methods give?
Show the worked solution
Ridge gives 0.5 and 0.125; lasso gives 0.4 and exactly 0. With orthonormal regressors each coefficient is shrunk on its own. Ridge divides each by 1 + 0.2 = 1.2, so it scales both down and never reaches zero. Lasso subtracts 0.2 from each size and stops at zero, so the small coefficient, 0.15, is removed. This uses the scaling of half the residual sum of squares plus 0.2 times the penalty.
Why do orthonormal regressors make this a one-line problem?
When the regressors are uncorrelated and scaled to unit length, the fit for each coefficient does not depend on the others, so the penalised problem splits into separate one-variable problems. Think of adjusting the volume on two speakers that are not wired together: turning one down does not change the other. For each coefficient you minimise half of (beta minus b) squared plus the penalty, where b is its least squares value, and the answer depends only on that one number. Ridge's penalty is half of beta squared times 0.2; lasso's is the size of beta times 0.2.
Ridge scales every least squares coefficient by 1/1.2, so 0.6 becomes 0.5 and 0.15 becomes 0.125, while lasso subtracts 0.2 and stops at zero, so 0.6 becomes 0.4 and 0.15, inside the dead zone below 0.2, becomes exactly 0. What does each penalty do to a coefficient?
Ridge's squared penalty pulls hard on big coefficients and gently on small ones. Setting the slope to zero gives beta x (1 + 0.2) = b, so ridge divides by 1.2: 0.6 becomes 0.5 and 0.15 becomes 0.125, shrunk but never zero. Lasso's absolute-value penalty pulls with the same force, 0.2, whatever the size. Its answer is the sign of b times the larger of (the size of b minus 0.2) and zero: 0.6 becomes 0.4, and 0.15, weaker than the pull of 0.2, lands exactly on 0.
The relationshipb the least squares coefficient, 0.6 or 0.15 lambda the penalty weight, 0.2 What it says in wordsRidge divides every coefficient by the same factor; lasso takes the same amount off every coefficient and never goes past zero.Why does only lasso select variables?
At zero, the squared penalty is flat: its slope is zero, so any small coefficient still earns its place by improving the fit a little. The absolute-value penalty has a corner at zero with a slope of 0.2 on each side, so a coefficient stays at zero unless the fit improves by more than 0.2 per unit, and 0.15 does not. That is why lasso gives sparseHaving many coefficients exactly equal to zero, so the model uses only a few of the available variables. models and ridge does not. The limitation to state: with correlated regressors lasso tends to keep one of a group arbitrarily and drop the rest, which is why desks often blend the two penalties in an elastic net.
Where candidates lose it
The common slip is to swap the two, saying ridge sets small coefficients to zero because it penalises harder. Ridge's penalty is heavy on big coefficients and almost nothing on small ones, which is exactly why it never zeroes them.
The second loss is getting the lasso numbers off by a factor of two. Written as the full residual sum of squares plus 0.2 times the absolute values, without the half, the threshold is 0.1, giving 0.5 and 0.05. State your scaling before you give numbers.
What the interviewer asks next
- At what penalty does lasso set the 0.6 coefficient to zero as well?
- The two regressors now have a correlation of 0.9. How do ridge and lasso behave differently?
- How would you choose the penalty in practice without fitting it to noise?
Asked at Citadel, Quantitative Research, London, 2026 (Wall Street Oasis):
very detailed and difficult questions about regularisation ridge and lasso
099You borrow in yen at 1% and invest in rupee deposits at 7% for a year. By how much can the rupee's value in yen fall over the year before the trade loses money?Long-short equity fundsGlobal macro funds
Try it first
How big a fall in the rupee can the trade absorb?
Show the worked solution
About 5.6%. Borrow 1 of yen value, convert it, and the rupee deposit grows to 1.07 in yen terms at today's rate. You owe 1.01. If the rupee's yen value falls by d, the deposit is worth 1.07 x (1 - d), which equals 1.01 when d = 1 - 1.01/1.07, about 5.6%. The cushion is slightly below the 6-point gap because the fall hits the interest as well as the principal.
Where does the carry trade's income come from?
Think of borrowing from a relative at 1% to put the money in a deposit paying 7%, except that the relative wants to be repaid in a different currency. The 6 points look like easy income until the currency moves. A carry trade earns the interest gap only if the currency you hold does not fall by more than about that gap over the holding period. The rate gap is the cushion; the exchange rate is the risk.
Per Rs 100 of yen borrowed, the trade makes Rs 6 if the rupee holds its value, breaks even at a 5.6% fall, already loses Rs 0.42 at a 6% fall, and loses Rs 4.70 at a 10% fall. Why is the cushion 5.6% and not 6%?
Work in yen per unit borrowed. At the end you hold 1.07 of rupee deposit, measured at today's exchange rate, and you owe 1.01 of yen; the rupee can lose value until 1.07 x (1 - d) = 1.01, so d = 5.6%. The fall applies to the whole 1.07, including the 0.07 of interest, which is why the cushion is a little smaller than 6. Borrow Rs 100 crore worth of yen and a 5.6% fall leaves you exactly even; a 10% fall loses about Rs 4.7 crore.
The relationship1.07 what each unit becomes in the rupee deposit at 7% 1.01 what you owe on each unit of yen borrowed at 1% d the fall in the rupee's value in yen over the year What it says in wordsThe trade breaks even when the grown rupee deposit, after the currency fall, buys back exactly the yen you owe.What does the forward market say about this?
Covered interest parityThe rule that the forward exchange rate must offset the interest rate gap between two currencies, or else borrowing in one, lending in the other and locking in the forward would be an arbitrage. sets the one-year forward so that the rupee buys about 5.6% fewer yen forward than today, exactly the fall that wipes out the carry. So an unhedged carry trade is a bet that the rupee will fall by less than the forward already implies. The limitation to state: carry trades tend to earn small amounts steadily and then lose sharply when many holders unwind at once, so the spread of outcomes has a fat left tail that the breakeven alone does not show.
Where candidates lose it
The quick answer is 6%, the gap between the two rates. It misses that the currency loss applies to the interest earned as well as to the principal, so the true cushion is slightly smaller, and at exactly a 6% fall the trade is already losing.
The second loss is treating the 6-point gap as a return and forgetting the currency altogether. Name the exchange rate as the risk before doing any arithmetic.
What the interviewer asks next
- The rupee weakens 10% against the yen over the year. What is the loss on Rs 100 crore of yen borrowed?
- What one-year forward rate does covered interest parity imply, and what does it say about the trade?
- Why do carry trades tend to lose money suddenly rather than gradually?
100You have n cars, each with enough fuel to drive 1,000 miles, and cars can transfer fuel to each other on the road. How far can one car get? What happens as n grows?Millennium ManagementLondon · 2024
Try it first
With 4 cars, how far can one car get?
Show the worked solution
1,000 x (1 + 1/2 + 1/3 + ... + 1/n) miles, which grows without limit but only like 1,000 x ln n. Drive all n cars together for 1,000/n miles: together they have burned one tankful, so one car tops up the others and stops. Then n - 1 cars drive 1,000/(n - 1), and so on, until the last car drives a full 1,000. With 4 cars that is about 2,083 miles; with 100, about 5,187.
Why can the fuel not simply be pooled into one car?
Each tank holds exactly 1,000 miles of fuel, so one car can never carry more than that at once. Think of porters carrying water across a desert: the helpers walk part of the way, hand over what they can spare, and drop back. The helper cars exist to keep the lead car's tank full for as long as possible, and they can only do that by travelling with it and burning fuel themselves.
Four cars drive 250 miles together before one refills the other three and stops, three drive 333 more, two drive 500 more and the last drives a full 1,000, reaching 2,083 miles; ten cars reach 2,929 and a hundred reach 5,187. How long is each leg?
With k cars travelling together on full tanks, drive until the group has burned exactly one tankful, which takes 1,000/k miles. Each tank is then 1/k empty, so the k - 1 cars that continue have (k - 1)/k of a tank of space between them, and the car that stops has exactly (k - 1)/k of a tank left to fill it. The legs are 1,000/n, then 1,000/(n - 1), and so on to 1,000 for the last car alone. With four cars: 250 + 333.3 + 500 + 1,000 = 2,083.3 miles. Dropping each helper the moment its fuel can refill the rest keeps as few cars as possible burning fuel at every mile.
The relationshipH_n the harmonic number, 1 + 1/2 + ... + 1/n 1000/k the leg driven while k cars are still moving 0.577 Euler's constant, the gap between H_n and ln n for large n What it says in wordsThe distance is 1,000 miles times the sum of one over each number of cars still driving, which grows like the natural logarithm of the number of cars.What happens as n grows?
The harmonic series never stops growing, so with enough cars there is no ceiling on the distance. But it grows only like the logarithm of n: 10 cars reach about 2,929 miles, 100 cars about 5,187, and each further tenfold increase in cars adds only about 2,303 miles. Reaching 5,000 miles takes 83 cars. That is the pattern worth naming in the room: unlimited in principle, very expensive in practice, the same diminishing return you meet whenever each extra unit of effort adds less than the one before.
Where candidates lose it
The quick wrong answer is n x 1,000 miles, pooling all the fuel, which ignores that no tank holds more than 1,000 and that helpers burn fuel just keeping up. The opposite slip is 1,000 miles, forgetting that fuel can be passed forward at all.
The second loss is reaching the harmonic series and then saying the distance levels off, or that it grows in proportion to n. Name the growth rate: like ln n, unbounded but slow.
What the interviewer asks next
- With 3 cars, exactly how far can one car get?
- Cars may now turn back and refuel at the start. Can the lead car get further?
- Roughly how many cars do you need for one car to travel 5,000 miles?
Asked at Millennium Management, Investments, London, 2024 (Wall Street Oasis):
Suppose you have n cars, each fueled so that they can drive for 1000 miles.
