Hedge Funds puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 38
- Topics
- 14
- Hard
- 30
003X and Y are independent and each uniform on 0 to 1. What is the probability that X + Y is less than 1.5, and what shape is the density of X + Y?CitadelChicago · 2025
Try it first
Pick before you draw anything.
Show the worked solution
The probability is 7/8, and the density of X + Y is a triangle, a tent peaking at 1. Because X and Y are independent and uniform, every point of the unit square is equally likely, so probability is area. The line x + y = 1.5 slices off a corner triangle with legs of 0.5, area 1/8. The sum's density rises in a straight line from 0 to 1 and falls back to 0 at 2.
Why does probability become area here?
Throw a dart at a square board so that every point is equally likely to be hit. The chance it lands in a region is that region's share of the board. Two independent uniforms are exactly that dart: the pair (X, Y) lands evenly on the unit square, so any question about X and Y becomes a question about an area. The condition X + Y below 1.5 is everything under the line x + y = 1.5, which is the whole square except one corner.
The line x + y = 1.5 removes a corner triangle of area 1/8 from the unit square, so X + Y is below 1.5 with probability 7/8, and the density of X + Y is a tent on 0 to 2 whose tail beyond 1.5 also has area 1/8. How do you get the shape of the sum's density?
Slide the line x + y = s across the square and watch how long it is inside. Near s = 0 it barely clips the corner; at s = 1 it runs corner to corner, the longest it gets; past 1 it shortens again. The density of the sum at s is proportional to the length of that line inside the square, which gives a triangle rising from 0 to a peak at 1 and falling to 2. This is the convolutionThe density of a sum of independent variables, found by adding up every way the two parts can combine to the same total. of two flat densities, and the same reason two dice most often total 7.
The relationshipf_{X+Y}(s) the density of the sum at the value s f_Y(s - x) equal to 1 when s - x lies between 0 and 1, otherwise 0 What it says in wordsAdd up every split of s into an x and a y that both lie in 0 to 1; the count of splits rises to s = 1 and then falls.Check the first answer with the tent. The area beyond 1.5 is a triangle with base 0.5 and height 0.5, which is 1/8 again. Two routes that agree is the check an interviewer wants to hear before you commit. Add a third uniform and the density becomes three joined curved pieces; add many and the sum looks normal, which is the central limit theorem arriving in slow motion.
Where candidates lose it
Candidates reach for a double integral before drawing, set the limits wrongly, and spend two minutes on what is a one-line area argument. Draw the square first; the corner triangle is visible at a glance.
The second loss is saying the sum of two uniforms is uniform on 0 to 2. It is not: there is only one way to get a sum near 0 and many ways to get a sum near 1, which is why the density is a tent and not a flat line.
What the interviewer asks next
- What is the probability that X + Y is less than 0.5?
- What is the probability that the larger of X and Y is below 0.5, and how does the picture change?
- What does the density of X + Y + Z look like?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
041Two orders arrive one after the other. The first arrives after a wait that is exponential with a mean of one minute; the second arrives after a further, independent exponential wait with the same mean. What is the probability that both have arrived within one minute?CitadelChicago · 2025
Try it first
Your estimate:
Show the worked solution
1 minus 2/e, about 26.4%. The total wait is the sum of two independent exponential waits. Convolving the two densities gives t e^-t, a gamma shape that starts at zero because two steps cannot both be instant. Its area from 0 to 1 is 1 minus e^-1 (1 + 1), which is 1 minus 2/e. A second route: it is the chance that a Poisson process with rate 1 produces at least two arrivals in one minute.
Why is this not the chance of one wait, squared?
Squaring would be right if both orders were racing from the same start line. Here they queue: the second clock only starts when the first order lands. It is like two buses where you must take the first to reach the stop for the second. The event is that the sum of the two waits is under one minute, which is stricter than each wait being under one minute. The chance one wait is under a minute is 1 minus 1/e, about 63%; squaring gives 40.0%, which answers a different question.
How do you get the density of the sum?
Add up every way to split the total t between the two waits. The density of a sum of independent waits is the convolutionThe density of a sum of two independent variables, found by integrating one density against the other shifted over every possible split of the total. of their densities, and for two exponentials it is t e^-t. Every split of t into s and t minus s has density e^-s times e^-(t minus s), which is e^-t whatever s is, and there is a length t of possible splits. Integrate t e^-t from 0 to 1 by parts and you get 1 minus 2/e.
The total of two independent one-minute exponential waits has density t e^-t, which starts at zero and peaks at one minute, so only 26.4% of its area, 1 minus 2/e, lies below one minute. The relationshipX1, X2 the two independent exponential waits, each with mean one minute t e^-t the density of their sum, from convolving the two exponential densities What it says in wordsThe chance the total wait is under a minute is the area under the gamma density up to one minute.Check it with counting. Exponential waits are the gaps of a Poisson process, so both orders arrive within a minute exactly when the process makes at least two arrivals in that minute. With one arrival expected per minute, the chance of zero is e^-1 and of exactly one is e^-1, so at least two is 1 minus 2/e, the same number. Say both routes and the interviewer will usually skip ahead.
Where candidates lose it
The common loss is squaring the single-wait probability, which answers the question of two independent orders racing in parallel. Read the setup again: one after the other means the waits add.
The second loss is freezing on the convolution integral. If the integral will not come, switch to the Poisson count: at least two arrivals in one minute. Candidates who know one route and not the other are the ones interviewers push hardest.
What the interviewer asks next
- What is the probability that three orders in sequence all arrive within two minutes?
- Given that both orders arrived within one minute, what is the expected arrival time of the first?
- The two waits have means of one and two minutes. What is the density of their sum?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
if i knew this was about convolutions, i would have answered better.
067Daily returns are normal with 1% volatility on 80% of days and normal with 4% volatility on the other 20%, both with zero mean. What are the overall daily volatility and the kurtosis of this mixture?Two SigmaNew York · 2025
Try it first
What is the overall daily volatility of the mixture?
Show the worked solution
Overall volatility is 2% a day and the kurtosis is 9.75, against 3 for a normal distribution. Variances mix in proportion: 0.8 x 1 + 0.2 x 16 = 4, so volatility is 2%. Fourth moments mix the same way, and each normal contributes 3 times its volatility to the fourth: 0.8 x 3 + 0.2 x 768 = 156. Dividing by the variance squared, 16, gives 9.75.
Why is a mixture of normals not normal?
Think of a commute that takes 30 minutes on most days and two hours on strike days. The average trip hides the shape: most days cluster tightly and a few days sit far out. Mixing a calm regime with a wild one gives more small moves than a normal with the same overall spread, fewer medium ones, and far more big ones. The peak is taller, the shoulders thinner and the tails fatter, which is what {term('kurtosis', 'The fourth moment of a distribution divided by the variance squared; 3 for a normal distribution, higher when tails are fatter.')} measures.
How do you get the two numbers?
Work with moments, because they mix in proportion to the weights. The variance is 0.8 x 1 + 0.2 x 16 = 4, so volatility is 2%; the fourth moment is 0.8 x 3 x 1 + 0.2 x 3 x 256 = 2.4 + 153.6 = 156, and kurtosis is 156 / 4 squared = 9.75. Look at where the 156 comes from: 153.6 of it is the wild days, which occur only one day in five. The fourth power makes rare large moves dominate.
The relationshipw_i the share of days in each regime, 0.8 and 0.2 sigma_i the volatility in each regime, 1% and 4% 3 sigma_i^4 the fourth moment of a zero-mean normal with volatility sigma_i What it says in wordsAverage the second and fourth moments across regimes, then divide the fourth moment by the variance squared.Against a normal with the same 2% volatility, the mixture has a taller peak, thinner shoulders and fatter tails: a daily move beyond 6% either way happens on 2.7% of days under the mixture but only 0.27% under the normal, about 10 times as often. What does this mean for a risk model?
A model that fits a normal to the 2% volatility is right about the average day and wrong about the days that matter. It says a move beyond 6% happens on about 0.27% of days, roughly once in 370 trading days; the mixture says 2.7%, roughly once in 37. That is the usual story of market returns: calm stretches and volatile stretches, each close to normal, adding up to fat tails. A model that lets volatility change over time captures much of it.
Where candidates lose it
The first slip is averaging the volatilities, 0.8 x 1% + 0.2 x 4% = 1.6%. Variances average in a mixture, not standard deviations, so the answer is 2%.
The second is guessing that a mixture of normals has kurtosis 3 because each piece does. Mixing different variances always pushes kurtosis above 3, and here the fourth-power weight on the wild days takes it to 9.75.
What the interviewer asks next
- What mixture weight on the 4% regime maximises the kurtosis?
- What is the probability density of the mixture at zero, compared with the normal?
- If the two regimes had different means but the same volatility, what would happen to skew and kurtosis?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis):
They asked a couple questions involving Mixture Gaussians (e.g., probability density and moments).
