Quant puzzles, solved step by step
- Puzzles
- 100
- Traced to a firm
- 71
- Topics
- 12
- Hard
- 30
004You keep drawing independent random numbers, each uniform on 0 to 1, until their running total exceeds 1. What is the expected number of draws?Citadel SecuritiesChicago · 2025
Try it first
What is your instinct for the answer?
Show the worked solution
e, about 2.718. The chance that n uniforms still add to at most 1 is 1/n!, the volume of a corner of the n-dimensional cube. The number of draws N exceeds n exactly when that happens, and an expected count is the sum of the chances of exceeding each n. So E[N] is 1 + 1 + 1/2 + 1/6 + 1/24 and so on, which is e.
Why is two draws the wrong answer?
Fill a one litre jug with cups of random size, each somewhere between empty and full. On average two cups make a litre, but you stop at the first cup that overflows, and some pairs of cups fall short. Averages of the draws do not tell you the average stopping time: you need the chance that you are still short after each draw. After two draws you are still at or below 1 exactly half the time, so a third draw is needed often, and occasionally a fourth.
The chance of still being at or below 1 after n draws is 1/n!, so the bars run 1, 1, 1/2, 1/6, 1/24, and their running sum, the expected number of draws, closes in on e, about 2.718. Where does 1/n! come from?
For two draws, the pairs with a total at or below 1 fill the triangle under the line x + y = 1 in the unit square, area 1/2. For three, they fill a corner of the unit cube, volume 1/6. In general the region where n uniforms add to at most 1 is a corner of the n-dimensional cube with volume 1/n!, because the n! orderings of the coordinates carve the cube into equal pieces. You can also build it by convolutionThe density of a sum of independent variables, found by combining every way the parts can add up to the same total.: the density of the sum below 1 is s to the power n-1 over (n-1)!, and integrating from 0 to 1 gives 1/n!.
The relationshipN the number of draws needed P(N > n) the chance that n draws were not enough U_i the uniform draws What it says in wordsThe expected count equals the sum over n of the chance that n draws were still not enough, and those chances are 1/n!.How do you check an answer this surprising?
Check the pieces. N is at least 2 always, since one draw never exceeds 1, so the answer must be above 2; the bars for n = 0 and n = 1 are both 1 for that reason. A simulation of 200,000 runs gives an average of 2.721 draws, within a whisker of 2.718. Saying that you would simulate it, and roughly what you expect to see, is a good close in a research interview.
Where candidates lose it
The instinctive answer is 2, because two draws average exactly 1. It confuses the average of the draws with the average stopping time, and it ignores that the stopping rule waits for the total to pass 1, not reach it on average.
The second loss is knowing the answer is e without being able to say why. The tail-sum formula for an expected count, plus the 1/n! volume, is the whole argument, and it takes three sentences.
What the interviewer asks next
- What is the expected number of draws to exceed 2?
- What is the expected value of the total at the moment it first exceeds 1?
- What is the probability that exactly two draws are needed?
Asked at Citadel Securities, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
016You need to sample a point uniformly at random from a triangle with vertices A, B and C, using two independent uniform numbers u and v on 0 to 1. How do you do it, and why does the formula A + u(B - A) + v(C - A) fail on its own?Two SigmaNew York · 2023
Try it first
What goes wrong with A + u(B - A) + v(C - A) for u, v uniform on 0 to 1?
Show the worked solution
Draw u and v; if u + v is above 1, replace them with 1 - u and 1 - v; then return A + u(B - A) + v(C - A). The plain formula maps the unit square onto a parallelogram twice the size of the triangle, so half the draws, 50.0% in a simulation, land outside. Reflecting through the square's centre folds that half exactly onto the other, keeping the density flat and wasting no draws.
Why does the plain formula give a parallelogram?
Think of a tiled floor where each tile is a parallelogram and you want to pick a spot on one triangular half of a tile. Pick any spot on the tile and half the time you are on the wrong half. A + u(B - A) + v(C - A) with u and v each free on 0 to 1 walks up to one full step along AB and one full step along AC, which covers the parallelogram with corners A, B, C and D = B + C - A, not the triangle. The triangle is exactly the part where u + v is at most 1.
Two uniforms fill a unit square that the linear map turns into a parallelogram twice the size of triangle ABC, so draws with u + v above 1 land outside; reflecting such a draw from (0.8, 0.6) to (0.2, 0.4) brings it back inside at a uniformly distributed spot. Why does reflecting keep the distribution uniform?
Two facts. The map (u, v) to (1 - u, 1 - v) is a half turn about the square's centre, so it carries the upper triangle onto the lower one without stretching any area; and an affine mapA linear map followed by a shift, such as A + u(B - A) + v(C - A); it scales every area by the same factor. scales every area by the same factor, so a flat density stays flat. Put together, each small patch of the triangle receives draws from exactly two equal patches of the square. A simulation that splits the triangle into four equal pieces finds 24.9%, 25.1%, 25.0%, 25.0% of the points in them, each a quarter.
The relationshipu, v independent uniforms on 0 to 1 (1-u, 1-v) the reflection of a draw through the square's centre P the sampled point, uniform on triangle ABC What it says in wordsFold the unwanted half of the square onto the wanted half, then map it linearly onto the triangle.What other methods would an interviewer accept, and which fail?
Rejection works: throw away draws with u + v above 1. It is correct but wastes half the random numbers. A popular wrong method draws three uniforms and divides each by their sum to get weights on A, B and C; the weights add to 1, but the points pile up near the centre, so the result is not uniform. A correct closed form uses a square root: with r1 and r2 uniform, take (1 - root r1)A + root r1 (1 - r2)B + root r1 r2 C. Name one fast method, prove it, then name the tempting wrong one.
Where candidates lose it
The trap is writing the linear formula and stopping, because it looks like a weighted average of the vertices. Half the points leave the triangle, and the candidate who does not draw the square never sees it.
The second loss is fixing the problem in a way that breaks uniformity, such as normalising random weights to sum to 1 or clamping u + v to 1. Both keep points inside but crowd them into part of the triangle. Say why your fix preserves area.
What the interviewer asks next
- Prove the square-root method gives a uniform point.
- How would you sample uniformly from a convex polygon with n vertices?
- How would you sample uniformly from the surface of a sphere?
Asked at Two Sigma, Research, New York, 2023 (Wall Street Oasis):
Biased gamblers ruin problems; Markov Chain problems; sampling uniformly from triangle
032n points are placed independently and uniformly on a circle of circumference 1, with n at least 3. Each point colours the arc between itself and its nearest neighbour. What is the expected total length that gets coloured?Susquehanna International GroupLondon · 2026
Try it first
Which is closest to the expected coloured length?
Show the worked solution
7/18, about 0.389, for every n from 3 upwards. A gap is left uncoloured only when it is longer than both gaps beside it, because then neither endpoint has it as its nearest. For three points that gap is simply the longest of three pieces, which averages 11/18, so 7/18 is coloured. For larger n the same 11/18 comes out, so the answer does not depend on n.
When is a gap left uncoloured?
Picture people standing round a circular table, each turning to talk to whoever is closer, left or right. A stretch of table between two people stays silent only if both of them turned away, which means each had a closer person on their other side. A gap is uncoloured exactly when it is longer than both of its neighbouring gaps. A gap coloured from both ends is still coloured once, so the question becomes: what is the expected total length of gaps that are local maxima?
Each gap is coloured if it is the shorter gap for at least one endpoint and left uncoloured if it is longer than both neighbours; this sample of ten points colours 0.57 of the circle, and the average over all placements is 7/18, about 0.389, for any n of 3 or more. How do you get 11/18 for the uncoloured part?
Start with n = 3, the case you can finish in the room. With three gaps, every gap's two neighbours are the other two gaps, so the only uncoloured gap is the longest one. Three random points cut the circle like a stick broken into three, and the longest of three pieces averages (1/3)(1 + 1/2 + 1/3) = 11/18. So the coloured length is 7/18.
For larger n, use the fact that the n gaps behave like n independent exponentialA random length whose chance of ending is the same at every instant; waiting times between random arrivals follow it. lengths rescaled to add up to 1, and that the rescaling is independent of the shape. For three unit exponentials X, Y and Z, the expected value of X counted only when X is the largest is 1 - 2/4 + 1/9 = 11/18. Each of the n gaps contributes that, divided by the expected total of n, and the n gaps sum to 11/18 again. The uncoloured share is 11/18 whatever n is, so the coloured share is always 7/18.
The relationshipx e^(-x) a gap's length times its density, in the exponential picture (1 - e^(-x))^2 the chance both neighbouring gaps are shorter 1/n rescaling so the n gaps add to a circle of length 1 What it says in wordsThe expected length of gaps longer than both neighbours is 11/18, and the rest of the circle is coloured.Say the check: a seeded simulation of 40,000 random circles gives 0.389 for n = 3, 0.388 for n = 5 and 0.389 for n = 10. The limitation is that the exponential step is a known result you should name, not derive, in an interview; the n = 3 case is the part you prove on the spot.
Where candidates lose it
The usual loss is counting gaps instead of measuring them. One gap in three is a local maximum, so candidates answer 2/3 coloured. The uncoloured gaps are selected for being long, which is why their share of length, 11/18, is far above one third.
The second is double counting a gap that both endpoints colour. It is coloured once. Frame the problem around uncoloured gaps and both mistakes disappear.
What the interviewer asks next
- What is the expected number of uncoloured gaps?
- What if each point colours the arc to its farther neighbour instead?
- Does the answer change for points on a line segment rather than a circle?
Asked at Susquehanna International Group, Quantitative Research, London, 2026 (Wall Street Oasis):
if n points are placed on a circle and each point colours in the arc to its nearest neighbour
044A stick of length 1 is broken at three independent uniform points into four pieces. What is the expected length of the longest piece?Hudson River TradingNew York · 2024
Try it first
What is the expected length of the longest piece?
Show the worked solution
25/48, about 0.521. For a stick broken into n pieces, the expected k-th smallest piece is (1/n)(1/n + 1/(n - 1) + ... ) with k terms. For n = 4 the sorted pieces average 3/48, 7/48, 13/48 and 25/48, which add to 1. The longest is (1/4)(1 + 1/2 + 1/3 + 1/4) = 25/48, more than twice the average piece of 1/4.
Why is the longest piece so much longer than a quarter?
Cut a sheet of dough at three random spots and the pieces are rarely even: one is usually a big slab and one a sliver. Random breaks produce uneven pieces, and the longest piece collects the unevenness, so its average sits far above the average piece. The average piece is always 1/4; the question asks about the largest of four correlated lengths, which is an {term('order statistic', 'The k-th smallest value in a sample, for example the minimum, the median or the maximum.')}.
Sorted by length, the four pieces of a randomly broken stick average 3/48, 7/48, 13/48 and 25/48 of its length, so the longest piece averages about 0.52, more than twice the naive quarter, and a 100,000-stick simulation agrees to three decimals. Where does the harmonic formula come from?
Start with the shortest piece. The chance that all four pieces exceed x is (1 - 4x) cubed: take x off every piece and the three breaks must fit into the remaining length 1 - 4x. Integrating that from 0 to 1/4 gives an expected shortest piece of 1/16. Then the key fact: the step from each sorted piece to the next adds on average (1/n) times 1 over the number of pieces still longer. After the shortest, three pieces remain longer, so the next piece averages 1/16 + (1/4)(1/3); then add (1/4)(1/2), then (1/4)(1). The longest piece is (1/4)(1/4 + 1/3 + 1/2 + 1) = 25/48.
The relationshipL_(4) the longest of the four pieces 1/4 one over the number of pieces 1 + 1/2 + 1/3 + 1/4 the harmonic sum up to the number of pieces What it says in wordsThe longest piece averages one quarter of the fourth harmonic number.The step rule comes from the fact that the pieces behave like independent exponential lengths scaled to total 1, and the gap between successive minima of exponentials is memoryless. You can name that in the room rather than prove it. The check that the formula is right: the four sorted averages add to exactly 1, and a simulation of 100,000 sticks gives 0.521 for the longest. For n pieces in general, the longest averages (1/n) times the n-th harmonic number, which grows like (ln n)/n.
Where candidates lose it
The common loss is answering 1/4, the average piece. The question asks for the average of the largest piece, and the largest of four uneven pieces is usually more than half the stick.
The second is trying to integrate the maximum directly over the three break points, which gets messy fast. Start from the minimum, use the step rule, and check that the four averages add to 1.
What the interviewer asks next
- What is the expected length of the shortest piece for n pieces?
- What is the probability the four pieces can form a quadrilateral?
- Break the stick at two points instead. What is the expected longest piece?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
I was asked an expected value question involving order statistics.
051You and a friend agree to meet at a spot some time between 5 and 6 pm. Each of you arrives at an independent, uniformly random time in that hour and waits 20 minutes for the other before leaving (or until 6 pm, whichever comes first). What is the probability you meet?Jane StreetNew York · 2026
Try it first
Before you draw anything: what is the chance you meet?
Show the worked solution
5/9, about 55.6%. Put your arrival time on one axis and your friend's on the other, so every outcome is a point in a unit square. You meet when the two times differ by at most a third of an hour, a band along the diagonal. The two corner triangles outside the band each have legs of 2/3, area 2/9, so the band is 1 - 4/9 = 5/9.
Why turn two arrival times into a square?
Think of two people trying to catch each other at a tea stall with no phones. Nothing about the answer depends on who is you and who is the friend; it depends only on the pair of times. With two independent uniform times, every pair is equally likely, so the pair is a point spread evenly over a square and any probability is simply an area. The event you meet becomes a region: the set of points where the two times are within 20 minutes of each other. That region is a diagonal band, because the line x = y is where you arrive together.
Plotting your arrival against your friend's, the meeting region is the diagonal band where the times differ by 20 minutes or less; the two white corner triangles, each 2/9 of the square, are the misses, so you meet with probability 5/9. How do you get the area without any integration?
Count the region you do not want. The two corners where one person arrives more than 20 minutes after the other are right triangles with both legs 40 minutes long, which is 2/3 of the side. Each has area (2/3) x (2/3) / 2 = 2/9, together 4/9, so the band is 5/9. Complements are the fastest route here, as they are for most geometric probability questions, because the leftover pieces are usually triangles.
The relationshipX, Y the two arrival times as fractions of the hour, independent and uniform on 0 to 1 w the waiting time as a fraction of the hour, here 20 of 60 minutes What it says in wordsThe chance of meeting is one minus the two corner triangles, whose legs are each one minus the waiting time.What does the general formula tell you that the number does not?
Read 2w - w squared term by term. The 2w is the naive answer of either person waiting, and the minus w squared removes the double count and the clipping at the edges of the hour. It also tells you how waiting time buys certainty: to meet half the time each person must wait about 17.6 minutes, and to be sure each must wait the whole hour. The same picture prices any tolerance between two random arrivals, such as two orders landing in the same matching window of an auction.
Where candidates lose it
The common answer is 1/3, from reading 20 minutes as a third of the hour. It ignores that either person can be the one who waits, and it has no way to handle the edges of the hour, where a person arriving at 5:55 can only wait five minutes.
The second trap is trying to integrate over one person's arrival time case by case near the edges. It works but wastes three minutes. Draw the square first and subtract the two triangles out loud.
What the interviewer asks next
- How long would each person need to wait for a 50% chance of meeting?
- You wait 10 minutes and your friend waits 30. What is the chance now?
- Three people arrive at random in the hour and each waits 20 minutes. What is the chance all three are together at some moment?
Asked at Jane Street, Technology, New York, 2026 (Wall Street Oasis):
1v1 math problems. bus stop. two people meeting probelm
063X and Y are independent random variables, each uniform on 0 to 1. What is the density of X + Y, and what is the probability that X + Y is less than 1.5?CitadelChicago · 2025
Try it first
What is P(X + Y < 1.5)?
Show the worked solution
The density is a triangle, f(s) = s for s up to 1 and 2 - s from 1 to 2, and P(X + Y < 1.5) = 7/8. Convolving two flat densities gives a tent peaking at 1. The part above 1.5 is a triangle with base 0.5 and height 0.5, area 1/8. In the unit square it is the same corner: the line x + y = 1.5 cuts off a triangle with legs of 0.5.
Why is the sum not uniform on 0 to 2?
Roll two dice: a total of 7 can be made six ways, a total of 12 only one way. Continuous uniforms behave the same. A sum near the middle can be made from many pairs, a sum near either end from very few, so the density of the sum rises to a peak and falls again. The mechanism that builds it is {term('convolution', 'The density of a sum of independent variables: for each possible total, add up the density of every pair of values that makes it.')}: the density at s is the length of the set of x values for which both x and s - x lie between 0 and 1.
In the unit square the line x + y = 1.5 cuts off a corner triangle of area 1/8, and in the triangular density of the sum the tail above 1.5 is the same 1/8, so X + Y is below 1.5 with probability 7/8. How does the convolution give the triangle?
Fix a total s. You need x between 0 and 1 and also s - x between 0 and 1, so x must lie between max(0, s - 1) and min(1, s). For s below 1 that interval has length s, and for s above 1 it has length 2 - s, so the density is a tent with its peak of 1 at s = 1. Check that the area is 1: a triangle with base 2 and height 1. The mean is 1 and the variance is 1/12 + 1/12 = 1/6, both of which you can read from symmetry and independence.
The relationshipf(s) the density of the sum at the value s the indicator 1 when s - x is a valid value of Y, otherwise 0 What it says in wordsThe density of the sum at s is how many ways of splitting s are allowed, which rises linearly to 1 and falls back.Why give the square picture as well?
It is a check that costs ten seconds. Because the pair is uniform on the unit square, any probability about X + Y is an area, and the event X + Y at least 1.5 is the corner triangle above the line x + y = 1.5. Its legs run from 0.5 to 1 on each axis, so its area is 1/8 and the answer is 7/8 again. The density gives you the whole distribution; the square gives you any single probability fast. Keep both, because the next question is usually three uniforms, where the density becomes piecewise quadratic and the square becomes a cube: P(X + Y + Z < 1) = 1/6.
Where candidates lose it
The common slip is 3/4, from assuming a sum of uniforms is uniform. Sums are never uniform unless one of the pieces is degenerate; they pile up in the middle.
The second loss is getting the convolution limits wrong and producing a density that does not integrate to 1. Write the two constraints on x out loud, and check the triangle's area before you use it.
What the interviewer asks next
- What is the density of X - Y?
- What is P(X + Y + Z < 1) for three independent uniforms?
- What is the expected value of max(X, Y), and of X + Y given that X + Y > 1?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
He was asking some questions about the probability, especially on the convolution.
074A point is dropped uniformly at random in a unit square. What is the expected distance from the point to the nearest edge of the square?Hudson River TradingNew York · 2024
Try it first
What is the expected distance to the nearest edge?
Show the worked solution
1/6. Let D be the distance to the nearest edge. D exceeds d exactly when the point lies in the inner square of side 1 - 2d, so P(D > d) = (1 - 2d)^2 for d up to 1/2. The expected value of a non-negative variable is the integral of its tail, and the integral of (1 - 2d)^2 from 0 to 1/2 is 1/6.
Why work with the chance of being far rather than the distance itself?
Think of a sandpit where a child stands at a random spot and the question is how far they are from the nearest edge. Writing the distance as min(x, 1 - x, y, 1 - y) and integrating a minimum of four things means splitting the square into four triangles. Asking instead when the point is farther than d from every edge has a one-picture answer: the point must lie in a smaller square, shrunk by d on every side. That square has side 1 - 2d, so its area, (1 - 2d)^2, is the tail probability. One formula replaces four cases.
A point is more than d from every edge only inside the inner square of side 1 - 2d, so the chance of being farther than 0.1, 0.2, 0.3 and 0.4 is 0.64, 0.36, 0.16 and 0.04, and the area under that tail curve is the expected distance, 1/6. How does the tail give the expectation?
For any non-negative random variable, the expected value equals the integral of the chance that it exceeds each level, E[D] = integral of P(D > d). Here that is the integral of (1 - 2d)^2 from 0 to 1/2. Substitute u = 1 - 2d and it becomes half the integral of u^2 from 0 to 1, which is 1/6. A simulation with 200,000 random points gives 0.1669, against the exact 0.1667.
The relationshipD the distance from the random point to the nearest edge P(D > d) the area of the inner square of side 1 - 2d What it says in wordsAdd up the chance of being farther than each distance, and the total is the expected distance.How do you sanity-check 1/6 against simpler cases?
Build up the number of edges. The distance to one fixed edge averages 1/2, to the nearer of two opposite edges averages 1/4, and to the nearest of all four it falls to 1/6, so each added constraint pulls the minimum closer. The density of D is the slope of the tail, 4(1 - 2d), largest at the edge, which says most random points are near the boundary. That is the same reason most of the volume of a high-dimensional cube sits near its surface, a fact that matters when sampling scenarios in many risk factors at once.
Where candidates lose it
The trap answers are 1/2 and 1/4, from handling one edge or one axis and forgetting that the nearest of four edges is a minimum. A candidate who integrates min(x, 1 - x, y, 1 - y) directly often splits the square wrongly and lands on a different number.
Draw the inner square and say tail integral; the whole calculation is then one line.
What the interviewer asks next
- What is the expected distance to the nearest edge in a unit cube?
- What is the expected distance to the nearest corner of the square?
- What is the density of the distance to the nearest edge, and where is it highest?
Asked at Hudson River Trading, Campus Algo Dev Interview, New York, 2024 (Wall Street Oasis):
I was asked a expected value question involving the expected value among distance to an edge, with a randomly placed object.
082Two independent waiting times are each exponentially distributed with a mean of one minute. What is the probability that their total is less than one minute?CitadelChicago · 2025
Try it first
Pick the closest value.
Show the worked solution
1 - 2/e, about 26.4%. Convolving two exponential densities gives the total the density x e^-x, which starts at zero and peaks at one minute. Its area below one minute is 1 - 2/e. The same number drops out of the Poisson view: the total is under a minute exactly when at least two arrivals land in the first minute of a rate-one Poisson process.
Why does adding two waits change the shape so much?
Suppose you need two buses, one after the other, and each arrives on average a minute after you reach its stop. Catching the first bus quickly is common; catching both quickly is rare, because both have to cooperate. A single exponential wait is most likely near zero, but a sum of two is almost never near zero, so its density starts at zero and rises into a hump. That shift of mass away from zero is why the answer is much smaller than the 63.2% chance that one wait is under a minute.
The single exponential wait puts most of its mass near zero, but the total of two waits has density x e^-x, which starts at zero and peaks at one minute, so only 26.4% of its area, shaded, falls below one minute. The relationshipS the total of the two waits x the first wait, which can be anything from 0 to s e^{-x} the exponential density with mean 1 What it says in wordsTo land on a total of s, the first wait takes any value x and the second makes up the rest; adding over all x gives s e^-s.Is there a way to get 1 - 2/e without integrating?
Yes, and it is the cleaner answer to give aloud. Exponential waits with mean one are the gaps between arrivals of a Poisson process with rate one per minute. The second arrival comes before one minute exactly when at least two arrivals land in the first minute, and the Poisson chance of zero or one arrival is e^-1 + e^-1 = 2/e. So the answer is 1 - 2/e, about 26.4%, with no calculus at all.
Sanity-check the size. Both waits being under a minute has probability (1 - 1/e)^2, about 40.0%, and the total being under a minute is a stricter event, so the answer must be smaller: 26.4% is. The limitation is the independence assumption; if the two waits were driven by the same traffic, they would move together and the total would be more spread out.
Where candidates lose it
The frequent wrong answer is (1 - 1/e)^2, about 40%, which is the chance that each wait is under a minute. The question asks about the total, and two waits of 0.7 minutes each pass that test while failing this one.
The second loss is starting a convolution integral and getting lost in the limits. Say the Poisson route first: at least two arrivals in the first minute, one minus the chance of zero or one.
What the interviewer asks next
- What is the probability that the sum of three such waits is under one minute?
- Given the total is exactly 2 minutes, what is the distribution of the first wait?
- What is the probability that the first wait is shorter than the second?
Asked at Citadel, Quant Research Interview, Chicago, 2025 (Wall Street Oasis):
if i knew this was about convolutions, i would have answered better
094Two points are chosen independently and uniformly on the surface of a unit sphere. What is the expected distance between them measured along the surface, that is, the great-circle distance?Tower Research CapitalNew York · 2019
Try it first
What is the expected great-circle distance?
Show the worked solution
pi/2, about 1.571. Rotate the sphere so the first point sits at the north pole; nothing changes, because the second point is uniform. On a unit sphere the surface distance is the polar angle theta of the second point. The northern and southern hemispheres are mirror images, so theta is as likely to be pi/2 - t as pi/2 + t, and its mean is pi/2.
Why can you put the first point at the pole?
Ask how far apart two random towns are on a perfectly round planet, and you can simply stand in one of them: the globe looks the same from every spot on it. Symmetry lets you fix one point anywhere, so the problem shrinks to one random point and its angle from the pole. On a sphere of radius 1, the distance along the surface between the pole and a point at polar angle theta is theta itself, measured in radians, so the question becomes: what is the average polar angle of a uniform point?
With the first point at the pole, the surface distance is the polar angle theta of the second point, whose density (1/2) sin theta is symmetric about pi/2, so the expected distance is pi/2; a uniform angle, the dashed line, puts too many points near the poles. The relationshiptheta the polar angle of the second point, equal to the surface distance on a unit sphere f(theta) the density of that angle sin theta the relative size of the band of latitude at angle theta What it says in wordsThere is more surface near the equator than near the poles, in proportion to sin theta, and that density is symmetric about pi/2.Where does sin theta come from? The circle of latitude at angle theta from the pole has circumference 2 pi sin theta, so a thin band there holds surface in proportion to sin theta: almost none near the poles, the most at the equator. Archimedes put it more neatly: the area of a band is proportional to its height along the axis, so cos theta is uniform between -1 and 1. The density (1/2) sin theta is a mirror image about pi/2, so the mean is pi/2 without doing the integral. Integration by parts confirms it, a numerical integral gives 1.5708, and a seeded simulation of 100,000 pairs gives 1.570.
If a uniform angle gives the same mean, why does the shape matter?
Because the mean survives by luck of symmetry and almost nothing else does. Choosing theta uniformly on 0 to pi crowds points near the poles. Ask for the chance the two points are within 60 degrees of each other and the correct answer is (1 - cos 60 degrees)/2 = 0.25, while the uniform angle says 0.33. Ask for the expected straight-line chord, 2 sin(theta/2), and the correct density gives 4/3, about 1.333, while the uniform angle gives 4/pi, about 1.273. The simulation gives 1.333 for the chord. This is the limitation of the shortcut: it answers this one question and must not be reused for the next.
Where candidates lose it
The commonest wrong answer is 4/3, the expected straight-line chord, which some candidates remember from a related puzzle. The question asks for distance along the surface, which on a unit sphere is the angle itself.
The second loss is the right answer for the wrong reason: picking the angle uniformly between 0 and pi. The mean comes out right by symmetry, but any follow-up on the chord or on the chance of being close gives the wrong number. Say that the band of latitude grows like sin theta.
What the interviewer asks next
- What is the expected straight-line distance between the two points?
- What is the probability that the two points are within 60 degrees of each other?
- Four points are chosen uniformly on a sphere. What is the chance they all lie in one hemisphere?
Asked at Tower Research Capital, Quantitative Research, New York, 2019 (Wall Street Oasis):
a 3d geometry question about the surface distance between points chosen randomly on the surface of a sphere

