Hedge Funds interview preparation
Long-short equity, macro, event-driven, distressed, multi-manager platforms and the Indian Category III landscape. Every question is either traced to a named firm from a public candidate report, or tagged at desk level when we could not trace it — answers lead with the point, then the mechanism, then the limitation.
100 questions, mapped to the firms that asked them
- Questions
- 100
- Traced to a firm
- 39
- Firms
- 16
- Updated
- September 2026
079Explain the construction of a factor. Why that method, and how would you optimise it?AQR Capital ManagementInvestment Research · New York · 2021
Say this
Take value as the example. Define the signal, in this case book to price or a composite of several value measures; clean and winsorise it; standardise it cross-sectionally within industry; then build a long-short portfolio from the ranks, usually top minus bottom quintile, weighted and rebalanced on a defined schedule. Every one of those steps is a choice, and the choices matter as much as the signal.
Then walk it
- Signal definition first, and use a composite rather than a single ratio. Book to price, earnings to price, cash flow to price and sales to enterprise value capture the same idea with different noise, so the average is more robust than any one. That is the main argument for composites over single metrics.
- Then the cleaning: point-in-time data with the correct reporting lag so you are not using numbers before they were published, delisted returns included so you are not survivorship biased, winsorise or rank-transform the outliers, and handle negative book values explicitly.
- Then neutralisation. Standardise within industry, because a raw value screen just buys banks and sells software. Neutralise size too, or the factor becomes a small-cap bet. The choice of what to neutralise defines what the factor actually measures.
- Then portfolio construction: quintile or decile spreads, equal weight versus value weight, rebalance monthly or quarterly. Equal weight shows a stronger factor premium and is much harder to trade. Say that trade-off out loud, because it is where academic factors and investable factors part company.
- On optimisation, define the objective honestly: maximise net-of-cost information ratio, not gross return, with constraints on turnover, capacity and exposure to other factors. Then use cross-validation across time and across regions rather than optimising a single sample.
- And say the limitation before being asked, because this is the real question inside the question. With enough parameters you can produce any backtest you like. The defences are economic priors before data mining, a small number of specification choices, out-of-sample and out-of-region testing, sensitivity analysis showing the result is not knife-edge, and a documented count of how many specifications you tried. A factor that only works with one lookback and one weighting scheme is a coincidence.
Where candidates lose it
Describing the signal and skipping the construction choices. Neutralisation, point-in-time data and the equal-versus-value weighting decision are where the real work is. And on 'how would you optimise it', a candidate who does not immediately raise overfitting has failed the question at a firm built on factor research.
Expect next
- How would you know you had overfitted?
- Why neutralise by industry?
- How would you test whether your new factor is distinct from momentum?
Reported by candidates at AQR Capital Management (Investment Research, New York, 2021). Source: Wall Street Oasis.
080What are the assumptions of linear regression?Squarepoint CapitalHedge Fund · Montreal · 2024
Say this
Linearity in the parameters, exogenous errors with zero conditional mean, no perfect multicollinearity, homoscedastic and uncorrelated errors, and for exact small-sample inference, normally distributed errors. The first two give you unbiasedness; the rest are about whether your standard errors mean anything.
Then walk it
- Separate the tiers, because that is what distinguishes someone who has used regression from someone who memorised a list. Linearity and exogeneity are needed for the coefficients to be unbiased. Homoscedasticity and no autocorrelation are needed for the usual standard errors to be correct. Normality is only needed for exact t and F inference in small samples.
- So a violation of homoscedasticity does not bias your beta, it biases your confidence in it. That distinction matters enormously in practice: you can still use the estimate, you just cannot trust the t-statistic.
- In financial time series the assumptions that actually break are autocorrelation and heteroscedasticity, because volatility clusters and returns overlap. The standard fixes are Newey-West or White standard errors, and clustered errors in panel data.
- Endogeneity is the serious one. If a regressor is correlated with the error, the coefficient is biased and no standard error fix helps. In finance this usually arises from omitted variables or from a feedback loop where price affects the supposed predictor.
- Multicollinearity does not bias anything, it just inflates variances, so coefficients become unstable and flip sign between samples. That is very common with factor exposures, and the tell is a large R-squared with no individually significant coefficient.
- Practical additions I would name: outliers dominate least squares because it minimises squared errors, so winsorise or use robust regression; and out-of-sample performance matters more than any in-sample diagnostic, because for a trading signal I care about prediction, not about the p-value.
Where candidates lose it
Reciting the list without saying what each assumption buys you. Tiering them into unbiasedness versus valid inference is the differentiator. Also, do not claim normality of the dependent variable is required; it is normality of the errors, and only for small-sample inference.
Expect next
- Which assumption is most often violated in financial data, and what do you do about it?
- What is the consequence of multicollinearity?
- How would you detect endogeneity?
Reported by candidates at Squarepoint Capital (Hedge Fund, Montreal, 2024). Source: Wall Street Oasis.
081Two series can be negatively correlated within each month but positively correlated over a full year. How?Squarepoint CapitalHedge Fund · Montreal · 2024
Say this
Because correlation measured within groups and correlation measured across the pooled data answer different questions. If both series share a common upward trend across months, the between-month variation is positive and can dominate the negative within-month relationship. It is Simpson's paradox in a time series.
Then walk it
- Decompose the covariance into within-group and between-group parts. Total covariance equals the average within-month covariance plus the covariance of the monthly means. Those two terms can have opposite signs, and whichever has more variance wins the pooled number.
- Concrete picture: every month, A and B move in opposite directions day to day, so within-month correlation is negative. But each month both drift higher, so the monthly averages rise together. Pool the daily data over a year and the shared drift dominates.
- The generic driver is a common slow-moving factor. Both series load positively on something persistent, such as inflation, liquidity or a market trend, while their high-frequency innovations offset. Long-horizon correlation is dominated by the common factor and short-horizon correlation by the idiosyncratic part.
- There is also a pure measurement version of this: correlation of returns is horizon dependent when returns are autocorrelated. Compute correlation on daily returns and on annual returns for the same pair and you generally get different numbers, and neither is wrong.
- Why it matters practically, which is what the interviewer is really testing: hedge ratios and diversification estimated at one horizon do not hold at another. A pair that looks hedged on daily data can be a directional bet over a year, which is exactly how a relative value book acquires an unintended factor exposure.
- So the answer to 'which correlation is right' is neither. You choose the horizon that matches your holding period and your rebalancing frequency, and you look at both to know which part of the relationship you are actually trading.
Where candidates lose it
Treating it as a paradox to be resolved rather than a decomposition to be stated. Write down the within-plus-between covariance split and the answer is immediate. And do not stop at the maths: the reason they ask is the practical consequence for hedge ratios at different horizons.
Expect next
- Which correlation would you use to set a hedge ratio?
- How does return autocorrelation affect measured correlation?
- Give me another example of Simpson's paradox in markets.
Reported by candidates at Squarepoint Capital (Hedge Fund, Montreal, 2024). Source: Wall Street Oasis.
082Talk me through your research process for a systematic signal. How do you avoid fooling yourself?Balyasny Asset ManagementQuantitative Trading · London · 2025
Say this
Start with an economic reason the signal should work, then test it in a way that can fail. Hypothesis first, then data preparation, then a simple specification, then out-of-sample and out-of-region validation, then costs, then capacity. The discipline is that the hypothesis comes before the backtest, not after it.
Then walk it
- State the economic mechanism first and write it down before running anything. Who is on the other side, and why are they there? A signal with no story about who is losing money is almost certainly a data artifact.
- Then the data work, which is most of the time and all of the risk. Point-in-time data with correct reporting lags, restated figures handled properly, delisted and merged companies included, corporate actions adjusted, and survivorship bias eliminated. Look-ahead bias is the most common silent killer and it always flatters the result.
- Then the simplest possible specification. One parameter, no optimisation, sensible defaults. If the effect does not show up in the naive version, it probably is not there. Elaboration after validation, never before.
- Then validation that can actually fail: hold out a period you never look at, test in other regions and other asset classes, test across sub-periods, and check that the result is not driven by a handful of stocks or one month. Report the number of specifications you tried, because a t-statistic loses its meaning after the twentieth attempt.
- Then costs and capacity, which kill more signals than statistics do. Model spread and impact, compute net-of-cost performance at realistic size, and check whether the signal survives a one-day implementation lag. A signal requiring same-second execution is not a signal for a fundamental-horizon book.
- Then the honest self-checks: decide the kill criteria before the test, keep a research log of everything tried including the failures, and have someone else reproduce the pipeline. The uncomfortable truth is that most published anomalies do not replicate out of sample, so my prior on my own new signal should be low.
Where candidates lose it
Describing a backtest rather than a research process. The order of operations is the answer: hypothesis, then data hygiene, then a naive test, then validation, then costs. A candidate who does not mention point-in-time data, look-ahead bias and the multiple-testing problem will not get through a quant research interview.
Expect next
- How many specifications did you try on your last project?
- How do you handle restated financials in a backtest?
- What is your kill criterion for a signal?
Reported by candidates at Balyasny Asset Management (Quantitative Trading, London, 2025). Source: Wall Street Oasis.
083What is the angle between the hands of a clock at 3:15?Man GroupEquity Hedge · London · 2016
Say this
7.5 degrees. The minute hand is exactly at 90 degrees, but the hour hand has moved a quarter of the way from 3 towards 4, which is a quarter of 30 degrees, so it sits at 97.5. The difference is 7.5.
Then walk it
- Set up the units once and the whole family of these questions becomes trivial. The hour hand moves 360 degrees in 12 hours, so 0.5 degrees per minute. The minute hand moves 360 in 60 minutes, so 6 degrees per minute.
- Positions from 12 o'clock: minute hand is 15 times 6, which is 90. Hour hand is 3 times 30 plus 15 times 0.5, which is 90 plus 7.5, so 97.5.
- Difference is 7.5 degrees, and it is the smaller of the two angles, which is what the question means unless it says otherwise.
- The general formula worth memorising: the angle equals the absolute value of 30 times hours minus 5.5 times minutes. At 3:15 that is 90 minus 82.5, which is 7.5.
- The whole trap is the hour hand. Candidates say zero because they picture the hour hand parked on the 3. It is not; it moves continuously, and that is the entire point of the question.
- Say the answer, then say the setup in one line. In a phone screen this question is testing whether you can be quick and precise about a small thing, so do not over-narrate.
Where candidates lose it
Answering zero. It is by far the most common response and it comes from forgetting that the hour hand moves continuously. Also, say which angle you are giving, the smaller one, and do not spend ninety seconds deriving a formula the interviewer already knows.
Expect next
- When is the next time the hands overlap exactly?
- How many times a day do the hands form a right angle?
- What is the angle at 9:45?
Reported by candidates at Man Group (Equity Hedge, London, 2016). Source: Wall Street Oasis.
084I roll two fair dice. What is the probability the sum is 7, and what is the probability of at least one six?Wolverine TradingEquity Hedge · Chicago · 2025
Say this
A sum of 7 is 6 out of 36, so one in six. At least one six is 1 minus the probability of no sixes, which is 1 minus 25 over 36, so 11 out of 36, a bit under a third.
Then walk it
- Count the sample space first: 36 equally likely ordered outcomes. Ordered matters, and treating the dice as indistinguishable is the classic way to get these wrong.
- Sum of 7 has six combinations: 1-6, 2-5, 3-4, 4-3, 5-2, 6-1. So 6 over 36, which is one in six. Worth knowing that 7 is the most likely sum, and the distribution of sums is a triangle peaking at 7.
- For at least one six, use the complement. No six on either die is 5 over 6 times 5 over 6, which is 25 over 36. So at least one six is 11 over 36, about 30.6 percent.
- Note why it is not 2 over 6. Adding the two individual probabilities double counts the double six, so you subtract it: 6 over 36 plus 6 over 36 minus 1 over 36 equals 11 over 36. Inclusion-exclusion gives the same answer and it is worth saying both ways.
- The general rule that follows: for at least one of anything, go to the complement. It converts a messy union into a product, and it is the single most useful reflex in dice and coin questions.
- Then the standard follow-up they are setting up: given the sum is 7, the probability that one die is a 6 is 2 out of 6, so one third, because conditioning restricts the sample space to the six ordered pairs. Answer these fast and cleanly, and say the fraction before the decimal.
Where candidates lose it
Saying 2 over 6 for at least one six, which double counts the double six. And treating the dice as unordered, which wrecks the sample space. Say 'complement' out loud and do the arithmetic in fractions. At a prop shop these are timed, so speed and a clean statement of the sample space matter as much as the answer.
Expect next
- Given the sum is 7, what is the probability one die shows a 6?
- What is the expected number of rolls until you see a six?
- I pay you the sum of the dice. What would you pay to play?
Reported by candidates at Wolverine Trading (Equity Hedge, Chicago, 2025). Source: Wall Street Oasis.
Firm tags come from public, anonymous candidate reports on Wall Street Oasis: strong signal, not sworn testimony. Firms are named as the places a question was reported, not as partners of Fin Maverick. Answers are written for this page to show how to think out loud; they are not scripts to recite.
