Case 087Systematic research and dataCore
Open-ended: would monsoon rainfall help predict quarterly sales for Bhumora Agro and four other farm-input makers? You have 15 years of district rainfall and quarterly sales. How would you build and test the signal, and how many independent observations do you really have?
1The situation
Bhumora Agro sells seeds, fertiliser and crop protection, mostly in six states, and four listed peers sell similar products in overlapping regions. A researcher proposes that rainfall during the monsoon, which is observed day by day before the companies report, predicts their sales for the season better than analysts' forecasts do.
You have 15 seasons of district-level rainfall, each company's quarterly sales and the consensus forecast made before each season. The researcher has already computed that across all 75 company-seasons the rainfall anomaly and the sales surprise are positively correlated.
2Your task
How would you build the signal, how would you test it, and how many independent observations does the test really have?
Quick check
Five companies over 15 seasons: how many independent observations of the weather effect do you have?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
You have about 15 independent observations, one per monsoon, not 75, so the test is weak. Build a rainfall index for each company weighted by where it sells, compare it with the sales surprise against consensus, and fix the method before looking at results. Across 15 seasons the correlation is 0.47, below the 0.51 needed to pass a standard test. Pooling companies adds rows but not independent weather.
Step 1How do you turn rainfall into a signal for one company?
Think of an umbrella seller with stalls in five towns. Total rainfall across the country tells you little; rain in the five towns, weighted by how much each stall sells, tells you a lot. So weight each district's rainfall anomaly, its departure from the long-run average, by the share of the company's sales that come from that district. Bhumora, strong in six states, gets a different index from a peer selling mostly in the south. Then predict the sales surprise against the consensus made before the season, not raw sales, because the market already expects some weather effect.
Step 2How do you test it without fooling yourself?
Write the method down before you run it: which districts, which weights, which quarters, which measure of surprise. Every choice made after seeing results is a hidden extra test, and with 15 seasons a few such choices are enough to manufacture a pattern. Hold out the last five seasons and test only once. Check that the effect has the same sign for each company, and that the timing works: rainfall is known before the results, but ask whether the stock already moved during the monsoon, in which case the market has the signal too.
Step 3Why are 75 rows only about 15 observations?
Because the five companies share one sky. When a monsoon fails, all five miss together; the second, third and fourth company add little that the first did not already say. If their surprises in a season correlate at about 0.7, the 75 rows carry the information of roughly 20 independent ones, and the weather itself happened only 15 times. A t-statistic computed as if the 75 rows were independent would look about twice as convincing as it should.
| N | rows in the dataset, 75 |
| m | companies per season, 5 |
| rho | correlation of the companies' surprises within a season, about 0.7 |
| r crit | correlation needed to pass a two-sided 5% test with 15 observations |
Step 4So is the signal worth pursuing?
Possibly, but not on this evidence alone. The mechanism is sound, since farmers buy more inputs after good rain, but a correlation of 0.47 over 15 seasons is what chance produces often enough that a research committee would reject it. Three ways to strengthen it: go to finer detail within a season, such as monthly rainfall against monthly dispatches where companies disclose them; add companies whose regions do not overlap, which brings more independent weather; and wait, adding seasons live and out of sample. Say the limitation plainly: a signal this thin is a small position until more monsoons arrive.
Where candidates lose it
The frequent miss is counting 75 observations, computing a t-statistic that looks significant, and declaring a signal. Rows that share the same weather are not independent, and the interviewer asks the question to see whether you notice.
The second is predicting raw sales instead of the surprise against consensus. Analysts already watch the monsoon; a signal is only useful if it predicts what they missed.
What the interviewer asks next
- How would you check whether the stock prices already move with rainfall during the season?
- Would adding district-level data within each season fix the sample size problem?
- How would you size a position on a signal whose t-statistic is just under 2?
Company names and figures are illustrative.
