Case 040Signal research and data tasksCore
Daily card-spend data covers 8% of transactions for 60 listed retailers and arrives with a 3-day lag; results come out 45 days after quarter end. How long is the information window, and how precise can the estimate be?
1The situation
Bhumrika Signals is offered an anonymised, aggregated panel of card transactions covering about 8% of all card spending at 60 listed retailers, delivered daily with a 3-day lag. The retailers report quarterly revenue about 45 days after each quarter ends. Analyst consensus for quarterly revenue growth has historically missed the reported figure by about 3 percentage points, as a typical error.
Back-testing on past quarters suggests that even with perfect sampling, the panel's growth differs from reported growth by about 1.8 points, because the panel's customers, payment methods and online share differ from each retailer's full base.
2Your task
Estimate the information window, the error that comes from sampling only 8% of transactions, which retailers the data can say anything useful about, and what you would check before buying it.
Quick check
What mainly sets the sampling error for one retailer's quarterly spend estimate?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
The clean window is about 42 days, from 3 days after quarter end to the results, and the data is precise enough to beat consensus only for the larger retailers. Sampling error falls with the square root of transactions captured, from about 3.8 points for a small chain to 0.24 for a large one, but a panel bias of about 1.8 points never goes away. Against consensus errors of 3 points, the edge is real for big names and absent for small ones.
Step 1When does the data know something the market does not?
Draw the calendar. Spend builds up through the 91 days of the quarter and arrives three days late, so the full quarter's data is in hand three days after quarter end, while the official number arrives 45 days after it: a 42-day window in which you know roughly what the results will say. Inside the quarter there is a partial window too; two thirds of the quarter is visible by about day 64. But the window is not all yours. Company trading updates, other data vendors and channel checks leak the same information, so the value is highest early in the window and decays towards the results date.
Step 2How precise can an 8% sample be?
Think of an exit poll. Interviewing 2,000 voters gives a similar error whether the constituency has one lakh voters or ten lakh; what matters is the count of people asked, not the share. For card spend, the sampling error of a retailer's total is the variability of a single transaction divided by the square root of the number of transactions in the panel. With a transaction value spread of about 1.5 times the average, a large chain with 50 lakh card transactions a quarter puts 4 lakh in the panel and gets a sampling error of about 0.24 points; a small chain with 20,000 puts 1,600 in and gets about 3.8 points.
Then add the error sampling cannot fix. The panel's cardholders are not the retailer's whole customer base: cash buyers, other card networks and a different online share all sit outside it. That panel biasThe gap between what a data panel measures and the full population it stands for, caused by who is in the panel rather than by how many are in it. More data from the same panel does not shrink it. of about 1.8 points sets a floor, so even the largest chain's total error is about 1.8 points. For a small chain the two combine to about 4.2 points, worse than consensus. Setting the combined error equal to the 3-point consensus error gives the break-even size: about 49 thousand card transactions a quarter.
Step 3What would you check before buying the data?
Three things, each testable. First, the history: how many past quarters exist, so the bias can be measured per retailer and not assumed; with 60 retailers and four quarters a year there are only 240 results a year to learn from. Second, stability: banks join and leave card panels, so check that coverage of each retailer has not jumped, which would look like a sales surge. Third, the legal basis: how the data was collected and anonymised, and whether compliance is satisfied that it is not inside information. Then estimate the economics: the trade only exists for the larger retailers, inside a window that shrinks as more funds buy the same feed.
Where candidates lose it
The usual loss is treating 8% coverage as the precision: saying an 8% sample is too small, or that it is plenty, without converting it to a count of transactions per retailer. The same share is precise for a large chain and noisy for a small one.
The second is assuming more data fixes everything. Sampling error shrinks with size; panel bias does not. A candidate who separates the two, and measures the bias from history, has answered the question the interviewer cares about.
What the interviewer asks next
- How would you adjust for a bank joining the panel halfway through a quarter?
- How would you turn the spend estimate into a trading signal around the results date?
- Which retailers would you expect the panel to track worst, and why?
Company names and figures are illustrative.
