Case 065Signal research and data tasksCore
You are asked to build a model of monthly rents for 5,000 Mumbai flats from carpet area, floor, distance to the nearest station and building age. Choose the target, the features and the validation scheme, and interpret a coefficient of 0.9 on log area.
1The situation
Awasiya Analytics has a dataset of 5,000 rented flats across Mumbai, each with its monthly rent, carpet area in square feet, floor number, distance to the nearest railway or metro station, building age, and the locality it sits in. Several flats often come from the same building. The listings span two years.
In a model design interview you are asked how you would build a model to predict rents for flats coming onto the market next quarter. Your first fit puts a coefficient of 0.9 on log carpet area, and the interviewer asks what it means.
2Your task
Choose the target variable, the features and their form, and the validation scheme, explain the 0.9, and name what would make the model's error estimate too optimistic.
Quick check
With log rent as the target and log area as a feature, what does a coefficient of 0.9 mean?
Worked solution
Try it on paper, then open one step at a time.
30-second answerThe answer to give first
Model log rent, use log area, floor, log distance, age and locality, validate by holding out whole buildings and the latest months, and read 0.9 as an elasticity: a 10% larger flat rents for about 9% more. Logs turn a fanning spread into even percentage errors. Random splits would leak building quality between training and test and flatter the error. Convert predictions back with a small correction, since exponentiating predicted log rent understates the average.
Step 1Why model log rent rather than rent?
Rents in Mumbai run from about Rs 15,000 to several lakh, and the errors grow with the rent: a model that is Rs 5,000 off on a Rs 25,000 flat is badly wrong, while the same miss on a Rs 2 lakh flat is excellent. Taking logs makes errors roughly proportional, so the model is judged on percentage misses, and it turns the multiplicative way prices work, location times size times condition, into a sum a linear model can fit. The left panel below shows the raw spread fanning out as flats get bigger; the right shows the same points with an even spread around a straight line once both axes are logged.
Step 2What does the coefficient of 0.9 on log area mean?
In a model with logs on both sides, the coefficient is an elasticityThe percentage change in one quantity that goes with a 1% change in another.. A 0.9 means a 1% larger flat rents for about 0.9% more, so 10% more area gives 1.1 to the power 0.9 = 1.0896, about 9.0% more rent. A 500 square foot flat at Rs 40,000 becomes about Rs 43,583 at 550 square feet, other things equal; doubling the area raises rent by 87%, not 100%. Because 0.9 is below 1, rent per square foot falls slightly with size, by about 0.9% for each 10% more area, which matches how larger flats are actually let.
Step 3Which features, and in what form?
Locality is the feature that matters most and the one candidates forget. Rent per square foot can differ several times between two Mumbai localities, so without locality effects the area coefficient soaks up location, because bigger flats cluster in some areas. Use distance to the station in logs, since the first kilometre matters far more than the fifth. Floor is not straight-line either: ground floors and very high floors in old buildings without lifts often rent lower, so use bands or a curve, and add a lift flag if the data has it. Age enters in bands, because a new building and a 40-year-old one differ in kind as well as degree.
| Choice | What to use | Why |
|---|---|---|
| Target | Log of monthly rent | Errors in percentages; multiplicative pricing becomes additive |
| Area | Log carpet area | Coefficient is an elasticity; curve becomes a line |
| Location | Locality fixed effects | Biggest driver; otherwise area absorbs it |
| Station distance | Log distance, km | First kilometre matters most |
| Floor and age | Bands, plus a lift flag | Effects are not straight lines |
| Validation | Hold out whole buildings and the latest quarter | Flats in one building share unseen quality; the model predicts forward |
Step 4How should the model be validated?
Ask what the model will face: new flats next quarter, many in buildings it has never seen. A random split puts flats from the same building in both training and test sets, so the model is graded partly on memorised building quality and its error looks smaller than it will be. Hold out whole buildings in cross-validation, and keep the most recent quarter as a final test so time trends in rent are not leaked backward. Report median absolute percentage error, which a client understands. Finally, when converting back to rupees, exponentiating predicted log rent gives a median, not a mean; with residual spread of about 0.25 in logs, multiply by about 1.032 if you need the average rent.
Where candidates lose it
Candidates often read 0.9 as Rs 0.9 per square foot or as 90% of variance explained. In a log-log model it is an elasticity, and saying so in percentage terms is the core of the answer.
The second miss is a random train-test split. Flats from the same building share quality the features do not capture, and only a building-level split gives an honest error for flats in new buildings.
What the interviewer asks next
- How would you handle a locality with only five flats in the data?
- The model's error is twice as large for flats above Rs 1 lakh. What would you check?
- Would you try a tree-based model here, and how would you compare it fairly with the linear one?
Asked at Two Sigma, Quantitative Research, New York, 2025 (Wall Street Oasis): Three rounds of tech interviews. One is like model design- predict rent prices in Manhattan
Company names and figures are illustrative.
