Fin Maverick
Foundations VocabularyAccounting & ReportingEconomics & MacroQuant Methods & ProgrammingBusiness & Company AnalysisCorporate Finance & ValuationBehavioural Finance
Banking & Market InfrastructureFixed Income & RatesDerivatives & Structured ProductsPublic EquitiesTransactions & DealsPortfolio ConstructionFunds & AMCs
Private Markets & AlternativesRisk, Treasury & ControlAI & Digital FinanceStochastic Calculus & PricingWealth & Personal FinanceIndian Markets & RegulationProfessional Practice
CalculatorComparison
Frameworks
Explore Bootcamps
Equity ResearchPortfolio ManagementMutual Fund MasteryInvestment Banking Analyst
Private Equity AnalystQuant & Hedge Fund AnalystBreaking Into VCFinancial Analyst Program
Risk Management ProgramPrivate Wealth ManagementDebt Capital MarketsDerivatives Foundation
Explore Free Courses

Equity Research6

Writing an Investment ThesisBuilding a Discounted Cash FlowReading an Annual Report FastReading a Sector Before a CompanySpotting Quality of Earnings Red FlagsBuilding a Revenue Forecast From Drivers

Portfolio Management3

Rebalancing: When, Why and What It CostsStrategic and Tactical Asset AllocationMeasuring Risk in a Portfolio

Mutual Fund Mastery3

Comparing Funds Without Being FooledHow a NAV Is Struck and Which Day You GetReading a Fund Factsheet Properly

Derivatives Unlocked4

Hedging a Real ExposureThe Greeks, PracticallyFutures, the Basis and What Moves ItReading an Option Payoff

AI For Finance2

Retrieval and Grounding for FinanceDocument Extraction in Finance

Breaking Into Quants4

Backtesting a StrategyHypothesis TestingCleaning Financial DataRegression for Finance

Breaking Into VC3

Sizing a MarketReading a Term Sheet as a FounderHow a Venture Round Actually Works

Financial Analyst Program4

Common Size and Trend AnalysisReading a Cash Flow StatementRatio Analysis That Says SomethingBuilding a Working Capital Schedule

Risk Management Program2

Credit Exposure and How It Is ReducedValue at Risk and What It Hides

Investment Banking Analyst3

Precedent Transactions and Why They DifferReading a Term Sheet StructurallyBuilding a Comparable Companies Table

Private Wealth Management3

Tax Aware Portfolio DecisionsBuilding a Client Risk ProfileGoal Based Planning Arithmetic

Debt Capital Markets3

Analysing an Issuer's CreditDuration and What It Does Not Tell YouBond Pricing and Yield Mechanics

Private Equity Analyst2

Fund Waterfalls and CarryThe LBO in Structure

Hedge Funds Analyst2

Short Selling MechanicsLong Short Mechanics
QuarksCourses
Explore Interview Preparation
Investment BankingEquity ResearchVenture CapitalistPrivate EquityHedge Funds
QuantFinancial AnalysisPrivate Wealth ManagementDebt Capital MarketsRisk Management
Derivatives FoundationPortfolio ManagementMutual Fund Mastery
PartnershipsShowdown
Log inSign up
Interview tracksAll
1Investment Banking
Question bankPuzzlesCase studies
2Equity Research
Question bankPuzzlesCase studies
3Venture Capital
Question bankPuzzlesCase studies
4Private Equity
Question bankPuzzlesCase studies
5Hedge Funds
Question bankPuzzlesCase studies
6Quant
Question bankPuzzlesCase studies
7Financial Analysis
Question bankPuzzlesCase studies
8Private Wealth Management
Question bankPuzzlesCase studies
9Debt Capital Markets
Question bankPuzzlesCase studies
10Risk Management
Question bankPuzzlesCase studies
11Derivatives Foundation
Question bankPuzzlesCase studies
12Portfolio Management
Question bankPuzzlesCase studies
13Mutual Fund Mastery
Question bankPuzzlesCase studies
022

Case 022Signal research and data tasksWarm up

A one-day tick file has 2% of rows at price zero, 5% duplicated timestamps and one print of 8,120 between prints of 812. Identify each error, fix it, and show how each would distort daily realised volatility.

HRHudson River TradingAnonymous interview candidate in · 2024

1The situation

Chandrakoop Trading's data round gives you one trading day of one-minute prints for Nirjhari Steel, a stock near Rs 812, in a file that should have 375 rows from 9:15 to 15:30. A quick scan shows three faults: about 2% of rows (8 of them) carry a price of exactly 0; about 5% of rows (19) repeat a timestamp that already appears; and one row prints 8,120 between a print of 812 and another of 812 a minute later.

The desk's standard statistic is daily realised volatility, computed from one-minute log returns, and on a clean day for this stock the minute returns have a standard deviation of about 0.08%.

2Your task

Name each error and its likely cause, say how you would fix it, and show what each one does to the daily realised volatility if left in.

Quick check

Before computing: which single fault does the most damage to the volatility statistic?

Worked solution

Try it on paper, then open one step at a time.

30-second answerThe answer to give first

Clean before you compute, and log every change. The zeros are missing prints, not prices: drop them, because a log of zero or a return of minus 100% makes the statistic undefined. The duplicated timestamps are a feed replay: keep one row per minute, and they would otherwise understate volatility by about 2%. The 8,120 is a decimal shift that reverses the next minute: correct or drop it, because left in it lifts the day's realised volatility from 1.55% to about 326%, 210 times the truth.

Step 1What is each fault, and where does it come from?

Name the cause, because the cause tells you the fix. A price of exactly 0 is never a trade; it is a field the feed filled with its default when no print arrived, or a parser reading a blank as a number. Treat it as missing. A repeated timestamp is usually a replayed packet or two feeds merged without de-duplication; check whether the repeated rows carry the same price, because identical rows are harmless to drop while different prices at the same stamp mean the file's ordering cannot be trusted. A print of exactly ten times the price that reverses a minute later is a decimal shift, a dropped or added digit somewhere between the exchange and your file, and the factor of exactly 10 with a full reversal is the signature that separates it from a real move. A real jump that does not reverse, say a stock going from 812 to 406 and staying there, could be a split or news, and that one you investigate rather than delete.

Three kinds of dirt in one tick file, each flagged8008108208308 rows at price 0 (off scale)one print at 8,120 (off scale, 10x)next minute back at 81219 duplicated timestampssame minute printed twice9:1511:2013:2515:30Price in rupees by minute; the trace is illustrative, the three faults are the ones in the file
The 8 zero prices sit off the bottom of any sensible scale, the 19 duplicated timestamps repeat minutes already in the file, and the single print at 8,120 sits off the top and reverses the next minute, which is the signature of a decimal shift rather than a real trade.
Step 2How much does each one distort the volatility?

Work the clean number first so there is something to compare with. Minute returns with a standard deviation of 0.08% over 375 minutes give a daily realised volatility of 0.08% times the square root of 375, about 1.55%. The zeros do not distort the statistic; they destroy it: a log return into a zero price is the log of zero, and a simple return is minus 100% followed by a division by zero on the way out. Code that silently drops the resulting NaN values will also drop the two returns around each zero and report a number that looks fine, which is worse than an error message.

The duplicates are the quiet fault. If the repeated rows carry the same price, each adds a zero return, so the sum of squared returns is unchanged but there are 394 returns instead of 375; a volatility computed as the standard deviation of minute returns times the square root of 375 falls to about 1.51%, an understatement of about 2%. Small, but it is a bias that appears every day in the same direction, and anything built on the row count, such as volume per minute, is 5% wrong.

The relationship
RVspike=∑ri2+2 (ln⁡10)2=0.00024+10.60≈3.26\text{RV}_{\text{spike}} = \sqrt{\textstyle\sum r_i^2 + 2\,(\ln 10)^2} = \sqrt{0.00024 + 10.60} \approx 3.26
r_ithe clean one-minute log returns
\ln 10the log return of a tenfold price jump, 2.303
\text{RV}daily realised volatility, the square root of the summed squared returns
What it says in wordsTwo squared returns of 2.30 each outweigh the whole day of real returns by a factor of tens of thousands, so one bad print sets the volatility on its own.

The decimal shift is the loud fault: left in, the day's realised volatility is about 326% instead of 1.55%, 210 times too high, and with simple returns instead of log returns the figure is about 904%. A single row has done that. A spreadsheet that says one of your cells is a thousand times the others does not change the average much if you are summing values, but a volatility squares the differences, so one outlier is not diluted by 374 good minutes; it replaces them.

What each fault does to daily realised volatility (log scale)Clean file1.55%Duplicates left in1.51%Decimal shift left in326% (210x)Zeros left inundefined: log of zero, or a return of minus 100%1%10%100%1,000%
On a log axis the clean day's volatility of 1.55% and the duplicate-inflated 1.51% sit side by side, the decimal-shift version of 326% is 210 times higher, and the version with zeros is not a number at all, so one tick decides the statistic unless the file is cleaned first.
Step 3What is the cleaning routine, and what do you never do?

Fix in the order that keeps each step checkable. Drop the zero prices and count them. Sort by timestamp, drop exact duplicate rows, and if two rows share a stamp with different prices keep the last and flag the minute. Then run a plausibility filter on returns: any minute return beyond a threshold, say ten clean standard deviations, that reverses fully within one or two prints and matches a factor of 10 or 100 is a decimal shift, corrected by dividing or dropped, and logged. Never fix silently: write the counts of rows dropped and corrected next to the result, and assert at the end that prices are positive, timestamps unique and increasing, and no return exceeds the threshold, so the next file that breaks one of those stops the pipeline instead of poisoning the statistic. Say the limitation too: a filter that deletes every large return will also delete the genuine crash, so the rule needs the reversal test, not just the size test.

FaultRowsLikely causeFixVolatility if left in
Price 08missing print filled with defaultdrop, countundefined
Duplicate timestamp19replayed packets or merged feedskeep one per stamp, flag1.51% vs 1.55%
8,120 between 812s1decimal shiftcorrect to 812 or drop, log it326% vs 1.55%
The two faults that look like data errors move the statistic by a few per cent or break it; the one that looks like a price moves it 210-fold, so the plausibility filter is the step a data round is really testing.

Where candidates lose it

The common loss is reaching for the volatility formula before looking at the file. The interviewer has planted faults that make the formula either crash or return nonsense, and wants to see the scan, the counts and the fixes come first.

The second is deleting every large return. That removes the decimal shift and also the real crash that comes some other day; the test for a bad print is a reversal and a round factor, not size alone.

What the interviewer asks next

  • How would you detect a decimal shift that does not fully reverse because the stock also moved?
  • If the duplicated timestamps carry different prices, which one do you keep and why?
  • How would you build the same checks so they run automatically on every new day's file?

Asked at Hudson River Trading, Quantitative Research, Anonymous interview candidate in, 2024 (Wall Street Oasis): Final Onsite consists of 4-5 interviews including data analysis, coding and math

← Case 021Two bank stocks have 0.92 daily return correlation but their price ratio drifted from 1.0 to 1.6 over three years; a second pair has 0.55 correlation but a stationary spread with an ADF p-value of 0.01. Which pair do you trade?Case 023 →An options book shows delta 2,000 shares, gamma 300 shares per rupee, vega Rs 2 lakh per vol point and theta minus Rs 1.5 lakh a day. The stock rises Rs 4 and implied volatility rises 1 point; reported P&L is Rs 1 lakh. Attribute the P&L and size the unexplained residual.

Company names and figures are illustrative.

Fin Maverick Free CoursesExplore Free Courses
Fin Maverick BootcampsExplore Bootcamps
Fin Maverick

Finance education that ends in a job, not a certificate that gathers dust. Built for young India.

LEARN
CalculatorsFrameworksComparisonsInterview RoadmapsShowdown
RESOURCES
All CoursesFree CoursesBootcampsInternships
COMPANY
AboutJob openingPartnership
LEGAL
Privacy PolicyTerms & ConditionsContent LicenseReturn & Refund Policy
© 2026 FIN MAVERICK / BUILT FOR INDIA.DO FINANCE, DO NOT JUST READ ABOUT IT.