State Variables: The Quantities That Describe the System
A state variable is a quantity that has to be carried forward for the future of a model to be described. The test is sufficiency: once the state is known, the rest of the past adds nothing. How many a model needs is settled by that test rather than by taste, and every extra one widens what the model can say while multiplying what it costs to solve.
Everything in this guide rests on one sentence: the state is whatever the test says it is. The sentence sounds unremarkable until it is set against what it replaces. Choosing which quantities a model carries looks like an aesthetic decision, the sort of thing a modeller settles by habit or by what was in the last model built. It is not. There is a question that can be put to any candidate quantity, the answer is either yes or no, and the collection of quantities that answer yes is the state. The same sentence also explains why two careful people can look at the same thing and honestly reach different counts: they are not disagreeing about style, they are disagreeing about what the future depends on.
What makes a quantity a state variable?
A lift is the simplest illustration. Suppose the only thing known about it is which floor it is on, and it is on floor four. The floor number alone is not enough to say where the lift will be in ten seconds, and the reason is worth being precise about. Two lifts can both be sitting at floor four with completely different futures: one is on its way up and one is on its way down. Knowing the floor is not enough. Now suppose the floor and the direction of travel are both known. Two lifts that agree on both have the same future, in the sense that matters here, and nothing about how either one arrived at floor four changes that.
The lift exercise is the whole test, and it is a test rather than a definition. The test does not ask whether a quantity feels important. The test takes two situations that agree on the candidate quantity today and disagree about everything else in the past, and asks whether their futures are the same. If the futures agree for every such pair, the candidate is sufficient and the rest of the history can be thrown away. If a single pair can be found whose futures differ, the candidate is not enough, and something has to be added to it.
Notice how narrow the test is. The test does not ask whether the quantity is interesting, whether it is easy to measure, whether it is quoted anywhere, or whether the modeller likes it. The test asks one thing, and that thing is sufficiencyThe condition that once a quantity is known, the rest of the history makes no difference to what happens next.. A quantity that passes the test earns its place by making the past irrelevant. A quantity that fails does not get in by being important.
Written formally, the test says that the chance of any future event, worked out from the whole history, is the same as the chance worked out from the candidate quantity alone. The Markov property, set out separately, is the statement that the present is a sufficient summary. State variables are the other half of that idea: given that the present has to be sufficient, what exactly has to be in the present?
| \(X_t\) | the candidate state at time \(t\), which may be one quantity or several written together |
| \(S_u\) | the standard process at a later time \(u\) |
| \(\mathcal{F}_t\) | everything known at time \(t\), the whole history and not just the latest reading |
| \(A\) | any set of future values that might be asked about |
| \(\mathbb{P}\) | the physical measure, the rule under which the chances are taken here |
Two situations agree on today's level of the standard process and disagree on where it stood last month. In the one variable model, do they have the same future?
How many state variables does a model need?
The honest answer is: as many as the test forces, and not one more. The answer is a real procedure rather than a slogan. The procedure starts with the shortest candidate imaginable, usually the single quantity the model is about, and puts the test to it. If it passes, the work is finished and the model carries one variable. If it fails, the failure hands over a pair of situations that agree on the candidate and have different futures, and the difference between those two situations is what has to be added. The difference is added, and the test is put to the widened candidate. The cycle repeats until nothing fails.
The count comes out of that procedure rather than going into it, and nowhere in the procedure does taste enter. Two people can still reach different counts without either of them being careless. Two counts differ when the two people disagree about what the future depends on. The disagreement is a claim about the thing being modelled rather than a preference about models. One person thinks a level standing still tells the whole story; the other thinks how violently it has been moving lately also matters. The two modellers will build models of different widths, and the disagreement is worth having in those terms rather than as an argument about elegance.
There is a squeeze on both sides. Too few variables and the model simply cannot express the thing asked of it: no amount of careful solving will get an answer out of a description that never held the relevant quantity. Too many and the model becomes something nobody can solve. Failing to solve it is a different way of getting no answer. Every model in this subject sits somewhere between those two failures, and the number of state variables is precisely where that position is set.
Is the number of state variables in a model a matter of taste?
How is a state variable told apart from a parameter and from an output?
Most of the confusion about state variables sits here, and it clears up quickly once it is seen that there are three roles rather than two. The future depends on a state variable, so it moves as the model runs and has to be carried forward. A parameterA number the model holds fixed, which does not move while the model runs and therefore never needs carrying forward. is a number the model holds fixed. A parameter shapes how the state moves and does not move itself, so there is nothing to carry. An outputA quantity computed from the state whenever it is wanted, which never has to be carried separately. is computed from the state whenever it is wanted. Carrying an output separately would be carrying the same information twice.
Two questions separate the three cleanly. Does it move while the model runs? If no, it is a parameter. If yes, can it be computed from the other things already being carried? If yes, it is an output. If no, it is a state variable and the model is stuck with it.
Run those questions on the standard process. The invented drift of 8 per cent a year does not move, so it is a parameter. The level of the process moves and cannot be recovered from anything else, so it is a state variable. The at-the-money contract at Rs 100/- prices at Rs 10.45/- on the one variable model. Its value moves constantly. The level and the elapsed time are what moved, and the value can be recomputed from those two at any moment. The contract value is an output. Carrying an output as though it were a state variable is a common and expensive mistake. The width of the model doubles, and not a single thing the model can express is added.
Now for the sharp part. The volatility of 20 per cent a year is a parameter in the one variable model and a state variable in the two variable one. The symbol is the same, the number is the same, the process it describes is the same. The model changed, not the volatility. The role belongs to the model rather than to the quantity, and a quantity has no role at all until a model has been specified. This is why an argument about whether volatility "is" a state variable never resolves: it is not a question about volatility.
When several state variables are carried together they are written as one object and called the state vectorThe full set of state variables written together as a single object, so the model can be described as a function of one thing.. The state vector is more than notation. Once the state is one object, everything the model produces is a function of that object and of time, and the whole apparatus of the subject can be written down without knowing how many components the object has.
| \(X_t\) | the state vector at time \(t\), holding every quantity the model must carry |
| \(d\) | how many state variables there are, which is the width of the model |
| \(V_t\) | any quantity the model produces, such as the value of a contract |
| \(V(t,\cdot)\) | the function that turns the time and the state into that quantity |
In the one variable model, is the volatility of 20 per cent a year a state variable or a parameter?
The at-the-money contract prices at Rs 10.45/- on the one variable model, and that value changes every time the level moves. Is the value a state variable?
What does an extra state variable cost?
Two things, and they behave very differently. The first is what has to be carried, and it grows gently: one more number in the state vector, one more line in every description, one more thing to keep track of. If that were the whole story, nobody would think twice about adding a variable. The second is what has to be solved, and that is where the trouble is.
Most methods for solving a model of this kind lay a grid over the state and work on the grid. Suppose 100 points along each state variable, a modest choice and one used throughout this guide. With one state variable the grid is a line of 100 points, so there are 100 cells. With two, every one of those 100 points has to be paired with 100 points of the second variable, so there are 10,000 cells. With three, every one of those 10,000 has to be paired with 100 more, so there are 10,00,000 cells, ten lakh of them. With four, 10,00,00,000, ten crore.
Each extra state variable multiplies the work by 100 rather than adding 100 to it, and that single fact decides which models actually get built. Here is a counting example with no finance in it at all. From a list of 100 names, writing down every name takes 100 lines. Writing down every pair takes 10,000 lines. Writing down every triple takes ten lakh lines. Nobody thinks the third task is three times the first. The grid works exactly like that, and so does the memory it needs and the time each sweep of it takes.
| \(m\) | grid points along each state variable, taken as 100 throughout this guide |
| \(d\) | the number of state variables, which is the width of the model |
| \(m^{\,d}\) | the number of cells the grid holds, which is what has to be swept |
The name for this is the curse of dimensionalityThe name for the way work multiplies as each extra state variable is added, coined by Bellman in the setting of dynamic programming., a phrase coined by Bellman in the setting of dynamic programming and used ever since. The dimensionHow many state variables a model carries, which is what sets the cost of solving it. of a model is simply how many state variables it carries. No other single number says more about what solving that model will take.
| State variables | Model in this subject | Grid cells | Written out | Numbers to hold |
|---|---|---|---|---|
| 1 | the level alone | 100 | one hundred | 800 bytes |
| 2 | the level with its variance | 10,000 | ten thousand | 80,000 bytes |
| 3 | the level, its variance and the rate | 10,00,000 | ten lakh | 80,00,000 bytes |
| 4 | beyond what this subject builds | 10,00,00,000 | ten crore | 80,00,00,000 bytes |
The last column takes 8 bytes to hold one number, the ordinary choice, and the two ends of that column are worth reading together. The one variable grid fits in less space than this paragraph. The four variable grid needs eight hundred megabytes just to hold one copy of itself, before anything has been computed and before any of the several copies a solver typically keeps have been allocated. The 100 points a variable and the 8 bytes a number are stated choices. A finer grid or a wider number type moves every figure in the table. The multiplication by 100 at each step does not move.
One state variable on a grid of 100 points takes 100 cells. Before the control below is moved: how many cells for three?
Add a state variable and watch the grid stop being drawable
One control: how many state variables the model carries, from one to four. Each band below holds exactly 100 marks, and each mark stands for a whole copy of the band above it. The default of two state variables is the middle model in the worked instance, holding 10,000 cells.
Name the two things an extra state variable buys and costs.
Which quantities are the state variables in the models this subject builds?
Three models on the same standard process, the invented traded quantity that all of this reading runs on: it starts at Rs 100/-, drifts at 8 per cent a year under the physical measure, carries a volatility of 20 per cent a year, and is watched over one year against a risk-free rate of 5 per cent. Later reading builds each of these three models properly. Counting their state variables is the lesson here, so the counting is all that happens now.
Model one, the level alone
One state variable: the level of the process. Neither the drift of 8 per cent nor the volatility of 20 per cent moves as the model runs. Both are parameters. Look at the locked path, the published twelve step path this reading order draws whenever a path is needed. At month six it reads Rs 111.08/-. In this model, that one number describes the entire future of the process. The fact that it came down from Rs 100/- through Rs 97.64/- and up through Rs 107.63/- adds nothing whatsoever, and neither does the month nine reading of Rs 93.74/- that has not happened yet. Two paths standing at Rs 111.08/- have identical futures here, however differently they arrived.
| \(S_t\) | the standard process at time \(t\), the only state variable here |
| \(\mu\) | the drift, a parameter fixed at 8 per cent a year |
| \(\sigma\) | the volatility, a parameter fixed at 20 per cent a year |
| \(W_t\) | standard Brownian motion under the physical measure |
| \(X_t\) | the state vector, which here has one component |
Model two, the level with its variance
Now let the variance move as well. The model that does so is named after Heston in 1993. The variance starts at 0.04, whose square root is exactly the locked volatility of 20 per cent, and it is pulled toward a long-run level of 0.04, the same number, at a speed of 2.0 a year, with a volatility of volatility of 0.30 and a correlation with the process of minus 0.7. The positivity condition holds with room: twice the speed times the long-run variance is 0.16, against a squared volatility of volatility of 0.09.
The level alone now fails the test, and it fails on exactly the pair of situations the test asks for. Two paths both standing at Rs 111.08/- can have different variances, one quiet and one violent, and their futures are plainly different: one has a narrow spread of next month's outcomes, the other a wide one. Agreeing on the level is no longer enough to make the futures agree. The state has to hold the pair, and it is a pair rather than a single number for a reason that can now be stated precisely rather than asserted.
| \(v_t\) | the variance process, the second state variable, starting at 0.04 |
| \(\theta_v\) | the long-run variance, a parameter, also 0.04 |
| \(\kappa_v\) | the speed of reversion, a parameter, 2.0 a year |
| \(\xi\) | the volatility of the variance, a parameter, 0.30 |
| \(W^{(1)},W^{(2)}\) | two Brownian motions with correlation \(\rho\), a parameter, minus 0.7 |
Model three, the level, its variance and the rate
Let the interest rate move too. The model that does so is named after Vasicek in 1977. The short rate starts at 5 per cent, deliberately the same 5 per cent as the constant risk-free rate, and is pulled toward a long-run 6 per cent at a speed of 0.5 a year with a rate volatility of 1 percentage point a year. The reversion gives a half-life of 1.386294 years, an expected level of 5.393469 per cent after one year, and a standard deviation of that level of 0.795060 percentage points, against a stationary standard deviation of 1.000000 percentage points exactly.
Why carry the rate at all? Because of a computed gap rather than an assertion. Under this model the one year zero-coupon bond is 0.949216, a zero rate of 5.211896 per cent. Under a flat 5 per cent the bond is 0.951229, a zero rate of 5 per cent exactly. The 0.211896 percentage point gap between those two zero rates is the entire reason a moving rate has to be carried as a state variable rather than held as a parameter, and it is a number that can be checked rather than a claim that has to be accepted. The pull toward 6 per cent and the variance around it both move the bond, and a model that holds the rate fixed cannot produce either effect at any setting of any parameter.
| \(r_t\) | the short rate, the third state variable, starting at 5 per cent |
| \(\theta\) | the long-run mean of the rate, a parameter, 6 per cent |
| \(\kappa\) | the speed of reversion, a parameter, 0.5 a year |
| \(\sigma_r\) | the rate volatility, a parameter, 1 percentage point a year |
| \(W^{(3)}\) | the Brownian motion driving the rate |
The word for that pull toward a long-run level, used in both the second and the third model, is mean reversionThe tendency of a quantity to be drawn back toward a long-run level, at a speed the model sets., and notice that it is a property of how a state variable moves rather than a reason to stop carrying it. A quantity that reverts still has to be carried. Where it stands today still changes where it will be tomorrow.
Two paths of the standard process both stand at Rs 111.08/- and carry different variances. Is the level alone a sufficient state?
What happens when a state variable is missing?
The failure is hard to catch for one reason: a model missing a state variable does not stop working. The model keeps producing numbers. The model disagrees with something instead, quietly and repeatedly, and each disagreement is small enough to fix by hand.
Picture the adjustment log. On the first week the volatility parameter is nudged from 20.0 to 21.4 because the model was reading a little low. On the second week from 21.4 to 22.6. On the third it comes back to 21.9, then 20.8, then out to 22.7, then 23.5. Every one of those looks like a small correction, the kind of routine maintenance any working model needs. Nobody writes down the sequence as a sequence.
The sequence is the missing state variable, and the pattern of the adjustments is the shape of the quantity the model refused to carry. A parameter that has to be re-set every week is not behaving like a parameter. The parameter is moving. Something that moves and that the future depends on is a state variable by the test, and the only question left is whether it gets carried inside the model where it can be reasoned about, or outside it in somebody's habit of nudging.
The error that gets made, and what it costs
Treating a parameter as though it were a state variable, or a state variable as though it were a parameter. The second direction is the damaging one. The volatility of 20 per cent is a parameter in the one variable model and a state variable in the two variable one, and the same symbol appears in both, so nothing in the written statement of either model announces which role it is playing.
A model that carries the volatility as a parameter while its user believes it is a state will be adjusted by hand every time it disagrees with something. Each adjustment looks like a small correction rather than like evidence that a variable is missing. The cost is a model kept alive by repeated manual intervention, where the pattern of the interventions is the missing variable and nobody reads it that way.
The intervention is invisible to every check the model runs on itself, and that invisibility is what makes the mistake expensive rather than merely untidy. The arithmetic inside is correct. The parameters are within their stated ranges. Because a model has no way of noticing what is done to it from outside, nothing anywhere reports that a fixed number has been re-set six times in six weeks. The only place the missing variable is recorded is the adjustment log, usually the one artefact nobody reviews.
A model needs a manual adjustment most weeks, and each one is small. What does that pattern suggest?
How does somebody checking a model rather than building one use this?
Three questions, in order, and none of them needs any access to the workings.
- Ask for the state vector, and count it
The quantities a model carries forward, once counted, give the width of the model, and that is the single most informative thing one question can reveal about a model. It says what the model can possibly express, and it says roughly what solving it costs.
A model whose builder cannot list its state variables has not been specified, whatever else has been written down.
- Ask which quantities are parameters, and whether any of them gets re-set
A parameter is a number held fixed. Ask when each one was last changed and why. A parameter that gets re-set on a schedule is a state variable wearing the wrong label, and this question finds it in a sentence.
The tell is a calendar. Anything re-set weekly is moving, whatever the model calls it.
- Ask what is an output rather than a state
Anything the model carries that could be recomputed from the rest is redundant width. It costs the same as a real state variable and buys nothing, and on a grid it costs a factor of one hundred.
Redundant width is expensive in exactly the same way as necessary width.
One more thing is worth knowing when these models are met in practice. Some state variables cannot be read off anything directly. The level of the process is visible. The variance is not: a variance is a property of a distribution rather than a reading that can be taken. A state variable that has to be inferred rather than observed is called a latent stateA state variable that cannot be read directly and has to be inferred from the things that can be observed., and the whole business of inferring one is covered separately. Whether a quantity can be observed has nothing to do with whether it belongs in the state. The test asks whether the future depends on it, not whether anybody can see it. A model that leaves out a variable because it is hard to observe has not simplified anything; it has just moved the difficulty into the adjustment log.
The everyday version is a weighing scale that reports the same number whether the object on it is settling or already still. The reading alone does not say which, and if the quantity of interest is the number in three seconds, then whether the object is still moving is a quantity that is needed even though the scale never displays it. Hard to observe, and still part of the state.
References
| Source | Document | Where |
|---|---|---|
| arXiv Quantitative Finance | Preprint repository for state space formulations and grid methods | arxiv.org |
| Social Science Research Network | Working paper repository for the same material | ssrn.com |
| Bellman | Dynamic Programming, where the multiplying cost of extra variables is named | Princeton University Press, 1957 |
| Heston, 1993 | The model in which the variance is carried as a second state variable | Review of Financial Studies, volume 6 |
| Vasicek, 1977 | The model in which the short rate is carried as a state variable | Journal of Financial Economics, volume 5 |
| Hull, Shreve and Wilmott | Standard book-length treatments of derivatives and stochastic calculus | Pearson, Springer and Wiley |
The standard process, its parameters, the two wider models and the adjustment log are invented.
Educational material. Not advice on any investment, tax, budget or market position.
