Random Vectors: Several Uncertain Quantities at Once
A random vector is several uncertain quantities read off one draw and described by a single joint rule, not by separate rules placed side by side. The joint rule carries what the separate ones cannot, namely how the quantities move together. Two random vectors can match on every average and every spread and still disagree about how often both go badly at once.
Once the outcome set is fixed, a random vectorSeveral uncertain quantities read off the same single outcome. is nothing more exotic than several functions of the same outcome. The shared outcome is the whole idea, and it is why the joint behaviour is settled before anybody measures anything. The components are not separate experiments that happen to be reported together. They are different readings taken off one draw.
What makes several uncertain quantities one vector rather than a list?
Consider something concrete. One day happens. A thermometer in the yard reads its temperature, a rain gauge beside it reads its rainfall, and a clock records the hour the first shower arrived. Three numbers. One day. The day that was hot is the same day that was dry, and there is only one day, so the three numbers cannot be shuffled independently. Whatever relationship holds between heat and rain is already sitting inside the single thing that happened.
The formal picture is exactly that. A random variable is a function on the outcome set: hand it an outcome and it returns a number. A random vector is several such functions defined on the same outcome set. Hand it one outcome and it returns several numbers at once, in a fixed order. Nothing else is being claimed. There is no extra machinery.
A list of separate rules is missing something, and the gap shows up immediately. Take three separate rules, one for each quantity, on three unrelated outcome sets. Three separate rules say what each quantity does on its own and say nothing about how the quantities behave together. Saying anything about the three at once requires an assumption added from outside. The components of a random vector already share the draw, so no such addition is ever needed.
The standard process is the case this sequence runs on throughout, and it makes the point cleanly. One path is drawn. The value at three months, the value at six months, the value at the horizon, the highest point reached along the way: every one of those is a function of that single path. The four numbers are not four experiments. Each one is a reading taken off a single path.
| \(\omega\) | one outcome, a single complete description of one way things went |
| \(X_i\) | the \(i\)th component, itself a function of the outcome |
| \(d\) | how many quantities are being read off, a whole number of at least two |
What makes several random variables a random vector rather than a list of separate quantities?
What does the joint rule carry that the separate rules cannot?
The primary object attached to a random vector is its joint distributionThe rule giving the probability of every combination of values at once.: the rule that gives a probability to every combination of values at once. Not to each component in turn. To the combination. Everything else, including each component considered on its own, is recovered from it.
Recovering a component on its own is a matter of letting the other components take any value at all. The rule that comes back is that component's marginal distributionWhat is left of the joint rule when the other components are ignored.. So the direction of travel is fixed and it only runs one way. The joint rule always determines the marginals, and the marginals never determine the joint rule. The asymmetry runs one way only, and every difficulty in this subject follows from it.
| \(F\) | the joint distribution function of the whole vector |
| \(F_1\) | the distribution function of the first component alone |
| \(\mathbb{P}\) | the physical measure P, the rule assigning probabilities on this outcome set |
| \(x_i\) | a threshold for the \(i\)th component, an ordinary number |
Here is the picture worth carrying. Think of the joint rule as a cloud of dust hanging in the middle of a room, denser where the combination is more likely. Light it from one side and it throws a shadow on one wall. Light it from the other side and it throws a shadow on the adjacent wall. The two shadows are the two separate rules. They are real, they are useful, and they are flat.
Then the awkward question. Given the two shadows, can the cloud be rebuilt? It cannot. A shadow discards the depth in the direction of the light, so a great many differently shaped clouds throw exactly the same pair of shadows. The failure to rebuild is not a defect of the analogy. It is precisely what is lost when a joint rule is replaced by its components.
What does a covariance matrix actually say?
Since the whole joint rule is a heavy object, practice reaches for a summary of it. The summary almost always chosen is the covarianceThe average product of two quantities departures from their own averages. between each pair of components: the average of the product of their departures from their own averages. If the two tend to sit above their averages together, the products are mostly positive and the covariance is positive. If one tends to be high when the other is low, the products are mostly negative.
Covariance carries the units of both quantities multiplied together, and that makes it awkward to read. Dividing by the two standard deviations strips the units out and gives the correlationCovariance rescaled by the two standard deviations, so it sits between minus one and one., a pure number that cannot leave the range from minus one to one. The bound is not a convention. The bound falls straight out of the fact that no combination of the quantities can have a negative variance, and that fact returns below as the one condition a covariance matrix must satisfy.
Collecting every pairwise covariance into a square table gives the covariance matrixThe table of every pairwise covariance, with the variances down its diagonal.. The covariance of a quantity with itself is its variance, so the diagonal holds each component's own variance. Swapping the two quantities in the product changes nothing, so the matrix is symmetric. For a vector of ten components the covariance matrix holds ten variances and forty five distinct covariances, and that is the entire summary the working model usually carries.
| \(\Sigma\) | the covariance matrix, square, symmetric, with \(d\) rows and \(d\) columns |
| \(\mathbb{E}\) | expectation taken under the physical measure P |
| \(\Sigma_{ij}\) | the entry in row \(i\), column \(j\), the covariance of those two components |
| \(\Sigma_{ii}\) | a diagonal entry, holding the variance of that component alone |
The matrix says something genuinely useful, and it is worth stating precisely. The matrix fixes every average, every variance, and the average co-movement of every pair. From those alone the average and the variance of any weighted sum of the components can already be computed, which is more than it sounds. It is enough to answer a large class of questions completely.
Two random vectors have the same averages, the same variances and the same correlation between their components. Do they have the same probability that both components fall a long way at once?
What does the covariance matrix leave completely open?
The matrix leaves the shape open, and the tail is part of the shape. Correlation is one number standing in for an entire two dimensional rule. Correlation reports how the two quantities move together on the average of all outcomes, weighting the quiet middle of the picture just as heavily as the far corners. Behaviour in one particular far corner is simply not a question that a single averaged number was ever built to answer.
Here are two constructed joint rules that make the gap concrete. Both have components with average zero and variance one. Both have correlation exactly one half. Every entry of their covariance matrices agrees to the last decimal place. Shape A is the ordinary bell shaped joint rule at that correlation. Shape B is built differently: half the time the two components are the same single draw, and the other half of the time they are two unrelated draws. The mixture reproduces the same averages, the same variances and the same correlation of one half exactly, by construction.
| The far corner, both components at or below | Shape A, the bell shape | Shape B, the half shared mixture | Shape B divided by shape A |
|---|---|---|---|
| two standard deviations | 0.004053 | 0.011634 | 2.87 times |
| three standard deviations | 0.000082 | 0.000676 | 8.25 times |
| Correlation, in both | 0.500000 | 0.500000 | identical |
Identical covariance matrices, and the chance that both quantities land three standard deviations down differs by a factor of eight. The direction in which the gap runs further out matters too. At two standard deviations the two shapes are within a factor of three of each other. At three they are more than eight apart. The disagreement is not a fixed offset that could be carried as a safety margin. The gap widens the deeper into the corner the question reaches, and the far corner is the one region where a wrong answer tends to matter.
What happens to a weighted sum of the components?
Almost everything anyone does with a random vector eventually collapses it to one number by taking a weighted sum of the components. The weighted sum is why the covariance matrix is worth carrying at all. Its two moments respond completely differently, and the difference is worth knowing by heart.
The average of a weighted sum is the weighted sum of the averages. Always. It does not matter what the correlations are, whether the components are related at all, or what shape the joint rule has. Averages simply add. The variance is another matter entirely: it picks up every pairwise covariance, each one counted twice with the product of its two weights attached.
| \(a\) | the column of weights, one per component, chosen by the reader |
| \(a^{\top}X\) | the weighted sum of the components, a single random quantity |
| \(\Sigma\) | the covariance matrix of the vector, as defined above |
| \(a^{\top}\Sigma a\) | the double sum written out on the right, the variance of that sum |
Consider the two readings of the standard process used throughout this guide: the logarithm of the value at six months and the logarithm of the value at one year. Their variances are 0.02 and 0.04 and their covariance is 0.02, figures the worked instance below derives rather than asserts. Adding the two readings together with weights of one each: the average of the sum is 9.300340 and it stays at 9.300340 whatever number sits in the off diagonal. The variance is 0.06 plus twice the covariance, so it lands at 0.10 exactly, and the standard deviation at 0.316228.
Now sweep the correlation through its whole permitted range while holding both variances fixed. At minus one the variance of the sum drops to 0.003431 and its standard deviation to 0.058579. At zero it is 0.06 and 0.244949. At plus one it is 0.116569 and 0.341421. The standard deviation of the sum moves by a factor of nearly six across that sweep while the average of the sum does not move at all. The split between the two moments is why the second term is the one that ever gets argued about.
One weighted sum on this vector is worth pausing over, because it is the whole reason this subject is built on increments. Take weights of minus one and plus one, which turns the pair into the change in the logarithm between six months and the horizon. Its variance is 0.02 plus 0.04 less twice 0.02, which is 0.02: exactly the variance rate multiplied by the half year the change covers. Its covariance with the earlier reading is 0.02 less 0.02, which is zero. The right weighted sum turns two heavily related readings into a quantity that carries no relationship with the earlier one at all, and that is the independence of increments falling out of the arithmetic.
The correlation between two components rises while both variances stay put. Which moves: the average of their sum, the variance of their sum, or both?
What does assuming a multivariate normal shape add?
The gap traced so far has a standard fix, and the fix is an assumption rather than a discovery. Assume the vector has the multivariate normalThe joint rule in which every linear combination of the components is itself normal. shape and the averages together with the covariance matrix stop being a summary. The averages and the matrix become the entire description. Nothing is left over to choose.
The cleanest definition of that shape is the one that names weighted sums directly, and it is the definition worth remembering because it is the one that gets used. A vector has the multivariate normal shape when every weighted sum of its components is an ordinary one dimensional normal quantity. Not some of them. Every single one, for every choice of weights.
| \(\mathcal{N}_d\) | the joint normal law in \(d\) dimensions; the script letter distinguishes it from N, the symbol reserved for the standard normal distribution function throughout this subject |
| \(\mathbf{m}\) | the column of averages, one per component; written m rather than mu because mu is reserved here for the drift of the standard process |
| \(\Sigma\) | the covariance matrix of the vector |
| \(a\) | any column of weights whatsoever |
The assumption buys something and costs something, and both are worth stating precisely. The assumption buys closure: the far corner, the near middle and every region in between are now pinned down by numbers already in hand. Because the covariance matrix is identical under this shape and under every other shape with the same summary, the assumption costs a claim about the world that no amount of covariance arithmetic can check. Shape A above was this assumption. Shape B was not. Their covariance matrices could not tell them apart.
One warning that catches people. Two components can each be normal on their own without the pair being jointly normal. The shadows can both be bell shaped while the cloud is not. Shape B is exactly that case: both of its components are perfectly ordinary normal quantities, and the pair together is not jointly normal at all. Normality of every component separately is a strictly weaker statement than joint normality, and only the joint version closes the shape.
What does assuming a multivariate normal shape buy that the covariance matrix alone does not?
When is a table of correlations not a covariance matrix at all?
Not every symmetric table of plausible looking numbers can be the covariance matrix of anything. There is one condition, it is not a matter of inspection, and it fails on tables that look entirely reasonable entry by entry.
The condition falls straight out of the previous section. Every weighted sum of the components has variance equal to the weights read through the matrix, and a variance can never be negative. So if any choice of weights produces a negative number when read through the table, that table describes nothing. The requirement that no choice of weights can do this is what makes a matrix positive semi-definiteThe condition that stops a covariance matrix implying a negative variance..
| \(a\) | any column of weights, including ones nobody would ever choose in practice |
| \(a^{\top}\Sigma a\) | the variance the table implies for that weighted sum |
| \(\ge 0\) | the requirement, holding for every single choice of weights without exception |
Work the standard counterexample and the failure becomes vivid. Three quantities, each with variance one, each correlated with the other two at minus 0.9. Every individual entry is perfectly legal: two things certainly can move strongly against each other. Now weight all three equally and read the table. The variance of the sum is three, from the diagonal, plus six times minus 0.9 from the six off diagonal entries. The total is three less 5.4, or minus 2.4.
A negative variance is not a small modelling blemish; it is a proof that no such three quantities exist. And the intuition behind the failure is worth holding, because it generalises. If the first quantity opposes the second and the second opposes the third, then the first and the third are being pushed toward each other, so they cannot also oppose each other strongly. For three quantities all sharing one common correlation, the floor is minus one half. Anything below that is impossible however reasonable each entry looks.
Three quantities are each correlated with the other two at minus 0.9. Can that be a real covariance matrix?
Before the worked instance. On one path of the standard process, the correlation between the reading at three months and the reading at one year: closer to 0.25 or to 0.5?
What does the standard process look like as a random vector?
The standard process starts at Rs 100/- exactly, it drifts at 8 per cent a year under the physical measure P, its volatility is 20 per cent a year and the horizon is one year. Every figure below is a computed consequence of those four numbers rather than an observation of anything.
Reading it at six months and again at one year gives a random vector with two components. Not two readings that happen to be filed together. Two functions of the one path that was drawn. The logarithm of each reading is normal, and here is the structure that matters: the variance of the logarithm accumulates at the variance rate of 0.04 a year, so at six months it is 0.02 and at one year it is 0.04.
The covariance is worth deriving rather than accepting. Write the later reading as the earlier reading plus the change between them. The change over the second half year is unrelated to everything that happened in the first half, so it contributes nothing to the covariance. The variance of the earlier reading alone is what is left. The covariance of the two readings is the variance rate multiplied by the shorter of the two times, and it is the shorter time because that is the only stretch of variance the two readings actually share.
| \(S_t\) | the standard process at time \(t\), the single traded quantity this subject works with throughout |
| \(\sigma\) | volatility, locked at 0.20 a year, so the variance rate is 0.04 |
| \(s,t\) | the two observation times in years, with \(s\) the earlier of the two |
| \(\min(s,t)\) | the smaller of the two times, the stretch of variance the two readings share |
Put the six month and one year readings into a table and check every entry against that rule. The correlation comes out as 0.02 divided by the square root of 0.02 times 0.04, which is 0.02 divided by 0.028284, which is 0.707107. The figure 0.707107 is the square root of one half, and one half is the fraction of the year the earlier reading has covered.
| The random vector | Reading at six months | Reading at one year |
|---|---|---|
| Average of the logarithm | 4.635170 | 4.665170 |
| Variance of the logarithm | 0.020000 | 0.040000 |
| Covariance with the other reading | 0.020000 | 0.020000 |
| Correlation with the other reading | 0.707107 | 0.707107 |
| Where the locked path stood | Rs 111.08/- | Rs 106.18/- |
The last row is the locked path, the twelve step path this subject publishes once and draws everywhere, so that a worked example can be checked rather than regenerated. The path stood at Rs 111.08/- at six months and Rs 106.18/- at the horizon, one above the other. The pair is one draw from the random vector described in the rows above, not the vector itself, and no property of the vector can be read off it.
Now sweep the earlier time across the year and the pattern is worth seeing whole. At three months the correlation is 0.5 exactly. At six months 0.707107. At nine months 0.866025. The correlation between the reading at time t and the reading at the one year horizon is the square root of t, so a quarter of the elapsed time already buys half the correlation. The curve climbs steeply and then flattens. The shape is the signature of the square root and the reason nobody guesses these numbers correctly from the times alone.
Take the reading at three months and the reading at nine months. What is the covariance of their logarithms?
The reading at six months and the reading at one year. Before the control below is moved: is their correlation above or below 0.5?
Move the earlier reading through the year
The horizon stays at one year and the volatility stays at 20 per cent a year, so the variance rate stays at 0.04 and the variance of the horizon reading stays at 0.04. Only the earlier time moves. The curve on the left is the square root relationship, with the marker at the month selected. The cloud on the right is a deterministic grid of the pair, drawn at midpoints of five equal probability slices in each direction and pushed through the correlation, so it reproduces identically on every reload. At six months the reading is 0.707107, the figure in the table above; three months gives 0.500000 and nine months gives 0.866025.
How does someone reviewing a model actually use this?
What gets checked, and in what order
The person who has to sign off a model rarely gets handed a joint rule. The reviewer gets handed a covariance matrix, and the first three things worth doing with it are all cheap. First, run the one test: pick some awkward weightings and confirm none of them produces a negative variance. A table assembled from separately sourced pieces frequently fails this, and a table that fails it will produce nonsense somewhere downstream without ever raising an error.
Second, ask where the joint shape was written down, since the matrix does not contain it. If the answer is that nobody wrote it down, then the shape is whatever the code happens to implement, and that is nearly always the bell shape because it is the easy one to generate from a matrix. The bell shape is then a decision, and it should be visible as one.
Third, check the structure against the model's own logic rather than against a data file. If a model reads one process at several times, its correlations must follow the square root of the ratio of times: 0.500000 at a quarter of the way, 0.707107 at half, 0.866025 at three quarters. An entry that does not sit on that curve is either a different model or a mistake, and the arithmetic says which without anybody needing a single observation.
The everyday version is a household with two electricity meters, one for each floor, read on the same day. The two readings are a random vector, because one day of weather and one household routine produced both. Averaging each meter over a year gives what each floor costs. Neither average says how often both floors run hot on the same evening, which is the only question that decides whether the main fuse trips. The summary answers the running cost and stays silent on the failure that actually bites.
Where does accepting the matrix as the description cost money?
The mistake is quiet, and that is what makes it expensive. Nobody writes down a false statement. A covariance matrix arrives, it is treated as the description of how the quantities move together, and every subsequent step of arithmetic is performed correctly. The stated probability of both quantities going badly at once then comes out wrong by a multiple, and nothing anywhere in the calculation is available to flag it, because the error was committed before the first line of arithmetic ran.
The error that gets made, and what it costs
Treating a covariance matrix as though it fixed the joint distribution. Two random vectors can share every average, every variance and every correlation and still disagree by a wide margin about the probability that both components fall together, because correlation is one number summarising an entire joint shape and the far corners are not in it.
On the two constructed shapes above, the stated chance that both quantities land three standard deviations down is 0.000082 under one and 0.000676 under the other. Same covariance matrix, same averages, same correlation of one half, a factor of 8.25 between the answers. The cost is a stated probability of a joint bad outcome wrong by a multiple rather than by a rounding, produced by a calculation in which every step is correct. The mistake was made at the moment the matrix was accepted as the description.
References
| Source | Document | Where |
|---|---|---|
| arXiv, Quantitative Finance | Preprints on multivariate models and dependence in finance | arxiv.org |
| Social Science Research Network | Working papers on covariance structure and joint distributions | ssrn.com |
| Shreve, Hull and Wilmott | Standard texts covering random vectors, the covariance matrix and the multivariate normal shape | Published books |
The standard process and the locked path are invented.
Educational material. Not advice on any investment, tax, budget or market position.
