What a Model Holds: Features In, Parameters Learned, Hyperparameters Chosen
A model holds three kinds of thing. Features are the columns it reads. Parameters are the numbers it learns from the record, two on the straight line fitted to the Nakshatra unit and the Vasant unit. Hyperparameters are settings a person chooses before fitting starts and the fitting never touches, and a reported result usually shows the parameters and hides the settings.
Three things are already settled. A fitted line and the numbers sitting inside it were worked out in earlier material, so what it means to draw a rule through a scatter of readings is already established. The idea of charging a fitting routine extra for large numbers, and the idea of a cut off at which a call gets made one way rather than the other, have both been met already. And the ten paired months of the Nakshatra unit and the Vasant unit, both of them invented for teaching, are carried in whole from where they were built. Three names go to the parts of a model, and each part is counted on that one short record. The next result somebody hands over can then be read for exactly what is missing from it.
What are the three kinds of thing a model holds?
The line between the three is not about importance and not about size: it is entirely about who decided the number. Features are what goes in. Parameters are what the fitting step works out. Hyperparameters are what a person fixes before the fitting step is allowed to run at all. Every one of those three sentences answers the question who, and none of them answers the question how big. A setting somebody typed in four seconds can move a headline further than any number the arithmetic laboured over, and that happens twice below.
The everyday version holds all the way through. Nothing later breaks it. A cook puts a tray in the oven. The cook fixes the oven at a temperature, by hand, before anything goes in. The tray then takes as long as it takes, and the cook watches and writes the timing on the recipe card. The recipe card is the thing to notice. Somebody worked the timing out, and the card carries it. Nobody worked the temperature out; somebody just set it. The card very often does not carry the temperature at all. And a person handed that card in a different kitchen, with a different oven, will get a different tray and will not know why.
The third kind of thing is the whole of the problem: it is easy to leave out of a report precisely because nothing computed it, so nothing printed it. A fitting routine is a machine for producing numbers and then reporting them. The routine has nothing to say about the numbers it was handed on the way in. From its point of view the incoming numbers were never in question, but simply the terms on which it was asked to work.
What actually separates a parameter from a hyperparameter?
What is a feature, and which four does this record offer?
A feature is a column the fitting step is handed to read, and that is the entire definition. Not a column that is interesting, not a column that turned out to be useful, not a column somebody believes in. A column that was handed over. If it went in, it is a feature; if it was left in the file and never passed across, it is not, however much it might have helped.
The ten-month record offers four features, and all four were built in earlier material. The first is the Nakshatra unit's monthly change, the plain reading. The second is the Chandana unit, the same reading nudged by five hundredths of a percentage point up in four months, down in four more, and left exactly alone in two. The third is the Sharad marker, simply on in five of the ten months and off in the other five. The fourth is the marker multiplied by the Nakshatra unit's change, an interaction columnA column made by multiplying two other columns together, so that what one of them says is allowed to depend on what the other says.: it reads zero in every month the marker is off, and repeats the Nakshatra reading in every month the marker is on.
The fourth column is built entirely out of two columns that were already there, and it is still its own feature. A feature is whatever column is handed to the fitting step, not whatever column carries new information. Think of a form at a lending desk. The form asks for monthly income in one box and rent paid in another. Somebody designing the form can add a third box reading income less rent, and that third box holds nothing the first two did not already hold between them. The third box is still a box: still filled in, still read, still counted, and the clerk reading the form treats it as a question in its own right. The fitting step is exactly that literal.
| Month | The Nakshatra unit | The Chandana unit | The Sharad marker | Marker times Nakshatra | The Vasant unit |
|---|---|---|---|---|---|
| 1 | 1.00 | 1.00 | off | 0.00 | 3.00 |
| 2 | 6.00 | 5.95 | off | 0.00 | 18.50 |
| 3 | minus 4.00 | minus 3.95 | on | minus 4.00 | 2.50 |
| 4 | 11.00 | 11.05 | off | 0.00 | 17.00 |
| 5 | 1.00 | 0.95 | on | 1.00 | minus 3.00 |
| 6 | minus 9.00 | minus 9.00 | off | 0.00 | minus 13.00 |
| 7 | 6.00 | 5.95 | on | 6.00 | 6.50 |
| 8 | 1.00 | 1.05 | on | 1.00 | minus 1.00 |
| 9 | minus 4.00 | minus 4.05 | off | 0.00 | minus 7.50 |
| 10 | 1.00 | 1.05 | on | 1.00 | minus 3.00 |
The last column is not a feature, and it is worth saying why. The Vasant unit is the column being explained rather than a column doing the explaining, so it is never handed across as an input. Which of two columns should sit on which side of that arrangement, and why swapping them is not a symmetric move, is covered separately. The count is what matters: four columns available to go in, one column waiting to be explained.
The Sharad marker multiplied by the Nakshatra unit's change is built from two columns that are already features. Is it a feature in its own right?
What is a parameter, and how many does each model here learn?
A parameter is a number the fitting step chooses for itself, by hunting for the values that make the misses small, and nobody types it in. On the straight line through these ten months there are exactly two of them. One is the interceptThe part of a fitted rule that does not depend on any column at all: what the rule says when every reading handed to it is zero., coming out at 0.5000, and one is the coefficientThe number a fitted rule multiplies a column by. Two boards for every crate is a coefficient of two. on the Nakshatra unit, coming out at 1.5000. Neither was chosen. Both were found. Reading a coefficient off a fitted line is covered separately. A coefficient is one kind of parameter and always has been.
Now hand the fitting step one more column and count again. Add the Sharad marker as a plain input and it learns three numbers instead of two: an intercept of 2.1000, a coefficient of 1.5000 on the Nakshatra unit, and a coefficient of minus 3.2000 on the marker. Look at the middle one. The coefficient on the Nakshatra unit has not shifted a hair, and there is a reason: the marker and the Nakshatra unit line up with each other at exactly zero on this record, so adding one of them leaves the other's number untouched. Such a clean split is a property of this made-up record rather than a rule about models, and it will not usually happen.
Add the interaction column as well and the count goes to four. The four numbers are 1.8800 on the intercept, 1.7200 on the Nakshatra unit, minus 1.8800 on the marker and minus 1.3200 on the interaction. Every one of those four moved when the fourth column arrived, including the two that were 0.5000 and 1.5000 a moment ago. Push further and let the rule bend: a curve of degreeHow curved a rule is allowed to be. Degree one is a straight line, degree two bends once, degree three twice, and so on. four learns five numbers, and on these ten months it runs the leftover misses all the way down to the floor those readings set. The floor itself, and why no rule reading this one column can get under it, is settled separately.
| What was handed to the fitting step | Numbers learned | Months on the record | Months per learned number |
|---|---|---|---|
| The Nakshatra unit alone, as a straight line | 2 | 10 | 5.00 |
| The Nakshatra unit and the Sharad marker | 3 | 10 | 3.33 |
| Both of those and the interaction column | 4 | 10 | 2.50 |
| The Nakshatra unit bent into a curve of degree four | 5 | 10 | 2.00 |
The count runs 2, 3, 4 and 5 on one record of ten months, and the record never grew by a single month: every step up that ladder is a choice somebody made. That is the sentence to carry away. The record is a fixed thing, sitting there with ten rows in it, and a person walked up to it four times and asked for a different number of learned numbers each time. Nothing in the ten months requested any of that. Whether pushing the count up is a good idea, and what happens to a short record when the count climbs, is covered separately and deliberately left alone here.
A model carrying the Sharad marker and its interaction learns four numbers from ten months. What question should that raise?
What is a hyperparameter, and who chooses it?
A hyperparameter is one of the terms the fitting was asked to work under. A person fixes it before the fitting step runs, and the fitting step never reads it off the record. Four of them turn up in this arithmetic, and listing them plainly makes the shape easy to recognise. There is the degree of the curve: how much bending the rule is allowed. There is the size of the penaltyAn extra charge added to the fitting arithmetic for large learned numbers, so the fitting settles for smaller ones. Met earlier in these notes., which is how hard large learned numbers are discouraged. There is the number of foldsOne of the equal parts a record is cut into so that each part can take a turn being held back while the rest is used for fitting. a record gets cut into. And there is the threshold at which a call is made one way rather than the other.
Watch what those four have in common. Not one of them can be discovered by staring harder at the ten months. The record has nothing to say about whether the rule should bend once or four times. The record has nothing to say about how much a large number should cost. The ten months have no view on being cut into five parts rather than ten, and none at all on where a cut off should sit. Every one of the four has to be supplied from outside, and every one of the four changes the answer that comes back. That combination, no opinion from the record plus a real effect on the result, is exactly what makes a setting worth reporting and exactly what makes it easy to forget.
The everyday shape is a shopkeeper packing rice. Before any rice is weighed, somebody decides the packet is a kilogram. The packet size is not read off the sack but from a person, and the packet size determines how many packets the sack yields and what each one costs. No amount of weighing the sack will say what size the packet should be. Weighing will only ever say, once the size is decided, how many packets the sack gives.
| The setting | What it controls | Read off the record? | What moves when it moves |
|---|---|---|---|
| The degree of the curve | How much the rule is allowed to bend | no | The count of learned numbers, and every one of their values |
| The size of the penalty | How hard large learned numbers are discouraged | no | The learned numbers themselves, sometimes enormously |
| The number of folds | How many parts the record is cut into for testing | no | The figure the testing reports back |
| The threshold a call is made at | Where a score stops meaning down and starts meaning up | no | Every call, and the score the calling is judged by |
The third column answers no four times over, and the repetition is the point of the table. How any one of these ought to be chosen, honestly rather than conveniently, is a whole subject of its own and is covered separately. Here they are named and counted, and none of them is a parameter.
How is a parameter told apart from a setting in practice?
Names are cheap, so here is a test that can be run on a real number, and this record allows it twice over with two very different settings.
Test one, and it is the loudest result here: with the calling threshold moved from 0.50 to 0.75, the accuracyThe share of months a rule called correctly, counting an up called up and a down called down alike. goes from 60.00 per cent to 80.00 per cent with the model completely untouched. The rule scoring these months is a straight line on the up or down label, reading 0.4500 plus 0.0500 times the Nakshatra change, and it was fitted once. Nothing about it is recomputed. The threshold is applied afterwards, to scores that already exist, and that is precisely why moving it refits nothing. Two numbers that nobody fitted sit between the record and the headline, and one of them just moved the headline by twenty percentage points.
The arithmetic here is full of figures that look alike and mean nothing to each other, so one caution comes before the sweep. A threshold of 0.50 and an accuracy of 50.00 per cent are counted in completely different things: one is a cut off on a score that runs from zero to one, the other is a share of ten months called correctly. The two can be written with the same digits, but that is a coincidence of where the decimal point sits, and the resemblance carries no meaning whatever.
| Calling threshold | Months called up | Called up and went up | Called down and went down | Accuracy |
|---|---|---|---|---|
| 0.00 | 10 | 5 | 0 | 50.00 per cent |
| 0.25 | 9 | 5 | 1 | 60.00 per cent |
| 0.50 | 7 | 4 | 2 | 60.00 per cent |
| 0.75 | 3 | 3 | 5 | 80.00 per cent |
| 1.00 | 1 | 1 | 5 | 60.00 per cent |
Every row of that table came out of the same two learned numbers. Not one refit happened anywhere in it. Two other ways of scoring the same calls pull apart differently as the threshold climbs, and both are covered separately.
Test two works from the opposite direction. Raise the penalty from 0.00 to 1.00, and the number sitting on the Nakshatra unit travels from minus 1.0000 all the way to 0.7314. Meanwhile R squaredA single figure saying how much of the movement in the column being explained a fitted rule accounts for. Higher means less left over. moves from 0.756019 to 0.755950, a shift of 0.000069. This one is a fit on two columns at once, the Nakshatra unit and the Chandana unit, which as noted above are near enough the same column twice. With no penalty at all the fitting step splits the work between them as minus 1.0000 and 2.5000, a gap of 3.5000. With a penalty of one it splits it as 0.7314 and 0.7661, a gap of 0.0347. The learned numbers have travelled a very long way and the quality of the fit has barely twitched.
So: the penalty was typed by a person before the fitting ran, and it changed what the fitting returned. The threshold was typed by a person after the fitting ran, and it changed what the result was called. Both are settings. Here is the test in three lines. If the fitting step produced it, it is a parameter. If changing it means fitting again, it is a setting. If changing it changes what the refit gives back, it is a setting, and the fact that it moved a learned number is what proves it sits above the fitting step rather than inside it.
The calling threshold is about to move from 0.50 to 0.75. What happens to the two learned numbers?
Slide the setting nobody fitted, and watch the numbers somebody did fit refuse to move.
The scoring rule is fixed for the whole of this panel at 0.4500 plus 0.0500 times the Nakshatra change, and it is never fitted again. Only the calling threshold moves. Each of the ten months sits on the score line at its own score, filled if the month actually went up and hollow if it went down, and a month is called up when its marker sits at or to the right of the red line. The panel opens at a threshold of 0.50 and reads an accuracy of 60.00 per cent. Push it to 0.75 and the same untouched model reads 80.00 per cent.
Moving the penalty from 0.00 to 1.00 changes the Nakshatra coefficient from minus 1.0000 to 0.7314. Is the penalty a parameter or a setting?
What does a printed result leave out?
A fitting routine prints the parameters because it computed them, and it prints nothing about the settings because somebody typed those before it was called. There is no conspiracy in this and usually no carelessness either. Ask a machine to report on its work and it reports on its work. The degree it was told to use was not its work. The penalty it was handed was not its work. The folds it was given and the threshold applied to its output afterwards were not its work. So a printout carrying two coefficients and a goodness figure is a complete and honest account of the fitting step, and an incomplete account of the model.
Put it beside the cook and the recipe card again and the shape is identical. The card is a complete account of what happened once the tray went in. The card leaves out the one decision made before the tray went in, and that decision is the one most likely to be different in somebody else's kitchen. A result quoting only parameters describes the half of the model the arithmetic chose and stays silent about the half a person chose, and the silent half is the half that will not travel.
Why does a printed model summary almost never show the settings it was run under?
What should be asked of a fitted model handed over by somebody else?
Four questions, and they take under a minute between them. Which columns went in? How many numbers were learned, and from how many records? Which settings were fixed before fitting, and what were they? And was any setting chosen after somebody had looked at the result it was going to be judged on?
The four questions are not an academic exercise. A person at a lending desk reading a scoring model, an analyst handed a fitted rule by a colleague, an investor sent a sheet of results, a household comparing two quotes worked out by two different calculators: every one of them is in the same position, holding an output and not the terms it was produced under. The first three questions describe the model. The fourth describes how the model came to be reported, and it is the one that does the real work.
A setting chosen after seeing the figure it was going to be judged on is not a setting any more; it is a parameter that somebody fitted by hand, without saying so and without counting it. On this record the room that leaves is visible exactly. Somebody who tries the threshold at five places and reports the best one has quietly fitted a number and then reported the accuracy as though nothing had been fitted. The two learned numbers are still 0.5000 and 1.5000 on the size model, or 0.4500 and 0.0500 on the label model, and every one of them is still honest. The headline is not.
The second question belongs beside the third. Together they show how much authorship is packed into how little evidence. Four learned numbers fitted on ten months is two and a half rows for every number the arithmetic had to choose. Ten months is a short record carrying a lot of decisions. Knowing the ratio is not the same as knowing what to do about it, and what to do about it is covered separately.
Which of the four questions turns a reported figure into an unreported choice?
An audit that checked every number nobody had chosen
A fitted model arrives with a short note attached. Two coefficients, a goodness figure, and a headline accuracy. The person receiving it does the responsible thing and checks the arithmetic. The coefficients are recomputed from the ten months and they land exactly where the note says. The goodness figure is recomputed and it lands exactly where the note says. Everything reconciles. The note is signed off.
Nothing was verified. The degree was chosen by somebody. The penalty was chosen by somebody. The threshold was chosen by somebody, and on this record moving that last one alone takes the accuracy from 60.00 per cent to 80.00 per cent without a single learned number changing by so much as a decimal place. The coefficients have nothing to do with the setting at all, so every coefficient in the note was correct at every one of those settings. The audit examined the half of the work the arithmetic did and never went near the half a person did, and the half a person did was where the whole of the twenty point swing lived.
The habit that fixes it costs one line. Before any number is checked, the settings are requested, in writing, with their values. Then comes the fourth question: was any of them chosen after somebody had already seen the figure being reported. A model whose settings arrive without argument is a model that can then be checked. A model whose settings arrive after some hesitation about which run this was, or which threshold got used in the end, has said something the arithmetic never could.
The coefficients of a reported model are audited and every one of them checks out. What has actually been verified?
Where did these figures come from?
Every number above was counted or recomputed from ten made-up months printed in full above, so the whole of it can be redone on paper. Naming what a fitted model holds needs no outside authority at all, and the table below has exactly one row because of that.
| Document | Site | Date consulted |
|---|---|---|
| None. The four columns, the learned numbers and the four settings were all made up for teaching and recounted here from the ten months printed above | None | Not applicable, because no maintained record was opened |
The Nakshatra unit, the Vasant unit, the Chandana unit and the Sharad marker are invented.
Educational material. Not advice on any investment, tax, budget or market position.
