Demonstration 1 of 4
Fit small runs, forecast a run that is not there yet
If a curve of the form in Equation (3.1) is fitted to small runs, how does the miss on a much larger run depend on how far away it is and on whether the runs share a recipe?
The exponent b is searched over a grid, and at each b the constants a and c are found by least squares; the best b is kept. With exact measurements the true curve is recovered. With a small wiggle the fit still looks good inside the measured range, but the fitted a and c trade off against the exponent b, so the forecast a + c at C = 1 shifts, and for this constructed pattern the shift grows with distance. An unrecorded difference between members is absorbed as if it were scale, which is why the family has to be documented before it is fitted.
Scroll sideways for the whole equation
L is loss, a graded measure of surprise, and C is training compute, scaled so the held-out run has C = 1. The constants a, b and c are fitted: b is negative, so loss falls as compute grows, and c is the level loss levels off at. The constructed true curve has a = 1.5, b = -0.15 and c = 1.0. Five measured runs span two powers of ten of compute; the first control sets how many powers of ten lie between the largest of them and the held-out run. The second control sets what the measured runs contain: exact values on the curve, a small wiggle, or a shift of 0.10 in the two largest runs that were trained differently in a way nobody recorded. The residual is observed minus forecast.
Predict first. With the wiggle of 0.04 (the default), set the distance to 4 powers of ten. Is the miss bigger or smaller than at 1 power of ten?
Choose an example
Scroll sideways for the whole figure
Constructed example: a loss curve of the form in Equation (3.1) with constants defined for this reader. It is a schematic of the shape of the chapter's GPT-4 loss forecast and is not that report's data.
Calculated values
- Fitted a, b, c
- 0.807, -0.197, 2.053
- Largest misfit inside the measured runs
- 0.036
- Forecast at the held-out run
- 2.860
- Observed later (constructed true value)
- 2.500
- Residual, observed minus forecast
- -0.360
At the held-out run C = 1, so C^b = 1 for any exponent b and the observed loss is a x 1 + c = 1.5 x 1 + 1.0 = 2.5. The forecast is the fitted a + c = 0.807 + 2.053 = 2.860. Residual = 2.500 - 2.860 = -0.360. The fit follows the measured runs to within 0.036, yet the forecast is off by 0.360. Across the four distances the miss grows for this wiggle pattern (compare the bars on the right). The forecast is a + c, so the miss comes from the fitted constants: a and c trade off against the exponent b, and here a = 0.807 and c = 2.053 against the true 1.5 and 1.0.
Worked steps
- Five measured runs at compute 10^-4.0 to 10^-2.0; the held-out run is at C = 1, 2 powers of ten beyond the largest.
- Search b from -1.000 to -0.010; solve a and c by least squares at each b and keep the best: a = 0.807, b = -0.197, c = 2.053.
- At C = 1 the power C^b is 1, so the forecast = a + c = 0.807 + 2.053 = 2.860.
- The constructed true value is 1.5 + 1.0 = 2.5.
- Residual = 2.500 - 2.860 = -0.360.
Use the idea
When a team reports that a fitted curve predicted a bigger run, ask how far the held-out run sat from the largest run used in the fit, whether the runs really belong to one documented family, and whether the forecast was recorded first. A small miss across one power of ten and the same miss across four are different achievements.
Where the conclusion applies
The true curve has exactly the form in Equation (3.1), which is assumed here and is not guaranteed for any real family; the book calls the relation measured, not derived. The wiggle and the shift are fixed constructed patterns, not random noise, and the exponent search covers b from -1.000 to -0.010 in steps of 0.001. The fitted a and c are rounded to three decimals before they are added. If the true curve had a different shape, the miss could be larger than shown.
Common wrong turn: A small miss over a short range is the same achievement as over a long one
What this does not settle
The GPT-4 result covers loss and one bounded HumanEval aggregate for one family.
Chapter 3 source: "What this does not settle".
Check your understanding: With exact measurements and a = 1.5, c = 1.0, what loss does the fit forecast at C = 1, and what is the residual at 3 powers of ten?
Chapter 3 source: section "The forecast model". Demonstration C03-D01.
Demonstration 2 of 4
The commitment boundary and the registered interval
Given three development measurements and a stated error bound, what point and interval does the team register, what does each possible result say, and what changes if the forecast is rewritten after the result is seen?
The forecast at x = 4 is a weighted sum of the three development values with weights -2/3, 1/3 and 4/3. Each observation can be off by at most the bound, so the forecast can be off by the sum of the weights in size times the bound, plus one more bound for the new observation. Figure 3.2 shows where the line falls: steps 1 to 6 are committed before the answer exists, steps 7 and 8 can be written only afterwards. A forecast or an interval rewritten after step 7 can no longer lose.
Scroll sideways for the whole equation
x is a declared family coordinate and the contrast is the measured interaction for the member at x. Three development members sit at x = 1, 2, 3. The slope and intercept come from a straight line fitted by least squares. The error bound is set at 0.02: the stipulated largest distance of any observation from one common true line. The residual is observed minus forecast. The first control is the observed contrast at x = 4; the second says whether the forecast stays as registered, is refitted through all four points after the result, or gets an interval written after the result.
Predict first. With the forecast kept as registered, an observed contrast of 0.36 is 0.03 below the forecast 0.39. Is it inside the interval?
Choose an example
Scroll sideways for the whole figure
Constructed example: the chapter's own worked forecast (points 1, 2, 3 with contrasts 0.10, 0.21, 0.29; interval half-width 0.066667 and the held-out result 0.43), with the other observed results and the after-the-fact choices defined for this reader.
Calculated values
- Fitted slope and intercept
- 0.095, 0.01
- Registered point forecast at x = 4
- 0.39
- Registered interval
- 0.323333 to 0.456667
- Observed at x = 4
- 0.43
- Residual against the registered forecast
- 0.04
- Inside the registered interval
- yes
Slope = [(1 - 2) x (0.10 - 0.20) + (3 - 2) x (0.29 - 0.20)] / 2 = 0.19 / 2 = 0.095, so the forecast at x = 4 is 0.01 + 0.095 x 4 = 0.39. The weights (-2/3, 1/3, 4/3) add up in size to 2/3 + 1/3 + 4/3 = 2.3333, so the inherited error is at most 2.3333 x 0.02 = 0.046667. Adding 0.02 for the fourth observation gives 0.046667 + 0.02 = 0.066667, so the interval runs from 0.323333 to 0.456667. Residual = 0.43 - 0.39 = 0.04. The observed value is inside the registered interval: one successful check of this declared forecast. It does not prove the line or the error bound, and one hit does not show they will keep holding. The forecast and interval were fixed at step 6, before the result existed, so the comparison at step 8 is a test.
Worked steps
- Fit a line to the three development points: slope 0.095, intercept 0.01.
- Forecast at x = 4: 0.01 + 0.095 x 4 = 0.39.
- Interval half-width = (2.3333 + 1) x 0.02 = 0.066667, so 0.323333 to 0.456667.
- Register the forecast and interval (step 6). Everything before the dashed line is committed.
- Observed 0.43; residual = 0.43 - 0.39 = 0.04; inside the registered interval.
Use the idea
Write the point forecast and an interval with a stated construction before the held-out member exists. A point alone hides how much uncertainty the team knew it had, and an interval that is declared in advance can be missed, which is what makes a hit mean something.
Where the conclusion applies
The guarantee is conditional: the true relation is a straight line and every observation is within the stipulated bound of it. The three development points do not prove either. This is a worst-case interval, not a claimed 95 percent coverage rate, and one hit does not establish that the assumptions will keep holding. A nominal 95 percent procedure, unlike this bound, is allowed to miss.
Common wrong turn: A refit that matches the result is a better forecast
What this does not settle
Equation (3.2) is a proposal; the cited studies do not report that prospective interaction-contrast test.
Chapter 3 source: "What this does not settle".
Check your understanding: Keep the forecast as registered. The observed contrast is 0.46. What is the residual, and is it inside the interval?
Chapter 3 source: section "A constructed forecast example". Demonstration C03-D02.
Demonstration 3 of 4
Smooth ability, abrupt score
When does a sudden jump in a reported pass-or-fail score reflect a sudden change in the model, and when only the cutoff?
Equation (3.3) turns the graded q into one bit. A small change in q next to the threshold produces a change of one whole point in the reported score. The same flip can accompany a gentle rise in q or a sharp one, so the thresholded score alone cannot say which story it is.
Scroll sideways for the whole equation
m numbers the members of a family in order of scale. q is the probability that one attempt by member m succeeds, a smooth quantity on the left. The threshold (tau) is a cutoff chosen in advance. The verdict on the right is 1 when q is at least tau and 0 otherwise. The four profiles are constructed curves for q.
Predict first. Choose the metric-created jump and compare the left and right panels. By how much does q change at the step where the score jumps from 0 to 1?
Choose an example
Scroll sideways for the whole figure
Constructed example: four success-probability curves and two thresholds defined for this reader to match the chapter's four profiles. No real model is measured.
Calculated values
- Profile
- Metric-created jump
- Threshold
- 0.50
- First member that passes
- m = 5
- Largest one-step change in q
- 0.100
- Largest one-step change in the verdict
- 1
Equation (3.3) with threshold 0.50: q(4) = 0.1 + 0.1 x 3 = 0.40 < 0.50 gives 0, then q(5) = 0.1 + 0.1 x 4 = 0.50 >= 0.50 gives 1. The score jumps by 1 while q changed by 0.500 - 0.400 = 0.100. q rises by the same small step at every member, so the sudden event belongs mostly to the scoring rule, not to the model. A fit to the verdict would project a step at a particular scale, which is a claim about where the cutoff sits. The passing member reaches the threshold exactly, and Equation (3.3) counts that as a pass because it uses at least.
Worked steps
- q(4) = 0.1 + 0.1 x 3 = 0.40, below the threshold 0.50: verdict 0.
- q(5) = 0.1 + 0.1 x 4 = 0.50, at least the threshold 0.50: verdict 1.
- The verdict jumps by 1; q changed by 0.100 between the two members.
- Largest one-step change in q over the range: 0.100.
Use the idea
For any new family, record q beside the thresholded verdict and read the pair. If only the verdict is kept, a flat early stretch can be mistaken for no progress, and a jump can be mistaken for a change inside the model.
Where the conclusion applies
q is treated as known; in practice it is estimated from repeated attempts and carries sampling error. The four curves are constructed illustrations, and a graded measure does not guarantee smooth capability in an arbitrary model. A sharp rise in q is evidence of concentrated change in measured behavior, not proof of a phase transition.
Common wrong turn: A sudden score jump is a sudden change inside the model
What this does not settle
Nothing here establishes a universal threshold or a maximum amount of cooperation.
Chapter 3 source: "What this does not settle".
Check your understanding: Take the metric-created jump with q = 0.1 + 0.1 x (m - 1) and a threshold of 0.65. Which member is the first to pass?
Chapter 3 source: section "Smooth ability, abrupt score". Demonstration C03-D03.
Demonstration 4 of 4
Declare the family before the outcomes arrive
Why does it matter that the shape of the fitted relation is fixed before the held-out results are visible?
The line or curve is fitted to the three development members only and then frozen. The outcomes at the test scales score it and never change its coefficients. A family picked after seeing the outcomes can always look at least as good, because that choice spends the evidence. With log-shaped development data the log family forecasts exactly and the straight line falls short.
Scroll sideways for the whole equation
In this demonstration the measured score for member m is the constructed number on the vertical axis, and the declared feature is the scale x, used directly (the straight-line family) or as ln(x), the natural logarithm of x (the log family). The fitted line is the function f_psi in Equation (3.2), where psi stands for the line's fitted intercept and slope, and the plotted score stands in for the quantity Gamma_m on the left of that equation. The leftover epsilon_m is the residual, observed minus forecast. RMSE is the square root of the mean squared residual. The laboratory subtracts the other way, so its bias has the opposite sign. The first control picks the development data: a straight rise forecast at scales 4 and 5, or log-shaped data forecast at scales 8 and 16.
Predict first. Declare the straight-line family and let the outcomes flatten off (0.45, 0.48). Which family would fit better after the fact, and does that change the registered forecast?
Choose an example
Scroll sideways for the whole figure
Constructed example: the laboratory's development points (1, 0.2), (2, 0.3), (3, 0.4) with test scales 4 and 5, and its log-shaped transfer data (1, 0), (2, 0.6931), (4, 1.3863) with test scales 8 and 16; the outcomes other than the transfer case's continuation are defined for this reader.
Calculated values
- Declared family
- linear
- Fitted slope and intercept
- 0.1000, 0.1000
- Frozen forecast at 4 and 5
- 0.500, 0.600
- Observed at 4 and 5
- 0.80, 0.95
- Residuals, observed minus forecast
- 0.300, 0.350
- Test RMSE
- 0.326
- Mean residual (bias)
- 0.325
- RMSE of the log family, which was not declared
- 0.418
- Family that fits best after the fact
- linear
Forecast at x = 4 is 0.1000 x 4 + 0.1000 = 0.500. Residual at 4 = 0.80 - 0.500 = 0.300; at 5 = 0.95 - 0.600 = 0.350. RMSE = sqrt((0.300 x 0.300 + 0.350 x 0.350) / 2) = 0.326. The declared linear family is also the better fit, so this time declaring it first cost nothing. It would still not have been a prospective result had the family been picked after the outcomes.
Worked steps
- Development points (1, 0.20), (2, 0.30), (3, 0.40); the family is declared before any outcome: linear.
- Least squares on the development points: slope 0.1000, intercept 0.1000.
- Frozen forecasts: at 4 = 0.500; at 5 = 0.600.
- Residuals, observed minus forecast: 0.300 and 0.350.
- RMSE = 0.326; the log family would score 0.418, which cannot change the registered forecast.
Use the idea
Before a held-out result exists, write down the family, the transformation and the rule for choosing among several shapes, and file that with the forecast. Afterwards report the residual whether or not it is flattering.
Where the conclusion applies
Three development points and two test points, with only two candidate shapes. Both families fit three points reasonably well, which is exactly why the choice between them is a free parameter worth declaring. The demonstration cannot detect a family chosen in private after the outcomes; that is a matter of procedure. The outcomes other than the laboratory's own are defined for this reader.
Common wrong turn: Choosing the best family afterwards is a normal analytic decision with no cost
What this does not settle
Equation (3.2) is a proposal; the cited studies do not report that prospective interaction-contrast test. The worked forecast numbers are constructed.
Chapter 3 source: "What this does not settle".
Check your understanding: Declare the log family on the first data set and let the outcomes flatten off. What is the frozen forecast at x = 4?
Chapter 3 source: section "The eight-step protocol". Demonstration C03-D04.