Identify the input space, transformation group, output type and desired transformation law. Return explicit matched inputs, expected scalar invariance or vector equivariance, and numerical discrepancy measures when outputs are available. Treat semantic invariance as a hypothesis to audit.
“Should a rotated point cloud preserve a scalar prediction or rotate a vector prediction?”
Use mathllms-ch05-symmetry with the companion's AI skill package. The illustrations below also work on their own.
Which answers stay put, and which turn?
From Chapter 5, The Thing That Commutes; What the Group Leaves Alone
When a shape is turned, which of its measurements should stay the same and which should turn with it?
A length is just a number, so turning the shape cannot change it. An arrow has a direction, so its coordinates must change by the same turn that was applied to the input. The identity R-transpose times R equals the identity says a turn never stretches anything.
Predict first: Turn the triangle by 90 degrees. Does the edge arrow keep pointing to the right, and does its length change?
Turn applied (degrees): 45 · Triangle size (scale): 1
What happens: No to both. The arrow now points straight up, (0, 2), and its length is still 2. The length is invariant; the arrow is equivariant, so it turns with the input.
After a 45 degree turn the edge length is still 2, but the arrow now points along (1.41, 1.41). The length ignores the turn (invariant); the arrow has to turn with the input (equivariant).
- Edge length before
- 2
- Edge length after
- 2
- Edge arrow after the turn
- (1.41, 1.41)
- Gap from R times the old arrow
- 0
A model that reads molecules or point clouds should give a turn-proof answer for a single number such as an energy, and a turning answer for a direction such as a force. Mixing the two up is a common design error.
Show the calculation
Before the turn the edge is (2, 0), so its length is sqrt(4) = 2. The rotation matrix R has cos 0.707 and sin 0.707. R times the arrow gives (1.41, 1.41), whose squared length is 4: the same.
The equations and symbols
- x
- the input: three corners of a triangle in the plane
- R
- a turn (rotation) by the chosen angle
- ell
- length of the first edge, a single number
- v
- the first edge drawn as an arrow, with a direction
- scale
- size multiplier for the whole triangle
Euclidean coordinates with the same units on both axes. Rotations about the origin preserve length; scaling is a separate change. The output arrow is expressed in the same coordinate frame as the input. The triangle is a toy input.
Check your understanding: The edge is the arrow (0, 3) and the shape is turned by 180 degrees. What are the new arrow and its length?
Book source: Chapter 5, The Thing That Commutes; What the Group Leaves Alone. Illustration C05-D01. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Average over four quarter turns
From Chapter 5, The Minimum Viable Theory
If we average a measurement over the four right-angle turns of its input, does it stop caring about rotation altogether?
Measure each of the four turned copies of the input, then average the four results. Turning the input by one more quarter turn only reorders the four copies, so the average cannot change. A 45 degree tilt is not one of the four copies, so nothing forces the average to match.
Predict first: Will the averaged curve be flat for every input angle, or will it only repeat every 90 degrees?
Input angle (degrees): 0 · Radius of the input: 1
What happens: It only repeats every 90 degrees. The average is 1.5 at 0 degrees and dips to 1.25 at 45 degrees, then returns to 1.5 at 90. Averaging over four turns fixes the answer for those four turns, not for angles in between.
The four turned copies measure 1, 2, 1, 2, and their mean is 1.5. A further quarter turn only reorders the copies, so the mean repeats every 90 degrees. It is not constant: it is 1.5 at 0 degrees and 1.25 at 45 degrees, so it is not fully rotation invariant.
- Input point
- (1, 0)
- Raw measurement
- 1
- Average of the four copies
- 1.5
- Average after a further 90 degree turn
- 1.5
Image, map and sensor-grid models are often made robust to right-angle turns, by averaging or by training on turned copies. This shows the guarantee covers only those turns, so tilted inputs still need their own test.
Show the calculation
f(x) = x1^4 + 2 x2^2. Sum of the four values = 1 + 2 + 1 + 2 = 6; divide by 4 to get 1.5. Closed form: (x1^4 + x2^4)/2 + x1^2 + x2^2 = (1 + 0)/2 + 1 = 1.5.
The equations and symbols
- f
- a measurement that depends on direction
- C4
- the four quarter turns: 0, 90, 180 and 270 degrees
- g
- one of those four turns; "g inverse x" is the input turned back by g
- R[f]
- the four-turn average of f
- x
- the input point, at the chosen angle and radius
Exact, deterministic averaging over the four elements of C4 with equal weights. The measurement f is a toy function. The point is x = (radius cos a, radius sin a) at input angle a. No claim is made about invariance at other angles.
Check your understanding: Would averaging over only two turns (0 and 180 degrees) make the answer repeat every 90 degrees?
Book source: Chapter 5, The Minimum Viable Theory. Illustration C05-D02. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Slide then filter, or filter then slide
From Chapter 5, The Thing That Commutes; The Group Convolution
On a ring of eight samples, does sliding the input and then filtering give the same result as filtering and then sliding?
A shared ring convention lets sliding move the input and the output together. The filter looks the same from every position, so it does not matter whether the slide comes before or after. At a real image border the algebra is no longer exact.
Predict first: Slide the bump far enough that it wraps from position 7 back to position 0. Will the two orders still agree?
Slide by (places): 4 · Filter: smooth
What happens: Yes. The bump and its smoothed copy wrap round the ring, and both orders give the same eight numbers (the diamonds sit exactly on the bars). On a ring there is no edge to fall off.
Sliding by 4 and filtering give the same eight numbers as filtering and then sliding (largest gap 0), and here the pattern runs past position 7 and wraps round to 0. The ring has no edge, so nothing falls off.
- Largest gap between the two orders
- 0
- Filter weights
- [0.25 0.5 0.25]
- Output at position 6
- 1.25
Convolutional layers in image, audio and text models rely on this: a feature found in one place is found the same way in another. Padding and cropping at real edges can break the match.
Show the calculation
Output at position 6 adds input[j] times filter[(6 - j) mod 8] over j. Nonzero terms: 0.5 + 0.75 = 1.25.
The equations and symbols
- f
- the eight input samples on the ring
- k
- the filter: smoothing (0.25, 0.5, 0.25) or difference (1, -1)
- T_s
- slide every sample s places round the ring (numpy roll)
- i, j
- positions on the ring; indices wrap modulo 8
Cyclic domain of length eight, with the same index convention and filter in both orders. The difference filter is a convolution with weights (1, -1), not a cross-correlation. The input bump (0, 1, 3, 1, 0, 0, 0, 0) is a toy signal.
Check your understanding: Why can a cropped, zero-padded image break exact equivariance?
Book source: Chapter 5, The Thing That Commutes; The Group Convolution. Illustration C05-D03. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
A set has no first item
From Chapter 5, DeepSets and Permutation Invariance
When a collection has no natural order, which calculations ignore the order, and which still notice repeated items?
Summing a value per item gives the same total however the items are listed, which is what makes it order-invariant. A repeated item is a new term in the sum, so multiplicity is kept. A formula that multiplies by position depends on the listing.
Predict first: Add another copy of 4 to the collection. Does the order-free total stay the same?
Order the items are listed in: original · Extra copies of the value 4: 0
What happens: No. The sum of squares goes from 21 to 37, because the extra 4 adds 16, even though the order is untouched. Ignoring order is not the same as ignoring repeats.
The sum of squares is 21 for every listing, so order is invisible to it. The position-weighted sum is 17 here and changes with the order. Each extra 4 adds 16 to the sum of squares, so repeats are still counted.
- Sum of the values
- 7
- Sum of squares (order does not matter)
- 21
- Position-weighted sum (order matters)
- 17
- Number of items
- 3
Graph networks and attention over unordered neighbours sum or average messages. Whether a repeated neighbour counts twice depends on that choice, and it changes what the model can tell apart.
Show the calculation
Sum of squares = 1^2 + 2^2 + 4^2 = 21. Position-weighted sum = 1 x 1 + 2 x 2 + 3 x 4 = 17.
The equations and symbols
- x_i
- the measurements: 1, 2, 4, plus optional extra copies of 4
- phi(t)
- turns each item into the pair (t, t squared)
- rho
- reads out the second number of the total, the sum of squares
- position-weighted sum
- each item times its place in the list (1, 2, 3, ...), added up
A finite collection with exact sums. This shows elementary order-invariance, not a universal-approximation theorem. The listings are the original (1, 2, 4, ...), its reverse, and a one-place cycle.
Check your understanding: How do sum, mean and maximum change when (1, 2, 4) becomes (1, 2, 4, 4)?
Book source: Chapter 5, DeepSets and Permutation Invariance. Illustration C05-D04. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Turning the input only turns the phases
From Chapter 5, The Fourier Analysis on Groups; Worked example: Peter-Weyl for Z/4Z
When data on a four-position cycle is moved round the cycle, which part of its Fourier coefficients changes?
The coefficient at frequency k measures how much of the pattern that repeats k times round the cycle is present. Moving the data round the cycle by s places multiplies that coefficient by a number of size one, so the size is untouched and the phase turns by k times s quarter turns.
Predict first: Turn the input by a quarter turn. Which part of each Fourier coefficient changes: its size, its phase, or both?
Turn the input by (degrees): 90 · Input pattern: ramp
What happens: Only the phase. The sizes stay 1.5, 0.707, 0.5 and 0.707. After a half turn the frequency 1 and 3 arrows flip to the opposite side, while frequency 2 comes back to where it started, because it turns by 2 x 180 = 360 degrees.
Turning the input by 90 degrees leaves the sizes of all four coefficients at 1.5, 0.707, 0.5, 0.707. Only the phases move: frequency k turns by -90 times k degrees, counted modulo a full turn, so some frequencies land back where they started.
- Energy in the positions
- 3.5
- Energy in the frequencies
- 3.5
- Largest change in any coefficient size
- 0
- Phase turn of frequency 1
- -90 degrees
Spectral layers and FFT-based convolutions work on exactly these coefficients. Because a shift changes only phases, sizes give shift-proof features, and a convolution becomes one multiplication per frequency.
Show the calculation
f_hat(1) = (1/4) sum f(n) e^(-2 pi i n/4) = -0.5 + 0.5i before and 0.5 + 0.5i after. Predicted factor e^(-2 pi i s/4) with s = 1 is -1i; size 0.707 both times. Frequency 2 factor: -1.
The equations and symbols
- f(n)
- value at position n = 0, 1, 2, 3 round the cycle
- k
- frequency: how many times a pattern repeats round the cycle (0 to 3)
- f_hat(k)
- Fourier coefficient: a complex number with a size and a phase (an angle)
- T_s
- turn the input by s quarter turns, i.e. shift by s places
- i
- the imaginary unit, i squared = -1
Real signals on the cyclic group of four positions, with the 1/4 normalization used in the book. A quarter turn of the input is a one-place cyclic shift. The four input patterns are companion illustrations. A coefficient of size zero has no phase; the step pattern has one at frequency 2 and the arrow is then omitted.
Check your understanding: A pattern has frequency 3 coefficient 0.5 + 0.5i. After shifting by one place, what is it, and what is its size?
Book source: Chapter 5, The Fourier Analysis on Groups; Worked example: Peter-Weyl for Z/4Z. Illustration C05-D05. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Lifting an image into rotation channels
From Chapter 5, The Group Convolution; Worked example: G-CNN forward pass for p4
If one filter is used in four orientations, what happens to the four response maps when the input is turned by a quarter turn?
Lifting applies the same nine filter weights four times, each time turned by a further 90 degrees. Turning the image by 90 degrees is the same as turning the filter by 90 degrees and then turning the result, so each map moves to the next channel and is turned along the way.
Predict first: Turn the input by a quarter turn. Do the four response maps become new numbers, or the same maps in a different order?
Turn the input by (degrees): 90 · Filter: edge
What happens: The same maps, rearranged. Each map is turned by 180 degrees and moved two places along the orientation list. The channel totals 1, 5, -1, -5 come back as -1, -5, 1, 5.
Turning the input by 90 degrees leaves the four channel maps as the same set, each turned by 90 degrees and moved 1 place along the orientation list (largest gap 0). Only the order of the channels and the layout of each map change.
- Channel totals, original
- 1, 5, -1, -5
- Channel totals, turned input
- -5, 1, 5, -1
- Places the pattern moved
- 1
- Largest gap from the predicted maps
- 0
Rotation-aware convolutional layers reuse one filter in four orientations. A turned input then gives turned features rather than unrelated ones, so a model needs fewer turned training examples to recognise the same pattern.
Show the calculation
The four channels come from one set of 9 filter weights (each rotated by 0, 90, 180, 270 degrees). Channel totals before: 1, 5, -1, -5. After the 90 degree turn: -5, 1, 5, -1, the same numbers moved 1 place. A plain CNN with 4 separate 3x3 filters also has 36 weights but no such link.
The equations and symbols
- f
- the 4 by 4 input image
- k
- one 3 by 3 filter (9 weights), reused in four orientations
- g
- which orientation of the filter: 0, 90, 180 or 270 degrees (e, r, r squared, r cubed)
- F(g, x)
- response at pixel x using the filter turned by g: one map per orientation
- r
- a quarter turn
Zero padding, a single input channel and the toy 4 by 4 integer image shown in the notebook code. Cross-correlation, as deep-learning libraries compute it, stands in for the flipped sum in the formula; the equivariance law is the same. Turns are exact quarter turns, so the match is exact.
Check your understanding: If the input is turned by 270 degrees instead of 90, how many places do the channels move?
Book source: Chapter 5, The Group Convolution; Worked example: G-CNN forward pass for p4. Illustration C05-D06. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Matching feature types shrinks the layer
From Chapter 5, Schur's Lemma and the Shape of Equivariant Layers; Theorem 5.2
How many weights does a linear layer need if it must respect rotations, compared with an unconstrained one?
Each feature type is its own kind of object, and a rotation-respecting map can only send a type to itself, scaled by one number per pair of copies. Counting those allowed weights gives the small total. The same idea for images: a per-position map with 64 channels in and out on a 32 by 32 grid has 64 x 64 x 32 x 32 = 4,194,304 weights, while a 3 by 3 convolution has 64 x 64 x 9 = 36,864.
Predict first: With 32 scalar and 16 vector channels the saving is 5 times. If richer feature types (degrees 2 and 3, 16 channels each) are added, does the saving grow or shrink?
Channels per richer type (v): 16 · Highest degree included (1 = vectors): 1
What happens: It grows to about 41 times. The input has 272 numbers, so a dense layer needs 73,984 weights, while the equivariant layer needs 1,792 (32 x 32 plus three blocks of 16 x 16). The book's sixty-four-fold figure uses 32 channels per type instead of 16 (set v to 32 to see 64).
A dense layer on 80 numbers needs 6,400 weights. Keeping only same-type connections needs 1,280, a saving of 5 times. Every weight outside the teal blocks must be zero for the layer to be equivariant.
- Numbers per input
- 80
- Dense layer weights
- 6,400
- Equivariant layer weights
- 1,280
- Saving factor
- 5
Equivariant networks for molecules, proteins and 3D point clouds are small next to dense ones for this reason. The cost is that one linear layer cannot mix scalar and directional features.
Show the calculation
Input size = 32 + 16 x (3) = 80, so a dense layer has 80^2 = 6,400 weights. Equivariant: 32^2 + 16^2 = 1,280. Ratio 6,400 / 1,280 = 5. The book example (32 scalars, 16 vectors) gives 6,400 / 1,280 = 5.
The equations and symbols
- rho, sigma
- feature types: degree 0 scalars, degree 1 vectors, and richer degrees
- m, n
- how many copies of a type appear in the input and the output
- v
- channels of each richer type (the book uses 16 vector channels)
- L
- highest degree included; a degree l feature has 2l + 1 numbers
- 32
- number of scalar channels (from the book example)
Counts follow Theorem 5.2 for rotations of 3D space (SO(3)); each type used here is its own real type, so the count holds over the real numbers too. Input and output have the same channel mix: 32 scalar channels and v channels of each degree 1 to L. No biases and no distance-dependent weights.
Check your understanding: A rotation-respecting layer has 10 scalar and 10 vector channels in and out. How many weights, and what would a dense layer need?
Book source: Chapter 5, Schur's Lemma and the Shape of Equivariant Layers; Theorem 5.2. Illustration C05-D07. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
A moved molecule: energy stays, forces turn
From Chapter 5, Worked Example: SE(3)-Equivariant Molecular Property Prediction (energy output and forces)
If a molecule is moved as a rigid body, what happens to its energy, to the force on each atom, and to the total force?
Distances between atoms do not change when the whole molecule is rotated or shifted, so an energy made from them cannot change either. Differentiating an invariant energy gives forces that rotate with the molecule, and because the pull between each pair of atoms is equal and opposite, they sum to zero.
Predict first: After a rigid move the energy should stay the same and each force should turn with the molecule. What goes wrong if the energy formula uses raw x-coordinates instead of distances?
Rotation applied (degrees): 60 · Energy built from: distances
What happens: All three break. The energy changes from -1.16 to 1.61, the recomputed forces no longer match the turned old forces (on the left the gold and terracotta arrows part by the angle shown), and the total force is (-1.2, 0, 0) instead of zero.
The energy is 0.285 before and after the move, and each recomputed force matches the old force turned with the molecule (gap 0). The forces add up to (0, 0, 0): the molecule has no net push.
- Energy before to after the move
- 0.285 to 0.285
- Largest gap: new force vs turned old force
- 0
- Angle between them on atom 1 (degrees)
- 0
- Total force on the molecule
- (0, 0, 0)
Molecular models predict one energy and get forces by differentiating it. If the energy is invariant, the forces turn correctly and conserve momentum with no extra training.
Show the calculation
Atom 1 force before: (-0.641, -0.472, -0.588). Turning it by 60 degrees about z gives (0.0881, -0.791, -0.588); the force recomputed from the moved positions is (0.0881, -0.791, -0.588). Energy sums 1/2 x 2 x (distance - 1.2)^2 over the 6 atom pairs: 0.285 before, 0.285 after.
The equations and symbols
- r_i
- position of atom i in 3D space
- R, t
- the rigid move: a rotation R followed by a shift t
- E
- energy: here a toy spring energy built from distances only
- F_i
- force on atom i: minus the slope of E with respect to its position
- kappa, r_0
- spring stiffness 2 and rest length 1.2, toy values
Four atoms and a toy spring energy stand in for the learned energy network; forces are the exact derivative, what automatic differentiation would return. The rigid move is a rotation about the z axis through the origin, then a shift by (2.4, -1.6, 0.8). The plot is the x-y view from above; the molecule is 3D and z enters every calculation. The "raw coordinates" option adds 0.3 times the sum of x-coordinates, which is not invariant.
Check your understanding: Atom 1 feels the force (0.4, 0, 0) and the molecule is turned by 90 degrees about z. What force does atom 1 feel afterwards?
Book source: Chapter 5, Worked Example: SE(3)-Equivariant Molecular Property Prediction (energy output and forces). Illustration C05-D08. Illustration. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.
Bring the idea to a question of your own
Identify the input space, transformation group, output type and desired transformation law. Return explicit matched inputs, expected scalar invariance or vector equivariance, and numerical discrepancy measures when outputs are available. Treat semantic invariance as a hypothesis to audit.
The chapter skill can adapt the calculations to your inputs. It should identify the assumptions, explain what the result supports, and show what still needs evidence.