The mathematical companion · Chapter 5
Explore · Calculate · Apply

Symmetry and Geometric Inductive Bias

Specify what should stay fixed and what should transform.

8 guided illustrations. Move a slider or choose a value, watch the mathematics change, and check your prediction. All calculations are included; no account or connection is required.

Identify the input space, transformation group, output type and desired transformation law. Return explicit matched inputs, expected scalar invariance or vector equivariance, and numerical discrepancy measures when outputs are available. Treat semantic invariance as a hypothesis to audit.

Try asking the chapter skill

“Should a rotated point cloud preserve a scalar prediction or rotate a vector prediction?”

Use mathllms-ch05-symmetry with the companion's AI skill package. The illustrations below also work on their own.

01 / 08

Which answers stay put, and which turn?

From Chapter 5, The Thing That Commutes; What the Group Leaves Alone

When a shape is turned, which of its measurements should stay the same and which should turn with it?

A length is just a number, so turning the shape cannot change it. An arrow has a direction, so its coordinates must change by the same turn that was applied to the input. The identity R-transpose times R equals the identity says a turn never stretches anything.

Predict first: Turn the triangle by 90 degrees. Does the edge arrow keep pointing to the right, and does its length change?

0270
Left: a triangle and its turned copy with the first edge as an arrow. Right: edge length is flat while the arrow parts swing with the turn.
A length ignores the turn (invariant), but an arrow must turn with the input (equivariant).
Turn applied (degrees): 45 · Triangle size (scale): 1

After a 45 degree turn the edge length is still 2, but the arrow now points along (1.41, 1.41). The length ignores the turn (invariant); the arrow has to turn with the input (equivariant).

Edge length before
2
Edge length after
2
Edge arrow after the turn
(1.41, 1.41)
Gap from R times the old arrow
0
Why it matters for language models

A model that reads molecules or point clouds should give a turn-proof answer for a single number such as an energy, and a turning answer for a direction such as a force. Mixing the two up is a common design error.

Show the calculation

Before the turn the edge is (2, 0), so its length is sqrt(4) = 2. The rotation matrix R has cos 0.707 and sin 0.707. R times the arrow gives (1.41, 1.41), whose squared length is 4: the same.

The equations and symbols

ℓ(Rx)=ℓ(x),v(Rx)=Rv(x) \ell(Rx)=\ell(x),\qquad v(Rx)=Rv(x)

R⊤R=I R^\top R=I

x
the input: three corners of a triangle in the plane
R
a turn (rotation) by the chosen angle
ell
length of the first edge, a single number
v
the first edge drawn as an arrow, with a direction
scale
size multiplier for the whole triangle
Where the conclusion applies

Euclidean coordinates with the same units on both axes. Rotations about the origin preserve length; scaling is a separate change. The output arrow is expressed in the same coordinate frame as the input. The triangle is a toy input.

Check your understanding: The edge is the arrow (0, 3) and the shape is turned by 180 degrees. What are the new arrow and its length?
The new arrow is (0, -3) and the length is still 3: a half turn flips the direction and keeps the size.

Book source: Chapter 5, The Thing That Commutes; What the Group Leaves Alone. Illustration C05-D01. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

02 / 08

Average over four quarter turns

From Chapter 5, The Minimum Viable Theory

If we average a measurement over the four right-angle turns of its input, does it stop caring about rotation altogether?

Measure each of the four turned copies of the input, then average the four results. Turning the input by one more quarter turn only reorders the four copies, so the average cannot change. A 45 degree tilt is not one of the four copies, so nothing forces the average to match.

Predict first: Will the averaged curve be flat for every input angle, or will it only repeat every 90 degrees?

0180
Left: raw and four-turn-average curves against angle, repeating every 90 degrees with a dip at 45. Right: four bars and their mean.
Averaging over four right-angle turns makes the answer repeat every 90 degrees, but not stay constant at every angle.
Input angle (degrees): 0 · Radius of the input: 1

The four turned copies measure 1, 2, 1, 2, and their mean is 1.5. A further quarter turn only reorders the copies, so the mean repeats every 90 degrees. It is not constant: it is 1.5 at 0 degrees and 1.25 at 45 degrees, so it is not fully rotation invariant.

Input point
(1, 0)
Raw measurement
1
Average of the four copies
1.5
Average after a further 90 degree turn
1.5
Why it matters for language models

Image, map and sensor-grid models are often made robust to right-angle turns, by averaging or by training on turned copies. This shows the guarantee covers only those turns, so tilted inputs still need their own test.

Show the calculation

f(x) = x1^4 + 2 x2^2. Sum of the four values = 1 + 2 + 1 + 2 = 6; divide by 4 to get 1.5. Closed form: (x1^4 + x2^4)/2 + x1^2 + x2^2 = (1 + 0)/2 + 1 = 1.5.

The equations and symbols

ℛ[f](x)=14∑g∈C4f(g−1x) \mathcal R[f](x)=\tfrac14\sum_{g\in C_4} f(g^{-1}x)

f(x)=x14+2x22 f(x)=x_1^4+2x_2^2

f
a measurement that depends on direction
C4
the four quarter turns: 0, 90, 180 and 270 degrees
g
one of those four turns; "g inverse x" is the input turned back by g
R[f]
the four-turn average of f
x
the input point, at the chosen angle and radius
Where the conclusion applies

Exact, deterministic averaging over the four elements of C4 with equal weights. The measurement f is a toy function. The point is x = (radius cos a, radius sin a) at input angle a. No claim is made about invariance at other angles.

Check your understanding: Would averaging over only two turns (0 and 180 degrees) make the answer repeat every 90 degrees?
No, only every 180 degrees. Here f(-x) = f(x), so the two-turn average is just f, which is 1 at angle 0 and 2 at angle 90 (radius 1).

Book source: Chapter 5, The Minimum Viable Theory. Illustration C05-D02. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

03 / 08

Slide then filter, or filter then slide

From Chapter 5, The Thing That Commutes; The Group Convolution

On a ring of eight samples, does sliding the input and then filtering give the same result as filtering and then sliding?

A shared ring convention lets sliding move the input and the output together. The filter looks the same from every position, so it does not matter whether the slide comes before or after. At a real image border the algebra is no longer exact.

Predict first: Slide the bump far enough that it wraps from position 7 back to position 0. Will the two orders still agree?

07
Left: an eight-position bump and its slid copy as paired bars. Right: output bars with diamonds on top showing both orders agree.
On a ring, filtering and sliding can be done in either order with the same result, even when the pattern wraps round.
Slide by (places): 4 · Filter: smooth

Sliding by 4 and filtering give the same eight numbers as filtering and then sliding (largest gap 0), and here the pattern runs past position 7 and wraps round to 0. The ring has no edge, so nothing falls off.

Largest gap between the two orders
0
Filter weights
[0.25 0.5 0.25]
Output at position 6
1.25
Why it matters for language models

Convolutional layers in image, audio and text models rely on this: a feature found in one place is found the same way in another. Padding and cropping at real edges can break the match.

Show the calculation

Output at position 6 adds input[j] times filter[(6 - j) mod 8] over j. Nonzero terms: 0.5 + 0.75 = 1.25.

The equations and symbols

(f*k)[i]=∑j=07f[j]k[(i−j)mod⁡8] (f*k)[i]=\sum_{j=0}^{7}f[j]k[(i-j)\bmod8]

Ts(f*k)=(Tsf)*k T_s(f*k)=(T_sf)*k

f
the eight input samples on the ring
k
the filter: smoothing (0.25, 0.5, 0.25) or difference (1, -1)
T_s
slide every sample s places round the ring (numpy roll)
i, j
positions on the ring; indices wrap modulo 8
Where the conclusion applies

Cyclic domain of length eight, with the same index convention and filter in both orders. The difference filter is a convolution with weights (1, -1), not a cross-correlation. The input bump (0, 1, 3, 1, 0, 0, 0, 0) is a toy signal.

Check your understanding: Why can a cropped, zero-padded image break exact equivariance?
Values slid past the crop are thrown away and zero padding brings in different border values. Compare interiors only, or use an explicitly cyclic domain.

Book source: Chapter 5, The Thing That Commutes; The Group Convolution. Illustration C05-D03. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

04 / 08

A set has no first item

From Chapter 5, DeepSets and Permutation Invariance

When a collection has no natural order, which calculations ignore the order, and which still notice repeated items?

Summing a value per item gives the same total however the items are listed, which is what makes it order-invariant. A repeated item is a new term in the sum, so multiplicity is kept. A formula that multiplies by position depends on the listing.

Predict first: Add another copy of 4 to the collection. Does the order-free total stay the same?

05
Left: bars for the listed values 1, 2, 4. Right: for three orders, sum-of-squares bars are equal and position-weighted bars differ.
Adding up a value for each item ignores order but still counts repeats; a result built from positions does not ignore order.
Order the items are listed in: original · Extra copies of the value 4: 0

The sum of squares is 21 for every listing, so order is invisible to it. The position-weighted sum is 17 here and changes with the order. Each extra 4 adds 16 to the sum of squares, so repeats are still counted.

Sum of the values
7
Sum of squares (order does not matter)
21
Position-weighted sum (order matters)
17
Number of items
3
Why it matters for language models

Graph networks and attention over unordered neighbours sum or average messages. Whether a repeated neighbour counts twice depends on that choice, and it changes what the model can tell apart.

Show the calculation

Sum of squares = 1^2 + 2^2 + 4^2 = 21. Position-weighted sum = 1 x 1 + 2 x 2 + 3 x 4 = 17.

The equations and symbols

f({xi})=ρ(∑iϕ(xi)),ϕ(t)=(t,t2) f(\{x_i\})=\rho\left(\sum_i\phi(x_i)\right),\quad\phi(t)=(t,t^2)

ρ(z1,z2)=z2 \rho(z_1,z_2)=z_2

x_i
the measurements: 1, 2, 4, plus optional extra copies of 4
phi(t)
turns each item into the pair (t, t squared)
rho
reads out the second number of the total, the sum of squares
position-weighted sum
each item times its place in the list (1, 2, 3, ...), added up
Where the conclusion applies

A finite collection with exact sums. This shows elementary order-invariance, not a universal-approximation theorem. The listings are the original (1, 2, 4, ...), its reverse, and a one-place cycle.

Check your understanding: How do sum, mean and maximum change when (1, 2, 4) becomes (1, 2, 4, 4)?
Sum goes from 7 to 11, mean from 7/3 to 11/4, and the maximum stays 4. All three ignore order, but they keep different information about repeats.

Book source: Chapter 5, DeepSets and Permutation Invariance. Illustration C05-D04. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

05 / 08

Turning the input only turns the phases

From Chapter 5, The Fourier Analysis on Groups; Worked example: Peter-Weyl for Z/4Z

When data on a four-position cycle is moved round the cycle, which part of its Fourier coefficients changes?

The coefficient at frequency k measures how much of the pattern that repeats k times round the cycle is present. Moving the data round the cycle by s places multiplies that coefficient by a number of size one, so the size is untouched and the phase turns by k times s quarter turns.

Predict first: Turn the input by a quarter turn. Which part of each Fourier coefficient changes: its size, its phase, or both?

Left: four values on a cycle before and after a turn. Right: complex-plane arrows for frequencies 0 to 3 keep their size and turn along circles.
Turning the input rotates the phase of frequency k by k quarter steps and leaves every coefficient size unchanged.
Turn the input by (degrees): 90 · Input pattern: ramp

Turning the input by 90 degrees leaves the sizes of all four coefficients at 1.5, 0.707, 0.5, 0.707. Only the phases move: frequency k turns by -90 times k degrees, counted modulo a full turn, so some frequencies land back where they started.

Energy in the positions
3.5
Energy in the frequencies
3.5
Largest change in any coefficient size
0
Phase turn of frequency 1
-90 degrees
Why it matters for language models

Spectral layers and FFT-based convolutions work on exactly these coefficients. Because a shift changes only phases, sizes give shift-proof features, and a convolution becomes one multiplication per frequency.

Show the calculation

f_hat(1) = (1/4) sum f(n) e^(-2 pi i n/4) = -0.5 + 0.5i before and 0.5 + 0.5i after. Predicted factor e^(-2 pi i s/4) with s = 1 is -1i; size 0.707 both times. Frequency 2 factor: -1.

The equations and symbols

f̂(k)=14∑n=03f(n)e−2πikn/4 \hat f(k)=\tfrac14\sum_{n=0}^{3}f(n)\,e^{-2\pi i kn/4}

(Tsf)(n)=f(n−s)⇒Tsf̂(k)=e−2πiks/4f̂(k) (T_sf)(n)=f(n-s)\ \Rightarrow\ \widehat{T_sf}(k)=e^{-2\pi i ks/4}\,\hat f(k)

f(n)
value at position n = 0, 1, 2, 3 round the cycle
k
frequency: how many times a pattern repeats round the cycle (0 to 3)
f_hat(k)
Fourier coefficient: a complex number with a size and a phase (an angle)
T_s
turn the input by s quarter turns, i.e. shift by s places
i
the imaginary unit, i squared = -1
Where the conclusion applies

Real signals on the cyclic group of four positions, with the 1/4 normalization used in the book. A quarter turn of the input is a one-place cyclic shift. The four input patterns are companion illustrations. A coefficient of size zero has no phase; the step pattern has one at frequency 2 and the arrow is then omitted.

Check your understanding: A pattern has frequency 3 coefficient 0.5 + 0.5i. After shifting by one place, what is it, and what is its size?
Multiply by e^(-2 pi i x 3/4) = i, giving -0.5 + 0.5i. The size is 0.707 both before and after.

Book source: Chapter 5, The Fourier Analysis on Groups; Worked example: Peter-Weyl for Z/4Z. Illustration C05-D05. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

06 / 08

Lifting an image into rotation channels

From Chapter 5, The Group Convolution; Worked example: G-CNN forward pass for p4

If one filter is used in four orientations, what happens to the four response maps when the input is turned by a quarter turn?

Lifting applies the same nine filter weights four times, each time turned by a further 90 degrees. Turning the image by 90 degrees is the same as turning the filter by 90 degrees and then turning the result, so each map moves to the next channel and is turned along the way.

Predict first: Turn the input by a quarter turn. Do the four response maps become new numbers, or the same maps in a different order?

Left: total response for four filter orientations before and after turning the input, with the peak moving. Right: four response maps.
Turning the input turns every response map and moves it to the next orientation channel; no new responses appear.
Turn the input by (degrees): 90 · Filter: edge

Turning the input by 90 degrees leaves the four channel maps as the same set, each turned by 90 degrees and moved 1 place along the orientation list (largest gap 0). Only the order of the channels and the layout of each map change.

Channel totals, original
1, 5, -1, -5
Channel totals, turned input
-5, 1, 5, -1
Places the pattern moved
1
Largest gap from the predicted maps
0
Why it matters for language models

Rotation-aware convolutional layers reuse one filter in four orientations. A turned input then gives turned features rather than unrelated ones, so a model needs fewer turned training examples to recognise the same pattern.

Show the calculation

The four channels come from one set of 9 filter weights (each rotated by 0, 90, 180, 270 degrees). Channel totals before: 1, 5, -1, -5. After the 90 degree turn: -5, 1, 5, -1, the same numbers moved 1 place. A plain CNN with 4 separate 3x3 filters also has 36 weights but no such link.

The equations and symbols

[Lkf](g,x)=∑yf(y)k(g−1(x−y)),g∈{e,r,r2,r3} [L_kf](g,x)=\sum_{y}f(y)\,k\big(g^{-1}(x-y)\big),\qquad g\in\{e,r,r^2,r^3\}

F(g,x)[T90∘f]=F(r−1g,r−1x)[f] F(g,x)[T_{90^\circ}f]=F(r^{-1}g,\,r^{-1}x)[f]

f
the 4 by 4 input image
k
one 3 by 3 filter (9 weights), reused in four orientations
g
which orientation of the filter: 0, 90, 180 or 270 degrees (e, r, r squared, r cubed)
F(g, x)
response at pixel x using the filter turned by g: one map per orientation
r
a quarter turn
Where the conclusion applies

Zero padding, a single input channel and the toy 4 by 4 integer image shown in the notebook code. Cross-correlation, as deep-learning libraries compute it, stands in for the flipped sum in the formula; the equivariance law is the same. Turns are exact quarter turns, so the match is exact.

Check your understanding: If the input is turned by 270 degrees instead of 90, how many places do the channels move?
Three places in the same direction (equivalently one place back), because 270 degrees is three quarter turns.

Book source: Chapter 5, The Group Convolution; Worked example: G-CNN forward pass for p4. Illustration C05-D06. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

07 / 08

Matching feature types shrinks the layer

From Chapter 5, Schur's Lemma and the Shape of Equivariant Layers; Theorem 5.2

How many weights does a linear layer need if it must respect rotations, compared with an unconstrained one?

Each feature type is its own kind of object, and a rotation-respecting map can only send a type to itself, scaled by one number per pair of copies. Counting those allowed weights gives the small total. The same idea for images: a per-position map with 64 channels in and out on a 32 by 32 grid has 64 x 64 x 32 x 32 = 4,194,304 weights, while a 3 by 3 convolution has 64 x 64 x 9 = 36,864.

Predict first: With 32 scalar and 16 vector channels the saving is 5 times. If richer feature types (degrees 2 and 3, 16 channels each) are added, does the saving grow or shrink?

048
Left: block matrix with free weights only on the diagonal blocks. Right: saving factor against channels for three choices of highest degree.
Schur's Lemma lets an equivariant layer connect only matching feature types, so most of a dense layer's weights are forced to zero.
Channels per richer type (v): 16 · Highest degree included (1 = vectors): 1

A dense layer on 80 numbers needs 6,400 weights. Keeping only same-type connections needs 1,280, a saving of 5 times. Every weight outside the teal blocks must be zero for the layer to be equivariant.

Numbers per input
80
Dense layer weights
6,400
Equivariant layer weights
1,280
Saving factor
5
Why it matters for language models

Equivariant networks for molecules, proteins and 3D point clouds are small next to dense ones for this reason. The cost is that one linear layer cannot mix scalar and directional features.

Show the calculation

Input size = 32 + 16 x (3) = 80, so a dense layer has 80^2 = 6,400 weights. Equivariant: 32^2 + 16^2 = 1,280. Ratio 6,400 / 1,280 = 5. The book example (32 scalars, 16 vectors) gives 6,400 / 1,280 = 5.

The equations and symbols

dim⁡HomG(V,W)=∑ρ≃σmρnσ \dim\mathrm{Hom}_G(V,W)=\sum_{\rho\simeq\sigma}m_\rho\,n_\sigma

dense=(32+v∑ℓ=1L(2ℓ+1))2,equivariant=322+Lv2 \text{dense}=\Big(32+v\sum_{\ell=1}^{L}(2\ell+1)\Big)^2,\qquad \text{equivariant}=32^2+L\,v^2

rho, sigma
feature types: degree 0 scalars, degree 1 vectors, and richer degrees
m, n
how many copies of a type appear in the input and the output
v
channels of each richer type (the book uses 16 vector channels)
L
highest degree included; a degree l feature has 2l + 1 numbers
32
number of scalar channels (from the book example)
Where the conclusion applies

Counts follow Theorem 5.2 for rotations of 3D space (SO(3)); each type used here is its own real type, so the count holds over the real numbers too. Input and output have the same channel mix: 32 scalar channels and v channels of each degree 1 to L. No biases and no distance-dependent weights.

Check your understanding: A rotation-respecting layer has 10 scalar and 10 vector channels in and out. How many weights, and what would a dense layer need?
Equivariant: 10 x 10 + 10 x 10 = 200. Dense: the input has 10 + 3 x 10 = 40 numbers, so 40 x 40 = 1,600, which is 8 times more.

Book source: Chapter 5, Schur's Lemma and the Shape of Equivariant Layers; Theorem 5.2. Illustration C05-D07. Identity. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

08 / 08

A moved molecule: energy stays, forces turn

From Chapter 5, Worked Example: SE(3)-Equivariant Molecular Property Prediction (energy output and forces)

If a molecule is moved as a rigid body, what happens to its energy, to the force on each atom, and to the total force?

Distances between atoms do not change when the whole molecule is rotated or shifted, so an energy made from them cannot change either. Differentiating an invariant energy gives forces that rotate with the molecule, and because the pull between each pair of atoms is equal and opposite, they sum to zero.

Predict first: After a rigid move the energy should stay the same and each force should turn with the molecule. What goes wrong if the energy formula uses raw x-coordinates instead of distances?

0270
Left: force on atom 1 turned with the molecule versus recomputed, from one point. Right: energy flat while atom 1 force parts swing with angle.
An energy built from distances is unchanged by a rigid move, so its forces turn with the molecule and add to zero; raw coordinates break all three.
Rotation applied (degrees): 60 · Energy built from: distances

The energy is 0.285 before and after the move, and each recomputed force matches the old force turned with the molecule (gap 0). The forces add up to (0, 0, 0): the molecule has no net push.

Energy before to after the move
0.285 to 0.285
Largest gap: new force vs turned old force
0
Angle between them on atom 1 (degrees)
0
Total force on the molecule
(0, 0, 0)
Why it matters for language models

Molecular models predict one energy and get forces by differentiating it. If the energy is invariant, the forces turn correctly and conserve momentum with no extra training.

Show the calculation

Atom 1 force before: (-0.641, -0.472, -0.588). Turning it by 60 degrees about z gives (0.0881, -0.791, -0.588); the force recomputed from the moved positions is (0.0881, -0.791, -0.588). Energy sums 1/2 x 2 x (distance - 1.2)^2 over the 6 atom pairs: 0.285 before, 0.285 after.

The equations and symbols

E=∑i<j12κ(‖ri−rj‖−r0)2 E=\sum_{i<j}\tfrac12\kappa\,(\lVert r_i-r_j\rVert-r_0)^2

Fi=−∇riE,E(Rr+t)=E(r)⇒Fi(Rr+t)=RFi(r) F_i=-\nabla_{r_i}E,\qquad E(Rr+t)=E(r)\ \Rightarrow\ F_i(Rr+t)=R\,F_i(r)

r_i
position of atom i in 3D space
R, t
the rigid move: a rotation R followed by a shift t
E
energy: here a toy spring energy built from distances only
F_i
force on atom i: minus the slope of E with respect to its position
kappa, r_0
spring stiffness 2 and rest length 1.2, toy values
Where the conclusion applies

Four atoms and a toy spring energy stand in for the learned energy network; forces are the exact derivative, what automatic differentiation would return. The rigid move is a rotation about the z axis through the origin, then a shift by (2.4, -1.6, 0.8). The plot is the x-y view from above; the molecule is 3D and z enters every calculation. The "raw coordinates" option adds 0.3 times the sum of x-coordinates, which is not invariant.

Check your understanding: Atom 1 feels the force (0.4, 0, 0) and the molecule is turned by 90 degrees about z. What force does atom 1 feel afterwards?
The force turns with the molecule, so it becomes (0, 0.4, 0). Its size, 0.4, is unchanged.

Book source: Chapter 5, Worked Example: SE(3)-Equivariant Molecular Property Prediction (energy output and forces). Illustration C05-D08. Illustration. Book equation with companion toy inputs stated in the assumptions; every plotted value is recomputed. No random numbers are used. v39 EPUB / v43 print.

Bring the idea to a question of your own

Identify the input space, transformation group, output type and desired transformation law. Return explicit matched inputs, expected scalar invariance or vector equivariance, and numerical discrepancy measures when outputs are available. Treat semantic invariance as a hypothesis to audit.

The chapter skill can adapt the calculations to your inputs. It should identify the assumptions, explain what the result supports, and show what still needs evidence.