Compute default epsilon.
Each row has half L1 distance 0.02; maximum is 0.02.
Compute horizon 20 finite bound.
With the common reward bound 1 and epsilon 0.02, the finite-horizon bound is 1*0.02*20*19/2=3.8.
What if the reward model also differs?
This transition-only bound is insufficient; add a declared reward-error term.