Chapter 14: separate solutions

Question 1

Compute default epsilon.

Each row has half L1 distance 0.02; maximum is 0.02.

Question 2

Compute horizon 20 finite bound.

With the common reward bound 1 and epsilon 0.02, the finite-horizon bound is 1*0.02*20*19/2=3.8.

Question 3

What if the reward model also differs?

This transition-only bound is insufficient; add a declared reward-error term.