The training algorithm is probably the wrong place to ensure the behavior, point taken. It probably requires some harness level intervention. This is essentially the rationale behind the first two laws of robotics, follow an order unless it harms a human.