Hmm, that example is really interesting. I think you're right once we formalize L_infinity, because the goal will be to minimize the maximum possible distance. But the other scores can be phrased as penalties and minimizing expected penalty or total penalty summed over the data points, and it's not clear how to phrase L_infinity this way because the "penalty" would be infinity for any outcome...
I might be missing some context here, but usually the L_infinity norm means the max norm, or the maximum absolute difference over all components (data points). This gives (min + max)/2 as GP suggested.
For what it's worth, I think the confusion of your comment and jules's below comes from the idea that the lp-norm is sum(x_i^p), when it's actually [sum(x_i^p)^1/p] (please forgive my notation). Since raising to the 1/p power is monotonic, it doesn't actually make a difference in minimizing or maximizing the norm, so people often use them interchangeably.
You're right and we're on the same page about the solution. The reasoning behind my comment of "interesting" is that formula you gave is not well-defined for p=infinity, although we can define it as the limit of that expression as p-->infinity. Furthermore, for p < infinity the lp-norm can assign a well-defined penalty to each x, and the goal is to minimize the sum of the penalties. But that's not really true for p=infinity.
In the limit it's indeed the max of the absolute differences, so that's how you conventionally define the infinity norm (even though the base formula itself is not well defined, just as with 0). And when you try to minimise the max of the distances, you get the midrange.
I just want to give a bit more intuition on the L_p norms, basically a measure of length of a vector. A norm then induces a distance: the distance between a and b is the length of (a-b).
We assume that for one dimension, we can tell the length: length of d is just |d|, the absolute value of d. But what if we have many dimensions i=1..N?
Then we can use this L_p norm, and the idea is basically just that we set the norm to
[ sum_i=1..N |d_i|^p ] ^(1/p)
Now, what does that give us? Note a few things:
* when all of the d_i are zero except for one, we get back just the value that's not zero (in absolute value), because we go to the p-th power and then back. So, we recover the intuition (and number) from the one-dimensional case.
* for p=1, we just add all the absolute values of the d_i. That just gives us the citiblock/Manhattan norm.
L1 norm of [3, 4] = 7.
L1 norm of [100, 1, 1] = 102.
* for p=2, we square all values, add, and take the square root. This, by repeated application of Phytagoras, gives us the usual "Euclidean" distance in space: if you're 3 m west and 4 m north, then you're 5 = sqrt(9+16) m away.
L2 norm of [3, 4] = 5.
L2 norm of [100, 1, 1] = 100.00999
* now, think about what happens if d_1, say, is very big, and all the others are tiny. If we add them (L_1), we just get the sum, a bit bigger than d_1. If we add the squares, and then take the square root, we'll be even closer to d_1, because the square of d_1 is going to be very big, and the other squares really small, and then we add and take the square root to basically get back d_1. See example.
* Now imagine p=4, or 100. We take a huge power of all d_i, then add, then "invert" the power. All the terms in the sum are going to be dwarfed, except the biggest one!
L4 norm of [3, 4] = 4.284572295
L4 norm of [100, 1, 1] = 100.0000005
* Thus, the infinity norm goes towards the maximum of the d_i. And that's what it's defined as: max_i=1..N |d_i|
* So, big p tend to exaggerate the differences between the d_i, and only the biggest counts. Conversely, small p tend to make the d_i more similar, until it only really matters whether a d_i is 0 or not. In other words, you just count how many are not zero. And that motivates the L0 norm.
Now, think about the unit circle in 2D.
Here some ASCII art:
L0:
|
-+-
|
L1:
/\
/ \
\ /
\/
L2:
-
/ \
( )
\ _ /
Ok, that's my best ASCII rendering of a circle...
L_infty:
___
| |
|___|
You see that the 4 points up/down and left/right are fixed (-1, 0), (0, 1), (1, 0), (0, -1), because of this property that if all are zero except one we fall back to the normal distance. But then, for others, the points we consider close enough to be within the unit circle "bulge out" as we increase p.
* L_0 -> mode
* L_1 -> median
* L_2 -> mean
* L_infinity -> midrange, I think
that is, (smallest observation + largest observation)/2
(BTW, the author is also a big contributor to the wonderful Julia language, I believe)