Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

please elaborate - are you thinking of the curse of dimensionality ?


There are many examples, one I came across recently is that the large majority of the probability mass of a high-dimensional gaussian distribution is in a shell at a distance from the mean, the mass at the center is actually quite low.

Also anything related to topology, which is important when you are looking at decision boundaries, becomes counterintuitive in high dimensions, because so many things can be adjacent at the same time.


Can you please try and explain why that is?

If true, you're very correct that lower-dimensional intuition does not transfer into higher-dimensional spaces: my intuition tells me that a Gaussian distribution drops off as you fall away from the mean, and it's quite easy for me to imagine that in 2 dimensions, 3 dimensions (e.g. by imagining a mound on a plane) and 4 dimensions (e.g. a cloud in 3-space with increased density around the mean).

Is my intuition wrong in any of those cases? If so, why? If not, how many dimensions do we need before it becomes wrong?


Because an outlier in any single dimension will put the point outside the "center" of the distribution, and as the number of dimensions increases there's more of a chance of that happening.

Say you have an N-dimensional gaussian where each dimension has mean 0 and standard deviation 1. Define the center as the N-dimensional cube whose edges go from -3 to +3 in each dimension. A normally distributed value is within 3 standard deviations of the mean with probability 0.9973, so the probability that an N-dimensional point being in the center is 0.9973^N. With N=4 that's 0.989 which matches your intuition, but at N=1000 it's 0.067 and at N=10000 it's 1.81e-12.


The center of the distribution always has the highest density, but the ratio of 'probability mass close to centroid' / 'total probability mass' drops off as number of dimensions grows.

This is somewhat related to another 'curse of dimensionality' observation, which is that the volume of a hyperball / volume of hyperspace tends towards zero as dimensions grow -- there's just a lot more volume that's in some sense 'far' from the center.


>If true, you're very correct that lower-dimensional intuition does not transfer into higher-dimensional spaces: my intuition tells me that a Gaussian distribution drops off as you fall away from the mean, and it's quite easy for me to imagine that in 2 dimensions, 3 dimensions (e.g. by imagining a mound on a plane) and 4 dimensions (e.g. a cloud in 3-space with increased density around the mean).

Density is different from mass. Namely, mass is the integral of density. So your intuition is roughly correct for density, but you need to make it accord with a good intuition for mass.

Since getting the mass requires an integral, getting the mass over N-dimensional distributions requires integrating an N-dimensional region, which means N integrations for N dimensions. Each integration is, intuitively, a kind of sum. Integrating out many dimensions happens recursively; looped or recursive addition is multiplication. So on some level, to take the probability mass of a region in N-dimensional space, you need to "multiply" a density.

Since the total probability mass is fixed (1.0), adding more dimensions means you need to "multiply" the density by a larger number to get the mass, which means you need to divide the mass by a larger number to get the density, which means that despite the density peaking at the mean, the available density at any given point gets smaller as the dimensionality rises.


> it's quite easy for me to imagine that in 2 dimensions

It starts to fail really badly when dimension grows.

Two simple examples:

1) Consider 3 dimensional unit sphere centered at origin and unit cube centered at origin. Cube is clearly completely inside the sphere. Now generalize to n-dimensions. Hyperdimensional volume of hypercube with side length 1 moves almost completely outside the n-sphere with radius 1 when n-grows.

2) Alternatively almost all volume of n-sphere is close to the surface.

These are all very counterintuitive, yet simple to check toy examples. When you start to integrate over more complex multidimensional function, things get weird really fast.


>Alternatively almost all volume of n-sphere is close to the surface.

How does this go against intuition?

Intuition from 1/2/3d tells me that the volume of an N-ball is O(r^N), and indeed it is the case in higher dimensions. Therefore it’s easy to see that the difference between the volume of an N-ball of radius r and an N-ball of radius (r + epsilon) will grow exponentially with N.


isn't this just because we're comparing n-dimensional objects by a 2-norm ? i.e. the dimension of the space grows but we're keeping the dimension of the norm fixed, but if we used the p-norm of the same dimension as the space, then maybe that would return intuitive results ?


This is the problem that makes my brain melt when I try to think about genetics and mutational load. The naive idea is that a species, S, has a correct genome, G, but mutations build up, increasing with each generation. Presumably mutations build up until they are common enough that back-mutations are a thing. Then there is an equilibrium. In a fecund species, each individual has many children, but most have a higher mutational load, many have the save mutational load, and a lucky few have a smaller mutational load, closer to G, the correct genome. Differential reproductive success then maintains the equilibrium.

I don't see how the numbers are supposed to work out for large mammals, with each female having under a dozen offspring. To have a decent chance of a back-mutation, the typical member of the species would need one twelfth of their genome to be deleterious mutations.

Meanwhile, people are thinking about using CRISPR to correct the human genome, creating unusually happy, healthy people. The underlying thought is that the correct genome is best. But why do we think that the correct genome works at all?

Most of the population is in a shell at a distance from the correct genome, the number at the center is actually quite low. Given the combinatorics, with two to the millions of possible genomes, but populations in the millions, the number at the center, or even close, is actually zero. Maybe the correct genome codes for a sickly, miserable individual?

My current guess is that the evolution of large mammals with few offspring is constrained by genetic load considerations. It is not sufficient, (or even necessary) for the correct genome to be any good. There needs to be a big blob of mediocrity in genome space. The species exists as a shell of individuals on the edge of the blob of mediocrity. The blob needs to be huge, so that individuals whose genome is one twelfth mutations are still in the blob. Then there can be an equilibrium between back-mutations, taking offspring towards the interior of the blob and other mutations, taking offspring out of the blob and out of the gene pool.

This potentially solves the Fermi paradox. Can creatures such as humans actually exist in this universe? It is not enough for natural selection to discover a good genome. Natural selection has to discover a huge blob of mediocrity. Such blobs might be vastly rarer than we realize.

This potentially shits on the CRISPR master race. There might be nothing special about the interior of the huge blob of mediocrity.


What is the "correct genome"? Seems like you could only define it as a local minimum in fitness space, or some kind of attractor.


I don't know.There is medical perspective which focuses on deleterious mutations causing disease. This is a black-and-white perspective which sees mutations as either wholly bad or entirely unimportant. A genome without any deleterious mutations is a correct genome.

But what happens if you step back from black-and-white thinking and ask about mutations with ambiguous effects. Which is the mutation and which is the correct genome? It becomes unclear.

An alternative perspective asks: how well separated are the local minima in fitness space? Perhaps the typical separation is as large as the gaps between species. Then each species has only its own local minimum, which defines its correct genome. Or perhaps fitness space is littered with local minima, such that a single species has genetically healthy individuals in several different minima plus other individuals, perhaps not quite so healthy, nearby.


Theoretically, yes. Can you give a more concrete example? Many hard high dimensional and general topology problems can be visualized through their 2D special cases.


Yeah, but the distance from the center is itself just a one-dimensional gaussian...


Related: Hamming's "The Art of Doing Science and Engineering" chapter 9, N-Dimensional Space

http://worrydream.com/refs/Hamming-TheArtOfDoingScienceAndEn...

(I assume Bret Victor has permission to host the PDF on his website, he is far from an anonymous pirate)




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: