Kalman Filters were developed outside the context of “Machine Learning” and found practical applications much earlier.
If you define ML as an algorithm that is “trained by input data”, than any statistical model would be ML. So is the ambiguity around the terms ML and AI in general.
Where a KF is really going to kick the pants of a multi-layer perception/neural network is how computationally efficient it is. A KF only takes a couple of matrices of size N^2, where N is the number of variables you’re trying to predict. Compare this to a NN with hundreds/thousands of nodes. Also, the KF is “online learning” in that it “is trained as you go” rather than some other ML models that require upfront training, a KF is very useful for live update-and-predict use cases. And, again, it’s extremely computationally efficient, and can run easily on embedded systems. The book ad, here, suggests tracking: so tracking an airplane with an air traffic control radar would be an effective use for a KF. (Where NNets have found any other uses I’m sure you’re aware of.)
Another huge benefit of a KF is that, unlike NNets, KFs are “explainable”, and in fact extremely well understood by many professionals. This means that a KF can be better tuned to suit a purpose with less fear of unexpected results that may be more common in other ML models. Like, KF(S, x) will always return an explainable new state, where NN(x) may result in a surprise state and no amount of analysis can reveal why (and require training a new model, the “retrain and pray” solution).
Because of "resume driven engineering". Even simple problems that you can solve with PID controllers or Kalman Filters but everyone wants to throw ML at it instead, so they can put "ML experience" on their LinkedIn because that's what's hot right now as recruiters probably never heard of Kalman filters or PID controllers.
A lot of technical decisions aren't based on "what's the quickest, cheapest and easiest solution to the problem?" but "what solution is most likely to get me hired at a pay bump when I jump ship?"
The thing is, to train a NN to estimate your output from your input, you need input-output pairs. KFs are a way of measuring that output in the first place. So they are not even the same class of solutions.
Where a KF is really going to kick the pants of a multi-layer perception/neural network is how computationally efficient it is. A KF only takes a couple of matrices of size N^2, where N is the number of variables you’re trying to predict. Compare this to a NN with hundreds/thousands of nodes. Also, the KF is “online learning” in that it “is trained as you go” rather than some other ML models that require upfront training, a KF is very useful for live update-and-predict use cases. And, again, it’s extremely computationally efficient, and can run easily on embedded systems. The book ad, here, suggests tracking: so tracking an airplane with an air traffic control radar would be an effective use for a KF. (Where NNets have found any other uses I’m sure you’re aware of.)
Another huge benefit of a KF is that, unlike NNets, KFs are “explainable”, and in fact extremely well understood by many professionals. This means that a KF can be better tuned to suit a purpose with less fear of unexpected results that may be more common in other ML models. Like, KF(S, x) will always return an explainable new state, where NN(x) may result in a surprise state and no amount of analysis can reveal why (and require training a new model, the “retrain and pray” solution).
There’s a couple of differences for you.