Not weird really. On the one hand, you have a flagging policy that lets any random man-child with some specific bias and enough karma completely eliminate a post for whatever bullshit emotional reaction of their own, and at the same time, you have a bunch of people here who suffer from AI-fanboy derangement and hate to see any criticism of their fantasies about a glorious future in which current LLMs are just a hop away from making all things wonderful.
Related to the article, are there actively updated benchmarks for plan recognition?