Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

This is not convincing though that gpt4 actually understands the problem. Here's a slight variation I asked and it fails miserably.

https://chat.openai.com/share/22a9027f-a2c1-428a-94a2-8fd918...

I wonder what lends itself it answer correct in one situation but not the other? Was your question previously asked already and it recognized it whereas my question is different enough?



Your link is not to GPT4, your link is to the free version of ChatGPT, aka gpt-3.5-turbo (you can tell because the icon is green, not purple).

GPT4 indeed understands your variant, as evidenced here: https://chat.openai.com/share/46916f21-c469-4e93-9bed-bbd18b...


It's a bit random, which doesn't help, and different interfaces have different system prompts.

I repeated your question a few times and it got it wrong once, and right the others. It repeatedly mixed up who was supposed to be the host.

Here's a quote

> In the scenario you've described, you've initially chosen door number one, which you know has a car behind it because the doors are made of transparent glass. Since you already know the contents behind each door, the classic Monty Hall problem's probability-based decision-making does not apply here.


> Was your question previously asked already and it recognized it

Given that LLMs training data consists to a large extent of "stuff people have written on the internet", and The Monty Hall Problem is something that comes up as a topic for discussion on the internet not entirely infrequently - as well as having a wikipedia page - yes, I suspect that the words describing the monty hall problem being followed by words describing the correct solution appeared often in the training set, so LLMs are likely to reproduce that.

Words describing a problem similar to the monty hall problem are going to be less common, and probably have a lot of discussion about whether they accurately match the monty hall problem, and disagreement about what the right answer is. LLMs will confabulate something that looks like a plausible answer based on the language used in those discussions, because that's how they work. Whether they get a right answer is probably going to be much more up to chance.


You could say it doesn't "understand" anything really.


That's what I like about this problem (and similar Dumb variants of classic brain teasers). It exposes that there's not understanding, there's just a statistically weighted answer space. A question that looks a lot like a know popular topic ends up trapped in the probability distribution of the popular question.


How do you explain them answering correctly and explaining why it's different from the classic puzzle.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: