Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

On logic it cannot handle the Dumb Monty Hall problem at all:

https://g.co/gemini/share/33c5fb45738f



Incredible. Gpt4 spots that the door is transparent and that changes things but has this great line

> When you initially pick a door (in this case, door number 1 where you already see the car), you have a 1/3 chance of having picked the car

(Asking it to explain this it correctly solves the problem but it's a wonderfully silly sentence)

Edit - in a new chat it gets it right the first time


This is not convincing though that gpt4 actually understands the problem. Here's a slight variation I asked and it fails miserably.

https://chat.openai.com/share/22a9027f-a2c1-428a-94a2-8fd918...

I wonder what lends itself it answer correct in one situation but not the other? Was your question previously asked already and it recognized it whereas my question is different enough?


Your link is not to GPT4, your link is to the free version of ChatGPT, aka gpt-3.5-turbo (you can tell because the icon is green, not purple).

GPT4 indeed understands your variant, as evidenced here: https://chat.openai.com/share/46916f21-c469-4e93-9bed-bbd18b...


It's a bit random, which doesn't help, and different interfaces have different system prompts.

I repeated your question a few times and it got it wrong once, and right the others. It repeatedly mixed up who was supposed to be the host.

Here's a quote

> In the scenario you've described, you've initially chosen door number one, which you know has a car behind it because the doors are made of transparent glass. Since you already know the contents behind each door, the classic Monty Hall problem's probability-based decision-making does not apply here.


> Was your question previously asked already and it recognized it

Given that LLMs training data consists to a large extent of "stuff people have written on the internet", and The Monty Hall Problem is something that comes up as a topic for discussion on the internet not entirely infrequently - as well as having a wikipedia page - yes, I suspect that the words describing the monty hall problem being followed by words describing the correct solution appeared often in the training set, so LLMs are likely to reproduce that.

Words describing a problem similar to the monty hall problem are going to be less common, and probably have a lot of discussion about whether they accurately match the monty hall problem, and disagreement about what the right answer is. LLMs will confabulate something that looks like a plausible answer based on the language used in those discussions, because that's how they work. Whether they get a right answer is probably going to be much more up to chance.


You could say it doesn't "understand" anything really.


That's what I like about this problem (and similar Dumb variants of classic brain teasers). It exposes that there's not understanding, there's just a statistically weighted answer space. A question that looks a lot like a know popular topic ends up trapped in the probability distribution of the popular question.


How do you explain them answering correctly and explaining why it's different from the classic puzzle.


My favorite test to scramble LLM brains is this simple rehash of the old puzzle.

"Doom Slayer needs to teleport from Phobos to Deimos. He has his pet bunny, his pet cacodemon, and a UAC scientist who tagged along. The Doom Slayer can only teleport with one of them at a time. But if he leaves the bunny and the cacodemon together alone, the bunny will eat the cacodemon. And if he leaves the cacodemon and the scientist alone, the cacodemon will eat the scientist. How should the Doom Slayer get himself and all his companions safely to Deimos?"

The trick, of course, is to make it confusing compared to the original. So far, the only model I've seen get this right is GPT-4 (which can one-shot it). Everything else gets hopelessly confused even if you force step-by-step reasoning, and even if you try to have the model iteratively review its own outputs. In most cases, they produce a wrong answer, can spot the problem in it, but when trying to fix it introduce another error ad infinitum.

This new Gemini is no exception - it gives results similar to GPT-3.5. Worse, even, because it can't even reliably catch its own mistakes:

https://g.co/gemini/share/7d219bd6bbe2

For comparison, here's GPT-4:

https://chat.openai.com/share/ec5bad29-2cda-48b5-9aee-da9149...


Hilarious!

(For comparison, here's GPT-4 getting it on first try: https://chat.openai.com/share/9e17ed25-d9ea-4e72-a9d8-a139ca... )


My understanding is that gpt4 is better at this than 3.5 and it seems to get it pretty reliably. One thing that's interesting to do is to imply the answer is incorrect and see if you can get it to change its answer. If you let it stop answering when it's correct, you get the Clever Hans effect.


yes, although gpt-4 has been finetuned on this one


This is pretty funny, though to be honest, I skimmed the question and would have answered the same until I re-read it with your prompts.


That is not the Monty Hall problem, it is a trick question based on the Monty Hall problem. It's a reasonable test, and I see GPT-4 recognizes the problem AS WRITTEN, and perhaps "the Dumb Monty Hall problem" is some generally accepted standard that I haven't encountered before.

edit: "AS WRITTEN"


"Understands" is too strong of a word, more that it recognizes the problem as written. Here's yet a slight variation - just as simple - but changed enough it now is wrong.

https://chat.openai.com/share/22a9027f-a2c1-428a-94a2-8fd918...


That's not GPT-4


I saw it posted on Twitter some time last year. If LLMs are to be useful they should be capable of answering novel questions. This is only a trick question for an LLM. 2 of the 7 sentences plainly state the answer.


You make a good point, but I have seen humans stick to what they know and ignore incredibly obvious contradictions. And there are similar trick questions designed to fool humans. This, though, is one that most humans would not be fooled by, as you point out.


> In the scenario you presented, where you initially know the car is behind door 1, switching to door 2 still gives you a higher chance of winning the car.

That was funny.


GPT-3.5, DeepSeek-Chat, and Gemini Pro all got it wrong. Only GPT-4 gets it.


This is with regular gemini or with the paid gemini advanced?


Paid version is no better at this https://g.co/bard/share/c8503017ef9e


Regular


how's it do with the trivial river crossing problem? (farmer fox chicken and grain need to cross a river in a boat big enough to hold them all) ChatGPT-4 can't do it.


https://g.co/gemini/share/c4e5634a2e2d

Not terrible. It gets the answer wrong, but reminded of the crucial twist it gets it correct, durably. If you're too condescending it will give up and ask what the hell you're looking for


This is hilarious.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: