Hacker Newsnew | past | comments | ask | show | jobs | submit | tpkm's commentslogin

The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation


This is not true. They claim not to train by default for business and enterprise agreements, but for plus and pro plans they enable it by default and you can allegedly turn it off (I don’t trust them very much though, I’m sure there is something in the T&C saying they can modify that deal any time)


Absolutely 100% not true. I have colleagues in 4 "FAANG adjacent" companies plus the one I work at - zero of them have any faith in Enterprise Agreements from either OpenAI or Anthropic.

There's a reason why people spend more $$$ with Data Bricks, Palantir, AWS Bedrock etc.. and don't even consider using Anthropic or OpenAI APIs directly - it's because those guarantees provide very little in the way of data-discovery, audit requirements, or liquidated damages should it ever be discovered there was data leakage.

At least with these other companies, while the LD is likewise not great (typically limited to the amount of money you paid them) - you at least have some data-governance guarantees around running on dedicated hardware - no multi-tenancy, no third-party access outside of the AWS operators who keep the HW running - but are very much not in the business of looking at your data.

I think this is mostly a function of what's at risk - when company valuations get into the 10s of billions of dollars, the risk of IP leaking into what could be seen as competitive companies (OpenAI/Anthropic would be happy to take over the world - I don't sense that AWS or Azure, are as ruthless in stepping on their customers business, unless of course they are a SAAS provider) is just too significant a liability to take - particularly when you can de-risk.


That's interesting - I just tried on my Ryzen 5 3600 and got the following:

10k using 48 processes: 0.82 seconds (using the original single threaded script this was actually faster at 0.65)

100k using 48 processes: 28.41s

200k using 48 processes: 99.68s

--- EDIT - looking at resource monitor python.exe is only using 7-8% of total available CPU resources

--- EDIT 2 - switching to ThreadPool from multiprocessing.dummy brought the 10k result down from 0.8 seconds to 0.3 seconds, but didn't impact the 100k or 200k results


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: