How much does it bother you that OpenAI is trained on your data? What can we do about it?

duncesplayed@lemmy.one · edit-2 1 year ago

How much does it bother you that OpenAI is trained on your data? What can we do about it?

duncesplayed@lemmy.one · 1 year ago

I have a similar kind of idea. I think if it had been a free/open source/community project that made the headlines I would have been all like “this is so awesome”.

I guess what I don’t like is the economic system that makes that impractical. In order to build one of those giant GPTs, you need tonnes of hardware (capital), so the community projects are always going to be playing catchup, and I think quite serious catchup in this arena. So the economic system requires that instead of our posts going to a “collective hive mind” that aid human knowledge, they go to some walled garden owned by OpenAI, which filters and controls it for us, and gives us little bits of access to our own data, as long as we used it only in approved ways (i.e., ways that benefit them).

tst123@lemmy.world · edit-2 1 year ago

deleted by creator

IDe@lemmy.one · edit-2 1 year ago

Most of the data used in training GPT4 has been gathered through open initiatives like Wikipedia and CommonCrawl. Both are freely accessible by anyone. As for building datasets and models, there are many non-profits like LAION and EleutherAI involved that release their models for free for others to iterate on.

While actually running the larger models at a reasonable scale will always require expensive computational resources, you really only need to do the expensive base model training once. So the cost is not nearly as expensive as one might first think.

Any headstart OpenAI may have gotten is quickly diminishing, and it’s not like they actually have any super secret sauce behind the scenes. The situation is nowhere as bleak as you make it sound.

Fighting against the use of publicly accessible data is ultimately as self-sabotaging ludditism as fighting against encryption.