OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole From Us

ForgottenFlux@lemmy.world · 4 days ago

OpenAI Furious DeepSeek Might Have Stolen All the Data OpenAI Stole From Us

mechoman444@lemmy.world · 3 days ago

The core infrastructure issue is distinguishing between queries made by individuals and those made by programs scraping the internet for AI training data. The answer is that you can’t. The way data is presented online makes such differentiation impossible.

Either all data must be placed behind a paywall, or none of it should be. Selective restriction is impractical. Copyright is not the central issue, as AI models do not claim ownership of the data they train on.

If information is freely accessible to everyone, then by definition, it is free to be viewed, queried, and utilized by any application. The copyrighted material used in AI training is not being stored verbatim—it is being learned.

In the same way, an artist drawing inspiration from Michelangelo or Raphael does not need to compensate their estates. They are not copying the work but rather learning from it and creating something new.

Lifter@discuss.tchncs.de · 11 hours ago

I disagree. Machines aren’t “learning”. You are anthropomorphising theem. They are storing the original works, just in a very convoluted way which makes it hard to know which works were used when generating a new one.

I tend to see it as they used “all the works” they trained on.

For the sake of argument, assume I could make an “AI” mesh together images but then only train it on two famous works of art. It would spit out a split screen of half the first one to the left and half of the other to the right. This would clearly be recognized as copying the original works but it would be a “new piece of art”, right?

What if we add more images? At some point it would just be a jumbled mess, but still consist wholly of copies of original art. It would just be harder to demonstrate.

Morally - not practically - is the sophistication of the AI in jumbling the images together really what should constitute fair use?

mechoman444@lemmy.world · 11 hours ago

That’s literally not remotely what llms are doing.

And they most certainly do learn in the common sense of the term. They even use neural nets which mimic the way neurons function in the brain.

Lifter@discuss.tchncs.de · 10 hours ago

Mimic, perhaps inspired but neural nets in machine learning doesn’t work at all like real neural nets. They are just variables in a huge matrix multiplication.

FYI, I do have a Master’s degree in Machine Learning.

mechoman444@lemmy.world · 2 hours ago

Yes I also have a master’s and a PhD in machine learning as well which automatically qualifies me as an authority figure.

And I can clearly say that you are wrong.