No idea. The whole copyright topic is a clusterfuck when it comes to LLMs. The matter of the fact is: most AI companies steal data to train their models and they don't care about licenses or copyrights.
Who the fuck gives Anthropic the right to use all that training data for their model?! And if other companies distill their models, THEY are to be protected by laws?! Cry me a fucking river.
I think it’s important to note that these models are open-WEIGHT and not open-SOURCE. Open-source would mean we could explore the training data itself. Open-weight models are great, but please let’s not call them open-source.
What the fuck? I can’t believe this headline. Does anyone have this article without paywall?