Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)G
Posts
93
Comments
1020
Joined
3 yr. ago

  • Stupid human could've said thank you.

  • But the article later does back it up

    The CEO of Cloudflare did not assert that. I was surprised that he would claim such a thing, and that should have made me read more carefully. Elon Musk notwithstanding, neither incompetence nor conspiracy theorizing are common at that level, publicly anyway.

    You can believe whatever you like, of course. Freedom of opinion is nothing if not the right to be wrong.

  • It would be a lot to write, if you had to say what something does not do rather than what it does.

    I looked at what the Cloudflare CEO said again. To be fair to him, he is not actually backing you up. He's saying that Google makes no difference between the AI overview and the other search results. That is true. The AI overview is a search feature. I'm not sure why someone would want their link listed in search but not appear much more prominently in the AI overview.

  • You look up what Googlebot does. No AI.

    You want to know what crawlers do AI? Just search for "AI", or "training", or some such, or skim through. It's not long. Google-Extended collects training data. Note that Google-Extended is explicitly not used to rank pages.

    Did that help?

  • I'm not really sure what you are asking here. Did you notice that you can scroll down and see a list of their crawlers?

  • that is not how general news media has been talking about robots.txt.

    Ahh, yes. I think there is a lesson there.

  • Ok. That quotes a tweet by Cloudflare's CEO. IDK what his qualifications are, but his conflict of interest is obvious enough. Real quality journalism there.

    ETA: I looked at what the Cloudflare CEO said again. To be fair to him, he is not actually claiming that Googlebot collects AI training data. He's talking about the AI overview, which is a search feature. The data for search features is collected by Googlebot. I'm not sure why someone would want their link listed in search but not appear much more prominently in the AI overview.

    Here's Google technical documentation on its crawlers: https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers

  • That's very different from what I called false.

    What you describe may happen, but probably not as much as you think. Much of that stuff is just not that valuable. Some personal, colloquial writing is necessary, but Google already pays Reddit. Other stuff is better obtained from torrents or shadow libraries like Anna's Archive.

  • Googlebot if enabled won’t just list you for search, but will also scrape your contents for Google’s AI.

    False.

  • What did he think a crawler is? Why was he surprised that not allowing companies to use his data lead to them not using his data? Looks like he has another surprise coming when he notices that search engines no longer index his blog.

  • There is a lot of disinformation being spread on copyright because major rights-holders hope to gain a lot of money for nothing.

    US fair use has always worked like this. Other countries without fair use had to make laws to enable AI training. I know about Japan and the EU.

    It is precisely because of these new laws that AI training in the EU is possible at all (eg by Mistral AI or by various universities/research institutions). But because of lobbying by rights-holders, this is quite limited. It's not possible to train AIs in the EU that are as capable as those from the US, where Fair Use comes directly from the constitution and can't be easily lobbied aside by monied interests.

  • I don't see how that would be fair use or what the argument is supposed to be.

    Let me warn you that Lemmy is full of disinformation on copyright. If you picked the idea up here, then it probably is absolutely bonkers.

    In any case, fair use is a US thing. In the EU, it would still be yoink.

  • To save everyone a click: It's a non-commercial license (with a very rude yoink clause, if anyone is foolish enough to build something on it.)

    By the by, there's a good chance that AI models are not copyrightable under US law; making the license moot in the US. In other regions, such as the EU, it likely holds.

    3.3 Use Limitation. The Work and any derivative works thereof only may be used or intended for use non-commercially. Notwithstanding the foregoing, NVIDIA Corporation and its affiliates may use the Work and any derivative works commercially. As used herein, “non-commercially” means for non-commercial research and educational purposes only.

  • He was arrested in September 2022 on allegations of bribing an executive related to the Tokyo Olympics in exchange for KADOKAWA receiving preferential sponsorship treatment.

    He was later charged by prosecutors and stepped down as chairman of the company on October 4, 2022. He denies the charge. KADOKAWA’s current CEO, Takeshi Natsuno, confirmed as recently as March 2025 that Tsuguhiko is barred from meeting with him and is not involved in the company (Toyo Keizai). Despite Tsuguhiko’s lack of involvement with KADOKAWA, which is active in conventional and short anime while also “actively investing” in AI for production, his words underscore a growing trend.

    From the article.

  • reason based robots

    What's that?

  • This is mandated by UK law. If you created a node so that UK users can bypass this, you would be doing something illegal. You'd probably get defederated.

  • smh

    That guy should be happy that no AI will ever be trained on their work. It's ok to contribute to progress, but only if it's progress the cool kids approve of. Know your place, nerds.

  • I didn't need that mental image. But since you have installed it in my head, I am honor-bound to upvote.

    sigh