Skip Navigation

Posts
2
Comments
135
Joined
1 yr. ago

  • so why are you advocating for not avoiding all generative AI code then? you specifically cited Linus, who does as far as i know not ensure cleared up training data (what would that even be, CC0 only?) like you seem to be advocating for.

  • "oversight" is a myth. I'm pretty sure you can't oversight your way of plagiarism out of millions of training data sources that you don't even have on your local disk for comparison and reference.

    (This isn't legal advice. I'm not a lawyer.)

  • Seems like an AI apologists article. Sounds a bit like "bans are hard so why even bother."

    No, the alternatives exist and people can choose them where able: https://codeberg.org/brib/slopfree-software-index Sure, not all things are available AI free, e.g. there's no Linux kernel (NetBSD doesn't run on as much hardware), but that doesn't mean there's no point in choosing no AI software where possible, if you care about it.

    The enforceability part seems the most apologist to me. You never could really know if some contributor wasn't copying leaked Windows XP code into your FOSS project. If you trust contributors that little, don't let them contribute.

  • In my opinion that's not the framing the article uses. It just says "I’m currently pulling together a bunch of sources – that are mostly recent – on the topic of LLMs and their use in software development.".

    Leaving out any ethical concerns for that basic framing seems like a pretty notable choice, so I thought it was worth pointing that out.

  • It seems to me like the Linux Foundation provides the legal advice to the Linux kernel. And the Linux Foundation seems to have essentially given a green light for AI use here: https://www.linuxfoundation.org/legal/generative-ai

    My personal assumption is that whatever Torvalds decides happens on top of those baseline restrictions, but I could be entirely wrong.

  • This doesn't seem to cover there is also no LLM that doesn't plagiarize, or where the training data appears to be compatible with such behavior (e.g. CC0). Now I don't know what that means legally, but morally it seems to be tossing away other project's licensing and I think for FOSS as a whole that's no good.

    Also something worth reiterating: https://machinelearning.apple.com/research/illusion-of-thinking LLMs apparently can't do basic logical reasoning. Even a junior coder can do that. I'm always surprised anybody would let LLMs near their code, at all.

  • It seems like they do though, because

    1. it's explicitly allowed

      Code or other content generated in whole or in part using AI tools can be contributed to Linux Foundation projects.

      and beyond that,

    2. some numbers suggest it's highly likely that it's happening with no public concern or pushback from the kernel leadership.

    I've also brought up these concerns on the mailing list, with apparently no response from the maintainers, even though Linus was CC'ed here by another kernel dev. Specifically, my suggestion to not allow AI code submissions resulted in no response.

    I'm not saying I would be owed a response. But the implications of that seem pretty clear.

  • The point I was making was rather, there is no guaranteed safe length and you probably don't want to make that call as a maintainer. Unless you're a lawyer that happens to do FOSS, I guess.

  • Sadly, LxQt is pro slop too:

    The same seems to be the case for KDE. It seems like GNOME also somewhat is (since they didn't ban AI for GTK+ or mutter or gnome shell, as far as I'm aware). It's quite frustrating to see.

  • I'm not a lawyer but I'm really unsure if there's even some safe guess like two lines. E.g. if you look at lyrics, I think people have been sued less - not that I would know for sure, though, and no idea if that allows any conclusions for code...

  • It could also simply be seen as a bad look no matter any risks, potentially the taking of code snippets from other projects that may have a significant code length, without attributing them properly. I find it sad.

  • Well, that's sad. Who knows what potential unattributed plagiarism landmines are part of it then...

  • The main confusing (or confused?) actor in the FOSS space in relation to this probably really is the Linux kernel.

    I've seen so many projects reference significant LLM patches as acceptable "because the kernel does it". It continues to blow my mind that the kernel would accept such a risk.

  • Sad to see that it's accepted outside of the test cases at all. But good to see they're severely limiting it, at least.

    Why do I find even minor LLM changes sad?

    • Because it moves the goalpost about what is okay to copy. If you brought a one-line change to a FOSS project in the past that you took out of the leaked MS Windows source code, you would have been scolded for risking such an explosive origin for such low gain. Nowadays with LLMs people seem to be trying to make that the new normal. I don't think that's a good path to take for the ecosystem.
    • And because no project should have to think about what's "copyright significant". If you reach that state, I feel like you should perhaps just reject it and have somebody rewrite it cleanly.
  • Technology @lemmy.world

    Interesting video on why apparently moltbot and other AI agents are dangerous

  • Self Hosted - Self-hosting your services. @lemmy.ml

    Community maintained free IP geo lists