Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)H
Posts
6
Comments
1506
Joined
5 yr. ago

A software developer and Linux nerd, living in Germany. I'm usually a chill dude but my online persona doesn't always reflect my true personality. Take what I say with a grain of salt, I usually try to be nice and give good advice, though.

I'm into Free Software, selfhosting, microcontrollers and electronics, freedom, privacy and the usual stuff. And a few select other random things as well.

MRZ LCK 00

  • I don't think this is an easy debate to have. Even those kinds of models have effects on society, on people, on the environment. A lot of resources are needed for research and training. Same for inference even if you own the computer. And we'd need to agree on what's ethical. Usually that means treat the entire world like a self-service outlet and just take everything unless someone calls you and explicitly said no. At that point it'll become a nasty technical problem because there's no easy way to get rid of data after the fact.

    Models could in theory be ethically sourced, but used for evil. Which is gonna be a problem to society. And then all of the debate is hypothetical because that's not even what's in use out there...

    My post was more concerned with the facts. What kinds of models we do have. Not what we should do or an universally accepted definition of ethical.

  • What would be an application for this? More transparency for court cases? Or enable people to search for an alternative source for the data, like a torrent?

  • Sudoku and Minesweeper could very well be some internal project names and have nothing to do with what the words mean. I have no clue.

    Seems to me their announcements ooze with marketing speech. I bet they put it all in one sentence to mislead about the datasets. Put a comma there and omitted the word "partly". It's technically borderline correct. I'll have to read up on it, whether that's a good amount of data, or just some tiny fraction. It's not looking good, though.

  • Hmmh. I sometimes struggle to emphasize. I don't live in the US. Culture and politics are a bit different here. We certainly have those dynamics as well. But I don't think we're (yet) at a majority of the population living in constant fear. We Germans tend to be less open towards new technology, though. For other reasons. We always sell our souls to American big tech eventually. But it's always met with some initial reluctance/backlash.

    I'm not sure about your point about fascism. I mean there's some inflation with how people use the word... But proper fascism is really, really bad. All my grandparents had to suffer through it. Somehow they survived. But it cost like 70 million lives. And most cities here were turned into burning piles of trash. You kinda want to avoid it. Even at a substantial cost. And my place isn't the only example. We have plenty wars and fascists on earth. Constantly. Fear isn't the right tool to avoid it, though. Being clever and education are good strategies.

    Ultimately it's certainly hard to judge some hypothetical future. I still don't see any reason to completely sell out to some specific big companies. I mean they sometimes do come up with good products. Sometimes their path to profit is scattered with dead bodies. Could be anything unless there's accountability and regulation. And it's not either/or... We could as well have some less nasty company do the research. Or do it a bit more slowly... Maybe that'll take 10 years longer but we get less people killed and less permanent harm done to the environment. There's not just one path forwards.

    I fully agree on the atmosphere and sentiment,. It's really hard to talk about AI. Especially online. Everyone is yelling and doing identity politics. Most people don't have a clue how it works. And/or they believe in some falsehoods. I think the best way is to just tell it how it is. Quite a lot of people won't listen. But they're not listening anyways.

    Be a bit careful with the evil companies, though. I don't think Meta, Anthropic, OpenAI, Elon Musk or the Chinese government act in your, or my best interest. I'd advise not to be naïve about that. They always tell nice stories. Promise all kinds of stuff. Politicians do the same thing. It's their job. It doesn't have to come true, though, by any means. In fact the more they promise, the more likely it is they're trying to manipulate you.

    This isn't really connected to the technology and science though. Whether AI is going to succeed and improve from a scientific perspective is yet another can of worms. I still strongly doubt LLMs are going to be the direct predecessor to strong AI or something properly "clever", but we'll see about that.

  • Huh. It doesn't come with all the datasets, though. Does it? The model card contains a long list of "Private Non-publicly Accessible Datasets", both by NVIDIA and by third parties. And some of the stuff just says "Undisclosed".

  • I know I'm way too nitpicky... But I feel I should point out there's some confusion hidden here...

    There's stuff which is moral to do. And stuff which is legal to do. Those aren't the same! When talking about who can mention what, what copyright law demands people to do and those blurry lines, we're concerned with legality. That doesn't have anything to do with ethics. (At least not directly.)

  • Thanks! Yes, definitely deserves a mention.

  • Uh, that one looks nice. Thanks!

  • Yeah, good question! I bet nobody even knows their names. I had hoped to get a reply by @popcar2@programming.dev or @sims@lemmy.ml who seem to know plenty of them?!

    I was talking about the RedPajama project who recreated the dataset of Meta's LLaMA and proceedingly trained a model on it. And Apertus by the ETH Zürich. They also factored in ethics when compiling the dataset.

    I made a post asking on !localllama: https://programming.dev/post/55905594In case anyone knows about the existence of such models, please educate me over there!

  • Hmmh. I'm really unsure about how to compare timeframes. Those were medieval times. Everything was slow to impossible. Nowadays I can communicate information from here to America within a fraction of a second. Back then they weren't even aware of the Americas. And there wasn't any hardware store to buy the material needed. They had to invent the oil-based paint, have a bunch of workers dig up the metal, cut down trees, carpenters and goldsmiths... Nonetheless within 40 years they went from invention to 110 printing presses all across Europe. I think that's crazy fast given the situation. And it really took off after that. Not long after we got it to carry the Reformation. And some century later the Scientific Revolution. (And none of those we're good things for the kings.)

    And 3 years ago is an arbitrary point in time. Neural networks were invented 70 years ago. It's been 50 years since the first AI hype with expert systems. 40 years since the second AI winter. And now 4 years after the breakthrough of LLMs. It's really difficult to compare the two. And I personally don't think LLMs are anywhere close to proper AI. They're great and all. But also fairly limited in what they can do. For example they can't update their weights while running. So there's no learning nor adapting to the environment. Even remembering something isn't really possible and has to be added on top. And I wonder if we can make them reliable enough for some tasks. Seems impossible to me with current tech. We still need some time to develop the first customer service chatbot which isn't vulnerable to prompt injection or which doesn't on occasion "hallucinate" terms and services and procedures. And just lie to the customers.

    I don't really get what you mean by fear. It's certainly a thing. But I don't think it is mainstream complaint #1 anymore. I come from a culture where we often meet new things with a healthy(?) dose of skepticism. I don't see anything wrong with that. And a lot of the criticism regarding AI isn't concerned with fear. It's about Elon Musk putting up 6 dirty (and slightly illegal) gas power plants next to your backyard. Sam Altman polluting the environment and making electricity, computer chips and everything more expensive for you. And some idiots starting an investment bubble that requires AI to displace 30% of the human workforce by 203x or it will pop and have catastrophic effects on the stock exchange, including your retirement fund. I mean I'm positive we'll manage to get by in some form... But I think it's natural to feel a lot of rejection towards all of that.

  • Some day I'm gonna need someone to tell me the name of such an LLM. I'm not aware of any full-libre model which is able to code on an acceptable level. I don't think they exist in this world.

  • I missed AerynOS. Thanks! Let's hope you're going to find another one. Given there's lots and lots of people who don't appreciate AI written code... I feel there has to be a market for a mainstream Linux distro with a no AI policy.

  • By the way, there's like two open-source LLMs, which are large enough to be usable in some form. And one of them claims not to contain pirated training data. It's pretty much the opposite of "plenty".

  • Yes. And it goes further than that: Maintainability should also be a concern to maintainers. Doing good in the world is what pretty much all Free Software community projects are trying to achieve. They have to deal with licensing issues... It's complicated.

  • That remains to be seen... As I said, they earned my trust over many years. Maybe they're able to tackle it, as they faced other challenges. I mean you're free to throw away democracy on a whim. I think it needs to be addressed and tackled. Votings have turned out with bad results before. It's ugly. But I'm still subscribed to the ideology. My red line is something bad actually happening not just on paper and in some hypothetical future. Until then, there's plenty time to repent. I still have my hopes up. They're good people and have been doing a phenomenal job until now. I don't think I want to ignore that. But I'm not stopping you from having some alternative plan ready. 😅 That's not a bad idea either. All I'm saying is: Don't borrow trouble. (And "Don’t throw the baby out with the bathwater." Democracy might indeed be a good thing, despite it's downsides.)

  • I think democracy is a good thing. Full stop.

    It ain't easy though. Not even for the Debian community. But what's the alternative? Some corpo distro? They're going to adopt AI as well. And benevolent dictator might be alright... Yet, I personally still prefer a democracy. Even if it's difficult. Sometimes voting results are sub par. But we need to deal with that, not drop it once the first difficulty arises?! What kind of strategy is that? I don't think anyone is getting anywhere if they just keep dropping everything once they're faced with problems.

  • Yeah. But my point is: They're not utilizing those tools. At least not yet.

    Looking back at my Linux experience, I have a lot of trust in the Debian community. Including to do the right thing. Even if it's difficult. They don't always get everything right. But they're pretty alright. I wouldn't just throw it all in the bin because of some hypotheticals. And democracy isn't easy. We judge people by what they do. (At least I do.)

  • And that's okay from your perspective?