Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)U
Posts
8
Comments
79
Joined
1 yr. ago

  • Don’t worry peeps, it’s just spicy autocomplete

  • Nothing can disprove that to that crowd, I’m afraid.

  • It was supposed to have no internet access, but the config was wrong. The report goes on to say that Opus and Mythos then proceeded on the premise that everything was a simulation, while the unnamed stronger model concluded after a while that it had real internet access and stopped the attack.

  • From what I've seen of AI autonomous capabilities, Occam would land on "AI did the hack", I think

  • ITT: jet fuel AI can't melt steel beams hack anything.

    Also I want to know the name of this company so I can avoid them:

    Claude believed the package registry it was using to be part of the simulation, but in reality the package was made freely available online for roughly one hour. During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access further infrastructure from this company.

    ETA: I thought I posted this top-level, not my intention to single out this comment specifically.

  • Okay - I don’t believe that, since there’s too much released detail. I can easily believe that they’ve put a spin on it where possible, like another comment proposed, but that’s around the why, not the how.

  • Yeah, this is the bit where it’s not hard to believe marketing would polish the narrative, at least if they can’t be caught in an outright lie.

  • It might be - but which parts? Do you suspect that huggingface and openai made the entire thing up? That's bound to become public at some point, and I can't see that the risk is worth the reward

  • Technology @lemmy.world

    Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

    huggingface.co /blog/agent-intrusion-technical-timeline
  • John Searle is the GOAT?

  • Because when you're old enough to remember what AIM chat it's could do 25 years ago, it stops being impressive what today's chatbots can do...

    C’mon, that’s just silly.

  • Technology @lemmy.world

    When AI builds itself

    www.anthropic.com /institute/recursive-self-improvement
  • Technology @lemmy.world

    the solution might be cancelling my AI subscription

    thoughts.hmmz.org /2026-05-31.html
  • I feel like it gets more intrinsically interesting the better it gets, even when the initial shock has faded a bit, but tastes vary of course.

    The LLM creators won’t shut up about what we can use it for and why. Some of those use cases actually work fairly well, like coding, so that part doesn’t really trigger any alarms.

    What I don’t see is how they intend to make actual money when open weight models catch up in the next months, but if we can lose the frontier labs and keep the current abilities available that’s fine by me (apart from the whole “possible collapse of Western economy”, that is)

  • State-of-the-art models rely on late-1800s and early-1900s print books for high-quality training data, and those books use ~30% more em-dashes than contemporary English prose. That’s why it’s so hard to get models to stop using em-dashes: because they learned English from texts that were full of them.

    That sounds really plausible -- I associate the em-dash with old books and stilted prose, like Sherlock Holmes stories

  • Yeah, I've had that existential crisis this Spring, and so have other devs I know. There's still a good way to go, but unless LLMs hit a hard limit on cognition I tend to share the author's feelings.

  • I’m sorry, but I don’t agree with your first point at all. Things can have negative sides and still be interesting.

    The Turing test, as I interpret it at least, is more of a philosophical than a technical thing, trying to provide a way to evaluate the thinking ability of someone or -thing without being able to look at its innards. I’ve always found it fascinating, but I can understand if people disagree (just don’t drag the Chinese room into it). However, if you don’t think a conversation with Claude is more interesting than a faux psychiatrist session with ELIZA, I don’t know where we could go from there 🤷

  • That remains to be seen. If open weight models get good enough and efficient enough, I don’t see what moat Anthropic, OpenAI, et al have. Maybe they can fade away so we can buy hardware again, and still reap the benefits of Turing-certified “autocomplete”.

  • Pricing is an issue, yes - the open-weight models aren't on par with claude and codex yet. I have hopes that six months to a year can bring them to the level of current frontier models, and if so I think that's probably good enough for most users, including me. How Anthropic and OpenAI intend to make money at that point, I couldn't tell you, but I don't see an actual downside there :)

  • To begin with, I wouldn't say I'm an enthusiast, but I do find the breakthroughs in LLM tech the recent years to be interesting. I sometimes wonder how we got so blasé that a computer acing the Turing test is passed off as "spicy autocomplete, ho hum".

    I also think you'll find that many people on Lemmy do hate AI to a worrying degree. Just look at the reception this and other posts about it get here, in a technology community, where you'd expect news about one of the most sci-fi-like (to me, at least) technologies to be welcome.

    To the rest of your comment, I must say I find it strange to come to this community and complain that you find news about LLMs (a technology) useful for coding (also a technology), arguing that it's not interesting to you. To each their own, I suppose.

  • Gpt 5.4 xhigh isn’t too bad for automated reviews and the like, and 5.5 is fairly efficient for interactive coding. I prefer those to Claude and opus, the Anthropic models feel like they’re trying to hard to be human to me, but that’s personal preference I guess.

    Yeah, it’s not free (or the free models aren’t good enough), but the consensus at work is that this is a potential game changer, and we need to experiment to see what works and what doesn’t. So, the budget is there until things settle, and afterwards if things work out.

  • I know no one here wants to hear it, but the newest models from Anthropic and OpenAI are not bad coders with proper direction. If used correctly they can be positive force multipliers for developers, and used incorrectly they can do a lot of damage.

    Note that this goes for developers with some experience. If you try to use an LLM in place of experience, or use it as a shortcut to try to gain experience, it turns into a negative multiplier really quickly, and you probably build bad habits that are hard to kick.

    I’m not sure what the future of coding looks like, but I’ll be very surprised if AI in its current or a future incarnation is not involved somehow. How to learn coding correctly for that I don’t know, but looking at the junior devs I know, I am sure they will figure it out and grow into AI-native senior devs in due time.

  • Technology @lemmy.world

    Introducing Claude Opus 4.8

    www.anthropic.com /news/claude-opus-4-8
  • Technology @lemmy.world

    The pressure

    daniel.haxx.se /blog/2026/05/26/the-pressure/
  • Technology @lemmy.world

    FTC to Require Cox Media Group, Two Other Firms to Pay Nearly $1 Million to Settle Charges They Deceived Customers About “Active Listening” AI-Powered Marketing Service

    www.ftc.gov /news-events/news/press-releases/2026/05/ftc-require-cox-media-group-two-other-firms-pay-nearly-1-million-settle-charges-they-deceived
  • Technology @lemmy.world

    An open-source spec for Codex orchestration: Symphony.

    openai.com /index/open-source-codex-orchestration-symphony/
  • Technology @lemmy.world

    Significant raise of reports

    lwn.net /Articles/1065620/