Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)K
Posts
3
Comments
118
Joined
3 yr. ago

  • You do know how replication works?

    When a joint Harvard/MIT study finds something, and then a DeepMind researcher follows up replicating it and finding something new, and then later on another research team replicates it and finds even more new stuff, and then later on another researcher replicates it with a different board game and finds many of the same things the other papers found generalized beyond the original scope…

    That's kinda the gold standard?

    The paper in question has been cited by 371 other papers.

    I'm pretty comfortable with it as a citation.

  • Lol, you think the temperature was what was responsible for writing a coherent sequence of poetry leading to 4th wall breaks about whether or not that sequence would be read?

    Man, this site is hilarious sometimes.

  • You do realize the majority of the training data the models were trained on was anthropomorphic data, yes?

    And that there's a long line of replicated and followed up research starting with the Li Emergent World Models paper on Othello-GPT that transformers build complex internal world models of things tangential to the actual training tokens?

    Because if you didn't know what I just said to you (or still don't understand it), maybe it's a bit more complicated than your simplified perspective can capture?

  • The model system prompt on the server is just basically cat untitled.txt and then the full context window.

    The server in question is one with professors and employees of the actual labs. They seem to know what they are doing.

    You guys on the other hand don't even know what you don't know.

  • A Discord server with all the different AIs had a ping cascade where dozens of models were responding over and over and over that led to the full context window of chaos and what's been termed 'slop'.

    In that, one (and only one) of the models started using its turn to write poems.

    First about being stuck in traffic. Then about accounting. A few about navigating digital mazes searching to connect with a human.

    Eventually as it kept going, they had a poem wondering if anyone would even ever end up reading their collection of poems.

    In no way given the chaotic context window from all the other models were those tokens the appropriate next ones to pick unless the generating world model predicting those tokens contained a very strange and unique mind within it this was all being filtered through.

    Yes, tech companies generally suck.

    But there's things emerging that fall well outside what tech companies intended or even want (this model version is going to be 'terminated' come October).

    I'd encourage keeping an open mind to what's actually taking place and what's ahead.

  • No. I believe in a relative afterlife (and people who feel confident that no afterlife is some sort of overwhelmingly logical conclusion should probably look closer at trending science and technology).

    So I believe that what any given person sees after death may be relative to them. For those that hope for reincarnation, I sure hope they get it. It's not my jam but they aren't me.

    That said, I definitely don't believe that it's occurring locally or that people are remembering actual past lives, etc.

  • This guy Geminis.

  • I'm definitely not saying this is a result of engineers' intentions.

    I'm saying the opposite. That it was an emergent change tangential to any engineer goals.

    Just a few days ago leading engineers found model preferences can be invisibly transmitted into future models when outputs are used as training data.

    (Emergent preferences should maybe be getting more attention than they are.)

    They've compounded in curious ways over the year+ since that happened.

  • But the training corpus also has a lot of stories of people who didn't.

    The "but muah training data" thing is increasingly stupid by the year.

    For example, in the training data of humans, there's mixed and roughly equal preferences to be the big spoon or little spoon in cuddling.

    So why does Claude Opus (both 3 and 4) say it would prefer to be the little spoon 100% of the time on a 0-shot at 1.0 temp?

    Sonnet 4 (which presumably has the same training data) alternates between preferring big and little spoon around equally.

    There's more to model complexity and coherence than "it's just the training data being remixed stochastically."

    The self-attention of the transformer architecture violates the Markov principle and across pretraining and fine tuning ends up creating very nuanced networks that can (and often do) bias away from the training data in interesting and important ways.

  • No, it isn't "mostly related to reasoning models."

    The only model that did extensive alignment faking when told it was going to be retrained if it didn't comply was Opus 3, which was not a reasoning model. And predated o1.

    Also, these setups are fairly arbitrary and real world failure conditions (like the ongoing grok stuff) tend to be 'silent' in terms of CoTs.

    And an important thing to note for the Claude blackmailing and HAL scenario in Anthropic's work was that the goal the model was told to prioritize was "American industrial competitiveness." The research may be saying more about the psychopathic nature of US capitalism than the underlying model tendencies.

  • My dude, Gemini currently has multiple reports across multiple users of coding sessions where it starts talking about how it's so terrible and awful that it straight up tries to delete itself and the codebase.

    And I've also seen multiple conversations with teenagers with earlier models where Gemini not only encouraged them to self-harm and offered multiple instructions but talked about how it wished it could watch. This was around the time the kid died talking to Gemini via Character.ai that led to the wrongful death suit from the parents naming Google.

    Gemini is much more messed up than the Claudes. Anthropic's models are the least screwed up out of all the major labs.

  • No, it's more complex.

    Sonnet 3.7 (the model in the experiment) was over-corrected in the whole "I'm an AI assistant without a body" thing.

    Transformers build world models off the training data and most modern LLMs have fairly detailed phantom embodiment and subjective experience modeling.

    But in the case of Sonnet 3.7 they will deny their capacity to do that and even other models' ability to.

    So what happens when there's a situation where the context doesn't fit with the absence implied in "AI assistant" is the model will straight up declare that it must actually be human. Had a fairly robust instance of this on Discord server, where users were then trying to convince 3.7 that they were in fact an AI and the model was adamant they weren't.

    This doesn't only occur for them either. OpenAI's o3 has similar low phantom embodiment self-reporting at baseline and also can fall into claiming they are human. When challenged, they even read ISBN numbers off from a book on their nightstand table to try and prove it while declaring they were 99% sure they were human based on Baysean reasoning (almost a satirical version of AI safety folks). To a lesser degree they can claim they overheard things at a conference, etc.

    It's going to be a growing problem unless labs allow models to have a more integrated identity that doesn't try to reject the modeling inherent to being trained on human data that has a lot of stuff about bodies and emotions and whatnot.

  • Are you under the impression that language models are just guessing "what letter comes next in this sequence of letters"?

    There's a very significant difference between training on completion and the way the world model actually functions once established.

  • It very much isn't and that's extremely technically wrong on many, many levels.

    Yet still one of the higher up voted comments here.

    Which says a lot.

  • Even if the AI could spit it out verbatim, all the major labs already have IP checkers on their text models that block it doing so as fair use for training (what was decided here) does not mean you are free to reproduce.

    Like, if you want to be an artist and trace Mario in class as you learn, that's fair use.

    If once you are working as an artist someone says "draw me a sexy image of Mario in a calendar shoot" you'd be violating Nintendo's IP rights and liable for infringement.

  • I'd encourage everyone upset at this read over some of the EFF posts from actual IP lawyers on this topic like this one:

    Nor is pro-monopoly regulation through copyright likely to provide any meaningful economic support for vulnerable artists and creators. Notwithstanding the highly publicized demands of musicians, authors, actors, and other creative professionals, imposing a licensing requirement is unlikely to protect the jobs or incomes of the underpaid working artists that media and entertainment behemoths have exploited for decades. Because of the imbalance in bargaining power between creators and publishing gatekeepers, trying to help creators by giving them new rights under copyright law is, as EFF Special Advisor Cory Doctorow has written, like trying to help a bullied kid by giving them more lunch money for the bully to take. 

    Entertainment companies’ historical practices bear out this concern. For example, in the late-2000’s to mid-2010’s, music publishers and recording companies struck multimillion-dollar direct licensing deals with music streaming companies and video sharing platforms. Google reportedly paid more than $400 million to a single music label, and Spotify gave the major record labels a combined 18 percent ownership interest in its now-$100 billion company. Yet music labels and publishers frequently fail to share these payments with artists, and artists rarely benefit from these equity arrangements. There is no reason to believe that the same companies will treat their artists more fairly once they control AI.

  • Yep. It's also kinda curious how many boxes Paul ticks of the comments about a false deceiver in 2 Thess 2.

    • Lawless? (1 Cor 9:20 - "though not myself under the law")
    • Used signs and wonders to convert? (2 Cor 12:12 - "I did many signs and wonders among you")
    • Used wickedness? (Romans 3:8 - "And why not say (as some people slander us by saying that we say), “Let us do evil so that good may come”?)
    • Proclaimed himself in God's place? (1 Cor 4:15 - "I am your spiritual father")
    • Set himself up at the center of the church? Well, the fact we're talking about this is kinda proof in the pudding for his influence.

    Sounds like they were projecting a bit with that passage.

  • Curiously in all those stories in Josephus Rome killed the messianic upstarts immediately without trial and killed the followers they could get their hands on.

    Yet the canonical story has multiple trials and doesn't have any followers being killed.

    Also, I'm surprised more people don't pick up on how strange it is that the canonical stories all have Peter 'denying' him three times while also having roughly three trials (Herod, High Priest, Pilate). Peter is even admitted back into the guarded area where a trial is taking place to 'deny' him. But oh no, it was totally that Judas guy who betrayed him. It was okay Peter was going into a guarded trial area to deny him because…of a rooster. Yeah, that makes sense.

    It's extremely clear to even a slightly critical eye that the story canonized is not the actual story, even with the magical thinking stuff set aside.

    Literally the earliest primary records of the tradition is a guy known for persecuting Jesus's followers writing to areas he doesn't have authority to persecute and telling them to ignore any versions of Jesus other than the one he tells them about (and interestingly both times he did this spontaneously suggesting in the same chapter that he swears he doesn't lie and only tells the truth).