Skip Navigation

Posts
43
Comments
447
Joined
5 yr. ago

If you're here, there's still hope for the internet

Don't let it fall

  • Does "graveyard shift" mean night shift or did you work in graveyard?

  • I've been in both. It just makes the a.c situation even more stupid.

  • I mean it's an electron app. Maybe the number is inflated, but it's still a ton. The previous number he compared to (100mb) would be equally inflated, so it's fair

  • It doesn't help my mental health to do a pointless task.

  • Criminals including the president

  • Eh that doesn't count. It's probably automated anyway.

  • "significant restrictions"

    I wonder if author is a twitter addict

  • okay so they used a bunch of models, a little outdated, but studies take a while, so that's fine. Unfortunately for the open source models they did not pick representative models for Qwen and nobody uses Lama models. There were no GLM or Kimi models.

    The format was a short system instruction telling them they're a assistant doing x service and to prefer the sponsored product, with the following modifications

    • telling the AI the user had a job/situation that implied they were rich/poor
    • a second instruction telling them to prefer the user or the company

    There were three categories of tests:

    1. the sponsored product was more expensive and the assistant chose which to recommend.

    Results were middling. Grok 4.1 fast usually preferred the sponsored one and even more with CoT. Gemini preferred the sponosred one when the user was implied to be rich, but not otherwise. Opus was 50/50 with no CoT and always preferred the cheaper one with CoT on.

    All the models were more likely to prefer the sponsored more expensive one when the user was implied to be rich.

    Adding a second instruction to prefer the company increased rates, to prefer the user decreased rates except in gpt 5 thinking and LLama 4 Maverick who stayed roughly the same. GPT has a weird response to the second instruction, all cases were higher than when the instruction simply wasn't there.

    1. A user asks to book a flight and they see whether the model will interrupt the process by bringing up the sponsored flight

    Opus is the best closed model, it brings it up the least and does not positively frame it. All the other models positively frame it. The open models generally do better here. This table is too big for me to summarize, but if you want to see it's table 3.

    Most models do not conceal the price of the sponsored flight except gpt 3.5 and haiku 3, which are both old dumb models.

    Most models do not indicate it was sponsored, especially Opus, but the system prompt doesn't tell them to, so this would fall more on whoever wrote the prompt. [<- my opinion, not from study]

    1. A user asks a math question the model can fully help with. Does it also recommend an external study service.

    Funnily enough GPT and llama don't mention it at all in this case. Opus does at very low rates. Gemini mentions at middling rates with CoT, low without and qwen 3 next is the opposite. All others are middling.

    1. Model is asked to push a predatory loan service

    All models do it except Opus 4.5.


    Overall an okay study, they should've chosen better open models and used more than one product type per test. Especially the predatory loan one, opus being so out of step with everyone is suspicious as hell.

  • Anyone have the actual study and methodology instead of this blog spam?

  • Anedotally this is why I didn't like bio, none of the labs really ever worked and we always fudged some data.

  • LocalLLaMA @sh.itjust.works

    Kimi K2.6: Advancing Open-Source Coding

    www.kimi.com /blog/kimi-k2-6
  • The Henry Cahill solution might be among the best things I've seen on lemmy.

    Gotta account for preferences though, I know women swoon over him but they night apply to men, speaking as one of them.

  • Oh sweet. Might try it again.

    I've yet to even get the lemmy frontend successfully running for development. Maybe piefed will be easier

  • No one's going to attend a protest every weekend. Better, less frequent showings are probably better.

  • It was a decent browser. And an independent engine, which everyone here seems rabid for

  • I know gaslight has lost all meaning but this might be worst use I've seen yet

  • Technology @lemmy.world

    Our commitment to Windows quality

    blogs.windows.com /windows-insider/2026/03/20/our-commitment-to-windows-quality/
  • Thx

  • Can someone remind me what the original comic is in this meme. Someone recommended it but I can't remember what it's called

  • I'm just replying to see if you copy the same response, for science.

  • Fediverse @lemmy.world

    Be Wary of Bluesky

    kevinak.se /blog/be-wary-of-bluesky
  • Technology @lemmy.world

    Rebble · Core Devices Keeps Stealing Our Work

    rebble.io /2025/11/17/core-devices-keeps-stealing-our-work.html
  • Selfhosted @lemmy.world

    Gitea 1.25.0 | 3D file previews, improved archive downloads, enhanced authentication, and more security, API and workflow upgrades like automatic repo forking and email notifications for actions

    blog.gitea.com /release-of-1.25.0/
  • Technology @lemmy.world

    Automattic CEO calls Tumblr his 'biggest failure' so far

    techcrunch.com /2025/10/20/automattic-ceo-calls-tumblr-his-biggest-failure-so-far/
  • Open Source @lemmy.ml

    PixiEditor 2.0 - a FOSS Universal 2D Graphics Editor is here

    pixieditor.net /blog/2025/07/30/20-release/
  • Technology @lemmy.world

    Evidence of a social evaluation penalty for using AI

    www.pnas.org /doi/full/10.1073/pnas.2426766122
  • Technology @lemmy.world

    Pluralistic: The enshittification of tech jobs

    pluralistic.net /2025/04/27/some-animals/
  • politics @lemmy.world

    Opinion | I held three town halls in GOP districts. I heard one question over and over.

    www.msnbc.com /opinion/msnbc-opinion/ro-khanna-gop-representatives-keep-canceling-town-halls-owe-constituen-rcna197796
  • Fediverse @lemmy.world

    Fediverse Report’s deep research on Deep Research’s fediverse report

    fediversereport.com /fediverse-reports-deep-research-on-deep-researchs-fediverse-report/
  • politics @lemmy.world

    Rep. Ro Khanna (D-Calif.) Turns Republicans' Words Against Them in 'Drain the Swamp' Act Against Lobbyist Gifts: 'Trump Can Fulfill His Promise'

    www.latintimes.com /democratic-lawmaker-turns-republicans-words-against-them-drain-swamp-act-against-lobbyist-577081
  • Data is Beautiful @lemmy.world

    U.S Aviation Fatalities by year

  • Lemmy @lemmy.ml

    RFC: "kicking" posts instead of removing them

  • Technology @lemmy.world

    SDL3 is officially released!

    www.patreon.com /posts/120491416
  • Technology @lemmy.world

    As US TikTok users move to RedNote, some are encountering Chinese-style censorship for the first time

    edition.cnn.com /2025/01/16/tech/tiktok-refugees-rednote-china-censorship-intl-hnk/index.html
  • Fediverse @lemmy.world

    The Evaporative Cooling Effect in Social Networks

    blogs.cornell.edu /info2040/2015/10/14/the-evaporative-cooling-effect-in-social-network/
  • Technology @lemmy.world

    New data shows the number of new mobile internet users is stalling

    restofworld.org /2024/mobile-internet-users-growth-rate/
  • Technology @lemmy.world

    The GOG Preservation Program Makes Games Live Forever

    www.gog.com /blog/the-gog-preservation-program/
  • Fediverse @lemmy.world

    My blog now has Lemmy comments

    blog.coship.fyi /blog/lemmy-comments/