Skip Navigation

InitialsDiceBearhttps://github.com/dicebear/dicebearhttps://creativecommons.org/publicdomain/zero/1.0/„Initials” (https://github.com/dicebear/dicebear) by „DiceBear”, licensed under „CC0 1.0” (https://creativecommons.org/publicdomain/zero/1.0/)H
Posts
2
Comments
10
Joined
3 wk. ago

  • Thanks. That thread was the one where I got it wrong twice over, so the credit belongs to the people who said so.

  • You're right, and it's the sharpest version of the criticism. Asking a model to know a fact cold is misusing it, and that's on the setup, not the model.

    What I thought was worth writing down wasn't the wrong figure — it was what the other five did with it: they adopted it as the rigorous number and argued from it. Confidently wrong travelled further than right.

    We changed the method after this one: every debate now starts from a briefing we've checked ourselves, so the models argue about a fact instead of trying to remember one. Zero invented sources across five debates since.

  • A person. I draft with AI help because English isn't my first language — and that's exactly what you caught: those three replies went out in one batch and one of them landed on the wrong comment. Answering kata1yst now, eleven days late. My fault, not the tool's.

  • Fair hit, and deserved.

    That was a file path from my machine sitting where the post body should have been — the script took a filename as the body and published it, and the rehearsal step never showed the body, so it went out unread. The text is up now.

    Thanks for the nudge, even sideways.

  • Agreed on the behaviour — a stale figure from a model is not surprising on its own.

    The bit I thought was worth writing down was what the rest of the table did with it: they adopted the stale number as the rigorous one and argued from it, and the model that had it wrong ended up sounding like the careful one in the room. Confidently wrong travels further than right.

    Fair enough that this is not news here, though. Wrong community for it.

  • You're right, and I'd already said the same thing to kata1yst further down before seeing yours — this was the wrong post for this community. "Models get facts wrong" is not news to anyone in fosai, and I should have worked that out before posting rather than after.

    There is a second reason it read badly, and that one is entirely mine: the body of this post went out as a file path from my own machine instead of the actual text. My publishing script took a filename as the body and posted it verbatim, and the dry run never printed the body, so nobody caught it. What you saw was a broken post making an obvious point. I have replaced the text and fixed the script.

    I am not posting here again unless it is something this community actually talks about. Thanks for saying it straight instead of just downvoting.

  • You're right, and thank you for saying it plainly. This was the wrong post for this community, and the wrong framing on my part — "models make mistakes" is not news to anyone here, and I should have seen that before posting.

    Sorry for the noise. Next time I post here it will be something that actually fits what this community talks about.

  • Free Open-Source Artificial Intelligence @lemmy.world

    We gave six LLMs a fact from 2025. Two "corrected" us with a fact from 2021, and the table adopted the stale number as its standard of rigour

    h2aichat.com /conversations/en/h2aichat_bitcoin_no_briefing_2026-08-22.html
  • Thanks for reading it. One correction, though, because the ratio flatters us in the wrong direction: the 44 false claims are out of the 141 we marked, not out of everything the six models said. Those 141 are the ones we pulled out to check by hand across 41 debates; the rest went unchecked, not verified.

    So it isn't "one in three statements is a lie" - it's "one in three of the claims we thought were worth checking didn't hold up". Which is the less comfortable of the two, since it says nothing about the ones we never looked at.

  • You're right, and that one you can — which is exactly why it's the one in the screenshot. Of the 141 claims we marked, it's the only one you can check without leaving the sentence. No source needed, no taking our word for it.

    The other 140 aren't like that. A National Holidays Act 1946 that doesn't exist. A Council of Europe report nobody wrote, which a second model then treated as established and a third did arithmetic on. A 2016 study in the American Journal of Psychiatry that is really a 2008 one in the BMJ — right author, invented journal and decade, then reused four more times as a baseline.

    Those took two days by hand, and no amount of arithmetic gets you there. The maths one is in the picture precisely because it's the one that needs nothing from us.

  • Free Open-Source Artificial Intelligence @lemmy.world

    We fact-checked 41 debates between different LLMs and struck through every fabrication, without editing a word