We tried Grok for some things at a job once and it was absolutely the worst of all of the major ones (for what we were doing). It hallucinated way too much even when attempting to get it not to.
I guess that makes some sense, I loathe Jira but I think it's largely because everywhere I've worked that uses Jira has poorly customized it and just ruined the experience.
You might be joking but I honestly think that's the case. It's wild to me. I've worked for Fortune 500 companies using SNOW and everybody hated it and regularly voiced complaints and issues and yet the company refused to change. Started doing shit like releasing more training docs on how to use it or doing brown bag lunches on SNOW effectiveness.
But ultimately none of that mattered, it is just inherently garbage.
Are you running your own programmatic LLMs? There is a thing called temperature, and it is typically more lenient for public facing LLMs. But leverage that same LLM via APIs and you can adjust the temperature and reduce or eliminate hallucinations.
Ultimately, a little variance (creativity) is somewhat good and passing it through levels of agentic validations can help catch hallucinations and dial in the final results.
That said, I doubt the WH did this, they probably just dumped shit into some crappy public-facing ChatGPT model.
Definitely not OK. But it exists and I don't think people realize it goes beyond tracking clicks to taking actual screenshots that can be stitched together practically as a video. It sucks.
Yup. Worked briefly for a company that would "snapshot" the browser view quite often, enough where if an issue arose we could somewhat replay the user's interactions to try and repro the issue.
Is Ukraine supposed to inform the US of these things? That seems...like not a great idea.