My favorite example was when they gave all the top LLMs a task to figure out how many circles in an image were touching. Every single one failed miserably at this task that would have been trivial for a 3 year old.
Unless the answer was five. Because of the Olympics logo.
Disregarding sampling bias and other effects, sure, 1,245 is enough for a statistically significant representative sample. You can’t do a lot of subsampling or cohort analysis from it, but if you just want to infer the behavior and attitudes of “Americans” and you sample that many Americans effectively enough, you’re fine.
But practically speaking, no. You’re not going to sample that many Americans effectively enough. Polling is a shitshow these days.
Big brain move, asking an LLM to do economics math. They’re large LANGUAGE models, not large MATHEMATICAL models. They forgo rigorous methodologies in favor of stochastically predicting correlations based on intractably large datasets, which can lead to such impressive mathematical feats as failing to count to three. Great job, Yahoo.
Both Control and the dogshit Avengers game had these upgrade systems where you were constantly bombarded with pickups that offered inane benefits like “2.5% increase to headshot damage for 3 seconds after taking damage while in midair” and you spent half the game managing your goddamn upgrades and the limited upgrade slots instead of having fun. It got to the point where I was relieved when I DIDN’T get any upgrades after a battle.
I take it as shorthand for “without known harmful chemicals, probably, also, we think you’re an idiot”.