It would have to be a stroke during the time there was an ongoing scenario that resulted in higher than typical call volume about the scenario, and a stroke victim calling in who could only say "yes".
Sure, usually, but compare like to like. If a human operator asked if the call was about an already known scenario and transferred people to an automated service if they said "yes", then the same exact thing could happen.
You attemtped to craft a very specific scenario in which harm would result, when in reality it is wildly unlikely that someone would call in at that precise time and only respond with the word "yes".
The requirements for cloud gaming revolve around having the hardware to drive the local monitor, and having fast enough internet with low enough latency.
There was recently a study using human short stories and LLM generated short stories prompted to write a similar story. The participants where told which stories were AI generated, but sometimes they were lied to, and told a human-written story was AI generated.
The study showed that people liked the AI stories more. It also showed that if someone has a negative bias against AI generated content, they'd claim the AI story was worse, even if it were actually the human story and they were just told it was AI.
Also, another test was performed where they provided a story and asked them if it was AI generated or human generated. The participants could only tell the AI written story 45% to 50% of the time. Essentially, they were guessing.
I think many people have taken a mental snapshot of the quality of AI generated content from a couple years ago and just assumed it hasn't improved since then.
The models are tested, among other things, on their ability to turn vulnerabilities into exploits. The OpenAI scenario was exactly this.
It is a very wise practice to test these things in isolation, especially when you're telling it to hack.
I'm not completely sold on it being a publicity stunt, personally. The law was broken by these models, and I don't believe these companies want to start people and politicians asking the question about who is culpable when an AI breaks the law.
It's not clear what you mean when you say "on their own". It wasn't like the LLM was idle and randomly decided to start hacking. At least for the OpenAI one, it was being tested and given a task, and it determined that part of accomplishing that task was hacking another server. It was supposed to be isolated in a secure "sandbox" not connected to the internet, but found a vulnerability in some software running in the sandbox and broke out.
Edit: I should add that there are credible accusations that these companies are intentionally making it possible to break out of their test environments for publicity.
I actually really like the technology that has been collectively lumped together as "AI". I think it's fairly useful now, and suspect we're at the beginning of the curve for this technology, and it will only continue to improve, even as it becomes more efficient. I understand this is a wildly unpopular stance in these parts, judging by the avalanche of downvotes I get for having this opinion.
What I don't like is feeding tons of information about myself to a giant tech company. I run a local LLM on my phone and another on my (fairly decent) computer, and sometimes I use duckduckgo's front end, which anonymizes prompts.
As for agents, it's pretty much as you say. I just am not ready to have an LLM take action without my direct supervision, because I'm confident that between flaws in the LLM and flaws in how I give it direction, something will go sideways.
Back when search engines were becoming a thing, people would search more or less using natural language, e.g., "I want to see cat pictures and also cute kittens." but early search engines performed poorly with that kind of query; it was much better to just type keywords, e.g. "cat kitten pictures". Exacerbating the situation was the fact that search engines weren't very good at the time, even if you used carefully considered keywords.
It might seem silly now, but using a search engine effectively was a skill people had to learn, and a skill not everyone had.
Using chatbots is the same thing, I think. They're not super reliable to begin with, and on top of that, I don't think people understand the importance of a well crafted prompt. The combination can make using them very frustrating.
If an ICE officer is found breaking the newly enacted PROTECT Act, they face exposure to state-level civil lawsuits, formal prosecution threats by state officials, and extensive auditing, but they will not be arrested by local police.
Because of the Supremacy Clause of the U.S. Constitution and federal immunity doctrines, a local police department does not have the legal authority to physically handcuff and jail a federal agent who is carrying out an immigration enforcement action, even if that action violates state law.
But but but what if the person is stroking out and presses 1 when they meant 2???