My locally hosted Qwen3 30b said “Walk” including this awesome line:
Why you might hesitate (and why it’s wrong):
- X “But it’s a car wash!” -> No, the car doesn’t need to drive there—you do.
Note that I just asked the Ollama app, I didn’t alter or remove the default system prompt nor did I force it to answer in a specific format like in the article.
EDIT: after playing with it a bit more, qwen3:30b sometimes gives the correct answer for the correct reasoning, but it’s pretty rare and nothing I’ve tried has made it more consistent.
I keep hearing “oh this new model is better!”
I have a test case I’ve been using. A real-world piece of code I needed, that isn’t something super original but has a few tricky steps in it.
The first time I’ve seen AI able to get even close to finishing the task was recently. Its code worked, and there were only a few minor tweaks I had to make before it was in a condition I’d consider acceptable to merge into my own work.
It took about 30 minutes to do its task, maybe longer. I spent 5 minutes reviewing it but it would have been longer if I hadn’t previously done the task myself and knew exactly what I was looking for. I think it took me about an hour when I did that task myself the first time. Ultimately using AI for this might have saved me about 15 minutes?
I guess it might be borderline useful at that rate. I might look into using it more in the future, but I still don’t really expect it to become a tool I regularly use.