You need a model compiled for the architecture. I saw some for the RK35xx devices when shopping for hardware. I do not think there is software made to split up or run models in general on a NPU. The models must be configured for the physical hardware topology. The stuff that runs on most devices is very small, and these either need a ton of custom fine tuning or they are barely capable of simple tasks.
On the other hand, segmentation models are small, and that makes layers, object identification, and background removal stuff work. Looking at your CPU speed, and available memory, it is unlikely to make much difference. You are also memory constrained for running models, though you could use deepspeed to load from a disk drive too.
Most of the major developments that lead to the public stuff happened between 2017-2021. Transformers was the big one that made scaling a thing. Altman pushed in a stupid direction that caused a lot of the nonsense, like turning the name "Open AI" into an oxymoron.
There are some aspects of alignment that point at political corruption and planning with nefarious intent that fits in with the present political bullshit too, but that is very complicated to explain in any depth. If you were to search the token vocabulary, you will find dubious elements are present in compound multi word tokens that disproportionately represent a single political camp, likewise with religious media, and science denialism. Much of that stuff dates from 2019 or before.