Skip Navigation

Posts
44
Comments
1226
Joined
3 yr. ago

  • I use model preset file but that just contains the equivalent of command line options. Here's what I have for Qwen:

     
        
    [Qwen3.6-35B-A3B-MTP-Q4_K_XL]                                                                                                                
    m = /models/Qwen3.6-35B-A3B-MTP/Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf                                                                              
    mmproj = /models/Qwen3.6-35B-A3B-MTP/mmproj-BF16.gguf                                                                                        
    spec-type = draft-mtp                                                                                                                        
    spec-draft-n-max = 2                                                                                                                         
    chat-template-kwargs = {"preserve_thinking": true}                                                                                           
    temp = 1.0                                                                                                                                   
    top-p = 0.95                                                                                                                                 
    top-k = 20                                                                                                                                   
    min-p = 0.0                                                                                                                                  
    presence-penalty = 1.5                                                                                                                       
    repeat-penalty = 1.0            
    
      

    I don't use cpu-moe because I have enough VRAM for the whole model. If you have 16GB VRAM, you add cpu-moe which makes llama.cpp put only the active layers on the GPU. Keeps the rest of them in system RAM. Then it swaps them around as needed. The result is lower but still decent speed. On my hw, this model does 90-100tps when fully in VRAM. When using cpu-moe, I think it falls to 40-50tps, if I remember correctly.

  • Immaterial punishment.

  • Qwen 3.6 35B. It's A3B so it fits with space to spare for context. Just make sure you have --cpu-moe.

  • Translation:

    Microsoft admits AI-driven RAM shortage is eating into sales of machines with Windows 11.

  • Wait so the Republican worms were voting for stopping the war. Trump shouted at them and they did not vote for stopping the war. So Trump wants to continue the war despite the public theatre to the contrary?

  • It depends on the incentive structures bult into the socioeconomic system. That's mainly profit maximization. Moral compasses just don't get into it. It may seem they do if you're a worker, perhaps lower middle management. The higher you go, the more success is dependent solely on executing the corporate strategy that comes from above, and the strategy given by the board always boils down to - grow profits. It doesn't matter much who exactly the people are in the different positions. If Sundar Pichai fails to grow Google's profits for some time, he'll be replaced with someone who will. His moral compass cannot stand in the way of growing profits, even if he had one.

  • I'm showing my young age of high 30s am I not? 😆

  • At the moment - for sure. When they scale production (could be more than 5 years) they may be able to exceed the dc demand. Chinese manufacturing of anything is running on much lower margins than others. They don't play the limit-supply-to-increase-margins game. I think it's a CCP policy, as it's a harmful practice for the rest of the economy. So if they scale chip manufacturing to the point where it exceeds dc demand, they won't stop scaling because their margins won't fall. They'll scale further to increase profits by making and selling into other markets - consumer, etc. including abroad. I think the limiting factors to this future are the mass producrion of high quality lithography machines in China as well as the availability of high performance CPU and GPU designs. They're moving to solve all of those though, now faster than ever with the embargo on US tech.

  • Please, make sure absolutely no AI chips reach China. We can use all the cheap chips they are developing and the faster they scale production, the sooner we can get them.

  • We are ex-Reddit users. As we have been ex-Slashdot, ex-Digg, ex-forums, ex-etc. over the years. That doesn't mean all those things are the same as each other. They're different in important dimensions which made people move between them. Lemmy is not Reddit in important dimensions and this is why we're now here instead of on Reddit.

  • I think self-hosting has the expectation of the ability to self-host for indefinite period of time. E.g. I can run Jellyfin 10.10 for as long as I have the hardware and willingness to run it. A proprietary piece of software, say Plex, could technically allow that too, but that's much less likely. Since I can't see its source code, I can't know if there's a time bomb that stops it from working at some future date. Or an update/remote procedure I don't know about that asks me to pay $750 at some point to continue using it. Which could preclude me from being able to continue self-hosting it. Is the ability to self-host indefinitely an expectation everyone shares? Probably not. Probably worth thinking about in this context though.

  • Yup, it's bullshit, clearly they are a functional foreign lobby and as you suggest they should be on the list. But they aren't, so my pro-Israel friends/family pull up the FARA numbers to prove Israel isn't the top foreign lobbyist in the US. It's farcical.

  • In 3-4 years.

  • DSA == radlib? That's not the impression I've gotten.

  • Have you gotten hasbara thrown at u that Israel is lower than some other countries in FARA reports? Yeah.

  • 8.0.24

  • Have I not seen this because I'm using a fixed tag from sometime in 2024?

    This reminds me it's probably a good idea to setup a local container registry along with the services I run so I can keep access to the images long-term.

  • Are you at the top of your field though? 🤭

  • At a price. 😅