Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
Orthrus-Qwen3: up to 7.8×tokens/forward on Qwen3, identical output distribution
github.com
GitHub - chiennv2000/orthrus: Fast, lossless LLM inference via dual-view diffusion decoding.

Crossposted from https://lemmy.ml/post/47429470