Post
  1. Portrait of Onur Solmaz
    @onusoz · /2026/07/30 · View on
    Here is your daily reminder to use @UnslothAI quantizations of Qwen3.6-35B-A3B llama.cpp is giving the best performance now, 61 tok/s with 64k context length There is also an issue with vLLM recipes or builds for NVFP4 quants unsloth/Qwen3.6-35B-A3B-NVFP4 is supposed to be even better than the GGUF, but I'm getting: ValueError: moe_backend='flashinfer_b12x' is not supported for FP8 MoE Is anyone else getting this as well?