32b q4_k_m drops to 3 tok/s at 16k context on a 24gb 3090 vram-fit
At 4k context I get 28 tok/s and everything is on the card. At 16k the same model gives me 3.1 tok/s. No error, no warning, it just gets slow. shows 23.9GB used, and the server log says it offloaded 51 of 65 layers to…