Curated developer articles, tutorials, and guides — auto-updated hourly


Error feedback makes a biased gradient compressor unbiased over time, and under SGD it restored the ...


A palette study on synthetic blocks said evenly-spaced INT4 beats NVIDIA's FP4 grid once a Hadamard ...


A per-block scale that cuts FP4 gradient-quantization error on 45 of 45 tensors, 14% against the pub...


The importance matrix used to compress Qwen3.8-27B contains no entries for block 64, the model's mul...


A supercomputing-conference study that injected more than 13 million simulated hardware faults into ...


Support for Qwen3.8-Flash-Next landed in llama.cpp on August 27, adding a sparse-attention graph, vi...


TielCoder, a 4-bit re-quantization of the Ornith-1.5 mixture-of-experts model, fits in 22.4 GB and f...