AI
MoQ GGUFs and GSQ: Low-Bit GGUFs Are About to Get Much Better
🧱data structures Content type: News Content type: BlogLaunch HN: General Instinct (YC P26) – Frontier models on edge devices
🧩lisp Content type: Discussionbigattichouse/packed-twin-inference: PTI achieves ~2× throughput using a single quantized model (Q5_K_M or better) by running 4 generation streams in one batched decode call. The GPU loads model weights once per step and produces 4 predictions simultaneously. KV cache overhead is ~0.8 GiB total for all 4 streams. No draft model. No quality loss
🧱data structures Content type: CodeQwen3.6 + MTP: Calculated context size is smaller when I use `--spec-draft-type-* q4_0`. is this normal? · ggml-org llama.cpp · Discussion #24102
🧩lisp Content type: Discussion Content type: CodeNo more posts from tionis's subscribed feeds.