The open-weights story fed straight into a hardware conversation, with people showing a 2.78-trillion-parameter model actually running on machines they own. Unsloth's head-to-head — 1-bit K3 against Claude Opus 5 and GPT-5.6 on an identical physics sim — gave the day a concrete quality bar, and the political undertone was explicit: guardrails and pricing on closed models are pushing developers toward self-hosted stacks.
tinygrad framed 42 tok/s K3 + 120 tok/s GLM-5.2 on AMD as a ~$600k frontier setup, one claim ran the full unmodified model off NVMe on a laptop, DeepSeek V4 Flash hit 197 tok/s aggregate on 2× DGX Spark, and the recurring argument was that closed-model guardrails are driving people to self-hosted OSS.