---
format: "aidr-story-markdown/v1"
id: "66c10c1bec5a9cb0a03c6a4180d8bb19fef3a39271d4e6871a71c59e31872c3f"
canonical_url: "https://aidr.today/66c10c1b?lang=en"
title: "Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-09T12:53:21.000Z"
category: "Research"
topics: ["qwen","deepseek","quantization","benchmark","inference","chips"]
source_urls: ["https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/","https://lobste.rs/s/lxiabq/benchmarking_qwen3_8_27b_quantizations_4"]
summary: "I test Unsloth GGUFs of Qwen3.8 27B (Q4_K_M, UD-Q2_K_XL, UD-IQ1_S) with llama.cpp on GPQA Diamond, IFBench, Terminal-Bench 2.1. Q4_K_M matches BF16 abd fits an RTX 4090."
---

# Benchmarking Qwen3\.8 27B quantizations: 4\-bit holds up, 1\-bit collapses

> [Open the canonical story](<https://aidr.today/66c10c1b?lang=en>)

**Published:** 2026-09-09T12:53:21.000Z
**Category:** Research
**Topics:** qwen, deepseek, quantization, benchmark, inference, chips

## Summary

I test Unsloth GGUFs of Qwen3\.8 27B \(Q4\_K\_M, UD\-Q2\_K\_XL, UD\-IQ1\_S\) with llama\.cpp on GPQA Diamond, IFBench, Terminal\-Bench 2\.1\. Q4\_K\_M matches BF16 abd fits an RTX 4090\.

## Sources

- [Story source](<https://quesma.com/blog/qwen38-27b-quantizations-benchmarked/>)
- [Discussion](<https://lobste.rs/s/lxiabq/benchmarking_qwen3_8_27b_quantizations_4>)

