---
format: "aidr-story-markdown/v1"
id: "ab3fa2da22740ccb93d6f85e3aa0d9412268166e9f456379b2a4eaebb2eb74b1"
canonical_url: "https://aidr.today/ab3fa2da?lang=en"
title: "Show HN: Shoehorn – Quantize any model down to run on your machine"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-18T14:29:20.000Z"
category: "Models"
topics: ["llm","quantization","open-source","chips","inference"]
source_urls: ["https://notactuallytreyanastasio.github.io/shoehorn/","https://news.ycombinator.com/item?id=49346135"]
summary: "Solves a per-tensor quantization that uses 99.99% of your exact VRAM budget — every spare megabyte spent where it buys the most quality. macOS, Linux, and Windows."
---

# Show HN: Shoehorn – Quantize any model down to run on your machine

> [Open the canonical story](<https://aidr.today/ab3fa2da?lang=en>)

**Published:** 2026-08-18T14:29:20.000Z
**Category:** Models
**Topics:** llm, quantization, open\-source, chips, inference

## Summary

Solves a per\-tensor quantization that uses 99\.99% of your exact VRAM budget — every spare megabyte spent where it buys the most quality\. macOS, Linux, and Windows\.

## Sources

- [Story source](<https://notactuallytreyanastasio.github.io/shoehorn/>)
- [Discussion](<https://news.ycombinator.com/item?id=49346135>)

