#quantization
Live, measured metrics for the hashtag #quantization from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #quantization
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · mastodon.online (Mastodon public tags API) · fetched 2026-07-27 00:11 UTC2 uses by 2 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · mastodon.online (Mastodon public search API) · fetched 2026-07-27 00:11 UTCLive pulse
measured · mastodon.online (Mastodon tag timeline) · fetched 2026-07-27 00:11 UTCEverything below is measured over the latest 40 public posts (spanning ~690 hours).
Posting hours (UTC) — busiest: 13:00
Languages: English (40)
Avg boosts / post: 0
Top of the latest posts
Joel - coffee/acc (@JoelDeTeves) 한 사용자가 특정 모델의 최근 품질이 개선됐다고 평가하며, KV 캐시 양자화에는 민감하므로 피하는 것을 권장했습니다. RTX A6000에서 빠르고 똑똑하게 동작했으며, Q4_K_XL 양자화는 잘 작동했고 VRAM 여유가 생기면 Q6/Q8을 사용할 계획이라고 밝혔습니다. https://x.com/JoelDeTeves/status/2081398003475308747 #ll
Akshay (@akshay_pachaar) 단일 GPU에서 70B LLM을 구동하기 위한 양자화 기법을 정리한 글이다. FP16 가중치만 약 140GB가 필요한 70B 모델을 4비트 양자화로 약 35GB까지 줄일 수 있지만, 단순 반올림은 대형 모델 품질을 크게 훼손할 수 있어 적절한 양자화 전략이 필요하다고 설명한다. https://x.com/akshay_pachaar/status/2079946312318095852 #ll
Sudo su (@sudoingX) 6GB·8GB·24GB VRAM GPU에서 27B급 LLM을 로컬 추론할 때 가능한 구성과 한계를 정리한 가이드입니다. GTX 1660 Super 6GB에서는 약 8K 컨텍스트의 채팅 중심 사용이 가능하지만 서버 구동은 어렵다고 제시합니다. 제한된 VRAM에서 양자화 모델·컨텍스트 길이·처리량을 판단하려는 로컬 LLM 개발자에게 참고가 됩니다. https://x.com/sudoingX/sta
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/quantization