#llminference
Live, measured metrics for the hashtag #llminference from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #llminference
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 05:33 UTC5 uses by 4 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 05:33 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 05:33 UTCEverything below is measured over the latest 25 public posts (spanning ~7644 hours).
Posting hours (UTC) — busiest: 13:00
Languages: English (24)
Avg boosts / post: 0.2
Top of the latest posts
The AI world is buzzing over TurboQuant, Google Research’s new answer to the AI Memory Wall. This isn't just an incremental update; it’s a fundamental shift in how we think about hardware efficiency. By combining two new methods—PolarQuant
TonoKen3Local LLM&Infrastructure Architectとのけん (@Tono_Ken3) GLM의 CPU 추론 속도가 7 TPS에서 13.5 TPS로 개선된 상황에서, 추론(thinking)은 영어 기반 표현으로 처리하고 실제 작업은 영어에 강한 Laguna-S-2.1로 분리하는 다중 모델 파이프라인을 실험하는 아이디어를 제시했다. 언어별 모델 강점과 병렬 처리량을 조합하려는 에이전트 설계 관점의 제안이다.
Karl Freund (@karlfreund) AMD와 Cerebras가 협력해 대규모 컴퓨팅과 저지연 추론을 결합한 AI 인프라 대안을 제공한다고 발표했다. 특히 LLM 추론의 prefill과 decode 단계 성능을 겨냥하며, Nvidia 및 Groq에 대한 경쟁 구도로 제시됐다. 올해 후반 제공 예정이다. https://x.com/karlfreund/status/2080349408303153199 #aiinfrastruc
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/llminference