#llminference
Live, measured metrics for the hashtag #llminference from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #llminference
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · mastodon.online (Mastodon public tags API) · fetched 2026-10-02 13:12 UTC6 uses by 3 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · mastodon.online (Mastodon public search API) · fetched 2026-10-02 13:12 UTCLive pulse
measured · mastodon.online (Mastodon tag timeline) · fetched 2026-10-02 13:12 UTCEverything below is measured over the latest 40 public posts (spanning ~975 hours).
Posting hours (UTC) — busiest: 06:00
Languages: English (40)
Avg boosts / post: 0
Top of the latest posts
NVIDIA (@nvidia) NVIDIA가 OpenAI의 GPT-6 Astra Ultrafast가 NVIDIA 인프라에서 동작하며 최대 8배 빠른 성능을 제공한다고 소개했습니다. 구체적인 비교 기준과 워크로드는 트윗만으로 확인되지 않지만, 대규모 LLM 추론에서 가속기·서빙 스택 최적화가 지연시간과 처리량에 미치는 영향을 보여주는 사례입니다. https://x.com/nvidia/status/210580944236542391
fly51fly (@fly51fly) LeapQuant는 선형 어텐션 모델의 recurrent state를 정확하게 양자화하는 방법을 다룬 연구입니다. 긴 컨텍스트 추론에서 메모리 사용량과 상태 전송 비용을 줄이면서도 정확도를 유지하려는 접근으로, 선형 어텐션 기반 LLM의 효율적 서빙·온디바이스 추론에 관련성이 있습니다. https://x.com/fly51fly/status/2105409982435090571 #lineara
Together AI (@togethercompute) Together AI가 OpenRouter의 상위 오픈 코딩 모델 토큰 점유율에서 높은 비중을 차지한다고 발표했다. GLM 5.3 Flash(29.2%), DeepSeek V4.1 Flash(25.6%), Kimi K3(18.9%)를 통해 코딩 에이전트 워크로드를 제공한다는 내용으로, 오픈 모델 기반 코딩 에이전트의 실제 사용 분포를 참고할 수 있다. https://x.c
#llminference across platforms
every network with a public tag surfaceFollow #llminference straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/llminference