#agentbenchmark
Live, measured metrics for the hashtag #agentbenchmark from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #agentbenchmark
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 05:04 UTC3 uses by 2 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 05:04 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 05:04 UTCEverything below is measured over the latest 4 public posts (spanning ~1662 hours).
Posting hours (UTC)
Languages: English (4)
Avg boosts / post: 0.8
Top of the latest posts
Artificial Analysis (@ArtificialAnlys) Artificial Analysis의 AA-Briefcase 평가에서 Kimi K3는 태스크당 평균 56.4분이 걸렸습니다. 높은 턴 수(평균 83회), 태스크당 약 12만 출력 토큰 사용, 퍼스트파티 Kimi API의 낮은 처리 속도가 원인으로 제시됐습니다. 에이전트형 벤치마크에서 모델 성능뿐 아니라 토큰 비용·지연시간·도구 호출 반복이 운영 비용에 미치는
Artificial Analysis (@ArtificialAnlys) Moonshot의 2.8T 파라미터 모델 Kimi K3가 에이전트형 지식 노동 벤치마크 AA-Briefcase에서 Fable 5 다음 순위를 기록했지만, 작업당 평균 약 1시간이 걸리고 실행 비용은 Opus 4.8보다 높다는 평가다. 에이전트 모델 선택 시 단일 점수뿐 아니라 지연 시간과 총 추론비용을 함께 비교해야 한다는 사례다. https://x.com/
Arena.ai (@arena) Kimi K3가 Agent Arena 리더보드 4위에 올라 Claude Opus 4.8 및 GPT-5.6 Sol과 비슷한 에이전트 성능을 보였다는 주장이다. 7월 27일 예정대로 가중치가 공개될 경우, 공개 가중치 모델 중 최고 수준의 에이전트 모델이 될 수 있다는 전망을 제시했다. https://x.com/arena/status/2079253211077300736 #kimi #agentbenc
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/agentbenchmark