#ModelEvaluation
Live, measured metrics for the hashtag #ModelEvaluation from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #modelevaluation
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 04:31 UTC9 uses by 4 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 04:31 UTCNo related tags with measured usage found for #modelevaluation.
Live pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 04:31 UTCEverything below is measured over the latest 40 public posts (spanning ~5275 hours).
Posting hours (UTC) — busiest: 09:00
Languages: English (38) · Korean (2)
Avg boosts / post: 0.3
Top of the latest posts
RootCo (@rootlesscp) 현재 사용 중인 26B·31B 모델은 일부 멀티홉 에이전트 작업에서 루프에 빠지는 반면, Qwen 27B는 정상 동작한다는 실사용 비교입니다. 에이전트 모델 선택 시 단순 벤치마크 외에 장기 실행·루프 안정성을 검증해야 함을 보여줍니다. https://x.com/rootlesscp/status/2081434893595066604 #agents #qwen #multihop #modeleval
Mikeysee (@mikeysee) Anthropic의 Opus 5가 Convex 환경에서 기대 수준의 성능을 보인다는 초기 평가입니다. 구체적 벤치마크나 기능 비교는 없지만, 차세대 Claude 계열 모델의 실사용 성능에 대한 현장 관찰로 볼 수 있습니다. https://x.com/mikeysee/status/2080969759018148134 #anthropic #opus5 #llm #modelevaluation
Design Arena (@DesignArena) Anthropic의 Claude Opus 5가 Design Arena에 추가됐다. 게시물은 Fable급 성능에 근접하면서 비용은 절반 수준이며, 코딩과 지식 노동 전반에서 일상적 사용 효율을 높였다고 주장한다. 실제 모델 선택 시 Arena 결과와 비용 대비 성능을 추가 검증할 만하다. https://x.com/DesignArena/status/208070444668281667
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/modelevaluation