#evaluation
Live, measured metrics for the hashtag #evaluation from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #evaluation
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · mas.to (Mastodon public tags API) · fetched 2026-07-27 00:50 UTC14 uses by 14 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · mas.to (Mastodon public search API) · fetched 2026-07-27 00:50 UTCLive pulse
measured · mas.to (Mastodon tag timeline) · fetched 2026-07-27 00:50 UTCEverything below is measured over the latest 40 public posts (spanning ~706 hours).
Posting hours (UTC) — busiest: 23:00
Languages: English (29) · French (10) · German (1)
Avg boosts / post: 0.6
Top of the latest posts
Wes Roth (@WesRoth) Opus 5가 ARC-AGI 3에서 큰 성과를 냈지만, 유사한 퍼즐 유형을 겨냥해 학습됐기 때문으로 보인다는 지적입니다. 벤치마크 점수 해석 시 학습 데이터·과제 분포와의 유사성, 평가 오염 가능성을 함께 검토해야 한다는 시사점이 있습니다. https://x.com/WesRoth/status/2081179929090396201 #llm #evaluation #benchmark #arcagi
Avi Chawla (@_avichawla) LLM 평가에서 BLEU, BERTScore 등 지표마다 동일한 두 모델의 순위가 정반대로 나올 수 있음을 설명하며, AI 엔지니어가 알아야 할 11가지 평가 방법을 소개한다. 참조 답변과의 표면적 일치뿐 아니라 의미적 유사성 등 목적에 맞는 평가 지표 선택이 중요하다는 내용이다. https://x.com/_avichawla/status/2080924298571813101 #llm
Both Opus 5 and 4.8 generated far more tokens than the cross-model average during Artificial Analysis testing, labeled 'very verbose.' Token efficiency gains may reflect prompt engineering rather than fundamental model improvements. https:/
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/evaluation