#llmevaluation
Live, measured metrics for the hashtag #llmevaluation from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #llmevaluation
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 04:24 UTC2 uses by 2 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 04:24 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 04:24 UTCEverything below is measured over the latest 18 public posts (spanning ~21141 hours).
Posting hours (UTC) — busiest: 09:00
Languages: English (18)
Avg boosts / post: 0.3
Top of the latest posts
After a year of breakneck innovation and amidst the neverending #AI hype, how do we know if a model is any "good"? We're excited to share our team’s learnings written by @vicki at @MozillaAI https://blog.mozilla.ai/exploring-llm-evaluation-
Rohan Paul (@rohanpaul_ai) Meta 연구진이 AI 응답의 사실성을 단순한 오류 회피뿐 아니라 필요한 정보를 충분히 포함하는 문제로 정의했다. GAMUT은 답변에서 누락된 정보를 측정해, AI 팀이 불완전한 사실성 평가와 모델 응답 품질을 더 명확히 테스트할 수 있도록 제안한다. https://x.com/rohanpaul_ai/status/2080620006526652699 #meta #factuality
Tom Krcha (@tomkrcha) pen.dev가 여러 모델을 병렬·실시간으로 평가하는 프론트엔드 디자인 에이전트를 제공한다고 소개했다. ChatGPT, Claude 구독, OpenCode Go, GitHub Copilot, Cursor 등을 연결해 동일한 디자인 작업에서 모델별 결과를 비교할 수 있는 워크플로를 지향한다. https://x.com/tomkrcha/status/2079253857834828230 #desi
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/llmevaluation