#llmevals
Live, measured metrics for the hashtag #llmevals from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #llmevals
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-28 20:32 UTC2 uses by 2 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-28 20:32 UTCNo related tags with measured usage found for #llmevals.
Live pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-28 20:32 UTCEverything below is measured over the latest 3 public posts (spanning ~1564 hours).
Posting hours (UTC)
Languages: English (2) · French (1)
Avg boosts / post: 0
Top of the latest posts
Heikolino | Nyxia AI Labs (@heikolino_) Anthropic의 harness 기반 평가에서는 Opus 5가 큰 성능 도약으로 보이지만, Artificial Analysis의 원시 모델 평가에서는 Anthropic 모델의 우위가 관측되지 않는다는 지적이다. 에이전트·도구 사용 환경에서의 평가 하네스가 모델 순위와 출시 인식을 크게 바꿀 수 있음을 보여준다. https://x.com/heikolino_
Yes, you can copy our eval setup Langfuse가 자사 문서 챗봇에 적용한 평가·관측성 루프를 공개했다. 사용자 메시지 단위 트레이싱과 세션·사용자 ID를 기반으로 검색, MCP 도구 호출, 최종 생성 결과를 분리해 분석하고, 사용자 반박·범위 이탈·불만·대문자 사용 등을 LLM-as-a-judge 및 코드 기반 평가기로 감지한다. 검토한 프로덕션 트레이스를 주석과 함께 데이터셋으로 전환한 뒤, 정답
🤖 How do you actually know if your AI agent is any good? Great practical read on evaluating AI agent performance metrics, methods & the traps to avoid. A must for anyone moving into LLM evals. 👉 https://tinyurl.com/26pfmobc #AITesting #LL
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/llmevals