#llmjudge
Live, measured metrics for the hashtag #llmjudge from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #llmjudge
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · mastodon.online (Mastodon public tags API) · fetched 2026-10-04 23:53 UTC2 uses by 1 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · mastodon.online (Mastodon public search API) · fetched 2026-10-04 23:53 UTCLive pulse
measured · mastodon.online (Mastodon tag timeline) · fetched 2026-10-04 23:53 UTCEverything below is measured over the latest 8 public posts (spanning ~6290 hours).
Posting hours (UTC)
Languages: English (6) · Russian (2)
Avg boosts / post: 0.1
Top of the latest posts
fly51fly (@fly51fly) University of Pennsylvania 연구진이 JEV와 LLM 기반 루브릭 평가자를 비교하며, JEV가 더 저렴하고 빠르지만 LLM 평가자와 유사한 지점에서 오답·편향을 보인다는 결과를 제시합니다. LLM-as-a-Judge를 비용 절감형 자동 평가기로 대체할 때 오류 상관관계와 검증 체계를 함께 점검해야 함을 시사합니다. https://x.com/fly51fly/status/2
Akshay 🚀 (@akshay_pachaar) CMU 논문은 LLM-as-a-judge 평가에서 모든 응답을 세밀하게 채점하는 대신, ‘근거 기반인가’, ‘지시를 따랐는가’처럼 경계가 명확한 이진·제한적 판단으로 평가를 구성하는 접근을 다룬다. LLM 생성 결과의 자동 평가 신뢰성과 비용을 개선하려는 개발자에게 참고할 만하다. https://x.com/akshay_pachaar/status/210421411699927887
fly51fly (@fly51fly) Meta Superintelligence Labs 연구진이 침묵, 압박, 반복 질문 조건에서 LLM 평가자(judge)의 인식론적 안정성을 다루는 논문 「Jagged Judges」를 공개했다. LLM-as-a-Judge를 학습·벤치마크·에이전트 평가에 사용하는 경우, 프롬프트 조건 변화에 따른 판정 일관성과 신뢰성 검증에 관련된 연구다. https://x.com/fly51fly/status
#llmjudge across platforms
every network with a public tag surfaceFollow #llmjudge straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/llmjudge