#rlhf
Live, measured metrics for the hashtag #rlhf from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #rlhf
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 02:52 UTC0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 02:52 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 02:52 UTCEverything below is measured over the latest 40 public posts (spanning ~4344 hours).
Posting hours (UTC) — busiest: 16:00
Languages: English (21) · Russian (11) · German (1) · Chinese (Taiwan) (1) · Korean (1)
Avg boosts / post: 0.1
Top of the latest posts
Mojofull (@furoku) AI 세션 평가가 매일 자동으로 이뤄지는 환경이라면, 기업은 ZentouAI를 직원에게 쓰게 하는 것만으로도 1년 뒤 대부분의 업무를 AI가 대체할 수 있다고 말한다. 매일 최적의 업무 프로세스와 산출물이 선택되며, RLHF 수준을 넘어서는 변화라는 주장이다. https://x.com/furoku/status/2040235067197501917 #ai #automation #rlhf #ente
Saeed Anwar (@saen_dev) GRPO의 레퍼런스 모델 제거는 정렬 학습 비용을 줄이지만 안정성 트레이드오프를 만들 수 있고, DPO는 설계상 이를 피한다는 주장이다. RLHF·선호 최적화 방식을 선택할 때 방법론의 이론적 우열보다 학습에 투입 가능한 하드웨어와 컴퓨트 예산이 실질적 결정 요인이 될 수 있음을 짚는다. https://x.com/saen_dev/status/2078745370448839060 #grp
Avi Chawla (@_avichawla) RLHF, DPO, GRPO를 시각적으로 비교하며 각 정렬·선호학습 기법이 모델 행동을 학습시키는 방식의 차이를 설명하는 자료다. LLM 파인튜닝 및 정렬 기법 면접·학습용 개요로 참고할 만하지만, 새로운 연구 결과나 구현 세부사항은 본문에 없다. https://x.com/_avichawla/status/2078381748199833987 #rlhf #dpo #grpo #llm #a
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/rlhf