#rlhf
Live, measured metrics for the hashtag #rlhf from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #rlhf
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · mastodon.online (Mastodon public tags API) · fetched 2026-09-28 01:41 UTC5 uses by 4 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · mastodon.online (Mastodon public search API) · fetched 2026-09-28 01:41 UTCLive pulse
measured · mastodon.online (Mastodon tag timeline) · fetched 2026-09-28 01:41 UTCEverything below is measured over the latest 40 public posts (spanning ~4300 hours).
Posting hours (UTC) — busiest: 12:00
Languages: English (24) · Russian (12) · Czech (1) · Japanese (1) · German (1)
Avg boosts / post: 0.2
Top of the latest posts
Sedm dolarů za hodinu O stroji, který hraje Doom desetkrát za sekundu a za celou dobu neřekne ani slovo, o muži, který naučil počítače mluvit a teď je prosí, aby zmlkly, o bičníku na prvním automobilu, o tom, proč plynulá lež vždycky dostal
Emily (@IamEmily2050) Tencent RL Team의 연구를 소개하는 Gemini Notebook 영상으로, LLM 강화학습(RL)에서 배치 크기 확장이 학습 효율에 미치는 영향을 다룬다. 대규모 RL 파인튜닝 시 배치 크기와 스케일링 전략을 검토하려는 개발자에게 관련 연구 동향으로 참고할 만하다. https://x.com/IamEmily2050/status/2103339972430270484 #llm #rei
https://youtube.com/shorts/VfWwn168ZnM?feature=share Former OpenAI researcher Diogo Almeida shares a hilarious demonstration of RLHF optimization using a viral test on ChatGPT. The example perfectly illustrates how models are optimized to f
What “rlhf” means
WikipediaIn machine learning, reinforcement learning from human feedback (RLHF) is a technique to align an intelligent agent with human preferences. It involves training a reward model to represent preferences, which can then be used to train other models through reinforcement learning.
“Reinforcement learning from human feedback” on Wikipedia (CC BY-SA) →#rlhf across platforms
every network with a public tag surfaceFollow #rlhf straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/rlhf