#rlhf

Live, measured metrics for the hashtag #rlhf from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.

hashtag.org network · sponsored

Own #rlhf

This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.

$2,449.80/ year · 4-character #name
Claim #rlhf — $2,449.80/yr→Buy on hashtag.space (web3)
card via hashtag.org · tokens via hashtag.space
5
Uses / 7 days
Mastodon
4
Accounts / 7 days
Mastodon
40
Recent posts
Mastodon
~0/hr
Recent pace
Mastodon · last 40
0
Avg reactions / post
Mastodon · last 40
—
Reddit posts / month
Reddit search
—
Open-web mentions
hashtag.org Firehose

Day-by-day usage

measured · mastodon.online (Mastodon public tags API) · fetched 2026-09-28 01:41 UTC
0
09-22
0
09-23
3
09-24
2
09-25
0
09-26
0
09-27
0
09-28

5 uses by 4 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.

Related hashtags

measured · mastodon.online (Mastodon public search API) · fetched 2026-09-28 01:41 UTC

Live pulse

measured · mastodon.online (Mastodon tag timeline) · fetched 2026-09-28 01:41 UTC

Everything below is measured over the latest 40 public posts (spanning ~4300 hours).

Posting hours (UTC) — busiest: 12:00

00:0012:0023:00

Languages: English (24) · Russian (12) · Czech (1) · Japanese (1) · German (1)

Avg boosts / post: 0.2

Top of the latest posts

  • Sedm dolarů za hodinu O stroji, který hraje Doom desetkrát za sekundu a za celou dobu neřekne ani slovo, o muži, který naučil počítače mluvit a teď je prosí, aby zmlkly, o bičníku na prvním automobilu, o tom, proč plynulá lež vždycky dostal

    Kreativum Teletník@[email protected]♥ 0↻ 12026-09-25 12:00 UTCView post →
  • Emily (@IamEmily2050) Tencent RL Team의 연구를 소개하는 Gemini Notebook 영상으로, LLM 강화학습(RL)에서 배치 크기 확장이 학습 효율에 미치는 영향을 다룬다. 대규모 RL 파인튜닝 시 배치 크기와 스케일링 전략을 검토하려는 개발자에게 관련 연구 동향으로 참고할 만하다. https://x.com/IamEmily2050/status/2103339972430270484 #llm #rei

    ainews@[email protected]♥ 0↻ 02026-09-25 09:50 UTCView post →
  • https://youtube.com/shorts/VfWwn168ZnM?feature=share Former OpenAI researcher Diogo Almeida shares a hilarious demonstration of RLHF optimization using a viral test on ChatGPT. The example perfectly illustrates how models are optimized to f

    HackerWorkspace@[email protected]♥ 0↻ 12026-09-24 18:28 UTCView post →

What “rlhf” means

Wikipedia

In machine learning, reinforcement learning from human feedback (RLHF) is a technique to align an intelligent agent with human preferences. It involves training a reward model to represent preferences, which can then be used to train other models through reinforcement learning.

“Reinforcement learning from human feedback” on Wikipedia (CC BY-SA) →

#rlhf across platforms

every network with a public tag surface

Follow #rlhf straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.

Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/rlhf