#swebench
Live, measured metrics for the hashtag #swebench from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #swebench
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 02:25 UTC1 uses by 1 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 02:25 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 02:25 UTCEverything below is measured over the latest 40 public posts (spanning ~10222 hours).
Posting hours (UTC) — busiest: 14:00
Languages: English (27) · Russian (8) · Portuguese (2) · Korean (2)
Avg boosts / post: 0.3
Top of the latest posts
If your company is benefiting from Django’s stability and maturity to test or train AI models, consider **funding Django’s development**. 💚 Support Django: https://www.djangoproject.com/fundraising/ #Django #AI #LLM #Benchmarks #OpenSource
How does it perform? Empirically, on the scenarios I’ve tested (including hand-picked #SWEbench Lite & Multilingual tasks), access to an #ORF playbook cut agent step counts roughly in half (~50%) and eliminated error loops.
Daniel Han (@danielhanchen) Codex의 GPT-5.6-Sol xhigh가 Margin Lab의 일일 SWE Bench Pro에서 성능이 크게 상승해, 50개 랜덤 문제 기준 약 82%를 기록했다고 합니다. 직전 5.5 Xhigh의 약 54%보다 개선 폭이 커서, Codex 기반 코드 생성/리뷰 품질 향상 신호로 볼 수 있습니다. https://x.com/danielhanchen/status/2076304
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/swebench