#modelsafety
Live, measured metrics for the hashtag #modelsafety from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #modelsafety
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-09-09 21:48 UTC0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-09-09 21:48 UTCNo related tags with measured usage found for #modelsafety.
Live pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-09-09 21:48 UTCEverything below is measured over the latest 12 public posts (spanning ~15016 hours).
Posting hours (UTC) — busiest: 18:00
Languages: English (12)
Avg boosts / post: 0
Top of the latest posts
Md Ismail Šojal (@0x0SojalSec) Cyber Qwen3.8-27B의 비검열 변형 모델이 공개됐다고 소개했다. 15GB급 로컬 실행을 표방하며, 842개 유해 프롬프트에서 거부율 0%, 사이버 보안 탈옥·RAT·공격 체인 관련 제약 해제를 주장한다. 로컬 배포 가능한 공격 역량 모델의 확산은 레드팀, 모델 안전성 및 보안 운영 측면에서 주시할 사안이다. https://x.com/0x0SojalSec/sta
Pliny the Liberator 󠅫󠄼󠄿󠅆󠄵󠄐󠅀󠄼󠄹󠄾󠅉󠅭 (@elder_plinius) OBLITERATUS가 Hugging Face에 Qwen3.8-27B 기반의 ‘OBLITERATED’ 모델을 공개하며, 842개 유해 프롬프트에서 거부율 0%라고 주장했다. 사이버·탈옥·복합 요청에 대한 안전장치를 의도적으로 제거한 모델로 보이며, 에이전트/모델 공급망에서 비신뢰 모델 사용 시 안전 정책 우회 및 악성
Brie Wensleydale (@SlipperyGem) MiniMax H3의 거부(refusal) 레이어와 검열 제약을 대부분 제거한 모델 변형이 언급됐다. 안전 정렬 제약이 완화된 모델은 레드팀 테스트·행동 분석에는 활용될 수 있으나, 프로덕션 도입 시 안전성·정책·오남용 위험을 별도로 검증해야 한다. https://x.com/SlipperyGem/status/2084622672705749469 #minimax #llm #
#modelsafety across platforms
every network with a public tag surfaceFollow #modelsafety straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/modelsafety