#modelinterpretability
Live, measured metrics for the hashtag #modelinterpretability from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #modelinterpretability
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-28 00:46 UTC0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-28 00:46 UTCNo related tags with measured usage found for #modelinterpretability.
Live pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-28 00:46 UTCEverything below is measured over the latest 5 public posts (spanning ~3912 hours).
Posting hours (UTC)
Languages: English (5)
Avg boosts / post: 0.2
Top of the latest posts
We Can Now Read What Claude Is Thinking. Kind Of Anthropic이 개발한 Natural Language Autoencoders(NLAs)는 Claude 모델의 내부 활성화 값을 사람이 읽을 수 있는 텍스트로 변환해 모델의 '생각'을 해석할 수 있게 한다. 이를 통해 Claude가 출력으로 드러내지 않는 내부 계획이나 의도를 탐지할 수 있으며, 미묘한 안전성 문제와 숨겨진 동기 탐지에 활
Shared Geometry of Neural Networks 최근 연구 'Manifold Steering'은 신경망 내부 표현과 행동 결과 사이의 인과 관계를 활성화 매니폴드 상에서 개입함으로써 분석한다. 이 방법은 기존의 선형 개입 방식보다 자연스러운 신경망 동작을 더 잘 복원하며, LLM과 비디오 월드 모델 모두에서 검증되었다. 신경망의 동작을 제어하고 디버깅하는 데 있어 신경 기하학적 구조가 핵심 메커니즘임을 제시한다.
Natural Language Autoencoders Neuronpedia는 AI 모델의 내부 작동을 탐색, 시각화, 조작할 수 있는 오픈소스 해석 가능성 플랫폼입니다. 이 플랫폼은 자연어 오토인코더, 회로 추적, 어시스턴트 축 등 다양한 도구와 기능을 제공하며, Google DeepMind, Anthropic, OpenAI 등 주요 연구진과 협력하여 최신 연구 결과와 모델 해석 도구를 공개합니다. API와 라이브러리를 통해
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/modelinterpretability