#commoncrawl
Live, measured metrics for the hashtag #commoncrawl from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.
Own #commoncrawl
This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.
Day-by-day usage
measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 14:47 UTC0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.
Related hashtags
measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 14:47 UTCLive pulse
measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 14:47 UTCEverything below is measured over the latest 40 public posts (spanning ~24459 hours).
Posting hours (UTC) — busiest: 16:00
Languages: English (36) · French (2) · German (2)
Avg boosts / post: 2
Top of the latest posts
Most generative AI models were trained on Common Crawl, a massive archive of web crawl data. Yet most people never heard of it. My new research studies Common Crawl in-depth and highlights its influence on LLM research and development #comm
Woohoo! #CommonCrawl has bumped the truncation limit to 5MiB from 1MiB, which means that the number of truncated PDFs has gone from ~26% down to 7%. This is critical for other binary formats as well. Thank you #CommonCrawl! https://commoncr
Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web. Sharp, well considered insights from the #CommonCrawl team #POSAIS26
Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/commoncrawl