#commoncrawl

Live, measured metrics for the hashtag #commoncrawl from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.

hashtag.org network · sponsored

Own #commoncrawl

This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.

$10.85/ year · 11-character #name
Claim #commoncrawl$10.85/yr
Annual, renews each year
Buy on hashtag.space (web3)
one-timepay once, yours for life
card via hashtag.org · tokens via hashtag.space
0
Uses / 7 days
Mastodon
0
Accounts / 7 days
Mastodon
40
Recent posts
Mastodon
~0/hr
Recent pace
Mastodon · last 40
0.1
Avg reactions / post
Mastodon · last 40

Day-by-day usage

measured · fosstodon.org (Mastodon public tags API) · fetched 2026-07-27 14:47 UTC
0
07-21
0
07-22
0
07-23
0
07-24
0
07-25
0
07-26
0
07-27

0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.

Related hashtags

measured · fosstodon.org (Mastodon public search API) · fetched 2026-07-27 14:47 UTC

Live pulse

measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-07-27 14:47 UTC

Everything below is measured over the latest 40 public posts (spanning ~24459 hours).

Posting hours (UTC) — busiest: 16:00

00:0012:0023:00

Languages: English (36) · French (2) · German (2)

Avg boosts / post: 2

Top of the latest posts

  • Most generative AI models were trained on Common Crawl, a massive archive of web crawl data. Yet most people never heard of it. My new research studies Common Crawl in-depth and highlights its influence on LLM research and development #comm

    Stefan Baack (OLD)@[email protected]4512024-02-06 16:18 UTCView post →
  • Woohoo! #CommonCrawl has bumped the truncation limit to 5MiB from 1MiB, which means that the number of truncated PDFs has gone from ~26% down to 7%. This is critical for other binary formats as well. Thank you #CommonCrawl! https://commoncr

    Tim Allison@[email protected]132025-04-08 11:58 UTCView post →
  • Part of the trust model for AI, is users understanding that the data is a broad representation of the real world web. Sharp, well considered insights from the #CommonCrawl team #POSAIS26

    o lаvrоvsky@[email protected]002026-07-10 12:50 UTCView post →

Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/commoncrawl