#heritrix

Live, measured metrics for the hashtag #heritrix from the open social web. Every number carries a named source and the time it was fetched. Nothing is estimated.

hashtag.org network · sponsored

Own #heritrix

This #name is available to claim. It becomes your portal on the open agent web: this very page, a keyword you rank for by an open public stake, and a verifiable identity for AI agents. Nobody else sells a page like this for every #name.

$110.68/ year · 8-character #name
Claim #heritrix$110.68/yrBuy on hashtag.space (web3)
card via hashtag.org · tokens via hashtag.space
0
Uses / 7 days
Mastodon
0
Accounts / 7 days
Mastodon
2
Recent posts
Mastodon
~0/hr
Recent pace
Mastodon · last 2
0.5
Avg reactions / post
Mastodon · last 2
Reddit posts / month
Reddit search
Open-web mentions
hashtag.org Firehose

Day-by-day usage

measured · fosstodon.org (Mastodon public tags API) · fetched 2026-09-12 17:47 UTC
0
09-06
0
09-07
0
09-08
0
09-09
0
09-10
0
09-11
0
09-12

0 uses by 0 unique accounts across the window. Real per-day counts, not estimates. Newest bar is today so far.

Related hashtags

measured · fosstodon.org (Mastodon public search API) · fetched 2026-09-12 17:47 UTC

No related tags with measured usage found for #heritrix.

Live pulse

measured · fosstodon.org (Mastodon tag timeline) · fetched 2026-09-12 17:47 UTC

Everything below is measured over the latest 2 public posts (spanning ~550 hours).

Top of the latest posts

  • What are your favorite / the best #WebCrawlers for broad / #WebScale #crawling? I've built a list but am looking for anything I missed: https://github.com/davidshq/awesome-search-engines/blob/main/WebCrawlers.md Main options I've found incl

    Dave Mackey@[email protected]152023-04-16 12:34 UTCView post →
  • @elan also see #Heritrix and of course #StormCrawler as alternatives to #ApacheNutch

    Tim Allison@[email protected]002023-03-24 14:44 UTCView post →

What “heritrix” means

Wikipedia

Heritrix is a web crawler designed for web archiving. It was originally written in collaboration between the Internet Archive, National Library of Norway and National Library of Iceland. Heritrix is available under a free software license and written in Java. The main interface is accessible using a web browser, and there is a command-line tool that can optionally be used to initiate crawls.

Heritrix” on Wikipedia (CC BY-SA) →

#heritrix across platforms

every network with a public tag surface

Follow #heritrix straight to each platform’s own tag page. Where a platform publishes open data we measure it above; the rest lock their numbers behind paid APIs, so we link rather than guess.

Every number above is measured from a named public API at the shown fetch time. Nothing is estimated or extrapolated. Platforms that lock their data behind paid APIs are not shown. Agents: the same numbers, as JSON, at /api/hashtags/heritrix