How this engine is built

No crawler of our own. A topic slice of Common Crawl, plus the feeds and APIs of sites that block crawlers entirely.

1

Harvest

The per-crawl CDX index is binary-searched, then only the needed WARC ranges are fetched — politely, five requests a second, days not minutes.

2

Index

A deterministic BM25 build: same pages in, byte-identical index out. Ranking adds source weight, page shape and, for jobs, freshness.

3

Serve

Title, a short attributed snippet, and a link out. Never a cached page. Removal requests are honoured on the next build.

The rules
Zero runtime dependencies · at most three results per domain, so one loud site can't own page one · honest coverage, gaps published · works with JavaScript disabled