No description
When the multi-crawler fetch can't get article text (FT, WSJ, etc. return 403/JS-only shells to every crawler), fall back to the Internet Archive: resolve the closest snapshot via the availability API, fetch the raw id_ version, and extract THAT into the Radical Reader with a 'via Wayback Machine, archived DATE' source note. archive.today can't be used server-side (429 + captcha on datacenter IPs) so it stays as the user-side button. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|---|---|---|
| src | ||
| .gitignore | ||
| package-lock.json | ||
| package.json | ||
| wrangler.toml | ||