collect_content() now caches fetched pages. A cache hit replays the stored provenance – the time the page was really fetched, the URL it really resolved to – rather than synthesising a fresh timestamp, so a cached row still says when it is from. Entries expire (cache_ttl), the cache is pruned to a size budget (cache_max_size), and a corrupt entry is a miss rather than an error. Nothing persists past the session unless you pass rdomains_cache_dir().rdomains_cache_dir() and cache_clear().collect_content() and pie_cat() take an archive_date, fetching each domain as it was on that date rather than as it is now. Set against a live run, that is how you measure whether a label has gone stale – cnn.com classifies as news from a 2020 capture and news today, at 0.966 and 0.709 confidence. Rows carry the realised snapshot_timestamp, which is the capture found and rarely the date asked for.collect_content() recovers pages from the Internet Archive when a host serves an anti-bot interstitial – detection plus a fallback rather than an evasion arms race. Recovered rows carry source = "archive" and a snapshot_timestamp, so the vintage is never hidden. A domain that no longer resolves is deliberately not recovered this way: answering “what was this” while looking like “what is this” is the staleness this package exists to surface.pie_cat(): fetches a domain’s homepage and classifies its current content with the piedomains model, rather than looking the domain up in a list written years ago. It reaches the model through the Python package via reticulate, so the four fragile pieces of the model input – the domain prefix, temperature scaling, label projection and the text cleaner – stay on the Python side rather than becoming a second place they can drift. parked and unavailable pages are labelled from the page itself and cost no model call.collect_content()’s max_bytes default is now 10 MB, matching piedomains. At 2 MB it rejected cnn.com – roughly 6 MB of homepage – as content_too_large while the server returned 200.ut1_cat() and get_ut1_data(): the UT-Capitole blacklists, the maintained successor to Shallalist. Where shalla_cat() answers from a list that stopped in January 2022, this one is updated continuously. Categories are UT1’s own and are reported verbatim, with ut1_usage saying whether UT1 maintains the list for blocking or for allowing – both describe content.collect_content() honours an explicit http:// or https:// in the input, keeping its port and path. Forcing https:// meant http-only hosts were unreachable.content_too_large rather than connection_error. curl aborts the transfer, so the limit surfaced as a request failure – and connection_error is retryable, so callers would have retried a too-large page forever.webfakes), covering redirects to private addresses, robots.txt refusal, oversized bodies, bot walls, and whether the cache actually prevents a second request.Two of the four lookup sources are no longer published: DMOZ closed in March 2017 and Shallalist stopped in 2022. Their labels were correct when assigned, but domains expire and change hands, so a lookup today can return the previous registrant’s category. Until now that answer was presented identically to one from a list updated last week.
source_vintage() reports every category source, when it was last published, whether it is still maintained, and its successor where one exists.shalla_cat(), dmoz_cat() and stevenblack_cat() now return a source_last_published column. For the two dead lists this is a constant; for Steven Black’s actively-maintained hosts file it is the fetched file’s own date.The sibling project piedomains measured what this confusion costs: its worst class disagreed with its own page content 71% of the time, and the cause was not bad annotation but roughly a decade between the label and the page. Only 60% of the domains it trained on still resolve.
New collect_content() fetches homepage HTML and text, so a domain can be classified on what it says today rather than on what a list said years ago. It returns one row per requested domain, never dropped, each carrying status, stage, error_code and retryable – so a transient failure is distinguishable from a permanent one and only the right rows get retried. fetch_error_codes() documents the closed set of reasons; fetch_report() summarises a run.
Supporting functions, all usable on their own if you already hold HTML:
page_signals() reports whether a page is an anti-bot interstitial, a domain-parking placeholder, a server’s “nothing here” page, or too thin to classify. Vendor presence alone is not a block: reddit, walmart and quora all serve real pages while embedding reCAPTCHA.html_text_content() extracts text, title, description and language.The crawler identifies itself as rdomains/<version> with a contact URL, obeys robots.txt including Crawl-delay, spaces requests to the same host, caps the response body, follows redirects by hand so every hop is re-validated, and refuses hosts resolving to private or link-local addresses.
Static HTML only – no headless browser, so a JavaScript-rendered page comes back thin and says so.
get_dmoz_data() built its output path with paste0(), so the documented default outdir = "." produced a hidden .dmoz_domain_category.csv that dmoz_cat() would then fail to find. It now uses file.path(), matching its two siblings.stevenblack_cat(use_file = NULL) re-downloaded roughly 4 MB on every call. The file is now cached for the session.virustotal_cat() to use VirusTotal API v3 (previously v2.0)virustotal_cat() implementation to properly extract categories from v3 API response structureclean_domains() - standardized domain cleaningvalidate_domains() - comprehensive input validationvalidate_data_file() - consistent file validationget_api_key() - unified API key retrievalbuild_categorization_prompt() - LLM prompt constructionapply_rate_limit() - rate limiting logic:: notation for imported functions (cleaner code, consistent with @importFrom)get_alexa_data() has been removed (service discontinued)virustotal_cat() parameter renamed from domain to domains for consistencyopenai_cat() and claude_cat() functions*_cat() functions for seamless integration