Two things are worth establishing before any discussion of how to scrape a website.
The first is legal, and it is the reason most guidance on this subject is unsafe reading for a Canadian business. Almost every web scraping article, tutorial and vendor page assumes American law. In the United States, fair use is an open-ended standard that courts apply case by case, and several recent decisions have gone in favour of large-scale data collection. Canada has no equivalent. Fair dealing under the Copyright Act is a closed list of specific purposes, and Canada has no text and data mining exception at all. The federal government has consulted on introducing one; it has not done so. A Canadian business scraping copyrighted material for a commercial purpose that does not fall within the enumerated categories is on considerably weaker ground than an American business doing the same thing, and no amount of "web scraping is legal" content written in California changes that.
The second is about the word "better." A large part of the scraping industry defines better as harder to detect: rotate through residential IP addresses, spoof browser fingerprints, defeat the CAPTCHA, make your bot indistinguishable from a person. That is the business model of several of the tools you will find recommended, and this article does not cover it. Not out of squeamishness, but because it is a bad engineering strategy as well as a legally exposed one. It produces systems that are expensive, fragile, permanently one step from breaking, and dependent on infrastructure of dubious provenance.
The better ways are duller and they work. Check whether the data is already available in a form you are welcome to take. Read the terms before you write the code. Identify yourself honestly. Make fewer requests. Cache what you already have. Fail gracefully. Store less than you could. These practices are more reliable, cheaper to maintain, and far less likely to end in a letter from someone's lawyer.
This guide covers both, because for most readers here both apply. The first half is about collecting data from other people's sites. The second half is about the other side of the same activity: what to do when your own site is the one being scraped, which for a business running a website is by now a near-certainty rather than a possibility.
What web scraping actually is
Web scraping is the automated extraction of data from web pages. A program requests a page as a browser would, receives the HTML, and pulls specified values out of it.
Three related terms get used interchangeably and should not be:
- Crawling is the discovery of pages by following links or reading sitemaps. Search engines crawl. A crawler finds URLs; a scraper extracts data from them. Most real systems do both.
- Scraping is extraction from the retrieved page.
- Using an API is requesting data from an interface the publisher built for that purpose, in a structured format, under stated terms. This is not scraping, and where it exists it is almost always the better option.
The distinction that matters most in practice is between taking data a publisher has chosen to make available for programmatic use, and taking data by parsing pages built for human readers. The first is cooperative. The second is unilateral, and everything difficult about scraping — the legal exposure, the fragility, the blocking — flows from that.
Start further up the ladder
This is the single most useful section in this guide, and it is the one most tutorials skip entirely because it does not involve writing any code.
Before building a scraper, work down this list. Stop at the first option that gets you what you need.
- An official API. Structured, documented, stable, and covered by terms you can actually read. It will not break when the site is redesigned. Many organisations have one and do not advertise it prominently; check the developer or documentation section, and check whether the site's own front end calls a JSON endpoint you could use directly.
- A bulk download or data dump. Some publishers offer the whole dataset as a file. Wikipedia is the well-known example, but plenty of government and research bodies do the same. One download beats ten thousand requests for everyone involved.
- An open data portal. For Canadian businesses this is genuinely underused. The federal Open Government Portal at open.canada.ca carries tens of thousands of datasets under an open licence. Statistics Canada publishes extensive data tables with programmatic access. Every province and most large municipalities run open data portals. Business registries, census and demographic data, economic indicators, geographic and property data, transit data, health statistics — a surprising amount of what businesses try to scrape is already sitting in a portal, cleaner, licensed for reuse, and free.
- A licensed commercial feed. If the data has real commercial value to you, someone may sell it properly. Paying for a licensed feed removes the legal question, removes the maintenance burden, and is often cheaper than the engineering time a scraper consumes over a year.
- Ask. This one gets dismissed and it should not. A short email explaining who you are, what data you need, and why, has a better success rate than people expect — particularly with smaller organisations, research bodies, and anyone who would rather give you a clean export than have you hammering their server. Sometimes the answer is a CSV by return email.
- Scraping. Last, because it is the only option on this list where you are taking rather than receiving, and the only one that carries the full set of legal, technical and maintenance costs.
A useful way to hold this: scraping is what you do when cooperation has failed or was never possible. Treating it as the default rather than the fallback is how businesses end up with an expensive, brittle system extracting data they could have downloaded.
The Canadian legal position
This section is longer than the equivalent in most guides because the position is genuinely different here and genuinely consequential. It is general information, not legal advice, and the closing part of this section covers when to actually get advice.
Copyright: no text and data mining exception, and a closed list
Copyright is the first exposure, and the Canadian position is more restrictive than most people assume.
What copyright does and does not cover. Facts are not protected. A price, a date, a temperature, an address — these are information, not expression, and copyright does not reach them. What is protected is the expression: the text of a product description, a photograph, an article, a review. Also potentially protected is the selection and arrangement of a compilation, where skill and judgment went into assembling it, even if the individual items are unprotected facts.
So scraping a page and extracting the numeric price is a different act from scraping the page and republishing the description.
Fair dealing is not fair use. This is the crucial difference. American fair use is an open-ended, four-factor standard that a court can apply to any purpose. Canadian fair dealing under section 29 of the Copyright Act is available only for enumerated purposes: research, private study, education, parody, satire, criticism or review, and news reporting. Canadian courts interpret those purposes broadly, but a use that does not fall within one of them is simply outside the exception, however reasonable it might seem.
There is no TDM exception. The Copyright Act does not address text and data mining. Innovation, Science and Economic Development Canada consulted on a modern copyright framework for AI, including whether to introduce a TDM exception, and submissions from legal academics and the Canadian Bar Association went both ways. No exception has been enacted. Section 30.71 provides a narrow exception for temporary reproductions that are an essential part of a technological process, which is not a general licence for building datasets.
The practical consequence: a Canadian business scraping copyrighted expression for a commercial product has a weaker defence than an American one. If your purpose is genuine research or private study, fair dealing may well cover you. If your purpose is assembling a commercial dataset or training a commercial model, the enumerated purposes may not fit, and there is no TDM exception to fall back on.
Privacy: "publicly available" is a narrow legal term
If what you are collecting includes personal information — names, contact details, photographs, profiles, reviews attributable to individuals — you are in privacy law as well, and this is where Canadian businesses most often get a nasty surprise.
The Personal Information Protection and Electronic Documents Act requires consent to collect personal information in the course of commercial activity, with substantially similar provincial legislation in Quebec, British Columbia and Alberta. There is an exception for "publicly available" information.
That exception does not mean what it sounds like. "Publicly available" is not a common-sense description of anything findable online. It is a term defined by regulation, listing specific categories — information in a publication such as a magazine, book or newspaper, certain public registries, and similar. The internet is not among the enumerated sources.
This was tested directly. In 2021 the federal Privacy Commissioner, together with the Quebec, British Columbia and Alberta commissioners, published the findings of a joint investigation into Clearview AI, which had built a facial recognition database by scraping billions of images from public websites and social media. Clearview's central argument was that the information was publicly available and therefore exempt from consent. The commissioners rejected it, concluding the exception must be read narrowly and does not extend to personal information gathered from online sources including social media. They also found the purpose was not one a reasonable person would consider appropriate. The provincial commissioners, which have order-making powers the federal commissioner lacks under PIPEDA, ordered Clearview to stop.
But the law here is genuinely unsettled, and honesty requires saying so. Clearview sought judicial review. In 2025 the Alberta Court of King's Bench upheld the Commissioner's narrow interpretation of "publicly available" as reasonable — and then found that the resulting provisions of Alberta's Personal Information Protection Act unjustifiably infringe the freedom of expression guaranteed by section 2(b) of the Charter, on the basis that where obtaining consent is impractical, as in mass online collection, the consent requirement operates as a complete prohibition. Clearview did not pursue the constitutional argument in British Columbia. A Quebec judicial review appears to remain outstanding.
So the position as it stands: regulators read the exception narrowly, an Alberta court has found the resulting prohibition constitutionally overbroad under that province's statute, other provinces have not ruled the same way, and the federal position under PIPEDA has not been resolved by a court. Anyone telling you confidently that scraping personal information in Canada is either clearly fine or clearly prohibited is overstating the state of the law.
What follows practically is not paralysis but proportion. If your scraping touches personal information, this is the area to take advice on, and the area where collecting less is the cheapest risk reduction available.
Quebec's Law 25 raises the bar further for organisations handling information about Quebec residents, including assessment requirements before transferring personal information outside the province.
Contract, and the terms of service question
Most websites have terms of service, and most prohibit automated collection.
Whether those terms bind you depends on whether a contract was formed. Terms you actively agreed to by creating an account are on much firmer ground than terms sitting in a footer link on a page you merely visited. Canadian courts have enforced online terms in various circumstances, and the analysis is fact-specific.
The practical points:
- Scraping behind a login is materially riskier than scraping public pages, because you almost certainly accepted terms to get the account, and you may be using credentials in a way those terms forbid.
- Terms of service breach is a contract matter, which is a different and generally lesser exposure than copyright infringement or a privacy complaint — but it is a real one, and it is the most common basis for the cease-and-desist letters scrapers actually receive.
- Read them. It takes five minutes and it tells you what you are dealing with. Some sites explicitly permit non-commercial research crawling. Some publish a crawling policy. You cannot make a considered decision about terms you have not looked at.
Other considerations, briefly
The Criminal Code. Section 342.1 addresses unauthorised use of a computer. It is aimed at conduct considerably more serious than reading public web pages, and it is not a provision that ordinary scraping of publicly accessible content would be expected to engage. It is worth knowing it exists, particularly if any part of your approach involves circumventing access controls or authentication, which moves you into genuinely different territory.
Volume and disruption. A scraper that degrades a site's availability for its actual users is a different proposition from one that fetches a page every few seconds. Aggressive collection that amounts to a denial of service is both an obvious legal risk and simply antisocial.
Competition and misleading representation. If you republish scraped data in a way that misrepresents its source or suggests an affiliation, the Competition Act's provisions on misleading representations come into view.
Risk, in proportion
Not all scraping carries similar exposure. A rough tiering:
| What you collect | What you do with it | Exposure |
| Non-personal facts (prices, availability, specifications) | Internal analysis only | Low. Facts are not copyrightable; no personal information involved |
| Non-personal facts | Commercial product built on the data | Moderate. Compilation rights and terms of service become live questions |
| Copyrighted expression (text, images, reviews) | Internal research | Moderate. May fall within fair dealing for research, depending on the use |
| Copyrighted expression | Republication or a commercial product | High. No TDM exception, closed-list fair dealing, likely infringement |
| Personal information | Any purpose | High and unsettled. The Clearview line of authority applies and the law is in flux |
| Anything behind authentication | Any purpose | High. Contract almost certainly formed; access control questions arise |
When to get advice rather than reading an article
Speak to a Canadian lawyer before proceeding if any of the following apply:
- The data includes personal information about identifiable individuals.
- You intend to build a commercial product or service on the collected data.
- You will republish content rather than derive facts or aggregates from it.
- The collection is at significant scale, or ongoing rather than one-off.
- You are scraping behind a login, or from a site whose terms explicitly prohibit it.
- You are collecting material to train a model.
- The source is a competitor.
If none of those apply — you are pulling public, non-personal facts at modest volume for your own internal use — you are in the lowest-risk category on that table, and the rest of this guide is mostly about doing it competently.
Doing it well: the technical practice
The engineering advice below is what actually distinguishes a scraper that runs quietly for two years from one that breaks every fortnight and gets blocked. The code is Python, because that is what most of this is written in, and deliberately short — these are patterns rather than a library.
Before you write anything, look at four files
/robots.txt — what the site asks crawlers not to fetch, and sometimes how fast.
/sitemap.xml — often listed in robots.txt. If it exists, it is a complete list of the URLs the publisher wants found, which is almost always better than discovering pages by crawling links. It frequently carries last-modified dates too, which tells you what has changed since your last run.
The terms of service. Five minutes.
The page source. Specifically, look for a <script type="application/ld+json"> block and check whether the page's own front end fetches data from a JSON endpoint. Both give you structured data instead of parsed HTML, and both are far more stable.
Respect robots.txt properly
robots.txt was standardised as RFC 9309 in 2022, after twenty-eight years as a de facto convention. It is a request rather than an access control, and ignoring it is not itself unlawful — but it is the clearest available statement of what the site owner wants, and disregarding it undermines any argument that your collection was reasonable.
Python has a parser in the standard library:
| from urllib.robotparser import RobotFileParser
UA = "AcmeResearchBot/1.0 (+https://acme.ca/bot)"
rp = RobotFileParser() rp.set_url("https://example.ca/robots.txt") rp.read()
if not rp.can_fetch(UA, target_url): # Disallowed. Do not fetch it. return None
# Honour Crawl-delay if the site sets one; otherwise pick a sane default. delay = rp.crawl_delay(UA) or 2.0 |
Two notes. Fetch robots.txt once per run and cache it, rather than before every request. And Crawl-delay is a widely-honoured extension rather than part of the standard — Google ignores it, but if a site has set it, that is an explicit statement of tolerance and you should follow it.
Identify yourself honestly
The User-Agent string is where you say who you are. The industry norm is to lie: copy a real Chrome string so you look like a person.
Do the opposite. Name your bot, include a version, and include a URL or email where the site owner can reach you:
| UA = "AcmeResearchBot/1.0 (+https://acme.ca/bot)"
session.headers.update({ "User-Agent": UA, "Accept": "text/html,application/xhtml+xml", "Accept-Language": "en-CA,en;q=0.9", "From": "data@acme.ca", }) |
That page at /bot should say who you are, what you collect, why, and how to ask you to stop.
This feels counterintuitive and it pays off in three ways. A site owner investigating unusual traffic can contact you instead of blocking you. Some sites allowlist identified, well-behaved bots and block unidentified ones. And if you ever have to explain your collection to a regulator or a court, "we identified ourselves and provided a contact address" is a materially better position than "we disguised our traffic as a human browser."
Rate limit yourself before anyone does it for you
Almost all blocking is a response to request volume. Slow down and most of the problem disappears.
Sensible defaults: one request every one to two seconds to a single host, one concurrent connection per host, and run overnight in the target's local time zone if the job is large and not time-critical.
| import time
last_request = 0.0
def throttled_get(url, session, delay): global last_request elapsed = time.monotonic() - last_request if elapsed < delay: time.sleep(delay - elapsed) last_request = time.monotonic() return session.get(url, timeout=20) |
If you are collecting across many different sites, parallelise across hosts rather than within them. Ten sites at one request per second each is polite. One site at ten requests per second is not.
Handle failure the way the protocol intends
Most scrapers treat any non-200 response as an error and either retry immediately or give up. Both are wrong, and the difference between a scraper that gets blocked and one that does not is often just this.
The two responses that matter: 429 Too Many Requests means slow down. 503 Service Unavailable means the server is struggling. Either may carry a Retry-After header telling you exactly how long to wait, and honouring it is the single most useful thing you can do.
| import time, random
def polite_get(url, session, ua, max_tries=5): for attempt in range(max_tries): r = session.get(url, headers={"User-Agent": ua}, timeout=20)
if r.status_code == 200: return r
if r.status_code in (429, 503): retry_after = r.headers.get("Retry-After", "") if retry_after.isdigit(): wait = float(retry_after) else: wait = (2 ** attempt) + random.random() time.sleep(wait) continue
if 400 <= r.status_code < 500: # 404, 403, 410: the page is not coming. Don't hammer it. return None
time.sleep((2 ** attempt) + random.random())
return None |
The random component matters. Without jitter, a fleet of workers that all back off by the same amount retries in synchronised waves, which looks like an attack.
Use conditional requests, which almost nobody does
This is the highest-value technique in this guide and it is absent from every tutorial in the reference list.
HTTP has a built-in mechanism for asking "has this changed since I last looked?" If the server supports it, and most do, an unchanged page costs you a 304 Not Modified with no body — a few hundred bytes instead of a few hundred kilobytes.
Store the ETag and Last-Modified values the server sent, then send them back:
| def fetch_if_changed(url, session, ua, cache): headers = {"User-Agent": ua} entry = cache.get(url, {})
if entry.get("etag"): headers["If-None-Match"] = entry["etag"] if entry.get("last_modified"): headers["If-Modified-Since"] = entry["last_modified"]
r = session.get(url, headers=headers, timeout=20)
if r.status_code == 304: return entry["body"], False # unchanged
if r.status_code == 200: cache[url] = { "etag": r.headers.get("ETag"), "last_modified": r.headers.get("Last-Modified"), "body": r.text, } return r.text, True # changed
return None, False |
On a recurring job over a site where most pages rarely change, this typically cuts transferred data by an order of magnitude. It reduces your bandwidth, reduces the target's load, makes your job faster, and makes you dramatically less likely to be noticed as a problem. If a site publishes a sitemap with <lastmod> dates, combine the two: skip anything whose date has not moved, and use conditional requests for the rest.
Parse defensively, and prefer structured data
CSS selectors break. A site redesign, or even a minor template change, silently turns your extraction into empty strings. The most common scraper failure is not being blocked; it is quietly collecting nothing for three weeks.
Where a page carries JSON-LD structured data — which most e-commerce, article and event pages now do, because search engines reward it — take that instead. It is a stable, documented format that changes far less often than markup:
| import json from bs4 import BeautifulSoup
def extract_jsonld(html): soup = BeautifulSoup(html, "html.parser") blocks = [] for tag in soup.find_all("script", type="application/ld+json"): try: data = json.loads(tag.string or "") except (json.JSONDecodeError, TypeError): continue blocks.extend(data if isinstance(data, list) else [data]) return blocks |
Whatever you parse, validate it. If a field you expect is missing or a value falls outside a plausible range, log it loudly rather than writing a null and moving on:
| def validate(record): problems = [] if not record.get("name"): problems.append("missing name") price = record.get("price") if price is None or not (0 < float(price) < 1_000_000): problems.append(f"implausible price: {price}") return problems |
Then alert on the rate. If the proportion of records failing validation jumps from one per cent to forty, something changed and you want to know today rather than at the end of the quarter.
If you get blocked, the constructive options
Since this guide does not cover evasion, it should say what to do instead, because being blocked is the most common practical problem readers actually have.
Work through these in order.
Check whether you deserved it. Look at your own logs before anything else. How many requests did you make, over what period, and did you ignore a 429 or a Crawl-delay? A surprising proportion of blocks are the direct, proportionate response to a scraper that was hammering a small server. Fix the behaviour and the block frequently lifts on its own, because most rate-based blocks are temporary.
Reduce scope rather than volume alone. Ask what you actually need. Fetching an entire catalogue nightly when you need forty products weekly is the real problem; slowing it down only makes it a slower version of the same problem. Narrowing the target list is usually a larger reduction than any timing change.
Contact the site owner. This is the option people skip and it has a better success rate than expected — which is exactly why an honest User-Agent with a contact URL is worth having, because it means they may reach out to you first. Explain who you are, what you need, how often, and offer to work within whatever limit suits them. Smaller organisations in particular would often rather give you a scheduled export than have you guessing at their tolerance.
Ask about an API or a data agreement. If the data matters commercially to you, it may be worth something to them. Plenty of API programmes and data licences exist because someone asked.
Accept that the answer may be no. A site owner who has said no, technically or explicitly, has made a decision they are entitled to make. Routing around it moves you from a technical disagreement into a position that is harder to defend on every dimension — legally, commercially, and in terms of the engineering time you will spend maintaining the workaround.
The pattern worth noticing: every option on that list is cheaper than building and maintaining an evasion layer, and none of them stops working when the target changes its detection.
The hard part nobody warns you about: data quality
Ask anyone who has run a scraping project for a year what consumed the time, and it will not be fetching pages. It will be making the collected data usable.
This is genuinely the least-covered aspect of the subject and the most likely to determine whether the project delivers anything.
Entity matching. If you are collecting the same items across multiple sources, you have to decide which records refer to the same thing. Product names differ between retailers. Model numbers are formatted inconsistently or omitted. One site lists a variant as a separate product, another as an option on one page. There is no clean solution — the workable approach is to match on the most stable identifier available (a GTIN, an ISBN, a manufacturer part number), fall back to fuzzy matching on normalised names, and keep a manually curated mapping table for the cases that matter most. Expect the mapping table to be permanent.
Normalisation. The same value arrives in many shapes. Prices with and without tax, in different currencies, with different thousands separators. Dates in three formats. Units mixed between metric and imperial. Whitespace, HTML entities, and non-breaking spaces embedded in numbers. Normalise at ingest, store the normalised value alongside the raw string, and never discard the raw string — when a figure looks wrong six months later, the original text is how you find out why.
Deduplication. The same page reachable at several URLs, pagination that repeats items across pages, and re-runs that re-insert what you already have. Key your storage on something stable rather than on the URL you happened to fetch, and make inserts idempotent so a partially failed run can simply be repeated.
Change detection versus churn. Distinguishing a real change from noise is harder than it sounds. A price that appears to move by a cent every run is more likely a rounding or tax-calculation difference than an actual repricing. Decide what magnitude of change is meaningful and only record those, or your change log becomes useless.
Provenance and timestamps. Every record should carry the source URL and the time it was fetched. Without that you cannot audit a number, cannot explain a discrepancy, cannot honour a deletion request, and cannot tell stale data from current.
A realistic budget for a data collection project is roughly a third of the effort on fetching and two thirds on everything after it. Projects that assume the reverse are the ones that end up with a large database nobody trusts.
Avoid headless browsers unless you have to
Driving a real browser through Playwright or Selenium handles JavaScript-rendered content, and it is sometimes the only option.
It is also, compared with plain HTTP requests, roughly one to two orders of magnitude more expensive in memory and CPU, considerably slower, and much more fragile. Before reaching for it, check whether the data is available through the JSON endpoint the page's own JavaScript is calling. Open your browser's network tab, watch what the page fetches, and quite often you will find a clean API response you can request directly — smaller, faster, structured, and stable.
Use a headless browser when there is genuinely no other route. Not as the default.
Where you run it matters, and this is where hosting comes in
This is a practical point that almost no scraping tutorial makes, and it causes real problems for small businesses.
Do not run scrapers from the same hosting account as your website. Three reasons, and any one of them is sufficient:
Shared hosting terms generally prohibit it. Sustained outbound automated requests are resource-intensive in a way shared plans are not provisioned for, and most acceptable-use policies say so. Check yours before assuming otherwise.
You are sharing an IP address, and reputation with it. On shared hosting your outbound traffic leaves from an address shared with other accounts. If your scraper attracts blocks or lands on a reputation list, that affects everyone on the address — and, more to the point, it affects you. The same shared IP handles your outbound mail. A scraper that gets your server's address flagged can take your transactional email down with it.
Resource contention hits your own site first. A scraping job consuming CPU and memory on the same account serving your website degrades the thing that actually makes you money. The symptom is a slower time to first byte on your own pages, which affects both visitors and how consistently crawlers can fetch you.
The right arrangement is separation. Run collection on infrastructure distinct from the infrastructure serving your website — VPS hosting with root access gives you a separate environment with its own IP, its own resources, and the control to install whatever you need without a shared-hosting policy in the way. For larger or continuous collection, dedicated server hosting removes co-tenancy from the picture entirely. Either way, the principle is that your data collection should not be able to damage your public presence.
Two related considerations. Proximity still matters: if you are collecting predominantly from Canadian sources, running from Canadian infrastructure means shorter round trips and faster jobs. And where the collected data comes to rest is a data residency question, particularly if it includes personal information — hosting the storage in Canadian data centres in Vancouver and Toronto keeps that question simple to answer.
Store less than you can
Scrapers make it trivially easy to hoard. Resist it, because every field you keep is a field you are accountable for.
- Collect only what you need for the stated purpose. This is not merely tidiness; data minimisation is a principle of Canadian privacy law and the most effective risk reduction available to you.
- Strip personal information at ingest if you do not need it. Do not store it and plan to filter later.
- Set a retention period and enforce it with a scheduled job, not an intention.
- Keep provenance. Record the source URL and fetch timestamp with every record. When someone asks where a figure came from, or asks you to delete their information, you will need it.
- Secure it. A scraped dataset sitting in an unsecured bucket or an unencrypted database is a breach waiting to be reported.
Expect maintenance
A scraper is not a finished artefact. Sites change, and the honest planning assumption is that any given scraper will need attention several times a year.
Budget for it, monitor for silent failure rather than only for crashes, keep the extraction logic separate from the fetching logic so a markup change is a small fix, and review annually whether the thing is still worth maintaining. A surprising number of scrapers outlive the question they were built to answer.
The other side: when your site is the one being scraped
If you run a website, this half is the one that applies to you. Unwanted automated traffic is now a routine operational cost, and it arrives whether or not you have thought about it.
How to tell it is happening
The signals, roughly in order of how obvious they are:
- Bandwidth that does not match your visitor numbers. Analytics counts humans; bandwidth counts everything. A widening gap is the clearest early indicator.
- Requests in a pattern no human produces — perfectly regular intervals, sequential URL traversal, an entire category fetched in ninety seconds, activity at a constant rate through the night.
- A single address or narrow range accounting for a disproportionate share of requests.
- Requests for pages no human navigates to, such as deep pagination or expired product URLs.
- Missing or unusual headers. No Accept-Language, no referrer, a User-Agent that is either absent, obviously a library default, or an implausibly old browser version.
- Server load with no corresponding business activity. Elevated CPU, slower responses, and no increase in orders or enquiries.
Your access logs hold all of this. A quick look at requests per IP over the last day is usually enough to see whether you have a problem.
What it actually costs
Worth being concrete, because it is easy to dismiss as harmless.
Bandwidth and resources, which on a metered plan is a direct bill, and on any plan is capacity not available to customers.
Performance for real visitors. Bot load consumes the same CPU, memory and database connections your customers need. It shows up as a slower time to first byte, which sets the floor under your Core Web Vitals field data and affects how reliably search crawlers can fetch your pages.
Skewed analytics. Bot traffic that gets counted corrupts the numbers you make decisions on.
Competitive exposure, if a competitor is monitoring your prices, stock levels or new listings in real time.
Content taken and republished, which at minimum creates duplicate content competing with your originals and at worst is straightforward infringement.
Discovery of things you did not mean to expose. Systematic crawling finds the staging subdomain, the old directory, the file nobody remembered.
robots.txt: necessary, and not a control
Publish a sensible robots.txt. Understand what it does: it tells cooperative crawlers what you would prefer. It is a signal, not a lock. Well-behaved bots honour it; bots you actually want to stop will ignore it.
| # Search engines: full access, with a courtesy delay for the rest User-agent: * Disallow: /cart/ Disallow: /checkout/ Disallow: /my-account/ Disallow: /search Crawl-delay: 10
# AI training crawlers User-agent: GPTBot Disallow: /
User-agent: CCBot Disallow: /
Sitemap: https://example.ca/sitemap.xml |
Three things to know about that file. Crawl-delay is not part of the standard and Google ignores it, though many other crawlers honour it. Blocking AI training crawlers by name only works for the ones that publish a name and respect the file, and the list changes, so it needs periodic review. And never use `robots.txt` to hide anything sensitive — it is a public file listing the paths you would rather people did not visit, which is an invitation to the sort of person who reads robots.txt files.
You will also see llms.txt proposed as a convention for guiding AI systems. It is a proposal with partial and inconsistent adoption rather than a standard, and it is not currently a reliable control. Publishing one costs nothing; relying on it would be a mistake.
Rate limiting at the reverse proxy
The most effective single measure, and it belongs in front of your application rather than inside it, so that limited requests never consume application resources at all. In Nginx:
| # Define the zone once, in the http block limit_req_zone $binary_remote_addr zone=general:10m rate=60r/m; limit_req_zone $binary_remote_addr zone=expensive:10m rate=10r/m;
server { location / { limit_req zone=general burst=20 nodelay; limit_req_status 429; # ... }
# Search and other costly endpoints get a tighter limit location /search { limit_req zone=expensive burst=5 nodelay; limit_req_status 429; # ... } } |
Points worth noting. Returning 429 rather than 403 is correct and useful: it tells a well-built client to slow down, and a well-built client will. The burst parameter allows short legitimate spikes, which matters because a real visitor loading a page fires many requests at once. Set limits generously at first and tighten while watching for false positives — a limit that catches real customers is worse than the bots. And configuring this requires access to the server layer, which shared hosting generally does not provide; it is one of the practical reasons businesses move to a VPS. If you are not sure what sits in front of your application, our guide to what a reverse proxy actually does covers that layer.
Layers beyond rate limiting
A web application firewall inspects requests before your application sees them and can filter on patterns, reputation and behaviour. Most managed hosting and all CDNs offer one.
Bot management services score traffic on behavioural signals rather than simple rules. They are more capable than rate limiting and more expensive, and they carry a false-positive risk worth testing before enforcing.
Graduated responses. Blocking outright is one option among several, and often not the best. You can slow a suspected bot down, serve it a cached response so it costs you almost nothing, require a session before serving expensive endpoints, or return a reduced version of the page. Slowing is frequently better than blocking, because a blocked scraper adapts while a throttled one often just leaves.
Do not put your prices in a JSON endpoint with no rate limit. A great deal of scraping is easy because a site's own front end exposes a convenient unauthenticated API. Rate limit those endpoints as tightly as you can.
Do not break search engines
This deserves its own warning, because it is the way this work most commonly goes wrong.
Aggressive bot blocking that catches Googlebot will damage your indexation, which is a far larger problem than the scraping you were trying to stop.
The safe way to distinguish a real search crawler from something claiming to be one is not to check the User-Agent string, which anyone can set. It is forward-confirmed reverse DNS: look up the requesting IP address, confirm it resolves to a hostname in the crawler's documented domain, then resolve that hostname and confirm it returns the same IP.
| $ host 66.249.66.1 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com.
$ host crawl-66-249-66-1.googlebot.com crawl-66-249-66-1.googlebot.com has address 66.249.66.1 |
Both directions must agree. Google publishes its crawler IP ranges as well, and Bing offers an equivalent verification tool.
Practically: allowlist verified search crawlers before applying any rate limits, and after any change to your bot rules, check Search Console's crawl stats for a rise in failed fetches. A blocking rule that quietly degrades your crawl rate is expensive in a way that takes months to show up in traffic.
If your content is being republished
Technical measures do not help once content is already taken. What does:
- Document it. Screenshots with dates, archived copies, the URLs involved.
- Prove originality. Your publication dates, your Search Console history, your version control.
- Ask first. A polite note to the site owner resolves a meaningful share of cases, particularly where a contractor or a plugin is the actual culprit.
- Escalate to the host. Every hosting provider has an abuse contact, and most act on well-documented complaints about infringing content.
- Search engine removal. Google and Bing both accept copyright removal requests for infringing copies.
- Take advice before sending anything that reads as a legal threat.
And a practical note on the SEO fear: a scraped copy of your page does not usually outrank the original. Search engines are generally competent at identifying the canonical source, and your original has age, links and a crawl history the copy does not. It is worth addressing, but it is rarely the emergency it feels like.
Five situations, and what each should actually do
A market research consultancy in Toronto
Collecting: industry pricing and product availability across several dozen public websites, weekly.
The right approach: check for open data and APIs first, because a good deal of industry data is published. For the rest, one request every two seconds, conditional requests so unchanged pages cost nothing, an honest bot identity, and extraction restricted to factual values rather than copying descriptions.
The risk position: low, and the reason is worth naming. Non-personal facts, internal analysis, no republication. This is the tier where scraping is genuinely uncontroversial. Keep it there by not drifting into collecting reviews with author names attached.
An online retailer monitoring competitor prices
Collecting: competitor prices on a few hundred matched products, daily.
The right approach: narrow and shallow. Match the product list first and fetch only those pages rather than crawling entire catalogues. Conditional requests. Overnight. And run it from separate infrastructure, because a retailer whose scraper gets its hosting IP flagged has traded a pricing advantage for an email deliverability problem.
The risk position: moderate, and higher than the consultancy's for one reason — the source is a competitor. Terms of service become a live question, and a competitor is the party most likely to actually pursue it. Take advice if this is more than incidental to your business.
The mirror image: you are also the target. Your own prices are being monitored. Rate limit your product endpoints.
A publisher whose articles are being taken by AI crawlers
The situation: original editorial content appearing in AI-generated answers, and bandwidth rising without traffic rising.
The right approach: decide the policy before the technology. Some publishers want the visibility and the citations; some want to be excluded; some want to license. These are commercially different positions and the technical implementation follows from whichever you pick.
If exclusion is the answer: name the crawlers that publish names in robots.txt, review that list quarterly because it changes, rate limit at the proxy for the ones that ignore it, and accept that neither is complete. If licensing is the answer, that is a commercial conversation rather than a configuration change.
The trap: blocking so broadly that search crawlers get caught. Verify by reverse DNS and allowlist before enforcing.
A design agency hosting client sites
The situation: managing bot traffic across many small sites, none individually high-value.
The right approach: standardise. One sensible robots.txt template, one set of rate limit rules applied at the shared proxy layer, one monitoring approach so you notice a problem on any client site without watching all of them. Staging environments locked down and excluded from indexing — systematic crawling finds staging subdomains, and a client discovering their unreleased redesign in a search result is an awkward conversation.
A research group at a Canadian university
Collecting: publicly posted text for a linguistics study.
The right approach: the strongest fair dealing position of anyone on this list, since research is an enumerated purpose. That is a real advantage, and it is not unlimited — it covers the copyright question, not the privacy one.
The critical distinction: if the corpus contains personal information, PIPEDA and provincial legislation apply regardless of the research purpose, and the Clearview line of authority is directly relevant. University research ethics review exists for this and should be engaged early rather than after collection. Aggregate and de-identify at ingest where the research design permits.
Mistakes worth avoiding
On the collecting side:
- Assuming American law applies. Canada has no TDM exception and fair dealing is a closed list. The confident "scraping is legal" content is written for a different jurisdiction.
- Treating "publicly available" as a plain-English phrase. It is a narrowly defined statutory term, and Canadian regulators have found it does not extend to online sources.
- Skipping the ladder. Building a scraper for data sitting in an open data portal, or behind an API, or available as a download.
- Disguising the bot. Fake browser user agents remove the site owner's ability to contact you and remove your ability to say you were transparent.
- Never using conditional requests. Refetching unchanged pages indefinitely, wasting your bandwidth and theirs.
- Retrying immediately on 429. The server asked you to slow down and you sped up.
- Running it on the same hosting as your website. Shared IP reputation, resource contention, and probably an acceptable-use breach.
- Silent extraction failure. No validation, no alerting, three weeks of empty fields.
- Keeping every field because storage is cheap, which converts a low-risk collection into a high-risk one.
- Reaching for a headless browser first, when the page's own JSON endpoint would have done.
On the receiving side:
- Assuming `robots.txt` stops anything. It is a request to cooperative clients.
- Putting sensitive paths in `robots.txt`, which publishes a list of what you would rather people did not find.
- Blocking on `User-Agent` alone, which is trivially forged in both directions.
- Catching Googlebot. The single most damaging mistake in this half of the article. Verify by reverse DNS and allowlist first.
- An unrate-limited internal JSON endpoint serving your prices to anyone who opens the network tab.
- Not looking at the logs. Most sites with a bot problem have not checked.
- Blocking when throttling would work better. A blocked scraper adapts; a throttled one frequently gives up.
Two checklists
If you are collecting data:
- Confirm there is no API, bulk download, open data source, or licensed feed that would serve.
- Read the terms of service.
- Fetch and parse txt; honour Disallow and any Crawl-delay.
- Check for a sitemap and use it instead of link crawling.
- Set an honest User-Agent with a contact URL, and publish a page at that URL.
- Rate limit to one request per one to two seconds per host, one connection at a time.
- Implement Retry-After, and exponential backoff with jitter.
- Implement conditional requests with ETag and Last-Modified.
- Prefer JSON-LD or a JSON endpoint over CSS selectors.
- Validate every record and alert on the failure rate.
- Run it on infrastructure separate from your website.
- Define what you collect, minimise it, set a retention period, and record provenance.
- Assess the risk tier honestly, and take legal advice if personal information, republication, commercial products, or a competitor is involved.
If you are being scraped:
- Read your access logs. Requests per IP over 24 hours is enough to start.
- Compare bandwidth against analytics visitor numbers; the gap is your bot volume.
- Publish a sensible txt, including AI crawler directives if that is your policy.
- Verify search engine crawlers by forward-confirmed reverse DNS and allowlist them.
- Rate limit at the proxy layer, returning 429, with generous limits first.
- Apply tighter limits to expensive endpoints: search, filtered listings, internal JSON APIs.
- Enable a web application firewall if you have access to one.
- Check Search Console crawl stats after every rules change.
- Confirm staging and development environments are access-restricted and not indexed.
- Decide your AI crawler policy deliberately rather than by default.
Where this is heading
Cautious forecasting, marked as such.
Canadian copyright reform is overdue and possible. A TDM exception was recommended by the 2019 statutory review, has been consulted on, and has support from legal academics — and opposition from parts of the cultural sector and from the Canadian Bar Association's IP section, which argues existing fair dealing suffices. A future Copyright Act review is due. If an exception arrives it would materially change the analysis in this article, which is a reason to check the position rather than rely on a guide.
The privacy position will be resolved by courts, not commentary. The Alberta Charter finding and the outstanding Quebec review mean the "publicly available" question is live. Expect clarification, and expect it to differ by province in the interim.
Crawling is becoming a commercial transaction. Infrastructure providers have begun building mechanisms to charge AI crawlers for access, and publishers have begun signing licensing agreements. The direction is away from crawling as a free-by-default activity, which changes the economics on both sides.
Detection and evasion will keep escalating, and the businesses that stay out of that arms race by collecting cooperatively will spend less and break less often.
Open data will keep growing. The most likely improvement in most organisations' data access over the next few years is not better scraping technique but noticing what is already published.
FAQ
Is web scraping legal in Canada?
There is no law that makes scraping itself illegal, but several bodies of law can apply depending on what you collect and what you do with it. Copyright is the main one, and Canada's position is more restrictive than the United States': fair dealing is limited to enumerated purposes such as research, private study, education, criticism and news reporting, and there is no text and data mining exception. If your collection includes personal information, privacy legislation applies. Terms of service may create contractual obligations. Scraping public, non-personal facts at modest volume for internal use sits at the low end of the risk range; building a commercial product on scraped copyrighted content or personal information sits at the high end.
Can I scrape information that is publicly available?
Be careful with that phrase, because in Canada it is a defined legal term rather than a description. The exception in privacy legislation covers specific enumerated categories such as certain publications and registries, and the internet is not among them. In the 2021 joint Clearview AI investigation, the federal, Quebec, BC and Alberta privacy commissioners found the exception must be read narrowly and does not extend to personal information collected from online sources including social media. The position is genuinely unsettled — an Alberta court has since found the resulting consent requirement in that province's legislation constitutionally overbroad — but "it was on a public website" is not the safe harbour most people assume.
Does Canada have a text and data mining exception?
No. The Copyright Act does not address text and data mining. Innovation, Science and Economic Development Canada has consulted on introducing one, the 2019 statutory review recommended it, and no exception has been enacted. Section 30.71 provides a narrow exception for temporary reproductions essential to a technological process, which is not a general permission to build datasets. This is a substantive difference from jurisdictions that have such exceptions, and from US fair use, which is open-ended.
Do I have to obey robots.txt?
It is not a legal obligation and it is not an access control. It is the clearest available statement of what a site owner wants, standardised as RFC 9309, and ignoring it undermines any argument that your collection was reasonable and considerate. Follow it. If a site sets Crawl-delay, follow that too — it is an explicit statement of how much traffic they will tolerate, and honouring it makes you dramatically less likely to be blocked.
How fast can I scrape without getting blocked?
There is no universal number, but one request every one to two seconds per host, with a single concurrent connection, is rarely a problem. What matters as much as the rate is how you handle refusal: honour Retry-After when you receive a 429 or 503, back off exponentially with a random component, and never retry immediately. Add conditional requests and most of your traffic disappears entirely, because unchanged pages return a 304 with no body.
What is the single most effective way to reduce my scraper's footprint?
Conditional requests. Store the ETag and Last-Modified values the server returns and send them back as If-None-Match and If-Modified-Since on the next run. Unchanged pages then cost a few hundred bytes instead of the full page. On a recurring job over a mostly-static site this commonly cuts transferred data by an order of magnitude, which speeds up your job, reduces the target's load, and makes you far less likely to be noticed.
Can I run a scraper on my web hosting account?
You generally should not, and on shared hosting it usually breaches the acceptable-use policy. Three practical problems: sustained automated requests consume resources your website needs, so your own site slows down; your outbound traffic shares an IP address with other accounts, so blocks and reputation damage spread; and that same shared address handles your outbound email, meaning a scraper that gets the IP flagged can break your transactional mail. Run collection on separate infrastructure — a VPS with its own IP and resources is the usual answer.
How do I know if my site is being scraped?
Compare your bandwidth against your analytics visitor count; a widening gap is bot traffic. Then look at your access logs for requests per IP over a day. The patterns are distinctive: perfectly regular intervals, sequential URL traversal, deep pagination no human visits, constant overnight activity, and missing headers such as Accept-Language or referrer. Server load with no corresponding orders or enquiries is another indicator.
How do I block scrapers without hurting my search rankings?
Never block on User-Agent string alone, in either direction — it is trivially forged. Verify legitimate search crawlers using forward-confirmed reverse DNS: look up the requesting IP, confirm it resolves to a hostname in the crawler's documented domain, then resolve that hostname and confirm it returns the same IP. Both directions must match. Allowlist verified crawlers before applying rate limits, and check Search Console crawl stats after any change to your rules. A blocking rule that quietly reduces your crawl rate is a much bigger problem than the scraping it prevented.
Can I stop AI crawlers from training on my content?
Partially. You can name the crawlers that publish names in robots.txt — GPTBot, CCBot and others — and the ones that respect the file will comply. That does not cover crawlers that do not identify themselves or do not honour the file, so rate limiting and a web application firewall are the backstop. The list of named crawlers changes, so review it periodically. You will also see llms.txt suggested; it is a proposal with inconsistent adoption rather than a standard, and it is not currently a reliable control.
Someone is republishing my content. What should I do?
Document it with dated screenshots and archived copies, establish your originality through publication dates and Search Console history, and start with a polite note to the site owner, which resolves more cases than people expect. If that fails, every hosting provider has an abuse contact and most act on well-documented complaints, and both Google and Bing accept copyright removal requests. Take advice before sending anything that reads as a legal threat. One reassurance: a scraped copy rarely outranks the original, because your page has age, links and crawl history the copy does not.
Should I use a proxy service to avoid being blocked?
This guide does not cover block evasion, and there is an engineering reason as well as a legal one. Systems built to avoid detection are expensive, fragile and permanently one change away from breaking, and the residential proxy networks commonly sold for the purpose route traffic through ordinary people's home connections under consent arrangements worth examining before you buy. If you are being blocked, the more durable fix is usually to request less, more slowly, with an honest identity — or to ask the site for access.
Key Takeaways
- Canada has no text and data mining exception, and fair dealing is a closed list of purposes. US fair use is open-ended; Canadian scraping guidance written for that jurisdiction does not transfer.
- "Publicly available" is a narrow statutory term in Canada. Regulators found in the Clearview AI investigation that it does not extend to personal information gathered online. The law is genuinely unsettled following a 2025 Alberta Charter ruling, which is a reason for advice rather than confidence.
- Work down the ladder before writing code: API, bulk download, open data portal, licensed feed, ask. Canadian open data portals hold a great deal of what businesses try to scrape.
- Better scraping is not harder-to-detect scraping. Fewer requests, honest identification, and cooperation are cheaper, more reliable, and less exposed.
- Conditional requests are the highest-value technique available and are absent from nearly every tutorial. ETag and Last-Modified turn unchanged pages into a few hundred bytes.
- Honour `Retry-After` and back off with jitter. Most blocking is a response to volume and to ignoring refusal.
- Do not run scrapers on your website's hosting account. Shared IP reputation, resource contention, and a likely acceptable-use breach — and a flagged IP can break your outbound email.
- Collect less. Data minimisation is both a legal principle and the cheapest risk reduction available.
- On the receiving side, rate limiting at the proxy is the most effective single measure, returning 429 and applied most tightly to expensive endpoints and internal JSON APIs.
- Verify search crawlers by forward-confirmed reverse DNS, never by `User-Agent`. Blocking Googlebot by accident is a far worse outcome than the scraping you were preventing.
Conclusion
Web scraping is one of those subjects where the available advice and the sensible practice have drifted a long way apart.
The advice, overwhelmingly, is American in its legal assumptions and adversarial in its engineering. It tells you fair use probably covers you, which is not a doctrine Canada has. It tells you public data is fair game, which misreads a defined statutory term. And it frames the technical problem as a contest against detection, which sells proxy subscriptions and produces systems that break constantly.
The sensible practice is quieter. Most of the data businesses try to scrape is either already published somewhere they are welcome to take it, or available through an interface built for the purpose, or obtainable by asking. Where scraping genuinely is the only route, the techniques that make it sustainable are the ones that reduce your footprint rather than disguise it: read the file that tells you what the owner wants, say who you are, request less, cache what has not changed, and stop when you are asked to.
That approach happens to align with the legal position rather than testing it. A collector who identified themselves, honoured robots.txt, took only factual values, kept the minimum, and stayed within a rate the target could absorb is in a materially better position than one who did none of those things — both practically, in that they are less likely to be blocked, and if it ever comes to it, in explaining what they did and why.
And if you run a website, the same subject arrives from the other direction whether you engage with it or not. Bot traffic is a cost you are already paying, in bandwidth, in performance for real visitors, and in analytics you cannot fully trust. Reading your logs and putting a rate limit in front of your expensive endpoints is an afternoon's work that most sites have never done.
Both halves come back to the same unglamorous conclusion: the better way is usually to do less, more carefully.
Two practical next steps, depending on which side of this you are on.
If you are collecting data, the first thing to fix is where it runs. A scraper sharing a hosting account with your website is competing with your own visitors for resources and sharing an IP address whose reputation you cannot fully control — including the address your outbound email leaves from. VPS hosting with root access gives you a separate environment with its own IP and its own resources, and the freedom to install what you need without a shared-hosting policy in the way. For continuous or larger collection, dedicated server hosting removes co-tenancy from the equation. If the data you are storing includes personal information, keeping it in Canadian data centres in Vancouver and Toronto keeps the residency question straightforward to answer.
If you are being scraped, start by reading your access logs — requests per IP over the last day, compared against your analytics visitor count. Most site owners who do this for the first time find something. Rate limiting at the proxy layer is then the highest-value fix, and it requires server-level access that shared plans generally do not offer. On managed WordPress hosting much of the caching and filtering layer is handled for you, which absorbs a good deal of bot load before it reaches your application.
Either way, run the relevant checklist above first. Both are free, both take an afternoon, and both tend to surface something worth knowing before you spend anything.
Not sure which applies, or want a second look at what your logs are telling you? Get in touch and we will go through it with you.









