
Two things are worth establishing before any discussion of how to scrape a website. The first is legal, and it is the reason most guidance on this subject is unsafe reading for a Canadian business. Almost every web scraping article, tutorial and vendor page assumes American law. In the United States, fair use is an open-ended standard that courts apply case by case, and several recent decisions have gone in favour of large-scale data collection. Canada has no equivalent. Fair dealing under the Copyright Act is a closed list of specific purposes, and Canada has no text and data mining exception at all. The federal government has consulted on introducing one; it has not done so. A Canadian business scraping copyrighted material for a commercial purpose that does not fall within the enumerated categories is on considerably weaker ground than an American business doing the same thing, and no amount of "web scraping is legal" content written in California changes that. The second is about the word "better." A large part of the scraping industry defines better as harder to detect: rotate through residential IP addresses, spoof browser fingerprints, defeat the CAPTCHA, make your bot indistinguishable from a person. That is the business model of several of the tools you will find recommended, and this article does not cover it. Not out of squeamishness, but because it is a bad engineering strategy as well as a legally exposed one. It produces systems that are expensive, fragile, permanently one step from breaking, and dependent on infrastructure of dubious provenance. The better ways are duller and they work. Check whether the data is already available in a form you are welcome to take. Read the terms before you write the code. Identify yourself honestly. Make fewer requests. Cache what you already have. Fail gracefully. Store less than you could. These practices are more reliable, cheaper to maintain, and far less likely to end in a letter from someone's lawyer. This guide covers both, because for most readers here both apply. The first half is about collecting data from other people's sites. The second half is about the other side of the same...










