Public Web Scraping under the CFAA
Tags: Regulatory
Overview
In LinkedIn v. hiQ, a U.S. federal circuit court held that scraping publicly available information did not violate the Computer Fraud and Abuse Act (CFAA), which criminalizes unauthorized access to “protected computers.” The decision has significant implications for data acquisition practices in AI, including training pipelines and retrieval augmented generation (RAG).
Key Holding
- Public data scraping: Accessing and scraping information that is publicly accessible on the web was found not to constitute a CFAA violation.
Relevance to AI and Advertising
- RAG and scraping risk: The ruling highlights that while CFAA exposure may be limited for public data scraping, other legal risks remain (e.g., copyright, contract, and terms‑of‑service breaches). Advertisers and publishers should consider both legal and technical controls to manage scraping.
- Contractual safeguards: Agreements with AI vendors often address third‑party data sourcing. Common covenants include adherence to websites’ robots.txt or contractual terms, proper identification of bots via user agent strings, and commitments not to bypass technical protections.
Practical Considerations
- Policy and enforcement: Even where scraping may not trigger the CFAA, publishers can implement layered defenses and licensing schemes to manage access and monetize use, particularly amid increased AI-driven ingestion.
- Due diligence: Buyers of AI tools should assess vendors’ scraping practices and data provenance, including compliance with applicable licenses and terms.
Bottom Line
LinkedIn v. hiQ narrows one avenue of potential liability (CFAA) for scraping public information, but it does not resolve broader IP and contractual issues that remain central to AI deployments in digital advertising.
Sources:- IAB_AI_Intellectual_Property_and_Transactions_Digital_Advertising_Playbook_December_2025.pdf
← Back to Knowledge Base