Build1 publisher3 min readPublished
The AI-training bans live on the big infrastructure blogs, not the small publications
A developer read the current Terms of Service for every feed in his news digest. The clauses that would break a summarization pipeline clustered in VC-data sites and well-lawyered corporate blogs.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- The developer's product, devdigest, is a daily tech digest that pulls from RSS feeds, has a model read each excerpt to categorise and score it, writes an original summary, and emails subscribers a list of titles, summaries and links back to the original articles.
- Full article text is never reproduced or stored beyond a short-lived per-run cache; to read anything, the subscriber clicks through to the publisher.
- Before charging money for the digest, the developer read the actual current Terms of Service for all of his sources, not privacy policies or summaries, quoting the clause or explicitly recording "no ToS found."
- He expected restrictiveness to track company size, assuming big corporate sites with their own RSS feeds, press APIs and developer-relations teams would be relaxed and small scrappy publications would be strict; he found the most restrictive terms came from VC-data companies and large corporate infrastructure blogs.
- Crunchbase News bans using content "to train models (including generative artificial intelligence technologies)," and separately bans anything that "'Crawls,' 'scrapes,' or 'spiders' any page, data, or portion of" the content.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A developer who runs devdigest, a daily tech digest built on RSS, read the actual current Terms of Service for every source in his pipeline before charging money for it, quoting the operative clause or recording "no ToS found" [s1c3]. His finding, published on dev.to, is that the restrictive clauses cluster in VC-data companies and large corporate infrastructure blogs rather than in small publications [s1c4] - which is the opposite of the assumption most people building on feeds are working from.
The pipeline in question is deliberately modest: pull RSS, have a model read each excerpt to categorise and score it, write an original summary, email subscribers titles, summaries and links back to the publisher [s1c1]. Full article text is never reproduced or stored beyond a short-lived per-run cache, and reading anything requires a click through to the source [s1c2].
The clauses that break that design, per his reading: Crunchbase News bans using content "to train models (including generative artificial intelligence technologies)" and separately bans anything that "'Crawls,' 'scrapes,' or 'spiders'" the content [s1c5]. HPCwire/AIwire bans any "robot, spider, or other automatic device" without prior written permission, plus "any form of data extraction or data mining, or other commercial exploitation of any kind" [s1c6]. TechRadar, owned by Future plc, prohibits text or data mining and web scraping "for any purpose, including the development, training, fine-tuning or validation of AI systems or models" [s1c7]. Cloudflare's blog bars automated bots from scraping or data mining content "for developing, training, fine-tuning, or otherwise contributing to or improving" a machine learning model or AI system [s1c8]. Sifted, the European startup publication, carries near-identical language at section 7.6 [s1c9].
The permissive end reads like a list of places you would expect to be precious. InfoQ's terms state: "We permit the posting of a summary and then a link back to the InfoQ landing page" [s1c10]. MIT News says in its Terms of Use that it "offers RSS feeds for syndication purposes" [s1c11]. GitHub's terms say they "do not restrict lawful access to or use of the contents of public repositories by third parties" [s1c12]. arXiv's API terms name RSS-based discovery and notification tools as a permitted use case [s1c13]. TechCrunch maintains dedicated RSS terms, separate from its general ToS, explicitly permitting display of feed content with attribution and a link [s1c14]. Five named sources on each side of the line [s1c19].
His explanation is a guess, and he labels it as one: the restrictive sites have a data-licensing business to protect or expect to have one, so an AI-training ban is an asset to be sold later, while the permissive ones live on distribution and treat the feed as a front door [s1c15].
The more operationally useful part is that the obvious check fails in more than one direction [s1c20]. He originally cleared Towards Data Science because it is Medium-hosted and Medium's ToS contain no RSS, scraping, commercial-use or AI-training restriction [s1c16]. Medium's robots.txt disallows ClaudeBot, GPTBot and every other major AI crawler by name; his pipeline runs on an OpenAI model, so he removed the source [s1c17]. In the other direction, Sifted surfaces an RSS feed link on its own homepage while its terms prohibit both data mining and any automated "robot", "bot", "spider" or "scraper" [s1c18]. A published feed is not a licence, and a clean ToS read is a clean answer to the wrong question [s1c20].
What to watch: whether the AI-training clause becomes standard boilerplate at the large commercial publishers, in which case the safest sources for a summarization pipeline are university news offices, engineering blogs and syndication-native sites - the ones with no licensing revenue to defend [s1c15].