2026
4 篇文章讓搜尋爬蟲進來、訓練爬蟲出去:Cloudflare 把混合用途爬蟲拆成三種行為
Cloudflare 推出 Disallow AI Training,讓站長在保留搜尋收錄的同時拒絕 AI 訓練,並把爬蟲控制拆成 Search、Training、Agent 三類。
閱讀文章 ↗BotBase for Operators:讓網站主與機器人營運者不再互相猜測
Cloudflare 推出 BotBase for Operators,把機器人提交流程從黑箱變成透明可追蹤的儀表板,並用新的行為與內容使用分類法加速審核。
閱讀文章 ↗Bot Preference Sync:Cloudflare 把 AI 機器人政策寫回 robots.txt
Cloudflare 推出 Bot Preference Sync,把 AI 搜尋、代理、訓練的政策自動同步到 robots.txt。文章拆解運作方式、出版者預設,以及透明度要求。
閱讀文章 ↗Firecrawl Crawl Endpoint 實戰:從單頁 scrape 到全站 crawl 的效率取捨
Firecrawl 的 /v2/crawl 用一個 job 自動發現並爬取整個網站、回傳乾淨 Markdown。本文用實戰走查整理核心參數(limit、path 過濾、sitemap 模式)、三種非同步交付模式與計費公式,以及接進 RAG 管線的落地建議。
閱讀文章 ↗
2026
4 ARTICLESCloudflare's Disallow AI Training Setting: What Changes for Your Crawl Policy
Cloudflare's new setting lets mixed-use crawlers keep indexing your site while refusing AI training use.
READ POST ↗BotBase for Operators: Cloudflare Gives Bot Operators a Clearer Path to the Directory
Cloudflare's BotBase for Operators brings transparency to bot submissions: status tracking, editable entries, and a new taxonomy.
READ POST ↗Bot Preference Sync: Cloudflare Syncs Your AI Bot Policies to robots.txt
Cloudflare's Bot Preference Sync writes your AI bot settings to robots.txt, keeping stated preferences and enforced rules aligned. Learn how it works, publisher defaults, and…
READ POST ↗Mastering Firecrawl's Crawl Endpoint: From Single Page Scrape to Full Site Crawl
Firecrawl /v2/crawl discovers and scrapes whole sites into clean markdown in one job. Core parameters, three async delivery modes, the credit formula, and the RAG handoff.
READ POST ↗