NoticiasNews

Hoy entra en vigor el bloqueo por defecto de Cloudflare a los rastreadores de IA: qué cambia para su sitio y sus agentesCloudflare’s Default AI Crawler Block Takes Effect Today: What Changes for Your Site and Your Agents

2026-09-15

La fecha que Cloudflare anunció el 1 de julio de 2026 llegó: desde hoy, 15 de septiembre de 2026, la compañía bloquea por defecto los rastreadores de entrenamiento y de agentes de IA en las páginas que muestran publicidad. Los rastreadores de búsqueda siguen permitidos. El cambio no es un ajuste técnico menor: Cloudflare está frente a aproximadamente una quinta parte de la web, y el nuevo comportamiento por defecto redefine quién puede leer contenido publicado y bajo qué condiciones.

De un interruptor a una taxonomía de propósitos

Hasta ahora la decisión era binaria: bloquear bots de IA, o no. Cloudflare reemplazó ese interruptor por un sistema basado en el propósito del bot. Ya no pregunta si el visitante es un bot, sino para qué viene:

  • Search: recopila o indexa contenido para poder responder preguntas sobre él después. Sigue permitido por defecto en todas las páginas.
  • Agent: comportamiento automatizado que actúa, normalmente en tiempo real, en nombre de una persona para resolver algo ahora. Bloqueado por defecto en páginas con publicidad.
  • Training: rastreo que toma contenido para entrenar o ajustar un modelo. Bloqueado por defecto en páginas con publicidad.

La taxonomía va más allá de esas tres categorías: Cloudflare también rastrea propósitos como Transact, Data Collection, Security Testing, SEO, Ads Verification, Social/Link Preview, Feed Fetching y Monitoring & Operations. Eso significa que el monitoreo de precios, la verificación publicitaria y las herramientas de disponibilidad ahora tienen carriles propios en lugar de quedar mezcladas con el entrenamiento de modelos.

El criterio de «páginas con publicidad» tiene una lógica explícita. En palabras de Cloudflare: un anuncio es la señal de que el dueño del sitio esperaba que una persona llegara ahí y lo viera. Si la página se monetiza con atención humana, un bot que consume el contenido sin entregar un humano es, por defecto, indeseado.

El problema de los rastreadores de uso mixto

El detalle más consecuente es cómo se tratan los rastreadores de uso mixto: bots que hacen indexación de búsqueda y entrenamiento de modelos bajo el mismo user agent. Googlebot, Applebot y Bingbot son los ejemplos que la propia Cloudflare cita. Bajo las nuevas reglas, un rastreador multipropósito será bloqueado por los clientes que hayan elegido bloquear Training: el rastreador hereda el trato de su propósito más restringido. Google discute esa caracterización y señala su token Google-Extended, que permite excluirse del entrenamiento sin afectar el posicionamiento en búsqueda. Cómo se resuelva ese pulso es la mayor incógnita abierta del cambio.

A quién le aplica hoy

Los nuevos valores por defecto no cambian todas las zonas existentes de un golpe. Aplican a:

  • Dominios nuevos que se incorporen a Cloudflare desde hoy.
  • Sitios nuevos creados por clientes existentes.
  • Zonas existentes del plan gratuito (free-tier).

Los clientes de pago que no hayan tocado nada conservan su comportamiento actual, y los controles de tráfico de IA llevan semanas disponibles en la configuración de seguridad de cada panel. Pero como las zonas gratuitas sí entran en el cambio, una porción muy grande de la web de cola larga pasa hoy a ser bloqueo-por-defecto para tráfico de entrenamiento y de agentes.

Content Signals: robots.txt con consecuencias

Junto al bloqueo, Cloudflare extendió su robots.txt gestionado con una línea de preferencia legible por máquinas, del tipo Content-Signal: search=yes,ai-train=no,use=reference. El campo use declara uno de tres niveles de uso permitido: immediate (el bot puede interactuar pero no almacenar ni reutilizar), reference (el valor por defecto: indexar, citar fragmentos y enlazar de vuelta) y full (resumir y reproducir el contenido).

Dos cosas convierten esto en algo más que una sugerencia cortés. Primero, Cloudflare lo ata a la aplicación: un bot que reproduce contenido completo no puede alcanzar el estatus de Verified, y los bots no verificados caen en el grupo bloqueado por defecto. Segundo, el contexto legal cambió: con las obligaciones para IA de propósito general del Reglamento de IA de la UE exigibles desde el 2 de agosto de 2026, las exclusiones legibles por máquina adquieren peso jurídico real para quien entrena modelos con datos rastreados. Una línea en robots.txt ya no es etiqueta; ignorarla empieza a ser una responsabilidad documentada y fechada.

Identidad verificable en lugar de reputación de IP

La verificación también cambió de significado. Antes, un bot Verified quedaba permitido por defecto: la verificación era el permiso. Ahora verificación y permiso están desacoplados — «Verified» significa que un bot es admisible dentro de su categoría. Un rastreador de entrenamiento verificado sigue bloqueado en una página donde el entrenamiento está bloqueado.

Cloudflare propuso además un mecanismo de confianza transitiva para arrastrar la identidad y el propósito declarado de un bot a través de intermediarios —capas de proxy, infraestructura de agentes— usando la cabecera estándar Forwarded del RFC 7239, sobre Web Bot Auth, el esquema criptográfico de identidad de bots basado en firmas de mensajes HTTP. La señal estratégica es difícil de ignorar: los porteros de la web están migrando de la reputación de IP («¿de dónde viene esta solicitud?») a la identidad declarada y verificable («¿quién es y qué hará con el contenido?»).

Y el acceso licenciado como tercera vía

El mercado Pay Per Crawl —que permitía a los sitios cobrar por solicitud— evoluciona hacia Pay Per Use, donde los editores cobran cuando su contenido genera valor aguas abajo, no solo cuando es descargado. Para los equipos de datos, la conclusión es que «bloqueado» y «gratis» ya no son los únicos dos estados posibles de una página: el acceso licenciado a un precio se convierte en una opción de primera clase, y será cada vez más el camino conforme los valores por defecto amurallen el resto.

Qué revisar esta semana

Si su organización opera rastreadores, agentes o cualquier recolección automatizada, vale la pena un inventario rápido: qué dominios objetivo están detrás de Cloudflare, cómo se clasifica cada carga de trabajo propia dentro de la taxonomía, si un solo rastreador está mezclando varios propósitos bajo un user agent (el «problema Googlebot» en su propia casa), y si conviene declararse y solicitar verificación o presupuestar acceso licenciado. Si además opera un sitio detrás de Cloudflare, los controles ya están en su panel: decidir por categoría es hoy una decisión de negocio, no de infraestructura.

En TEKFENIX construimos software empresarial a medida e integraciones donde la identidad de quien consume una API o un contenido —humano, sistema o agente de IA— tiene que ser explícita, auditable y gobernable. Es el mismo principio que aplicamos en Servigo365, nuestra mesa de ayuda multicanal con IA, donde cada acción automatizada queda atribuida y trazable, y en CumplimientoControl, donde ese rastro se convierte en evidencia para auditoría regulatoria. Si su empresa depende de integraciones, scraping legítimo o agentes que consumen contenido de terceros, conversemos sobre cómo adaptar su arquitectura a un internet que ya no pregunta si usted es un bot, sino para qué vino.

The date Cloudflare announced on July 1, 2026 has arrived: as of today, September 15, 2026, the company blocks AI training and agent crawlers by default on pages that carry ads. Search crawlers remain allowed. This is not a minor technical tweak: Cloudflare fronts roughly a fifth of the web, and the new default behavior redefines who may read published content and under what conditions.

From a switch to a taxonomy of purposes

Until now the decision was binary: block AI bots, or don’t. Cloudflare replaced that switch with a system based on the bot’s purpose. It no longer asks whether the visitor is a bot, but what the bot is for:

  • Search: collects or indexes content so it can answer questions about it later. Still allowed by default on all pages.
  • Agent: automated behavior acting, usually in real time, on a person’s behalf to get something done right now. Blocked by default on ad-supported pages.
  • Training: a crawler taking content to train or fine-tune a model. Blocked by default on ad-supported pages.

The taxonomy goes beyond those three: Cloudflare also tracks purposes such as Transact, Data Collection, Security Testing, SEO, Ads Verification, Social/Link Preview, Feed Fetching and Monitoring & Operations. That means price monitoring, ad verification and uptime tooling now have named lanes of their own rather than being lumped in with model training.

The “ad-supported pages” criterion has explicit reasoning. In Cloudflare’s words: an ad is a signal that a website owner meant for a person to land there and see it. If a page is monetized by human attention, a bot that consumes the content without delivering a human is, by default, unwelcome.

The mixed-use crawler problem

The most consequential detail is how mixed-use crawlers are treated: bots that perform search indexing and model training under the same user agent. Googlebot, Applebot and Bingbot are Cloudflare’s own examples. Under the new rules, a multi-purpose crawler will be blocked by customers who have selected to block Training — the crawler inherits the treatment of its most-restricted purpose. Google disputes that characterization, pointing to its Google-Extended token, which lets sites opt out of training use without affecting Search rankings. How that standoff resolves is the single biggest open question of the change.

Who it applies to today

The new defaults do not flip every existing zone at once. They apply to:

  • New domains onboarding to Cloudflare from today.
  • New sites created by existing customers.
  • Existing free-tier zones.

Paid customers who changed nothing keep their current behavior, and the AI traffic controls have been live in every dashboard’s security settings for weeks. But because free zones are included, a very large share of the long-tail web becomes block-by-default for training and agent traffic today.

Content Signals: robots.txt with consequences

Alongside the blocking, Cloudflare extended its managed robots.txt with a machine-readable preference line such as Content-Signal: search=yes,ai-train=no,use=reference. The use field declares one of three permitted content-use levels: immediate (a bot may interact but store and reuse nothing), reference (the default: index, excerpt and link back) and full (summarize and reproduce the content).

Two things make this more than a polite suggestion. First, Cloudflare ties it to enforcement: a bot that reproduces content in full cannot achieve Verified status, and unverified bots fall into the default-blocked pool. Second, the legal context shifted: with the EU AI Act’s general-purpose AI obligations enforceable since August 2, 2026, machine-readable opt-outs carry real legal weight for anyone training models on scraped data. A robots.txt line is no longer etiquette; ignoring it is becoming a documented, timestamped liability.

Verifiable identity instead of IP reputation

Verification also changed meaning. Previously a Verified bot was allowed by default: verification was the permission. Now verification and permission are decoupled — “Verified” means a bot is allowable within its relevant category. A verified training crawler is still blocked on a page where training is blocked.

Cloudflare also proposed a transitive-trust mechanism to carry a bot’s identity and declared purpose through intermediaries — proxy layers, agent infrastructure — using the standard Forwarded header from RFC 7239, on top of Web Bot Auth, the cryptographic bot-identity scheme built on HTTP Message Signatures. The strategic signal is hard to miss: the web’s gatekeepers are migrating from IP reputation (“where does this request come from?”) to declared, verifiable identity (“who is this and what will they do with the content?”).

And licensed access as a third path

The Pay Per Crawl marketplace — which let sites charge per request — is evolving into Pay Per Use, where publishers earn when their content creates value downstream, not just when it is fetched. For data teams, the takeaway is that “blocked” and “free” are no longer a page’s only two states: licensed access at a price becomes a first-class option, and increasingly the compliant path to content the defaults now wall off.

What to review this week

If your organization operates crawlers, agents or any automated collection, a quick inventory is worth the hour: which target domains sit behind Cloudflare, how each of your own workloads classifies within the taxonomy, whether a single crawler is blending several purposes under one user agent (the “Googlebot problem” in your own house), and whether to declare yourself and apply for verification or budget for licensed access. If you also run a site behind Cloudflare, the controls are already in your dashboard: deciding per category is now a business decision, not an infrastructure one.

At TEKFENIX we build custom enterprise software and integrations where the identity of whoever consumes an API or a piece of content — human, system or AI agent — has to be explicit, auditable and governable. It is the same principle we apply in Servigo365, our AI-powered multichannel help desk, where every automated action is attributed and traceable, and in CumplimientoControl, where that trail becomes evidence for regulatory audit. If your business depends on integrations, legitimate scraping or agents consuming third-party content, let’s talk about adapting your architecture to an internet that no longer asks whether you are a bot, but what you came for.

← Volver al blog← Back to blog