Cloudflare Opens BotBase to Operators: What Data-Collection Teams Should Record Now
Cloudflare announced BotBase for Operators on August 28, 2026. The new workflow gives bot operators a dedicated place to submit a bot, see whether a submission is waiting, accepted, or rejected, read rejection reasons, and update an existing submission when identification details change.
Cloudflare also says its intake now describes a bot across three dimensions: what the bot does, how it uses content, and who operates the traffic. Verification checks can examine an IP list, reverse DNS, or a Web Bot Auth signature. The update does not grant permission to crawl any site, and inclusion in a directory does not override a site owner's policy. It does make a broader operational point clear: crawler identity must be specific, current, and auditable.

Why this matters to proxy-backed data collection
A crawler's software name is only one layer of identity. The traffic seen by a destination also has an egress IP, network owner, reverse-DNS state, user-agent pattern, authentication method, declared purpose, rate profile, and accountable operator.
Proxy rotation can make those layers drift. A new provider, pool, region, or session policy can introduce prefixes that are absent from an allowlist. A broad user-agent can overlap another operator. An old reverse-DNS record can point at retired infrastructure. A project originally used for search indexing can later add model training or agent actions without updating its content-use declaration.
The practical response is not to disguise the drift. It is to manage crawler identity with the same discipline used for credentials and production releases.
Build one identity record per crawler purpose
Do not maintain a single document called “our bot.” Create a record for each materially different purpose and operator boundary. At minimum, record:
- stable crawler name and a narrow user-agent pattern;
- legal or organizational operator and a monitored contact route;
- direct operator versus intermediary role;
- behaviors such as search indexing, user-initiated retrieval, monitoring, research, or collection;
- content uses, including whether content is indexed, summarized, retained, referenced, or used for training;
- proxy provider and product, requested regions, IPv4 or IPv6, and rotating or sticky session mode;
- approved egress prefixes or the controlled endpoint that publishes them;
- reverse-DNS convention or cryptographic verification method;
- owner, reviewer, last verification time, and next review date;
- destination approvals, rate ceilings, retention limits, and emergency-disable control.
If two workloads cannot honestly share these values, they should not share one identity record.
Make proxy egress evidence reproducible
“Uses residential proxies” is not verification evidence. An audit should be able to reproduce what was observed for a defined sample and time window.
- Export the exact pool, product, geography, protocol, address family, and session settings used by the job.
- Sample exits at controlled intervals without forcing rotation around destination controls.
- Record the observed IP, prefix, ASN, network label, reverse DNS, region result, and timestamp.
- Compare the sample with the operator's published or registered identification method.
- Separate expected churn from unapproved network changes.
- Quarantine new prefixes until their ownership and intended use have been reviewed.
Never publish proxy credentials, session tokens, customer identifiers, or a raw internal inventory in a public crawler profile. Publish only the evidence required by the chosen verification model.
Treat user-agent specificity as a testable control
Cloudflare says automatic review checks whether a user-agent pattern is specific enough and does not overlap a bot already tracked. Teams can test the same risk before submission.
Create positive cases for every supported client version and negative cases for common browsers, generic libraries, and other internal crawlers. Run the declared pattern against both sets. A good pattern identifies the intended crawler without claiming unrelated traffic.
Then compare the declared pattern with production request logs. A perfect form entry is useless if a fallback client sends a generic library user-agent or if a browser workflow replaces the header. Alert on undeclared variants rather than silently accepting them.
Keep behavior and content use separate
Behavior answers what the crawler does. Content use answers what happens after retrieval. These can change independently.
A price-monitoring job may collect a small set of public fields for a short retention period. A search crawler may index pages and retain references. A user-triggered agent may fetch one page on demand. These workloads should not inherit one another's declarations merely because they share a proxy gateway.
Add a release gate when a product changes retrieval purpose, storage duration, downstream recipient, model use, or operator role. Require the identity record and destination permission review to be updated before the new behavior reaches production.
Add identity checks to every proxy change
Before promoting a new pool, region, or provider, run four gates:
- Network gate: sampled prefixes, ASN, reverse DNS, and address family match the reviewed evidence.
- Protocol gate: the user-agent and any supported verification signature survive the actual HTTP, SOCKS5, browser, and retry paths.
- Policy gate: purpose, content use, destinations, rate limits, retention, and privacy handling remain approved.
- Operations gate: contacts, ownership, change history, monitoring, and emergency stop are current.
Fail closed when identity evidence disappears. Do not rotate into another address simply to bypass a denial or rate limit.
Monitor acceptance without treating it as permission
The new submission history can show waiting, accepted, or rejected states and provide reasons when a submission needs changes. Store those states internally with timestamps and the exact evidence version that was submitted.
Acceptance means a directory has classified and tracked the crawler. It does not mean every website must allow it. Site-specific terms, robots directives, content signals, authentication, contractual limits, rate limits, privacy duties, and applicable law still control the job.
Deployment checklist
- [ ] Every crawler purpose has an owner and monitored contact.
- [ ] User-agent matching passes positive and negative tests.
- [ ] Proxy pool, geography, IP family, protocol, and session mode are recorded.
- [ ] IP-list, reverse-DNS, or signature evidence is current.
- [ ] Behavior and content-use declarations match the production workflow.
- [ ] Direct versus intermediary operation is explicit.
- [ ] New prefixes are reviewed before promotion.
- [ ] Destination permission and rate limits are enforced separately.
- [ ] Rejections and classification changes enter the change log.
- [ ] An emergency stop disables the crawler without waiting for a deploy.
FAQ
Does directory acceptance authorize scraping?
No. It is identity and classification evidence, not permission from a destination. The destination's rules and applicable law still apply.
Can a rotating residential proxy pool be verified?
It can be documented, but the evidence must match the verification model and remain current as the pool changes. Sampled network classification alone does not prove consent, purpose, or identity.
Should every exit IP be placed in a public list?
Not automatically. Follow the verification method and protect sensitive internal data. A controlled prefix list, reverse DNS, or signed identity may be more appropriate depending on the operator and platform.
What should trigger a resubmission or update?
Changes to the user-agent, verification endpoint, network ranges, reverse DNS, operator role, behavior, content use, or accountable organization should trigger review.
Compliance note
Use proxies and automated collection only for systems and data you are authorized to access. Respect destination terms, robots directives, declared content preferences, rate limits, privacy requirements, security controls, contracts, and applicable law. Identification must not be used to imply permission that has not been granted.
Related 98IP resources: residential proxy provenance auditing, proxy trial acceptance testing, and proxy concurrency ramp testing.
Related Recommendations
- Explore the benefits of proxy servers for online privacy
- What is family home IP? Comprehensively analyze its definition and application
- A must-see for TikTok operations: Why residential IP is the key to account stability
- What are the API interfaces provided by proxy IP products?
- How do foreign questionnaires make money? Do you need to use overseas residential IP?
- Crawler Platform Agents: Types and Key Factors for Selection
- Understand the value of purchasing a US IP address
- Why does an error occur when crawlers use a proxy?
- Search localization: How does Google search country-specific results?
- Overseas News IP Agents Help Mailbox Batch Registration